high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » QGTC: Accelerating Quantized GNN via GPU Tensor Core

QGTC: Accelerating Quantized GNN via GPU Tensor Core

Yuke Wang, Boyuan Feng, Yufei Ding

U.S.A. University of California, Santa Barbara

arXiv:2111.09547 [cs.DC], (18 Nov 2021)

@misc{wang2021qgtc,

title={QGTC: Accelerating Quantized GNN via GPU Tensor Core},

author={Yuke Wang and Boyuan Feng and Yufei Ding},

year={2021},

eprint={2111.09547},

archivePrefix={arXiv},

primaryClass={cs.DC}

}

Download (PDF)

View

Source

1359

views

Over the most recent years, quantized graph neural network (QGNN) attracts lots of research and industry attention due to its high robustness and low computation and memory overhead. Unfortunately, the performance gains of QGNN have never been realized on modern GPU platforms. To this end, we propose the first Tensor Core (TC) based computing framework, QGTC, to support any-bitwidth computation for QGNNs on GPUs. We introduce a novel quantized low-bit arithmetic design based on the low-bit data representation and bit-decomposed computation. We craft a novel TC-tailored CUDA kernel design by incorporating 3D-stacked bit compression, zero-tile jumping, and non-zero tile reuse technique to improve the performance systematically. We incorporate an effective bandwidth-optimized subgraph packing strategy to maximize the transferring efficiency between CPU host and GPU device. We integrate QGTC with Pytorch for better programmability and extensibility. Extensive experiments demonstrate that QGTC achieves an average of 3.17x speedup compared with the state-of-the-art Deep Graph Library framework across diverse settings.

Tags: Computer science, CUDA, Deep learning, Neural networks, nVidia, nVidia GeForce RTX 3090, TPU

November 21, 2021 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5

DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels

KernelGYM & Dr. Kernel: A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations

Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations

* * *

high performance computing on graphics processing units: hgpu.org

QGTC: Accelerating Quantized GNN via GPU Tensor Core

Your response

Recent source codes

A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5

DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels

KernelGYM & Dr. Kernel: A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations

Vortex-Optimized Light-weight Toolchain (VOLT)

SciDef: Automated Definition Extraction from Scientific Literature

bioagent-bench: Benchmark for evaluating LLM agents in bioinformatics

Benchmark suite for LLM inference on NVIDIA consumer GPUs

Theorizer: from the paper Generating Literature-Driven Scientific Discoveries at Scale

Nsight Python: a Python kernel profiling interface based on NVIDIA Nsight Tools

Awesome LLM-Driven Kernel Generation

Most viewed papers (last 30 days)

QGTC: Accelerating Quantized GNN via GPU Tensor Core

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)