hgpu.org » nVidia GeForce RTX 3090
Fabian Knorr, Peter Thoman, Thomas Fahringer
Tags: Compression, Computer science, CUDA, nVidia, nVidia GeForce RTX 2070, nVidia GeForce RTX 3090, Package, SYCL, Tesla V100
August 8, 2021 by hgpu
Boyuan Feng, Yuke Wang, Tong Geng, Ang Li, Yufei Ding
Tags: Algorithms, Computer science, CUDA, Deep learning, Neural networks, nVidia, nVidia GeForce RTX 3090, Package, Precision, Python, Tesla A100
June 27, 2021 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
- Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
- Concurrency Response of Plain Global Loads on the NVIDIA H100
- Hardware-Aware FP4 FlashAttention-4
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- Stencil Computation at the Intersection of AI and HPC
- An HPC Approach to Accelerate Tensor Decompositions
- Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification
- MaxKernel: Agentic Kernel Generation for TPUs
* * *




