hgpu.org » Network communication
Michael Mandulak, Sayan Ghosh, S M Ferdous, Mahantesh Halappanavar, George Slota
October 19, 2025 by hgpu
Zhiyi Hu, Siyuan Shen, Tommaso Bonato, Sylvain Jeaugey, Cedell Alexander, Eric Spada, Jeff Hammond, Torsten Hoefler
July 13, 2025 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
- Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
- Concurrency Response of Plain Global Loads on the NVIDIA H100
- Hardware-Aware FP4 FlashAttention-4
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- Stencil Computation at the Intersection of AI and HPC
- An HPC Approach to Accelerate Tensor Decompositions
- Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification
- MaxKernel: Agentic Kernel Generation for TPUs
* * *



