hgpu.org » nVidia H200
Mohammad Firas Sada, John J. Graham, Elham E Khoda, Mahidhar Tatineni, Dmitry Mishin, Rajesh K. Gupta, Rick Wagner, Larry Smarr, Thomas A. DeFanti, Frank Würthwein
Tags: AMD Radeon Instinct MI300A, ATI, Cloud, Computer science, Deep learning, HPC, nVidia, nVidia A100, nVidia H200
July 13, 2025 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
- Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
- DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering
- Concurrency Response of Plain Global Loads on the NVIDIA H100
- Hardware-Aware FP4 FlashAttention-4
- Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
- PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
* * *


