hgpu.org » nVidia H100
Wenqing Wu
Tags: Computer science, CUDA, nVidia, nVidia A100, nVidia GeForce RTX 3090, nVidia GeForce RTX 4080 Ti, nVidia H100, Task scheduling
November 27, 2023 by hgpu
Tal Kadosh, Niranjan Hasabnis, Vy A. Vo, Nadav Schneider, Neva Krien, Abdul Wasay, Nesreen Ahmed, Ted Willke, Guy Tamir, Yuval Pinter, Timothy Mattson, Gal Oren
Tags: Code generation, Computer science, Deep learning, HPC, nVidia, nVidia A40, nVidia H100, OpenMP, Package
September 6, 2023 by hgpu
Phuong Nguyen, Pratik Nayak, Hartwig Anzt
Tags: Computer science, CUDA, Intel, Intel Data Center GPU Max 1550, nVidia, nVidia A100, nVidia H100, Package, performance portability, Physics, SYCL
August 20, 2023 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
- Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
- DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering
- Concurrency Response of Plain Global Loads on the NVIDIA H100
- Hardware-Aware FP4 FlashAttention-4
- Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
- Stencil Computation at the Intersection of AI and HPC
* * *



