hgpu.org » nVidia H100
Phuong Nguyen, Pratik Nayak, Hartwig Anzt
Tags: Computer science, CUDA, Intel, Intel Data Center GPU Max 1550, nVidia, nVidia A100, nVidia H100, Package, performance portability, Physics, SYCL
August 20, 2023 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
- Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight
- MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
- Agentic Code Optimization via Compiler-LLM Cooperation
- Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
- FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
- DVM: Real-Time Kernel Generation for Dynamic AI Models
- ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
- Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
- A Human–Machine Collaborative Tuning Framework for Triton Kernel Optimization on SIMD Platforms
* * *




