hgpu.org » nVidia Tesla GP100
Sungho Shin, Youngmin Jo, Jungwook Choi, Swagath Venkataramani, Vijayalakshmi Srinivasan, Wonyong Sung
Tags: Artificial intelligence, Computer science, Deep learning, Neural networks, nVidia, nVidia DGX-1, nVidia GeForce GTX Titan XP, nVidia Tesla GP100
November 11, 2018 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization
- AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning
- UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization
- Real FP4 Tensor-Core Code in Pure Rust on a Gaming GPU - with NVIDIA's Own Compiler
- SpecGen: Accelerating Agentic Kernel Optimization with Speculative Generation
- Probe-and-Refine Tuning of Repository Guidance for Coding Agents
- The Correctness Illusion in LLM-Generated GPU Kernels
- CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries
- Enhancing the Performance Analysis of NCCL GPU Collectives
- Augmenting LLM Code Translation with Compiler Analysis for C to Triton Kernel Generation
* * *



