hgpu.org » Streaming
Mehrzad Samadi, Amir Hormati, Mojtaba Mehrara, Janghaeng Lee, and Scott Mahlke
Tags: Compiler, CUDA, GPU, nVidia, nVidia GeForce GTX 285, Optimization, Portability, Streaming, Tesla C2050
March 31, 2012 by Moaddeli
Recent source codes
* * *
Most viewed papers (last 30 days)
- Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight
- CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
- MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
- Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
- Agentic Code Optimization via Compiler-LLM Cooperation
- FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
- DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
- DVM: Real-Time Kernel Generation for Dynamic AI Models
- ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
- Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
* * *



