hgpu.org » nVidia GeForce RTX 4060
Erel Kaplan, Tomer Bitan, Lian Ghrayeb, Le Chen, Tom Yotam, Niranjan Hasabnis, Gal Oren
Tags: Code generation, Computer science, CUDA, LLM, nVidia, nVidia GeForce RTX 4060, OpenMP, Package
January 12, 2026 by hgpu
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
Evelyne Ringoot, Rabab Alomairy, Valentin Churavy, Alan Edelman
Tags: AMD Radeon Instinct MI250, Apple M1 Pro, ATI, Computer science, HIP, Intel, Intel Ponte Vecchio Max 1100, Kokkos, Linear Algebra, Machine learning, nVidia, nVidia A100, nVidia GeForce RTX 4060, nVidia H100, OpenCL, SYCL
August 17, 2025 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4
- Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization
- Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
- Concurrency Response of Plain Global Loads on the NVIDIA H100
- Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
- Harness Engineering for LLM-Driven GPU Kernel Generation
- Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code
- CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
* * *




