hgpu.org » CUBLAS
Hua Zhou, Kenneth Lange, Marc A. Suchard
Tags: Block relaxation, CUBLAS, CUDA, EM and MM algorithms, Mathematics, Multidimensional scaling, Nonnegative matrix factorization, nVidia, nVidia GeForce GTX 280, PET scanning, Statistics
October 30, 2010 by hgpu
Nguyen-Khang Pham, Annie Morin, Patrick Gros
October 29, 2010 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4
- Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization
- Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
- Harness Engineering for LLM-Driven GPU Kernel Generation
- Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code
- CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
- FlashPDE: A Drop-In Fused Triton Operator Library for Neural PDE Solvers
- A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- PortLBM: A Portable Lattice Boltzmann Tool Leveraging SYCL on AMD, NVIDIA, and Intel GPUs
* * *




