hgpu.org » AVX
Matteo Croci, Garth N. Wells
Tags: AVX, Computer science, Finite element method, Floating point error, Intel, Matrix multiplication, Mixed precision, Package
October 27, 2024 by hgpu
Alice Lasserre, Raymond Namyst, Pierre-André Wacrenier
Tags: Algorithms, AVX, Computer science, Education, MPI, OpenCL, OpenMP, Package, Pthreads
February 16, 2020 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
- Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight
- MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
- Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
- Agentic Code Optimization via Compiler-LLM Cooperation
- FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
- DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
- DVM: Real-Time Kernel Generation for Dynamic AI Models
- ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
- Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
* * *




