hgpu.org » Apple M2 Pro
Dahua Feng, Zhiming Xu, Rongxiang Wang, Felix Xiaozhu Lin
Tags: AI, Apple M2 Max, Apple M2 Pro, Apple M2 Ultra, Computer science, CUDA, Linear Algebra, LLM, Machine learning, nVidia, nVidia GeForce RTX 4090, nVidia GeFroce RTX 2080 Ti, nVidia Quadro RTX 4000, nVidia RTX A6000, Performance, PyTorch
February 3, 2025 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
- Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
- Concurrency Response of Plain Global Loads on the NVIDIA H100
- Hardware-Aware FP4 FlashAttention-4
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- Stencil Computation at the Intersection of AI and HPC
- An HPC Approach to Accelerate Tensor Decompositions
- Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification
- MaxKernel: Agentic Kernel Generation for TPUs
* * *



