hgpu.org » nVidia A10
Dipankar Sarkar
Tags: Benchmarking, Code generation, Computer science, CUDA, LLM, nVidia, nVidia A10, nVidia A100, nVidia GeForce RTX 3060, nVidia H100, nVidia L40s, Triton
June 28, 2026 by hgpu
Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, Minlan Yu
Tags: Artificial intelligence, Computer science, CUDA, LLM, Memory, nVidia, nVidia A10, nVidia H100, Tesla T4
November 10, 2024 by hgpu
Carl Andersson, Jonathan Nilsson
Tags: Benchmarking, Computer science, CUDA, Databases, nVidia, nVidia A10, nVidia M60, Performance, Tesla T4, Thesis
December 24, 2023 by hgpu
Muyang Du, Chuan Liu, Jiaxing Qi, Junjie Lai
December 4, 2022 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering
- AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
- Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws
- PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
- Hardware-Aware FP4 FlashAttention-4
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- pytest-gpu-proof: Enabling Cloud-CPU Continuous Integration for GPU Code with Local GPU Attestation
- Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures
- Stencil Computation at the Intersection of AI and HPC
- Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X
* * *


