hgpu.org » Stream Computing
Gabriele Mencagli, Patrizio Dazzi, Massimo Coppola
Tags: Computer science, CUDA, DSP, nVidia, nVidia A30, Package, Stream Computing
August 4, 2024 by hgpu
I. Buck, T. Foley, D. Horn, J. Sugerman, K. Mike, H. Pat
Tags: ATI, ATI Radeon 9800 XT, ATI Stream, Brook, Computer science, High-level Languages, nVidia, nVidia GeForce FX 5900 Ultra, OpenGL, Stream Computing
November 3, 2010 by hgpu
Recent source codes
* * *
Most viewed papers (last 30 days)
- DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering
- AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
- Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws
- PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
- Hardware-Aware FP4 FlashAttention-4
- Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
- pytest-gpu-proof: Enabling Cloud-CPU Continuous Integration for GPU Code with Local GPU Attestation
- Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures
- Stencil Computation at the Intersection of AI and HPC
- Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X
* * *



