Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X
author={Kothalkar, Kartik},
title={Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X},
month={sep},
year={2026},
publisher={Zenodo},
version={v1.0.0},
doi={10.5281/zenodo.22898895},
url={https://doi.org/10.5281/zenodo.22898895}
}
The commercial and scientific utility of Large Language Models (LLMs) hinges directly on the efficiency of hyperscale inference serving infrastructure. LLM inference operates across two fundamentally distinct operational regimes: a compute-bound Prefill phase that processes prompt tokens via dense General Matrix Multiplications (GEMMs), and a memory-bandwidth-bound Decode phase that autoregressively generates tokens one-by-one, constrained by the need to reload the extensive Key-Value (KV) cache from High Bandwidth Memory (HBM) for every single token step. As context windows expand from 4,096 to 128,000 tokens, memory capacity and bandwidth saturation become the absolute determinants of inference throughput and tail latency. Currently, enterprise AI deployments face an architectural duopoly between NVIDIA's Hopper H100 SXM5 (80 GB HBM3, 3.35 TB/s bandwidth, 4th Gen Tensor Cores with Transformer Engine) and AMD's Instinct MI300X (192 GB HBM3, 5.30 TB/s bandwidth, CDNA 3 Matrix Cores). Concurrently, specialized serving runtimes—specifically vLLM (with PagedAttention virtual memory), TensorRT-LLM (with In-Flight Batching and fused FP8 GEMMs), and FlashAttention-3 (with Hopper warp-specialized asynchronous memory scheduling)—attempt to maximize hardware utilization. This paper presents the first exhaustive, peer-review-grade empirical benchmark comparing the NVIDIA H100 SXM5 and AMD Instinct MI300X across modern open-weights frontier architectures, including Llama-3-70B-Instruct, Mixtral-8x22B (Sparse MoE), and Qwen-2.5-72B. Across multi-node clusters interconnected via 900 GB/s NVLink 4 and 896 GB/s AMD Infinity Fabric, we profile Time-to-First-Token (TTFT), Inter-Token Latency (ITL, p50 to p99.9), sustained tokens-per-second-per-GPU, HBM3 bus saturation percentage, KV-cache memory fragmentation, and electrical power consumption (Joules per Token). Key Empirical Findings:• Prefill Latency & Compute Dominance: NVIDIA H100 with FlashAttention-3 achieves a 22.4% lower Time-to-First-Token (TTFT) during compute-bound prefill for prompt lengths <=4k tokens due to warp-specialized Tensor Core pipelines.• High-Concurrency Decode Scaling: AMD Instinct MI300X's massive 192 GB HBM3 capacity and 5.30 TB/s bandwidth deliver 41.2% higher aggregate decode throughput at high concurrent request loads (256-512 concurrent streams) and fit Llama-3-70B entirely within a single GPU without requiring multi-GPU tensor parallelism.• KV-Cache Compression Efficiency: 8-bit FP8 KV-cache quantization recovers 88.4% of effective HBM bandwidth on H100 and reduces p99 tail latency variance by 3.2x under bursty context allocation churn.• Power & Infrastructure TCO: At maximum batch saturation, MI300X demonstrates an 18.6% lower energy expenditure per generated million tokens, altering the Total Cost of Ownership (TCO) calculus for enterprise generative AI deployments. The repository includes the full 31-page primary academic treatise, complete novelty and systems defense dossiers, and full Hindi and Marathi translations.
Your response
You must be logged in to post a comment.



