Hardware-Aware FP4 FlashAttention-4
arXiv:2609.04105 [cs.LG], (3 Sep 2026)
@misc{hu2026hardwareaware,
title={Hardware-Aware FP4 FlashAttention-4},
author={Robert Hu},
year={2026},
eprint={2609.04105},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.04105}
}
Blackwell’s 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink. We address this with Direct-P for noncausal inference and a causal path that passes the forward quantization directly into backward. Direct-P maps scores directly to FP4 probabilities and reaches up to 2.13x the bfloat16 (BF16) forward throughput on an NVIDIA GB200. The causal path reconstructs probabilities from saved quantized queries and keys and uses 8-bit floating-point (FP8) gradient operands, accelerating a complete single-GPU 8-billion-parameter update by up to 1.14x. Matched distributed training retains FP8 probabilities and values; every tested MXFP4 probability/value training trajectory diverges.
September 14, 2026 by hgpu
Your response
You must be logged in to post a comment.





