31196

Hardware-Aware FP4 FlashAttention-4

Robert Hu
arXiv:2609.04105 [cs.LG], (3 Sep 2026)

@misc{hu2026hardwareaware,

   title={Hardware-Aware FP4 FlashAttention-4},

   author={Robert Hu},

   year={2026},

   eprint={2609.04105},

   archivePrefix={arXiv},

   primaryClass={cs.LG},

   url={https://arxiv.org/abs/2609.04105}

}

Blackwell’s 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink. We address this with Direct-P for noncausal inference and a causal path that passes the forward quantization directly into backward. Direct-P maps scores directly to FP4 probabilities and reaches up to 2.13x the bfloat16 (BF16) forward throughput on an NVIDIA GB200. The causal path reconstructs probabilities from saved quantized queries and keys and uses 8-bit floating-point (FP8) gradient operands, accelerating a complete single-GPU 8-billion-parameter update by up to 1.14x. Matched distributed training retains FP8 probabilities and values; every tested MXFP4 probability/value training trajectory diverges.
No votes yet.
Please wait...

You must be logged in to post a comment.

* * *

* * *

HGPU group © 2010-2026 hgpu.org

All rights belong to the respective authors

Contact us: