https://hgpu.org/?p=16556
Efficient softmax approximation for GPUs