https://hgpu.org/?p=4250
Accelerating GPU kernels for dense linear algebra