https://hgpu.org/?p=19118
Efficient Interleaved Batch Matrix Solvers for CUDA