https://hgpu.org/?p=16401
A Comparison of Potential Interfaces for Batched BLAS Computations