https://hgpu.org/?p=23623
Flexible Performant GEMM Kernels on GPUs