https://hgpu.org/?p=11187
Opportunities for Parallelism in Matrix Multiplication