https://hgpu.org/?p=16893
A Framework for Dense Triangular Matrix Kernels on Various Manycore Architectures