https://hgpu.org/?p=4680
Acceleration of Streamed Tensor Contraction Expressions on GPGPU-Based Clusters