https://hgpu.org/?p=3223
Implementing the Himeno benchmark with CUDA on GPU clusters