https://hgpu.org/?p=8194
Systematic Approach in Optimizing Numerical Memory-Bound Kernels on GPU