https://hgpu.org/?p=3339
Optimization and Implementation of LBM Benchmark on Multithreaded GPU