https://hgpu.org/?p=10513
Histogram Computations on GPUs Kernel using Global and Shared Memory Atomics