https://hgpu.org/?p=8746
Efficient Weighted Histogramming on GPUs with CUDA