https://hgpu.org/?p=8602
Fast Parallel Sorting Algorithms on GPUs