https://hgpu.org/?p=1203
Designing efficient sorting algorithms for manycore GPUs