https://hgpu.org/?p=4439
Case study: Runtime reduction of a buffer insertion algorithm using GPU parallel programming