https://hgpu.org/?p=5784
A parallel error diffusion implementation on a GPU