https://hgpu.org/?p=8205
An Optimized Parallel IDCT on Graphics Processing Units