https://hgpu.org/?p=11389
Multi-tier Dynamic Vectorization for Translating GPU Optimizations into CPU Performance