https://hgpu.org/?p=16755
Deep Tensor Convolution on Multicores