high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » TTC: A Tensor Transposition Compiler for Multiple Architectures

TTC: A Tensor Transposition Compiler for Multiple Architectures

Paul Springer, Aravind Sankaran, Paolo Bientinesi

AICES, RWTH Aachen University, Germany

arXiv:1607.01249 [cs.MS], (5 Jul 2016)

DOI:10.1145/2935323.2935328

@article{springer2016tensor,

title={TTC: A Tensor Transposition Compiler for Multiple Architectures},

author={Springer, Paul and Sankaran, Aravind and Bientinesi, Paolo},

year={2016},

month={jul},

archivePrefix={"arXiv"},

primaryClass={cs.MS},

doi={10.1145/2935323.2935328}

}

Download (PDF)

View

Source

Source codes

Package:

TTC: A high-performance Compiler for Tensor Transpositions

2289

views

We consider the problem of transposing tensors of arbitrary dimension and describe TTC, an open source domain-specific parallel compiler. TTC generates optimized parallel C++/CUDA C code that achieves a significant fraction of the system’s peak memory bandwidth. TTC exhibits high performance across multiple architectures, including modern AVX-based systems (e.g.,~Intel Haswell, AMD Steamroller), Intel’s Knights Corner as well as different CUDA-based GPUs such as NVIDIA’s Kepler and Maxwell architectures. We report speedups of TTC over a meaningful baseline implementation generated by external C++ compilers; the results suggest that a domain-specific compiler can outperform its general purpose counterpart significantly: For instance, comparing with Intel’s latest C++ compiler on the Haswell and Knights Corner architecture, TTC yields speedups of up to $8times$ and $32times$, respectively. We also showcase TTC’s support for multiple leading dimensions, making it a suitable candidate for the generation of performance-critical packing functions that are at the core of the ubiquitous BLAS 3 routines.

Tags: BLAS, Compilers, Computer science, CUDA, Intel Xeon Phi, Linear Algebra, Mathematical Software, nVidia, nVidia GeForce 840 M, Package, Performance, Tesla K40

July 8, 2016 by hgpu

Rating: 2.5/5. From 6 votes.

Please wait...

Your response

You must be logged in to post a comment.

high performance computing on graphics processing units: hgpu.org

TTC: A Tensor Transposition Compiler for Multiple Architectures

Package:

Your response

Recent source codes

Agentic Code Optimization via Compiler-LLM Cooperation

MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU

Device Virtual Machine (DVM)

AutoKernel: Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels

SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits

Triton-Sanitizer: A Fast and Device-Agnostic Memory Sanitizer for Triton with Rich Diagnostic Context

LLM.Q: Quantized LLM training in pure CUDA/C++

True 4-Bit Quantized CNN Training on CPU

cuFuzz: A GPU-oriented coverage-guided fuzzer for userland CUDA application

KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization

Most viewed papers (last 30 days)

TTC: A Tensor Transposition Compiler for Multiple Architectures

Package:

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)