high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » TransAxx: Efficient Transformers with Approximate Computing

TransAxx: Efficient Transformers with Approximate Computing

Dimitrios Danopoulos, Georgios Zervakis, Dimitrios Soudris, Jörg Henkel

School of Electrical and Computer Engineering, National Technical University of Athens, Athens, 15780, Greece

arXiv:2402.07545 [cs.LG], (12 Feb 2024)

DOI:10.48550/arXiv.2402.07545

@misc{danopoulos2024transaxx,

title={TransAxx: Efficient Transformers with Approximate Computing},

author={Dimitrios Danopoulos and Georgios Zervakis and Dimitrios Soudris and Jörg Henkel},

year={2024},

eprint={2402.07545},

archivePrefix={arXiv},

primaryClass={cs.LG}

}

Download (PDF)

View

Source

Source codes

Package:

TransAxx: Fast Emulation of Approximate ViT models in PyTorch

1067

views

Vision Transformer (ViT) models which were recently introduced by the transformer architecture have shown to be very competitive and often become a popular alternative to Convolutional Neural Networks (CNNs). However, the high computational requirements of these models limit their practical applicability especially on low-power devices. Current state-of-the-art employs approximate multipliers to address the highly increased compute demands of DNN accelerators but no prior research has explored their use on ViT models. In this work we propose TransAxx, a framework based on the popular PyTorch library that enables fast inherent support for approximate arithmetic to seamlessly evaluate the impact of approximate computing on DNNs such as ViT models. Using TransAxx we analyze the sensitivity of transformer models on the ImageNet dataset to approximate multiplications and perform approximate-aware finetuning to regain accuracy. Furthermore, we propose a methodology to generate approximate accelerators for ViT models. Our approach uses a Monte Carlo Tree Search (MCTS) algorithm to efficiently search the space of possible configurations using a hardware-driven hand-crafted policy. Our evaluation demonstrates the efficacy of our methodology in achieving significant trade-offs between accuracy and power, resulting in substantial gains without compromising on performance.

Tags: Computer science, CUDA, Machine learning, Neural networks, nVidia, nVidia V100, Package, PyTorch

February 18, 2024 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

KernelGYM & Dr. Kernel: A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations

Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations

* * *

high performance computing on graphics processing units: hgpu.org

TransAxx: Efficient Transformers with Approximate Computing

Package:

Your response

Recent source codes

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

CUDABench: Benchmarking LLMs for Text-to-CUDA Generation

CL4SE: A Context Learning Benchmark For Software Engineering Tasks

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Execution-Free Reward Models

A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5

DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels

KernelGYM & Dr. Kernel: A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations

Vortex-Optimized Light-weight Toolchain (VOLT)

SciDef: Automated Definition Extraction from Scientific Literature

bioagent-bench: Benchmark for evaluating LLM agents in bioinformatics

Most viewed papers (last 30 days)

TransAxx: Efficient Transformers with Approximate Computing

Package:

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)