high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs

Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs

Yifan Zhao, Egan Johnson, Prasanth Chatarasi, Vikram Adve, Sasa Misailovic

University of Illinois Urbana-Champaign

arXiv:2510.08726 [cs.PL], (9 Oct 2025)

DOI:10.48550/arXiv.2510.08726

@misc{zhao2025neptuneadvancedmloperator,

title={Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs},

author={Yifan Zhao and Egan Johnson and Prasanth Chatarasi and Vikram Adve and Sasa Misailovic},

year={2025},

eprint={2510.08726},

archivePrefix={arXiv},

primaryClass={cs.PL},

url={https://arxiv.org/abs/2510.08726}

}

Download (PDF)

View

Source

Source codes

Package:

Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs

1566

views

Operator fusion has become a key optimization for deep learning, which combines multiple deep learning operators to improve data reuse and reduce global memory transfers. However, existing tensor compilers struggle to fuse complex reduction computations involving loop-carried dependencies, such as attention mechanisms. The paper introduces Neptune, a tensor compiler for advanced operator fusion for sequences of reduction operators. Neptune presents a new approach for advanced operator fusion, which intentionally breaks some existing dependencies and compensates by constructing algebraic correction expressions that allow the kernel to produce the correct result. On ten attention-based benchmarks, Neptune, starting from simple attention code and a high-level scheduling template, outperforms existing compilers like Triton, TVM, and FlexAttention, including Triton-based implementations of FlashAttention. Across four different GPU architectures from NVIDIA and AMD, Neptune-generated kernels have average speedup of 1.35x over the next best alternative, demonstrating its effectiveness for deep learning workloads.

Tags: AMD Radeon Instinct MI300X, ATI, Computer science, CUDA, Deep learning, nVidia, nVidia A100, nVidia RTX A5000, nVidia RTX A6000, Package, Performance, ROCm

October 19, 2025 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

* * *

high performance computing on graphics processing units: hgpu.org

Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs

Package:

Your response

Recent source codes

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Probe-and-Refine Tuning of Repository Guidance for AI Coding Agents

CUDAnalyst (CUDA + Analyst)

CodegenBench

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels

CUDA Kernel Fusion Benchmarks

IntelliKit: Agent-first tooling for AMD hardware

DITRON: Distributed Compiler based on Triton for Parallel Systems

Most viewed papers (last 30 days)

Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs

Package:

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)