high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » Accelerating Krylov Subspace Solvers on Graphics Processing Units

Accelerating Krylov Subspace Solvers on Graphics Processing Units

Hartwig Anzt, Stanimire Tomov, Piotr Luszczek, Ichitaro Yamazaki, Jack Dongarra, William Sawyer

Innovative Computing Lab, University of Tennessee, Knoxville, USA

28th IEEE International Parallel & Distributed Processing Symposium, 2014

@article{anzt2014accelerating,

title={Accelerating Krylov Subspace Solvers on Graphics Processing Units},

author={Anzt, Hartwig and Tomov, Stanimire and Luszczek, Piotr and Yamazaki, Ichitaro and Dongarra, Jack and Sawyer, William},

year={2014}

}

Download (PDF)

View

Source

2550

views

Krylov subspace solvers are often the method of choice when solving sparse linear systems iteratively. At the same time, hardware accelerators such as graphics processing units (GPUs) continue to offer significant floating point performance gains for matrix and vector computations through easy-to-use libraries of computational kernels. However, as these libraries are usually composed of a well optimized but limited set of linear algebra operations, applications that use them often fail to leverage the full potential of the accelerator. In this paper we target the acceleration of the BiCGSTAB solver for GPUs, showing that significant improvement can be achieved by reformulating the method and developing application-specific kernels instead of using the generic CUBLAS library provided by NVIDIA. We propose an implementation that benefits from a significantly reduced number of kernel launches and GPU-host communication events, by means of increased data locality and a simultaneous reduction of multiple scalar products. Using experimental data, we show that, depending on the dominance of the untouched sparse matrix vector products, significant performance improvements can be achieved compared to a reference implementation based on the CUBLAS library. We feel that such optimizations are crucial for the subsequent development of highlevel sparse linear algebra libraries.

Tags: Computer science, CUBLAS, CUDA, Linear Algebra, nVidia, Sparse matrix, Tesla K20

August 3, 2014 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

* * *

high performance computing on graphics processing units: hgpu.org

Accelerating Krylov Subspace Solvers on Graphics Processing Units

Your response

Recent source codes

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Probe-and-Refine Tuning of Repository Guidance for AI Coding Agents

CUDAnalyst (CUDA + Analyst)

CodegenBench

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels

CUDA Kernel Fusion Benchmarks

IntelliKit: Agent-first tooling for AMD hardware

DITRON: Distributed Compiler based on Triton for Parallel Systems

Most viewed papers (last 30 days)

Accelerating Krylov Subspace Solvers on Graphics Processing Units

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)