high performance computing on graphics processing units: hgpu.org

hgpu.org » Programming » Algorithms » A CUDA implementation of the High Performance Conjugate Gradient benchmark

A CUDA implementation of the High Performance Conjugate Gradient benchmark

Everett Phillips, Massimiliano Fatica

NVIDIA Corporation, Santa Clara, CA 95050, USA

5th International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS14), 2014

@article{phillips2014cuda,

title={A CUDA implementation of the High Performance Conjugate Gradient benchmark},

author={Phillips, Everett and Fatica, Massimiliano},

year={2014}

}

Download (PDF)

View

Source

2691

views

The High Performance Conjugate Gradient (HPCG) benchmark has been recently proposed as a complement to the High Performance Linpack (HPL) benchmark currently used to rank supercomputers in the Top500 list. This new benchmark solves a large sparse linear system using a multigrid preconditioned conjugate gradient (PCG) algorithm. The PCG algorithm contains the computational and communication patterns prevalent in the numerical solution of partial differential equations and is designed to better represent modern application workloads which rely more heavily on memory system and network performance than HPL. GPU accelerated supercomputers have proved to be very effective, especially with regard to power efficiency, for accelerating compute intensive applications like HPL. This paper will present the details of a CUDA implementation of HPCG, and the results obtained at full scale on the largest GPU supercomputers available: the Cray XK7 at ORNL and the Cray XC30 at CSCS. The results indicate that GPU accelerated supercomputers are also very effective for this type of workload.

Tags: Algorithms, Benchmarking, Computer science, CUDA, Differential equations, nVidia, Partial differential equations, PDEs, Tesla K20

November 29, 2014 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

* * *

high performance computing on graphics processing units: hgpu.org

A CUDA implementation of the High Performance Conjugate Gradient benchmark

Your response

Recent source codes

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Probe-and-Refine Tuning of Repository Guidance for AI Coding Agents

CUDAnalyst (CUDA + Analyst)

CodegenBench

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels

CUDA Kernel Fusion Benchmarks

IntelliKit: Agent-first tooling for AMD hardware

DITRON: Distributed Compiler based on Triton for Parallel Systems

Most viewed papers (last 30 days)

A CUDA implementation of the High Performance Conjugate Gradient benchmark

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)