high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » 3D finite difference computation on GPUs using CUDA

3D finite difference computation on GPUs using CUDA

Paulius Micikevicius

NVIDIA, 2701 San Tomas Expressway, Santa Clara, CA 95050

In GPGPU-2: Proceedings of 2nd Workshop on General Purpose Processing on Graphics Processing Units (2009), pp. 79-84.

DOI:10.1145/1513895.1513905

@conference{micikevicius20093d,

title={3D finite difference computation on GPUs using CUDA},

author={Micikevicius, P.},

booktitle={Proceedings of 2nd Workshop on General Purpose Processing on Graphics Processing Units},

pages={79–84},

year={2009},

organization={ACM}

}

Download (PDF)

View

Source

3288

views

In this paper we describe a GPU parallelization of the 3D finite difference computation using CUDA. Data access redundancy is used as the metric to determine the optimal implementation for both the stencil-only computation, as well as the discretization of the wave equation, which is currently of great interest in seismic computing. For the larger stencils, the described approach achieves the throughput of between 2,400 to over 3,000 million of output points per second on a single Tesla 10-series GPU. This is roughly an order of magnitude higher than a 4-core Harpertown CPU running a similar code from seismic industry. Multi-GPU parallelization is also described, achieving linear scaling with GPUs by overlapping inter-GPU communication with computation.

Tags: Computer science, CUDA, Finite difference, nVidia, Tesla S1070, Wave equation

November 2, 2010 by hgpu

No votes yet.

Please wait...

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

gpu_tracker: Python package for tracking and profiling GPU utilization in both desktop and high-performance computing environments

* * *

high performance computing on graphics processing units: hgpu.org

3D finite difference computation on GPUs using CUDA

Recent source codes

QArray

Celerity: High-level C++ for Accelerator Clusters

CIFAR-10 Airbench: 94% on CIFAR-10 in 3.29 second

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

LOOPer: a polyhedral compiler for expressing fast and portable data parallel algorithms

OpenMC Monte Carlo Code

Polygeist: C/C++ frontend for MLIR

Parallel Gaussian process with kernel approximation in CUDA

Optical flow algorithms for SYCL

OpenMP5-Offload-OpenMC-Intel-PVC

Most viewed papers (last 30 days)

3D finite difference computation on GPUs using CUDA

Share this:

Recent source codes

Most viewed papers (last 30 days)