high performance computing on graphics processing units: hgpu.org

hgpu.org » Programming » Algorithms » Parallel Spherical Harmonic Transforms on heterogeneous architectures (GPUs/multi-core CPUs)

Parallel Spherical Harmonic Transforms on heterogeneous architectures (GPUs/multi-core CPUs)

Mikolaj Szydlarski, Pierre Esterie, Joel Falcou, Laura Grigori, Radek Stompor

Institut National de Recherche en Informatique et en Automatique

arXiv:1106.0159v2 [cs.DC] (5 Jun 2012)

@article{2011arXiv1106.0159S,

author={Szydlarski}, M. and {Esterie}, P. and {Falcou}, J. and {Grigori}, L. and {Stompor}, R.},

title={"{Spherical harmonic transform on heterogeneous architectures using hybrid programming}"},

journal={ArXiv e-prints},

archivePrefix={"arXiv"},

eprint={1106.0159},

primaryClass={"cs.DC"},

keywords={Computer Science – Distributed, Parallel, and Cluster Computing, Astrophysics – Cosmology and Extragalactic Astrophysics, Physics – Atmospheric and Oceanic Physics, Physics – Computational Physics, Physics – Geophysics},

year={2011},

month={jun},

adsurl={http://adsabs.harvard.edu/abs/2011arXiv1106.0159S},

adsnote={Provided by the SAO/NASA Astrophysics Data System}

}

Download (PDF)

View

Source

1671

views

Spherical Harmonic Transforms (SHT) are at the heart of many scientific and practical applications ranging from climate modelling to cosmological observations. In many of these areas new, cutting-edge science goals have been recently proposed requiring simulations and analyses of experimental or observational data at very high resolutions and of unprecedented volumes. Both these aspects pose formidable challenge for the currently existing implementations of the transforms. This paper describes parallel algorithms for computing the SHTs with two variants of intra-node parallelism appropriate for novel supercomputer architectures, multi-core processors and Graphic Processing Units (GPU) and discusses their performance tests, alone and embedded within a top-level, MPI-based parallelization layer ported from the S$^2$HAT library, in terms of their accuracy, overall efficiency and scalability. We show that our inverse SHTs with GeForce 400 Series GPUs equipped with latest CUDA architecture ("Fermi") outperforms the state of the art implementation for a multi-core processor executed on a current Intel Core i7-2600K. Furthermore, we show that an MPI/CUDA version of the inverse transform run on a cluster of 128 NVIDIA Tesla S1070 is as much as 3 times faster than the hybrid MPI/OpenMP version executed on the same number of quad-core processors Intel Nahalem for problem sizes motivated by our target applications. For the direct transforms, the performance is however found to be at the best comparable. Here we discuss in detail optimizations of two major steps involved in the transforms calculation, demonstrating how the overall performance efficiency can be obtained, and elucidating the sources of the dichotomy between the direct and the inverse operations

Tags: Algorithms, Astrophysics, Cluster computing, Computational Physics, Cosmology, Cosmology and Extragalactic Astrophysics, CUDA, Heterogeneous systems, MPI, nVidia, nVidia GeForce GTX 460, nVidia GeForce GTX 480, Tesla S1070

June 6, 2012 by hgpu

No votes yet.

Please wait...

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

gpu_tracker: Python package for tracking and profiling GPU utilization in both desktop and high-performance computing environments

high performance computing on graphics processing units: hgpu.org

Parallel Spherical Harmonic Transforms on heterogeneous architectures (GPUs/multi-core CPUs)

Recent source codes

SimSYCL: Synchronous, single-threaded, library-only SYCL implementation for debugging and verification

GPU plugin for PySCF

QArray

Celerity: High-level C++ for Accelerator Clusters

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

CIFAR-10 Airbench: 94% on CIFAR-10 in 3.29 second

LOOPer: a polyhedral compiler for expressing fast and portable data parallel algorithms

OpenMC Monte Carlo Code

Polygeist: C/C++ frontend for MLIR

Parallel Gaussian process with kernel approximation in CUDA

Most viewed papers (last 30 days)

Parallel Spherical Harmonic Transforms on heterogeneous architectures (GPUs/multi-core CPUs)

Share this:

Recent source codes

Most viewed papers (last 30 days)