high performance computing on graphics processing units: hgpu.org

hgpu.org » Programming » Algorithms » Improving Locality of Unstructured Mesh Algorithms on GPUs

Improving Locality of Unstructured Mesh Algorithms on GPUs

Andras Attila Sulyok, Gabor Daniel Balogh, Istvan Zoltan Reguly, Gihan R. Mudalige

Faculty of Information Technology and Bionics, Pazmany Peter Catholic University, Hungary

arXiv:1802.03749 [cs.MS], (11 Feb 2018)

BibTeX

Download (PDF)

View

Source

1753

views

To most efficiently utilize modern parallel architectures, the memory access patterns of algorithms must make heavy use of the cache architecture: successively accessed data must be close in memory (spatial locality) and one piece of data must be reused as many times as possible (temporal locality). In this work we analyse the performance of unstructured mesh algorithms on GPUs, specifically the use of the shared memory and two-layered colouring to cache the data. We also look at different block layouts to analyse the trade-off between data reuse and the amount of synchronisation. We developed a standalone library that can transparently reorder the operations done and data accessed by a kernel, without modifications to the algorithm by the user. Using this, we performed measurements on relevant scientific kernels from different applications, such as Airfoil, Volna, Bookleaf, Lulesh and miniAero; using Nvidia Pascal and Volta GPUs. We observed significant speedups (1.2-2.5x).

Tags: Algorithms, Computer science, CUDA, Finite element method, Mathematical Software, nVidia, Performance, Tesla P100

February 15, 2018 by hgpu

No votes yet.

Please wait...

hpcbench: A set of benchmarking utilities for biomolecular simulation tools

Engineering Supercomputing Platforms for Biomolecular Applications

high performance computing on graphics processing units: hgpu.org

Improving Locality of Unstructured Mesh Algorithms on GPUs

Recent source codes

hpcbench: A set of benchmarking utilities for biomolecular simulation tools

HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration

chemtrain: Training Molecular Dynamics Potentials in JAX

microSYCL: SYCL micro-benchmarks repository

XaaS containers

CASS: Cuda-Amd aSSembly

Cluser of smartphones for edge computing application using TensorFlow

SYCL Container

Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration

Can Large Language Models Predict Parallel Code Performance?

Most viewed papers (last 30 days)

Improving Locality of Unstructured Mesh Algorithms on GPUs

Share this:

Recent source codes

Most viewed papers (last 30 days)