high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » Portability and Scalability of OpenMP Offloading on State-of-the-art Accelerators

Portability and Scalability of OpenMP Offloading on State-of-the-art Accelerators

Yehonatan Fridman, Guy Tamir, Gal Oren

Department of Computer Science, Ben-Gurion University of the Negev, Israel

arXiv:2304.04276 [cs.DC], (9 Apr 2023)

DOI:10.48550/arXiv.2304.04276

@misc{fridman2023portability,

title={Portability and Scalability of OpenMP Offloading on State-of-the-art Accelerators},

author={Yehonatan Fridman and Guy Tamir and Gal Oren},

year={2023},

eprint={2304.04276},

archivePrefix={arXiv},

primaryClass={cs.DC}

}

Download (PDF)

View

Source

Source codes

Package:

ScalSALE: Scalable SALE Benchmark Framework for Supercomputers

1039

views

Over the last decade, most of the increase in computing power has been gained by advances in accelerated many-core architectures, mainly in the form of GPGPUs. While accelerators achieve phenomenal performances in various computing tasks, their utilization requires code adaptations and transformations. Thus, OpenMP, the most common standard for multi-threading in scientific computing applications, introduced offloading capabilities between host (CPUs) and accelerators since v4.0, with increasing support in the successive v4.5, v5.0, v5.1, and the latest v5.2 versions. Recently, two state-of-the-art GPUs – the Intel Ponte Vecchio Max 1100 and the NVIDIA A100 GPUs – were released to the market, with the oneAPI and GNU LLVM-backed compilation for offloading, correspondingly. In this work, we present early performance results of OpenMP offloading capabilities to these devices while specifically analyzing the potability of advanced directives (using SOLLVE’s OMPVV test suite) and the scalability of the hardware in representative scientific mini-app (the LULESH benchmark). Our results show that the vast majority of the offloading directives in v4.5 and 5.0 are supported in the latest oneAPI and GNU compilers; however, the support in v5.1 and v5.2 is still lacking. From the performance perspective, we found that PVC is up to 37% better than the A100 on the LULESH benchmark, presenting better performance in computing and data movements.

Tags: Benchmarking, Computer science, Intel, Intel Ponte Vecchio Max 1100, nVidia, nVidia A100, oneAPI, OpenMP, Package, performance portability

April 16, 2023 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

high performance computing on graphics processing units: hgpu.org

Portability and Scalability of OpenMP Offloading on State-of-the-art Accelerators

Package:

Your response

Recent source codes

Interleaved Learning and Exploration: A Self-Adaptive Fuzz Testing Framework for MLIR

Pinocchio: PINpointing Orbit Crossing Collapsed Hierarchical Objects

KernelCoder: trained on a curated dataset of reasoning traces and CUDA kernel pairs

VibeCodeHPC - Multi Agentic Vibe Coding for HPC

Compile-Time Resource Safety for GPU APIs: A Low-Overhead Typestate Framework

exa-AMD: Exascale Accelerated Materials Discovery

TRUST: a thermalhydraulic software package for CFD simulations

Modular: The Modular Platform (includes MAX & Mojo)

Allo: Accelerator Design Language

Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization

Most viewed papers (last 30 days)

Portability and Scalability of OpenMP Offloading on State-of-the-art Accelerators

Package:

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)