high performance computing on graphics processing units: hgpu.org

Posts

Aug, 27

Optimization of Data-Parallel Scientific Applications on Highly Heterogeneous Modern HPC Platforms

Over the past decade, the design of microprocessors has been shifting to a new model where the microprocessor has multiple homogeneous processing units, aka cores, as a result of heat dissipation and energy consumption issues. Meanwhile, the demand for heterogeneity increases in computing systems due to the need for high performance computing in recent years. […]

CUDA

Aug, 27

Surface Normal Integration for Convex Space-time Multi-view Reconstruction

We show that surface normal information allows to significantly improve the accuracy of a spatio-temporal multi-view reconstruction. On one hand, normal information can improve the quality of photometric matching scores. On the other hand, the same normal information can be employed to drive an adaptive anisotropic surface regularization process which better preserves fine details and […]

CUDA

Aug, 27

High Performance Financial Simulation Using Randomized Quasi-Monte Carlo Methods

GPU computing has become popular in computational finance and many financial institutions are moving their CPU based applications to the GPU platform. Since most Monte Carlo algorithms are embarrassingly parallel, they benefit greatly from parallel implementations, and consequently Monte Carlo has become a focal point in GPU computing. GPU speed-up examples reported in the literature […]

CUDA

Aug, 27

A Framework for Lattice QCD Calculations on GPUs

Computing platforms equipped with accelerators like GPUs have proven to provide great computational power. However, exploiting such platforms for existing scientific applications is not a trivial task. Current GPU programming frameworks such as CUDA C/C++ require low-level programming from the developer in order to achieve high performance code. As a result porting of applications to […]

CUDA

Aug, 27

Algorithms for Solving Non-Stationary Heat Conduction Problem for Design of a Technical Device

A model of a multilayer device with non-trivial geometrical and material structure and its working process is suggested. The thermal behavior of the device as one principle characteristic is simulated. The algorithm for solving the non-stationary heat conduction problem with a time-dependent periodical heating source is suggested. The algorithm is based on finite difference explicit–implicit […]

OpenCL

Aug, 26

HSApriori: High Speed Association Rule Mining using Apriori Based Algorithm for GPU

Apriori-Based algorithms are widely used for association rule mining. However, these algorithms cannot exploit the parallel processing power of modern GPU (Graphics Processing Unit). To make an algorithm to be compatible with GPU, it needs to be changed in representation of data, parallel processing and also in support count. In this paper we propose an […]

OpenCL

Aug, 26

Bandwidth Requirements of GPU Architectures

A new trend in chip multiprocessor (CMP) design is to incorporate graphics processing unit (GPU) cores, making them heterogeneous. GPU cores have a higher bandwidth requirement than CPU cores, as they tend to generate much more memory requests. In order to achieve good performance, there must be sufficient bandwidth between the GPU shader cores and […]

CUDA

Aug, 26

An Investigation of Unified Memory Access Performance in CUDA

Managing memory between the CPU and GPU is a major challenge in GPU computing. A programming model, Unified Memory Access (UMA), has been recently introduced by Nvidia to simplify the complexities of memory management while claiming good overall performance. In this paper, we investigate this programming model and evaluate its performance and programming model simplifications […]

CUDA

Aug, 26

Acceleration of Various Direct/Iterative Solvers for MoM by GPU and Its Computational Cost

Various guidelines for acceleration of MoM by GPU computing are summarized. Acceleration of direct/iterative solver for MoM by using GPU is realized. Quantitative study of computing time shows the performance of each guideline.

CUDA

Aug, 26

Speedup of Type-1 Fuzzy Logic Systems on Graphics Processing Units Using CUDA

Parallelcomputing is one of significant components of the High Performance Computing (HPC) and is being used to solve problems, which are large and complex in nature. Fuzzy Logic System (FLS) is a problem that becomes computationally intensive with increase in number of inputs and/or fuzzy rules. Running an FLS is highly parallel in nature, therefore, […]

CUDA

Aug, 23

Structured Orthogonal Inversion of Block p-Cyclic Matrices on Multicore with GPU Accelerators

We present a block structured orthogonal factorization (BSOF) algorithm and its parallelization for computing the inversion of block p-cyclic matrices.We aim at the high performance on multicores with GPU accelerators. We provide a quantitative performance model for optimal host-device load balance, and validate the model through numerical tests. Benchmarking results show that the parallel BSOF […]

CUDA

Aug, 23

GPU Virtualization for High Performance General Purpose Computing on the ESX Hypervisor

Graphics Processing Units (GPU) have become important components in high performance computing (HPC) systems for their massively parallel computing capability and energy efficiency. Virtualization technologies are increasingly applied to HPC to reduce administration costs and improve system utilization. However, virtualizing the GPU to support general purpose computing presents many challenges because of the complexity of […]

CUDA

HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration

chemtrain: Training Molecular Dynamics Potentials in JAX

chemtrain-deploy: A parallel and scalable framework for machine learning potentials in million-atom MD simulations

microSYCL: SYCL micro-benchmarks repository

Exploring SYCL as a Portability Layer for High-Performance Computing on CPUs

See all packages

* * *

high performance computing on graphics processing units: hgpu.org

Posts

Optimization of Data-Parallel Scientific Applications on Highly Heterogeneous Modern HPC Platforms

Surface Normal Integration for Convex Space-time Multi-view Reconstruction

High Performance Financial Simulation Using Randomized Quasi-Monte Carlo Methods

A Framework for Lattice QCD Calculations on GPUs

Algorithms for Solving Non-Stationary Heat Conduction Problem for Design of a Technical Device

HSApriori: High Speed Association Rule Mining using Apriori Based Algorithm for GPU

Bandwidth Requirements of GPU Architectures

An Investigation of Unified Memory Access Performance in CUDA

Acceleration of Various Direct/Iterative Solvers for MoM by GPU and Its Computational Cost

Speedup of Type-1 Fuzzy Logic Systems on Graphics Processing Units Using CUDA

Structured Orthogonal Inversion of Block p-Cyclic Matrices on Multicore with GPU Accelerators

GPU Virtualization for High Performance General Purpose Computing on the ESX Hypervisor

Recent source codes

Efficient GPU Implementation of Multi-Precision Integer Division

ParEval: A Parallel Code Evaluation Benchmark

FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores

exa-AMD: Exascale Accelerated Materials Discovery

WiLLM: An Open Wireless LLM Communication System

Vcc: the Vulkan Clang Compiler

hpcbench: A set of benchmarking utilities for biomolecular simulation tools

HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration

chemtrain: Training Molecular Dynamics Potentials in JAX

microSYCL: SYCL micro-benchmarks repository

Most viewed papers (last 30 days)