high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » An Automatic Input-Sensitive Approach for Heterogeneous Task Partitioning

An Automatic Input-Sensitive Approach for Heterogeneous Task Partitioning

Klaus Kofler, Ivan Grasso, Biagio Cosenza, Thomas Fahringer

Institute of Computer Science, University of Innsbruck, Austria

27th ACM international conference on Supercomputing, 2013

@inproceedings{Kofler13,

author={Kofler, Klaus and Grasso, Ivan and Cosenza, Biagio and Fahringer, Thomas},

title={An Automatic Input-Sensitive Approach for Heterogeneous Task Partitioning},

booktitle={Proceedings of the 27th ACM international conference on Supercomputing},

series={ICS ’13},

year={2013},

location={Eugene, Oregon, USA},

publisher={ACM},

address={New York, NY, USA},

keywords={heterogeneous computing, compilers, GPU, task partitioning, code analysis, machine learning, runtime system}

}

Download (PDF)

View

Source

1974

views

Unleashing the full potential of heterogeneous systems, consisting of multi-core CPUs and GPUs, is a challenging task due to the difference in processing capabilities, memory availability, and communication latencies of different computational resources. In this paper we propose a novel approach that automatically optimizes task partitioning for different (input) problem sizes and different heterogeneous architectures. We use the Insieme source-to-source compiler to translate a single-device OpenCL program into a multi-device OpenCL program. The Insieme Runtime System then performs dynamic task partitioning based on an offline-generated prediction model. In order to derive the prediction model, we use a machine learning approach based on Artificial Neural Networks (ANN) that incorporates static program features as well as dynamic, input sensitive features. Principal component analysis have been used to further improve the task partitioning. Our approach has been evaluated over a suite of 23 programs and respectively achieves a performance improvement of 22% and 25% compared to an execution of the benchmarks on a single CPU and a single GPU which is equal to 87.5% of the optimal performance.

Tags: ATI, ATI Radeon HD 5870, Code generation, Compilers, Computer science, Heterogeneous systems, Machine learning, Neural networks, nVidia, nVidia GeForce GTX 480, OpenCL, Performance, Task scheduling

April 21, 2013 by hgpu

Rating: 2.5/5. From 1 vote.

Please wait...

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

gpu_tracker: Python package for tracking and profiling GPU utilization in both desktop and high-performance computing environments

high performance computing on graphics processing units: hgpu.org

An Automatic Input-Sensitive Approach for Heterogeneous Task Partitioning

Recent source codes

SimSYCL: Synchronous, single-threaded, library-only SYCL implementation for debugging and verification

GPU plugin for PySCF

QArray

Celerity: High-level C++ for Accelerator Clusters

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

CIFAR-10 Airbench: 94% on CIFAR-10 in 3.29 second

LOOPer: a polyhedral compiler for expressing fast and portable data parallel algorithms

OpenMC Monte Carlo Code

Polygeist: C/C++ frontend for MLIR

Parallel Gaussian process with kernel approximation in CUDA

Most viewed papers (last 30 days)

An Automatic Input-Sensitive Approach for Heterogeneous Task Partitioning

Share this:

Recent source codes

Most viewed papers (last 30 days)