Iterative Statistical Kernels on Contemporary GPUs

hgpu.org » Programming » Algorithms » Iterative Statistical Kernels on Contemporary GPUs

Iterative Statistical Kernels on Contemporary GPUs

Thilina Gunarathne, Bimalee Salpitikorala, Arun Chauhan, Geoffrey Fox

School of Informatics and Computing, Indiana University, Bloomington, IN 47405, USASchool of Informatics and Computing, Indiana University, Bloomington, IN 47405, USA

International Journal of Computational Science and Engineering, Vol.8, 58-77, 2013

DOI:10.1504/IJCSE.2013.052118

BibTeX

Download (PDF)

View

Source

2069

views

We present a study of three important kernels that occur frequently in iterative statistical applications: Multi-Dimensional Scaling (MDS), PageRank, and K-Means. We implemented each kernel using OpenCL and evaluated their performance on NVIDIA Tesla and NVIDIA Fermi GPGPU cards using dedicated hardware, and in the case of Fermi, also on the Amazon EC2 cloud-computing environment. By examining the underlying algorithms and empirically measuring the performance of various components of the kernels we explored the optimization of these kernels by four main techniques: (1) caching invariant data in GPU memory across iterations, (2) selectively placing data in different memory levels, (3) rearranging data in memory, and (4) dividing the work between the GPU and the CPU. We also implemented a novel algorithm for MDS and a novel data layout scheme for PageRank. Our optimizations resulted in performance improvements of up to 5X to 6X, compared to naive OpenCL implementations and up to 100X improvement over single-core CPU. We believe that these categories of optimizations are also applicable to other similar kernels. Finally, we draw several lessons that would be useful in not only implementing other similar kernels with OpenCL, but also in devising code generation strategies in compilers that target GPGPUs through OpenCL.

Tags: Algorithms, Cloud, Clustering, Computer science, nVidia, OpenCL, Sparse matrix, Tesla C1060, Tesla C2070, Tesla M2050

March 15, 2012 by hgpu

Rating: 2.5/5. From 2 votes.

Please wait...

high performance computing on graphics processing units: hgpu.org

Iterative Statistical Kernels on Contemporary GPUs

Recent source codes

HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration

chemtrain: Training Molecular Dynamics Potentials in JAX

microSYCL: SYCL micro-benchmarks repository

XaaS containers

SYCL Container

CASS: Cuda-Amd aSSembly

Cluser of smartphones for edge computing application using TensorFlow

CFAL-bench

Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration

Can Large Language Models Predict Parallel Code Performance?

Most viewed papers (last 30 days)

Iterative Statistical Kernels on Contemporary GPUs

Share this:

Recent source codes

Most viewed papers (last 30 days)