high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Image and Signal Processing » Image processing » Dynamic Partitioning-based JPEG Decompression on Heterogeneous Multicore Architectures

Dynamic Partitioning-based JPEG Decompression on Heterogeneous Multicore Architectures

Wasuwee Sodsong, Jingun Hong, Seongwook Chung, Shin-Dug Kim, Bernd Burgstaller

Department of Computer Science, Yonsei University, Seoul, South Korea

arXiv:1311.5304 [cs.DC], (21 Nov 2013)

BibTeX

Download (PDF)

View

Source

2277

views

With the emergence of social networks and improvements in computational photography, billions of JPEG images are shared and viewed on a daily basis. Desktops, tablets and smartphones constitute the vast majority of hardware platforms used for displaying JPEG images. Despite the fact that these platforms are heterogeneous multicores, no approach exists yet that is capable of joining forces of a system’s CPU and GPU for JPEG decoding. In this paper we introduce a novel JPEG decoding scheme for heterogeneous architectures consisting of a CPU and an OpenCL-programmable GPU. We employ an offline profiling step to determine the performance of a system’s CPU and GPU with respect to JPEG decoding. For a given JPEG image, our performance model uses (1) the CPU and GPU performance characteristics, (2) the image entropy and (3) the width and height of the image to balance the JPEG decoding workload on the underlying hardware. Our run-time partitioning and scheduling scheme exploits task, data and pipeline parallelism by scheduling the non-parallelizable entropy decoding task on the CPU, whereas inverse cosine transformations (IDCTs), color conversions and upsampling are conducted on both the CPU and the GPU. Our kernels have been optimized for GPU memory hierarchies. We have implemented the proposed method in the context of the libjpeg-turbo library, which is an industrial-strength JPEG encoding and decoding engine. Libjpeg-turbo’s hand-optimized SIMD routines for ARM and x86 constitute a competitive yardstick for the comparison to the proposed approach. Retro-fitting our method with libjpeg-turbo provided insights on the software-engineering aspects of re-engineering legacy code for heterogeneous multicores.

Tags: Compression, Heterogeneous systems, Image processing, nVidia, nVidia GeForce GT 430, nVidia GeForce GTX 560 Ti, nVidia GeForce GTX 680, OpenCL, Performance

November 23, 2013 by hgpu

No votes yet.

Please wait...

HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration

* * *

high performance computing on graphics processing units: hgpu.org

Dynamic Partitioning-based JPEG Decompression on Heterogeneous Multicore Architectures

Recent source codes

WiLLM: An Open Wireless LLM Communication System

Vcc: the Vulkan Clang Compiler

hpcbench: A set of benchmarking utilities for biomolecular simulation tools

HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration

chemtrain: Training Molecular Dynamics Potentials in JAX

microSYCL: SYCL micro-benchmarks repository

XaaS containers

CASS: Cuda-Amd aSSembly

Cluser of smartphones for edge computing application using TensorFlow

SYCL Container

Most viewed papers (last 30 days)

Dynamic Partitioning-based JPEG Decompression on Heterogeneous Multicore Architectures

Share this:

Recent source codes

Most viewed papers (last 30 days)