high performance computing on graphics processing units: hgpu.org

Posts

Jul, 15

8th International Conference on Bioscience, Biochemistry and Bioinformatics (ICBBB), 2018

ICBBB conference series held annually to provide an interactive forum for presentation and discussion on Bioscience, Biochemistry and Bioinformatics. The conference welcomes participants from all over the world who are interested in developing professional ties to and/or exploring career opportunities in the region. The conference should serve as an ideal forum to establish relationships from […]

Jul, 15

The 10th International Conference on Computer Modeling and Simulation (ICCMS), 2018

The 10th International Conference on Computer Modeling and Simulation is the main annual research conference aims to bring together researchers around the world to exchange research results and address open issues in all aspects of Computer Modeling and Simulation. The ICCMS 2009-2017 were held in Macau, Sanya, Mumbai, Hong Kong, Rome, Barcelona, Amsterdam, Brisbane, and […]

Jul, 15

4th International Conference on Virtual Reality (ICVR), 2018

2018 4th International Conference on Virtual Reality (ICVR 2018) will be held during February 24-26, 2018 in Hong Kong. ICVR 2018 will bring together top researchers from Asian Pacific areas, North America, Europe and all around the world to exchange research results and address open issues in all aspects of Virtual Reality. Publication Accepted papers […]

Jul, 15

10th International Conference on Machine Learning and Computing (ICMLC), 2018

The ICMLC 2018: 2018 10th International Conference on Machine Learning and Computing aims to bring together leading academic scientists, researchers and research scholars to exchange and share their experiences and research results on all aspects of Machine Learning and Computing. It also provides a premier interdisciplinary platform for researchers, practitioners and educators to present and […]

Jul, 14

Cooperative Kernels: GPU Multitasking for Blocking Algorithms

There is growing interest in accelerating irregular data-parallel algorithms on GPUs. These algorithms are typically blocking, so they require fair scheduling. But GPU programming models (e.g. OpenCL) do not mandate fair scheduling, and GPU schedulers are unfair in practice. Current approaches avoid this issue by exploiting scheduling quirks of today’s GPUs in a manner that […]

OpenCL

Jul, 14

Multikernel Data Partitioning With Channel on OpenCL-Based FPGAs

Recently, field-programmable gate array (FPGA) vendors (such as Altera) have started to address the programmability issues of FPGAs via OpenCL SDKs. In this paper, we analyze the performance of relational database applications on FPGAs using OpenCL. In particular, we study how to improve the performance of data partitioning, which is a very important building block […]

OpenCL

Jul, 14

Benchmarking Data Analysis and Machine Learning Applications on the Intel KNL Many-Core Processor

Knights Landing (KNL) is the code name for the second-generation Intel Xeon Phi product family. KNL has generated significant interest in the data analysis and machine learning communities because its new many-core architecture targets both of these workloads. The KNL many-core vector processor design enables it to exploit much higher levels of parallelism. At the […]

Jul, 14

A Similarity Measure for GPU Kernel Subgraph Matching

Accelerator architectures specialize in executing SIMD (single instruction, multiple data) in lockstep. Because the majority of CUDA applications are parallelized loops, control flow information can provide an in-depth characterization of a kernel. CUDAflow is a tool that statically separates CUDA binaries into basic block regions and dynamically measures instruction and basic block frequencies. CUDAflow captures […]

CUDA

Jul, 14

DeepProf: Performance Analysis for Deep Learning Applications via Mining GPU Execution Patterns

Deep learning applications are computation-intensive and often employ GPU as the underlying computing devices. Deep learning frameworks provide powerful programming interfaces, but the gap between source codes and practical GPU operations make it difficult to analyze the performance of deep learning applications. In this paper, through examing the features of GPU traces and deep learning […]

CUDA

Jul, 5

OpenCL-Based Implementation of an FPGA Accelerator for Molecular Dynamics Simulation

Molecular dynamics (MD) simulations are very important to studyphysical properties of the atoms and molecules. However, a huge amount of processing time is required to simulate a few nano-seconds of an actual experiment. Although the hardware accelerationusing FPGAs provides promising results, huge design time and hardware design skills are required to implement an accelerator successfully. […]

OpenCL

Jul, 5

Real-time colouring and filtering with graphics shaders

Despite the popularity of the Graphics Processing Unit (GPU) for general purpose computing, one should not forget about the practicality of the GPU for fast scientific visualisation. As astronomers have increasing access to three dimensional (3D) data from instruments and facilities like integral field units and radio interferometers, visualisation techniques such as volume rendering offer […]

Jul, 5

A Fast Method For Computing Principal Curvatures From Range Images

Estimation of surface curvature from range data is important for a range of tasks in computer vision and robotics, object segmentation, object recognition and robotic grasping estimation. This work presents a fast method of robustly computing accurate metric principal curvature values from noisy point clouds which was implemented on GPU. In contrast to existing readily […]

CUDA

* * *

high performance computing on graphics processing units: hgpu.org

Posts

8th International Conference on Bioscience, Biochemistry and Bioinformatics (ICBBB), 2018

The 10th International Conference on Computer Modeling and Simulation (ICCMS), 2018

4th International Conference on Virtual Reality (ICVR), 2018

10th International Conference on Machine Learning and Computing (ICMLC), 2018

Cooperative Kernels: GPU Multitasking for Blocking Algorithms

Multikernel Data Partitioning With Channel on OpenCL-Based FPGAs

Benchmarking Data Analysis and Machine Learning Applications on the Intel KNL Many-Core Processor

A Similarity Measure for GPU Kernel Subgraph Matching

DeepProf: Performance Analysis for Deep Learning Applications via Mining GPU Execution Patterns

OpenCL-Based Implementation of an FPGA Accelerator for Molecular Dynamics Simulation

Real-time colouring and filtering with graphics shaders

A Fast Method For Computing Principal Curvatures From Range Images

Recent source codes

SYCL Container

CASS: Cuda-Amd aSSembly

Cluser of smartphones for edge computing application using TensorFlow

CFAL-bench

Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration

Can Large Language Models Predict Parallel Code Performance?

PELSI: Power-Efficient Layer-Switched Inference

Ouroboros: Virtualized Queues for dynamic memory management

MSCCL++: A GPU-driven communication stack for scalable AI applications

Benchmark compute shader of Unity against InteropUnityCUDA

Most viewed papers (last 30 days)