high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » Importance of Data Loading Pipeline in Training Deep Neural Networks

Importance of Data Loading Pipeline in Training Deep Neural Networks

Mahdi Zolnouri, Xinlin Li, Vahid Partovi Nia

Huawei Noah’s Ark Lab, Montreal, QC H3N 1X9, Canada

arXiv:2005.02130 [cs.CV], (21 Apr 2020)

@misc{zolnouri2020importance,

title={Importance of Data Loading Pipeline in Training Deep Neural Networks},

author={Mahdi Zolnouri and Xinlin Li and Vahid Partovi Nia},

year={2020},

eprint={2005.02130},

archivePrefix={arXiv},

primaryClass={cs.CV}

}

Download (PDF)

View

Source

Source codes

Package:

DALI: a library containing both highly optimized building blocks and an execution engine for data pre-processing in deep learning applications

2037

views

Training large-scale deep neural networks is a long, time-consuming operation, often requiring many GPUs to accelerate. In large models, the time spent loading data takes a significant portion of model training time. As GPU servers are typically expensive, tricks that can save training time are valuable.Slow training is observed especially on real-world applications where exhaustive data augmentation operations are required. Data augmentation techniques include: padding, rotation, adding noise, down sampling, up sampling, etc. These additional operations increase the need to build an efficient data loading pipeline, and to explore existing tools to speed up training time. We focus on the comparison of two main tools designed for this task, namely binary data format to accelerate data reading, and NVIDIA DALI to accelerate data augmentation. Our study shows improvement on the order of 20% to 40% if such dedicated tools are used.

Tags: Computer science, CUDA, Deep learning, Neural networks, nVidia, nVidia GeForce GTX Titan, Package, Tesla V100

May 10, 2020 by hgpu

Rating: 2.0/5. From 1 vote.

Please wait...

Your response

You must be logged in to post a comment.

* * *

high performance computing on graphics processing units: hgpu.org

Importance of Data Loading Pipeline in Training Deep Neural Networks

Package:

Your response

Recent source codes

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Probe-and-Refine Tuning of Repository Guidance for AI Coding Agents

CUDAnalyst (CUDA + Analyst)

CodegenBench

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels

CUDA Kernel Fusion Benchmarks

IntelliKit: Agent-first tooling for AMD hardware

DITRON: Distributed Compiler based on Triton for Parallel Systems

Most viewed papers (last 30 days)

Importance of Data Loading Pipeline in Training Deep Neural Networks

Package:

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)