high performance computing on graphics processing units: hgpu.org

hgpu.org » Applications » Computer science » Computer vision » Memory-Efficient Implementation of DenseNets

Memory-Efficient Implementation of DenseNets

Geoff Pleiss, Danlu Chen, Gao Huang, Tongcheng Li, Laurens van der Maaten, Kilian Q. Weinberger

Cornell University

arXiv:1707.06990 [cs.CV], (21 Jul 2017)

@article{pleiss2017memoryefficient,

title={Memory-Efficient Implementation of DenseNets},

author={Pleiss, Geoff and Chen, Danlu and Huang, Gao and Li, Tongcheng and Maaten, Laurens van der and Weinberger, Kilian Q.},

year={2017},

month={jul},

archivePrefix={"arXiv"},

primaryClass={cs.CV}

}

Download (PDF)

View

Source

Source codes

Package:

DenseNet: Densely Connected Convolutional Networks

2172

views

The DenseNet architecture is highly computationally efficient as a result of feature reuse. However, a naive DenseNet implementation can require a significant amount of GPU memory: If not properly managed, pre-activation batch normalization and contiguous convolution operations can produce feature maps that grow quadratically with network depth. In this technical report, we introduce strategies to reduce the memory consumption of DenseNets during training. By strategically using shared memory allocations, we reduce the memory cost for storing feature maps from quadratic to linear. Without the GPU memory bottleneck, it is now possible to train extremely deep DenseNets. Networks with 14M parameters can be trained on a single GPU, up from 4M. A 264-layer DenseNet (73M parameters), which previously would have been infeasible to train, can now be trained on a single workstation with 8 NVIDIA Tesla M40 GPUs. On the ImageNet ILSVRC classification dataset, this large DenseNet obtains a state-of-the-art single-crop top-1 error of 20.26%.

Tags: Computer science, Computer vision, CUDA, Deep learning, nVidia, nVidia GeForce GTX Titan X, Package, Tesla M40, Torch

July 25, 2017 by hgpu

Rating: 1.8/5. From 3 votes.

Please wait...

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

gpu_tracker: Python package for tracking and profiling GPU utilization in both desktop and high-performance computing environments

high performance computing on graphics processing units: hgpu.org

Memory-Efficient Implementation of DenseNets

Package:

Recent source codes

SimSYCL: Synchronous, single-threaded, library-only SYCL implementation for debugging and verification

GPU plugin for PySCF

QArray

Celerity: High-level C++ for Accelerator Clusters

gpu_tracker: Context manager and CLI that tracks the computational-resource-usage of a code block or shell command, particularly the GPU usage

CIFAR-10 Airbench: 94% on CIFAR-10 in 3.29 second

LOOPer: a polyhedral compiler for expressing fast and portable data parallel algorithms

OpenMC Monte Carlo Code

Polygeist: C/C++ frontend for MLIR

Parallel Gaussian process with kernel approximation in CUDA

Most viewed papers (last 30 days)

Memory-Efficient Implementation of DenseNets

Package:

Share this:

Recent source codes

Most viewed papers (last 30 days)