Compressing DNN Parameters for Model Loading Time Reduction
Journal
2019 IEEE International Conference on Consumer Electronics - Asia, ICCE-Asia 2019
Pages
78-79
Date Issued
2019
Author(s)
Abstract
Deep neural network (DNN) has been applied to a variety of computer vision tasks these days. However, DNN often suffers from its enormous execution time even with the aid of GPU. In this paper, we argue that the bandwidth bottleneck between GPU and GDRAM has to be addressed. To reduce loading time, we propose a DNN acceleration approach which compresses DNN parameters before loading model information to GPU and performs decompressing on GPU. Using JPEG compression as an example, the loss of the test accuracy can be kept within 4%, while an 8 × parameter-size reduction is achieved for VGG16.
SDGs
Type
conference paper
