LUTIN: Efficient Neural Network Inference with Table Lookup
Part Of
Proceedings of the 29th International Symposium on Low Power Electronics and Design, ISLPED 2024
Start Page
1
End Page
6
ISBN (of the container)
979-840070688-2
Date Issued
2024-08-05
Author(s)
Abstract
DNN models are becoming increasingly large and complex, but they are also being deployed on commodity devices that require low power and latency but lack specialized accelerators. We introduce LUTIN (LUT-based INference), which reduces the amount of matrix multiplication in DNN inference by converting it into table lookups. LUTIN's innovation is its use of hyperparameter optimization to refine the quantization process and vector partitioning, allowing it to run efficiently on a variety of hardware. By reducing off-chip memory lookups and designing a cache-efficient data layout, LUTIN reduces energy consumption while increasing the use of available CPU cache, even on devices with limited processing power. Our approach goes beyond the traditional limitations of 8-bit quantization, investigating lower bit-widths to further reduce LUT size while meeting accuracy requirements. Experimental results show that LUTIN achieves up to a 2.34x speedup in latency and a 2.04x improvement in energy efficiency over full-precision models.
Event(s)
29th ACM/IEEE International Symposium on Low Power Electronics and Design, ISLPED 2024
Publisher
ACM
Type
conference paper
