A 28-nm 141.4TOPS/W Scalable Reconfigurable Deep Learning SoC for Large-Scale Neural Networks
Part Of
2024 IEEE Asian Solid-State Circuits Conference, A-SSCC 2024
Start Page
1
End Page
3
ISBN
979-835037632-6
Date Issued
2024-11-18
Author(s)
I-Ting Lin
Yueh-Feng Tsai
Jeng-Sheng Hsieh
Vincent Chen
Chuei-Tang Wang
Douglas C. H. Yu
Abstract
Devices empowered by Al have brought tremendous benefits to humanity. Fig. 1 shows diverse applications, such as convolutional neural network (CNN) for environmental sensing, graph convolutional network (GCN) for social network analysis, and transformer for natural language processing. Al accelerators have been developed to tackle the increased computational complexity and versatile network structures. Hardware parallelism is applied to improve the throughput for neural network (NN) processing [1–3]. However, hardware parallelism for a specific network structure leads to decreased chip utilization. There has been a surge in the scale of NNs to improve the AI performance. Multi-chip solutions are promising to support large-scale NNs [1–4]. In the distributive multichip designs [1] [4], data for one single layer are partitioned into tensors along a single dimension and processed across multiple chips. This results in high data movement across chips. A chip may also receive insufficient data due to improper distribution, making it infeasible to support one complex layer (with more neurons). In the cascaded multi-chip designs [2] [3], a chip with uni-directional dataflow must await completion of operations by the subsequent chips. This causes low hardware utilization for the multi-chip system, which makes mapping deeper neural networks (with more layers) on multiple chips inefficient (or even infeasible), given limited on-chip memories.
Event(s)
2024 IEEE Asian Solid-State Circuits Conference, A-SSCC 2024
Publisher
Institute of Electrical and Electronics Engineers Inc.
Type
conference paper
