A 99.2TOPS/W Transformer Learning Processor with Approximated Attention Score Gradient Computation and Ternary Vector-Based Speculation
Part Of
Digest of Technical Papers - Symposium on VLSI Technology
ISBN (of the container)
979-835036146-9
Date Issued
2024-06-16
Author(s)
Abstract
This work presents the first Transformer learning processor supporting both inference and training acceleration. Byapplying algorithm-architecture optimizations, including approxi-mated gradient computation and ternary vector-based speculation, the training complexity is reduced by up to 94.2%. Adoption of the 8-bit block floating-point (Block-FP) format enables a 39-to-60% power reduction for multiply-accumulate (MAC) operations. The chip delivers a peak energy efficiency of 99. , outperforming the state-of-the-art Transformer inference-only processors by .
Event(s)
2024 IEEE Symposium on VLSI Technology and Circuits, VLSI Technology and Circuits
Publisher
IEEE
Type
conference paper
