A Dual-Mode High Efficient Hardware Architecture Design for Diffusion Models
Journal
Proceedings - 2025 21st IEEE Asia Pacific Conference on Circuits and Systems, APCCAS 2025
Start Page
1
End Page
5
ISBN
[9798331589073]
Date Issued
2025-10-12
Author(s)
Abstract
Diffusion models have achieved strong performance in generative vision tasks but suffer from high compute and memory demands due to their iterative U-Net structure. This paper presents a hardware-efficient diffusion accelerator with three key architectural optimizations. First, a Winograd-Enhanced DualMode Dispatcher improves convolution throughput by eliminating im2col overhead and maintaining high PE utilization. Second, a Dynamic Sparse Attention Engine predicts and prunes lowimportance attention scores at runtime, reducing unnecessary multiplications. Third, a Reconfigurable Mixed-Precision Processor adapts different bit-width workload demands to avoid resource waste. The proposed accelerator is implemented in TSMC 28nm CMOS technology and achieves a peak energy efficiency of 120.59 TOPS/W under 90% sparsity during attention computation, demonstrating its suitability for real-time and lowpower diffusion inference. © 2025 IEEE.
Event(s)
2025 21st IEEE Asia Pacific Conference on Circuits and Systems, APCCAS 2025
Subjects
Deep Learning
Diffusion
Sparsity
Publisher
IEEE
Type
conference paper
