Practicalizing Tree-Based Model Acceleration with CAM through Model Pruning and Data Placement Optimization
Journal
Proceedings - 2025 International Conference on Hardware/Software Codesign and System Synthesis, CODES+ISSS 2025
Start Page
11
End Page
12
ISBN (of the container)
979-840071992-9
ISBN
[9798400719929]
Date Issued
2025-12-09
Author(s)
Abstract
Tree-based model remains state-of-the-art for many tasks involving tabular data. While these models are favored in resource-constrained environments, the inherent characteristics result in inefficiency during inference, posing significant challenges for conventional accelerators. Recent research has achieved unprecedented acceleration with content-addressable memory (CAM), yet at the cost of overwhelming memory consumption with low utilization, which is impractical for numerous real-world applications. This work addresses these issues by introducing an end-to-end framework RETENTION. RETENTION incorporates (1) a pruning algorithm to minimize model complexity under a user-specified accuracy loss tolerance, and (2) two data placement strategies to enhance memory utilization and further reduce capacity requirement. Experiment results show that space efficiency can be improved from 4.35× to 207.12× with less than 3% accuracy loss.
Event(s)
2025 International Conference on Hardware/Software Codesign and System Synthesis, CODES+ISSS 2025
Subjects
content-addressable memory
in-memory computing
pruning algorithm
tree-based machine learning
Publisher
Association for Computing Machinery, Inc
Type
conference paper
