HintOcc: Enhancing Bev-to-3D Reconstruction in Occupancy Prediction with Spatial-Awareness and Dynamic Class Balancing
Journal
Proceedings - 2025 International Conference on Digital Image Computing: Techniques and Applications, DICTA 2025
Start Page
1
End Page
8
ISBN (of the container)
979-833157145-0
Date Issued
2025-12-29
Author(s)
Abstract
With the rapid advancement of autonomous driving, 3D perception has become essential for intelligent vehicles. In complex and dynamic traffic scenes, accurate and efficient object detection is critical for ensuring safety. Recent methods leverage bird's-eye view (BEV) representations for their computational efficiency, but lifting 2D BEV features back into 3D voxel space remains a fundamental challenge due to the loss of vertical information during the encoding process. This often leads to degraded reconstruction performance. Additionally, class imbalance in real-world 3D occupancy datasets-where safety-critical classes such as pedestrians and bicycles are severely underrepresented compared to dominant categories like roads and buildings-significantly hinders the model's performance on these rare classes. To address these challenges, we propose HintOcc, an efficient framework that improves 2 D BEV-to-3D reconstruction and alleviates class imbalance in 3D occupancy prediction. HintOcc introduces (1) a vertical-view branch to recover height information lost during BEV encoding, (2) a deformable depthwise separable head for flexible and lightweight decoding, and (3) a batch-wise dynamic weighting strategy to better emphasize rare classes during training. Evaluated on the Occ3D-NuScenes benchmark, HintOcc achieves competitive performance under similar computational budgets, especially improving accuracy on underrepresented classes.
Event(s)
2025 International Conference on Digital Image Computing: Techniques and Applications, DICTA 2025
Publisher
IEEE
Type
conference paper
