Open-Vocabulary Panoptic Segmentation Using Bert Pre-Training of Vision-Language Multiway Transformer Model
Part Of
Proceedings - International Conference on Image Processing, ICIP
Start Page
2494
End Page
2500
ISSN
15224880
ISBN (of the container)
979-835034939-9
ISBN
979-835034939-9
Date Issued
2024-10-27
Author(s)
DOI
10.1109/ICIP51287.2024.10647459
Abstract
Open-vocabulary panoptic segmentation remains a challenging problem. One of the biggest difficulties lies in training models to generalize to an unlimited number of classes using limited categorized training data. Recent popular methods involve large-scale vision-language pre-trained foundation models, such as CLIP. In this paper, we propose OMTSeg for open-vocabulary segmentation using another large-scale vision-language pre-trained model called BEiT-3 and leveraging the cross-modal attention between visual and linguistic features in BEiT-3 to achieve better performance. Experiments result demonstrates that OMTSeg performs favorably against state-of-the-art models. Code is available at https://github.com/AI-Application-and-IntegrationLab/OMTSeg.
Event(s)
31st IEEE International Conference on Image Processing, ICIP 2024
SDGs
Publisher
IEEE
Type
conference paper
