A Lightweight Enhancement Approach for Real-Time Semantic Segmentation by Distilling Rich Knowledge from Pre-Trained Vision-Language Model
Journal
APSIPA Transactions on Signal and Information Processing
Journal Volume
13
Journal Issue
5
ISSN
2048-7703
Date Issued
2024
Author(s)
Abstract
In this work, we propose a lightweight approach to enhance real-time semantic segmentation by leveraging the pre-trained vision-language models, specifically utilizing the text encoder of Contrastive Language-Image Pretraining (CLIP) to generate rich semantic embeddings for text labels. Then, our method distills this textual knowledge into the segmentation model, integrating the image and text embeddings to align visual and textual information. Additionally, we implement learnable prompt embeddings for better class-specific semantic comprehension. We propose a two-stage training strategy for efficient learning: the segmentation backbone initially learns from fixed text embeddings and subsequently optimizes prompt embeddings to streamline the learning process. The extensive evaluations and ablation studies validate our approach’s ability to effectively improve the semantic segmentation model’s performance over the compared methods.
SDGs
Publisher
Now Publishers
Type
journal article
