Swin Transformer for Pedestrian and Occluded Pedestrian Detection
Part Of
Proceedings - IEEE International Symposium on Circuits and Systems
ISBN (of the container)
979-835033099-1
Date Issued
2024-05-19
Author(s)
Jung-An Liang
Abstract
Pedestrian recognition is crucial for computer vision and self-driving system design. In this work, the Swin Transformer (SwinT), which can capture global contextual information and handle long-range dependencies, is adopted to perform pedestrian detection in the complex scene, including the heavy occlusion scenario. The SwinT is capable to capture multi-scale features and spatial relationships in images, making it well-suited for the challenging task of occluded pedestrian detection. We also apply a two-stage detector based on the faster R-CNN framework, which consists of a cascade region proposal network (RPN) and a region of interest (ROI) head, and use anchors and the focal loss during the RPN training process. The experiments conducted on Euro City Persons and CityPersons datasets demonstrate the outstanding performance of the proposed architecture in detecting heavily occluded pedestrians, highlighting its ability to handle challenging scenarios that traditional methods may struggle with.
Event(s)
2024 IEEE International Symposium on Circuits and Systems (ISCAS)
SDGs
Publisher
IEEE
Type
conference paper
