Audiovisual Transformer with Instance Attention for Audio-Visual Event Localization
Journal
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Journal Volume
12627 LNCS
Pages
274-290
Date Issued
2021
Author(s)
Lin Y.-B
Abstract
Audio-visual event localization requires one to identify the event label across video frames by jointly observing visual and audio information. To address this task, we propose a deep learning framework of cross-modality co-attention for video event localization. Our proposed audiovisual transformer (AV-transformer) is able to exploit intra and inter-frame visual information, with audio features jointly observed to perform co-attention over the above three modalities. With visual, temporal, and audio information observed across consecutive video frames, our model achieves promising capability in extracting informative spatial/temporal features for improved event localization. Moreover, our model is able to produce instance-level attention, which would identify image regions at the instance level which are associated with the sound/event of interest. Experiments on a benchmark dataset confirm the effectiveness of our proposed framework, with ablation studies performed to verify the design of our propose network model. ? 2021, Springer Nature Switzerland AG.
Subjects
Audiovisual; Deep learning; Audio features; Audio information; Benchmark datasets; Cross modality; Event localizations; Learning frameworks; Network modeling; Visual information; Computer vision
Type
conference paper
