Empower Typed Descriptions by Large Language Models for Speech Emotion Recognition
Part Of
APSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024
Start Page
1
End Page
6
ISBN
979-835036733-1
Date Issued
2024-12-03
Author(s)
Haibin Wu
Huang-Cheng Chou
Kai-Wei Chang
Lucas Goncalves
Jiawei Du
Chi-Chun Lee
DOI
10.1109/APSIPAASC63619.2025.10848758
Abstract
Training speech emotion recognition (SER) requires human-annotated labels and speech data. However, emotion perception is complex. The pre-defined emotion categories are not enough for annotators to describe their emotion perception. Devoted annotators will use natural language rather than traditional emotion labels when annotating data, resulting in typed descriptions (e.g., "Slightly Angry, calm" to notify the intensity of emotion). While these descriptions are highly valuable, SER models, designed as classification models, cannot process natural languages and thus discard them. To leverage the valuable typed descriptions, we propose a novel way to prompt ChatGPT to mimic annotators, comprehend natural language typed descriptions, and subsequently adjust the given label of the input data. By utilizing labels generated by ChatGPT, we consistently achieve an average relative gain of 3.08% across all settings using 15 speech self-surprised learning models on the SUPERB, which provides a potential way to integrate the power of LLMs to improve the performances of SER.
Event(s)
2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2024
SDGs
Publisher
IEEE
Type
conference paper
