Skip to main navigation Skip to search Skip to main content

Transformer-based fusion of acoustic and textual cues with proportional augmentation for emotion recognition

  • Institut Teknologi Sepuluh Nopember

Research output: Contribution to journalArticlepeer-review

Abstract

Emotion recognition plays a crucial role in human-computer interaction, sentiment analysis, and affective computing, enabling more natural and intuitive communication between machines and users. Although multimodal approaches that leverage both acoustic and textual cues offer a more comprehensive understanding of emotional states, they remain challenging owing to the complexity of human emotions, modality-specific variations, and class imbalance in emotional datasets. Recent transformer-based models have advanced feature extraction in the speech and text domains. However, many existing methods rely on complex fusion architectures or struggle with underrepresented emotions. To address these challenges, we propose a probability-level multimodal fusion framework that integrates a Vision Transformer (ViT) trained on spectrogram-based acoustic representations and Bidirectional Encoder Representations from Transformers (BERT) for textual modeling. The modality-wise emotion posterior probabilities were combined using a lightweight fusion classifier to exploit complementary confidence patterns without feature-level entanglement. Additionally, we introduced a proportional augmentation technique to mitigate class imbalances, ensuring a fairer representation of minority emotions without compromising data integrity. Evaluated on a benchmark dataset for conversational emotion recognition, the proposed approach consistently outperformed unimodal baselines and yielded competitive performance compared with recent multimodal methods, particularly in terms of weighted recall and F1-score for underrepresented emotions.

Original languageEnglish
Article number35
JournalInternational Journal of Speech Technology
Volume29
Issue number1
DOIs
Publication statusPublished - Mar 2026

Keywords

  • Acoustic-textual modalities
  • Affective computing
  • Class imbalance
  • Emotion recognition
  • Transformer models

Fingerprint

Dive into the research topics of 'Transformer-based fusion of acoustic and textual cues with proportional augmentation for emotion recognition'. Together they form a unique fingerprint.

Cite this