Projects
Cross-attentive emotion fusion
Studying the relationship between audio, text and video through cross-attention and feature fusion.

active
The Canadian AI 2025 paper examines a multimodal architecture for emotion recognition on IEMOCAP. It combines encoders for audio, text and video with cross-attention and uses focal loss to address class imbalance.
The publisher links Mental-Health-Monitoring-MMER as the paper’s code. A separate Multi-Modal-Emotion-Recognition repository describes submitted work; it is retained as research software without inventing a publication or acceptance.
Research directions
- 01
Research programs
Interpretable & multimodal learning
Understanding acoustic features and examining how audio, language and visual cues interact in emotion recognition.