Projects

Cross-attentive emotion fusion

Studying the relationship between audio, text and video through cross-attention and feature fusion.

Multimodal model with separate video, text and audio encoders. Audio and video features interact through cross-modal attention; their outputs and text features enter weighted fusion and an emotion classifier.
Audio, text and video meet through cross-attention and weighted feature fusion.Alaa Nfissi and coauthors · multimodal emotion recognition

Research figure

Audio, text and video meet through cross-attention and weighted feature fusion.

active

The Canadian AI 2025 paper examines a multimodal architecture for emotion recognition on IEMOCAP. It combines encoders for audio, text and video with cross-attention and uses focal loss to address class imbalance.

The publisher links Mental-Health-Monitoring-MMER as the paper’s code. A separate Multi-Modal-Emotion-Recognition repository describes submitted work; it is retained as research software without inventing a publication or acceptance.

Research directions

  1. 01

    Research programs

    Interpretable & multimodal learning

    Understanding acoustic features and examining how audio, language and visual cues interact in emotion recognition.

Sign in to save this record →