Research programs
Interpretable & multimodal learning
Understanding acoustic features and examining how audio, language and visual cues interact in emotion recognition.
Research direction
Predictive performance is only one part of a useful model. Iterative feature boosting uses Shapley-based analysis to examine the contribution of acoustic features and refine feature sets.
A complementary strand studies cross-attention and feature fusion across audio, text and video. The aim is to examine how different signals support an emotion-recognition task, while keeping evaluation, limitations and reproducibility visible.
- Which acoustic features contribute to a prediction?
- When does cross-modal information improve an emotion model?