Projects

Emotion from the raw waveform

CNN-n-GRU and related speech-analysis studies connect local acoustic patterns with temporal dependencies.

CNN-n-GRU pipeline from waveform input through convolution and pooling layers, adaptive average pooling, fully connected layers and stacked GRUs to a softmax emotion output.
CNN-n-GRU learns local acoustic structure and temporal dependencies from the waveform.Alaa Nfissi and coauthors · cnn-n-gru

Research figure

CNN-n-GRU learns local acoustic structure and temporal dependencies from the waveform.

active

This dossier brings together the original CNN-n-GRU conference paper, its later journal study, and the ICSC study of emotional states in crisis-helpline recordings. Convolutional layers learn acoustic representations and gated recurrent units model their temporal evolution.

The papers describe different datasets and evaluation settings. Results should be read with those conditions; research code is provided for investigation and reproducibility.

Research directions

  1. 01

    Research programs

    Speech, emotion & human context

    Learning from speech while keeping the speaker, the recording conditions and the limits of interpretation in view.

Sign in to save this record →