Publications / 2026

Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition

Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara

Abstract

LFST learns a time-frequency representation directly with a speech-emotion task. Starting from DC-corrected Morlet wavelets, it combines responses through a weighted geometric mean and learns fractional-order weights, a monotone log-frequency grid and frequency-dependent cycle counts. Learnable asymmetric hard thresholding produces sparse, denoised activations. A compact Spectro-Temporal Emotion Encoder (STEE) processes magnitude and phase-congruency channels through multiscale blocks, adaptive FiLM gating, axial attention and attentive pooling. The study evaluates the system on IEMOCAP, EMO-DB and the private NSPL-CRISE corpus, and examines frequency grids, cycle schedules and fractional orders through ablations. Mathematical analysis addresses admissibility, continuity, approximate analyticity and conditions for boundedness and stability. The compact encoder reduces parameter requirements relative to large self-supervised models, while the learned front end introduces additional computation compared with STFT- and LEAF-based baselines.

The central contribution is to make the time-frequency front end adaptive without abandoning its mathematical structure. Fractional orders are represented by normalized mixtures over discrete orders, allowing the representation to change continuously during end-to-end training.

STEE uses both magnitude and phase-congruency information. Its architecture combines residual temporal and depthwise-frequency processing with Adaptive FiLM gating and time-axis self-attention. Focal loss and optional class rebalancing support the emotion-classification objective.

The paper and public implementation provide the experimental protocol and configurations. NSPL-CRISE is a private research corpus; the repository’s availability does not make those recordings public.

LFST flow: learn frequency grid and cycle counts, compute fractional-order weights, apply Morlet wavelets, aggregate magnitudes and phase congruency, then threshold and stack the two channels.
LFST front end: learnable frequencies, fractional-order weighting and geometric aggregation.Alaa Nfissi and coauthors · LFST

Research figure

LFST front end: learnable frequencies, fractional-order weighting and geometric aggregation.

Download .bib ↓

Citation and BibTeX

Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara. (2026). Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition. ICLR. https://proceedings.iclr.cc/paper_files/paper/2026/hash/b2fb36504c53a6e55f0bcc6951c472cb-Abstract-Conference.html

@inproceedings{nfissi-lfst-2026,
  title = {Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition},
  author = {Nfissi, Alaa and Bouachir, Wassim and Bouguila, Nizar and Mishara, Brian L.},
  year = {2026},
  booktitle = {International Conference on Learning Representations},
  url = {https://proceedings.iclr.cc/paper_files/paper/2026/hash/b2fb36504c53a6e55f0bcc6951c472cb-Abstract-Conference.html}
}

Research projects

  1. 01

    Projects

    Adaptive time-frequency learning

    A research file connecting learned wavelets, wavelet packets, denoising and fractional superlets.

Research programs

  1. 01

    Research programs

    Learnable signal representations

    Wavelets, wavelet packets and fractional superlets as adaptive representations for emotion analysis and speech enhancement.

Code

  1. 01

    Repositories

    LFST for speech emotion recognition

    Research code associated with adaptive time-frequency learning.

Further reading

Sign in to save this record →