Publications / 2026
Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition
Abstract
LFST learns a time-frequency representation directly with a speech-emotion task. Starting from DC-corrected Morlet wavelets, it combines responses through a weighted geometric mean and learns fractional-order weights, a monotone log-frequency grid and frequency-dependent cycle counts. Learnable asymmetric hard thresholding produces sparse, denoised activations. A compact Spectro-Temporal Emotion Encoder (STEE) processes magnitude and phase-congruency channels through multiscale blocks, adaptive FiLM gating, axial attention and attentive pooling. The study evaluates the system on IEMOCAP, EMO-DB and the private NSPL-CRISE corpus, and examines frequency grids, cycle schedules and fractional orders through ablations. Mathematical analysis addresses admissibility, continuity, approximate analyticity and conditions for boundedness and stability. The compact encoder reduces parameter requirements relative to large self-supervised models, while the learned front end introduces additional computation compared with STFT- and LEAF-based baselines.
The central contribution is to make the time-frequency front end adaptive without abandoning its mathematical structure. Fractional orders are represented by normalized mixtures over discrete orders, allowing the representation to change continuously during end-to-end training.
STEE uses both magnitude and phase-congruency information. Its architecture combines residual temporal and depthwise-frequency processing with Adaptive FiLM gating and time-axis self-attention. Focal loss and optional class rebalancing support the emotion-classification objective.
The paper and public implementation provide the experimental protocol and configurations. NSPL-CRISE is a private research corpus; the repository’s availability does not make those recordings public.

Citation and BibTeX
Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara. (2026). Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition. ICLR. https://proceedings.iclr.cc/paper_files/paper/2026/hash/b2fb36504c53a6e55f0bcc6951c472cb-Abstract-Conference.html
@inproceedings{nfissi-lfst-2026,
title = {Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition},
author = {Nfissi, Alaa and Bouachir, Wassim and Bouguila, Nizar and Mishara, Brian L.},
year = {2026},
booktitle = {International Conference on Learning Representations},
url = {https://proceedings.iclr.cc/paper_files/paper/2026/hash/b2fb36504c53a6e55f0bcc6951c472cb-Abstract-Conference.html}
}Research projects
- 01
Projects
Adaptive time-frequency learning
A research file connecting learned wavelets, wavelet packets, denoising and fractional superlets.
Research programs
- 01
Research programs
Learnable signal representations
Wavelets, wavelet packets and fractional superlets as adaptive representations for emotion analysis and speech enhancement.
Code
- 01
Repositories
LFST for speech emotion recognition
Research code associated with adaptive time-frequency learning.