Publications / 2025

Neural Wavelet Packet-Based Bidirectional Autoencoder for Multi-Resolution Speech Enhancement

Alaa Nfissi, Wassim Bouachir, Nizar Bouguila

Abstract

Speech enhancement is a critical challenge in signal processing, particularly in noisy environments where preserving intelligibility and perceptual quality is essential. Unlike conventional deep learning-based models that operate exclusively in either the time or frequency domain, we present an adaptive multi-resolution approach that enables superior noise suppression while meticulously preserving critical speech structures across diverse frequency bands. To this end, we introduce the Neural Wavelet Packet-Based Bidirectional Autoencoder (NWPA), a novel framework for multi-resolution speech enhancement. NWPA leverages the Fast Discrete Wavelet Packet Transform with trainable filters that jointly decompose both approximation and detail sub-bands, capturing richer time-frequency features than traditional fixed-wavelet approaches. A bidirectional autoencoder design reduces parameter overhead by unifying the encoding and decoding stages, while an improved Learnable Asymmetric Hard Thresholding function adaptively suppresses noise in the wavelet domain. Furthermore, a Sparsity-Enforcing Loss Function balances reconstruction fidelity with wavelet sparsity, preserving critical speech components across multiple resolutions. Comprehensive evaluations on the VoiceBank-DEMAND dataset demonstrate NWPA’s state-of-the-art performance, underscoring its effectiveness in both noise reduction and intelligibility preservation. These results highlight NWPA’s potential as a robust and scalable solution for speech enhancement under diverse noise conditions. The source code is available at: https://github.com/alaaNfissi/Neural-Wavelet-Packet-Based-Bidirectional-Autoencoder-for-Multi-Resolution-Speech-Enhancement.

Download .bib ↓

Citation and BibTeX

Alaa Nfissi, Wassim Bouachir, Nizar Bouguila. (2025). Neural Wavelet Packet-Based Bidirectional Autoencoder for Multi-Resolution Speech Enhancement. Canadian Conference on Artificial Intelligence (Canadian AI). https://doi.org/10.21428/594757db.eead12b1

@inproceedings{nfissi-nwpa-speech-enhancement-2025,
  title = {Neural Wavelet Packet-Based Bidirectional Autoencoder for Multi-Resolution Speech Enhancement},
  author = {Nfissi, Alaa and Bouachir, Wassim and Bouguila, Nizar},
  year = {2025},
  booktitle = {Canadian Conference on Artificial Intelligence (Canadian AI)},
  doi = {10.21428/594757db.eead12b1},
  url = {https://doi.org/10.21428/594757db.eead12b1}
}

Research projects

  1. 01

    Projects

    Adaptive time-frequency learning

    A research file connecting learned wavelets, wavelet packets, denoising and fractional superlets.

Research programs

  1. 01

    Research programs

    Learnable signal representations

    Wavelets, wavelet packets and fractional superlets as adaptive representations for emotion analysis and speech enhancement.

Code

  1. 01

    Repositories

    NWPA

    Research code associated with adaptive time-frequency learning.

Further reading

Sign in to save this record →