Sarah Verhulst

dblp:222/1906 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-6498-7719ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021
YearPublicationVenuePosition
2025 Speech stimulus design to study the neural coding of speech and the impact of cochlear synaptopathy
abstract
International audience
Etienne Gaudrain, Sarah Verhulst, Deniz Baskent
INTERSPEECH2
2025 Individualized speech enhancement for hearing-impaired listeners
abstract
Despite significant progress in speech enhancement (SE), improving speech intelligibility and perceptual quality in noisy environments for hearing-impaired individuals remains challenging. This paper presents an individualized speech-enhancement (ISE) framework that integrates noise reduction (NR) and sound amplification for hearing loss compensation within a unified system. Our fully differentiable, closed-loop design incorporates biophysically realistic auditory models. The framework features two pathways: one simulating the auditory response of a normal-hearing (NH) system to denoised speech and the other modeling a hearing-impaired (HI) system's response to noisy speech. The ISE model is trained by minimizing the difference between NH and HI auditory responses. Experimental results show that the ISE model enhances speech intelligibility and perceptual quality in noisy conditions for HI listeners, offering a promising foundation for advancing personalized noise reduction strategies.
Chuan Wen, Sarah Verhulst
INTERSPEECH2
2023 A DNN-Based Hearing-Aid Strategy For Real-Time Processing: One Size Fits All
abstract
Although hearing aids (HAs) can compensate for elevated hearing thresholds using sound amplification, they often fail to restore auditory perception in adverse listening conditions. To achieve robust treatment outcomes for diverse HA users, we use a differentiable framework that can compensate for impaired auditory processing based on a biophysically realistic and personalisable auditory model. Here, we present a deep-neural-network (DNN) HA processing strategy that can provide individualised sound processing for the audiogram of a listener using a single model architecture. The DNN architecture was trained to compensate for different audiogram inputs and was able to enhance simulated responses and intelligibility even for audiograms that were not part of training. Our multi-purpose HA model can be used for different individuals and can process audio inputs of 3.2 ms in <0.5 ms, thus paving the way for precise DNN-based treatments of hearing loss that can be embedded in hearing devices.
Fotios Drakopoulos, Arthur Van Den Broucke, Sarah Verhulst
ICASSP3
2023 Biophysically-inspired single-channel speech enhancement in the time domain
abstract
Deep neural networks (DNN) based speech enhancement approaches have recently achieved great performance. There are a numerous applications that benefit from the speech enhancement model, including automatic speech recognition (ASR) and hearing aids. The majority of these previous methods were developed in the time-frequency (T-F) domain. However, the T-F domain approach has some limitations, including a high minimum delay in reconstructing the signal from T-F domain representation, poor generalizability in unseen noise, and bad performance for negative signal-to-noise ratio’s (SNRs). To address these problems, we propose a biophysically inspired end-to-end time-domain neural network, which adopts bio-inspired features from CoNNear, a neural network that accurately simulates biophysical properties of the human auditory system such as sharp and, level-dependent filter tuning. We first generated biophysical speech feature using CoNNear were subsequently fed into the U-Net-based speech enhancement module. The latter module consisted of generator network without the discriminator from SERGAN (Baby and Verhulst, 2019) and the training dataset we use was INTERSPEECH 2021 DNS Challenge dataset. An objective evaluation was performed using perceptual evaluation of speech quality (PESQ), segmental SNR (segSNR), cepstral distance (CD) and log-likelihood ratio (LLR) with unseen samples from DNS challenge, which is different from training noise scenarios. Results of objective evaluation reveal that bio-inspired features show comparable performance with T-F features at positive SNRs, with improved generalizability in negative SNR and for mismatched noises. Additionally, our time-domain CoNNear features dramatically decreased the minimum latency of the whole system towards 4ms, making it suitable for real-time applications with high constraints on signal delay. The good generalizability in adverse noise conditions and unseen noise, as well as the low latency of our DNN-based model ensure the promising applicability in hearing aids.
Chuan Wen, Sarah Verhulst
INTERSPEECH2
2023 A Neural-Network Framework for the Design of Individualised Hearing-Loss Compensation
abstract
Sound processing in the human auditory system is complex and highly non-linear, whereas hearing aids (HAs) still rely on simplified descriptions of auditory processing or hearing loss to restore hearing. Even though standard HA amplification strategies succeed in restoring audibility of faint sounds, they still fall short of providing targeted treatments for complex sensorineural deficits and adverse listening conditions. These shortcomings of current HA devices demonstrate the need for advanced hearing-loss compensation strategies that can effectively leverage the non-linear character of the auditory system. Here, we propose a differentiable deep-neural-network (DNN) framework that can be used to train DNN-based HA models based on biophysical auditory-processing differences between normal-hearing and hearing-impaired systems. We investigate different loss functions to accurately compensate for impairments that include outer-hair-cell (OHC) loss and cochlear synaptopathy (CS), and evaluate the benefits of our trained DNN-based HA models for speech processing in quiet and in noise. Our results show that auditory-processing enhancement was possible for all considered hearing-loss cases, with OHC loss proving easier to compensate than CS. Several objective metrics were considered to estimate the expected speech intelligibility after processing, and these simulations hold promise in yielding improved understanding of speech-in-noise for hearing-impaired listeners who use our DNN-based HA processing. Since our framework can be tuned to the hearing-loss profiles of individual listeners, we enter an era where truly individualised and DNN-based hearing-restoration strategies can be developed and be tested experimentally.
Fotios Drakopoulos, Sarah Verhulst
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 A Differentiable Optimisation Framework for The Design of Individualised DNN-based Hearing-Aid Strategies
abstract
Current hearing aids mostly provide sound amplification fittings based on individual hearing thresholds or perceived loudness, even though it is known that sensorineural hearing damage is functionally complex, and requires different treatment strategies. To meet this demand, we propose an optimisation framework for the design of individualised hearingaid signal processing based on simulated (hearing-impaired) auditory-nerve responses. The framework is fully differentiable, thus the backpropagation algorithm can be used to train DNN-based hearing-aid models that optimally process sound to restore hearing in impaired cochleae. The auditory models within the framework can be tuned to the precise hearing-loss profile of a listener to yield trully individualised restoration strategies. Our simulations show that the trained hearing-aid models were able to enhance the auditory-nerve responses of hearing-impaired cochleae, and this provides a promising outlook for embedding our framework within future hearing aids and augmented-hearing applications.
Fotios Drakopoulos, Sarah Verhulst
ICASSP2
2020 Hearing-Impaired Bio-Inspired Cochlear Models for Real-Time Auditory Applications
abstract
Biophysically realistic models of the cochlea are based on cascaded transmission-line (TL) models which capture longitudinal coupling, cochlear nonlinearities, as well as the human frequency selectivity. However, these models are slow to compute (order of seconds/minutes) while machine-hearing and hearing-aid applications require a real-time solution. Consequently, real-time applications often adopt more basic and less time-consuming descriptions of cochlear processing (gamma-tone, dual resonance nonlinear) even though there are clear advantages in using more biophysically correct models. To overcome this, we recently combined nonlinear Deep Neural Networks (DNN) with analytical TL cochlear model descriptions to build a real-time model of cochlear processing which captures the biophysical properties associated with the TL model. In this work, we aim to extend the normal-hearing DNN-based cochlear model (CoNNear) to simulate frequency-specific patterns of hearing sensitivity loss, yielding a set of normal and hearing-impaired auditory models which can be computed in real-time and are differentiable. They can hence be used in backpropagation networks to develop the next generation of hearing-aid and machine hearing applications.
Arthur Van Den Broucke, Deepak Baby, Sarah Verhulst
INTERSPEECH3
2019 Sergan: Speech Enhancement Using Relativistic Generative Adversarial Networks with Gradient Penalty
abstract
Popular neural network-based speech enhancement systems operate on the magnitude spectrogram and ignore the phase mismatch between the noisy and clean speech signals. Recently, conditional generative adversarial networks (cGANs) have shown promise in addressing the phase mismatch problem by directly mapping the raw noisy speech waveform to the underlying clean speech signal. However, stabilizing and training cGAN systems is difficult and they still fall short of the performance achieved by spectral enhancement approaches. This paper introduces relativistic GANs with a relativistic cost function at its discriminator and gradient penalty to improve time-domain speech enhancement. Simulation results show that relativistic discriminators provide a more stable training of cGANs and yield a better generator network for improved speech enhancement performance.
Deepak Baby, Sarah Verhulst
ICASSP2
2018 Biophysically-inspired Features Improve the Generalizability of Neural Network-based Speech Enhancement Systems
abstract
Recent advances in neural network (NN)-based speech enhancement schemes are shown to outperform most conventional techniques.However, the performance of such systems in adverse listening conditions such as negative signal-to-noise ratios and unseen noises is still far from that of humans.Motivated by the remarkable performance of humans under these challenging conditions, this paper investigates whether biophysicallyinspired features can mitigate the poor generalization capabilities of NN-based speech enhancement systems.We make use of features derived from several human auditory periphery models for training a speech enhancement system that employs long short-term memory (LSTM), and evaluate them on a variety of mismatched testing conditions.The results reveal that biophysically-inspired auditory models such as nonlinear transmission line models improve the generalizability of LSTMbased noise suppression systems in terms of various objective quality measures, suggesting that such features lead to robust speech representations that are less sensitive to the noise type.
Deepak Baby, Sarah Verhulst
INTERSPEECH2