Christian Schüldt

dblp:10/4977 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0003-3439-0468ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
abstract
Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each capturing different levels of representation. While prior studies explored their layer-wise representations for efficiency and performance, speech quality assessment (SQA) models predominantly rely on last-layer features, leaving intermediate layers underexamined. In this work, we systematically evaluate different layers of multiple SSL models for predicting mean-opinion-score (MOS). Features from each layer are fed into a lightweight regression network to assess effectiveness. Our experiments consistently show early-layers features outperform or match those from the last layer, leading to significant improvements over conventional approaches and state-of-the-art MOS prediction models. These findings highlight the advantages of early-layer selection, offering enhanced performance and reduced system complexity.
Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee
ASRU5
2025 Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
abstract
In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many types of impairments are clustered. While DNN-based SQA models are not trained for impairment classification, our experiments show good impairment classification results in an appropriate SQA latent representation. We investigate the clustering of impairments using various kinds of audio degradations that include different types of noises, waveform clipping, gain transition, pitch shift, compression, reverberation, etc. To visualize the clusters we perform classification of impairments in the SQA-latent representation domain using a standard k-nearest neighbor (kNN) classifier. We also develop a new DNN-based SQA model, named DNSMOS+, to examine whether an improvement in SQA leads to an improvement in impairment classification. The classification accuracy is 94% for LibriAugmented dataset with 16 types of impairments and 54% for ESC-50 dataset with 50 types of real noises.
Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee
ICASSP5
2025 Multivariate Probabilistic Assessment of Speech Quality
Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee
INTERSPEECH5
2024 DNSMOS Pro: A Reduced-Size DNN for Probabilistic MOS of Speech
Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee
INTERSPEECH5
2023 DeePMOS: Deep Posterior Mean-Opinion-Score of Speech
Fredrik Cumlin, Christian Schüldt, Saikat Chatterjee
INTERSPEECH3
2020 Performance Study of a Convolutional Time-Domain Audio Separation Network for Real-Time Speech Denoising
abstract
Time-domain audio separation networks based on dilated temporal convolutions have recently been shown to perform very well compared to methods that are based on a time-frequency representation in speech separation tasks, even outperforming an oracle binary time-frequency mask of the speakers. This paper investigates the performance of such a time-domain network (Conv-TasNet) for speech denoising in a real-time setting, comparing various parameter settings. Most importantly, different amounts of lookahead are evaluated and compared to the baseline of a fully causal model. We show that a large part of the increase in performance between a causal and non-causal model is achieved with a lookahead of only 20 milliseconds, demonstrating the usefulness of even small lookaheads for many real-time applications.
Samuel Sonning, Christian Schüldt, Hakan Erdogan, Scott Wisdom
ICASSP2
2019 Trigonometric Interpolation Beamforming for a Circular Microphone Array
abstract
Polynomial beamforming has previously been proposed for addressing the non-trivial problem of integrating acoustic echo cancellation with adaptive microphone beamforming. This paper demonstrates a design example for a circular array where traditional polynomial beamforming approaches exhibit severe (over 10 dB) directivity index (DI) oscillations at the edges of the design interval, leading to severe DI degradation for certain look directions. A solution, based on trigonometric interpolation, is proposed that stabilizes the oscillations significantly, resulting in a DI that deviates only about 1 dB from that of a fixed beamformer over all look directions.
Christian Schüldt
ICASSP1
2015 Noise robust integration for blind and non-blind reverberation time estimation
abstract
The estimation of the decay rate of a signal section is an integral component of both blind and non-blind reverberation time estimation methods. Several decay rate estimators have previously been proposed, based on, e.g., linear regression and maximum-likelihood estimation. Unfortunately, most approaches are sensitive to background noise, and/or are fairly demanding in terms of computational complexity. This paper presents a low complexity decay rate estimator, robust to stationary noise, for reverberation time estimation. Simulations using artificial signals, and experiments with speech in ventilation noise, demonstrate the performance and noise robustness of the proposed method.
Christian Schüldt, Peter Händel
ICASSP1
2014 Decay Rate Estimators and Their Performance for Blind Reverberation Time Estimation
abstract
Several approaches for blind estimation of reverberation time have been presented in the literature and decay rate estimation is an integral part of many, if not all, of such approaches. This paper provides both an analytical and experimental comparison, in terms of the bias and variance of three common decay rate estimators; a straight-forward linear regression approach as well as two maximum-likelihood based methods. Situations with and without interfering additive noise are considered. It is shown that the linear regression based approach is unbiased if no smoothing is applied, and that the estimation variance in the absence of noise is constantly about twice that of the maximum-likelihood based methods. It is shown that the methods that do not take possible noise into account suffer from similar estimation bias in the presence of noise. Further, a hybrid method, combining the noise robustness and low computational complexity advantages of the two different maximum-likelihood based methods, is presented.
Christian Schüldt, Peter Händel
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Voice radio communication, pedestrian localization, and the tactical use of 3D audio
abstract
The relation between voice radio communication and pedestrian localization is studied. 3D audio is identified as a linking technology which brings strong mutual benefits. Voice communication rendered with 3D audio provides a potential low secondary task interference user interface to the localization information. Vice versa, location information in the 3D audio provides spatial cues in the voice communication, improving speech intelligibility. An experimental setup with voice radio communication, cooperative pedestrian localization, and 3D audio is presented and we discuss high level tactical possibilities that the 3D audio brings. Finally, results of an initial experiment, demonstrating the effectiveness of the setup, are presented.
John-Olof Nilsson, Christian Schüldt, Peter Händel
IPIN2
2012 Robust low-complexity transfer logic for two-path echo cancellation
abstract
A well used approach for echo cancellation is the two-path method, where two adaptive filters in parallel are utilized. Typically, one filter is continuously updated, and when this filter is considered better adjusted to the echo-path than the other filter, the coefficients of the better adjusted filter is transferred to the other filter. When this transfer should occur is controlled by the transfer logic. This paper proposes transfer logic that is both more robust and more simple to tune, owing to fewer parameters, than the conventional approach. Extensive simulations show the advantages of the proposed method.
Christian Schüldt, Fredric Lindström, Ingvar Claesson
ICASSP1
2012 A Delay-Based Double-Talk Detector
abstract
When an adaptive filter is used for echo cancellation, it is essential to prevent the filter from diverging in situations when the echo signal is contaminated with near-end disturbance, i.e., during double-talk. This paper presents an extension of a previously proposed double-talk detector for improved performance. It is shown that the computational complexity of the proposed detector is lower than that of the well-used normalized cross correlation (NCC) double-talk detector, at the cost of performance. Further, it is shown that there can be a significant performance difference, in terms of detecting double-talk, between having a fixed echo cancellation filter, which is a common strategy in objective evaluation techniques, and an adaptive filter, which is more close to realistic conditions.
Christian Schüldt, Fredric Lindström, Ingvar Claesson
IEEE Trans. Speech Audio Process.1
2011 Low-Complexity Network Echo Cancellation Approach for Systems Equipped With External Memory
abstract
Long delays and sparseness characterize impulse responses in telecommunication networks and a vast number of solutions for network echo cancellation have been proposed over the years. In this paper, an approach for detecting dispersive regions of a sparse impulse response and a proportionate normalized least mean square (PNLMS)-based selective updating approach are combined with an adaptive double-talk detector to form a complete solution for echo cancellation. The proposed solution has low computational complexity and is targeted for systems equipped with external memory.
Magnus Berggren, Markus Borgh, Christian Schüldt, Fredric Lindström, Ingvar Claesson
IEEE ACM Trans. Audio Speech Lang. Process.3
2010 An improved deviation measure for two-path echo cancellation
abstract
Parallel adaptive filters have been proposed for echo cancellation to solve the dead-lock problem, occurring when the echo is detected as near-end speech after a severe echo-path change; causing the updating of the adaptive filter to halt. To control the parallel filters and monitor their performance, estimates of the filter deviation (i.e. the squared norm of the filter mismatch vector) are typically used. This paper presents a modification of a filter mismatch estimator. The proposed modification requires slightly more computational resources than the original measure, but provides a significant improvement in terms of robustness during double-talk. This is shown both analytically and through simulations.
Christian Schüldt, Fredric Lindström, Ingvar Claesson
ICASSP1
2009 Adaptive filter length selection for acoustic echo cancellation
Christian Schüldt, Fredric Lindström, Haibo Li 0001, Ingvar Claesson
Signal Process.1
2007 Local velocity-adapted motion events for spatio-temporal recognition
Ivan Laptev, Barbara Caputo, Christian Schüldt, Tony Lindeberg
Comput. Vis. Image Underst.3
2007 A hybrid acoustic echo canceller and suppressor
Fredric Lindström, Christian Schüldt, Ingvar Claesson
Signal Process.2
2007 An Improvement of the Two-Path Algorithm Transfer Logic for Acoustic Echo Cancellation
abstract
Adaptive filters for echo cancellation generally need update control schemes to avoid divergence in case of significant disturbances. The two-path algorithm avoids the problem of unnecessary halting of the adaptive filter when the control scheme gives an erroneous output. Versions of this algorithm have previously been presented for echo cancellation. This paper presents a transfer logic which improves the convergence speed of the two-path algorithm for acoustic echo cancellation, while retaining the robustness. Results from simulations show an improved performance, and a fixed-point DSP implementation verifies the performance in real-time
Fredric Lindström, Christian Schüldt, Ingvar Claesson
IEEE Trans. Speech Audio Process.2