Marvin Sach

dblp:356/4025 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-2909-6539ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Less is More: Data Curation Matters in Scaling Speech Enhancement
abstract
The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in “clean” training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500 -hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
Chenda Li, Wangyou Zhang, Wei Wang 0010, Robin Scheibler, Kohei Saijo, Samuele Cornell, Yihui Fu, Marvin Sach, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian
ASRU8
2025 URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
abstract
The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK’s superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data.
Chenda Li, Wei Wang 0010, Wangyou Zhang, Samuele Cornell, Marvin Sach, Robin Scheibler, Kohei Saijo, Yihui Fu, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian
ASRU6
2025 Interspeech 2025 URGENT Speech Enhancement Challenge
Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Yihui Fu, Wei Wang 0010, Tim Fingscheidt, Shinji Watanabe 0001
INTERSPEECH8
2025 Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
Wangyou Zhang, Kohei Saijo, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Wei Wang 0010, Yihui Fu, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian
INTERSPEECH8
2024 Employing Real Training Data for Deep Noise Suppression
abstract
Most deep noise suppression (DNS) models are trained with reference-based losses requiring access to clean speech. However, sometimes an additive microphone model is insufficient for real-world applications. Accordingly, ways to use real training data in supervised learning for DNS models promise to reduce a potential training/inference mismatch. Employing real data for DNS training requires either generative approaches or a reference-free loss without access to the corresponding clean speech. In this work, we propose to employ an end-to-end non-intrusive deep neural network (DNN), named PESQ-DNN, to estimate perceptual evaluation of speech quality (PESQ) scores of enhanced real data. It provides a reference-free perceptual loss for employing real data during DNS training, maximizing the PESQ scores. Furthermore, we use an epoch-wise alternating training protocol, updating the DNS model on real data, followed by PESQ-DNN updating on synthetic data. The DNS model trained with the PESQ-DNN employing real data outperforms all reference methods employing only synthetic training data. On synthetic test data, our proposed method excels the Inter-speech 2021 DNS Challenge baseline by a significant 0.32 PESQ points. Both on synthetic and real test data, the proposed method beats the baseline by 0.05 DNSMOS points – although PESQ-DNN optimizes for a different perceptual metric.
Marvin Sach, Jan Pirklbauer, Tim Fingscheidt
ICASSP2
2024 URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Jan Pirklbauer, Marvin Sach, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian
INTERSPEECH8
2023 An Efficient and Noise-Robust Audiovisual Encoder for Audiovisual Speech Recognition
Chenwei Liang, Timo Lohrenz, Marvin Sach, Björn Möller, Tim Fingscheidt
INTERSPEECH4
2023 EffCRN: An Efficient Convolutional Recurrent Network for High-Performance Speech Enhancement
Marvin Sach, Jan Franzen, Bruno Defraene, Kristoff Fluyt, Maximilian Strake, Wouter Tirry, Tim Fingscheidt
INTERSPEECH1