VLDB 2026 Research / reviewers in the wild / expert
Enno Hermann
dblp:217/2123
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-9945-1140ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
Karl El Hajal, Enno Hermann, Sevada Hovsepyan, Mathew Magimai-Doss |
INTERSPEECH | 2 |
| 2025 | Children's Voice Privacy: First Steps and Emerging ChallengesabstractInternational audience Ajinkya Kulkarni, Francisco Teixeira, Enno Hermann, Thomas Rolland, Isabel Trancoso, Mathew Magimai-Doss |
INTERSPEECH | 3 |
| 2024 | Towards interfacing large language models with ASR systems using confidence measures and prompting
Maryam Naderi, Enno Hermann, Alexandre Nanchen, Sevada Hovsepyan, Mathew Magimai-Doss |
INTERSPEECH | 2 |
| 2023 | Few-shot Dysarthric Speech Recognition with Text-to-Speech Data Augmentation
Enno Hermann, Mathew Magimai-Doss |
INTERSPEECH | 1 |
| 2023 | Using Commercial ASR Solutions to Assess Reading Skills in Children: A Case ReportabstractReading is an acquired skill that is essential for integrating and participating in today's society.Yet, becoming literate can be particularly laborious for some children.Identifying reading difficulties early enough is the first, necessary step toward remediation.Here we investigate the opportunities and limitations of integrating commercial, off-the-shelf automatic speech recognition (ASR) services from IBM Watson to ease the administration and evaluation of children's reading assessment tests in French and Italian. Timothy Piton, Enno Hermann, Angela Pasqualotto, Marjolaine Cohen, Mathew Magimai-Doss, Daphne Bavelier |
INTERSPEECH | 2 |
| 2021 | Handling Acoustic Variation in Dysarthric Speech Recognition Systems Through Model CombinationabstractDeveloping automatic speech recognition (ASR) systems that recognise dysarthric speech as well as control speech from unimpaired speakers remains challenging. Including more highly variable dysarthric speech during training can also negatively affect the performance on control speakers, which is not desirable when developing speech recognisers for a wider audience. In this work, we analyse how the acoustic variability of dysarthric speech affects ASR systems and propose the combination of multiple acoustic models trained on different subsets of speakers to mitigate this effect. This approach shows improvements for both dysarthric and control speakers on the Torgo and UA-Speech corpora. Enno Hermann, Mathew Magimai-Doss |
Interspeech | 1 |
| 2021 | Multilingual and unsupervised subword modeling for zero-resource languages
Enno Hermann, Herman Kamper, Sharon Goldwater |
Comput. Speech Lang. | 1 |
| 2020 | Dysarthric Speech Recognition with Lattice-Free MMIabstractRecognising dysarthric speech is a challenging problem as it differs in many aspects from typical speech, such as speaking rate and pronunciation. In the literature the focus so far has largely been on handling these variabilities in the framework of HMM/GMM and cross-entropy based HMM/DNN systems. This paper focuses on the use of state-of-the-art sequence-discriminative training, in particular lattice-free maximum mutual information (LF-MMI), for improving dysarthric speech recognition. Through a systematic investigation on the Torgo corpus we demonstrate that LF-MMI performs well on such atypical data and compensates much better for the low speaking rates of dysarthric speakers than conventionally trained systems. This can be attributed to inherent aspects of current speech recognition training regimes, like frame subsampling and speed perturbation, which obviate the need for some techniques previously adopted specifically for dysarthric speech. Enno Hermann, Mathew Magimai-Doss |
ICASSP | 1 |
| 2018 | Multilingual Bottleneck Features for Subword Modeling in Zero-resource LanguagesabstractHow can we effectively develop speech technology for languages where no transcribed data is available? Many existing approaches use no annotated resources at all, yet it makes sense to leverage information from large annotated corpora in other languages, for example in the form of multilingual bottleneck features (BNFs) obtained from a supervised speech recognition system. In this work, we evaluate the benefits of BNFs for subword modeling (feature extraction) in six unseen languages on a word discrimination task. First we establish a strong unsupervised baseline by combining two existing methods: vocal tract length normalisation (VTLN) and the correspondence autoencoder (cAE). We then show that BNFs trained on a single language already beat this baseline; including up to 10 languages results in additional improvements which cannot be matched by just adding more data from a single language. Finally, we show that the cAE can improve further on the BNFs if high-quality same-word pairs are available. Enno Hermann, Sharon Goldwater |
INTERSPEECH | 1 |