EDBT 2026 Demo / reviewers in the wild / expert
Bence Mark Halpern
dblp:271/4266
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-8787-359XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancerabstractMeaningful speech assessment is vital in clinical phonetics and therapy monitoring. This study examined the link between perceptual speech assessments and objective acoustic measures in a large head and neck cancer (HNC) dataset. Trained listeners provided ratings of intelligibility, articulation, voice quality, phonation, speech rate, nasality, and background noise on speech. Strong correlations were found between subjective intelligibility, articulation, and voice quality, likely due to a shared underlying cause of speech symptoms in our speaker population. Objective measures of intelligibility and speech rate aligned with their subjective counterpart. Our results suggest that a single intelligibility measure may be sufficient for the clinical monitoring of speakers treated for HNC using concomitant chemoradiation. Bence Mark Halpern, Thomas Tienkamp, Teja Rebernik, R. J. J. H. van Son, Martijn Wieling 0001, Defne Abur, Tomoki Toda |
INTERSPEECH | 1 |
| 2024 | Quantifying the effect of speech pathology on automatic and human speaker verificationabstractThis study investigates how surgical intervention for speech pathology (specifically, as a result of oral cancer surgery) impacts the performance of an automatic speaker verification (ASV) system. Using two recently collected Dutch datasets with parallel pre and post-surgery audio from the same speaker, NKI-OC-VC and SPOKE, we assess the extent to which speech pathology influences ASV performance, and whether objective/subjective measures of speech severity are correlated with the performance. Finally, we carry out a perceptual study to compare judgements of ASV and human listeners. Our findings reveal that pathological speech negatively affects ASV performance, and the severity of the speech is negatively correlated with the performance. There is a moderate agreement in perceptual and objective scores of speaker similarity and severity, however, we could not clearly establish in the perceptual study, whether the same phenomenon also exists in human perception. Bence Mark Halpern, Thomas Tienkamp, Wen-Chin Huang, Lester Phillip Violeta, Teja Rebernik, Sebastiaan A. H. J. de Visscher, Max J. H. Witjes, Martijn Wieling 0001, Defne Abur, Tomoki Toda |
INTERSPEECH | 1 |
| 2024 | Towards inclusive automatic speech recognitionabstractPractice and recent evidence show that state-of-the-art (SotA) automatic speech recognition (ASR) systems do not perform equally well for all speaker groups. Many factors can cause this bias against different speaker groups. This paper, for the first time, systematically quantifies and finds speech recognition bias against gender, age, regional accents and non-native accents, and investigates the origin of this bias by investigating bias cross-lingually (i.e., Dutch and Mandarin) and for two different SotA ASR architectures (a hybrid DNN-HMM and an attention based end-to-end (E2E) model) through a phoneme error analysis. The results show that only a fraction of the bias can be explained by pronunciation differences between speaker groups, and that in order to mitigate bias, language- and architecture specific solutions need to be found. Siyuan Feng 0001, Bence Mark Halpern, Olya Kudina, Odette Scharenborg |
Comput. Speech Lang. | 2 |
| 2023 | Improving Severity Preservation of Healthy-to-Pathological Voice Conversion With Global Style TokensabstractIn healthy-to-pathological voice conversion (H2P-VC), healthy speech is converted into pathological while preserving the identity. The paper improves on previous two-stage approach to H2 P-VC where (1) speech is created first with the appropriate severity, (2) then the speaker identity of the voice is converted while preserving the severity of the voice. Specifically, we propose improvements to (2) by using phonetic posteriorgrams (PPG) and global style tokens (GST). Furthermore, we present a new dataset that contains parallel recordings of pathological and healthy speakers with the same identity which allows more precise evaluation. Listening tests by expert listeners show that the framework preserves severity of the source sample, while modelling target speaker’s voice. We also show that (a) pathology impacts x-vectors but not all speaker information is lost, (b) choosing source speakers based on severity labels alone is insufficient. Bence Mark Halpern, Wen-Chin Huang, Lester Phillip Violeta, R. J. J. H. van Son, Tomoki Toda |
ASRU | 1 |
| 2023 | Automatic evaluation of spontaneous oral cancer speech using ratings from naive listenersabstractIn this paper, we build and compare multiple speech systems for the automatic evaluation of the severity of a speech impairment due to oral cancer, based on spontaneous speech. To be able to build and evaluate such systems, we collected a new spontaneous oral cancer speech corpus from YouTube consisting of 124 utterances rated by 100 non-expert listeners and one trained speech-language pathologist, which we made publicly available. We evaluated the systems in two scenarios: a scenario where transcriptions were available (reference-based) and a scenario where transcriptions might not be available (reference-free). The results of extensive experiments showed that (1) when transcriptions were available, the highest correlation with the human severity ratings was obtained using an automatic speech recognition (ASR) retrained with oral cancer speech. (2) When transcriptions were not available, the best results were achieved by a LASSO model using modulation spectrum features. (3) We found that naive listeners’ ratings are highly similar to the speech pathologist’s ratings for speech severity evaluation. (4) The use of binary labels led to lower correlations of the automatic methods with the human ratings than using severity scores. Bence Mark Halpern, Siyuan Feng 0001, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg |
Speech Commun. | 1 |
| 2022 | Towards Identity Preserving Normal to Dysarthric Voice ConversionabstractWe present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical decision making processes and alleviation of patient stress, (2) data augmentation for dysarthric speech recognition. This is an especially challenging task since the converted samples should capture the severity of dysarthric speech while being highly natural and possessing the speaker identity of the normal speaker. To this end, we adopted a two-stage framework, which consists of a sequence-to-sequence model and a nonparallel frame-wise model. Objective and subjective evaluations were conducted on the UASpeech dataset, and results showed that the method was able to yield reasonable naturalness and capture severity aspects of the pathological speech. On the other hand, the similarity to the normal source speaker’s voice was limited and requires further improvements. Wen-Chin Huang, Bence Mark Halpern, Lester Phillip Violeta, Odette Scharenborg, Tomoki Toda |
ICASSP | 2 |
| 2022 | The Effectiveness of Time Stretching for Enhancing Dysarthric Speech for Improved Dysarthric Speech RecognitionabstractIn this paper, we investigate several existing and a new state-of-the-art generative adversarial network-based (GAN) voice conversion method for enhancing dysarthric speech for improved dysarthric speech recognition. We compare key components of existing methods as part of a rigorous ablation study to find the most effective solution to improve dysarthric speech recognition. We find that straightforward signal processing methods such as stationary noise removal and vocoder-based time stretching lead to dysarthric speech recognition results comparable to those obtained when using state-of-the-art GAN-based voice conversion methods as measured using a phoneme recognition task. Additionally, our proposed solution of a combination of MaskCycleGAN-VC and time stretching is able to improve the phoneme recognition results for certain dysarthric speakers compared to our time stretched baseline. Luke Prananta, Bence Mark Halpern, Siyuan Feng 0001, Odette Scharenborg |
INTERSPEECH | 2 |
| 2022 | Mitigating bias against non-native accentsabstractAutomatic Speech Recognition (ASR) systems have seen substantial improvements in the past decade; however, not for all speaker groups. Recent research shows that bias exists against different types of speech, including non-native accents, in state-of-the-art (SOTA) ASR systems. To attain inclusive speech recognition, i.e., ASR for everyone irrespective of how one speaks or the accent one has, bias mitigation is essential and necessary. In this thesis, two SOTA ASR systems (one is based on the recurrent neural network (RNN) and the other is based on the transformer architecture) are built to uncover and quantify the bias against non-native accents. Here I focus on bias mitigation against non-native accents using two different approaches: data augmentation and by using more effective training methods. For data augmentation, an autoencoder-based cross-lingual voice conversion (VC) model is used to increase the amount of non-native accented speech training data in addition to data augmentation through speed perturbation. Moreover, I investigate two training methods, i.e., fine-tuning and Domain Adversarial Training (DAT), to see whether they can utilize the available non-native accented speech data more effectively than a standard training approach. Experimental results show for the transformer-based ASR model: (1) adding VC-generated and speed-perturbed data to train the ASR model gives the best bias mitigation performance and the lowest word error rate (WER); (2) fine-tuning reduces the bias against non-native accents but at the cost of native accent performance; and (3) compared with the standard training method, DAT does not leads to further bias reduction. While for the RNN-based ASR model, all the 4 bias mitigation approaches do not show obvious benefits. Bence Mark Halpern, Tanvina Patel, Odette Scharenborg |
INTERSPEECH | 3 |
| 2022 | Low-resource automatic speech recognition and error analyses of oral cancer speechabstractIn this paper, we introduce a new corpus of oral cancer speech and present our study on the automatic recognition and analysis of oral cancer speech. A two-hour English oral cancer speech dataset is collected from YouTube. Formulated as a low-resource oral cancer ASR task, we investigate three acoustic modelling approaches that previously have worked well with low-resource scenarios using two different architectures; a hybrid architecture and a transformer-based end-to-end (E2E) model: (1) a retraining approach; (2) a speaker adaptation approach; and (3) a disentangled representation learning approach (only using the hybrid architecture). The approaches achieve a (1) 4.7% (hybrid) and 7.5% (E2E); (2) 7.7%; and (3) 2.0% absolute word error rate reduction, respectively, compared to a baseline system which is not trained on oral cancer speech. A detailed analysis of the speech recognition results shows that (1) plosives and certain vowels are the most difficult sounds to recognise in oral cancer speech — this problem is successfully alleviated by our proposed approaches; (3) however these sounds are also relatively poorly recognised in the case of healthy speech with the exception of/p/. (2) recognition performance of certain phonemes is strongly data-dependent; (4) In terms of the manner of articulation, E2E performs better with the exception of vowels — however, vowels have a large contribution to overall performance. As for the place of articulation, vowels, labiodentals, dentals and glottals are better captured by hybrid models, E2E is better on bilabial, alveolar, postalveolar, palatal and velar information. (5) Finally, our analysis provides some guidelines for selecting words that can be used as voice commands for ASR systems for oral cancer speakers. Bence Mark Halpern, Siyuan Feng 0001, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg |
Speech Commun. | 1 |
| 2020 | Detecting and Analysing Spontaneous Oral Cancer Speech in the WildabstractOral cancer speech is a disease which impacts more than half a million people worldwide every year. Analysis of oral cancer speech has so far focused on read speech. In this paper, we 1) present and 2) analyse a three-hour long spontaneous oral cancer speech dataset collected from YouTube. 3) We set baselines for an oral cancer speech detection task on this dataset. The analysis of these explainable machine learning baselines shows that sibilants and stop consonants are the most important indicators for spontaneous oral cancer speech detection. Bence Mark Halpern, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg |
INTERSPEECH | 1 |