VLDB 2026 Research / reviewers in the wild / expert
Odette Scharenborg
dblp:87/1365
· DBLP profile ↗
94ranked-venue papers
32as first author
38since 2021 · last 2025
0000-0003-0693-8852ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 29 first-author · 29 since 2021Artificial intelligence and machine learning · 67 · 27 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven ApplicationsabstractIn this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing loudly. The loudspeaker beamformer modifies the loudspeaker playback signals to create a low-acoustic-energy region around the device that implements automatic speech recognition for a voice driven application (VDA). The algorithm utilises a distortion measure based on human auditory perception to limit the distortion perceived by human listeners. Simulations and real-world experiments show that the proposed loudspeaker beamformer improves the speech recognition performance in all tested scenarios. Moreover, the algorithm allows to further reduce the acoustic energy around the VDA device at the expense of reduced objective audio quality at the listener’s location. Dimme de Groot, Baturalp Karslioglu, Odette Scharenborg, Jorge Martínez 0002 |
ICASSP | 3 |
| 2025 | The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Shilong Wu, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001, Shinji Watanabe 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg |
INTERSPEECH | 9 |
| 2025 | Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric SpeechabstractDysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance. Dimme de Groot, Tanvina Patel, Devendra Kayande, Odette Scharenborg, Zhengjun Yue |
INTERSPEECH | 4 |
| 2025 | Challenges and practical guidelines for atypical speech data collection, annotation, usage and sharing: A multi-project perspectiveabstractContains fulltext : 325867.pdf (Publisher’s version ) (Open Access) Zhengjun Yue, Mara Barberis, Tanvina Patel, Judith Dineley, Willemijn Doedens, Lottie Stipdonk, Elke De Witte, Erfan Loweimi, Hugo Van hamme, Djaina Satoer, Marina B. Ruiter, Laureano Moro-Velázquez, Nicholas Cummins, Odette Scharenborg |
INTERSPEECH | 15 |
| 2024 | Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation PipelineabstractThe growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-automated dataset creation pipeline that leverages large language models. We use this pipeline to generate a dataset of speakers identifying themself or another speaker as belonging to a particular race, ethnicity, or national origin group. We use OpenaAI’s GPT-4 to perform two complex annotation tasks- separating files relevant to our intended dataset from the irrelevant ones (filtering) and finding and extracting information on identifications within a transcript (tagging). By evaluating GPT-4’s performance using human annotations as ground truths, we show that it can reduce resources required by dataset annotation while barely losing any important information. For the filtering task, GPT-4 had a very low miss rate of 6.93%. GPT-4’s tagging performance showed a trade-off between precision and recall, where the latter got as high as 97%, but precision never exceeded 45%. Our approach reduces the time required for the filtering and tagging tasks by 95% and 80%, respectively. We also present an in-depth error analysis of GPT-4’s performance. Maliha Jahan, Helin Wang, Thomas Thebaud, Yinglun Sun, Giang Ha Le, Zsuzsanna Fagyal, Odette Scharenborg, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC/COLING | 7 |
| 2024 | The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker ExtractionabstractPrevious Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompted a shift in focus towards the Audio-Visual Target Speaker Extraction (AVTSE) task for the MISP 2023 challenge in ICASSP 2024 Signal Processing Grand Challenges. Unlike existing audio-visual speech enhancement challenges primarily focused on simulation data, the MISP 2023 challenge uniquely explores how front-end speech processing, combined with visual clues, impacts back-end tasks in real-world scenarios. This pioneering effort aims to set the first benchmark for the AVTSE task, offering fresh insights into enhancing the accuracy of back-end speech recognition systems through AVTSE in challenging and real acoustic environments. This paper delivers a thorough overview of the task setting, dataset, and baseline system of the MISP 2023 challenge. It also includes an in-depth analysis of the challenges participants may encounter. The experimental results highlight the demanding nature of this task, and we look forward to the innovative solutions participants will bring forward. Shilong Wu, Hang Chen 0001, Yusheng Dai, Chenyue Zhang, Ruoyu Wang 0029, Hongbo Lan, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg, Zhongqiu Wang 0001, Jianqing Gao |
ICASSP | 12 |
| 2024 | Using articulated speech EEG signals for imagined speech decodingabstractBrain-Computer Interfaces (BCIs) open avenues for communication among individuals unable to use voice or gestures. Silent speech interfaces are one such approach for BCIs that could offer a trans- formative means of connecting with the external world. Performance on imagined speech decoding however is rather low due to, amongst others, data scarcity and the lack of a clear starting point of the imagined speech in the brain signal. We investigate whether using electroencephalography (EEG) signals from articulated speech can be used to improve imagined speech decoding in two ways: we investigate whether articulated speech EEG signals can be used to predict the end point of the imagined speech and use the articulated speech EEG as extra training data for speaker-independent imagined vowel classification. Our results show that using EEG data from articulated speech did not improve classification of vowels in imagined speech, probably due to high variability in EEG signals amongst speakers. Chris Bras, Tanvina Patel, Odette Scharenborg |
INTERSPEECH | 3 |
| 2024 | Self-supervised Speech Representations Still Struggle with African American Vernacular English
Kalvin Chang, Yi-Hui Chou, Jiatong Shi, Hsuan-Ming Chen, Nicole R. Holliday, Odette Scharenborg, David R. Mortensen |
INTERSPEECH | 6 |
| 2024 | As Biased as You Measure: Methodological Pitfalls of Bias Evaluations in Speaker Verification Research
Wiebke Hutiri, Tanvina Patel, Aaron Yi Ding, Odette Scharenborg |
INTERSPEECH | 4 |
| 2024 | Improving child speech recognition with augmented child-like speech
Zhengjun Yue, Tanvina Patel, Odette Scharenborg |
INTERSPEECH | 4 |
| 2024 | Towards inclusive automatic speech recognitionabstractPractice and recent evidence show that state-of-the-art (SotA) automatic speech recognition (ASR) systems do not perform equally well for all speaker groups. Many factors can cause this bias against different speaker groups. This paper, for the first time, systematically quantifies and finds speech recognition bias against gender, age, regional accents and non-native accents, and investigates the origin of this bias by investigating bias cross-lingually (i.e., Dutch and Mandarin) and for two different SotA ASR architectures (a hybrid DNN-HMM and an attention based end-to-end (E2E) model) through a phoneme error analysis. The results show that only a fraction of the bias can be explained by pronunciation differences between speaker groups, and that in order to mitigate bias, language- and architecture specific solutions need to be found. Siyuan Feng 0001, Bence Mark Halpern, Olya Kudina, Odette Scharenborg |
Comput. Speech Lang. | 4 |
| 2023 | Improving Adaptive Learning Models Using Prosodic Speech Features
Thomas Wilschut, Florian Sense, Odette Scharenborg, Hedderik van Rijn |
AIED | 3 |
| 2023 | Improving Whispered Speech Recognition Performance Using Pseudo-Whispered Based Data AugmentationabstractWhispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and the scarcity of adequate training data leads to low automatic speech recognition (ASR) performance. To address the data scarcity issue, we use a signal processing-based technique that transforms the spectral characteristics of normal speech to those of pseudo-whispered speech. We augment an End-to-End ASR with pseudo-whispered speech and achieve an 18.2 % relative reduction in word error rate for whispered speech compared to the baseline. Results for the individual speaker groups in the wTIMIT database show the best results for US English. Further investigation showed that the lack of glottal information in whispered speech has the largest impact on whispered speech ASR performance. Zhaofeng Lin, Tanvina Patel, Odette Scharenborg |
ASRU | 3 |
| 2023 | Exploring Data Augmentation in Bias Mitigation Against Non-Native-Accented SpeechabstractAutomatic speech recognition (ASR) should serve every speaker, not only the majority “standard” speakers of a language. In order to build inclusive ASR, mitigating the bias against speaker groups who speak in a “non-standard” or “diverse” way is crucial. We aim to mitigate the bias against non-native-accented Flemish in a Flemish ASR system. Since this is a low-resource problem, we investigate the optimal type of data augmentation, i.e., speed/pitch perturbation, cross-lingual voice conversion-based methods, and SpecAugment, applied to both native Flemish and non-native-accented Flemish, for bias mitigation. The results showed that specific types of data augmentation applied to both native and non-native-accented speech improve non-native-accented ASR while applying data augmentation to the non-native-accented speech is more conducive to bias reduction. Combining both gave the largest bias reduction for human-machine interaction (HMI) as well as read-type speech. Aaricia Herygers, Tanvina Patel, Zhengjun Yue, Odette Scharenborg |
ASRU | 5 |
| 2023 | Summary on the Multimodal Information Based Speech Processing (MISP) 2022 ChallengeabstractThe Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visual diarization and recognition (AVDR). The training material was based on previous MISP 2021 recordings, but we have accurately synchronized audio and visual data. Additionally, a new evaluation set was provided. This paper gives an overview of the challenge setup, presents the results, and summarizes the effective techniques employed by the participants. We also analyze the current technical challenges and suggest directions for future research in AVSD and AVDR. Hang Chen 0001, Shilong Wu, Yusheng Dai, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 10 |
| 2023 | DAIS: The Delft Database of EEG Recordings of Dutch Articulated and Imagined SpeechabstractSilent speech interfaces could enable people who lost the ability to use their voice or gestures to communicate with the external world, e.g., through decoding the person’s brain signals when imagining speech. Only a few and small databases exist that allow for the development and training of brain computer interfaces (BCIs) that can decode imagined speech from recorded brain signals. Here, we present an open database consisting of electroencephalography (EEG) and speech data from 20 participants recorded during the covert (imagined) and actual articulation of 15 Dutch prompts. A validation speaker-independent classification experiment using a ResNet-50 model with spatial-spectral-temporal features extracted from the EEG signals obtained an average accuracy of 70.6% for the classification of rest vs. covert vs. articulated speech trials. This and observed structural differences in the EEG signals between covert and articulated speech demonstrate that the EEG signals in the three classes contain discriminative information. Bo Dekker, Alfred C. Schouten, Odette Scharenborg |
ICASSP | 3 |
| 2023 | The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And RecognitionabstractThe Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two tracks: 1) audio-visual speaker diarization (AVSD), aiming to solve "who spoken when" using both audio and visual data; 2) a novel audio-visual diarization and recognition (AVDR) task that focuses on addressing "who spoken what when" with audio-visual speaker diarization results. Both tracks focus on the Chinese language, and use far-field audio and video in real home-tv scenarios: 2-6 people communicating each other with TV noise in the background. This paper introduces the dataset, track settings, and baselines of the MISP2022 challenge. Our analyses of experiments and examples indicate the good performance of AVDR baseline system, and the potential difficulties in this challenge due to, e.g., the far-field video quality, the presence of TV noise in the background, and the indistinguishable speakers. Shilong Wu, Hang Chen 0001, Maokui He, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 10 |
| 2023 | Automatic evaluation of spontaneous oral cancer speech using ratings from naive listenersabstractIn this paper, we build and compare multiple speech systems for the automatic evaluation of the severity of a speech impairment due to oral cancer, based on spontaneous speech. To be able to build and evaluate such systems, we collected a new spontaneous oral cancer speech corpus from YouTube consisting of 124 utterances rated by 100 non-expert listeners and one trained speech-language pathologist, which we made publicly available. We evaluated the systems in two scenarios: a scenario where transcriptions were available (reference-based) and a scenario where transcriptions might not be available (reference-free). The results of extensive experiments showed that (1) when transcriptions were available, the highest correlation with the human severity ratings was obtained using an automatic speech recognition (ASR) retrained with oral cancer speech. (2) When transcriptions were not available, the best results were achieved by a LASSO model using modulation spectrum features. (3) We found that naive listeners’ ratings are highly similar to the speech pathologist’s ratings for speech severity evaluation. (4) The use of binary labels led to lower correlations of the automatic methods with the human ratings than using severity scores. Bence Mark Halpern, Siyuan Feng 0001, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg |
Speech Commun. | 5 |
| 2023 | AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary PersonsabstractAutomatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate synchronized speech and talking-head videos on the basis of text and a single face image of an arbitrary person as input. In contrast to previous text-driven talking head generation methods, which can only synthesize the voice of a specific person, the proposed method is capable of synthesizing speech for any person. Specifically, the proposed method decomposes the generation of synchronized speech and talking head videos into two stages, i.e., a text-to-speech (TTS) stage and a speech-driven talking head generation stage. The proposed TTS module is a face-conditioned multi-speaker TTS model that gets the speaker identity information from face images instead of speech, which allows us to synthesize a personalized voice on the basis of the input face image. To generate the talking head videos from the face images, a facial landmark-based method that can predict both lip movements and head rotations is proposed. Extensive experiments demonstrate that the proposed method is able to generate synchronized speech and talking head videos for arbitrary persons, in which the timbre of the synthesized voice is in harmony with the input face, and the proposed landmark-based talking head method outperforms the state-of-the-art landmark-based method on generating natural talking head videos. Qicong Xie, Jihua Zhu, Lei Xie 0001, Odette Scharenborg |
IEEE Trans. Multim. | 5 |
| 2022 | The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And ResultsabstractIn this paper we discuss the rational of the Multi-model Information based Speech Processing (MISP) Challenge, and provide a detailed description of the data recorded, the two evaluation tasks and the corresponding baselines, followed by a summary of submitted systems and evaluation results. The MISP Challenge aims at tack-ling speech processing tasks in different scenarios by introducing information about an additional modality (e.g., video, or text), which will hopefully lead to better environmental and speaker robustness in realistic applications. In the first MISP challenge, two bench-mark datasets recorded in a real-home TV room with two reproducible open-source baseline systems have been released to promote research in audio-visual wake word spotting (AVWWS) and audio-visual speech recognition (AVSR). To our knowledge, MISP is the first open evaluation challenge to tackle real-world issues of AVWWS and AVSR in the home TV scenario. Hang Chen 0001, Hengshun Zhou, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 8 |
| 2022 | Towards Identity Preserving Normal to Dysarthric Voice ConversionabstractWe present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical decision making processes and alleviation of patient stress, (2) data augmentation for dysarthric speech recognition. This is an especially challenging task since the converted samples should capture the severity of dysarthric speech while being highly natural and possessing the speaker identity of the normal speaker. To this end, we adopted a two-stage framework, which consists of a sequence-to-sequence model and a nonparallel frame-wise model. Objective and subjective evaluations were conducted on the UASpeech dataset, and results showed that the method was able to yield reasonable naturalness and capture severity aspects of the pathological speech. On the other hand, the similarity to the normal source speaker’s voice was limited and requires further improvements. Wen-Chin Huang, Bence Mark Halpern, Lester Phillip Violeta, Odette Scharenborg, Tomoki Toda |
ICASSP | 4 |
| 2022 | Audio-Visual Speech Recognition in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we present the updated Audio-Visual Speech Recognition (AVSR) corpus of MISP2021 challenge, a large-scale audio-visual Chinese conversational corpus consisting of 141h audio and video data collected by far/middle/near microphones and far/middle cameras in 34 real-home TV rooms. To our best knowledge, our corpus is the first distant multi-microphone conversational Chinese audio-visual corpus and the first large vocabulary continuous Chinese lip-reading dataset in the adverse home-tv scenario. Moreover, we make a deep analysis of the corpus and conduct a comprehensive ablation study of all audio and video data in the audio-only/video-only/audiovisual systems. Error analysis shows video modality supplement acoustic information degraded by noise to reduce deletion errors and provide discriminative information in overlapping speech to reduce substitution errors. Finally, we also design a set of experiments such as frontend, data augmentation and end-to-end models for providing the direction of potential future work. The corpus and the code are released to promote the research not only in speech area but also for the computer vision area and cross-disciplinary research. Hang Chen 0001, Jun Du 0002, Yusheng Dai, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen |
INTERSPEECH | 7 |
| 2022 | Using cross-model learnings for the Gram Vaani ASR Challenge 2022abstractIn the diverse and multilingual land of India, Hindi is spoken as a first language by a majority of its population. Efforts are made to obtain data in terms of audio, transcriptions, dictionary, etc. to develop speech-technology applications in Hindi. Similarly, the Gram-Vaani ASR Challenge 2022 provides spontaneous telephone speech, with natural back-ground and regional variations in Hindi. The challenge provides: 100 hours of labeled train-set, 5 hours of labeled dev-set and 1000 hours of unlabeled data-set. For the 'Closed Challenge', we trained an End-to-End (E2E) Conformer model using speed perturbations, SpecAugment techniques and use VTLN to handle any unknown speaker groups in the blind evaluation set. On the dev-set, we achieved a 30.3% WER compared to the 34.8% WER by the Challenge E2E baseline. For the 'Self Supervised Closed Challenge', a semi-supervised learning approach is used. We generate pseudo-transcripts for the unlabeled data using a hybrid TDNN-3gram LM model and trained an E2E model. This is then used as a seed for retraining the E2E model with high confidence data. Cross-model learning and refining of the E2E model gave 25.3% WER on the dev-set compared to ∼33-35% WER by the Challenge baseline that use wav2vec models. Tanvina Patel, Odette Scharenborg |
INTERSPEECH | 2 |
| 2022 | The Effectiveness of Time Stretching for Enhancing Dysarthric Speech for Improved Dysarthric Speech RecognitionabstractIn this paper, we investigate several existing and a new state-of-the-art generative adversarial network-based (GAN) voice conversion method for enhancing dysarthric speech for improved dysarthric speech recognition. We compare key components of existing methods as part of a rigorous ablation study to find the most effective solution to improve dysarthric speech recognition. We find that straightforward signal processing methods such as stationary noise removal and vocoder-based time stretching lead to dysarthric speech recognition results comparable to those obtained when using state-of-the-art GAN-based voice conversion methods as measured using a phoneme recognition task. Additionally, our proposed solution of a combination of MaskCycleGAN-VC and time stretching is able to improve the phoneme recognition results for certain dysarthric speakers compared to our time stretched baseline. Luke Prananta, Bence Mark Halpern, Siyuan Feng 0001, Odette Scharenborg |
INTERSPEECH | 4 |
| 2022 | Mitigating bias against non-native accentsabstractAutomatic Speech Recognition (ASR) systems have seen substantial improvements in the past decade; however, not for all speaker groups. Recent research shows that bias exists against different types of speech, including non-native accents, in state-of-the-art (SOTA) ASR systems. To attain inclusive speech recognition, i.e., ASR for everyone irrespective of how one speaks or the accent one has, bias mitigation is essential and necessary. In this thesis, two SOTA ASR systems (one is based on the recurrent neural network (RNN) and the other is based on the transformer architecture) are built to uncover and quantify the bias against non-native accents. Here I focus on bias mitigation against non-native accents using two different approaches: data augmentation and by using more effective training methods. For data augmentation, an autoencoder-based cross-lingual voice conversion (VC) model is used to increase the amount of non-native accented speech training data in addition to data augmentation through speed perturbation. Moreover, I investigate two training methods, i.e., fine-tuning and Domain Adversarial Training (DAT), to see whether they can utilize the available non-native accented speech data more effectively than a standard training approach. Experimental results show for the transformer-based ASR model: (1) adding VC-generated and speed-perturbed data to train the ASR model gives the best bias mitigation performance and the lowest word error rate (WER); (2) fine-tuning reduces the bias against non-native accents but at the cost of native accent performance; and (3) compared with the standard training method, DAT does not leads to further bias reduction. While for the RNN-based ASR model, all the 4 bias mitigation approaches do not show obvious benefits. Bence Mark Halpern, Tanvina Patel, Odette Scharenborg |
INTERSPEECH | 5 |
| 2022 | Audio-Visual Wake Word Spotting in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we describe and release publicly the audio-visual wake word spotting (WWS) database in the MISP2021 Challenge, which covers a range of scenarios of audio and video data collected by near-, mid-, and far-field microphone arrays, and cameras, to create a shared and publicly available database for WWS. The database and the code 2 are released, which will be a valuable addition to the community for promoting WWS research using multi-modality information in realistic and complex conditions. Moreover, we investigated the different data augmentation methods for single modalities on an end-to-end WWS network. A set of audio-visual fusion experiments and analysis were conducted to observe the assistance from visual information to acoustic information based on different audio and video field configurations. The results showed that the fusion system generally improves over the single-modality (audio- or video-only) system, especially under complex noisy conditions. Hengshun Zhou, Jun Du 0002, Gongzhen Zou, Zhaoxu Nian, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen, Shifu Xiong, Jianqing Gao |
INTERSPEECH | 8 |
| 2022 | Discovering phonetic inventories with crosslingual automatic speech recognition
Piotr Zelasko, Siyuan Feng 0001, Laureano Moro-Velázquez, Ali Abavisani, Saurabhchand Bhati, Odette Scharenborg, Mark Hasegawa-Johnson, Najim Dehak |
Comput. Speech Lang. | 6 |
| 2022 | Low-resource automatic speech recognition and error analyses of oral cancer speechabstractIn this paper, we introduce a new corpus of oral cancer speech and present our study on the automatic recognition and analysis of oral cancer speech. A two-hour English oral cancer speech dataset is collected from YouTube. Formulated as a low-resource oral cancer ASR task, we investigate three acoustic modelling approaches that previously have worked well with low-resource scenarios using two different architectures; a hybrid architecture and a transformer-based end-to-end (E2E) model: (1) a retraining approach; (2) a speaker adaptation approach; and (3) a disentangled representation learning approach (only using the hybrid architecture). The approaches achieve a (1) 4.7% (hybrid) and 7.5% (E2E); (2) 7.7%; and (3) 2.0% absolute word error rate reduction, respectively, compared to a baseline system which is not trained on oral cancer speech. A detailed analysis of the speech recognition results shows that (1) plosives and certain vowels are the most difficult sounds to recognise in oral cancer speech — this problem is successfully alleviated by our proposed approaches; (3) however these sounds are also relatively poorly recognised in the case of healthy speech with the exception of/p/. (2) recognition performance of certain phonemes is strongly data-dependent; (4) In terms of the manner of articulation, E2E performs better with the exception of vowels — however, vowels have a large contribution to overall performance. As for the place of articulation, vowels, labiodentals, dentals and glottals are better captured by hybrid models, E2E is better on bilabial, alveolar, postalveolar, palatal and velar information. (5) Finally, our analysis provides some guidelines for selecting words that can be used as voice commands for ASR systems for oral cancer speakers. Bence Mark Halpern, Siyuan Feng 0001, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg |
Speech Commun. | 5 |
| 2021 | The Effects of Onset and Offset Masking on the Time Course of Non-Native Spoken-Word Recognition in Noise
Florian Hintz, Cesko Voeten, James M. McQueen, Odette Scharenborg |
CogSci | 4 |
| 2021 | How Phonotactics Affect Multilingual and Zero-Shot ASR PerformanceabstractThe idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcriptions of languages presented during training. However, the representations it learned were not successful in zero-shot transfer to unseen languages. Because that model lacks an explicit factorization of the acoustic model (AM) and language model (LM), it is unclear to what degree the performance suffered from differences in pronunciation or the mismatch in phono-tactics. To gain more insight into the factors limiting zero-shot ASR transfer, we replace the encoder-decoder with a hybrid ASR system consisting of a separate AM and LM. Then, we perform an extensive evaluation of monolingual, multilingual, and crosslingual (zero-shot) acoustic and language models on a set of 13 phonetically diverse languages. We show that the gain from modeling crosslingual phonotactics is limited, and imposing a too strong model can hurt the zero-shot transfer. Furthermore, we find that a multilingual LM hurts a multilingual ASR system’s performance, and retaining only the target language’s phonotactic data in LM training is preferable. Siyuan Feng 0001, Piotr Zelasko, Laureano Moro-Velázquez, Ali Abavisani, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 6 |
| 2021 | Show and Speak: Directly Synthesize Spoken Description of ImagesabstractThis paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the need for any text or phonemes. The basic structure of SAS is an encoder-decoder architecture that takes an image as input and predicts the spectrogram of speech that describes this image. The final speech audio is obtained from the predicted spectrogram via WaveNet. Extensive experiments on the public benchmark database Flickr8k demonstrate that the proposed SAS is able to synthesize natural spoken descriptions for images, indicating that synthesizing spoken descriptions for images while bypassing text and phonemes is feasible. Siyuan Feng 0001, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg |
ICASSP | 5 |
| 2021 | Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image RetrievalabstractMultimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conducting retrieval and word discovery experiments on MSCOCO and Flickr8k, and empirically demonstrate that both neural MT with self-attention and statistical MT achieve word discovery scores that are superior to those of a state-of-the-art neural retrieval system, outperforming it by 2% and 5% alignment F1 scores respectively. Liming Wang 0003, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 4 |
| 2021 | Unsupervised Acoustic Unit Discovery by Leveraging a Language-Independent Subword Discriminative Feature RepresentationabstractThis paper tackles automatically discovering phone-like acoustic units (AUD) from unlabeled speech data. Past studies usually proposed single-step approaches. We propose a two-stage approach: the first stage learns a subword-discriminative feature representation and the second stage applies clustering to the learned representation and obtains phone-like clusters as the discovered acoustic units. In the first stage, a recently proposed method in the task of unsupervised subword modeling is improved by replacing a monolingual out-of-domain (OOD) ASR system with a multilingual one to create a subword-discriminative representation that is more language-independent. In the second stage, segment-level k-means is adopted, and two methods to represent the variable-length speech segments as fixed-dimension feature vectors are compared. Experiments on a very low-resource Mboshi language corpus show that our approach outperforms state-of-the-art AUD in both normalized mutual information (NMI) and F-score. The multilingual ASR improved upon the monolingual ASR in providing OOD phone labels and in estimating the phone boundaries. A comparison of our systems with and without knowing the ground-truth phone boundaries showed a 16% NMI performance gap, suggesting that the current approach can significantly benefit from improved phone boundary estimation. Siyuan Feng 0001, Piotr Zelasko, Laureano Moro-Velázquez, Odette Scharenborg |
Interspeech | 4 |
| 2021 | Learning to Recognise Words Using Visually Grounded SpeechabstractWe investigated word recognition in a Visually Grounded Speech model. The model has been trained on pairs of images and spoken captions to create visually grounded embeddings which can be used for speech to image retrieval and vice versa. We investigate whether such a model can be used to recognise words by embedding isolated words and using them to retrieve images of their visual referents. We investigate the time- course of word recognition using a gating paradigm and perform a statistical analysis to see whether well known word competition effects in human speech processing influence word recognition. Our experiments show that the model is able to recognise words, and the gating paradigm reveals that words can be recognised from partial input as well and that recognition is negatively influenced by word competition from the word initial cohort. Sebastiaan Scholten, Danny Merkx, Odette Scharenborg |
ISCAS | 3 |
| 2021 | Learning Fine-Grained Semantics in Spoken Language Using Visual GroundingabstractIn the case of unwritten languages, acoustic models cannot be trained in the standard way, i.e., using speech and textual transcriptions. Recently, several methods have been proposed to learn speech representations using images, i.e., using visual grounding. Existing studies have focused on scene images. Here, we investigate whether fine-grained semantic information, reflecting the relationship between attributes and objects, can be learned from spoken language. To this end, a Fine-grained Semantic Embedding Network (FSEN) for learning semantic representations of spoken language grounded by fine-grained images is proposed. For training, we propose an efficient objective function, which includes a matching constraint, an adversarial objective, and a classification constraint. The learned speech representations are evaluated using two tasks, i.e., speech-image cross-modal retrieval and speech-to-image generation. On the retrieval task, FSEN outperforms other state-of-the-art methods on both a scene image dataset and two fine-grained datasets. The image generation task shows that the learned speech representations can be used to generate high-quality and semantic-consistent fine-grained images. Learning fine-grained semantics from spoken language via visual grounding is thus possible. Jihua Zhu, Odette Scharenborg |
ISCAS | 4 |
| 2021 | The effect of intermittent noise on lexically-guided perceptual learning in native and non-native listeningabstractThere is ample evidence that both native and non-native listeners deal with speech variation by quickly tuning into a speaker and adjusting their phonetic categories according to the speaker’s ambiguous pronunciation. This process is called lexically-guided perceptual learning. Moreover, the presence of noise in the speech signal has previously been shown to change the word competition process by increasing the number of candidate words competing for recognition and slowing down the recognition process. Given that reliable lexical information should be available quickly to induce lexically-guided perceptual learning and that word recognition is slowed down in the presence of noise, and especially so for non-native listeners, the present study investigated whether noise interferes with lexically-guided perceptual learning in native and non-native listening. Native English and Dutch listeners were exposed to a story in English in clean speech or with stretches of noise. All the /l/ and /ɹ/ sounds in the story were replaced with an ambiguous sound half-way between /l/ and /ɹ/. Although noise altered the pattern of responses for the non-native listeners in a subsequent phonetic categorization task, both native and non-native listeners demonstrated lexically-guided perceptual learning in both clean and noisy listening conditions. We argue that the robustness of perceptual learning in the presence of intermittent noise for both native and non-native listeners is additional evidence for the remarkable flexibility of native and non-native perceptual systems even in adverse listening conditions. Polina Drozdova, Roeland van Hout, Sven L. Mattys, Odette Scharenborg |
Speech Commun. | 4 |
| 2021 | Synthesizing Spoken Descriptions of ImagesabstractImage captioning technology has great potential in many scenarios. However, current text-based image captioning methods cannot be applied to approximately half of the world's languages due to these languages lack of a written form. To solve this problem, recently the image-to-speech task was proposed, which generates spoken descriptions of images bypassing any text via an intermediate representation consisting of phonemes (image-to-phoneme). Here, we present a comprehensive study on the image-to-speech task in which, 1) several representative image-to-text generation methods are implemented for the image-to-phoneme task, 2) objective metrics are sought to evaluate the image-to-phoneme task, and 3) an end-to-end image-to-speech model that is able to synthesize spoken descriptions of images bypassing both text and phonemes is proposed. Extensive experiments are conducted on the public benchmark database Flickr8k. Results of our experiments demonstrate that 1) State-of-the-art image-to-text models can perform well on the image-to-phoneme task, and 2) several evaluation metrics, including BLEU3, BLEU4, BLEU5, and ROUGE-L can be used to evaluate image-to-phoneme performance. Finally, 3) end-to-end image-to-speech bypassing text and phonemes is feasible. Justin van der Hout, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Generating Images From Spoken DescriptionsabstractText-based technologies, such as text translation from one language to another, and image captioning, are gaining popularity. However, approximately half of the world's languages are estimated to be lacking a commonly used written form. Consequently, these languages cannot benefit from text-based technologies. This paper presents 1) a new speech technology task, i.e., a speech-to-image generation (S2IG) framework which translates speech descriptions to photo-realistic images 2) without using any text information, thus allowing unwritten languages to potentially benefit from this technology. The proposed speech-to-image framework, referred to as S2IGAN, consists of a speech embedding network and a relation-supervised densely-stacked generative model. The speech embedding network learns speech embeddings with the supervision of corresponding visual information from images. The relation-supervised densely-stacked generative model synthesizes images, conditioned on the speech embeddings produced by the speech embedding network, that are semantically consistent with the corresponding spoken descriptions. Extensive experiments are conducted on four public benchmark databases: two databases that are commonly used in text-to-image generation tasks, i.e., CUB-200 and Oxford-102 for which we created synthesized speech descriptions, and two databases with natural speech descriptions which are often used in the field of cross-modal learning of speech and images, i.e., Flickr8k and Places. Results on these databases demonstrate the effectiveness of the proposed S2IGAN on synthesizing high-quality and semantically-consistent images from the speech signal, yielding a good performance and a solid baseline for the S2IG task. Tingting Qiao, Jihua Zhu, Alan Hanjalic, Odette Scharenborg |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Unsupervised Subword Modeling Using Autoregressive Pretraining and Cross-Lingual Phone-Aware ModelingabstractThis study addresses unsupervised subword modeling, i.e., learning feature representations that can distinguish subword units of a language. The proposed approach adopts a two-stage bottleneck feature (BNF) learning framework, consisting of autoregressive predictive coding (APC) as a front-end and a DNN-BNF model as a back-end. APC pretrained features are set as input features to a DNN-BNF model. A language-mismatched ASR system is used to provide cross-lingual phone labels for DNN-BNF model training. Finally, BNFs are extracted as the subword-discriminative feature representation. A second aim of this work is to investigate the robustness of our approach's effectiveness to different amounts of training data. The results on Libri-light and the ZeroSpeech 2017 databases show that APC is effective in front-end feature pretraining. Our whole system outperforms the state of the art on both databases. Cross-lingual phone labels for English data by a Dutch ASR outperform those by a Mandarin ASR, possibly linked to the larger similarity of Dutch compared to Mandarin with English. Our system is less sensitive to training data amount when the training data is over 50 hours. APC pretraining leads to a reduction of needed training material from over 5,000 hours to around 200 hours with little performance degradation. Siyuan Feng 0001, Odette Scharenborg |
INTERSPEECH | 2 |
| 2020 | Detecting and Analysing Spontaneous Oral Cancer Speech in the WildabstractOral cancer speech is a disease which impacts more than half a million people worldwide every year. Analysis of oral cancer speech has so far focused on read speech. In this paper, we 1) present and 2) analyse a three-hour long spontaneous oral cancer speech dataset collected from YouTube. 3) We set baselines for an oral cancer speech detection task on this dataset. The analysis of these explainable machine learning baselines shows that sibilants and stop consonants are the most important indicators for spontaneous oral cancer speech detection. Bence Mark Halpern, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg |
INTERSPEECH | 4 |
| 2020 | Evaluating Automatically Generated Phoneme Captions for ImagesabstractImage2Speech is the relatively new task of generating a spoken description of an image. This paper presents an investigation into the evaluation of this task. For this, first an Image2Speech system was implemented which generates image captions consisting of phoneme sequences. This system outperformed the original Image2Speech system on the Flickr8k corpus. Subsequently, these phoneme captions were converted into sentences of words. The captions were rated by human evaluators for their goodness of describing the image. Finally, several objective metric scores of the results were correlated with these human ratings. Although BLEU4 does not perfectly correlate with human ratings, it obtained the highest correlation among the investigated metrics, and is the best currently existing metric for the Image2Speech task. Current metrics are limited by the fact that they assume their input to be words. A more appropriate metric for the Image2Speech task should assume its input to be parts of words, i.e. phonemes, instead. Justin van der Hout, Zoltán D'Haese, Mark Hasegawa-Johnson, Odette Scharenborg |
INTERSPEECH | 4 |
| 2020 | S2IGAN: Speech-to-Image Generation via Adversarial LearningabstractAn estimated half of the world's languages do not have a written form, making it impossible for these languages to benefit from any existing text-based technologies.In this paper, a speech-toimage generation (S2IG) framework is proposed which translates speech descriptions to photo-realistic images without using any text information, thus allowing unwritten languages to potentially benefit from this technology.The proposed S2IG framework, named S2IGAN, consists of a speech embedding network (SEN) and a relation-supervised densely-stacked generative model (RDG).SEN learns the speech embedding with the supervision of the corresponding visual information.Conditioned on the speech embedding produced by SEN, the proposed RDG synthesizes images that are semantically consistent with the corresponding speech descriptions.Extensive experiments on datasets CUB and Oxford-102 demonstrate the effectiveness of the proposed S2IGAN on synthesizing high-quality and semantically-consistent images from the speech signal, yielding a good performance and a solid baseline for the S2IG task. Tingting Qiao, Jihua Zhu, Alan Hanjalic, Odette Scharenborg |
INTERSPEECH | 5 |
| 2020 | That Sounds Familiar: An Analysis of Phonetic Representations Transfer Across LanguagesabstractOnly a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing in other languages to train a multilingual automatic speech recognition (ASR) model, which, intuitively, should learn some universal phonetic representations. In this work, we focus on gaining a deeper understanding of how general these representations might be, and how individual phones are getting improved in a multilingual setting. To that end, we select a phonetically diverse set of languages, and perform a series of monolingual, multilingual and crosslingual (zero-shot) experiments. The ASR is trained to recognize the International Phonetic Alphabet (IPA) token sequences. We observe significant improvements across all languages in the multilingual setting, and stark degradation in the crosslingual setting, where the model, among other errors, considers Javanese as a tone language. Notably, as little as 10 hours of the target language training data tremendously reduces ASR error rates. Our analysis uncovered that even the phones that are unique to a single language can benefit greatly from adding training data from other languages - an encouraging result for the low-resource speech community. Piotr Zelasko, Laureano Moro-Velázquez, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 4 |
| 2020 | Speech Technology for Unwritten LanguagesabstractSpeech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible. Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Study of the Performance of Automatic Speech Recognition Systems in Speakers with Parkinson's DiseaseabstractParkinson’s Disease (PD) affects motor capabilities of patients, who in some cases need to use human-computer assistive technologies to regain independence. The objective of this work is to study in detail the differences in error patterns from state-of-the-art Automatic Speech Recognition (ASR) systems on speech from people with and without PD. Two different speech recognizers (attention-based end-to-end and Deep Neural Network - Hidden Markov Models hybrid systems) were trained on a Spanish language corpus and subsequently tested on speech from 43 speakers with PD and 46 without PD. The differences related to error rates, substitutions, insertions and deletions of characters and phonetic units between the two groups were analyzed, showing that the word error rate is 27% higher in speakers with PD than in control speakers, with a moderated correlation between that rate and the developmental stage of the disease. The errors were related to all manner classes, and were more pronounced in the vowel /u/. This study is the first to evaluate ASR systems’ responses to speech from patients at different stages of PD in Spanish. The analyses showed general trends but individual speech deficits must be studied in the future when designing new ASR systems for this population. Laureano Moro-Velázquez, Shinji Watanabe 0001, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 5 |
| 2019 | Survey Talk: Reaching Over the Gap: Cross- and Interdisciplinary Research on Human and Automatic Speech Processing
Odette Scharenborg |
INTERSPEECH | 1 |
| 2019 | The Neural Correlates Underlying Lexically-Guided Perceptual LearningabstractThere is ample evidence showing that listeners are able to quickly adapt their phoneme classes to ambiguous sounds using a process called lexically-guided perceptual learning. This paper presents the first attempt to examine the neural correlates underlying this process. Specifically, we compared the brain’s responses to ambiguous [f/s] sounds in Dutch non-native listeners of English (N=36) before and after exposure to the ambiguous sound to induce learning, using Event-Related Potentials (ERPs). We identified a group of participants who showed lexically-guided perceptual learning in their phonetic categorization behavior as observed by a significant difference in /s/ responses between pretest and posttest and a group who did not. Moreover, we observed differences in mean ERP amplitude to ambiguous phonemes at pretest and posttest, shown by a reliable reduction in amplitude of a positivity over medial central channels from 250 to 550 ms. However, we observed no significant correlation between the size of behavioral and neural pre/posttest effects. Possibly, the observed behavioral and ERP differences between pretest and posttest link to different aspects of the sound classification task. In follow-up research, these differences will be further investigated by assessing their relationship to neural responses to the ambiguous sounds in the exposure phase. Odette Scharenborg, Jiska Koemans, Cybelle Smith, Mark Hasegawa-Johnson, Kara D. Federmeier |
INTERSPEECH | 1 |
| 2019 | The Representation of Speech in Deep Neural Networks
Odette Scharenborg, Nikki van der Gouw, Martha A. Larson, Elena Marchiori |
MMM (2) | 1 |
| 2019 | Why listening in background noise is harder in a non-native language than in a native language: A review
Odette Scharenborg, Marjolein van Os |
Speech Commun. | 1 |
| 2018 | The Effects of Background Noise on Native and Non-native Spoken-word Recognition: A Computational Modelling Approach
Themis Karaminis, Odette Scharenborg |
CogSci | 2 |
| 2018 | Bayesian Models for Unit Discovery on a Very Low Resource LanguageabstractDeveloping speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the I-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus. Lucas Ondel Yang, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Lukás Burget, François Yvon, Sanjeev Khudanpur |
ICASSP | 6 |
| 2018 | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 WorkshopabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech. Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux |
ICASSP | 1 |
| 2018 | Articulatory Feature Classification Using Convolutional Neural Networks
Danny Merkx, Odette Scharenborg |
INTERSPEECH | 2 |
| 2018 | The Conversation Continues: the Effect of Lyrics and Music Complexity of Background Music on Spoken-Word RecognitionabstractBackground music in social interaction settings can hinder conversation.Yet, little is known of how specific properties of music impact speech processing.This paper addresses this knowledge gap by investigating the effect of the 1) complexity of the background music, and 2) the presence versus absence of sung lyrics on spoken-word recognition in background music.To answer these questions, a word identification experiment was run in which Dutch participants listened to Dutch CVC words embedded in stretches of background music in four conditions: low/high complexity and with lyrics/music-only, and at three SNRs.Music stretches with and without lyrics were sampled from the same song in order to control for factors beyond the complexity of the music and the presence of lyrics.The results showed a clear negative impact of more complex music and the presence of lyrics in background music on spoken-word recognition.The results open a path for future work, and suggest that social spaces (e.g., restaurants, cafés and bars) should make careful choices of music to promote conversation. Odette Scharenborg, Martha A. Larson |
INTERSPEECH | 1 |
| 2018 | Visualizing Phoneme Category Adaptation in Deep Neural NetworksabstractBoth human listeners and machines need to adapt their sound categories whenever a new speaker is encountered. This perceptual learning is driven by lexical information. The aim of this paper is two-fold: investigate whether a deep neural network-based (DNN) ASR system can adapt to only a few examples of ambiguous speech as humans have been found to do; investigate a DNN’s ability to serve as a model of human perceptual learning. Crucially, we do so by looking at intermediate levels of phoneme category adaptation rather than at the output level. We visualize the activations in the hidden layers of the DNN during perceptual learning. The results show that, similar to humans, DNN systems learn speaker-adapted phone category boundaries from a few labeled examples. The DNN adapts its category boundaries not only by adapting the weights of the output layer, but also by adapting the implicit feature maps computed by the hidden layers, suggesting the possibility that human perceptual learning might involve a similar nonlinear distortion of a perceptual space that is intermediate between the acoustic input and the phonological categories. Comparisons between DNNs and humans can thus provide valuable insights into the way humans process speech and improve ASR technology. Odette Scharenborg, Sebastian Tiesmeyer, Mark Hasegawa-Johnson, Najim Dehak |
INTERSPEECH | 1 |
| 2016 | Processing and Adaptation to Ambiguous Sounds during the Course of Perceptual LearningabstractContains fulltext : 162019.pdf (Publisher’s version ) (Open Access) Polina Drozdova, Roeland van Hout, Odette Scharenborg |
INTERSPEECH | 3 |
| 2016 | The Effect of Background Noise on the Activation of Phonological and Semantic Information During Spoken-Word Recognitionabstract\n Contains fulltext :\n 162036.pdf (Publisher’s version ) (Open Access)\n Florian Hintz, Odette Scharenborg |
INTERSPEECH | 2 |
| 2016 | Does the Importance of Word-Initial and Word-Final Information Differ in Native versus Non-Native Spoken-Word Recognition?abstractContains fulltext : 161987.pdf (Publisher’s version ) (Open Access) Odette Scharenborg, Juul Coumans, Sofoklis Kakouros, Roeland van Hout |
INTERSPEECH | 1 |
| 2016 | The Effect of Sentence Accent on Non-Native Speech Perception in Noiseabstract\n Contains fulltext :\n 162038.pdf (Publisher’s version ) (Open Access)\n Odette Scharenborg, Elea Kolkman, Sofoklis Kakouros, Brechtje Post |
INTERSPEECH | 1 |
| 2015 | Durational information in word-initial lexical embeddings in spoken Dutch
Odette Scharenborg |
INTERSPEECH | 1 |
| 2014 | Non-native word recognition in noise: the role of word-initial and word-final informationabstractContains fulltext : 131628.pdf (Publisher’s version ) (Open Access) Juul Coumans, Roeland van Hout, Odette Scharenborg |
INTERSPEECH | 3 |
| 2014 | Phoneme category retuning in a non-native languageabstractItem does not contain fulltext Polina Drozdova, Roeland van Hout, Odette Scharenborg |
INTERSPEECH | 3 |
| 2014 | Collecting a corpus of Dutch noise-induced 'slips of the ear'abstractWhen trying to understand how listeners recognise words, listeners’ misperceptions, so-called ‘slips of the ear’, can reveal important aspects of the underlying mechanisms of normal word recognition. Such misperceptions shed light onto how inferences are made by listeners about acoustic details in the speech signal and how these interact with other sound sources in the background. On the other hand, if speech from a particular speaker is more prone to being misperceived than that from another speaker, these misperceptions may also shed light onto speaker characteristics. To study these phenomena, misperceptions that occur consistently are invaluable. Although such confusions are quite rare, within the Marie Curie INSPIRE project, software has been developed to efficiently collect such consistent confusions for different languages. Using this software, we have started to collect Dutch consistent confusions. Single words, embedded in five different types of noise at different SNRs, produced by four speakers were presented to Dutch listeners. In a preliminary analysis, consistent confusions were analysed in terms of phoneme substitutions, insertions, and deletions, reconstructions of words using background noise, and eccentric cases. Moreover, the number and types of consistent confusions obtained in the different noise types and from different speakers are compared. Odette Scharenborg, Eric Sanders, Bert Cranen |
INTERSPEECH | 1 |
| 2014 | Age, hearing loss and the perception of affective utterances in conversational speechabstract15th Annual Conference of the International Speech Communication Association, 14 september 2014 Juliane Schmidt, Esther Janse, Odette Scharenborg |
INTERSPEECH | 3 |
| 2013 | Changes in the role of intensity as a cue for fricative categorisationabstractContains fulltext : 116379.pdf (Publisher’s version ) (Open Access) Odette Scharenborg, Esther Janse |
INTERSPEECH | 1 |
| 2012 | Modeling Cue Trading in Human Word RecognitionabstractClassical phonetic studies have shown that acoustic-articulatory cues can be interchanged without affecting the resulting phoneme percept (‘cue trading’). Cue trading has so far mainly been investigated in the context of phoneme identification. In this study, we investigate cue trading during recognition of words, the units of speech through which we communicate. This paper aims to provide a method to quantify cue trading effects by using a computational model of human word recognition. This model takes the acoustic signal as input and represents speech using articulatory feature streams. Importantly, it allows cue trading and underspecification. Its set-up is inspired by the functionality of Fine-Tracker, a recent computational model of human word recognition. This approach makes it possible, for the first time, to quantify cue trading in terms of a trade-off between features and to investigate cue trading in the context of a word recognition task. Index Terms: cue trading, human word recognition, computational modeling, articulatory features. Louis ten Bosch, Odette Scharenborg |
INTERSPEECH | 2 |
| 2012 | Hearing Loss and the Use of Acoustic Cues in Phonetic Categorisation of FricativesabstractAging often affects sensitivity to the higher frequencies, which results in the loss of sensitivity to phonetic detail in speech. Hearing loss may therefore interfere with the categorisation of two consonants that have most information to differentiate between them in those higher frequencies and less in the lower frequencies, e.g., /f/ and /s/. We investigate two acoustic cues, i.e., formant transitions and fricative intensity, that older listeners might use to differentiate between /f/ and /s/. The results of two phonetic categorisation tasks on 38 older listeners (aged 60+) with varying degrees of hearing loss indicate that older listeners seem to use formant transitions as a cue to distinguish /s/ from /f/. Moreover, this ability is not impacted by hearing loss. On the other hand, listeners with increased hearing loss seem to rely more on intensity for fricative identification. Thus, progressive hearing loss may lead to gradual changes in perceptual cue weighting. Odette Scharenborg, Esther Janse |
INTERSPEECH | 1 |
| 2012 | Perceptual Learning of /f/-/s/ by Older ListenersabstractYoung listeners can quickly modify their interpretation of a speech sound when a talker produces the sound ambiguously. Young Dutch listeners rely mainly on the higher frequencies to distinguish between /f/ and /s/, but these higher frequencies are particularly vulnerable to age-related hearing loss. We therefore tested whether older Dutch listeners can show perceptual retuning given an ambiguous pronunciation in between /f/ and /s/. Results of a lexically-guided perceptual learning experiment showed that older Dutch listeners are still able to learn non-standard pronunciations of /f/ and /s/. Possibly, the older listeners have learned to rely on other acoustic cues, such as formant transitions, to distinguish between /f/ and /s/. However, the size and duration of the perceptual effect is influenced by hearing loss, with listeners with poorer hearing showing a smaller and a shorter-lived learning effect. Odette Scharenborg, Esther Janse, Andrea Weber |
INTERSPEECH | 1 |
| 2012 | Computational Modelling of the Recognition of Foreign-Accented SpeechabstractIn foreign-accented speech, pronunciation typically deviates from the canonical form to some degree. For native listeners, it has been shown that word recognition is more difficult for strongly-accented words than for less strongly-accented words. Furthermore recognition of strongly-accented words becomes easier with additional exposure to the foreign accent. In this paper, listeners ’ behaviour was simulated with Fine-tracker, a computational model of word recognition that uses real speech as input. The simulations showed that, in line with human listeners, 1) Fine-Tracker’s recognition outcome is modulated by the degree of accentedness and 2) it improves slightly after brief exposure with the accent. On the level of individual words, however, Fine-tracker failed to correctly simulate listeners’ behaviour, possibly due to differences in overall familiarity with the chosen accent (German-accented Dutch) between human listeners and Fine-Tracker. Index Terms: foreign-accented speech, accent strength, word recognition, computational modelling, German-accented Dutch Odette Scharenborg, Marijt J. Witteman, Andrea Weber |
INTERSPEECH | 1 |
| 2011 | Perceptual Learning of LiquidsabstractPrevious research on lexically-guided perceptual learning has focussed on contrasts that differ primarily in local cues, such as plosive and fricative contrasts. The present research had two aims: to investigate whether perceptual learning occurs for a contrast with non-local cues, the /l/-/r/ contrast, and to establish whether STRAIGHT can be used to create ambiguous sounds on an /l/-/r/ continuum. Listening experiments showed lexically-guided learning about the /l/-/r/ contrast. Listeners can thus tune in to unusual speech sounds characterised by non-local cues. Moreover, STRAIGHT can be used to create stimuli for perceptual learning experiments, opening up new research possibilities. Index Terms: perceptual learning, morphing, liquids, human word recognition, STRAIGHT. Odette Scharenborg, Holger Mitterer, James M. McQueen |
INTERSPEECH | 1 |
| 2010 | Native and non-native listeners' perception of English consonants in different types of noise
Mirjam Broersma, Odette Scharenborg |
Speech Commun. | 2 |
| 2010 | Language-independent processing in speech perception: Identification of English intervocalic consonants by speakers of eight European languages
Martin Cooke, María Luisa García Lecumberri, Odette Scharenborg, Wim A. van Dommelen |
Speech Commun. | 3 |
| 2009 | Using temporal information for improving articulatory-acoustic feature classificationabstractThis paper combines acoustic features with a high temporal and a high frequency resolution to reliably classify articulatory events of short duration, such as bursts in plosives. SVM classification experiments on TIMIT and SV Articulatory showed that articulatory-acoustic features (AFs) based on a combination of MFCCs derived from a long window of 25 ms and a short window of 5 ms that are both shifted with 2.5 ms steps (Both) outperform standard MFCCs derived with a window of 25 ms and a shift of 10 ms (Baseline). Finally, comparison of the TIMIT and SV Articulatory results showed that for classifiers trained on data that allows for asynchronously changing AFs (SV Articulatory) the improvement from Baseline to Both is larger than for classifiers trained on data where AFs change simultaneously with the phone boundaries (TIMIT). Barbara Schuppler, Joost van Doremalen, Odette Scharenborg, Bert Cranen, Lou Boves |
ASRU | 3 |
| 2009 | Using durational cues in a computational model of spoken-word recognitionabstractEvidence that listeners use durational cues to help resolve temporarily ambiguous speech input has accumulated over the past few years. In this paper, we investigate whether durational cues are also beneficial for word recognition in a computational model of spoken-word recognition. Two sets of simulations were carried out using the acoustic signal as input. The simulations showed that the computational model, like humans, takes benefit from durational cues during word recognition, and uses these to disambiguate the speech signal. These results thus provide support for the theory that durational cues play a role in spoken-word recognition. Index Terms: duration, spoken-word recognition, computational modelling Odette Scharenborg |
INTERSPEECH | 1 |
| 2009 | Lexical embedding in spoken dutchabstractA stretch of speech is often consistent with multiple words, e.g., the sequence /hæm/ is consistent with 'ham' but also with the first syllable of 'hamster', resulting in temporary ambiguity.However, to what degree does this lexical embedding occur?Analyses on two corpora of spoken Dutch showed that 11.9%-19.5% of polysyllabic word tokens have word-initial embedding, while 4.1%-7.5% of monosyllabic word tokens can appear word-initially embedded.This is much lower than suggested by an analysis of a large dictionary of Dutch.Speech processing thus appears to be simpler than one might expect on the basis of statistics on a dictionary. Odette Scharenborg, Stefanie Okolowski |
INTERSPEECH | 1 |
| 2008 | The interspeech 2008 consonant challengeabstractListeners outperform automatic speech recognition systems at every level, including the very basic level of consonant identification. What is not clear is where the human advantage originates. Does the fault lie in the acoustic representations of speech or in the recognizer architecture, or in a lack of compatibility between the two? Many insights can be gained by carrying out a detailed human-machine comparison. The purpose of the Interspeech 2008 Consonant Challenge is to promote focused comparisons on a task involving intervocalic consonant identification in noise, with all participants using the same training and test data. This paper describes the Challenge, listener results and baseline ASR performance. Index Terms: consonant perception, VCV, humanmachine performance comparisons Martin Cooke, Odette Scharenborg |
INTERSPEECH | 2 |
| 2008 | The non-native consonant challenge for european languagesabstractThis paper reports on a multilingual investigation into the effects of different masker types on native and non-native perception in a VCV consonant recognition task. Native listeners outperformed 7 other language groups, but all groups showed a similar ranking of maskers. Strong first language (L1) interference was observed, both from the sound system and from the L1 orthography. Universal acoustic-perceptual tendencies are also at work in both native and non-native sound identifications in noise. The effect of linguistic distance, however, was less clear: in large multilingual studies, listener variables may overpower other factors. María Luisa García Lecumberri, Martin Cooke, Francesco Cutugno, Mircea Giurgiu, Bernd T. Meyer, Odette Scharenborg, Wim A. van Dommelen, Jan Volín |
INTERSPEECH | 6 |
| 2008 | Modelling fine-phonetic detail in a computational model of word recognitionabstractThere is now considerable evidence that fine-grained acoustic-phonetic detail in the speech signal helps listeners to segment a speech signal into syllables and words.In this paper, we compare two computational models of word recognition on their ability to capture and use this finephonetic detail during speech recognition.One model, SpeM, is phoneme-based, whereas the other, newly developed Fine-Tracker, is based on articulatory features.Simulations dealt with modelling the ability of listeners to distinguish short words (e.g., 'ham') from the longer words in which they are embedded (e.g., 'hamster').The simulations with Fine-Tracker showed that it was, like human listeners, able to distinguish between short words from the longer words in which they are embedded.This suggests that it is possible to extract this fine-phonetic detail from the speech signal and use it during word recognition. Odette Scharenborg |
INTERSPEECH | 1 |
| 2008 | Preparing a corpus of dutch spontaneous dialogues for automatic phonetic analysisabstractThis paper presents the steps needed to make a corpus of Dutch spontaneous dialogues accessible for automatic phonetic research aimed at increasing our understanding of reduction phenomena and the role of fine phonetic detail. Since the corpus was not created with automatic processing in mind, it needed to be reshaped. The first part of this paper describes the actions needed for this reshaping in some detail. The second part reports the results of a preliminary analysis of the reduction phenomena in the corpus. For this purpose a phonemic transcription of the corpus was created by means of a forced alignment, first with a lexicon of canonical pronunciations and then with multiple pronunciation variants per word. In this study pronunciation variants were generated by applying a large set of phonetic processes that have been implicated in reduction to the canonical pronunciations of the words. This relatively straightforward procedure allows us to produce plausible pronunciation variants and to verify and extend the results of previous reduction studies reported in the literature. Barbara Schuppler, Mirjam Ernestus, Odette Scharenborg, Lou Boves |
INTERSPEECH | 3 |
| 2007 | Finding Maximum Margin Segments in SpeechabstractMaximum margin clustering (MMC) is a relatively new and promising kernel method. In this paper, we apply MMC to the task of unsupervised speech segmentation. We present three automatic speech segmentation methods based on MMC, which are tested on TIMIT and evaluated on the level of phoneme boundary detection. The results show that MMC is highly competitive with existing unsupervised methods for the automatic detection of phoneme boundaries. Furthermore, initial analyses show that MMC is a promising method for the automatic detection of sub-phonetic information in the speech signal. Yago Pereiro-Estevan, Vincent Wan, Odette Scharenborg |
ICASSP (4) | 3 |
| 2007 | Segmentation of speech: child's play?abstractThe difficulty of the task of segmenting a speech signal into its words is immediately clear when listening to a foreign language; it is much harder to segment the signal into its words, since the words of the language are unknown. Infants are faced with the same task when learning their first language.\nThis study provides a better understanding of the task that infants face while learning their native language. We employed an automatic algorithm on the task of speech segmentation without prior knowledge of the labels of the phonemes. An analysis of the boundaries erroneously placed inside a phoneme showed that the algorithm consistently placed additional boundaries in phonemes\nin which acoustic changes occur. These acoustic changes may be as great as the transition from the closure to the burst of a plosive or as subtle as the formant transitions in low or back vowels.\nMoreover, we found that glottal vibration may attenuate the\nrelevance of acoustic changes within obstruents. An interesting question for further research is how infants learn to overcome the natural tendency to segment these ‘dynamic’ phonemes. Odette Scharenborg, Mirjam Ernestus, Vincent Wan |
INTERSPEECH | 1 |
| 2007 | Can unquantised articulatory feature continuums be modelled?abstractArticulatory feature (AF) modelling of speech has received a considerable amount of attention in automatic speech recognition research. Although termed ‘articulatory’, previous definitions make certain assumptions that are invalid, for instance, that articulators ‘hop’ from one fixed position to the next. In this paper, we studied two methods, based on support vector classification (SVC) and regression (SVR), in which the articulation continuum is modelled without being restricted to using discrete AF value classes. A comparison with a baseline system trained on quantised values of the articulation continuum showed that both SVC and SVR outperform the baseline for two of the three investigated AFs, with improvements up to 5.6% absolute. Odette Scharenborg, Vincent Wan |
INTERSPEECH | 1 |
| 2007 | 'Early recognition' of polysyllabic words in continuous speech
Odette Scharenborg, Louis ten Bosch, Lou Boves |
Comput. Speech Lang. | 1 |
| 2007 | A two-pass approach for handling out-of-vocabulary words in a large vocabulary recognition task
Odette Scharenborg, Stephanie Seneff, Lou Boves |
Comput. Speech Lang. | 1 |
| 2007 | Reaching over the gap: A review of efforts to link human and automatic speech recognition research
Odette Scharenborg |
Speech Commun. | 1 |
| 2007 | Towards capturing fine phonetic variation in speech using articulatory features
Odette Scharenborg, Vincent Wan, Roger K. Moore |
Speech Commun. | 1 |
| 2006 | Acoustic Scores and Symbolic Mismatch Penalties in Phone LatticesabstractThis paper builds on previous work aimed at unraveling the structure of the speech signal using probabilistic representations. The context of this work is a multi-pass speech recognition system in which a phone lattice is created and used as a basis for a lexical decoding pass (search) that allows symbolic mismatches at certain costs. The focus is on the optimization of the costs of the phone insertions, deletions and substitutions that are used in the lexical decoding pass. Two optimization approaches are presented, one related to a multi-pass computational model for human speech recognition, the other based on a decoding that minimizes Bayes' risks. In the final section, the advantages of the two optimization methods are discussed and compared Louis ten Bosch, Annika Hämäläinen, Odette Scharenborg, Lou Boves |
ICASSP (1) | 3 |
| 2005 | ASR decoding in a computational model of human word recognitionabstractRecently, a computational model of human word recognition, called SpeM, has been developed. In contrast to most current models of human word recognition, SpeM is able to process actual acoustic speech input, and decodes the incoming speech stream into lexical and non-lexical items. This model makes the links between HSR and ASR as explicit as possible. In this paper, we focus on unravelling the structure of the complex search space that is used in SpeM and similar decoding strategies. To that end, it discusses a number of properties of phone lattices in relation to canonical phone representations. Furthermore, we elaborate on the close relation between distances in this search space, and distance measures in search spaces that are based on a combination of acoustic and phonetic features. 1. Louis ten Bosch, Odette Scharenborg |
INTERSPEECH | 2 |
| 2005 | Parallels between HSR and ASR: how ASR can contribute to HSRabstractIn this paper, we illustrate the close parallels between the research fields of human speech recognition (HSR) and automatic speech recognition (ASR) using a computational model of human word recognition, SpeM, which was built using techniques from ASR. We show that ASR has proven to be useful for improving models of HSR by relieving them of some of their shortcomings. However, in order to build an integrated computational model of all aspects of HSR, a lot of issues remain to be resolved. In this process, ASR algorithms and techniques definitely can play an important role. 1. Odette Scharenborg |
INTERSPEECH | 1 |
| 2005 | Two-pass strategy for handling OOVs in a large vocabulary recognition taskabstractThis paper addresses the issue of large-vocabulary recognition in a specific word class. We propose a two-pass strategy in which only major cities are explicitly represented in the first stage lexicon. An unknown word model encoded as a phone loop is used to detect OOV city names (referred to as rare city names). After which SpeM, a tool that can extract words and word-initial cohorts from phone graphs on the basis of a large fallback lexicon, provides an N-best list of promising city names on the basis of the phone sequences generated in the first stage. This N-best list is then inserted into the second stage lexicon for a subsequent recognition pass. Experiments were conducted on a set of spontaneous telephone-quality utterances each containing one rare city name. We tested the size of the N-best list and three types of language models (LMs). The experiments showed that SpeM was able to include nearly 85% of the correct city names into an N-best list of 3000 city names when a unigram LM, which also boosted the unigram scores of a city name in a given state, was used. Odette Scharenborg, Stephanie Seneff |
INTERSPEECH | 1 |
| 2003 | Recognising 'real-life' speech with spem: a speech-based computational model of human speech recognitionabstractIn this paper, we present a novel computational model of human speech recognition - called SpeM - based on the theory underlying Shortlist. We will show that SpeM, in combination with an automatic phone recogniser (APR), is able to simulate the human speech recognition process from the acoustic signal to the ultimate recognition of words. This joint model takes an acoustic speech file as input and calculates the activation flows of candidate words on the basis of the degree of fit of the candidate words with the input.\n\nExperiments showed that SpeM outperforms Shortlist on the recognition of 'real-life' input. Furthermore, SpeM performs only slightly worse than an off-the-shelf full-blown automatic speech recogniser in which all words are equally probable, while it provides a transparent computationally elegant paradigm for modelling word activations in human word recognition. Odette Scharenborg, Louis ten Bosch, Lou Boves |
INTERSPEECH | 1 |
| 2003 | Modelling human speech recognition using automatic speech recognition paradigms in speMabstractThe following full text is an author's version which may differ from the publisher's version. Odette Scharenborg, James M. McQueen, Louis ten Bosch, Dennis Norris |
INTERSPEECH | 1 |
| 2002 | ASR in a human word recognition model: generating phonemic input for shortlistabstractThe current version of the psycholinguistic model of human word recognition Shortlist suffers from two unrealistic constraints. First, the input of Shortlist must consist of a single string of phoneme symbols. Second, the current version of the search in Shortlist makes it difficult to deal with insertions and deletions in the input phoneme string. This research attempts to fully automatically derive a phoneme string from the acoustic signal that is as close as possible to the number of phonemes in the lexical representation of the word. We optimised an Automatic Phone Recogniser (APR) using two approaches, viz. varying the value of the mismatch parameter and optimising the APR output strings on the output of Shortlist. The approaches show that it will be very difficult to satisfy the input requirements of the present version of Shortlist with a phoneme string generated by an APR. Odette Scharenborg, Lou Boves, Johan de Veth |
INTERSPEECH | 1 |
| 2001 | Business listings in automatic directory assistanceabstractSo far most attempts to automate Directory Assistance services focused on private listings, because it is not known precisely how callers will refer to a business listings. The research described in this paper, carried out in the SMADA project, tries to fill this gap. The aim of the research is to model the expressions people use when referring to a business listing by means of rules, in order to automatically create a vocabulary, which can be part of an automated DA service. In this paper a rule-base procedure is proposed, which derives rules from the expressions people use. These rules are then used to automatically create expressions from directory listings. Two categories of businesses, viz. hospitals and the hotel and catering industry, are used to explain this procedure. Results for these two categories are used to discuss the problem of the over- and undergeneration of expressions. Odette Scharenborg, Janienke Sturm, Lou Boves |
INTERSPEECH | 1 |