Tanvina Patel

dblp:226/2014 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-9544-0038ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2025 Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
abstract
Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
Dimme de Groot, Tanvina Patel, Devendra Kayande, Odette Scharenborg, Zhengjun Yue
INTERSPEECH2
2025 Challenges and practical guidelines for atypical speech data collection, annotation, usage and sharing: A multi-project perspective
abstract
Contains fulltext : 325867.pdf (Publisher’s version ) (Open Access)
Zhengjun Yue, Mara Barberis, Tanvina Patel, Judith Dineley, Willemijn Doedens, Lottie Stipdonk, Elke De Witte, Erfan Loweimi, Hugo Van hamme, Djaina Satoer, Marina B. Ruiter, Laureano Moro-Velázquez, Nicholas Cummins, Odette Scharenborg
INTERSPEECH3
2024 Using articulated speech EEG signals for imagined speech decoding
abstract
Brain-Computer Interfaces (BCIs) open avenues for communication among individuals unable to use voice or gestures. Silent speech interfaces are one such approach for BCIs that could offer a trans- formative means of connecting with the external world. Performance on imagined speech decoding however is rather low due to, amongst others, data scarcity and the lack of a clear starting point of the imagined speech in the brain signal. We investigate whether using electroencephalography (EEG) signals from articulated speech can be used to improve imagined speech decoding in two ways: we investigate whether articulated speech EEG signals can be used to predict the end point of the imagined speech and use the articulated speech EEG as extra training data for speaker-independent imagined vowel classification. Our results show that using EEG data from articulated speech did not improve classification of vowels in imagined speech, probably due to high variability in EEG signals amongst speakers.
Chris Bras, Tanvina Patel, Odette Scharenborg
INTERSPEECH2
2024 As Biased as You Measure: Methodological Pitfalls of Bias Evaluations in Speaker Verification Research
Wiebke Hutiri, Tanvina Patel, Aaron Yi Ding, Odette Scharenborg
INTERSPEECH2
2024 Improving child speech recognition with augmented child-like speech
Zhengjun Yue, Tanvina Patel, Odette Scharenborg
INTERSPEECH3
2023 Improving Whispered Speech Recognition Performance Using Pseudo-Whispered Based Data Augmentation
abstract
Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and the scarcity of adequate training data leads to low automatic speech recognition (ASR) performance. To address the data scarcity issue, we use a signal processing-based technique that transforms the spectral characteristics of normal speech to those of pseudo-whispered speech. We augment an End-to-End ASR with pseudo-whispered speech and achieve an 18.2 % relative reduction in word error rate for whispered speech compared to the baseline. Results for the individual speaker groups in the wTIMIT database show the best results for US English. Further investigation showed that the lack of glottal information in whispered speech has the largest impact on whispered speech ASR performance.
Zhaofeng Lin, Tanvina Patel, Odette Scharenborg
ASRU2
2023 Exploring Data Augmentation in Bias Mitigation Against Non-Native-Accented Speech
abstract
Automatic speech recognition (ASR) should serve every speaker, not only the majority “standard” speakers of a language. In order to build inclusive ASR, mitigating the bias against speaker groups who speak in a “non-standard” or “diverse” way is crucial. We aim to mitigate the bias against non-native-accented Flemish in a Flemish ASR system. Since this is a low-resource problem, we investigate the optimal type of data augmentation, i.e., speed/pitch perturbation, cross-lingual voice conversion-based methods, and SpecAugment, applied to both native Flemish and non-native-accented Flemish, for bias mitigation. The results showed that specific types of data augmentation applied to both native and non-native-accented speech improve non-native-accented ASR while applying data augmentation to the non-native-accented speech is more conducive to bias reduction. Combining both gave the largest bias reduction for human-machine interaction (HMI) as well as read-type speech.
Aaricia Herygers, Tanvina Patel, Zhengjun Yue, Odette Scharenborg
ASRU3
2022 Using cross-model learnings for the Gram Vaani ASR Challenge 2022
abstract
In the diverse and multilingual land of India, Hindi is spoken as a first language by a majority of its population. Efforts are made to obtain data in terms of audio, transcriptions, dictionary, etc. to develop speech-technology applications in Hindi. Similarly, the Gram-Vaani ASR Challenge 2022 provides spontaneous telephone speech, with natural back-ground and regional variations in Hindi. The challenge provides: 100 hours of labeled train-set, 5 hours of labeled dev-set and 1000 hours of unlabeled data-set. For the 'Closed Challenge', we trained an End-to-End (E2E) Conformer model using speed perturbations, SpecAugment techniques and use VTLN to handle any unknown speaker groups in the blind evaluation set. On the dev-set, we achieved a 30.3% WER compared to the 34.8% WER by the Challenge E2E baseline. For the 'Self Supervised Closed Challenge', a semi-supervised learning approach is used. We generate pseudo-transcripts for the unlabeled data using a hybrid TDNN-3gram LM model and trained an E2E model. This is then used as a seed for retraining the E2E model with high confidence data. Cross-model learning and refining of the E2E model gave 25.3% WER on the dev-set compared to ∼33-35% WER by the Challenge baseline that use wav2vec models.
Tanvina Patel, Odette Scharenborg
INTERSPEECH1
2022 Mitigating bias against non-native accents
abstract
Automatic Speech Recognition (ASR) systems have seen substantial improvements in the past decade; however, not for all speaker groups. Recent research shows that bias exists against different types of speech, including non-native accents, in state-of-the-art (SOTA) ASR systems. To attain inclusive speech recognition, i.e., ASR for everyone irrespective of how one speaks or the accent one has, bias mitigation is essential and necessary. In this thesis, two SOTA ASR systems (one is based on the recurrent neural network (RNN) and the other is based on the transformer architecture) are built to uncover and quantify the bias against non-native accents. Here I focus on bias mitigation against non-native accents using two different approaches: data augmentation and by using more effective training methods. For data augmentation, an autoencoder-based cross-lingual voice conversion (VC) model is used to increase the amount of non-native accented speech training data in addition to data augmentation through speed perturbation. Moreover, I investigate two training methods, i.e., fine-tuning and Domain Adversarial Training (DAT), to see whether they can utilize the available non-native accented speech data more effectively than a standard training approach. Experimental results show for the transformer-based ASR model: (1) adding VC-generated and speed-perturbed data to train the ASR model gives the best bias mitigation performance and the lowest word error rate (WER); (2) fine-tuning reduces the bias against non-native accents but at the cost of native accent performance; and (3) compared with the standard training method, DAT does not leads to further bias reduction. While for the RNN-based ASR model, all the 4 bias mitigation approaches do not show obvious benefits.
Bence Mark Halpern, Tanvina Patel, Odette Scharenborg
INTERSPEECH4
2018 TDNN-based Multilingual Speech Recognition System for Low Resource Indian Languages
Noor Fathima, Tanvina Patel, Mahima C, Anuroop Iyengar
INTERSPEECH2
2018 Development of Large Vocabulary Speech Recognition System with Keyword Search for Manipuri
Tanvina Patel, Krishna D. N, Noor Fathima, Nisar Shah, Mahima C, Anuroop Iyengar
INTERSPEECH1
2018 An Automatic Speech Transcription System for Manipuri Language
Tanvina Patel, Krishna D. N, Noor Fathima, Nisar Shah, Mahima C, Anuroop Iyengar
INTERSPEECH1