Satwinder Singh

dblp:27/1164 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
abstract
In this paper, we present a speaker-independent dysarthric speech recognition system, with a focus on evaluating the recently released Speech Accessibility Project (SAP-1005) dataset, which includes speech data from individuals with Parkinson’s disease (PD). Despite the growing body of research in dysarthric speech recognition, many existing systems are speaker-dependent and adaptive, limiting their generalizability across different speakers and etiologies. Our primary objective is to develop a robust speaker-independent model capable of accurately recognizing dysarthric speech, irrespective of the speaker. Additionally, as a secondary objective, we aim to test the cross-etiology performance of our model by evaluating it on the TORGO dataset, which contains speech samples from individuals with cerebral palsy (CP) and amyotrophic lateral sclerosis (ALS). By leveraging the Whisper model, our speaker-independent system achieved a CER of 6.99% and a WER of 10.71% on the SAP-1005 dataset. Further, in cross-etiology settings, we achieved a CER of 25.08% and a WER of 39.56% on the TORGO dataset. These results highlight the potential of our approach to generalize across unseen speakers and different etiologies of dysarthria.
Satwinder Singh, Zihan Zhong, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdullah, Seyed Reza Shahamiri
ICASSP1
2025 Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition
abstract
Automatic Speech Recognition (ASR) holds immense potential to provide an effective interface for assistive technologies, but its performance remains unsatisfactory for people with speech impairments such as dysarthria. Existing ASR systems struggle to accurately recognize dysarthric speech due to the significant speaker variability in dysarthric speech and the scarcity of dysarthric datasets. In this study, we propose a two-phase adaptation pipeline based on the Conformer architecture that leverages typical speech to transfer to individualized ASR models for dysarthric speakers. ASR performance is evaluated for isolated words and continuous sentences, yielding an average Word Error Rate of 21.5% on the UASpeech dataset and 12.7% on the TORGO dataset. Selectively freezing decoder layers was more often successful than selectively freezing encoder layers, suggesting that optimal performance is achieved by focusing the adaptation on the acoustic information contained in the encoder.
Zihan Zhong, Satwinder Singh, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdullah, Seyed Reza Shahamiri
ICASSP3
2025 Empowering Māori Automatic Speech Recognition through EMD-Based Augmentation
Chengxi Lei, Sheng Li 0010, Satwinder Singh, Feng Hou, Huia Jahnke, Ruili Wang 0001
PRICAI3
2024 Mix-fine-tune: An Alternate Fine-tuning Strategy for Domain Adaptation and Generalization of Low-resource ASR
abstract
Self-supervised Learning (SSL) using extensive unlabeled speech data has significantly improved the performance of ASR models on datasets like LibriSpeech.However, few studies have addressed the issue of domain mismatch between the data used to pre-train and fine-tune ASR models.Moreover, the Empirical Risk Minimization (ERM) principle, commonly used to train deep learning models, often causes the trained models to exhibit undesirable behaviors such as memorizing training data and being sensitive to adversarial examples.Thus, in this paper, we propose an alternate fine-tuning strategy, called Mix-fine-tune, to address domain mismatch in ASR systems and the limitations of the ERM training principle.Mix-finetune use a data-driven weighted sum of two speech sequences as input and the corresponding text sequences are used to calculate a weighted audio-text alignment Connectionist Temporal Classification (CTC) loss for fine-tuning a pre-trained model.Additionally, Mix-fine-tune incorporates the masked Contrastive Predictive Coding (CPC) loss, previously used exclusively for pre-training, into the fine-tuning process.Our novel strategy alternates between minimizing the CTC loss and the CPC loss to address the domain mismatch between pre-training and fine-tuning.We validate our method by fine-tuning different sizes of the Wav2Vec model using the public Air Traffic Control (ATC) corpus.The experiments show that Mix-fine-tune efficiently adapts the models pre-trained on general speech corpora like LibriSpeech to a specific domain (e.g., the air traffic control domain) by fine-turning.
Chengxi Lei, Satwinder Singh, Feng Hou, Ruili Wang 0001
MMAsia2
2024 A systematic literature review on the significance of deep learning and machine learning in predicting Alzheimer's disease
Arshdeep Kaur, Meenakshi Mittal, Jasvinder Singh Bhatti, Suresh Thareja, Satwinder Singh
Artif. Intell. Medicine5
2023 A Novel Self-training Approach for Low-resource Speech Recognition
Satwinder Singh, Feng Hou, Ruili Wang 0001
INTERSPEECH1
2022 Improved Meta Learning for Low Resource Speech Recognition
abstract
We propose a new meta learning based framework for low resource speech recognition that improves the previous model agnostic meta learning (MAML) approach. The MAML is a simple yet powerful meta learning approach. However, the MAML presents some core deficiencies such as training instabilities and slower convergence speed. To address these issues, we adopt multi-step loss (MSL). The MSL aims to calculate losses at every step of the inner loop of MAML and then combines them with a weighted importance vector. The importance vector ensures that the loss at the last step has more importance than the previous steps. Our empirical evaluation shows that MSL significantly improves the stability of the training procedure and it thus also improves the accuracy of the overall system. Our proposed system outperforms MAML based low resource ASR system on various languages in terms of character error rates and stable training behavior.
Satwinder Singh, Ruili Wang 0001, Feng Hou
ICASSP1
2022 CyclicAugment: Speech Data Random Augmentation with Cosine Annealing Scheduler for Auotmatic Speech Recognition
Zhihan Wang, Feng Hou, Yuanhang Qiu, Zhizhong Ma, Satwinder Singh, Ruili Wang 0001
INTERSPEECH5
2021 DeepF0: End-To-End Fundamental Frequency Estimation for Music and Speech Signals
abstract
We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. f0estimation is important in various speech processing and music information retrieval applications. Existing deep learning models for pitch estimations have relatively limited learning capabilities due to their shallow receptive field. The proposed model addresses this issue by extending the receptive field of a network by introducing the dilated convolutional blocks into the network. The dilation factor increases the network receptive field exponentially without increasing the parameters of the model exponentially. To make the training process more efficient and faster, DeepF0 is augmented with residual blocks with residual connections. Our empirical evaluation demonstrates that the proposed model outperforms the baselines in terms of raw pitch accuracy and raw chroma accuracy even using 77.4% fewer network parameters. We also show that our model can capture reasonably well pitch estimation even under the various levels of accompaniment noise.
Satwinder Singh, Ruili Wang 0001, Yuanhang Qiu
ICASSP1
2021 Self-Supervised Learning Based Phone-Fortified Speech Enhancement
Yuanhang Qiu, Ruili Wang 0001, Satwinder Singh, Zhizhong Ma, Feng Hou
Interspeech3