Anil Kumar Vuppala

dblp:15/8468 · also Anil Vuppala · DBLP profile ↗
← Back
37ranked-venue papers
1as first author
18since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 17 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 12 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Typical vs. Atypical Disfluency Classification: Introducing the IIITH-TISA Corpus and Temporal Context-Based Feature Representations
abstract
Speech disfluencies in spontaneous communication can be categorized as either typical or atypical. Typical disfluencies, such as hesitations and repetitions, are natural occurrences in everyday speech, while atypical disfluencies are indicative of pathological disorders like stuttering. Distinguishing between these categories is crucial for improving voice assistants (VAs) for Persons Who Stutter (PWS), who often face premature cutoffs due to misidentification of speech termination. Accurate classification also aids in detecting stuttering early in children, preventing misdiagnosis as language development disfluency. This research introduces the IIITH-TISA dataset, the first Indian English stammer corpus, capturing atypical disfluencies. Additionally, we extend the IIITH-IED dataset with detailed annotations for typical disfluencies. We propose Perceptually Enhanced Zero-Time Windowed Cepstral Coefficients (PE-ZTWCC) combined with Shifted Delta Cepstra (SDC) as input features to a shallow Time Delay Neural Network (TDNN) classifier, capturing both local and wider temporal contexts. Our method achieves an average F1 score of 85.01% for disfluency classification, outperforming traditional features.
Priyanka Kommagouni, Vamshiraghusimha Narasinga, Purva Barche, Sai Akarsh C, Anil Kumar Vuppala
ICASSP5
2025 A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
abstract
Automatic detection and severity assessment of dysarthria are crucial for delivering targeted therapeutic interventions to patients. While most existing research focuses primarily on speech modality, this study introduces a novel approach that leverages both speech and text modalities. By employing cross-attention mechanism, our method learns the acoustic and linguistic similarities between speech and text representations. This approach assesses specifically the pronunciation deviations across different severity levels, thereby enhancing the accuracy of dysarthric detection and severity assessment. All the experiments have been performed using UA-Speech dysarthric database. Improved accuracies of 99.53% and 93.20% in detection, and 98.12% and 51.97% for severity assessment have been achieved when speaker-dependent and speaker-independent, unseen and seen words settings are used. These findings suggest that by integrating text information, which provides a reference linguistic knowledge, a more robust framework has been developed for dysarthric detection and assessment, thereby potentially leading to more effective diagnoses.
Anuprabha M, Krishna Gurugubelli, Kesavaraj V, Anil Kumar Vuppala
ICASSP4
2025 Enhancing Stutter Detection using Long-Term Average Spectrum Values
abstract
Stuttering is recognized as a prevalent speech disorder that significantly affects individuals worldwide. Identifying and diagnosing in the early stages enhances the quality of life for individuals experiencing atypical speech patterns. Traditional methods for classifying stuttering primarily depend on subjective assessments and short-term acoustic analysis. Although helpful, these methods face limitations in accurately capturing all stutter types due to their inherent subjectivity and temporal constraints. This study uses the long-term average spectrum (LTAS) values for stutter classification derived from various filter banks such as Constant Q, Gamma-tone, and Single-frequency filter banks. It also compares these LTAS-based methods with cepstral coefficients, such as MFCC and ZTWCC. Classifiers such as SVM, LSTM, and Bi-LSTM deep networks were used to study the effectiveness of these representations in accurately discerning stuttered speech from fluent speech and reported the results.
Vamshiraghusimha Narasinga, Priyanka Kommagouni, Sridhar Vanga, Kowshik Siva Sai Motepalli, Sai Akarsh C, Purva Barche, Anil Kumar Vuppala
ICASSP7
2025 Towards Classification of Typical and Atypical Disfluencies: A Self Supervised Representation Approach
Priyanka Kommagouni, Pragya Khanna, Vamshiraghusimha Narasinga, Anirudh Bocha, Anil Kumar Vuppala
INTERSPEECH5
2025 Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS
Anuprabha M, Krishna Gurugubelli, Anil Kumar Vuppala
INTERSPEECH3
2025 ProBiEM: Acoustic and Lexical Correlates of Prosodic Prominence in English-Malayalam Bilingual Speech
Anindita Mondal, Rahul Biju, Anil Kumar Vuppala, Reni K. Cherian, Chiranjeevi Yarra
INTERSPEECH3
2025 ExagTTS: An Approach Towards Controllable Word Stress Incorporated TTS for Exaggerated Synthesized Speech Aiding Second Language Learners
Anindita Mondal, Monica Surtani, Anil Kumar Vuppala, Parameswari Krishnamurthy, Chiranjeevi Yarra
INTERSPEECH3
2025 End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
Aishwarya Pothula, Bhavana Akkiraju, Srihari Bandarupalli, Charan Devarkonda, Santosh Kesiraju, Anil Kumar Vuppala
INTERSPEECH6
2024 End-to-End User-Defined Keyword Spotting Using Shifted Delta Coefficients
Kesavaraj V, Anuprabha M, Anil Kumar Vuppala
ICPR (22)3
2024 Stress transfer in speech-to-speech machine translation
Sai Akarsh C, Vamshiraghusimha Narasinga, Anil Kumar Vuppala
INTERSPEECH3
2024 Custom wake word detection
Kesavaraj V, Charan Devarkonda, Vamshiraghusimha Narasinga, Anil Kumar Vuppala
INTERSPEECH4
2023 An Investigation of Indian Native Language Phonemic Influences on L2 English Pronunciations
Shelly Jain, Priyanshi Pal, Anil Kumar Vuppala, Prasanta Kumar Ghosh, Chiranjeevi Yarra
INTERSPEECH3
2023 Stuttering Detection Application
Kowshik Siva Sai Motepalli, Vamshiraghusimha Narasinga, Harsha Pathuri, Hina Khan, Sangeetha Mahesh, Ajish K. Abraham, Anil Kumar Vuppala
INTERSPEECH7
2023 IIITH-CSTD Corpus: Crowdsourced Strategies for the Collection of a Large-scale Telugu Speech Corpus
abstract
Due to the lack of a large annotated speech corpus, many low-resource Indian languages struggle to utilize recent advancements in deep neural network architectures for Automatic Speech Recognition (ASR) tasks. Collecting large-scale databases is an expensive and time-consuming task. Current approaches lack extensive traditional expert-based data acquisition guidelines, as they are tedious and complex. In this work, we present the International Institute of Information Technology Hyderabad-Crowd Sourced Telugu Database (IIITH-CSTD), a Telugu corpus collected through crowdsourcing. In particular, our main objective is to mitigate the low-resource problem for Telugu. We also present the sources, crowdsourcing pipeline, and the protocols used to collect the corpus for a low-resource language, namely, Telugu. Data of approximately 2,000 hours of transcribed audio is presented and released in this article, covering three major regional dialects of the Telugu language in three different (i.e., read, conversational and spontaneous) speaking styles on topics such as politics, sports, and arts, science, and so on. 1 We also present the experimental results of the collected corpus on ASR tasks. We hope this work will motivate researchers to curate large-scale annotated speech data for other low-resource Indic languages.
Sai Ganesh Mirishkar, Vishnu Vidyadhara Raju Vegesna, Meher Dinesh Naroju, Sudhamay Maity, Prakash Yalla, Anil Kumar Vuppala
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2022 Multi-Task End-to-End Model for Telugu Dialect and Speech Recognition
Aditya Yadavalli, Sai Ganesh Mirishkar, Anil Kumar Vuppala
INTERSPEECH3
2022 How do Phonological Properties Affect Bilingual Automatic Speech Recognition?
abstract
Multilingual Automatic Speech Recognition (ASR) for Indian languages is an obvious technique for leveraging their similarities. We present a detailed analysis of how phonological similarities and differences between languages affect Time Delay Neural Network (TDNN) and End-to-End (E2E) ASR. To study this, we select genealogically similar pairs from five Indian languages and train bilingual acoustic models. We compare these against corresponding monolingual acoustic models and find similar phoneme distributions within speech to be the primary factor for improving model performance, with phoneme overlap being secondary. The influence of phonological properties on performance is visible in both cases. Word Error Rate (WER) of E2E decreased by a median of 2.35%, and upto 8.5% when the phonological similarity was greatest. WER of TDNN increased by 11.69% when the similarity was lowest. Thus, it is clear that the choice of supplementary language is important for model performance.
Shelly Jain, Aditya Yadavalli, Sai Ganesh Mirishkar, Anil Kumar Vuppala
SLT4
2021 Comparative Study of Different Epoch Extraction Methods for Speech Associated with Voice Disorders
abstract
Accurate detection of epoch locations is important in extracting the features from the speech signal for automatic detection and assessment of voice disorders. Therefore, this study aimed to compare the various algorithms for detecting epoch locations from the speech associated with voice disorders. In this regard, nine state-of-the-art epoch extraction algorithms were considered, and their performance for different categories of voice disorders was evaluated on the SVD dataset. Experimental results indicate that most of the epoch extraction methods showed better performance for healthy speech; however, their performance was degraded for speech associated with voice disorders. Furthermore, the performance of epoch extraction methods was degraded for the speech of structural and neurogenic disorders compared to the speech of psychogenic and functional disorders. Among the different epoch extraction algorithms, zero phase-zero frequency filtering showed the best performance in terms identification rate (90.37%) and identification accuracy (0.34ms), for speech associated with voice disorders.
Purva Barche, Krishna Gurugubelli, Anil Kumar Vuppala
ICASSP3
2021 Reed: An Approach Towards Quickly Bootstrapping Multilingual Acoustic Models
abstract
Multilingual automatic speech recognition (ASR) system is a single entity capable of transcribing multiple languages sharing a common phone space. Performance of such a system is highly dependent on the compatibility of the languages. State of the art speech recognition systems are built using sequential architectures based on recurrent neural networks (RNN) limiting the computational parallelization in training. This poses a significant challenge in terms of time taken to bootstrap and validate the compatibility of multiple languages for building a robust multilingual system. Complex architectural choices based on self-attention networks are made to improve the parallelization thereby reducing the training time. In this work, we propose Reed, a simple system based on 1D convolutions which uses very short context to improve the training time. To improve the performance of our system, we use raw time-domain speech signals directly as input. This enables the convolutional layers to learn feature representations rather than relying on handcrafted features such as MFCC. We report improvement on training and inference times by atleast a factor of 4× and 7.4× respectively with comparable WERs against standard RNN based baseline systems on SpeechOcean's multilingual low resource dataset.
Bipasha Sen, Aditya Agarwal, Sai Ganesh Mirishkar, Anil Kumar Vuppala
SLT4
2020 Single Frequency Filter Bank Based Long-Term Average Spectra for Hypernasality Detection and Assessment in Cleft Lip and Palate Speech
abstract
Hypernasality is an abnormality in speech production observed in subjects with craniofacial anomalies like cleft lip and palate (CLP). Detection and assessment of hypernasality is a primary step in the clinical diagnosis of individuals with CLP. Existing methods explore the short-term spectral information from speech to assess hy-pernasality. The present work examines long-term average spectral (LTAS) features obtained from speech to detect and assess hyper-nasality. This work proposes single frequency filter bank based long-term average spectral (SFFB-LTAS) features for hypernasality detection and assessment. The SFFB is used to extract long-term average spectra with a good spectral resolution. The experiments are carried out using NMCPC-CLP database collected from 41 speakers with CLP and 32 speakers without CLP. The experimental results show that, SFFB-LTAS features performed better compared to state-of-art spectral and prosody features. The proposed systems for the detection and assessment of hypernasality have shown classification accuracy of 89% and 82.1%, respectively.
Mohammad Hashim Javid, Krishna Gurugubelli, Anil Kumar Vuppala
ICASSP3
2020 Towards Automatic Assessment of Voice Disorders: A Clinical Approach
Purva Barche, Krishna Gurugubelli, Anil Kumar Vuppala
INTERSPEECH3
2020 Analytic phase features for dysarthric speech detection and intelligibility assessment
Krishna Gurugubelli, Anil Kumar Vuppala
Speech Commun.2
2020 Duration of the rhotic approximant /ɹ/ in spastic dysarthria of different severity levels
Krishna Gurugubelli, Anil Kumar Vuppala, N. P. Narendra, Paavo Alku
Speech Commun.2
2019 An Investigation of LSTM-CTC based Joint Acoustic Model for Indian Language Identification
abstract
In this paper, phonetic features derived from the joint acoustic model (JAM) of a multilingual end to end automatic speech recognition system are proposed for Indian language identification (LID). These features utilize contextual information learned by the JAM through long short-term memory-connectionist temporal classification (LSTM-CTC) framework. Hence, these features are referred to as CTC features. A multi-head self-attention network is trained using these features, which aggregates the frame-level features by selecting prominent frames through a parametrized attention layer. The proposed features have been tested on IIITH-ILSC database that consists of 22 official Indian languages and Indian English. Experimental results demonstrate that CTC features outperformed i-vector and phonetic temporal neural LID systems and produced an 8.70% equal error rate. The fusion of shifted delta cepstral and CTC feature-based LID systems at the model level and feature level further improved the performance.
Tirusha Mandava, Ravi Kumar Vuddagiri, Hari Krishna Vydana, Anil Kumar Vuppala
ASRU4
2019 Perceptually Enhanced Single Frequency Filtering for Dysarthric Speech Detection and Intelligibility Assessment
abstract
This paper proposes a new speech feature representation that improves the intelligibility assessment of dysarthric speech. The formulation of the feature set is motivated from the human auditory perception and high time-frequency resolution property of single frequency filtering (SFF) technique. The proposed features are named as perceptually enhanced single frequency cepstral coefficients (PE-SFCC). As a part of SFF technique implementation, speech signal passed through a single pole complex bandpass filter bank to obtain high-resolution time-frequency distribution. Then, the distribution is enhanced by using a set of auditory perceptual operators. Lastly, traditional homomorphic analysis has been carried out on the resulting signal to obtain PE-SFCC feature vector. The performance of proposed features in dysarthric speech detection and its intelligibility assessment has been reported on UASPEECH database. The PE-SFCC features outperformed the state-of-the-art features in dysarthric speech detection and intelligibility assessment.
Krishna Gurugubelli, Anil Kumar Vuppala
ICASSP2
2019 IIIT-H Spoofing Countermeasures for Automatic Speaker Verification Spoofing and Countermeasures Challenge 2019
K. N. R. K. Raju Alluri, Anil Kumar Vuppala
INTERSPEECH2
2019 Sound Privacy: A Conversational Speech Corpus for Quantifying the Experience of Privacy
abstract
With the growing popularity of social networks, cloud services and online applications, people are becoming concerned about the way companies store their data and the ways in which the data can be applied. Privacy with devices and services operated by the voice are of particular interest. To enable studies in privacy, this paper presents a database which quantifies the experience of privacy users have in spoken communication. We focus on the effect of the acoustic environment on that perception of privacy. Speech signals are recorded in scenarios simulating real-life situations, where the acoustic environment has an effect on the experience of privacy. The acoustic data is complemented with measures of the speakers’ experience of privacy, recorded using a questionnaire. The presented corpus enables studies in how acoustic environments affect peoples’ experience of privacy, which in turn, can be used to develop speech operated applications which are respectful of their right to privacy.
Pablo Pérez Zarazaga, Sneha Das, Tom Bäckström, Vishnu Vidyadhara Raju Vegesna, Anil Kumar Vuppala
INTERSPEECH5
2019 Application of Emotion Recognition and Modification for Emotional Telugu Speech Recognition
Vishnu Vidyadhara Raju Vegesna, Krishna Gurugubelli, Anil Kumar Vuppala
Mob. Networks Appl.3
2019 Stable Implementation of Zero Frequency Filtering of Speech Signals for Efficient Epoch Extraction
abstract
Epochs are the abrupt-closure events in vocal fold vibration during the production of voiced speech. Zero frequency filtering is a simple and effective technique used to estimate the glottal closure instants accurately from the speech signal. However, the zero frequency filter is an unstable system. Hence, it may not be suitable for practical implementation due to the requirements of high precision computation. In this letter, zero-phase zero frequency resonator is proposed as an alternative to zero frequency filter. The proposed approach provides a stable zero-phase response. The experimental results indicate that the performance of the proposed method outperformed the state-of-the-art methods in terms of identification rate 99.17% and provides comparable performance in terms of false alarm rate (0.41%), and identification accuracy (0.28 ms).
Krishna Gurugubelli, Anil Kumar Vuppala
IEEE Signal Process. Lett.2
2018 An Exploration towards Joint Acoustic Modeling for Indian Languages: IIIT-H Submission for Low Resource Speech Recognition Challenge for Indian Languages, INTERSPEECH 2018
Hari Krishna Vydana, Krishna Gurugubelli, Vishnu Vidyadhara Raju Vegesna, Anil Kumar Vuppala
INTERSPEECH4
2018 Curriculum learning based approach for noise robust language identification using DNN with attention
Ravi Kumar Vuddagiri, Hari Krishna Vydana, Anil Kumar Vuppala
Expert Syst. Appl.3
2018 Improved vowel region detection from a continuous speech using post processing of vowel onset points and vowel end-points
Ramakrishna Thirumuru, Suryakanth V. Gangashetty, Anil Kumar Vuppala
Multim. Tools Appl.3
2017 Significance of neural phonotactic models for large-scale spoken language identification
abstract
Language identification (LID) is vital frontend for spoken dialogue systems operating in diverse linguistic settings to reduce recognition and understanding errors. Existing LID systems which use low-level signal information for classification do not scale well due to exponential growth of parameters as the classes increase. They also suffer performance degradation due to the inherent variabilities of speech signal. In the proposed approach, we model the language-specific phonotactic information in speech using recurrent neural network for developing an LID system. The input speech signal is tokenized to phone sequences by using a common language-independent phone recognizer with varying phonetic coverage. We establish a causal relationship between phonetic coverage and LID performance. The phonotactics in the observed phone sequences are modeled using statistical and recurrent neural network language models to predict language-specific symbol from a universal phonetic inventory. Proposed approach is robust, computationally light weight and highly scalable. Experiments show that the convex combination of statistical and recurrent neural network language model (RNNLM) based phonotactic models significantly outperform a strong baseline system of Deep Neural Network (DNN) which is shown to surpass the performance of i-vector based approach for LID. The proposed approach outperforms the baseline models in terms of mean F1 score over 176 languages. Further we provide significant information-theoretic evidence to analyze the mechanism of the proposed approach.
Brij Mohan Lal Srivastava, Hari Krishna Vydana, Anil Kumar Vuppala, Manish Shrivastava 0001
IJCNN3
2017 SFF Anti-Spoofer: IIIT-H Submission for Automatic Speaker Verification Spoofing and Countermeasures Challenge 2017
K. N. R. K. Raju Alluri, Sivanand Achanta, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Anil Kumar Vuppala
INTERSPEECH5
2017 Detection of Replay Attacks Using Single Frequency Filtering Cepstral Coefficients
K. N. R. K. Raju Alluri, Sivanand Achanta, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Anil Kumar Vuppala
INTERSPEECH5
2016 An Investigation of Deep Neural Network Architectures for Language Recognition in Indian Languages
Mounika K. V., Sivanand Achanta, Lakshmi H. R., Suryakanth V. Gangashetty, Anil Kumar Vuppala
INTERSPEECH5
2013 Non-uniform time scale modification using instants of significant excitation and vowel onset points
K. Sreenivasa Rao, Anil Kumar Vuppala
Speech Commun.2
2012 Vowel Onset Point Detection for Low Bit Rate Coded Speech
abstract
In this paper, we propose a method for detecting the vowel onset points (VOPs) for low bit rate coded speech. VOP is the instant at which the onset of the vowel takes place in the speech signal. VOP plays an important role for the applications, such as consonant-vowel (CV) unit recognition and speech rate modification. The proposed VOP detection method is based on the spectral energy present in the glottal closure region of the speech signal. Speech coders considered to carry out this study are Global System for Mobile Communications (GSM) full rate, code-excited linear prediction (CELP), and mixed-excitation linear prediction (MELP). TIMIT database and CV units collected from the broadcast news corpus are used for evaluation. Performance of the proposed method is compared with existing methods, which uses the combination of evidence from the excitation source, spectral peaks energy, and modulation spectrum. The proposed VOP detection method has shown significant improvement in the performance compared to the existing method under clean as well as coded cases. The effectiveness of the proposed VOP detection method is analyzed in CV recognition by using VOP as an anchor point.
Anil Kumar Vuppala, Jainath Yadav, Saswat Chakrabarti, K. Sreenivasa Rao
IEEE Trans. Speech Audio Process.1