Chiranjeevi Yarra

dblp:158/4179 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-0574-8777ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LexiRep: Iterative One-Shot Linguistically Constrained Lexical Stress Representations Learning
abstract
Lexical stress detection is vital for effective Computer-Assisted Language Learning (CALL) systems and is typically modeled as a binary task, labeling syllables as stressed or unstressed. However, detecting stress in non-native speech is challenging due to deviations from native (canonical) patterns and limited annotated data. We propose LexiRep, a one-shot framework utilizing linguistic knowledge for learning lexical stress representations. It combines a linguistically-guided contrastive framework-based discriminative stress representation learning with Improved Deep Embedded Clustering (IDEC) and iteratively refines the representations under linguistic constraints. We evaluated German and Italian non-native English speech with transfer learning from native data and observed that LexiRep achieves a 3% absolute improvement over the respective state-of-the-art while reducing model complexity by 99%. Further, LexiRep outperforms by 7.62% in an unsupervised setting, demonstrating the effectiveness of linguistically constrained contrastive learning for low-resource stress detection.
Jhansi Mallela, Rangavajjala Sankara Bharadwaj, Chiranjeevi Yarra
IEEE Signal Process. Lett.3
2025 Post-Net2.0: An adaptive weighted loss function driven by linguistic constraint for automatic syllable stress detection
abstract
Automatic syllable stress detection is an essential component in Computer assisted language learning (CALL) systems to guide nonnative language learners. In English, each word typically contains only one primary stressed syllable. However, standard loss functions, such as Binary Cross-Entropy (BCE), often result in predictions where multiple syllables may be stressed or none at all. As a result, automatic syllable stress detection models frequently require an additional post-processing step to ensure that only one syllable is stressed per word. This reliance on post-processing suggests that the model is not fully capturing the stress patterns accurately. To address this issue, we propose an adaptive weighted loss function that builds upon the Stress Intensity Modulation Loss proposed in our recent work of Post-Net. This adaptive weighted loss function is designed to enforce the constraint of a single primary stressed syllable directly during model training. We integrate this loss function into the previously proposed Post-Net (PN_DNN) and on a new architecture which is a hybrid of Post-Net and LSTM (PN_DLSTM). Their performance is compared against the state-of-the-art models trained with standard BCE loss. Experiments conducted on the ISLE corpus reveal that both the models trained only with BCE loss show a significant accuracy gap between with and without post-processing. In contrast, when these models are trained on the proposed adaptive weighted loss function, the gap is narrowed in both the models. Between the two models, the highest reduction is observed in PN_DNN with a decrease from 3.87% to 2.45% & 4.6% to 3.75% for GER & ITA respectively. This indicates that the adaptive weighted loss function effectively captures the linguistic constraint during training, reducing the need for post-processing.
Sai Harshitha Aluru, Jhansi Mallela, Chiranjeevi Yarra
ICASSP3
2025 Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation
abstract
Automatic syllable stress detection is a crucial component in Computer-Assisted Language Learning (CALL) systems for language learners. Current stress detection models are typically trained on clean speech, which may not be robust in real-world scenarios where background noise is prevalent. To address this, speech enhancement (SE) models, designed to enhance speech by removing noise, might be employed, but their impact on preserving syllable stress patterns is not well studied. This study examines how different SE models, representing discriminative and generative modeling approaches, affect syllable stress detection under noisy conditions. We assess these models by applying them to speech data with varying signal-to-noise ratios (SNRs) from 0 to 20 dB, and evaluating their effectiveness in maintaining stress patterns. Additionally, we explore different feature sets to determine which ones are most effective for capturing stress patterns amidst noise. To further understand the impact of SE models, a human-based perceptual study is conducted to compare the perceived stress patterns in SE-enhanced speech with those in clean speech, providing insights into how well these models preserve syllable stress as perceived by listeners. Experiments are performed on English speech data from non-native speakers of German and Italian. And the results reveal that the stress detection performance is robust with the generative SE models when heuristic features are used. Also, the observations from the perceptual study are consistent with the stress detection outcomes under all SE models.
Rangavajjala Sankara Bharadwaj, Jhansi Mallela, Sai Harshitha Aluru, Chiranjeevi Yarra
ICASSP4
2025 SGED-Probe: Probing E2E ASR decoder and aligner for spoken grammar error detection under three speaking practice conditions
Chowdam Venkata Thirumala Kumar, Chiranjeevi Yarra
INTERSPEECH2
2025 SupraDoRAL: Automatic Word Prominence Detection Using Suprasegmental Dependencies of Representations with Acoustic and Linguistic Context
Jhansi Mallela, Upendra Vishwanath Y. S., Sankara Bharadwaj Rangavajjala, Bhaskar Bhatt, Chiranjeevi Yarra
INTERSPEECH5
2025 ProBiEM: Acoustic and Lexical Correlates of Prosodic Prominence in English-Malayalam Bilingual Speech
Anindita Mondal, Rahul Biju, Anil Kumar Vuppala, Reni K. Cherian, Chiranjeevi Yarra
INTERSPEECH5
2025 ExagTTS: An Approach Towards Controllable Word Stress Incorporated TTS for Exaggerated Synthesized Speech Aiding Second Language Learners
Anindita Mondal, Monica Surtani, Anil Kumar Vuppala, Parameswari Krishnamurthy, Chiranjeevi Yarra
INTERSPEECH5
2025 GoP2Vec: A few shot learning for pronunciation assessment with goodness of pronunciation (GoP) based representations from an i-vector framework and augmentation
Meenakshi Sirigiraju, Chiranjeevi Yarra
INTERSPEECH2
2024 Post-Net: A linguistically inspired sequence-dependent transformed neural architecture for automatic syllable stress detection
Sai Harshitha Aluru, Jhansi Mallela, Chiranjeevi Yarra
INTERSPEECH3
2024 A comparative analysis of sequential models that integrate syllable dependency for automatic syllable stress detection
Jhansi Mallela, Sai Harshitha Aluru, Chiranjeevi Yarra
INTERSPEECH3
2024 IIITH Ucchar e-Sudharak: an automatic English pronunciation corrector for school-going children with a teacher in the loop
Meenakshi Sirigiraju, Arjun Rajasekar, Abhishikth Meejuri, Chiranjeevi Yarra
INTERSPEECH4
2023 An Investigation of Indian Native Language Phonemic Influences on L2 English Pronunciations
Shelly Jain, Priyanshi Pal, Anil Kumar Vuppala, Prasanta Kumar Ghosh, Chiranjeevi Yarra
INTERSPEECH5
2022 mulEEG: A Multi-view Representation Learning on EEG Signals
Vamsi Kumar, Likith Reddy, Shivam Kumar Sharma, Kamalaker Dadi, Chiranjeevi Yarra, Raju S. Bapi, Srijithesh Rajendran
MICCAI (3)5
2022 Automatic syllable stress detection under non-parallel label and data condition
Chiranjeevi Yarra, Prasanta Kumar Ghosh
Speech Commun.1
2021 MUCS 2021: Multilingual and Code-Switching ASR Challenges for Low Resource Indian Languages
abstract
Recently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking advantage of low amounts of labeled corpora in multiple languages. With multilingualism becoming common in today's world, there has been increasing interest in code-switching ASR as well. In code-switching, multiple languages are freely interchanged within a single sentence or between sentences. The success of low-resource multilingual and code-switching ASR often depends on the variety of languages in terms of their acoustics, linguistic characteristics as well as the amount of data available and how these are carefully considered in building the ASR system. In this challenge, we would like to focus on building multilingual and code-switching ASR systems through two different subtasks related to a total of seven Indian languages, namely Hindi, Marathi, Odia, Tamil, Telugu, Gujarati and Bengali. For this purpose, we provide a total of ~600 hours of transcribed speech data, comprising train and test sets, in these languages including two code-switched language pairs, Hindi-English and Bengali-English. We also provide a baseline recipe for both the tasks with a WER of 30.73% and 32.45% on the test sets of multilingual and code-switching subtasks, respectively.
Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah, Ankita Singh, Srinivasa Raghavan K. M., Shreya Khare, Vinit Unni, Saurabh Vyas, Akash Rajpuria, Chiranjeevi Yarra, Ashish R. Mittal, Prasanta Kumar Ghosh, Preethi Jyothi, Kalika Bali, Vivek Seshadri, Sunayana Sitaram, Samarth Bharadwaj, Jai Nanavati, Raoul Nanavati, Karthik Sankaranarayanan
Interspeech10
2021 Noise Robust Pitch Stylization Using Minimum Mean Absolute Error Criterion
abstract
We propose a pitch stylization technique in the presence of pitch halving and doubling errors. The technique uses an optimization criterion based on a minimum mean absolute error to make the stylization robust to such pitch estimation errors, particularly under noisy conditions. We obtain segments for the stylization automatically using dynamic programming. Experiments are performed at the frame level and the syllable level. At the frame level, the closeness of stylized pitch is analyzed with the ground truth pitch, which is obtained using a laryngograph signal, considering root mean square error (RMSE) measure. At the syllable level, the effectiveness of perceptual relevant embeddings in the stylized pitch is analyzed by estimating syllabic tones and comparing those with manual tone markings using the Levenshtein distance measure. The proposed approach performs better than a minimum mean squared error criterion based pitch stylization scheme at the frame level and a knowledge-based tone estimation scheme at the syllable level under clean and 20dB, 10dB and 0dB SNR conditions with five noises and four pitch estimation techniques. Among all the combinations of SNR, noise and pitch estimation techniques, the highest absolute RMSE and mean distance improvements are found to be 6.49Hz and 0.23, respectively. Copyright © 2021 ISCA.
Chiranjeevi Yarra, Prasanta Kumar Ghosh
Interspeech1
2020 Pseudo Likelihood Correction Technique for Low Resource Accented ASR
abstract
With the availability of large data, ASRs perform well on native English but poorly for non-native English data. Training nonnative ASRs or adapting a native English ASR is often limited by the availability of data, particularly for low resource scenarios. A typical HMM-DNN based ASR decoding requires pseudo-likelihood of states given an acoustic observation, which changes significantly from native to non-native speech due to accent variation. In order to improve the performance of a native English ASR on non-native English data, we, in this work, propose a DNN-based pseudo-likelihood correction (PLC) technique, in which a non-native pseudo-likelihood vector is mapped to match its native counterpart. Instead of correcting all elements of a non-native pseudo-likelihood vector, a loss function is proposed to correct only top few of them. Experiments with one native and multiple Indian English corpora show an improvement of WER by ~12% and over ~5% using the proposed PLC technique unadapted and adapted native English ASR respectively, when recognition is performed on an Indian English corpus different from that used for both PLC and adaptation. Experiments with upto 2 hours of parallel native and non-native English data reveal that, PLC performs better than adaptation for all unseen cases considered.
Avni Rajpal, M. V. Achuth Rao, Chiranjeevi Yarra, Ritu Aggarwal, Prasanta Kumar Ghosh
ICASSP3
2019 ASR Inspired Syllable Stress Detection for Pronunciation Evaluation Without Using a Supervised Classifier and Syllable Level Features
abstract
ABSTRAKTBu məqalədə milli iqtisadiyyatın özünəməxsus xüsusiyyətlərini özündə cəmləşdirən DSÜT modelinin ekonometrik qiymətləndirməsi aparılır.Empirik qiymətləndirmə zamanı rüblük məlumatlara əsaslanılır və müşahidə sayının məhdud olması nəzərə alınıb Bayez metodlarına müraciət edilir.Qiymətləndirmələr bir sıra maraqlı məqamların ortaya çıxmasına şərait yaradır.İlk öncə məlum olur ki, milli iqtisadiyyat üçün qurulan əvvəlki yeni Keynezçi modellərdə istifadə edilən bir sıra vacib parametrlərin kalibrasiya praktikası əldə edilən empirik nəticələrlə uzlaşmır və ölkənin xüsusiyyətlərini əks etdirmir.İkincisi, aydın olur ki, əksər struktur parametrlər dövrü stabillik sərgiləsələr də, milli iqtisdiyyatı sarsan şokların strukturunda mühim dəyişiklər baş vermişdir.Bu tapıntı post-neft bumu dövründə proqnozlaşdırma işini çətinləşdirən amillərdən biri hesab oluna bilər.Üçüncüsü, pul kütləsi üzrə qiymətləndirilən parametrlərin identifikasiyasında problemlərin mövcud olduğu aşkardır.Bununla yanaşı, qiymətləndirilən model bir sıra adekvatlıq sınaqlarından uğurla keçir.Model siyasət qurumları tərəfindən müxtəlif ssenari analizlərinin aparılması və proqnozlaşdırma məqsədi üçün istifadə edilə bilər.Həmçinin, qurulan model ölkə iqtisadiyyatının özünəməxsusluğunu özündə ehtiva edən və ekonometrik qiymətləndirməsi aparılan ilk DSÜT modeli olması səbəbindən də maraq kəsb edir.
Manoj Kumar Ramanathi, Chiranjeevi Yarra, Prasanta Kumar Ghosh
INTERSPEECH2
2019 Low Resource Automatic Intonation Classification Using Gated Recurrent Unit (GRU) Networks Pre-Trained with Synthesized Pitch Patterns
abstract
Second language learners of British English (BE) are typically trained to learn four intonation classes - Glide-up, Glide-down, Dive and Take-off. We predict the intonation class in a learner's utterance by modeling the temporal dependencies in the pitch patterns with gated recurrent unit (GRU) networks. For these, we pre-train the GRU network using a set of synthesized pitch patterns representing each intonation class. For the synthesis, we propose to obtain pitch patterns from the tone sequences representing each intonation class obtained from domain knowledge. Experiments are conducted on speech data collected from experts in a spoken English training material for teaching BE intonation. The absolute improvements in the unweighted average recall (UAR) using the proposed scheme with pre-training are found to be 4.14 and 6.01 respectively over the proposed approach without pre-training and the baseline scheme that uses hidden Markov models (HMMs).
Atreyee Saha, Chiranjeevi Yarra, Prasanta Kumar Ghosh
INTERSPEECH2
2019 An Improved Goodness of Pronunciation (GoP) Measure for Pronunciation Evaluation with DNN-HMM System Considering HMM Transition Probabilities
abstract
Goodness of pronunciation (GoP) is typically formulated with Gaussian mixture model-hidden Markov model (GMM-HMM) based acoustic models considering HMM state transition probabilities (STPs) and GMM likelihoods of context dependent phonemes. On the other hand, deep neural network (DNN)HMM based acoustic models employed sub-phonemic (senone) posteriors instead of GMM likelihoods along with STPs. However, each senone is shared across many states; thus, there is no one-to-one correspondence between them. In order to circumvent this, most of the existing works have proposed modifications to the GoP formulation considering only posteriors neglecting the STPs. In this work, we derive a formulation for the GoP and it results in the formulation involving both senone posteriors and STPs. Further, we illustrate the steps to implement the proposed GoP formulation in Kaldi, a state-of-the-art automatic speech recognition toolkit. Experiments are conducted on English data collected from Indian speakers using acoustic models trained with native English data from LibriSpeech and Fisher-English corpora. The highest improvement in the correlation coefficient between the scores from the formulations and the expert ratings is found to be 14.89 (relative) better with the proposed approach compared to the best of the existing formulations that don't include STPs.
Sweekar Sudhakara, Manoj Kumar Ramanathi, Chiranjeevi Yarra, Prasanta Kumar Ghosh
INTERSPEECH3
2019 SPIRE-fluent: A Self-Learning App for Tutoring Oral Fluency to Second Language English Learners
Chiranjeevi Yarra, Aparna Srinivasan, Sravani Gottimukkala, Prasanta Kumar Ghosh
INTERSPEECH1
2018 Concatenative Articulatory Video Synthesis Using Real-Time MRI Data for Spoken Language Training
abstract
Spoken language training benefits from showing a video of native speakers' articulatory movements to train the second language learners. Typically, the articulatory video is prepared in conjunction with the audio which is collected simultaneously with the articulatory recording. Articulatory video recording requires specialized equipment and, hence, is expensive and time consuming. In this work, we propose a concatenative synthesis approach to obtain articulatory videos for an audio, which may not have a simultaneous articulatory recording. In the training stage of the proposed approach, we make a repository for phoneme specific articulatory image sequence from the available articulatory video. During testing, image sequences are selected from this repository to ensure a smooth transition across phonetic events. The selected image sequences are finally stitched to synthesize the articulatory video for the test audio. Articulatory videos are synthesized for 50 words randomly selected from the MRI-TIMIT database, not seen in the training data. Subjective evaluation on the quality of the synthesized videos using twelve subjects suggests that the videos are close to the original ones with a rating of 3.78 out of 5, where a score of 5 (1) indicates that there is no (great) difference in quality between the original and the synthesized videos.
Urvish Desai, Chiranjeevi Yarra, Prasanta Kumar Ghosh
ICASSP2
2018 Intonation tutor by SPIRE (In-SPIRE): An Online Tool for an Automatic Feedback to the Second Language Learners in Learning Intonation
Anand P. A, Chiranjeevi Yarra, N. K. Kausthubha, Prasanta Kumar Ghosh
INTERSPEECH2
2018 Automatic Visual Augmentation for Concatenation Based Synthesized Articulatory Videos from Real-time MRI Data for Spoken Language Training
Chandana Srinivasan, Chiranjeevi Yarra, Ritu Aggarwal, Sanjeev Kumar Mittal, N. K. Kausthubha, Raseena K. T, Astha Singh, Prasanta Kumar Ghosh
INTERSPEECH2
2018 SPIRE-SST: An Automatic Web-based Self-learning Tool for Syllable Stress Tutoring (SST) to the Second Language Learners
Chiranjeevi Yarra, Anand P. A, N. K. Kausthubha, Prasanta Kumar Ghosh
INTERSPEECH1
2017 Automatic detection of syllable stress using sonority based prominence features for pronunciation evaluation
abstract
Automatic syllable stress detection is useful in assessing and diagnosing the quality of the pronunciation of second language (L2) learners in an automated way. Typically, the syllable stress depends on three prominence measures - intensity level, duration, pitch - around the sound unit with the highest sonority in the respective syllable. Stress detection is often formulated as a binary classification task using cues from the feature contours representing the prominence measures. We observe that cues from a feature contour obtained by incorporating relative sonority levels in the prominence measures are more indicative of the syllable stress compared to those from the feature contours representing only the prominence measures. Based on this observation, we propose a new feature contour based on temporal correlation selected sub-band correlation with an optimal set of sub-bands, called sonorous sub-bands, to maximize the stress detection accuracy. Experiments on ISLE corpus show that, for German and Italian non-native English speakers, the syllable stress detection accuracies (87.53% and 86.26%) are higher when the proposed features are used compared to the baseline accuracies (85.81% and 83.17%) indicating the effectiveness of the sonority based prominence features.
Chiranjeevi Yarra, Om Deshmukh, Prasanta Kumar Ghosh
ICASSP1
2016 A robust speech rate estimation based on the activation profile from the selected acoustic unit dictionary
abstract
A typical solution for the speech rate estimation consists of two stages, which involves first computing a short-time feature contour such that most of peaks of the contour correspond to the syllable nuclei followed by the detection of the peaks of the contour corresponding to the syllable nuclei. Temporal correlation selected subband correlation (TCSSBC) is often used as a feature contour for the speech rate estimation in which correlation within and across a few selected sub-band energies are computed. In this work, instead of a fixed set of sub-bands, we learn them in a data-driven manner using a dictionary learning approach. Similarly, instead of the energy contours, we use the activation profile from the learned dictionary elements. We found that the peaks detected from the data-driven approach significantly improve the speech rate estimation when combined with the traditional TCSSBC approach using a proposed peak-merging strategy. Experiments are performed separately using Switchboard, TIMIT and CTIMIT corpora. Except Switchboard, the correlation coefficient for the speech rate estimation using the proposed approach is found to be higher than those by the TCSSBC technique - 3.1% and 5.2% (relative) improvements for TIMIT and CTIMIT respectively.
Supriya Nagesh, Chiranjeevi Yarra, Om Deshmukh, Prasanta Kumar Ghosh
ICASSP2
2016 A mode-shape classification technique for robust speech rate estimation and syllable nuclei detection
abstract
Acoustic feature based speech (syllable) rate estimation and syllable nuclei detection are important problems in automatic speech recognition (ASR), computer assisted language learning (CALL) and fluency analysis. A typical solution for both the problems consists of two stages. The first stage involves computing a short-time feature contour such that most of the peaks of the contour correspond to the syllabic nuclei. In the second stage, the peaks corresponding to the syllable nuclei are detected. In this work, instead of the peak detection, we perform a mode-shape classification, which is formulated as a supervised binary classification problem – mode-shapes representing the syllabic nuclei as one class and remaining as the other. We use the temporal correlation and selected sub-band correlation (TCSSBC) feature contour and the mode-shapes in the TCSSBC feature contour are converted into a set of feature vectors using an interpolation technique. A support vector machine classifier is used for the classification. Experiments are performed separately using Switchboard, TIMIT and CTIMIT corpora in a five-fold cross validation setup. The average correlation coefficients for the syllable rate estimation turn out to be 0.6761, 0.6928 and 0.3604 for three corpora respectively, which outperform those obtained by the best of the existing peak detection techniques. Similarly, the average F -scores (syllable level) for the syllable nuclei detection are 0.8917, 0.8200 and 0.7637 for three corpora respectively.
Chiranjeevi Yarra, Om Deshmukh, Prasanta Kumar Ghosh
Speech Commun.1
2014 Comparison of speech quality with and without sensors in electromagnetic articulograph AG 501 recording
abstract
In the recordings using electromagnetic articulograph AG 501, sensors are glued to subject’s articulators such as jaw, lips and tongue and both speech and articulatory movements are simultaneously recorded. In this work, we study the effect of the presence of the sensors on the quality of speech spoken by the subject. This is done by recording when a subject speaks a set of 19 VCV stimuli while sensors are attached to subject’s articulators. For comparison we also record the same set of stimuli spoken by the same subject but with no sensors attached to subject’s articulators. Both subjective and objective comparisons are made on the recorded stimuli in these two settings. Subjective evaluation is carried out using 16 evaluators. Listening experiments with recordings from five subjects show that the recordings with sensors attached are significantly different from those without sensors attached in terms of human recognition score as well as on a perceptual difference measure. This is also supported in the objective comparison which computes dissimilarity measure using the spectral shape information. Index Terms: Electromagnetic Articulography, speech quality, listening test
Nisha Meenakshi, Chiranjeevi Yarra, Yamini Belur, Prasanta Kumar Ghosh
INTERSPEECH2