Jhansi Mallela

dblp:263/4695 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0009-6988-2682ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 LexiRep: Iterative One-Shot Linguistically Constrained Lexical Stress Representations Learning
abstract
Lexical stress detection is vital for effective Computer-Assisted Language Learning (CALL) systems and is typically modeled as a binary task, labeling syllables as stressed or unstressed. However, detecting stress in non-native speech is challenging due to deviations from native (canonical) patterns and limited annotated data. We propose LexiRep, a one-shot framework utilizing linguistic knowledge for learning lexical stress representations. It combines a linguistically-guided contrastive framework-based discriminative stress representation learning with Improved Deep Embedded Clustering (IDEC) and iteratively refines the representations under linguistic constraints. We evaluated German and Italian non-native English speech with transfer learning from native data and observed that LexiRep achieves a 3% absolute improvement over the respective state-of-the-art while reducing model complexity by 99%. Further, LexiRep outperforms by 7.62% in an unsupervised setting, demonstrating the effectiveness of linguistically constrained contrastive learning for low-resource stress detection.
Jhansi Mallela, Rangavajjala Sankara Bharadwaj, Chiranjeevi Yarra
IEEE Signal Process. Lett.1
2025 Post-Net2.0: An adaptive weighted loss function driven by linguistic constraint for automatic syllable stress detection
abstract
Automatic syllable stress detection is an essential component in Computer assisted language learning (CALL) systems to guide nonnative language learners. In English, each word typically contains only one primary stressed syllable. However, standard loss functions, such as Binary Cross-Entropy (BCE), often result in predictions where multiple syllables may be stressed or none at all. As a result, automatic syllable stress detection models frequently require an additional post-processing step to ensure that only one syllable is stressed per word. This reliance on post-processing suggests that the model is not fully capturing the stress patterns accurately. To address this issue, we propose an adaptive weighted loss function that builds upon the Stress Intensity Modulation Loss proposed in our recent work of Post-Net. This adaptive weighted loss function is designed to enforce the constraint of a single primary stressed syllable directly during model training. We integrate this loss function into the previously proposed Post-Net (PN_DNN) and on a new architecture which is a hybrid of Post-Net and LSTM (PN_DLSTM). Their performance is compared against the state-of-the-art models trained with standard BCE loss. Experiments conducted on the ISLE corpus reveal that both the models trained only with BCE loss show a significant accuracy gap between with and without post-processing. In contrast, when these models are trained on the proposed adaptive weighted loss function, the gap is narrowed in both the models. Between the two models, the highest reduction is observed in PN_DNN with a decrease from 3.87% to 2.45% & 4.6% to 3.75% for GER & ITA respectively. This indicates that the adaptive weighted loss function effectively captures the linguistic constraint during training, reducing the need for post-processing.
Sai Harshitha Aluru, Jhansi Mallela, Chiranjeevi Yarra
ICASSP2
2025 Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation
abstract
Automatic syllable stress detection is a crucial component in Computer-Assisted Language Learning (CALL) systems for language learners. Current stress detection models are typically trained on clean speech, which may not be robust in real-world scenarios where background noise is prevalent. To address this, speech enhancement (SE) models, designed to enhance speech by removing noise, might be employed, but their impact on preserving syllable stress patterns is not well studied. This study examines how different SE models, representing discriminative and generative modeling approaches, affect syllable stress detection under noisy conditions. We assess these models by applying them to speech data with varying signal-to-noise ratios (SNRs) from 0 to 20 dB, and evaluating their effectiveness in maintaining stress patterns. Additionally, we explore different feature sets to determine which ones are most effective for capturing stress patterns amidst noise. To further understand the impact of SE models, a human-based perceptual study is conducted to compare the perceived stress patterns in SE-enhanced speech with those in clean speech, providing insights into how well these models preserve syllable stress as perceived by listeners. Experiments are performed on English speech data from non-native speakers of German and Italian. And the results reveal that the stress detection performance is robust with the generative SE models when heuristic features are used. Also, the observations from the perceptual study are consistent with the stress detection outcomes under all SE models.
Rangavajjala Sankara Bharadwaj, Jhansi Mallela, Sai Harshitha Aluru, Chiranjeevi Yarra
ICASSP2
2025 SupraDoRAL: Automatic Word Prominence Detection Using Suprasegmental Dependencies of Representations with Acoustic and Linguistic Context
Jhansi Mallela, Upendra Vishwanath Y. S., Sankara Bharadwaj Rangavajjala, Bhaskar Bhatt, Chiranjeevi Yarra
INTERSPEECH1
2024 Post-Net: A linguistically inspired sequence-dependent transformed neural architecture for automatic syllable stress detection
Sai Harshitha Aluru, Jhansi Mallela, Chiranjeevi Yarra
INTERSPEECH2
2024 A comparative analysis of sequential models that integrate syllable dependency for automatic syllable stress detection
Jhansi Mallela, Sai Harshitha Aluru, Chiranjeevi Yarra
INTERSPEECH1
2021 Effect of Noise and Model Complexity on Detection of Amyotrophic Lateral Sclerosis and Parkinson's Disease Using Pitch and MFCC
abstract
Dysarthria due to Amyotrophic Lateral Sclerosis (ALS) and Parkinson’s disease (PD) impacts both articulation and prosody in an individual’s speech. Complex deep neural networks exploit these cues for detection of ALS and PD. These are typically done using recordings in laboratory condition. This study aims to examine the robustness of these cues against background noise and model complexity, which has not been investigated before. We perform classification experiments with pitch and Mel-frequency cepstral coefficients (MFCC) using models of three different complexities and additive white Gaussian noise in four signal-to-noise-ratio (SNR) conditions. The findings are as follows: 1) In clean condition, pitch performs similar to MFCC across most model complexities considered, suggesting that one-dimensional pitch pattern provides discriminative cues for the classification to an extent equal to that of multi-dimensional MFCC, 2) Similar trend is observed in noisy cases when classifiers are trained and tested in matched noise and SNR conditions, 3) When the classifiers trained on clean data are applied in noisy cases, pitch based average classification accuracies are found to be 20.09% and 24.73% higher than those using MFCC for ALS vs. healthy and PD vs. healthy, respectively, suggesting robustness of pitch based classifier against noise and model complexity.
Tanuka Bhattacharjee, Jhansi Mallela, Yamini Belur, Nalini Atchayarcmf, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh
ICASSP2
2021 Source and Vocal Tract Cues for Speech-Based Classification of Patients with Parkinson's Disease and Healthy Subjects
Tanuka Bhattacharjee, Jhansi Mallela, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh
Interspeech2
2020 Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learning
abstract
In this paper, we consider 2-class and 3-class classification problems for classifying patients with Amyotrophic Lateral Sclerosis (ALS), Parkinson's Disease (PD), and Healthy Controls (HC) using a CNNLSTM network. Classification performance is examined for three different tasks, namely, Spontaneous speech (SPON), Diadochokinetic rate (DIDK) and Sustained phoneme production (PHON). Experiments are conducted using speech data recorded from 60 ALS, 60 PD, and 60 HC subjects. Classifications using SVM and DNN are considered as baseline schemes. Classification accuracy of ALS and HC (indicated by ALS/HC) using CNN-LSTM has shown an improvement of 10.40%, 4.22% and 0.08% for PHON, SPON and DIDK tasks, respectively over the best of the baseline schemes. Furthermore, the CNN-LSTM network achieves the highest PD/HC classification accuracy of 88.5% for the SPON task and the highest 3-class (ALS/PD/HC) classification accuracy of 85.24% for the DIDK task. Experiments using transfer learning at low resource training data show that data from ALS benefits PD/HC classification and vice-versa. Experiments with fine-tuning weights of 3-class (ALS/PD/HC) classifier for 2-class classification (PD/HC or ALS/HC) gives an absolute improvement of 2% classification accuracy in SPON task when compared with randomly initialized 2-class classifier.
Jhansi Mallela, Aravind Illa, Suhas B. N., Sathvik Udupa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh
ICASSP1
2020 Raw Speech Waveform Based Classification of Patients with ALS, Parkinson's Disease and Healthy Controls Using CNN-BLSTM
Jhansi Mallela, Aravind Illa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh
INTERSPEECH1
2019 Acoustic and Articulatory Feature Based Speech Rate Estimation Using a Convolutional Dense Neural Network
abstract
In this paper, we propose a speech rate estimation approach using a convolutional dense neural network (CDNN). The CDNN based approach uses the acoustic and articulatory features for speech rate estimation. The Mel Frequency Cepstral Coefficients (MFCCs) are used as acoustic features and the articulograms representing time-varying vocal tract profile are used as articulatory features. The articulogram is computed from a real-time magnetic resonance imaging (rtMRI) video in the midsagittal plane of a subject while speaking. However, in practice, the articulogram features are not directly available, unlike acoustic features from speech recording. Thus, we use an Acoustic-to-Articulatory Inversion method using a bidirectional long-short-term memory network which estimates the articulogram features from the acoustics. The proposed CDNN based approach using estimated articulatory features requires both acoustic and articulatory features during training but it requires only acoustic data during testing. Experiments are conducted using rtMRI videos from four subjects each speaking 460 sentences. The Pearson correlation coefficient is used to evaluate the speech rate estimation. It is found that the CDNN based approach gives a better correlation coefficient than the temporal and selected sub-band correlation (TCSSBC) based baseline scheme by 81.58 and 73.68 (relative) in seen and unseen subject conditions respectively.
Renuka Mannem, Jhansi Mallela, Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH2