Mittapalle Kiran Reddy

dblp:203/2309 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-7987-1735ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Classification of phonation types in singing and speaking voice using self-supervised learning models
Prathamesh Parasharam Patil, Mittapalle Kiran Reddy, Paavo Alku
Speech Commun.2
2025 Towards robust heart failure detection in digital telephony environments by utilizing transformer-based codec inversion
abstract
This study introduces the Codec Transformer Network (CTN) to enhance the reliability of automatic heart failure (HF) detection from coded telephone speech by addressing codec-related challenges in digital telephony. The study specifically addresses the codec mismatch between training and inference in HF detection. CTN is designed to map the mel-spectrogram representations of encoded speech signals back to their original, non-encoded forms, thereby recovering HF-related discriminative information. The effectiveness of CTN is demonstrated in conjunction with three HF detectors, based on Support Vector Machine, Random Forest, and K-Nearest Neighbors classifiers. The results show that CTN effectively retrieves the discriminative information between patients and controls, and performs comparably to or better than a baseline approach, based on multi-condition training.
Saska Tirronen, Farhad Javanmardi, Hilla Pohjalainen, Sudarsana Reddy Kadiri, Mittapalle Kiran Reddy, Pyry Helkkula, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku
Speech Commun.5
2024 Automatic classification of neurological voice disorders using wavelet scattering features
abstract
Neurological voice disorders are caused by problems in the nervous system as it interacts with the larynx. In this paper, we propose to use wavelet scattering transform (WST)-based features in automatic classification of neurological voice disorders. As a part of WST, a speech signal is processed in stages with each stage consisting of three operations–convolution, modulus and averaging–to generate low-variance data representations that preserve discriminability across classes while minimizing differences within a class. The proposed WST-based features were extracted from speech signals of patients suffering from either spasmodic dysphonia (SD) or recurrent laryngeal nerve palsy (RLNP) and from speech signals of healthy speakers of the Saarbruecken voice disorder (SVD) database. Two machine learning algorithms (support vector machine (SVM) and feed forward neural network (NN)) were trained separately using the WST-based features, to perform two binary classification tasks (healthy vs. SD and healthy vs. RLNP) and one multi-class classification task (healthy vs. SD vs. RLNP). The results show that WST-based features outperformed state-of-the-art features in all three tasks. Furthermore, the best overall classification performance was achieved by the NN classifier trained using WST-based features.
Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, Paavo Alku, K. Sreenivasa Rao, Pabitra Mitra
Speech Commun.2
2023 Classification of functional dysphonia using the tunable Q wavelet transform
abstract
Functional dysphonia (FD) refers to an abnormality in voice quality in the absence of an identifiable lesion. In this paper, we propose an approach based on the tunable Q wavelet transform (TQWT) to automatically classify two types of FD (hyperfunctional dysphonia and hypofunctional dysphonia) from a healthy voice using the acoustic voice signal. Using TQWT, voice signals were decomposed into sub-bands and the entropy values extracted from the sub-bands were utilized as features for the studied 3-class classification problem. In addition, the Mel-frequency cepstral coefficient (MFCC) and glottal features were extracted from the acoustic voice signal and the estimated glottal source signal, respectively. A convolutional neural network (CNN) classifier was trained separately for the TQWT, MFCC and glottal features. Experiments were conducted using voice signals of 57 healthy speakers and 113 FD patients (72 with hyperfunctional dysphonia and 41 with hypofunctional dysphonia) taken from the VOICED database. These experiments revealed that the TQWT features yielded an absolute improvement of 5.5% and 4.5% compared to the baseline MFCC features and glottal features, respectively. Furthermore, the highest classification accuracy (67.91%) was obtained using the combination of the TQWT and glottal features, which indicates the complementary nature of these features.
Mittapalle Kiran Reddy, Yagnavajjula Madhu Keerthana, Paavo Alku
Speech Commun.1
2023 Automatic Assessment of Parkinson's Disease Using Speech Representations of Phonation and Articulation
abstract
Speech from people with Parkinson's disease (PD) are likely to be degraded on phonation, articulation, and prosody. Motivated to describe articulation deficits comprehensively, we investigated 1) the universal phonological features that model articulation manner and place, also known as speech attributes, and 2) glottal features capturing phonation characteristics. These were further supplemented by, and compared with, prosodic features using a popular compact feature set and standard MFCC. Temporal characteristics of these features were modeled by convolutional neural networks. Besides the features, we were also interested in the speech tasks for collecting data for automatic PD speech assessment, like sustained vowels, text reading, and spontaneous monologue. For this, we utilized a recently collected Finnish PD corpus (PDSTU) as well as a Spanish database (PC-GITA). The experiments were formulated as regression problems against expert ratings of PD-related symptoms, including ratings of speech intelligibility, voice impairment, overall severity of communication disorder on PDSTU, as well as on the Unified Parkinson's Disease Rating Scale (UPDRS) on PC-GITA. The experimental results show: 1) the speech attribute features can well indicate the severity of pathologies in parkinsonian speech; 2) combining phonation features with articulatory features improves the PD assessment performance, but requires high-quality recordings to be applicable; 3) read speech leads to more accurate automatic ratings than the use of sustained vowels, but not if the amount of speech is limited to correspond to the sustained vowels in duration; and 4) jointly using data from several speech tasks can further improve the automatic PD assessment performance.
Yuanyuan Liu 0002, Mittapalle Kiran Reddy, Nelly Penttilä, Tiina Ihalainen, Paavo Alku, Okko Johannes Räsänen
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Exemplar-Based Sparse Representations for Detection of Parkinson's Disease From Speech
abstract
Parkinson's disease (PD) is a progressive neurological disorder which affects the motor system. The automatic detection of PD improves the diagnosis of the disease, and it can be done in a non-invasive manner from speech. In this paper, we investigate the use of an exemplar-based sparse representation (SR) classification approach for detecting PD from speech. Exemplars are speech feature vectors extracted from the training data. The idea is to formulate the detection task as a problem of finding sparse representations of test speech feature vectors with respect to training speech exemplars. The main advantage of using the SR approach instead of conventional machine learning (ML)-based approaches is that the training step–which is time-consuming and sometimes requires unorganized hyper-parameter tuning–is not needed. Furthermore, SRs are more robust to redundancy and noise in the data. In this work, we study SR classification approaches based on two sparse coding models, namely, l1-regularized least squares ($l_{1}$LS) and non-negative least squares (NNLS). We propose a strategy based on class-specific dictionaries for improving performance of the$l_{1}$LS- and NNLS-based SR classification. To investigate the detection performance, the$l_{1}$LS- and NNLS-based approaches are applied and compared with the traditional PD detection approach based on ML classification algorithms using the PC-GITA PD dataset and an openly available dataset consisting of mobile device voice recordings from healthy and PD patients. The results indicate that the proposed NNLS-based SR classification approach performs better than the traditional ML-based methods in discriminating PD patients from healthy subjects.
Mittapalle Kiran Reddy, Paavo Alku
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Glottal flow characteristics in vowels produced by speakers with heart failure
abstract
Heart failure (HF) is one of the most life-threatening diseases globally. HF is an under-diagnosed condition, and more screening tools are needed to detect it. A few recent studies have suggested that HF also affects the functioning of the speech production mechanism by causing generation of edema in the vocal folds and by impairing the lung function. It has not yet been studied whether these possible effects of HF on the speech production mechanism are large enough to cause acoustically measurable differences to distinguish speech produced in HF from that produced by healthy speakers. Therefore, the goal of the present study was to compare speech production between HF patients and healthy controls by focusing on the excitation signal generated at the level of the vocal folds, the glottal flow. The glottal flow was computed from speech using the quasi-closed phase glottal inverse filtering method and the estimated flow was parameterized with 12 glottal parameters. The sound pressure level (SPL) was measured from speech as an additional parameter. The statistical analyses conducted on the parameters indicated that most of the glottal parameters and SPL were significantly different between the HF patients and healthy controls. The results showed that the HF patients generally produced a more rounded glottal pulse and a lower SPL level compared to the healthy controls, indicating incomplete glottal closure and inappropriate leakage of air through the glottis. The results observed in this preliminary study indicate that glottal features are capable of distinguishing speakers with HF from healthy controls. Therefore, the study suggests that glottal features constitute a potential feature extraction approach which should be taken into account in future large-scale investigations in studying the automatic detection of HF from speech.
Mittapalle Kiran Reddy, Hilla Pohjalainen, Pyry Helkkula, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku
Speech Commun.1
2022 End-to-End Pathological Speech Detection Using Wavelet Scattering Network
abstract
In recent years, developing robust systems for automatic detection of pathological speech has attracted increasing interest among researchers and clinicians. This study proposes an end-to-end approach based on wavelet scattering network (WSN) for detection of pathological speech. In the proposed approach, the WSN (which involves no learning) extracts suitable information from the input raw speech signal and this information is then passed through a multi-layer perceptron (MLP) in order to classify the speech signal as either healthy or pathological. The results show that the proposed approach outperformed a convolutional neural network (CNN) based end-to-end system in distinguishing pathological speech from healthy speech. Furthermore, the proposed system achieved comparable performance with a state-of-the-art traditional system based on hand-crafted features for uncompressed speech, but gave better performance than the traditional system for compressed speech of low bit rates.
Mittapalle Kiran Reddy, Yagnavajjula Madhu Keerthana, Paavo Alku
IEEE Signal Process. Lett.1
2021 The automatic detection of heart failure using speech signals
abstract
Heart failure (HF) is a major global health concern and is increasing in prevalence. It affects the larynx and breathing – thereby the quality of speech. In this article, we propose an approach for the automatic detection of people with HF using the speech signal. The proposed method explores mel-frequency cepstral coefficient (MFCC) features, glottal features, and their combination to distinguish HF from healthy speech. The glottal features were extracted from the voice source signal estimated using glottal inverse filtering. Four machine learning algorithms , namely, support vector machine , Extra Tree , AdaBoost , and feed-forward neural network (FFNN), were trained separately for individual features and their combination. It was observed that the MFCC features yielded higher classification accuracies compared to glottal features. Furthermore, the complementary nature of glottal features was investigated by combining these features with the MFCC features. Our results show that the FFNN classifier trained using a reduced set of glottal + MFCC features achieved the best overall performance in both speaker-dependent and speaker-independent scenarios.
Mittapalle Kiran Reddy, Pyry Helkkula, Yagnavajjula Madhu Keerthana, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku
Comput. Speech Lang.1
2020 Excitation modelling using epoch features for statistical parametric speech synthesis
Mittapalle Kiran Reddy, K. Sreenivasa Rao
Comput. Speech Lang.1
2020 DNN-Based Cross-Lingual Voice Conversion Using Bottleneck Features
Mittapalle Kiran Reddy, K. Sreenivasa Rao
Neural Process. Lett.1
2020 Robust f0 extraction from monophonic signals using adaptive sub-band filtering
Pradeep Rengaswamy, Mittapalle Kiran Reddy, K. Sreenivasa Rao, Pallab Dasgupta
Speech Commun.2
2020 Multilingual and multimode phone recognition system for Indian languages
Kumud Tripathi, Mittapalle Kiran Reddy, K. Sreenivasa Rao
Speech Commun.2
2019 CWT-Based Approach for Epoch Extraction From Telephone Quality Speech
abstract
Epochs are the instants of significant excitation to vocal tract system. Existing methods can extract epochs accurately from clean speech signals. However, identification of epoch locations from band-limited telephonic speech is challenging due to the attenuation of fundamental frequency component and degradation caused by channel effect. This letter proposes an epoch extraction method that can accurately extract epochs from clean as well as telephonic speech signals. In the proposed method, the significant impulse-like discontinuities are extracted directly from the speech signal using continuous wavelet transform. The performance of the proposed method is evaluated using three speakers, namely, SLT, BDL, and JMK from CMU Arctic database. The clean speech is simulated using G.191 software tools to obtain telephonic speech. Experimental results show that the epoch identification rate of proposed method is significantly better than the state-of-the-art methods for the telephone quality speech.
Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, K. Sreenivasa Rao
IEEE Signal Process. Lett.2
2018 Inverse filter based excitation model for HMM-based speech synthesis system
abstract
Even today, the speech generated by hidden Markov model (HMM)‐based speech synthesis system (HTS) still has the buzziness due to the improper modelling of the excitation signal. This study proposes an efficient excitation modelling approach for improving the quality of HTS. In the proposed method, the residual signal obtained from inverse filter is parameterised as excitation features. HMMs are used to model these excitation parameters. During synthesis, the excitation signal is constructed by overlap adding the natural residual segments, and the excitation signal is further modified as per the target source features generated from HMMs. The proposed approach is incorporated in the HTS. Performance evaluation results indicate that the proposed method enhances the quality of synthesis, and is better than the state‐of‐the‐art approaches used for modelling the excitation signal.
Mittapalle Kiran Reddy, K. Sreenivasa Rao
IET Signal Process.1
2017 Robust Pitch Extraction Method for the HMM-Based Speech Synthesis System
abstract
This letter proposes an efficient method for extracting pitch from speech signals for the hidden Markov model (HMM)-based speech synthesis system (HTS). In the proposed method, voicing detection and pitch estimation is performed using the mean signal obtained from continuous wavelet transform coefficients. The proposed pitch extraction method is integrated in the HMM-based speech synthesis system. The Performance of the proposed method is evaluated on CMU Arctic and Keele databases. Both objective and subjective evaluation results show that the quality of speech synthesized with the proposed pitch estimation method is much better compared with HMM-based speech synthesis systems developed using the state-of-the-art pitch extraction methods, namely, robust algorithm for pitch tracking and speech transformation and representation using adaptive interpolation of weighted spectrum employed in the HTS.
Mittapalle Kiran Reddy, K. Sreenivasa Rao
IEEE Signal Process. Lett.1