K. Sreenivasa Rao

dblp:50/5583 · also Krothapalli S. Rao, Krothapalli Sreenivasa Rao, Rao Sreenivasa Krothapalli · DBLP profile ↗
← Back
67ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0001-6112-6887ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 39 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accent identification from emotional speech using classification fusion of multiple deep learning models
Priya Dharshini G, K. Sreenivasa Rao
Neural Comput. Appl.2
2025 Hierarchical speech emotion recognition using the valence-arousal model
Arijul Haque, K. Sreenivasa Rao
Multim. Tools Appl.2
2025 Two-stage pipeline based robust hand gesture recognition from Bharatanatyam dance images
Soumen Paul, Gautam Sagar, Partha Pratim Das 0001, K. Sreenivasa Rao
Multim. Tools Appl.4
2025 Self Supervised Prediction of Genetic Associations in Comorbid Diseases With Masked Autoencoder Using Hypergraph Representations
abstract
Comorbid disease association refers to the simultaneous occurrence of a disease with the coexistence of another primary disease. Due to the complex traits of these co-occurring multi-diseases, it is crucial to know the underlying genetic molecular basis of the prevalent diseases. The inference of common genetic association based on gene co-expression data helps to unveil the pathogenesis of comorbid diseases. There exist a few disease-specific gene co-expression-based analyses to predict the hub genes causing these diseases. However, works lack multi-relational biological data integration. In addition, there still does not exist any unified method to predict the common genetic associations from the co-expression graph across comorbid diseases. Hence, we introduce a generalized and novel approach to predict overlapping genetic associations from disease-specific gene co-expression networks with a self-supervised edge-masking technique catapult with a hypergraph-based pre-embedding learning approach. The advantage of hypergraph learning is that it induces higher-order rich biological information of candidate genes. In addition, we use the self-supervised-based edge masking strategy to attain model training over only a few numbers of edge labels. Our proposed approach outperforms the six baseline models for our case-study datasets and also predicts novel genetic associations across comorbid disease pairs.
Saikat Biswas, Vibhanshu Ranjan, Pabitra Mitra, K. Sreenivasa Rao
IEEE Trans. Comput. Biol. Bioinform.4
2024 Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
abstract
Deepfakes are a major security risk for biometric authentication. This technology creates realistic fake videos that can impersonate real people, fooling systems that rely on facial features and voice patterns for identification. Existing multimodal deepfake detectors rely on conventional fusion methods, such as majority rule and ensemble voting, which often struggle to adapt to changing data characteristics and complex patterns. In this paper, we introduce the Straight-through Gumbel-Softmax (STGS) framework, offering a comprehensive approach to search multimodal fusion model architectures. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Initially, crucial features were efficiently identified from backbone networks, whereas within the cell structure, a weighted fusion operation integrated information from various sources. An architecture that maximizes the classification performance is derived by varying parameters such as temperature and sampling time. The experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrated an impressive AUC value 94.4% achieved with minimal model parameters. https://github.com/Aravinda27/STGS-BMNAS
Aravinda Reddy P. N., Ramachandra Raghavendra, K. Sreenivasa Rao, Pabitra Mitra, Vinod Rathod
IJCB3
2024 NeuralMultiling: A Novel Neural Architecture Search for Smartphone Based Multilingual Speaker Verification
Aravinda Reddy P. N., Ramachandra Raghavendra, K. Sreenivasa Rao, Pabitra Mitra
ICPR (14)3
2024 Hierarchical emotion recognition from speech using source, power spectral and prosodic features
Arijul Haque, K. Sreenivasa Rao
Multim. Tools Appl.2
2024 Automatic classification of neurological voice disorders using wavelet scattering features
abstract
Neurological voice disorders are caused by problems in the nervous system as it interacts with the larynx. In this paper, we propose to use wavelet scattering transform (WST)-based features in automatic classification of neurological voice disorders. As a part of WST, a speech signal is processed in stages with each stage consisting of three operations–convolution, modulus and averaging–to generate low-variance data representations that preserve discriminability across classes while minimizing differences within a class. The proposed WST-based features were extracted from speech signals of patients suffering from either spasmodic dysphonia (SD) or recurrent laryngeal nerve palsy (RLNP) and from speech signals of healthy speakers of the Saarbruecken voice disorder (SVD) database. Two machine learning algorithms (support vector machine (SVM) and feed forward neural network (NN)) were trained separately using the WST-based features, to perform two binary classification tasks (healthy vs. SD and healthy vs. RLNP) and one multi-class classification task (healthy vs. SD vs. RLNP). The results show that WST-based features outperformed state-of-the-art features in all three tasks. Furthermore, the best overall classification performance was achieved by the NN classifier trained using WST-based features.
Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, Paavo Alku, K. Sreenivasa Rao, Pabitra Mitra
Speech Commun.4
2023 Self-Paced Pattern Augmentation for Spoken Term Detection in Zero-Resource
P. Sudhakar, K. Sreenivasa Rao, Pabitra Mitra
INTERSPEECH2
2023 Accent classification from an emotional speech in clean and noisy environments
Priya Dharshini G, K. Sreenivasa Rao
Multim. Tools Appl.2
2022 A novel approach to unsupervised pattern discovery in speech using Convolutional Neural Network
Kishore Kumar R., K. Sreenivasa Rao
Comput. Speech Lang.2
2021 Knowledge Distillation for Singing Voice Detection
abstract
Singing Voice Detection (SVD) has been an active area of research in music information retrieval (MIR). Currently, two deep neural network-based methods, one based on CNN and the other on RNN, exist in literature that learn optimized features for the voice detection (VD) task and achieve state-of-the-art performance on common datasets. Both these models have a huge number of parameters (1.4M for CNN and 65.7K for RNN) and hence not suitable for deployment on devices like smartphones or embedded sensors with limited capacity in terms of memory and computation power. The most popular method to address this issue is known as knowledge distillation in deep learning literature (in addition to model compression) where a large pre-trained network known as the teacher is used to train a smaller student network. Given the wide applications of SVD in music information retrieval, to the best of our knowledge, model compression for practical deployment has not yet been explored. In this paper, efforts have been made to investigate this issue using both conventional as well as ensemble knowledge distillation techniques.
Soumava Paul, Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001
Interspeech3
2021 Robust vowel region detection method for multimode speech
Kumud Tripathi, K. Sreenivasa Rao
Multim. Tools Appl.2
2021 Approaches for Multilingual Phone Recognition in Code-switched and Non-code-switched Scenarios Using Indian Languages
abstract
In this study, we evaluate and compare two different approaches for multilingual phone recognition in code-switched and non-code-switched scenarios. First approach is a front-end Language Identification (LID)-switched to a monolingual phone recognizer (LID-Mono), trained individually on each of the languages present in multilingual dataset. In the second approach, a common multilingual phone-set derived from the International Phonetic Alphabet (IPA) transcription of the multilingual dataset is used to develop a Multilingual Phone Recognition System (Multi-PRS). The bilingual code-switching experiments are conducted using Kannada and Urdu languages. In the first approach, LID is performed using the state-of-the-art i-vectors. Both monolingual and multilingual phone recognition systems are trained using Deep Neural Networks. The performance of LID-Mono and Multi-PRS approaches are compared and analysed in detail. It is found that the performance of Multi-PRS approach is superior compared to more conventional LID-Mono approach in both code-switched and non-code-switched scenarios. For code-switched speech, the effect of length of segments (that are used to perform LID) on the performance of LID-Mono system is studied by varying the window size from 500 ms to 5.0 s, and full utterance. The LID-Mono approach heavily depends on the accuracy of the LID system and the LID errors cannot be recovered. But, the Multi-PRS system by virtue of not having to do a front-end LID switching and designed based on the common multilingual phone-set derived from several languages, is not constrained by the accuracy of the LID system, and hence performs effectively on code-switched and non-code-switched speech, offering low Phone Error Rates than the LID-Mono system.
K. Manjunath, Srinivasa Raghavan K. M., K. Sreenivasa Rao, Dinesh Babu Jayagopi, V. Ramasubramanian 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Relation Prediction of Co-Morbid Diseases Using Knowledge Graph Completion
abstract
Co-morbid disease condition refers to the simultaneous presence of one or more diseases along with the primary disease. A patient suffering from co-morbid diseases possess more mortality risk than with a disease alone. So, it is necessary to predict co-morbid disease pairs. In past years, though several methods have been proposed by researchers for predicting the co-morbid diseases, not much work is done in prediction using knowledge graph embedding using tensor factorization. Moreover, the complex-valued vector-based tensor factorization is not being used in any knowledge graph with biological and biomedical entities. We propose a tensor factorization based approach on biological knowledge graphs. Our method introduces the concept of complex-valued embedding in knowledge graphs with biological entities. Here, we build a knowledge graph with disease-gene associations and their corresponding background information. To predict the association between prevalent diseases, we use ComplEx embedding based tensor decomposition method. Besides, we obtain new prevalent disease pairs using the MCL algorithm in a disease-gene-gene network and check their corresponding inter-relations using edge prediction task.
Saikat Biswas, Pabitra Mitra, K. Sreenivasa Rao
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Glottal Closure Instants Detection from EGG Signal by Classification Approach
Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001
INTERSPEECH2
2020 Excitation modelling using epoch features for statistical parametric speech synthesis
Mittapalle Kiran Reddy, K. Sreenivasa Rao
Comput. Speech Lang.2
2020 DNN-Based Cross-Lingual Voice Conversion Using Bottleneck Features
Mittapalle Kiran Reddy, K. Sreenivasa Rao
Neural Process. Lett.2
2020 Robust f0 extraction from monophonic signals using adaptive sub-band filtering
Pradeep Rengaswamy, Mittapalle Kiran Reddy, K. Sreenivasa Rao, Pallab Dasgupta
Speech Commun.3
2020 Multilingual and multimode phone recognition system for Indian languages
Kumud Tripathi, Mittapalle Kiran Reddy, K. Sreenivasa Rao
Speech Commun.3
2020 Children's Story Classification in Indian Languages Using Linguistic and Keyword-based Features
abstract
The primary objective of this work is to classify Hindi and Telugu stories into three genres: fable, folk-tale, and legend . In this work, we are proposing a framework for story classification (SC) using keyword and part-of-speech (POS) features. For improving the performance of SC system, feature reduction techniques and combinations of various POS tags are explored. Further, we investigated the performance of SC by dividing the story into parts depending on its semantic structure. In this work, stories are (i) manually divided into parts based on their semantics as introduction, main, and climax ; and (ii) automatically divided into equal parts based on number of sentences in a story as initial, middle, and end . We have also examined sentence increment model, which aims at determining an optimum number of sentences required to identify story genre by incremental selection of sentences in a story. Experiments are conducted on Hindi and Telugu story corpora consisting of 300 and 150 short stories, respectively. The performance of SC system is evaluated using different combinations of keyword and POS-based features, with three well-established machine learning classifiers: (i) Naive Bayes (NB), (ii) k-Nearest Neighbour (KNN), and (iii) Support Vector Machine (SVM). Performance of the classifier is evaluated using 10-fold cross-validation and effectiveness of classifier is measured using precision, recall, and F-measure. From the classification results, it is observed that adding linguistic information boosts the performance of story classification. In view of the structure of the story, main, and initial parts of the story have shown comparatively better performance. The results from the sentence incremental model have indicated that the first nine and seven sentences in Hindi and Telugu stories, respectively, are sufficient for better classification of stories. In most of the studies, SVM models outperformed the other models in classification accuracy.
Harikrishna D. M., K. Sreenivasa Rao
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 BOXREC: Recommending a Box of Preferred Outfits in Online Shopping
abstract
Fashionable outfits are generally created by expert fashionistas, who use their creativity and in-depth understanding of fashion to make attractive outfits. Over the past few years, automation of outfit composition has gained much attention from the research community. Most of the existing outfit recommendation systems focus on pairwise item compatibility prediction (using visual and text features) to score an outfit combination having several items, followed by recommendation of top-n outfits or a capsule wardrobe having a collection of outfits based on user’s fashion taste. However, none of these consider a user’s preference of price range for individual clothing types or an overall shopping budget for a set of items. In this article, we propose a box recommendation framework—BOXREC—which at first collects user preferences across different item types (namely, top-wear, bottom-wear, and foot-wear) including price range of each type and a maximum shopping budget for a particular shopping session. It then generates a set of preferred outfits by retrieving all types of preferred items from the database (according to user specified preferences including price ranges), creates all possible combinations of three preferred items (belonging to distinct item types), and verifies each combination using an outfit scoring framework—BOXREC-OSF. Finally, it provides a box full of fashion items, such that different combinations of the items maximize the number of outfits suitable for an occasion while satisfying maximum shopping budget. We create an extensively annotated dataset of male fashion items across various types and categories (each having associated price) and a manually annotated positive and negative formal as well as casual outfit dataset. We consider a set of recently published pairwise compatibility prediction methods as competitors of BOXREC-OSF. Empirical results show superior performance of BOXREC-OSF over the baseline methods. We found encouraging results by performing both quantitative and qualitative analysis of the recommendations produced by BOXREC. Finally, based on user feedback corresponding to the recommendations given by BOXREC, we show that disliked or unpopular items can be a part of attractive outfits.
Debopriyo Banerjee, K. Sreenivasa Rao, Shamik Sural, Niloy Ganguly
ACM Trans. Intell. Syst. Technol.2
2019 Glottal Closure Instants Detection from Speech Signal by Deep Features Extracted from Raw Speech and Linear Prediction Residual
Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001
INTERSPEECH2
2019 CWT-Based Approach for Epoch Extraction From Telephone Quality Speech
abstract
Epochs are the instants of significant excitation to vocal tract system. Existing methods can extract epochs accurately from clean speech signals. However, identification of epoch locations from band-limited telephonic speech is challenging due to the attenuation of fundamental frequency component and degradation caused by channel effect. This letter proposes an epoch extraction method that can accurately extract epochs from clean as well as telephonic speech signals. In the proposed method, the significant impulse-like discontinuities are extracted directly from the speech signal using continuous wavelet transform. The performance of the proposed method is evaluated using three speakers, namely, SLT, BDL, and JMK from CMU Arctic database. The clean speech is simulated using G.191 software tools to obtain telephonic speech. Experimental results show that the epoch identification rate of proposed method is significantly better than the state-of-the-art methods for the telephone quality speech.
Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, K. Sreenivasa Rao
IEEE Signal Process. Lett.3
2018 Robust Detection of Glottal Activity Using Unwrapped Phase Electroglottographic Signal
abstract
Glottal Activity is defined by the process of exciting the vocal tract system by the vibration of the vocal folds during speech processing. Detection of glottal activity refers to identifying the glottal instants present within a glottal cycle. GCI and GOI are the two important glottal instants of a glottal cycle. The DEGG signal provides significant peaks at those instants for normal voicing, but they are prone to error in detection for low strength EGG signal. In this paper, we mainly focus on the segments of EGG signal where the strength of the EGG signal is very poor and irregular in periodicity. The robust detection of the glottal instants in those segments of EGG signal will enhance the overall accuracy of the detection of glottal instants. In the proposed method, the unwrapped phase of the EGG signal has been used for the detection of the glottal instants. The phase of the signal has uniform characteristics throughout the signal which helps to detect the glottal instants robustly.
Tanumay Mandal, K. Sreenivasa Rao
ICASSP2
2018 Modifying LSTM Posteriors with Manner of Articulation Knowledge to Improve Speech Recognition Performance
abstract
The variant of recurrent neural networks (RNN) such as long short-term memory (LSTM) is successful in sequence modelling such as automatic speech recognition (ASR) framework. However the decoded sequence is prune to have false substitutions, insertions and deletions. We exploit the spectral flatness measure (SFM) computed on the magnitude linear prediction (LP) spectrum to detect two broad manners of articulation namely sonorants and obstruents. In this paper, we modify the posteriors generated at the output layer of LSTM according to the manner of articulation detection. The modified posteriors are given to the conventional decoding graph to minimize the false substitutions and insertions. The proposed method decreased the phone error rate (PER) by nearly 0.7 % and 0.3 % when evaluated on core TIMIT test corpus as compared to the conventional decoding involved in the deep neural networks (DNN) and the state of the art LSTM respectively.
Pradeep Rengaswamy, K. Sreenivasa Rao
ICMLA2
2018 Harmonic-Percussive Source Separation of Polyphonic Music by Suppressing Impulsive Noise Events
Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001
INTERSPEECH2
2018 Classification of Disorders in Vocal Folds Using Electroglottographic Signal
Tanumay Mandal, K. Sreenivasa Rao, Sanjay Kumar Gupta
INTERSPEECH2
2018 Indian Languages ASR: A Multilingual Phone Recognition Framework with IPA Based Common Phone-set, Predicted Articulatory Features and Feature fusion
K. Manjunath, K. Sreenivasa Rao, Dinesh Babu Jayagopi, V. Ramasubramanian 0001
INTERSPEECH2
2018 Analysis of sparse representation based feature on speech mode classification
Kumud Tripathi, K. Sreenivasa Rao
INTERSPEECH2
2018 One for the Road: Recommending Male Street Attire
Debopriyo Banerjee, Niloy Ganguly, Shamik Sural, K. Sreenivasa Rao
PAKDD (3)4
2018 Inverse filter based excitation model for HMM-based speech synthesis system
abstract
Even today, the speech generated by hidden Markov model (HMM)‐based speech synthesis system (HTS) still has the buzziness due to the improper modelling of the excitation signal. This study proposes an efficient excitation modelling approach for improving the quality of HTS. In the proposed method, the residual signal obtained from inverse filter is parameterised as excitation features. HMMs are used to model these excitation parameters. During synthesis, the excitation signal is constructed by overlap adding the natural residual segments, and the excitation signal is further modified as per the target source features generated from HMMs. The proposed approach is incorporated in the HTS. Performance evaluation results indicate that the proposed method enhances the quality of synthesis, and is better than the state‐of‐the‐art approaches used for modelling the excitation signal.
Mittapalle Kiran Reddy, K. Sreenivasa Rao
IET Signal Process.2
2018 A robust unsupervised pattern discovery and clustering of speech signals
Kishore Kumar R., Lokendra Birla, K. Sreenivasa Rao
Pattern Recognit. Lett.3
2018 Epoch detection from emotional speech signal using zero time windowing
Jainath Yadav, Md. S. Fahad, K. Sreenivasa Rao
Speech Commun.3
2017 Accurate Synchronization of Speech and EGG Signal Using Phase Information
S. B. Sunil Kumar, K. Sreenivasa Rao, Tanumay Mandal
INTERSPEECH2
2017 Implicit processing of LP residual for language identification
Dipanjan Nandi, Debadatta Pati, K. Sreenivasa Rao
Comput. Speech Lang.3
2017 Parametric representation of excitation source information for language identification
Dipanjan Nandi, Debadatta Pati, K. Sreenivasa Rao
Comput. Speech Lang.3
2017 Generation of creaky voice for improving the quality of HMM-based speech synthesis
N. P. Narendra, K. Sreenivasa Rao
Comput. Speech Lang.2
2017 Robust Pitch Extraction Method for the HMM-Based Speech Synthesis System
abstract
This letter proposes an efficient method for extracting pitch from speech signals for the hidden Markov model (HMM)-based speech synthesis system (HTS). In the proposed method, voicing detection and pitch estimation is performed using the mean signal obtained from continuous wavelet transform coefficients. The proposed pitch extraction method is integrated in the HMM-based speech synthesis system. The Performance of the proposed method is evaluated on CMU Arctic and Keele databases. Both objective and subjective evaluation results show that the quality of speech synthesized with the proposed pitch estimation method is much better compared with HMM-based speech synthesis systems developed using the state-of-the-art pitch extraction methods, namely, robust algorithm for pitch tracking and speech transformation and representation using adaptive interpolation of weighted spectrum employed in the HTS.
Mittapalle Kiran Reddy, K. Sreenivasa Rao
IEEE Signal Process. Lett.2
2016 A deterministic plus noise model of excitation signal using principal component analysis for parametric speech synthesis
abstract
This paper proposes a new approach of modeling the excitation signal as deterministic and noise components. Initially, a study on characteristics of excitation or residual signal around glottal closure instant (GCI) is performed using principal component analysis (PCA). Based on the study, the segment of residual signal around GCI is considered as the deterministic component and the remaining part of the residual signal is considered as the noise component. The deterministic component is parameterized using PCA coefficients, and the noise component can be represented in terms of spectral and amplitude envelopes. The proposed excitation modeling approach is incorporated in the HMM-based speech synthesis system. Subjective evaluation results show a significant improvement in the quality of speech synthesized by the proposed method, compared to three existing methods.
N. P. Narendra, K. Sreenivasa Rao
ICASSP2
2016 Predominant melody extraction from vocal polyphonic music signal by combined spectro-temporal method
abstract
A combined spectro-temporal based method is proposed to derive the predominant melody from vocal polyphonic music signals. In the proposed method, vocal (voiced) and non-vocal (unvoiced) segments are determined by strength of excitation. The vocal segments are further divided into voiced notes by detecting their onsets using transition cues present in spectral domain. The melody contour present in each of the voiced note segments is obtained by using an adaptive zero frequency filtering (ZFF) in time domain. The process of melody extraction is provided in more detail and the initial results showed the potential use of the proposed method for vocal melody extraction.
Gurunath Reddy M, K. Sreenivasa Rao
ICASSP2
2016 Enhanced Harmonic Content and Vocal Note Based Predominant Melody Extraction from Vocal Polyphonic Music Signals
Gurunath Reddy M, K. Sreenivasa Rao
INTERSPEECH2
2016 A Robust Non-Parametric and Filtering Based Approach for Glottal Closure Instant Detection
Pradeep Rengaswamy, Gurunath Reddy M, K. Sreenivasa Rao, Pallab Dasgupta
INTERSPEECH3
2016 Prosody modeling for syllable based text-to-speech synthesis using feedforward neural networks
V. Ramu Reddy, K. Sreenivasa Rao
Neurocomputing2
2016 Voice/non-voice detection using phase of zero frequency filtered speech signal
S. B. Sunil Kumar, K. Sreenivasa Rao
Speech Commun.2
2016 Time-domain deterministic plus noise model based hybrid source modeling for statistical parametric speech synthesis
N. P. Narendra, K. Sreenivasa Rao
Speech Commun.2
2015 Automatic detection of creaky voice using epoch parameters
N. P. Narendra, K. Sreenivasa Rao
INTERSPEECH2
2014 A novel boosting algorithm for improved i-vector based speaker verification in noisy environments
Sourjya Sarkar, K. Sreenivasa Rao
INTERSPEECH2
2013 High quality text-to-speech synthesis system with efficient duration models developed using coding schemes based on vowel production characteristics
abstract
This paper explores encoding schemes based on production characteristics of vowels. The performance of coding schemes is analyzed for accurate prediction of durations of syllables using neural network models. Linguistic and production constraints represented by positional, contextual, phonological and articulatory (PCPA) features are used for predicting the durations of syllables. These features are coded with distinct numerical values before feeding to the neural network for building models. The evaluation of coding schemes is carried out by means of objective and subjective measures. The quality of text-to-speech synthesis system is observed to be better by incorporating the duration model with vowels coded based on lip roundness.
V. Ramu Reddy, K. Sreenivasa Rao
ISDA2
2013 Two-stage intonation modeling using feedforward neural networks for syllable based text-to-speech synthesis
V. Ramu Reddy, K. Sreenivasa Rao
Comput. Speech Lang.2
2013 Non-uniform time scale modification using instants of significant excitation and vowel onset points
K. Sreenivasa Rao, Anil Kumar Vuppala
Speech Commun.1
2013 Detection of Vowel Offset Point From Speech Signal
abstract
Vowel regions play important role in various speech tasks, such as speech segmentation, speaker-verification, prosody modification and emotion conversion. The instants at which the onset and offset of vowel take place in the speech signal are known as vowel onset point and vowel offset point, respectively. Vowel regions start with the vowel onset point and end with the vowel offset point. In this letter, we have proposed two methods for determining the vowel offset points from the speech signal. The first method explores the combination of evidences from excitation source, spectral peaks and modulation spectrum for determining the vowel offset point. In the second method, spectral energy within glottal closure region is used for determining the vowel offset point. The proposed vowel offset point detection methods are evaluated on TIMIT database under clean and noisy environments.
Jainath Yadav, K. Sreenivasa Rao
IEEE Signal Process. Lett.2
2012 Vowel Onset Point Detection for Low Bit Rate Coded Speech
abstract
In this paper, we propose a method for detecting the vowel onset points (VOPs) for low bit rate coded speech. VOP is the instant at which the onset of the vowel takes place in the speech signal. VOP plays an important role for the applications, such as consonant-vowel (CV) unit recognition and speech rate modification. The proposed VOP detection method is based on the spectral energy present in the glottal closure region of the speech signal. Speech coders considered to carry out this study are Global System for Mobile Communications (GSM) full rate, code-excited linear prediction (CELP), and mixed-excitation linear prediction (MELP). TIMIT database and CV units collected from the broadcast news corpus are used for evaluation. Performance of the proposed method is compared with existing methods, which uses the combination of evidence from the excitation source, spectral peaks energy, and modulation spectrum. The proposed VOP detection method has shown significant improvement in the performance compared to the existing method under clean as well as coded cases. The effectiveness of the proposed VOP detection method is analyzed in CV recognition by using VOP as an anchor point.
Anil Kumar Vuppala, Jainath Yadav, Saswat Chakrabarti, K. Sreenivasa Rao
IEEE Trans. Speech Audio Process.4
2011 Recognition of emotions from video using neural network models
K. Sreenivasa Rao, V. K. Saroj, Sudhamay Maity, Shashidhar G. Koolagudi
Expert Syst. Appl.1
2010 Voice conversion by mapping the speaker-specific features using pitch synchronous approach
K. Sreenivasa Rao
Comput. Speech Lang.1
2009 Intonation modeling for Indian languages
K. Sreenivasa Rao, Bayya Yegnanarayana
Comput. Speech Lang.1
2009 Duration modification using glottal closure instants and vowel onset points
K. Sreenivasa Rao, Bayya Yegnanarayana
Speech Commun.1
2007 Modeling durations of syllables using neural networks
K. Sreenivasa Rao, Bayya Yegnanarayana
Comput. Speech Lang.1
2007 Determination of Instants of Significant Excitation in Speech Using Hilbert Envelope and Group Delay Function
abstract
This letter proposes a time-effective method for determining the instants of significant excitation in speech signals. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. The proposed method consists of two phases: the first phase determines the approximate epoch locations using the Hilbert envelope of the linear prediction residual of the speech signal. The second phase determines the accurate locations of the instants of significant excitation by computing the group delay around the approximate epoch locations derived from the first phase. The accuracy in determining the instants of significant excitation and the time complexity of the proposed method is compared with the group delay based approach.
K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana
IEEE Signal Process. Lett.1
2006 Prosody modification using instants of significant excitation
abstract
Prosody modification involves changing the pitch and duration of speech without affecting the message and naturalness. This paper proposes a method for prosody (pitch and duration) modification using the instants of significant excitation of the vocal tract system during the production of speech. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. Instants of significant excitation are computed from the linear prediction (LP) residual of speech signals by using the property of average group-delay of minimum phase signals. The modification of pitch and duration is achieved by manipulating the LP residual with the help of the knowledge of the instants of significant excitation. The modified residual is used to excite the time-varying filter, whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is good and is without any significant distortion. The proposed method is evaluated using waveforms, spectrograms, and listening tests. The performance of the method is compared with linear prediction pitch synchronous overlap and add (LP-PSOLA) method, which is another method for prosody manipulation based on the modification of the LP residual. The original and the synthesized speech signals obtained by the proposed method and by the LP-PSOLA method are available for listening at http://speech.cs.iitm.ernet.in/Main/result/prosody.html.
K. Sreenivasa Rao, Bayya Yegnanarayana
IEEE Trans. Speech Audio Process.1
2004 Modeling syllable duration in Indian languages using neural networks
abstract
We propose a neural network model for predicting the syllable duration in Indian languages. A four layer feedforward neural network trained with a backpropagation algorithm is used for modeling the syllable duration. Analysis is performed on broadcast news data in Hindi, Telugu and Tamil in order to predict the duration of syllables in these languages using a neural network model. The input to the neural network consists of a set of phonological, positional and contextual features extracted from the text. About 88% of the syllable durations are predicted within 25% of the actual duration. The relative importance of the positional and contextual features are examined separately.
K. Sreenivasa Rao, Bayya Yegnanarayana
ICASSP (5)1
2004 Two-Stage Duration Model for Indian Languages Using Neural Networks
K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana
ICONIP1
2004 Intonation modeling for indian languages
abstract
In this paper we propose models for predicting the intonation for the sequence of syllables present in the utterance.The term intonation refers to the temporal changes of the fundamental frequency ðF 0 Þ.Neural networks are used to capture the implicit intonation knowledge in the sequence of syllables of an utterance.We focus on the development of intonation models for predicting the sequence of fundamental frequency values for a given sequence of syllables.Labeled broadcast news data in the languages Hindi, Telugu and Tamil is used to develop neural network models in order to predict the F 0 of syllables in these languages.The input to the neural network consists of a feature vector representing the positional, contextual and phonological constraints.The interaction between duration and intonation constraints can be exploited for improving the accuracy further.From the studies we find that 88% of the F 0 values (pitch) of the syllables could be predicted from the models within 15% of the actual F 0 .The performance of the intonation models is evaluated using objective measures such as average prediction error ðlÞ, standard deviation ðrÞ and correlation coefficient ðcÞ.The prediction accuracy of the intonation models is further evaluated using listening tests.The prediction performance of the proposed intonation models using neural networks is compared with Classification and Regression Tree (CART) models.
K. Sreenivasa Rao, Bayya Yegnanarayana
INTERSPEECH1
2003 Prosodic manipulation using instants of significant excitation
abstract
The paper proposes a technique for prosodic (pitch and duration) manipulation using instants of significant excitation. Instants of significant excitation correspond to the instants of glottal closure (epochs) in voiced speech and to some random excitations like burst onset in the case of nonvoiced speech. Instants of significant excitation are computed from the average group delay of minimum phase signals. The manipulation of pitch and duration is achieved by modifying the linear prediction (LP) residual with the help of instants of significant excitation as pitch markers. The modified residual is used to excite the time-varying filter whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is found to be natural, and is without any distortion. The original and corresponding synthesized speech signals from the proposed approach are available at http://speech.cs.iitm.ernet.in/Main/Results/Prosody.html.
K. Sreenivasa Rao, Bayya Yegnanarayana
ICASSP (1)1
2003 Prosodic manipulation using instants of significant excitation
abstract
This paper proposes a technique for prosodic (pitch and duration) manipulation using instants of significant excitation. Instants of significant excitation correspond to the instants of glottal closure (epochs) in voiced speech and to some random excitations like burst onset in the case of nonvoiced speech. Instants of significant excitation are computed from the average group delay of minimum phase signals. The manipulation of pitch and duration is achieved by modifying the linear prediction (LP) residual with the help of instants of significant excitation as pitch markers. The modified residual is used to excite the time-varying filter whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is found to be natural, and is without any distortion. The original and corresponding synthesized speech signals from the proposed approach are available for listening at http://speech.cs.iitm.ernet.in/Main/Results/Prosody.html.
K. Sreenivasa Rao, Bayya Yegnanarayana
ICME1
2003 Combining evidence from multiple modular networks for recognition of consonant-vowel units of speech
abstract
In this paper, we present a method to combine evidence from multiple classifiers to recognize a large number of subword units of speech using small size training data sets. Grouping criteria based on phonetic description are considered, to build multiple modular networks for recognition of the large number of units. Nonlinear compression of feature vectors is carried out to obtain reduced dimensional patterns, and multiple classifiers are trained separately using the uncompressed feature vectors and compressed feature vectors. Evidence from multiple classifiers at different stages in the recognition system is combined using the sum rule. Effectiveness of the proposed method is demonstrated for recognition of isolated utterances of 145 consonant-vowel units of speech.
Suryakanth V. Gangashetty, K. Sreenivasa Rao, A. Nayeemulla Khan, Chellu Chandra Sekhar, Bayya Yegnanarayana
IJCNN2
2002 Speech enhancement using excitation source information
abstract
This paper proposes an approach for processing speech from multiple microphones to enhance speech degraded by noise and reverberation. The approach is based on exploiting the features of the excitation source in speech production. In particular, the characteristics of voiced speech can be used to derive a coherently added signal from the linear prediction (LP) residuals of the degraded speech data from different microphones. A weight function is derived from the coherently added signal. For coherent addition the time-delay between a pair of microphones is estimated using the knowledge of the source information present in the LP residual. The enhanced speech is generated by exciting the time varying all-pole filter with the weighted LP residual.
Bayya Yegnanarayana, S. R. Mahadeva Prasanna, K. Sreenivasa Rao
ICASSP3