VLDB 2026 Research / reviewers in the wild / expert
K. Sreenivasa Rao
dblp:50/5583 · also Krothapalli S. Rao, Krothapalli Sreenivasa Rao, Rao Sreenivasa Krothapalli
· DBLP profile ↗
67ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0001-6112-6887ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 39 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accent identification from emotional speech using classification fusion of multiple deep learning models
Priya Dharshini G, K. Sreenivasa Rao |
Neural Comput. Appl. | 2 |
| 2025 | Hierarchical speech emotion recognition using the valence-arousal model
Arijul Haque, K. Sreenivasa Rao |
Multim. Tools Appl. | 2 |
| 2025 | Two-stage pipeline based robust hand gesture recognition from Bharatanatyam dance images
Soumen Paul, Gautam Sagar, Partha Pratim Das 0001, K. Sreenivasa Rao |
Multim. Tools Appl. | 4 |
| 2025 | Self Supervised Prediction of Genetic Associations in Comorbid Diseases With Masked Autoencoder Using Hypergraph RepresentationsabstractComorbid disease association refers to the simultaneous occurrence of a disease with the coexistence of another primary disease. Due to the complex traits of these co-occurring multi-diseases, it is crucial to know the underlying genetic molecular basis of the prevalent diseases. The inference of common genetic association based on gene co-expression data helps to unveil the pathogenesis of comorbid diseases. There exist a few disease-specific gene co-expression-based analyses to predict the hub genes causing these diseases. However, works lack multi-relational biological data integration. In addition, there still does not exist any unified method to predict the common genetic associations from the co-expression graph across comorbid diseases. Hence, we introduce a generalized and novel approach to predict overlapping genetic associations from disease-specific gene co-expression networks with a self-supervised edge-masking technique catapult with a hypergraph-based pre-embedding learning approach. The advantage of hypergraph learning is that it induces higher-order rich biological information of candidate genes. In addition, we use the self-supervised-based edge masking strategy to attain model training over only a few numbers of edge labels. Our proposed approach outperforms the six baseline models for our case-study datasets and also predicts novel genetic associations across comorbid disease pairs. Saikat Biswas, Vibhanshu Ranjan, Pabitra Mitra, K. Sreenivasa Rao |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake DetectionabstractDeepfakes are a major security risk for biometric authentication. This technology creates realistic fake videos that can impersonate real people, fooling systems that rely on facial features and voice patterns for identification. Existing multimodal deepfake detectors rely on conventional fusion methods, such as majority rule and ensemble voting, which often struggle to adapt to changing data characteristics and complex patterns. In this paper, we introduce the Straight-through Gumbel-Softmax (STGS) framework, offering a comprehensive approach to search multimodal fusion model architectures. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Initially, crucial features were efficiently identified from backbone networks, whereas within the cell structure, a weighted fusion operation integrated information from various sources. An architecture that maximizes the classification performance is derived by varying parameters such as temperature and sampling time. The experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrated an impressive AUC value 94.4% achieved with minimal model parameters. https://github.com/Aravinda27/STGS-BMNAS Aravinda Reddy P. N., Ramachandra Raghavendra, K. Sreenivasa Rao, Pabitra Mitra, Vinod Rathod |
IJCB | 3 |
| 2024 | NeuralMultiling: A Novel Neural Architecture Search for Smartphone Based Multilingual Speaker Verification
Aravinda Reddy P. N., Ramachandra Raghavendra, K. Sreenivasa Rao, Pabitra Mitra |
ICPR (14) | 3 |
| 2024 | Hierarchical emotion recognition from speech using source, power spectral and prosodic features
Arijul Haque, K. Sreenivasa Rao |
Multim. Tools Appl. | 2 |
| 2024 | Automatic classification of neurological voice disorders using wavelet scattering featuresabstractNeurological voice disorders are caused by problems in the nervous system as it interacts with the larynx. In this paper, we propose to use wavelet scattering transform (WST)-based features in automatic classification of neurological voice disorders. As a part of WST, a speech signal is processed in stages with each stage consisting of three operations–convolution, modulus and averaging–to generate low-variance data representations that preserve discriminability across classes while minimizing differences within a class. The proposed WST-based features were extracted from speech signals of patients suffering from either spasmodic dysphonia (SD) or recurrent laryngeal nerve palsy (RLNP) and from speech signals of healthy speakers of the Saarbruecken voice disorder (SVD) database. Two machine learning algorithms (support vector machine (SVM) and feed forward neural network (NN)) were trained separately using the WST-based features, to perform two binary classification tasks (healthy vs. SD and healthy vs. RLNP) and one multi-class classification task (healthy vs. SD vs. RLNP). The results show that WST-based features outperformed state-of-the-art features in all three tasks. Furthermore, the best overall classification performance was achieved by the NN classifier trained using WST-based features. Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, Paavo Alku, K. Sreenivasa Rao, Pabitra Mitra |
Speech Commun. | 4 |
| 2023 | Self-Paced Pattern Augmentation for Spoken Term Detection in Zero-Resource
P. Sudhakar, K. Sreenivasa Rao, Pabitra Mitra |
INTERSPEECH | 2 |
| 2023 | Accent classification from an emotional speech in clean and noisy environments
Priya Dharshini G, K. Sreenivasa Rao |
Multim. Tools Appl. | 2 |
| 2022 | A novel approach to unsupervised pattern discovery in speech using Convolutional Neural Network
Kishore Kumar R., K. Sreenivasa Rao |
Comput. Speech Lang. | 2 |
| 2021 | Knowledge Distillation for Singing Voice DetectionabstractSinging Voice Detection (SVD) has been an active area of research in music information retrieval (MIR). Currently, two deep neural network-based methods, one based on CNN and the other on RNN, exist in literature that learn optimized features for the voice detection (VD) task and achieve state-of-the-art performance on common datasets. Both these models have a huge number of parameters (1.4M for CNN and 65.7K for RNN) and hence not suitable for deployment on devices like smartphones or embedded sensors with limited capacity in terms of memory and computation power. The most popular method to address this issue is known as knowledge distillation in deep learning literature (in addition to model compression) where a large pre-trained network known as the teacher is used to train a smaller student network. Given the wide applications of SVD in music information retrieval, to the best of our knowledge, model compression for practical deployment has not yet been explored. In this paper, efforts have been made to investigate this issue using both conventional as well as ensemble knowledge distillation techniques. Soumava Paul, Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001 |
Interspeech | 3 |
| 2021 | Robust vowel region detection method for multimode speech
Kumud Tripathi, K. Sreenivasa Rao |
Multim. Tools Appl. | 2 |
| 2021 | Approaches for Multilingual Phone Recognition in Code-switched and Non-code-switched Scenarios Using Indian LanguagesabstractIn this study, we evaluate and compare two different approaches for multilingual phone recognition in code-switched and non-code-switched scenarios. First approach is a front-end Language Identification (LID)-switched to a monolingual phone recognizer (LID-Mono), trained individually on each of the languages present in multilingual dataset. In the second approach, a common multilingual phone-set derived from the International Phonetic Alphabet (IPA) transcription of the multilingual dataset is used to develop a Multilingual Phone Recognition System (Multi-PRS). The bilingual code-switching experiments are conducted using Kannada and Urdu languages. In the first approach, LID is performed using the state-of-the-art i-vectors. Both monolingual and multilingual phone recognition systems are trained using Deep Neural Networks. The performance of LID-Mono and Multi-PRS approaches are compared and analysed in detail. It is found that the performance of Multi-PRS approach is superior compared to more conventional LID-Mono approach in both code-switched and non-code-switched scenarios. For code-switched speech, the effect of length of segments (that are used to perform LID) on the performance of LID-Mono system is studied by varying the window size from 500 ms to 5.0 s, and full utterance. The LID-Mono approach heavily depends on the accuracy of the LID system and the LID errors cannot be recovered. But, the Multi-PRS system by virtue of not having to do a front-end LID switching and designed based on the common multilingual phone-set derived from several languages, is not constrained by the accuracy of the LID system, and hence performs effectively on code-switched and non-code-switched speech, offering low Phone Error Rates than the LID-Mono system. K. Manjunath, Srinivasa Raghavan K. M., K. Sreenivasa Rao, Dinesh Babu Jayagopi, V. Ramasubramanian 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | Relation Prediction of Co-Morbid Diseases Using Knowledge Graph CompletionabstractCo-morbid disease condition refers to the simultaneous presence of one or more diseases along with the primary disease. A patient suffering from co-morbid diseases possess more mortality risk than with a disease alone. So, it is necessary to predict co-morbid disease pairs. In past years, though several methods have been proposed by researchers for predicting the co-morbid diseases, not much work is done in prediction using knowledge graph embedding using tensor factorization. Moreover, the complex-valued vector-based tensor factorization is not being used in any knowledge graph with biological and biomedical entities. We propose a tensor factorization based approach on biological knowledge graphs. Our method introduces the concept of complex-valued embedding in knowledge graphs with biological entities. Here, we build a knowledge graph with disease-gene associations and their corresponding background information. To predict the association between prevalent diseases, we use ComplEx embedding based tensor decomposition method. Besides, we obtain new prevalent disease pairs using the MCL algorithm in a disease-gene-gene network and check their corresponding inter-relations using edge prediction task. Saikat Biswas, Pabitra Mitra, K. Sreenivasa Rao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Glottal Closure Instants Detection from EGG Signal by Classification Approach
Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001 |
INTERSPEECH | 2 |
| 2020 | Excitation modelling using epoch features for statistical parametric speech synthesis
Mittapalle Kiran Reddy, K. Sreenivasa Rao |
Comput. Speech Lang. | 2 |
| 2020 | DNN-Based Cross-Lingual Voice Conversion Using Bottleneck Features
Mittapalle Kiran Reddy, K. Sreenivasa Rao |
Neural Process. Lett. | 2 |
| 2020 | Robust f0 extraction from monophonic signals using adaptive sub-band filtering
Pradeep Rengaswamy, Mittapalle Kiran Reddy, K. Sreenivasa Rao, Pallab Dasgupta |
Speech Commun. | 3 |
| 2020 | Multilingual and multimode phone recognition system for Indian languages
Kumud Tripathi, Mittapalle Kiran Reddy, K. Sreenivasa Rao |
Speech Commun. | 3 |
| 2020 | Children's Story Classification in Indian Languages Using Linguistic and Keyword-based FeaturesabstractThe primary objective of this work is to classify Hindi and Telugu stories into three genres: fable, folk-tale, and legend . In this work, we are proposing a framework for story classification (SC) using keyword and part-of-speech (POS) features. For improving the performance of SC system, feature reduction techniques and combinations of various POS tags are explored. Further, we investigated the performance of SC by dividing the story into parts depending on its semantic structure. In this work, stories are (i) manually divided into parts based on their semantics as introduction, main, and climax ; and (ii) automatically divided into equal parts based on number of sentences in a story as initial, middle, and end . We have also examined sentence increment model, which aims at determining an optimum number of sentences required to identify story genre by incremental selection of sentences in a story. Experiments are conducted on Hindi and Telugu story corpora consisting of 300 and 150 short stories, respectively. The performance of SC system is evaluated using different combinations of keyword and POS-based features, with three well-established machine learning classifiers: (i) Naive Bayes (NB), (ii) k-Nearest Neighbour (KNN), and (iii) Support Vector Machine (SVM). Performance of the classifier is evaluated using 10-fold cross-validation and effectiveness of classifier is measured using precision, recall, and F-measure. From the classification results, it is observed that adding linguistic information boosts the performance of story classification. In view of the structure of the story, main, and initial parts of the story have shown comparatively better performance. The results from the sentence incremental model have indicated that the first nine and seven sentences in Hindi and Telugu stories, respectively, are sufficient for better classification of stories. In most of the studies, SVM models outperformed the other models in classification accuracy. Harikrishna D. M., K. Sreenivasa Rao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | BOXREC: Recommending a Box of Preferred Outfits in Online ShoppingabstractFashionable outfits are generally created by expert fashionistas, who use their creativity and in-depth understanding of fashion to make attractive outfits. Over the past few years, automation of outfit composition has gained much attention from the research community. Most of the existing outfit recommendation systems focus on pairwise item compatibility prediction (using visual and text features) to score an outfit combination having several items, followed by recommendation of top-n outfits or a capsule wardrobe having a collection of outfits based on user’s fashion taste. However, none of these consider a user’s preference of price range for individual clothing types or an overall shopping budget for a set of items. In this article, we propose a box recommendation framework—BOXREC—which at first collects user preferences across different item types (namely, top-wear, bottom-wear, and foot-wear) including price range of each type and a maximum shopping budget for a particular shopping session. It then generates a set of preferred outfits by retrieving all types of preferred items from the database (according to user specified preferences including price ranges), creates all possible combinations of three preferred items (belonging to distinct item types), and verifies each combination using an outfit scoring framework—BOXREC-OSF. Finally, it provides a box full of fashion items, such that different combinations of the items maximize the number of outfits suitable for an occasion while satisfying maximum shopping budget. We create an extensively annotated dataset of male fashion items across various types and categories (each having associated price) and a manually annotated positive and negative formal as well as casual outfit dataset. We consider a set of recently published pairwise compatibility prediction methods as competitors of BOXREC-OSF. Empirical results show superior performance of BOXREC-OSF over the baseline methods. We found encouraging results by performing both quantitative and qualitative analysis of the recommendations produced by BOXREC. Finally, based on user feedback corresponding to the recommendations given by BOXREC, we show that disliked or unpopular items can be a part of attractive outfits. Debopriyo Banerjee, K. Sreenivasa Rao, Shamik Sural, Niloy Ganguly |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2019 | Glottal Closure Instants Detection from Speech Signal by Deep Features Extracted from Raw Speech and Linear Prediction Residual
Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001 |
INTERSPEECH | 2 |
| 2019 | CWT-Based Approach for Epoch Extraction From Telephone Quality SpeechabstractEpochs are the instants of significant excitation to vocal tract system. Existing methods can extract epochs accurately from clean speech signals. However, identification of epoch locations from band-limited telephonic speech is challenging due to the attenuation of fundamental frequency component and degradation caused by channel effect. This letter proposes an epoch extraction method that can accurately extract epochs from clean as well as telephonic speech signals. In the proposed method, the significant impulse-like discontinuities are extracted directly from the speech signal using continuous wavelet transform. The performance of the proposed method is evaluated using three speakers, namely, SLT, BDL, and JMK from CMU Arctic database. The clean speech is simulated using G.191 software tools to obtain telephonic speech. Experimental results show that the epoch identification rate of proposed method is significantly better than the state-of-the-art methods for the telephone quality speech. Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, K. Sreenivasa Rao |
IEEE Signal Process. Lett. | 3 |
| 2018 | Robust Detection of Glottal Activity Using Unwrapped Phase Electroglottographic SignalabstractGlottal Activity is defined by the process of exciting the vocal tract system by the vibration of the vocal folds during speech processing. Detection of glottal activity refers to identifying the glottal instants present within a glottal cycle. GCI and GOI are the two important glottal instants of a glottal cycle. The DEGG signal provides significant peaks at those instants for normal voicing, but they are prone to error in detection for low strength EGG signal. In this paper, we mainly focus on the segments of EGG signal where the strength of the EGG signal is very poor and irregular in periodicity. The robust detection of the glottal instants in those segments of EGG signal will enhance the overall accuracy of the detection of glottal instants. In the proposed method, the unwrapped phase of the EGG signal has been used for the detection of the glottal instants. The phase of the signal has uniform characteristics throughout the signal which helps to detect the glottal instants robustly. Tanumay Mandal, K. Sreenivasa Rao |
ICASSP | 2 |
| 2018 | Modifying LSTM Posteriors with Manner of Articulation Knowledge to Improve Speech Recognition PerformanceabstractThe variant of recurrent neural networks (RNN) such as long short-term memory (LSTM) is successful in sequence modelling such as automatic speech recognition (ASR) framework. However the decoded sequence is prune to have false substitutions, insertions and deletions. We exploit the spectral flatness measure (SFM) computed on the magnitude linear prediction (LP) spectrum to detect two broad manners of articulation namely sonorants and obstruents. In this paper, we modify the posteriors generated at the output layer of LSTM according to the manner of articulation detection. The modified posteriors are given to the conventional decoding graph to minimize the false substitutions and insertions. The proposed method decreased the phone error rate (PER) by nearly 0.7 % and 0.3 % when evaluated on core TIMIT test corpus as compared to the conventional decoding involved in the deep neural networks (DNN) and the state of the art LSTM respectively. Pradeep Rengaswamy, K. Sreenivasa Rao |
ICMLA | 2 |
| 2018 | Harmonic-Percussive Source Separation of Polyphonic Music by Suppressing Impulsive Noise Events
Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001 |
INTERSPEECH | 2 |
| 2018 | Classification of Disorders in Vocal Folds Using Electroglottographic Signal
Tanumay Mandal, K. Sreenivasa Rao, Sanjay Kumar Gupta |
INTERSPEECH | 2 |
| 2018 | Indian Languages ASR: A Multilingual Phone Recognition Framework with IPA Based Common Phone-set, Predicted Articulatory Features and Feature fusion
K. Manjunath, K. Sreenivasa Rao, Dinesh Babu Jayagopi, V. Ramasubramanian 0001 |
INTERSPEECH | 2 |
| 2018 | Analysis of sparse representation based feature on speech mode classification
Kumud Tripathi, K. Sreenivasa Rao |
INTERSPEECH | 2 |
| 2018 | One for the Road: Recommending Male Street Attire
Debopriyo Banerjee, Niloy Ganguly, Shamik Sural, K. Sreenivasa Rao |
PAKDD (3) | 4 |
| 2018 | Inverse filter based excitation model for HMM-based speech synthesis systemabstractEven today, the speech generated by hidden Markov model (HMM)‐based speech synthesis system (HTS) still has the buzziness due to the improper modelling of the excitation signal. This study proposes an efficient excitation modelling approach for improving the quality of HTS. In the proposed method, the residual signal obtained from inverse filter is parameterised as excitation features. HMMs are used to model these excitation parameters. During synthesis, the excitation signal is constructed by overlap adding the natural residual segments, and the excitation signal is further modified as per the target source features generated from HMMs. The proposed approach is incorporated in the HTS. Performance evaluation results indicate that the proposed method enhances the quality of synthesis, and is better than the state‐of‐the‐art approaches used for modelling the excitation signal. Mittapalle Kiran Reddy, K. Sreenivasa Rao |
IET Signal Process. | 2 |
| 2018 | A robust unsupervised pattern discovery and clustering of speech signals
Kishore Kumar R., Lokendra Birla, K. Sreenivasa Rao |
Pattern Recognit. Lett. | 3 |
| 2018 | Epoch detection from emotional speech signal using zero time windowing
Jainath Yadav, Md. S. Fahad, K. Sreenivasa Rao |
Speech Commun. | 3 |
| 2017 | Accurate Synchronization of Speech and EGG Signal Using Phase Information
S. B. Sunil Kumar, K. Sreenivasa Rao, Tanumay Mandal |
INTERSPEECH | 2 |
| 2017 | Implicit processing of LP residual for language identification
Dipanjan Nandi, Debadatta Pati, K. Sreenivasa Rao |
Comput. Speech Lang. | 3 |
| 2017 | Parametric representation of excitation source information for language identification
Dipanjan Nandi, Debadatta Pati, K. Sreenivasa Rao |
Comput. Speech Lang. | 3 |
| 2017 | Generation of creaky voice for improving the quality of HMM-based speech synthesis
N. P. Narendra, K. Sreenivasa Rao |
Comput. Speech Lang. | 2 |
| 2017 | Robust Pitch Extraction Method for the HMM-Based Speech Synthesis SystemabstractThis letter proposes an efficient method for extracting pitch from speech signals for the hidden Markov model (HMM)-based speech synthesis system (HTS). In the proposed method, voicing detection and pitch estimation is performed using the mean signal obtained from continuous wavelet transform coefficients. The proposed pitch extraction method is integrated in the HMM-based speech synthesis system. The Performance of the proposed method is evaluated on CMU Arctic and Keele databases. Both objective and subjective evaluation results show that the quality of speech synthesized with the proposed pitch estimation method is much better compared with HMM-based speech synthesis systems developed using the state-of-the-art pitch extraction methods, namely, robust algorithm for pitch tracking and speech transformation and representation using adaptive interpolation of weighted spectrum employed in the HTS. Mittapalle Kiran Reddy, K. Sreenivasa Rao |
IEEE Signal Process. Lett. | 2 |
| 2016 | A deterministic plus noise model of excitation signal using principal component analysis for parametric speech synthesisabstractThis paper proposes a new approach of modeling the excitation signal as deterministic and noise components. Initially, a study on characteristics of excitation or residual signal around glottal closure instant (GCI) is performed using principal component analysis (PCA). Based on the study, the segment of residual signal around GCI is considered as the deterministic component and the remaining part of the residual signal is considered as the noise component. The deterministic component is parameterized using PCA coefficients, and the noise component can be represented in terms of spectral and amplitude envelopes. The proposed excitation modeling approach is incorporated in the HMM-based speech synthesis system. Subjective evaluation results show a significant improvement in the quality of speech synthesized by the proposed method, compared to three existing methods. N. P. Narendra, K. Sreenivasa Rao |
ICASSP | 2 |
| 2016 | Predominant melody extraction from vocal polyphonic music signal by combined spectro-temporal methodabstractA combined spectro-temporal based method is proposed to derive the predominant melody from vocal polyphonic music signals. In the proposed method, vocal (voiced) and non-vocal (unvoiced) segments are determined by strength of excitation. The vocal segments are further divided into voiced notes by detecting their onsets using transition cues present in spectral domain. The melody contour present in each of the voiced note segments is obtained by using an adaptive zero frequency filtering (ZFF) in time domain. The process of melody extraction is provided in more detail and the initial results showed the potential use of the proposed method for vocal melody extraction. Gurunath Reddy M, K. Sreenivasa Rao |
ICASSP | 2 |
| 2016 | Enhanced Harmonic Content and Vocal Note Based Predominant Melody Extraction from Vocal Polyphonic Music Signals
Gurunath Reddy M, K. Sreenivasa Rao |
INTERSPEECH | 2 |
| 2016 | A Robust Non-Parametric and Filtering Based Approach for Glottal Closure Instant Detection
Pradeep Rengaswamy, Gurunath Reddy M, K. Sreenivasa Rao, Pallab Dasgupta |
INTERSPEECH | 3 |
| 2016 | Prosody modeling for syllable based text-to-speech synthesis using feedforward neural networks
V. Ramu Reddy, K. Sreenivasa Rao |
Neurocomputing | 2 |
| 2016 | Voice/non-voice detection using phase of zero frequency filtered speech signal
S. B. Sunil Kumar, K. Sreenivasa Rao |
Speech Commun. | 2 |
| 2016 | Time-domain deterministic plus noise model based hybrid source modeling for statistical parametric speech synthesis
N. P. Narendra, K. Sreenivasa Rao |
Speech Commun. | 2 |
| 2015 | Automatic detection of creaky voice using epoch parameters
N. P. Narendra, K. Sreenivasa Rao |
INTERSPEECH | 2 |
| 2014 | A novel boosting algorithm for improved i-vector based speaker verification in noisy environments
Sourjya Sarkar, K. Sreenivasa Rao |
INTERSPEECH | 2 |
| 2013 | High quality text-to-speech synthesis system with efficient duration models developed using coding schemes based on vowel production characteristicsabstractThis paper explores encoding schemes based on production characteristics of vowels. The performance of coding schemes is analyzed for accurate prediction of durations of syllables using neural network models. Linguistic and production constraints represented by positional, contextual, phonological and articulatory (PCPA) features are used for predicting the durations of syllables. These features are coded with distinct numerical values before feeding to the neural network for building models. The evaluation of coding schemes is carried out by means of objective and subjective measures. The quality of text-to-speech synthesis system is observed to be better by incorporating the duration model with vowels coded based on lip roundness. V. Ramu Reddy, K. Sreenivasa Rao |
ISDA | 2 |
| 2013 | Two-stage intonation modeling using feedforward neural networks for syllable based text-to-speech synthesis
V. Ramu Reddy, K. Sreenivasa Rao |
Comput. Speech Lang. | 2 |
| 2013 | Non-uniform time scale modification using instants of significant excitation and vowel onset points
K. Sreenivasa Rao, Anil Kumar Vuppala |
Speech Commun. | 1 |
| 2013 | Detection of Vowel Offset Point From Speech SignalabstractVowel regions play important role in various speech tasks, such as speech segmentation, speaker-verification, prosody modification and emotion conversion. The instants at which the onset and offset of vowel take place in the speech signal are known as vowel onset point and vowel offset point, respectively. Vowel regions start with the vowel onset point and end with the vowel offset point. In this letter, we have proposed two methods for determining the vowel offset points from the speech signal. The first method explores the combination of evidences from excitation source, spectral peaks and modulation spectrum for determining the vowel offset point. In the second method, spectral energy within glottal closure region is used for determining the vowel offset point. The proposed vowel offset point detection methods are evaluated on TIMIT database under clean and noisy environments. Jainath Yadav, K. Sreenivasa Rao |
IEEE Signal Process. Lett. | 2 |
| 2012 | Vowel Onset Point Detection for Low Bit Rate Coded SpeechabstractIn this paper, we propose a method for detecting the vowel onset points (VOPs) for low bit rate coded speech. VOP is the instant at which the onset of the vowel takes place in the speech signal. VOP plays an important role for the applications, such as consonant-vowel (CV) unit recognition and speech rate modification. The proposed VOP detection method is based on the spectral energy present in the glottal closure region of the speech signal. Speech coders considered to carry out this study are Global System for Mobile Communications (GSM) full rate, code-excited linear prediction (CELP), and mixed-excitation linear prediction (MELP). TIMIT database and CV units collected from the broadcast news corpus are used for evaluation. Performance of the proposed method is compared with existing methods, which uses the combination of evidence from the excitation source, spectral peaks energy, and modulation spectrum. The proposed VOP detection method has shown significant improvement in the performance compared to the existing method under clean as well as coded cases. The effectiveness of the proposed VOP detection method is analyzed in CV recognition by using VOP as an anchor point. Anil Kumar Vuppala, Jainath Yadav, Saswat Chakrabarti, K. Sreenivasa Rao |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Recognition of emotions from video using neural network models
K. Sreenivasa Rao, V. K. Saroj, Sudhamay Maity, Shashidhar G. Koolagudi |
Expert Syst. Appl. | 1 |
| 2010 | Voice conversion by mapping the speaker-specific features using pitch synchronous approach
K. Sreenivasa Rao |
Comput. Speech Lang. | 1 |
| 2009 | Intonation modeling for Indian languages
K. Sreenivasa Rao, Bayya Yegnanarayana |
Comput. Speech Lang. | 1 |
| 2009 | Duration modification using glottal closure instants and vowel onset points
K. Sreenivasa Rao, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 2007 | Modeling durations of syllables using neural networks
K. Sreenivasa Rao, Bayya Yegnanarayana |
Comput. Speech Lang. | 1 |
| 2007 | Determination of Instants of Significant Excitation in Speech Using Hilbert Envelope and Group Delay FunctionabstractThis letter proposes a time-effective method for determining the instants of significant excitation in speech signals. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. The proposed method consists of two phases: the first phase determines the approximate epoch locations using the Hilbert envelope of the linear prediction residual of the speech signal. The second phase determines the accurate locations of the instants of significant excitation by computing the group delay around the approximate epoch locations derived from the first phase. The accuracy in determining the instants of significant excitation and the time complexity of the proposed method is compared with the group delay based approach. K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 1 |
| 2006 | Prosody modification using instants of significant excitationabstractProsody modification involves changing the pitch and duration of speech without affecting the message and naturalness. This paper proposes a method for prosody (pitch and duration) modification using the instants of significant excitation of the vocal tract system during the production of speech. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. Instants of significant excitation are computed from the linear prediction (LP) residual of speech signals by using the property of average group-delay of minimum phase signals. The modification of pitch and duration is achieved by manipulating the LP residual with the help of the knowledge of the instants of significant excitation. The modified residual is used to excite the time-varying filter, whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is good and is without any significant distortion. The proposed method is evaluated using waveforms, spectrograms, and listening tests. The performance of the method is compared with linear prediction pitch synchronous overlap and add (LP-PSOLA) method, which is another method for prosody manipulation based on the modification of the LP residual. The original and the synthesized speech signals obtained by the proposed method and by the LP-PSOLA method are available for listening at http://speech.cs.iitm.ernet.in/Main/result/prosody.html. K. Sreenivasa Rao, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Modeling syllable duration in Indian languages using neural networksabstractWe propose a neural network model for predicting the syllable duration in Indian languages. A four layer feedforward neural network trained with a backpropagation algorithm is used for modeling the syllable duration. Analysis is performed on broadcast news data in Hindi, Telugu and Tamil in order to predict the duration of syllables in these languages using a neural network model. The input to the neural network consists of a set of phonological, positional and contextual features extracted from the text. About 88% of the syllable durations are predicted within 25% of the actual duration. The relative importance of the positional and contextual features are examined separately. K. Sreenivasa Rao, Bayya Yegnanarayana |
ICASSP (5) | 1 |
| 2004 | Two-Stage Duration Model for Indian Languages Using Neural Networks
K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICONIP | 1 |
| 2004 | Intonation modeling for indian languagesabstractIn this paper we propose models for predicting the intonation for the sequence of syllables present in the utterance.The term intonation refers to the temporal changes of the fundamental frequency ðF 0 Þ.Neural networks are used to capture the implicit intonation knowledge in the sequence of syllables of an utterance.We focus on the development of intonation models for predicting the sequence of fundamental frequency values for a given sequence of syllables.Labeled broadcast news data in the languages Hindi, Telugu and Tamil is used to develop neural network models in order to predict the F 0 of syllables in these languages.The input to the neural network consists of a feature vector representing the positional, contextual and phonological constraints.The interaction between duration and intonation constraints can be exploited for improving the accuracy further.From the studies we find that 88% of the F 0 values (pitch) of the syllables could be predicted from the models within 15% of the actual F 0 .The performance of the intonation models is evaluated using objective measures such as average prediction error ðlÞ, standard deviation ðrÞ and correlation coefficient ðcÞ.The prediction accuracy of the intonation models is further evaluated using listening tests.The prediction performance of the proposed intonation models using neural networks is compared with Classification and Regression Tree (CART) models. K. Sreenivasa Rao, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2003 | Prosodic manipulation using instants of significant excitationabstractThe paper proposes a technique for prosodic (pitch and duration) manipulation using instants of significant excitation. Instants of significant excitation correspond to the instants of glottal closure (epochs) in voiced speech and to some random excitations like burst onset in the case of nonvoiced speech. Instants of significant excitation are computed from the average group delay of minimum phase signals. The manipulation of pitch and duration is achieved by modifying the linear prediction (LP) residual with the help of instants of significant excitation as pitch markers. The modified residual is used to excite the time-varying filter whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is found to be natural, and is without any distortion. The original and corresponding synthesized speech signals from the proposed approach are available at http://speech.cs.iitm.ernet.in/Main/Results/Prosody.html. K. Sreenivasa Rao, Bayya Yegnanarayana |
ICASSP (1) | 1 |
| 2003 | Prosodic manipulation using instants of significant excitationabstractThis paper proposes a technique for prosodic (pitch and duration) manipulation using instants of significant excitation. Instants of significant excitation correspond to the instants of glottal closure (epochs) in voiced speech and to some random excitations like burst onset in the case of nonvoiced speech. Instants of significant excitation are computed from the average group delay of minimum phase signals. The manipulation of pitch and duration is achieved by modifying the linear prediction (LP) residual with the help of instants of significant excitation as pitch markers. The modified residual is used to excite the time-varying filter whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is found to be natural, and is without any distortion. The original and corresponding synthesized speech signals from the proposed approach are available for listening at http://speech.cs.iitm.ernet.in/Main/Results/Prosody.html. K. Sreenivasa Rao, Bayya Yegnanarayana |
ICME | 1 |
| 2003 | Combining evidence from multiple modular networks for recognition of consonant-vowel units of speechabstractIn this paper, we present a method to combine evidence from multiple classifiers to recognize a large number of subword units of speech using small size training data sets. Grouping criteria based on phonetic description are considered, to build multiple modular networks for recognition of the large number of units. Nonlinear compression of feature vectors is carried out to obtain reduced dimensional patterns, and multiple classifiers are trained separately using the uncompressed feature vectors and compressed feature vectors. Evidence from multiple classifiers at different stages in the recognition system is combined using the sum rule. Effectiveness of the proposed method is demonstrated for recognition of isolated utterances of 145 consonant-vowel units of speech. Suryakanth V. Gangashetty, K. Sreenivasa Rao, A. Nayeemulla Khan, Chellu Chandra Sekhar, Bayya Yegnanarayana |
IJCNN | 2 |
| 2002 | Speech enhancement using excitation source informationabstractThis paper proposes an approach for processing speech from multiple microphones to enhance speech degraded by noise and reverberation. The approach is based on exploiting the features of the excitation source in speech production. In particular, the characteristics of voiced speech can be used to derive a coherently added signal from the linear prediction (LP) residuals of the degraded speech data from different microphones. A weight function is derived from the coherently added signal. For coherent addition the time-delay between a pair of microphones is estimated using the knowledge of the source information present in the LP residual. The enhanced speech is generated by exciting the time varying all-pole filter with the weighted LP residual. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, K. Sreenivasa Rao |
ICASSP | 3 |