K. Sri Rama Murty

dblp:75/3459 · also Kodukula Sri Rama Murty, Sri Rama Murty Kodukula · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-6355-5287ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 25 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting
abstract
Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal supervision, disjoint optimization of audio-audio and audio-text alignment, and the need for task-specific models. To address these shortcomings, we propose a joint multimodal contrastive learning framework that unifies both acoustic and cross-modal supervision in a shared embedding space. Our approach simultaneously optimizes: (i) audio-text contrastive learning, inspired by the CLAP loss, to align audio and text representations and (ii) audio-audio contrastive learning, via Deep Word Discrimination (DWD) loss, to enhance intra-class compactness and inter-class separation. The proposed method outperforms existing AWE baselines on word discrimination task while flexibly supporting both STD and KWS. To our knowledge, this is the first comprehensive approach of its kind.
Gundluru Ramesh, K. Sri Rama Murty
ASRU3
2025 Stream-TTS: A Low-Latency Text-to-Speech using Kolmogorov-Arnold Networks for Streaming Speech Applications
abstract
The rise of conversational AI and multimodal streaming applications has led to a significant demand for low-latency Text-to-Speech (TTS) systems. This work presents a multilingual low-latency model that leverages the functional decomposition principles of Kolmogorov-Arnold Networks (KANs) in modeling the nonlinear functions with lesser parameters and smaller computational graphs than multi-layer perceptrons (MLPs). To enhance the learning capabilities of the proposed compact model, we use a supervised auxiliary learning approach with multiple subtasks whose target labels are extracted from the speech signal. We use a multispeaker nonlinear vocoder to reconstruct the natural-sounding speech from the Melspectrograms. The proposed model achieved a very low latency with a real-time factor (RTF) of 0.0795 and a mean opinion score (MOS) of 4.09 for naturalness in the multilingual streaming TTS challenge organized as part of ICASSP-2025.
Giridhar Pamisetty, Riya Ann Easow, Kaustubh Gupta, K. Sri Rama Murty
ICASSP4
2025 Test-Time Training for Speech Enhancement
abstract
This paper introduces a novel application of Test-Time Training (TTT) for Speech Enhancement, addressing the challenges posed by unpredictable noise conditions and domain shifts. This method combines a main speech enhancement task with a self-supervised auxiliary task in a Y-shaped architecture. The model dynamically adapts to new domains during inference time by optimizing the proposed self-supervised tasks like noise-augmented signal reconstruction or masked spectrogram prediction, bypassing the need for labeled data. We further introduce various TTT strategies offering a trade-off between adaptation and efficiency. Evaluations across synthetic and real-world datasets show consistent improvements across speech quality metrics, outperforming the baseline model. This work highlights the effectiveness of TTT in speech enhancement, providing insights for future research in adaptive and robust speech processing.
Avishkar Behera, Riya Ann Easow, Venkatesh Parvathala, K. Sri Rama Murty
INTERSPEECH4
2025 Exploiting Bispectral Features for Single-Channel Speech Enhancement
Venkatesh Parvathala, Gundluru Ramesh, Sreekanth Sankala, K. Sri Rama Murty
INTERSPEECH4
2025 Dynamic Layer Gating for Speech Enhancement
Venkatesh Parvathala, K. Sri Rama Murty
INTERSPEECH2
2025 MSFNet: A Nested Model for Multi-Sampling-Frequency Speech Enhancement
Venkatesh Parvathala, K. Sri Rama Murty
INTERSPEECH2
2025 Adversarial Attacks on Text-dependent Speaker Verification System
Sreekanth Sankala, Venkatesh Parvathala, Gundluru Ramesh, K. Sri Rama Murty
INTERSPEECH4
2023 Lightweight Prosody-TTS for Multi-Lingual Multi-Speaker Scenario
abstract
This work presents a lightweight end-to-end text-to-speech (TTS) synthesis for the multi-lingual multi-speaker (ML-MS) scenario. The proposed system uses nonautoregressive modular architecture with interconnected subnets for text-encoder, duration estimator, f0estimator, and acoustic decoder. The text encoder is conditioned with language embeddings, while the duration and f0estimators are conditioned with speaker embeddings. All the subnets are optimized in an end-to-end fashion using accumulated loss across the modules. The intermediate auxiliary loss functions help effectively capture the speech information with lesser data. The proposed model achieved a mean opinion score (MOS) of 4.40 and a speaker similarity score of 3.8 with just 4.89 million (M) parameters in LIMMITS grand challenge organized as part of ICASSP-23.
Giridhar Pamisetty, Sahukari Chaitanya Varun, K. Sri Rama Murty
ICASSP3
2022 Multi-Feature Integration for Speaker Embedding Extraction
abstract
The performance of the automatic speaker recognition system is becoming more and more accurate, with the advancement in deep learning methods. However, current speaker recognition system performances are subjective to the training conditions, thereby decreasing the performance drastically even on slightly varied test data. A lot of methods such as using various data augmentation structures, various loss functions, and integrating multiple features systems have been proposed and shown a performance improvement. This work focuses on integrating multiple features to improve speaker verification performance. Speaker information is commonly represented in the different kinds of features, where the redundant and irrelevant information such as noise and channel information will affect the dimensions of different features in a different manner. In this work, we intend to maximize the speaker information by reconstructing the extracted speaker information in one feature from the other features while at the same time minimizing the irrelevant information. The experiments with the multi-feature integration model demonstrate improved performance than the stand-alone models by significant margins. Also, the extracted speaker embeddings are found to be noise-robust.
Sreekanth Sankala, B. Shaik Mohammad Rafi, K. Sri Rama Murty
ICASSP3
2021 Self-Supervised Phonotactic Representations for Language Identification
abstract
Phonotactic constraints characterize the sequence of permissible phoneme structures in a language and hence form an important cue for language identification (LID) task. As phonotactic constraints span across multiple phonemes, the short-term spectral analysis (20-30 ms) alone is not sufficient to capture them. The speech signal has to be analyzed over longer contexts (100s of milliseconds) in order to extract features representing the phonotactic constraints. The supervised senone classifiers, aimed at modeling triphone context, have been used for extracting language-specific features for the LID task. However, it is difficult to get large amounts of manually labeled data to train the supervised models. In this work, we explore a selfsupervised approach to extract long-term contextual features for the LID task. We have used wav2vec architecture to extract contextualized representations from multiple frames of the speech signal. The contextualized representations extracted from the pre-trained wav2vec model are used for the LID task. The performance of the proposed features is evaluated on a dataset containing 7 Indian languages. The proposed self-supervised embeddings achieved 23% absolute improvement over the acoustic features and 3% absolute improvement over their supervised counterparts. Copyright © 2021 ISCA.
Gundluru Ramesh, C. Shiva Kumar, K. Sri Rama Murty
Interspeech3
2019 Zero Resource Speaking Rate Estimation from Change Point Detection of Syllable-like Units
abstract
Speaking rate is an important attribute of the speech signal which plays a crucial role in the performance of automatic speech processing systems. In this paper, we propose to estimate the speaking rate by segmenting the speech into syllable-like units using end point detection algorithms which do not require any training and fine-tuning. Also, there are no predefined constraints on the expected number of syllabic segments. The acoustic subword units are obtained only from speech signal to estimate the speaking rate without any requirement of transcriptions or phonetic knowledge of the speech data. A recent theta-rate oscillator based syllabification algorithm is also employed for speaking rate estimation. The performance is evaluated on TIMIT corpus and spontaneous speech from Switchboard corpus. The correlation results are comparable to recent algorithms which are trained with specific training set and/or make use of the available transcriptions.
Shekhar Nayak, Saurabhchand Bhati, K. Sri Rama Murty
ICASSP3
2019 Importance of Analytic Phase of the Speech Signal for Detecting Replay Attacks in Automatic Speaker Verification Systems
abstract
In this paper, the importance of analytic phase of the speech signal in automatic speaker verification systems is demonstrated in the context of replay spoof attacks. In order to accurately detect the replay spoof attacks, effective feature representations of speech signals are required to capture the distortion introduced due to the intermediate playback/recording devices, which is convolutive in nature. Since the convolutional distortion in time-domain translates to additive distortion in the phase-domain, we propose to use IFCC features extracted from the analytic phase of the speech signal. The IFCC features contain information from both clean speech and distortion components. The clean speech component has to be subtracted in order to highlight the distortion component introduced by the playback/recording devices. In this work, a dictionary learned from the IFCCs extracted from clean speech data is used to remove the clean speech component. The residual distortion component is used as a feature to build binary classifier for replay spoof detection. The proposed phase-based features delivered a 9% absolute improvement over the baseline system built using magnitude-based CQCC features.
B. Shaik Mohammad Rafi, K. Sri Rama Murty
ICASSP2
2019 Unsupervised Acoustic Segmentation and Clustering Using Siamese Network Embeddings
abstract
Unsupervised discovery of acoustic units from the raw speech signal forms the core objective of zero-resource speech processing. It involves identifying the acoustic segment boundaries and consistently assigning unique labels to acoustically similar segments. In this work, the possible candidates for segment boundaries are identified in an unsupervised manner from the kernel Gram matrix computed from the Mel-frequency cepstral coefficients (MFCC). These segment boundary candidates are used to train a siamese network, that is intended to learn embeddings that minimize intrasegment distances and maximize the intersegment distances. The siamese embeddings capture phonetic information from longer contexts of the speech signal and enhance the intersegment discriminability. These properties make the siamese embeddings better suited for acoustic segmentation and clustering than the raw MFCC features. The Gram matrix computed from the siamese embeddings provides unambiguous evidence for boundary locations. The initial candidate boundaries are refined using this evidence, and siamese embeddings are extracted for the new acoustic segments. A graph growing approach is used to cluster the siamese embeddings, and a unique label is assigned to acoustically similar segments. The performance of the proposed method for acoustic segmentation and clustering is evaluated on Zero Resource 2017 database.
Saurabhchand Bhati, Shekhar Nayak, K. Sri Rama Murty, Najim Dehak
INTERSPEECH3
2019 Unsupervised Universal Attribute Modeling for Action Recognition
abstract
A fixed dimensional representation for action clips of varying lengths has been proposed in the literature using aggregation models like bag-of-words and Fisher vector. These representations are high dimensional and require classification techniques for action recognition. In this paper, we propose a framework for unsupervised extraction of a discriminative low-dimensional representation called action-vector. To start with, local spatio-temporal features are utilized to capture the action attributes implicitly in a large Gaussian mixture model called the universal attribute model (UAM). To enhance the contribution of the significant attributes in each action clip, a maximum aposteriori adaptation of the UAM means is performed for each clip. This results in a concatenated mean vector called super action vector (SAV) for each action clip. However, the SAV is still high dimensional because of the presence of redundant attributes. Hence, we employ factor analysis to represent every SAV only in terms of the few important attributes contributing to the action clip. This leads to a low-dimensional representation called action-vector. This entire procedure requires no class labels and produces action-vectors that are distinct representations of each action irrespective of the inter-actor variability encountered in unconstrained videos. An evaluation on trimmed action datasets UCF101 and HMDB51 demonstrates the efficacy of action-vectors for action classification over state-of-the-art techniques. Moreover, we also show that action-vectors can adequately represent untrimmed videos from the THUMOS14 dataset and produce classification results comparable to existing techniques.
Debaditya Roy, K. Sri Rama Murty, C. Krishna Mohan
IEEE Trans. Multim.2
2018 Phoneme Based Embedded Segmental K-Means for Unsupervised Term Discovery
abstract
Identifying and grouping the frequently occurring word-like patterns from raw acoustic waveforms is an important task in the zero resource speech processing. Embedded segmental K-means (ES-KMeans) discovers both the word boundaries and the word types from raw data. Starting from an initial set of subword boundaries, the ES-Kmeans iteratively eliminates some of the boundaries to arrive at frequently occurring longer word patterns. Notice that the initial word boundaries will not be adjusted during the process. As a result, the performance of the ES-Kmeans critically depends on the initial subword boundaries. Originally, syllable boundaries were used to initialize ES-Kmeans. In this paper, we propose to use a phoneme segmentation method that produces boundaries closer to true boundaries for ES-KMeans initialization. The use of shorter units increases the number of initial boundaries which leads to a significant increment in the computational complexity. To reduce the computational cost, we extract compact lower dimensional embeddings from an auto-encoder. The proposed algorithm is benchmarked on Zero Resource 2017 challenge, which consists of 70 hours of unlabeled data across three languages, viz. English, French, and Mandarin. The proposed algorithm outperforms the baseline system without any language-specific parameter tuning.
Saurabchiand Bhati, Herman Kamper, K. Sri Rama Murty
ICASSP3
2018 Action Recognition Based on Discriminative Embedding of Actions Using Siamese Networks
abstract
Actions can be recognized effectively when the various atomic attributes forming the action are identified and combined in the form of a representation. In this paper, a low-dimensional representation is extracted from a pool of attributes learned in a universal Gaussian mixture model using factor analysis. However, such a representation cannot adequately discriminate between actions with similar attributes. Hence, we propose to classify such actions by leveraging the corresponding class labels. We train a Siamese deep neural network with a contrastive loss on the low-dimensional representation. We show that Siamese networks allow effective discrimination even between similar actions. The efficacy of the proposed approach is demonstrated on two benchmark action datasets, HMDB51 and MPII Cooking Activities. On both the datasets, the proposed method improves the state-of-the-art performance considerably.
Debaditya Roy, C. Krishna Mohan, K. Sri Rama Murty
ICIP3
2018 Speech Source Separation Using ICA in Constant Q Transform Domain
abstract
In order to separate individual sources from convoluted speech mixtures, complex-domain independent component analysis (ICA) is employed on the individual frequency bins of time-frequency representations of the speech mixtures, obtained using short term Fourier transform (STFT). The frequency components computed using STFT are separated by constant frequency di�erence with a constant frequency resolution. However, it is well known that the human auditory mechanism o�ers better resolution at lower frequencies. Hence, the perceptual quality of the extracted sources critically depends on the separation achieved in the lower frequency components. A method has been proposed to perform source separation on the time-frequency representation computed though constant Q transform, which o�ers non uniform logarithmic binning in the frequency domain. Complex-domain ICA is performed on the individual bins of the CQT in order to get separated components in each frequency bin which are suitably scaled and permuted to obtain separated sources in the CQT domain. The estimated sources are obtained by applying inverse Q transform to the scaled and permuted sources. In comparison with the STFT based frequency domain ICA methods, there has been a consistent improvement of 3dB or more in the Signal to Interference Ratios of the extracted sources. vi
Dheeraj Sai D. V. L. N, Kishor K. S, K. Sri Rama Murty
INTERSPEECH3
2017 Action-vectors: Unsupervised movement modeling for action recognition
abstract
Representation and modelling of movements play a significant role in recognising actions in unconstrained videos. However, explicit segmentation and labelling of movements are non-trivial because of the variability associated with actors, camera viewpoints, duration etc. Therefore, we propose to train a GMM with a large number of components termed as a universal movement model (UMM). This UMM is trained using motion boundary histograms (MBH) which capture the motion trajectories associated with the movements across all possible actions. For a particular action video, the MAP adapted mean vectors of the UMM are concatenated to form a fixed dimensional representation referred to as “super movement vector” (SMV). However, SMV is still high dimensional and hence, Baum-Welch statistics extracted from the UMM are used to arrive at a compact representation for each action video, which we refer to as an “action-vector”. It is shown that even without the use of class labels, action-vectors provide a more discriminatory representation of action classes translating to a 8 % relative improvement in classification accuracy for action-vectors based on MBH features over naïve MBH features on the UCF101 dataset. Furthermore, action-vectors projected with LDA achieve 93% accuracy on the UCF101 dataset which rivals state-of-the-art deep learning techniques.
Debaditya Roy, K. Sri Rama Murty, C. Krishna Mohan
ICASSP2
2017 Unsupervised Speech Signal to Symbol Transformation for Zero Resource Speech Applications
abstract
Zero resource speech processing refers to a scenario where no or minimal transcribed data is available. In this paper, we propose a three-step unsupervised approach to zero resource speech processing, which does not require any other information/dataset. In the first step, we segment the speech signal into phonemelike units, resulting in a large number of varying length segments. The second step involves clustering the varying-length segments into a finite number of clusters so that each segment can be labeled with a cluster index. The unsupervised transcriptions, thus obtained, can be thought of as a sequence of virtual phone labels. In the third step, a deep neural network classifier is trained to map the feature vectors extracted from the signal to its corresponding virtual phone label. The virtual phone posteriors extracted from the DNN are used as features in the zero resource speech processing. The effectiveness of the proposed approach is evaluated on both ABX and spoken term discovery tasks (STD) using spontaneous American English and Tsonga language datasets, provided as part of zero resource 2015 challenge. It is observed that the proposed system outperforms baselines, supplied along the datasets, in both the tasks without any task specific modifications
Saurabhchand Bhati, Shekhar Nayak, K. Sri Rama Murty
INTERSPEECH3
2017 IITG-Indigo System for NIST 2016 SRE Challenge
abstract
Orientador : Juarez Brandão Lopes
Nagendra Kumar 0004, Rohan Kumar Das, Sarfaraz Jelil, Dhanush B. K, H. Kashyap, K. Sri Rama Murty, Sriram Ganapathy, Rohit Sinha 0003, S. R. Mahadeva Prasanna
INTERSPEECH6
2016 Prosody Modification Using Allpass Residual of Speech Signals
abstract
In this paper, we attempt to signify the role of phase spectrum of speech signals in acquiring an accurate estimate of excitation source for prosody modification. The phase spectrum is parametrically modeled as the response of an all pass (AP) filter, and the filter coefficients are estimated by considering the linear prediction (LP) residual as the output of the AP filter. The resultant residual signal, namely AP residual, exhibits unambiguous peaks corresponding to epochs, which are chosen as pitch markers for prosody modification. This strategy efficiently removes ambiguities associated with pitch marking, required for pitch synchronous overlap-add (PSOLA) method. The prosody modification using AP residual is advantageous than time domain PSOLA (TD-PSOLA) using speech signals, as it offers fewer distortions due to its flat magnitude spectrum. Windowing centered around unambiguous peaks in AP residual is used for segmentation, followed by pitch/duration modification of AP residual by mapping of pitch markers. The modified speech signal is obtained from modified AP residual using synthesis filters. The mean opinion scores are used for performance evaluation of the proposed method, and it is observed that the AP residual-based method delivers equivalent performance as that of LP residual based method using epochs, and better performance than the linear prediction PSOLA (LP-PSOLA).
Karthika Vijayan, K. Sri Rama Murty
INTERSPEECH2
2016 Significance of analytic phase of speech signals in speaker verification
Karthika Vijayan, Raghavendra Reddy Pappagari, K. Sri Rama Murty
Speech Commun.3
2015 Feature selection using Deep Neural Networks
abstract
Feature descriptors involved in video processing are generally high dimensional in nature. Even though the extracted features are high dimensional, many a times the task at hand depends only on a small subset of these features. For example, if two actions like running and walking have to be identified, extracting features related to the leg movement of the person is enough. Since, this subset is not known apriori, we tend to use all the features, irrespective of the complexity of the task at hand. Selecting task-aware features may not only improve the efficiency but also the accuracy of the system. In this work, we propose a supervised approach for task-aware selection of features using Deep Neural Networks (DNN) in the context of action recognition. The activation potentials contributed by each of the individual input dimensions at the first hidden layer are used for selecting the most appropriate features. The selected features are found to give better classification performance than the original high-dimensional features. It is also shown that the classification performance of the proposed feature selection technique is superior to the low-dimensional representation obtained by principal component analysis (PCA).
Debaditya Roy, K. Sri Rama Murty, C. Krishna Mohan
IJCNN2
2015 Analysis of features from analytic representation of speech using MP-ABX measures
Raghavendra Reddy Pappagari, Karthika Vijayan, K. Sri Rama Murty
INTERSPEECH3
2015 Analysis of Phase Spectrum of Speech Signals Using Allpass Modeling
abstract
The phase spectrum of Fourier transform has received lesser prominence than its magnitude counterpart in speech processing. In this paper, we propose a method for parametric modeling of the phase spectrum, and discuss its applications in speech signal processing. The phase spectrum is modeled as the response of an allpass (AP) filter, whose coefficients are estimated from the knowledge of speech production process, especially the impulse-like nature of excitation source. A signal retaining only the phase spectral component of speech signal is derived by suppressing the magnitude spectral component, and is modeled as the output of an AP filter excited with a sequence of impulses. Entropy of energy of the input signal is minimized to estimate the coefficients of the AP filter. The resulting objective function, being nonconvex in nature, is minimized using particle swarm optimization. The group delay response of estimated AP filters can be used for accurate analysis of resonances of the vocal-tract system (VTS). The error signal associated with AP modeling provides unambiguous evidence about the instants of significant excitation of the VTS. The applications of the proposed AP modeling include, but not limited to, formant tracking, extraction of glottal closure instants, speaker verification and speech synthesis.
Karthika Vijayan, K. Sri Rama Murty
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Epoch extraction from allpass residual of speech signals
abstract
Identification of epochs from speech signals is a prominent task in speech processing. In this paper, epoch extraction is attempted from phase spectrum of speech signals. The phase spectrum of speech is modelled as an allpass (AP) filter by minimizing entropy of energy in the associated error signal. The AP residual thus obtained contains prominent unambiguous peaks at epoch locations. These peaks in AP residual constitute a set of candidate epoch locations from which appropriate ones are identified using a dynamic programming algorithm. The proposed method is evaluated on a subset of CMU Arctic database and it is observed that it delivered better epoch extraction performance than the prominent speech events estimation method-DYPSA. In case of telephone channel speech, the proposed method significantly outperformed zero frequency resonator based method also.
Karthika Vijayan, K. Sri Rama Murty
ICASSP2
2014 Novel speech duration modifier for packet based communication system
abstract
In this paper, we propose a real-time method for duration modification of speech for packet based communication system. While there is rich literature available on duration modification, it fails to clearly address the issues in real-time implementation of the same. Most of the duration modification methods rely on accurate estimation of pitch marks, which is not feasible in a real-time scenario. The proposed method modifies the duration of Linear Prediction residual of individual frames without using any look-ahead delay and knowledge of pitch marks. In this method, multiples of pitch period is repeated or removed from a frame depending on a scheduling algorithm. The subjective quality of the proposed method was found to be better than waveform similarity overlap and add (WSOLA) technique as well as Linear Prediction Pitch Synchronous Overlap and Add (LP-PSOLA) technique.
Senthil Kumar Mani, Jitendra Kumar Dhiman, K. Sri Rama Murty
INTERSPEECH3
2014 Unsupervised spoken word retrieval using Gaussian-bernoulli restricted boltzmann machines
abstract
The objective of this work is to explore a novel unsupervised framework, using Restricted Boltzmann machines, for Spoken Word Retrieval (SWR). In the absence of labelled speech data, SWR is typically performed by matching sequence of feature vectors of query and test utterances using dynamic time warping (DTW). In such a scenario, performance of SWR system critically depends on representation of the speech signal. Typical features, like mel-frequency cepstral coefficients (MFCC), carry significant speaker-specific information, and hence may not be used directly in SWR system. To overcome this issue, we propose to capture the joint density of the acoustic space spanned by MFCCs using Gaussian-Bernoulli restricted Boltzmann machine (GBRBM). In this work, we have used hidden activations of the GBRBM as features for SWR system. Since the GBRBM is trained with speech data collected from large number of speakers, the hidden activations are more robust to the speaker-variability compared to MFCCs. The performance of the proposed features is evaluated on Telugu broadcast news data, and an absolute improvement of 12% was observed compared to MFCCs.
Raghavendra Reddy Pappagari, Shekhar Nayak, K. Sri Rama Murty
INTERSPEECH3
2014 Feature extraction from analytic phase of speech signals for speaker verification
abstract
The objective of this work is to study the speaker-specific nature of analytic phase of speech signals. Since computation of analytic phase suffers from phase wrapping problem, we have used its derivative- The instantaneous frequency for feature extraction. The cepstral coefficients extracted from smoothed subband instantaneous frequencies (IFCC) are used as features for speaker verification. The performance of IFCC features is evaluated on NIST-2003 speaker recognition evaluation database and is compared with baseline mel-frequency cepstral coefficients (MFCC). The performance of IFCC features is observed to be comparable with MFCC features in terms of equal error rates and minimum detection cost function values. Different strategies for evaluating the speaker verification performance of IFCC and MFCC are explored and it is found that the evaluation based on cosine similarity delivers better performance than other strategies under consideration.
Karthika Vijayan, K. Sri Rama Murty
INTERSPEECH3
2012 Speaker recognition via sparse representations using orthogonal matching pursuit
abstract
The objective of this paper is to demonstrate the effectiveness of sparse representation techniques for speaker recognition. In this approach, each feature vector from unknown utterance is expressed as linear weighted sum of a dictionary of feature vectors belonging to many speakers. The weights associated with feature vectors in the dictionary are evaluated using orthogonal matching pursuit algorithm, which is a greedy approximation to l0 optimization. The weights thus obtained exhibit high level of sparsity, and only a few of them will have nonzero values. The feature vectors which belong to the correct speaker carry significant weights. The proposed method gives an equal error rate (EER) of 10.84% on NIST-2003 database, whereas the existing GMM-UBM system gives an EER of 9.67%. By combining evidence from both the systems an EER of 8.15% is achieved, indicating that both the systems carry complimentary information.
Vivek Boominathan, K. Sri Rama Murty
ICASSP2
2009 Analysis of laugh signals for detecting in continuous speech
abstract
Laughter is a nonverbal vocalization that occurs often in speech communication. Since laughter is produced by the speech production mechanism, spectral analysis methods are used mostly for the study of laughter acoustics. In this paper the significance of excitation features for discriminating laughter and speech is discussed. New features describing the excitation characteristics are used to analyze the laugh signals. The features are based on instantaneous pitch and strength of excitation at epochs. An algorithm is developed based on these features to detect laughter regions in continuous speech. The results are illustrated by detecting laughter regions in a TV broadcast program. Index Terms: Laughter detection, epoch, strength of excitation
K. Sudheer Kumar, Sri Harish Reddy Mallidi, K. Sri Rama Murty, Bayya Yegnanarayana
INTERSPEECH3
2009 Characterization of Glottal Activity From Speech Signals
abstract
The objective of this work is to characterize certain important features of excitation of speech, namely, detecting the regions of glottal activity and estimating the strength of excitation in each glottal cycle. The proposed method is based on the assumption that the excitation to the vocal-tract system can be approximated by a sequence of impulses of varying strengths. The effect due to an impulse in the time-domain is spread uniformly across the frequency-domain including at zero-frequency. We propose the use of a zero-frequency resonator to extract the characteristics of excitation source from speech signals by filtering out most of the time-varying vocal-tract information. The regions of glottal activity and the strengths of excitation estimated from the speech signal are in close agreement with those observed from the simultaneously recorded electro-glotto-graph signals. The performance of the proposed glottal activity detection is evaluated under different noisy environments at varying levels of degradation.
K. Sri Rama Murty, Bayya Yegnanarayana, Joseph M. Anand
IEEE Signal Process. Lett.1
2009 Event-Based Instantaneous Fundamental Frequency Estimation From Speech Signals
abstract
Exploiting the impulse-like nature of excitation in the sequence of glottal cycles, a method is proposed to derive the instantaneous fundamental frequency from speech signals. The method involves passing the speech signal through two ideal resonators located at zero frequency. A filtered signal is derived from the output of the resonators by subtracting the local mean computed over an interval corresponding to the average pitch period. The positive zero crossings in the filtered signal correspond to the locations of the strong impulses in each glottal cycle. Then the instantaneous fundamental frequency is obtained by taking the reciprocal of the interval between successive positive zero crossings. Due to filtering by zero-frequency resonator, the effects of noise and vocal-tract variations are practically eliminated. For the same reason, the method is also robust to degradation in speech due to additive noise. The accuracy of the fundamental frequency estimation by the proposed method is comparable or even better than many existing methods. Moreover, the proposed method is also robust against rapid variation of the pitch period or vocal-tract changes. The method works well even when the glottal cycles are not periodic or when the speech signals are not correlated in successive glottal cycles.
Bayya Yegnanarayana, K. Sri Rama Murty
IEEE Trans. Speech Audio Process.2
2009 Determining Mixing Parameters From Multispeaker Data Using Speech-Specific Information
abstract
In this paper, we propose an approach for processing multispeaker speech signals collected simultaneously using a pair of spatially separated microphones in a real room environment. Spatial separation of microphones results in a fixed time-delay of arrival of speech signals from a given speaker at the pair of microphones. These time-delays are estimated by exploiting the impulse-like characteristic of excitation during speech production. The differences in the time-delays for different speakers are used to determine the number of speakers from the mixed multispeaker speech signals. There is difference in the signal levels due to differences in the distances between the speaker and each of the microphones. The differences in the signal levels dictate the values of the mixing parameters. Knowledge of speech production, especially the excitation source characteristics, is used to derive an approximate weight function for locating the regions specific to a given speaker. The scatter plots of the weighted and delay-compensated mixed speech signals are used to estimate the mixing parameters. The proposed method is applied on the data collected in actual laboratory environment for an underdetermined case, where the number of speakers is more than the number of microphones. Enhancement of speech due to a speaker is also examined using the information of the time-delays and the mixing parameters, and is evaluated using objective measures proposed in the literature.
Bayya Yegnanarayana, R. Kumaraswamy 0001, K. Sri Rama Murty
IEEE Trans. Speech Audio Process.3
2008 Efficient representation of throat microphone speech
abstract
The objective of this work is to represent the information in the speech signal picked up by a throat microphone (TM) in an efficient manner in terms of number of bits required.Since the TM signal is unaffected by ambient noise, it is possible to extract the required information effectively under different environmental conditions.A spectral mapping technique is proposed from the TM speech to normal microphone (NM) speech to improve the perceptual quality.The mapping is done using vector quantization of pairwise spectral feature vectors derived from each frame of TM and the corresponding NM speech signals.Once the codebook is formed, the spectral features from a TM signal are represented as a sequence of codebook indices.The sequence of codebook indices, the pitch contour and the energy contour derived from the TM signal are used to store/transmit the TM speech information efficiently.From the received sequence of codebook indices, the NM spectral vectors are retrieved due to pairwise vector quantization of the feature vectors.A synthetic residual signal is generated at the receiver from prestored residual templates by incorporating the pitch and the energy.The synthetic residual signal is used to excite the system corresponding to the NM spectral vectors to generate the speech signal.
K. Sri Rama Murty, Saurav Khurana, Yogendra Umesh Itankar, M. R. Kesheorey, Bayya Yegnanarayana
INTERSPEECH1
2008 Epoch Extraction From Speech Signals
abstract
Epoch is the instant of significant excitation of the vocal-tract system during production of speech. For most voiced speech, the most significant excitation takes place around the instant of glottal closure. Extraction of epochs from speech is a challenging task due to time-varying characteristics of the source and the system. Most epoch extraction methods attempt to remove the characteristics of the vocal-tract system, in order to emphasize the excitation characteristics in the residual. The performance of such methods depends critically on our ability to model the system. In this paper, we propose a method for epoch extraction which does not depend critically on characteristics of the time-varying vocal-tract system. The method exploits the nature of impulse-like excitation. The proposed zero resonance frequency filter output brings out the epoch locations with high accuracy and reliability. The performance of the method is demonstrated using CMU-Arctic database using the epoch information from the electroglottograph as reference. The proposed method performs significantly better than the other methods currently available for epoch extraction. The interesting part of the results is that the epoch extraction by the proposed method seems to be robust against degradations like white noise, babble, high-frequency channel, and vehicle noise.
K. Sri Rama Murty, Bayya Yegnanarayana
IEEE Trans. Speech Audio Process.1
2007 Detection of instants of glottal closure using characteristics of excitation source
abstract
In this paper, we propose a method for detection of glottal clo-sure instants (GCI) in the voiced regions of speech signals. The method is based on periodicity of significant excitations of the vocal tract system. The key idea is the computation of coherent covariance sequence, which overcomes the effect of dynamic range of the excitation source signal, while preserv-ing the locations of significant excitations. The Hilbert enve-lope of linear prediction residual is used as an estimate of the source of excitation of the vocal tract system. Performance of the proposed method is evaluated in terms of the deviation be-tween true GCIs and hypothesized GCIs, using clean speech and degraded speech signals. The signal-to-noise ratio (SNR) of speech signals in the vicinity of GCIs has significant bear-ing on the performance of the proposed method. The proposed method is accurate and robust for detection of GCIs, even in the presence of degradations. Index Terms: glottal closure instants, excitation source, peri-odicity, coherent covariance sequence
Sunitha Guruprasad, Bayya Yegnanarayana, K. Sri Rama Murty
INTERSPEECH3
2007 Voice activity detection in degraded speech using excitation source information
abstract
This paper proposes a method for detection of voiced regions from speech signals collected in noisy environment.The proposed method is based on the characteristics of excitation source of speech production.The degraded speech signal is processed by linear prediction analysis for deriving the linear prediction residual.Hilbert envelope of the linear prediction residual is processed using covariance analysis to obtain coherentlyadded covariance signal.The periodicity property of the coherently added covariance signal is exploited to detect the voiced regions using autocorrelation analysis.The performance of the proposed voice activity detection algorithm is evaluated under different noise environments and at different levels of degradation.
K. Sri Rama Murty, Bayya Yegnanarayana, Sunitha Guruprasad
INTERSPEECH1
2007 Determining Number of Speakers From Multispeaker Speech Signals Using Excitation Source Information
abstract
In this letter, we address the issue of determining the number of speakers from multispeaker speech signals collected simultaneously using a pair of spatially separated microphones. The spatial separation of the microphones results in time delay of arrival of speech signals from a given speaker. The differences in the time delays for different speakers are exploited to determine the number of speakers from the multispeaker signals. The key idea is that for a given speaker, the relative spacings of the instants of significant excitation of the vocal tract system remain unchanged in the direct components of the speech signals at the two microphones. The time delays can be estimated from the cross-correlation of the Hilbert envelopes of the linear prediction residuals of the multispeaker signals collected at the two microphones.
R. Kumaraswamy 0001, K. Sri Rama Murty, Bayya Yegnanarayana
IEEE Signal Process. Lett.2
2006 Combining evidence from residual phase and MFCC features for speaker recognition
abstract
The objective of this letter is to demonstrate the complementary nature of speaker-specific information present in the residual phase in comparison with the information present in the conventional mel-frequency cepstral coefficients (MFCCs). The residual phase is derived from speech signal by linear prediction analysis. Speaker recognition studies are conducted on the NIST-2003 database using the proposed residual phase and the existing MFCC features. The speaker recognition system based on the residual phase gives an equal error rate (EER) of 22%, and the system using the MFCC features gives an EER of 14%. By combining the evidence from both the residual phase and the MFCC features, an EER of 10.5% is obtained, indicating that speaker-specific excitation information is present in the residual phase. This information is useful since it is complementary to that of MFCCs.
K. Sri Rama Murty, Bayya Yegnanarayana
IEEE Signal Process. Lett.1