EDBT 2026 Demo / reviewers in the wild / expert
Hema A. Murthy
dblp:38/5502
· DBLP profile ↗
89ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0003-3611-6550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 51 · 3 first-author · 9 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 13 |
| 2025 | LIMMITS'25: Multilingual Streaming TTS With Neural Codecs for Indian LanguagesabstractThis work provides a summary of the Multilingual streaming TTS with neural codecs for Indian languages challenge (LIMMITS’25), organized as part of the ICASSP 2025 signal processing grand challenge. Towards this, 278 hours of TTS data in 4 Indian languages - Gujarati, Indian English, Bhojpuri, and Kannada got released. The challenge focuses on advancing research in neural codec-based and streaming TTS systems. The top teams in the challenge attained high subjective scores on naturalness and similarity, thus contributing to the progress in text-to-speech generation systems. Philipp Olbrich, Hema A. Murthy, Pranaw Kumar, Shinji Watanabe 0001, Sheng Zhao 0002, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2025 | Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages
Utkarsh Pathak, Chandra Sai Krishna Gunda, Anusha Prakash 0001, Keshav Agarwal, Hema A. Murthy |
INTERSPEECH | 5 |
| 2025 | Enhancing Syllabic Recognition via Speech-EEG Phase Analysis and Non-Activity State Modeling
Rini A. Sharon, Hema A. Murthy |
INTERSPEECH | 2 |
| 2023 | Towards Developing State-of-The-Art TTS Synthesisers for 13 Indian Languages with Signal Processing Aided AlignmentsabstractEnd-to-end (E2E) systems synthesise high-quality speech, but this typically requires a large amount of data. As E2E synthesis progressed from Tacotron to FastSpeech2, it became evident that features representing prosody, particularly subword durations, are important for error-free synthesis. Variants of FastSpeech use a teacher model or forced alignments for training. This paper uses signal processing cues in tandem with forced alignment to produce accurate phone boundaries for the training data. As a result of better duration modelling, good-quality synthesisers are developed. Evaluations indicate that systems developed using the proposed signal processing-aided approach are better than systems developed using other alignment approaches, especially in low-resource scenarios. Our systems also outperform the existing best TTS systems available for 13 Indian languages. Anusha Prakash 0001, Srinivasan Umesh, Hema A. Murthy |
ASRU | 3 |
| 2023 | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-SpeechabstractThe Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research. Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
ICASSP | 9 |
| 2023 | Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages
Anusha Prakash 0001, Arun Kumar A, Ashish Seth, Bhagyashree Mukherjee, Ishika Gupta, Jom Kuriakose, Jordan Fernandes, K. V. Vikram, Mano Ranjith Kumar, Narla John Metilda Sagaya Mary, Mohammad Wajahat, Mohana N, Mudit Batra, Navina K, Nihal John George, Nithya Ravi, Pruthwik Mishra, Sudhanshu Srivastava 0001, Vasista Sai Lodagala, Vandan Mujadia, Kada Sai Venkata Vineeth, Vrunda N. Sukhadia, Dipti Misra Sharma, Hema A. Murthy, Pushpak Bhattacharyya, Srinivasan Umesh, Rajeev Sangal |
INTERSPEECH | 24 |
| 2023 | Exploring the Role of Language Families for Building Indic Speech SynthesisersabstractBuilding end-to-end speech synthesisers for Indian languages is challenging, given the lack of adequate clean training data and multiple grapheme representations across languages. This work explores the importance of training multilingual and multi-speaker text-to-speech (TTS) systems based onlanguage families. The objective is to exploit the phonotactic properties of language families, where small amounts of accurately transcribed data across languages can be pooled together to train TTS systems. These systems can then be adapted to new languages belonging to the same family in extremely low-resource scenarios. TTS systems are trained separately for Indo-Aryan and Dravidian language families, and their performance is compared to that of a combined Indo-Aryan+Dravidian voice. We also investigate the amount of training data required for a language in a multilingual setting. Same-family and cross-family synthesis and adaptation to unseen languages are analysed. The analyses show that language family-wise training of Indic systems is the way forward for the Indian subcontinent, where a large number of languages are spoken. Anusha Prakash 0001, Hema A. Murthy |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Analysis of Conversational Speech with Application to Voice AdaptationabstractConversational speech has always been challenging in the context of text-to-speech synthesis (TTS). Most speech synthesis systems are trained on read speech data recorded in a studio environment. But, the intelligibility of TTS systems degrades drastically when using conversational speech. The proposed work attempts to perform extensive analysis on the issues in dealing with conversational speech compared to read speech. As an application, we try to dub the lectures available in English into an Indian language (Hindi) in the original speaker's voice. The task is difficult as classroom lectures are extempore, with variations in speaking rate, and contain speaker mannerisms that lead to disfluencies. We analyze the capability of end-to-end TTS systems in modeling lecture-based data. Based on the analysis, an attempt is made to adapt “read speech TTS system” using conversational speech data to produce lectures in the original speaker's voice. Bhagyashree Mukherjee, Anusha Prakash 0001, Hema A. Murthy |
ASRU | 3 |
| 2021 | Dual Script E2E Framework for Multilingual and Code-Switching ASRabstractIndia is home to multiple languages, and training automatic speech recognition (ASR) systems for languages is challenging. Over time, each language has adopted words from other languages, such as English, leading to code-mixing. Most Indian languages also have their own unique scripts, which poses a major limitation in training multilingual and code-switching ASR systems. Inspired by results in text-to-speech synthesis, in this work, we use an in-house rule-based phoneme-level common label set (CLS) representation to train multilingual and code-switching ASR for Indian languages. We propose two end-to-end (E2E) ASR systems. In the first system, the E2E model is trained on the CLS representation, and we use a novel data-driven back-end to recover the native language script. In the second system, we propose a modification to the E2E model, wherein the CLS representation and the native language characters are used simultaneously for training. We show our results on the multilingual and code-switching tasks of the Indic ASR Challenge 2021. Our best results achieve 6% and 5% improvement (approx) in word error rate over the baseline system for the multilingual and code-switching tasks, respectively, on the challenge development data. Mari Ganesh Kumar, Jom Kuriakose, Anand Thyagachandran, Arun Kumar A, Ashish Seth, Lodagala Durga Prasad, Saish Jaiswal, Anusha Prakash 0001, Hema A. Murthy |
Interspeech | 9 |
| 2021 | Towards Zero-Shot Learning with Fewer Seen Class ExamplesabstractWe present a meta-learning based generative model for zero-shot learning (ZSL) towards a challenging setting when the number of training examples from each seen class is very few. This setup contrasts with the conventional ZSL approaches, where training typically assumes the availability of a sufficiently large number of training examples from each of the seen classes. The proposed approach leverages meta-learning to train a deep generative model that integrates variational autoencoder and generative adversarial networks. We propose a novel task distribution where meta-train and meta-validation classes are disjoint to simulate the ZSL behaviour in training. Once trained, the model can generate synthetic examples from seen and unseen classes. Synthesize samples can then be used to train the ZSL framework in a supervised manner. The meta-learner enables our model to generates high-fidelity samples using only a small number of training examples from seen classes. We conduct extensive experiments and ablation studies on four benchmark datasets of ZSL and observe that the proposed model outperforms state-of-the-art approaches by a significant margin when the number of examples per seen class is very small. Vinay Kumar Verma, Ashish Mishra 0001, Anubha Pandey, Hema A. Murthy, Piyush Rai |
WACV | 4 |
| 2021 | Functional parcellation of mouse visual cortex using statistical techniques reveals response-dependent clustering of cortical processing areasabstractThe visual cortex of the mouse brain can be divided into ten or more areas that each contain complete or partial retinotopic maps of the contralateral visual field. It is generally assumed that these areas represent discrete processing regions. In contrast to the conventional input-output characterizations of neuronal responses to standard visual stimuli, here we asked whether six of the core visual areas have responses that are functionally distinct from each other for a given visual stimulus set, by applying machine learning techniques to distinguish the areas based on their activity patterns. Visual areas defined by retinotopic mapping were examined using supervised classifiers applied to responses elicited by a range of stimuli. Using two distinct datasets obtained using wide-field and two-photon imaging, we show that the area labels predicted by the classifiers were highly consistent with the labels obtained using retinotopy. Furthermore, the classifiers were able to model the boundaries of visual areas using resting state cortical responses obtained without any overt stimulus, in both datasets. With the wide-field dataset, clustering neuronal responses using a constrained semi-supervised classifier showed graceful degradation of accuracy. The results suggest that responses from visual cortical areas can be classified effectively using data-driven models. These responses likely reflect unique circuits within each area that give rise to activity with stronger intra-areal than inter-areal correlations, and their responses to controlled visual stimuli across trials drive higher areal classification accuracy than resting state responses. Mari Ganesh Kumar, Aadhirai Ramanujan, Mriganka Sur, Hema A. Murthy |
PLoS Comput. Biol. | 5 |
| 2021 | Signal-to-signal neural networks for improved spike estimation from calcium imaging dataabstractSpiking information of individual neurons is essential for functional and behavioral analysis in neuroscience research. Calcium imaging techniques are generally employed to obtain activities of neuronal populations. However, these techniques result in slowly-varying fluorescence signals with low temporal resolution. Estimating the temporal positions of the neuronal action potentials from these signals is a challenging problem. In the literature, several generative model-based and data-driven algorithms have been studied with varied levels of success. This article proposes a neural network-based signal-to-signal conversion approach, where it takes as input raw-fluorescence signal and learns to estimate the spike information in an end-to-end fashion. Theoretically, the proposed approach formulates the spike estimation as a single channel source separation problem with unknown mixing conditions. The source corresponding to the action potentials at a lower resolution is estimated at the output. Experimental studies on the spikefinder challenge dataset show that the proposed signal-to-signal conversion approach significantly outperforms state-of-the-art-methods in terms of Pearson's correlation coefficient, Spearman's rank correlation coefficient and yields comparable performance for the area under the receiver operating characteristics measure. We also show that the resulting system: (a) has low complexity with respect to existing supervised approaches and is reproducible; (b) is layer-wise interpretable, and (c) has the capability to generalize across different calcium indicators. Jilt Sebastian, Mriganka Sur, Hema A. Murthy, Mathew Magimai-Doss |
PLoS Comput. Biol. | 3 |
| 2021 | Novel Architectures for Unsupervised Information Bottleneck Based Speaker Diarization of MeetingsabstractSpeaker diarization is an important problem that is topical, and is especially useful as a preprocessor for conversational speech related applications. The objective of this article is two-fold: (i) segment initialization by uniformly distributing speaker information across the initial segments, and (ii) incorporating speaker discriminative features within the unsupervised diarization framework. In the first part of the work, a varying length segment initialization technique for Information Bottleneck (IB) based speaker diarization system using phoneme rate as the side information is proposed. This initialization distributes speaker information uniformly across the segments and provides a better starting point for IB based clustering. In the second part of the work, we present a Two-Pass Information Bottleneck (TPIB) based speaker diarization system that incorporates speaker discriminative features during the process of diarization. The TPIB based speaker diarization system has shown improvement over the baseline IB based system. During the first pass of the TPIB system, a coarse segmentation is performed using IB based clustering. The alignments obtained are used to generate speaker discriminative features using a shallow feed-forward neural network and linear discriminant analysis. The discriminative features obtained are used in the second pass to obtain the final speaker boundaries. In the final part of the paper, variable segment initialization is combined with the TPIB framework. This leverages the advantages of better segment initialization and speaker discriminative features that results in an additional improvement in performance. An evaluation on standard meeting datasets shows that a significant absolute improvement of 3.9% and 4.7% is obtained on the NIST and AMI datasets, respectively. Nauman Dawalatabad, Srikanth R. Madikeri, Chellu Chandra Sekhar, Hema A. Murthy |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Evidence of Task-Independent Person-Specific Signatures in EEG Using Subspace TechniquesabstractElectroencephalography (EEG) signals are promising as alternatives to other biometrics owing to their protection against spoofing. Previous studies have focused on capturing individual variability by analyzing task/condition-specific EEG. This work attempts to model biometric signatures independent of task/condition by normalizing the associated variance. Toward this goal, the paper extends ideas from subspace-based text-independent speaker recognition and proposes novel modifications for modeling multi-channel EEG data. The proposed techniques assume that biometric information is present in the entire EEG signal and accumulate statistics across time in a high dimensional space. These high dimensional statistics are then projected to a lower dimensional space where the biometric information is preserved. The lower dimensional embeddings obtained using the proposed approach are shown to be task-independent. The best subspace system identifies individuals with accuracies of 86.4% and 35.9% on datasets with 30 and 920 subjects, respectively, using just nine EEG channels. The paper also provides insights into the subspace model's scalability to unseen tasks and individuals during training and the number of channels needed for subspace modeling. Mari Ganesh Kumar, Shri Narayanan, Mriganka Sur, Hema A. Murthy |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | State-Based Transcription of Components of Carnatic MusicabstractAutomatic Carnatic Music (CM) transcription is an open problem in need of a standardized descriptive notation. The level of detail needed in a descriptive transcription makes it tedious to obtain ground truth by manual means. In this paper, we propose a novel state-based representation of the pitch curve motivated by CM components called constant-pitch notes and stationary points. We also propose a novel transcription technique that uses the Viterbi algorithm to estimate the states and quantized pitch-values. The proposed technique adheres best to raga-notes compared to the existing critical-points technique and uniform quantization. In a listening test, clips synthesized from the proposed notation were rated significantly better (324 ratings, p <; 0.001) than those from critical-points. Further, speed-halving based on state information best matches the actual, observed CM component-duration ratios without losing raga-characteristics. Thus, the proposed transcription can be corrected manually to obtain ground truth and can enhance learning tools. Venkata Subramanian Viraraghavan, Arpan Pal 0001, Hema A. Murthy, Rangarajan Aravind |
ICASSP | 3 |
| 2020 | A Hybrid HMM-Waveglow Based Text-to-Speech Synthesizer Using Histogram Equalization for Low Resource Indian Languages
Mano Ranjith Kumar, Sudhanshu Srivastava 0001, Anusha Prakash 0001, Hema A. Murthy |
INTERSPEECH | 4 |
| 2020 | Generic Indic Text-to-Speech Synthesisers with Rapid Adaptation in an End-to-End FrameworkabstractBuilding text-to-speech (TTS) synthesisers for Indian languages is a difficult task owing to a large number of active languages. Indian languages can be classified into a finite set of families, prominent among them, Indo-Aryan and Dravidian. The proposed work exploits this property to build a generic TTS system using multiple languages from the same family in an end-to-end framework. Generic systems are quite robust as they are capable of capturing a variety of phonotactics across languages. These systems are then adapted to a new language in the same family using small amounts of adaptation data. Experiments indicate that good quality TTS systems can be built using only 7 minutes of adaptation data. An average degradation mean opinion score of 3.98 is obtained for the adapted TTSes. Extensive analysis of systematic interactions between languages in the generic TTSes is carried out. x-vectors are included as speaker embedding to synthesise text in a particular speaker's voice. An interesting observation is that the prosody of the target speaker's voice is preserved. These results are quite promising as they indicate the capability of generic TTSes to handle speaker and language switching seamlessly, along with the ease of adaptation to a new language. Anusha Prakash 0001, Hema A. Murthy |
INTERSPEECH | 2 |
| 2020 | Exploration of End-to-End Synthesisers for Zero Resource Speech Challenge 2020abstractA Spoken dialogue system for an unseen language is referred to as Zero resource speech. It is especially beneficial for developing applications for languages that have low digital resources. Zero resource speech synthesis is the task of building text-to-speech (TTS) models in the absence of transcriptions. In this work, speech is modelled as a sequence of transient and steady-state acoustic units, and a unique set of acoustic units is discovered by iterative training. Using the acoustic unit sequence, TTS models are trained. The main goal of this work is to improve the synthesis quality of zero resource TTS system. Four different systems are proposed. All the systems consist of three stages: unit discovery, followed by unit sequence to spectrogram mapping, and finally spectrogram to speech inversion. Modifications are proposed to the spectrogram mapping stage. These modifications include training the mapping on voice data, using x-vectors to improve the mapping, two-stage learning, and gender-specific modelling. Evaluation of the proposed systems in the Zerospeech 2020 challenge shows that quite good quality synthesis can be achieved. D. S. Karthik Pandia, Anusha Prakash 0001, Mano Ranjith Kumar, Hema A. Murthy |
INTERSPEECH | 4 |
| 2020 | The "Sound of Silence" in EEG - Cognitive Voice Activity DetectionabstractSpeech cognition bears potential application as a brain computer interface that can improve the quality of life for the otherwise communication impaired people. While speech and resting state EEG are popularly studied, here we attempt to explore a "non-speech"(NS) state of brain activity corresponding to the silence regions of speech audio. Firstly, speech perception is studied to inspect the existence of such a state, followed by its identification in speech imagination. Analogous to how voice activity detection is employed to enhance the performance of speech recognition, the EEG state activity detection protocol implemented here is applied to boost the confidence of imagined speech EEG decoding. Classification of speech and NS state is done using two datasets collected from laboratory-based and commercial-based devices. The state sequential information thus obtained is further utilized to reduce the search space of imagined EEG unit recognition. Temporal signal structures and topographic maps of NS states are visualized across subjects and sessions. The recognition performance and the visual distinction observed demonstrates the existence of silence signatures in EEG. Rini A. Sharon, Hema A. Murthy |
INTERSPEECH | 2 |
| 2020 | Stacked Adversarial Network for Zero-Shot Sketch based Image RetrievalabstractConventional approaches to Sketch-Based Image Retrieval (SBIR) assume that the data of all the classes are available during training. The assumption may not always be practical since the data of a few classes may be unavailable, or the classes may not appear at the time of training. Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) relaxes this constraint and allows the algorithm to handle previously unseen classes during the test. This paper proposes a generative approach based on the Stacked Adversarial Network (SAN) and the advantage of Siamese Network (SN) for ZS-SBIR. While SAN generates a high-quality sample, SN learns a better distance metric compared to that of the nearest neighbor search. The capability of the generative model to synthesize image features based on the sketch reduces the SBIR problem to that of an image-to-image retrieval problem. We evaluate the efficacy of our proposed approach on TU-Berlin, and Sketchy database in both standard ZSL and generalized ZSL setting. The proposed method yields a significant improvement in standard ZSL as well as in a more challenging generalized ZSL setting (GZSL) for SBIR. Anubha Pandey, Ashish Mishra 0001, Vinay Kumar Verma, Anurag Mittal, Hema A. Murthy |
WACV | 5 |
| 2020 | Zero-shot learning for action recognition using synthesized features
Ashish Mishra 0001, Anubha Pandey, Hema A. Murthy |
Neurocomputing | 3 |
| 2020 | Significance of spectral cues in automatic speech segmentation for Indian language speech synthesizers
Arun Baby, Jeena J. Prakash, Aswin Shanmugam Subramanian, Hema A. Murthy |
Speech Commun. | 4 |
| 2020 | Importance of Signal Processing Cues in Transcription Correction for Low-Resource Indian LanguagesabstractAccurate phonetic transcriptions are crucial for building robust acoustic models for speech recognition as well as speech synthesis applications. Phonetic transcriptions are not usually provided with speech corpora. A lexicon is used to generate phone-level transcriptions of speech corpora with sentence-level transcriptions. When lexical entries are not available, letter-to-sound (LTS) rules are used. Whether it is a lexicon or LTS, the rules for pronunciation are generic and may not match the spoken utterance. This can lead to transcription errors. The objective of this study is to address the issue of mismatch between the transcription and its acoustic realisation. In particular, the issue of vowel deletions is studied. Group-delay-based segmentation is used to determine insertion/deletion of vowels in the speech utterance. The transcriptions are corrected in the training data based on this. The corrected data are used in automatic speech recognition (ASR) and text to speech synthesis (TTS) systems. ASR and TTS systems built with the corrected transcriptions show improvements in the performance. Jeena J. Prakash, Rajan Golda Brunet, Hema A. Murthy |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Spoof Detection Using Time-Delay Shallow Neural Network and Feature SwitchingabstractDetecting spoofed utterances is a fundamental problem in voice-based biometrics. Spoofing can be performed either by logical accesses like speech synthesis, voice conversion or by physical accesses such as replaying the pre-recorded utterance. Inspired by the state-of-the-art x-vector based speaker verification approach, this paper proposes a time-delay shallow neural network (TD-SNN) for spoof detection for both logical and physical access. The novelty of the proposed TD-SNN system vis-a-vis conventional DNN systems is that it can handle variable length utterances during testing. Performance of the proposed TD-SNN systems and the baseline Gaussian mixture models (GMMs) is analyzed on the ASV-spoof-2019 dataset. The performance of the systems is measured in terms of the minimum normalized tandem detection cost function (min-t-DCF). When studied with individual features, the TD-SNN system consistently outperforms the GMM system for physical access. For logical access, GMM surpasses TD-SNN systems for certain individual features. When combined with the decision-level feature switching (DLFS) paradigm, the best TD-SNN system outperforms the best baseline GMM system on evaluation data with a relative improvement of 48.03% and 49.47% for both logical and physical access, respectively. Mari Ganesh Kumar, Suvidha Rupesh Kumar, M. S. Saranya, B. Bharathi 0001, Hema A. Murthy |
ASRU | 5 |
| 2019 | Incremental Transfer Learning in Two-pass Information Bottleneck Based Speaker Diarization System for MeetingsabstractThe two-pass information bottleneck (TPIB) based speaker diarization system operates independently on different conversational recordings. TPIB system does not consider previously learned speaker discriminative information while di-arizing new conversations. Hence, the real time factor (RTF) of TPIB system is high owing to the training time required for the artificial neural network (ANN). This paper attempts to improve the RTF of the TPIB system using an incremental transfer learning approach where the parameters learned by the ANN from other conversations are updated using current conversation rather than learning parameters from scratch. This reduces the RTF significantly. The effectiveness of the proposed approach compared to the baseline IB and the TPIB systems is demonstrated on standard NIST and AMI conversational meeting datasets. With a minor degradation in performance, the proposed system shows a significant improvement of 33.07% and 24.45% in RTF with respect to TPIB system on the NIST RT-04Eval and AMI-1 datasets, respectively. Nauman Dawalatabad, Srikanth R. Madikeri, Chellu Chandra Sekhar, Hema A. Murthy |
ICASSP | 4 |
| 2019 | An Empirical Study of Speech Processing in the Brain by Analyzing the Temporal Syllable Structure in Speech-input Induced EEGabstractClinical applicability of electroencephalography (EEG) is well established, however the use of EEG as a choice for constructing brain computer interfaces to develop communication platforms is relatively recent. To provide more natural means of communication, there is an increasing focus on bringing together speech and EEG signal processing. Quantifying the way our brain processes speech is one way of approaching the problem of speech recognition using brain waves. This paper analyses the feasibility of recognizing syllable level units by studying the temporal structure of speech reflected in the EEG signals. The slowly varying component of the delta band EEG(0.3-3Hz) is present in all other EEG frequency bands. Analysis shows that removing the delta trend in EEG signals results in signals that reveals syllable like structure. Using a 25 syllable framework, classification of EEG data obtained from 13 subjects yields promising results, underscoring the potential of revealing speech related temporal structure in EEG. Rini A. Sharon, Shri Narayanan, Mriganka Sur, Hema A. Murthy |
ICASSP | 4 |
| 2019 | Zero Resource Speech Synthesis Using Transcripts Derived from Perceptual Acoustic UnitsabstractZerospeech synthesis is the task of building vocabulary independent speech synthesis systems, where transcriptions are not available for training data. It is, therefore, necessary to convert training data into a sequence of fundamental acoustic units that can be used for synthesis during the test. This paper attempts to discover, and model perceptual acoustic units consisting of steady-state, and transient regions in speech. The transients roughly correspond to CV, VC units, while the steady-state corresponds to sonorants and fricatives. The speech signal is first preprocessed by segmenting the same into CVC-like units using a short-term energy-like contour. These CVC segments are clustered using a connected components-based graph clustering technique. The clustered CVC segments are initialized such that the onset (CV) and decays (VC) correspond to transients, and the rhyme corresponds to steady-states. Following this initialization, the units are allowed to re-organise on the continuous speech into a final set of AUs in an HMM-GMM framework. AU sequences thus obtained are used to train synthesis models. The performance of the proposed approach is evaluated on the Zerospeech 2019 challenge database. Subjective and objective scores show that reasonably good quality synthesis with low bit rate encoding can be achieved using the proposed AUs. D. S. Karthik Pandia, Hema A. Murthy |
INTERSPEECH | 2 |
| 2019 | Analysis of Inter-Pausal Units in Indian Languages and Its Application to Text-to-Speech SynthesisabstractLack of punctuation in Indian language text makes the analysis of phrases difficult. In this paper, inter-pausal units (IPUs) in read sentences are considered as phrases and are analyzed. A key observation from this analysis is that the length of the IPUs in read sentences follow uniformly (across all languages) a Gamma distribution. Additionally, an analysis of the scale and shape parameters suggest that these parameters are governed by the location of the IPU in an utterance. This information is used in text-to-speech (TTS) systems for four Indian languages leading to an improvement in naturalness. A novel IPU-based TTS system is proposed for better prosody modeling as well. A given text is parsed into IPUs, and an appropriate TTS system for different IPUs is used for synthesis. It is observed that there is a significant improvement in the naturalness of synthesized speech compared to that of a single TTS system being used for the entire sentence. Jeena J. Prakash, Hema A. Murthy |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Information Bottleneck Based Percussion Instrument Diarization System for Taniavartanam Segments of Carnatic Music Concerts
Nauman Dawalatabad, Jom Kuriakose, Chellu Chandra Sekhar, Hema A. Murthy |
INTERSPEECH | 4 |
| 2018 | Mobile Application for Learning Languages for the Unlettered
Gayathri G, N. Mohana, Radhika Pal, Hema A. Murthy |
INTERSPEECH | 4 |
| 2018 | Early Vocabulary Development Through Picture-based Software Solutions
G. R. Kasthuri, Prabha Ramanathan, Hema A. Murthy, Namita Jacob, Anil Prabhakar |
INTERSPEECH | 3 |
| 2018 | Resyllabification in Indian Languages and Its Implications in Text-to-speech Systems
Mahesh M, Jeena J. Prakash, Hema A. Murthy |
INTERSPEECH | 3 |
| 2018 | Brain-Computer Interface using Electroencephalogram Signatures of Eye Blinks
Srihari Maruthachalam, Sidharth Aggarwal, Mari Ganesh Kumar, Mriganka Sur, Hema A. Murthy |
INTERSPEECH | 5 |
| 2018 | Transcription Correction for Indian Languages Using Acoustic Signatures
Jeena J. Prakash, Rajan Golda Brunet, Hema A. Murthy |
INTERSPEECH | 3 |
| 2018 | Decision-level Feature Switching as a Paradigm for Replay Attack Detection
M. S. Saranya, Hema A. Murthy |
INTERSPEECH | 2 |
| 2018 | Denoising and Raw-waveform Networks for Weakly-Supervised Gender Identification on Noisy SpeechabstractThis paper presents a raw-waveform neural network and uses it along with a denoising network for clustering in weakly supervised learning scenarios under extreme noise conditions. Specifically, we consider language independent Automatic Gender Recognition (AGR) on a set of varied noise conditions and Signal to Noise Ratios (SNRs). We formulate the denoising problem as a source separation task and train the system using a discriminative criterion in order to enhance output SNRs. A denoising Recurrent Neural Network (RNN) is first trained on a small subset (roughly one-fifth) of the data for learning a speech specific mask. The denoised speech signal is then directly fed as input to a raw-waveform convolutional neural network (CNN) trained with denoised speech. We evaluate the standalone performance of denoiser in terms of various signal-to-noise measures and discuss its contribution towards robust AGR. An absolute improvement of 11.06% and 13.33% is achieved by the combined pipeline over the i-vector SVM baseline system for 0 dB and -5 dB SNR conditions, respectively. We further analyse the information captured by the first CNN layer in both noisy and denoised speech. Jilt Sebastian, Manoj Kumar 0007, Pavan Kumar D. S., Mathew Magimai-Doss, Hema A. Murthy, Shri Narayanan |
INTERSPEECH | 5 |
| 2018 | Code-switching in Indic Speech Synthesisers
Anju Leela Thomas, Anusha Prakash 0001, Arun Baby, Hema A. Murthy |
INTERSPEECH | 4 |
| 2017 | GDspike: An accurate spike estimation algorithm from noisy calcium fluorescence signalsabstractAccurate estimation of spike train from calcium (Ca2+) fluorescence signals is challenging owing to significant fluctuations of fluorescence level. This paper proposes a non-model-based approach for spike train inference using group delay (GD) analysis. It primarily exploits the property that change in Ca2+fluorescence corresponding to a spike has a notable onset location followed by a decaying transient. The proposed algorithm, GDspike, is compared with state-of-the-art systems on five datasets. F-measure is best for GDspike (41%) followed by STM (40%), MLspike (39%), and Vogelstein (35%). While existing methods are inspired by the physiology of neuronal responses, the proposed approach is inspired by GD-based high-resolution processing of the Ca2+fluorescence signal. GDspike is a fast and unsupervised algorithm. It is found to be unaffected when tested with five different GCaMP indicators and scanning rate varying from 15Hz to 60Hz. Jilt Sebastian, Mari Ganesh Kumar, Y. S. Sreekar, Rajeev Rikhye, Mriganka Sur, Hema A. Murthy |
ICASSP | 6 |
| 2017 | Deep Learning Techniques in Tandem with Signal Processing Cues for Phonetic Segmentation for Text to Speech Synthesis in Indian Languages
Arun Baby, Jeena J. Prakash, S. Rupak Vignesh, Hema A. Murthy |
INTERSPEECH | 4 |
| 2017 | TBT (Toolkit to Build TTS): A High Performance Framework to Build Multiple Language HTS Voice
Atish Shankar Ghone, Rachana Nerpagar, Pranaw Kumar, Arun Baby, Aswin Shanmugam Subramanian, M. Sasikumar 0001, Hema A. Murthy |
INTERSPEECH | 7 |
| 2017 | Discovering Language in Marmoset Vocalization
Sakshi Verma, K. L. Prateek, D. S. Karthik Pandia, Nauman Dawalatabad, Rogier Landman, Jitendra Sharma, Mriganka Sur, Hema A. Murthy |
INTERSPEECH | 8 |
| 2017 | Two-pitch tracking in co-channel speech using modified group delay functions
Rajeev Rajan, Hema A. Murthy |
Speech Commun. | 2 |
| 2017 | Feature-switching: Dynamic feature selection for an i-vector based speaker verification system
M. S. Saranya, R. Padmanabhan, Hema A. Murthy |
Speech Commun. | 3 |
| 2016 | Eigen and multimodal analysis for localizing moving sounding objectsabstractThis paper identifies moving objects in a video that are associated to the corresponding audio, by exploiting the correlation of audio and video features. The proposed technique is based on the correlation of motion features of eigen moving objects with audio mel frequency cepstral coefficients features using canonical correlation analysis. We propose two strategies to detect the eigen moving objects: (i) Per-frame mapped eigen moving object (PFEMO) and (ii) Temporally coherent eigen moving object (TCEMO). While PFEMO segments each frame using superpixel segmentation, TCEMO exploits supervoxel based video segmentation to identify eigen moving objects. Qualitative (mean-opinion score) and quantitative (precision, recall, area under the curve, hit ratio) analysis shows that the performance of the proposed techniques is superior to those of the state-of-the-art methods. Shreya Khare, Akshay Bhandari, Hema A. Murthy |
ICASSP | 3 |
| 2016 | Significance of Pseudo-syllables in building better acoustic models for Indian English TTSabstractSignal processing based landmark detection is precise compared to HMM based alignment, primarily because the location of the landmark is not factored in the estimation of parameters. Acoustic cues for syllable boundaries are usually obtained by exploiting the inherent sonority characteristics of a syllable. As syllabification of the text is based on generalized rules or lexicon definitions, there is a mismatch between the acoustical and the lexical segments for non-native syllabification. In this paper, an attempt is made to modify the syllabification rules for Indian English using acoustic cues obtained from syllable boundaries. The modified syllabifier is used to syllabify the text. Embedded re-estimation is performed using forced alignment at the modified syllable level to obtain refined phoneme boundaries. Indian English Text-to-Speech (TTS) systems are built using labels obtained after (i) embedded re-estimation at the sentence level and (ii) the aforementioned procedure. Reduction in the word error rates for both native Aryan and Dravidian speakers (relatively by 54.1% and 52.4% respectively), suggests that there is a significant synthesis quality improvement in the proposed system. S. Rupak Vignesh, Aswin Shanmugam Subramanian, Hema A. Murthy |
ICASSP | 3 |
| 2016 | Two-Pass IB Based Speaker Diarization System Using Meeting-Specific ANN Based FeaturesabstractIn this paper, we present a two-pass Information Bottleneck (IB) based system for speaker diarization which uses meetingspecific artificial neural network (ANN) based features.We first use IB based speaker diarization system to get the labelled speaker segments.These segments are re-segmented using Kullback-Leibler Hidden Markov Model (KL-HMM) based re-segmentation.The multi-layer ANN is then trained to discriminate these speakers using the re-segmented output labels and the spectral features.We then extract the bottleneck features from the trained ANN and perform principal component analysis (PCA) on these features.After performing PCA, these bottleneck features are used along with the different spectral features in the second pass using the same IB based system with KL-HMM re-segmentation.Our experiments on NIST RT and AMI datasets show that the proposed system performs better than the baseline IB system in terms of speaker error rate (SER) with a best case relative improvement of 28.6% amongst AMI datasets and 27.1% on NIST RT04eval dataset. Nauman Dawalatabad, Srikanth R. Madikeri, Chellu Chandra Sekhar, Hema A. Murthy |
INTERSPEECH | 4 |
| 2016 | Acoustic Analysis of Syllables Across Indian Languages
Anusha Prakash 0001, Jeena J. Prakash, Hema A. Murthy |
INTERSPEECH | 3 |
| 2016 | Organization-Level Control of Excessive Internet DownloadsabstractThe control of excessive downloads by rogue users in organizational LANs is the subject of this work. Two mechanisms have been used in order to accomplish this. The first mechanism, is TCP rate control (TCR), it is a receiver-based flow control technique that can be used to effectively rate limit rogue users' flows, making more bandwidth available to regular users. The second mechanism, admission control reduces the bandwidth wastage due to users disconnecting out of impatience when user goodputs are low. Using simulation-based experiments, it has been demonstrated that the composite technique, exclusive TCP rate and admission control (xTRAC) provides seamless control of rogue users, while improving response times and goodput by upto 58% during overload. In this way regular users are incentivized and rogue users are penalized leading to long-term control of users. Saad Y. Sait, Hema A. Murthy, Krishna M. Sivalingam |
LCN | 2 |
| 2016 | An analysis of the high resolution property of group delay function with applications to audio signal processing
Jilt Sebastian, Manoj Kumar 0007, Hema A. Murthy |
Speech Commun. | 3 |
| 2015 | A multi-level resilience framework for unified networked environmentsabstractNetworked infrastructures underpin most social and economical interactions nowadays and have become an integral part of the critical infrastructure. Thus, it is crucial that heterogeneous networked environments provide adequate resilience in order to satisfy the quality requirements of the user. In order to achieve this, a coordinated approach to confront potential challenges is required. These challenges can manifest themselves under different circumstances in the various infrastructure components. The objective of this paper is to present a multi-level resilience approach that goes beyond the traditional monolithic resilience schemes that focus mainly on one infrastructure component. The proposed framework considers four main aspects, i.e. users, application, network and system. The latter three are part of the technical infrastructure while the former profiles the service user. Under two selected scenarios this paper illustrates how an integrated approach coordinating knowledge from the different infrastructure elements allows a more effective detection of challenges and facilitates the use of autonomic principles employed during the remediation against challenges. Angelos K. Marnerides, Akshay Bhandari, Hema A. Murthy, Andreas Mauthe |
IM | 3 |
| 2014 | Feature Switching in the i-vector framework for speaker verificationabstractLIDIAP T. Asha, M. S. Saranya, D. S. Karthik Pandia, Srikanth R. Madikeri, Hema A. Murthy |
INTERSPEECH | 5 |
| 2014 | A hybrid approach to segmentation of speech using group delay processing and HMM based embedded reestimation
Aswin Shanmugam Subramanian, Hema A. Murthy |
INTERSPEECH | 2 |
| 2013 | Modal analysis and transcription of strokes of the mridangam using non-negative matrix factorizationabstractIn this paper we use a Non-negative Matrix Factorization (NMF) based approach to analyze the strokes of the mridangam, a South Indian hand drum, in terms of the normal modes of the instrument. Using NMF, a dictionary of spectral basis vectors are first created for each of the modes of the mridangam. The composition of the strokes are then studied by projecting them along the direction of the modes using NMF. We then extend this knowledge of each stroke in terms of its basic modes to transcribe audio recordings. Hidden Markov Models are adopted to learn the modal activations for each of the strokes of the mridangam, yielding up to 88.40% accuracy during transcription. Akshay Anantapadmanabhan, Ashwin Bellur, Hema A. Murthy |
ICASSP | 3 |
| 2013 | Group delay based melody monopitch extraction from musicabstractIn this paper, we propose a modified group delay based method for melodic pitch extraction from heterophonic music. The power spectrum of the music signal is first flattened in order that the system characteristics are annihilated, while the characteristics of the source are emphasized. The modified group delay function of this signal produces peaks at multiples of the pitch period. The first 3 peaks are used to determine the actual pitch period. The performance of the proposed system was evaluated on two datasets ADC-2004, and LabROSA. The performance is comparable to that of other magnitude spectrum based approaches. The algorithms are also applied to heterophonic music, namely Carnatic Music. As ground truth is not available for Carnatic Music, the pitch contours were used to synthesize the music, which was evaluated for correctness by a professional musician. Rajeev Rajan, Hema A. Murthy |
ICASSP | 2 |
| 2012 | Decoupling non-stationary and stationary components in long range network time series in the context of anomaly detectionabstractNetwork traffic characterisation and modeling using time series models is an area which has been extensively studied in the past. Coarse-grained (aggregated traffic) time series analysis using parametric approach, primarily carried out at the backbone network over a long time period (of the order of days to months), show strong deterministic cyclic trends, while the fine-grained (at the packet or flow level) counterpart, done mostly at edge network over small time period (of the order of few minutes), exhibit self-similar behaviour. This paper is an attempt to study the fine-grained time series characteristics of network traffic at an edge network, observed over a long period (of the order of days and weeks), using parametric approach. The analysis is carried out in the context of anomaly detection. Most of the earlier attempts in this direction followed a non-parametric approach, by either using adaptive or non-adaptive (i.e assuming stationarity) mechanisms, whose performance is found to be extremely sensitive towards empirically determined parameters of the model and hence difficult to determine. Also, the model parameters need to be recomputed at regular intervals of time (of the order of few seconds to minutes). To some extent, this make such algorithms less attractive in terms of generality and practical implementation. The first part of the paper discusses the statistical characteristics of such long range network time series. These are found to exhibit structural breaks apart from transient shocks and can be approximated by a stationary AR model, after an absolute first difference transformation (i.e decoupling stationary component from the non-stationary one). In the later part of the paper, the efficacy of the model proposed is evaluated, by conducting extensive trace driven simulations for the detection of low intensity TCP SYN flood Denial of Service (DoS) attacks. Performance is measured in terms of false positives, false alarm time, detection rate and detection delay. Experiments are performed on actual traffic traces collected from one of the edge networks over a period of three months and for various sampling intervals (10s, 60s, 120s). Comparative studies with adaptive and non-adaptive methods are carried out to demonstrate the relevance of the proposed model. It is observed that the proposed method gives better performance with 100% detection accuracy for false positive as low as 0.9%. Cyriac James, Hema A. Murthy |
LCN | 2 |
| 2010 | Inference Based Query Expansion Using User's Real Time Implicit Feedback
Sanasam Ranbir Singh, Hema A. Murthy, Timothy A. Gonsalves |
IC3K | 2 |
| 2010 | Acoustic feature diversity and speaker verificationabstractWe present a new method for speaker verification that uses the diversity of information from multiple feature representations. The principle behind the method is that certain features are better at recognising certain speakers. Thus, rather than using the same feature representation for all speakers, we use different features for different speakers. During training, we determine the optimal feature for each speaker from candidate features,
by measuring information-theoretic criteria. During evalua-
tion, verification is performed using the optimal feature of the claimed speaker. Experimental results with four candidate features show that the proposed system outperforms conventional systems that use a single feature or a combination of features.
Index Terms: speaker verification, feature selection R. Padmanabhan, Hema A. Murthy |
INTERSPEECH | 2 |
| 2009 | Robustness of phase based features for speaker recognitionabstractThis paper demonstrates the robustness of group-delay based features for speech processing. An analysis of group delay functions is presented which show that these features retain formant structure even in noise. Furthermore, a speaker verification task performed on the NIST 2003 database show lesser error rates, when compared with the traditional MFCC features. We also mention about using feature diversity to dynamically choose the feature for every claimed speaker. R. Padmanabhan, Sree Hari Krishnan Parthasarathi, Hema A. Murthy |
INTERSPEECH | 3 |
| 2008 | Significance of group delay based acoustic features in the linguistic search space for robust speech recognition
Rajesh M. Hegde, Hema A. Murthy |
INTERSPEECH | 3 |
| 2008 | Methods for improving the quality of syllable based speech synthesisabstractOur earlier work [1] on speech synthesis has shown that syllables can produce reasonably natural quality speech. Nevertheless, audible artifacts are present due to discontinuities in pitch, energy, and formant trajectories at the joining point of the units. In this paper, we present some minimal signal modification techniques for reducing these artifacts. Y. R. Venugopalakrishna, M. V. Vinodh, Hema A. Murthy, Coimbatore S. Ramalingam |
SLT | 3 |
| 2008 | Determining user's interest in real timeabstractMost of the search engine optimization techniques attempt to predict users interest by learning from the past information collected from different sources. But, a user's current interest often depends on many factors which are not captured in the past information. In this paper, we attempt to identify user's current interest in real time from the information provided by the user in the current query session. By identifying user's interest in real time, the engine could adapt differently to different users in real time. Experimental verification indicates that our approach is encouraging for short queries Sanasam Ranbir Singh, Hema A. Murthy, Timothy A. Gonsalves |
WWW | 2 |
| 2007 | Significance of the Modified Group Delay Feature in Speech RecognitionabstractSpectral representation of speech is complete when both the Fourier transform magnitude and phase spectra are specified. In conventional speech recognition systems, features are generally derived from the short-time magnitude spectrum. Although the importance of Fourier transform phase in speech perception has been realized, few attempts have been made to extract features from it. This is primarily because the resonances of the speech signal which manifest as transitions in the phase spectrum are completely masked by the wrapping of the phase spectrum. Hence, an alternative to processing the Fourier transform phase, for extracting speech features, is to process the group delay function which can be directly computed from the speech signal. The group delay function has been used in earlier efforts, to extract pitch and formant information from the speech signal. In all these efforts, no attempt was made to extract features from the speech signal and use them for speech recognition applications. This is primarily because the group delay function fails to capture the short-time spectral structure of speech owing to zeros that are close to the unit circle in the z-plane and also due to pitch periodicity effects. In this paper, the group delay function is modified to overcome these effects. Cepstral features are extracted from the modified group delay function and are called the modified group delay feature (MODGDF). The MODGDF is used for three speech recognition tasks namely, speaker, language, and continuous-speech recognition. Based on the results of feature and performance evaluation, the significance of the MODGDF as a new feature for speech recognition is discussed Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | A syllable based continuous speech recognizer for TamilabstractThis paper presents a novel technique for building a syllable based continuous speech recognizer when unannotated transcribed train data is available. We present two different segmentation algorithms to segment the speech and the corresponding text into comparable syllable like units. A group delay based two level segmentation algorithm is proposed to extract accurate syllable units from the speech data. A rule based text segmentation algorithm is used to automatically annotate the text corresponding to the speech into syllable units. Isolated style syllable models are built using multiple frame size (MFS) and multiple frame rate (MFR) for all unique syllables by collecting examples from annotated speech. Experiments performed on Tamil language show that the recognition performance is comparable to recognizers built using manually segmented train data. These experiments suggest that system development cost can be reduced by using minimum manual effort if sentence level transcription of the speech data is available. Index Terms: syllable based speech recognition, acoustic group delay segmentation, text segmentation, annotation. A. Lakshmi, Hema A. Murthy |
INTERSPEECH | 2 |
| 2006 | Language identification using acoustic log-likelihoods of syllable-like units
T. Nagarajan 0001, Hema A. Murthy |
Speech Commun. | 2 |
| 2005 | Speech Processing Using Joint Features Derived from the Modified Group Delay FunctionabstractThe paper discusses the significance of joint cepstral features derived from the modified group delay function and MFCC in speech processing. We start with a definition of cepstral features derived from the modified group delay function called the modified group delay feature (MODGDF) which is derived from the Fourier transform phase. Robustness issues like similarities of the MODGDF to RASTA and cepstral mean subtraction are discussed. The efficiency with which formants can be reconstructed for noisy cellular speech using joint features derived from early fusion is illustrated. The joint features are used for four speech processing tasks phoneme, syllable, speaker, and language recognition. Based on the results of analysis and performance evaluation, the significance of joint features derived from the MODGDF and MFCC are discussed. Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
ICASSP (1) | 2 |
| 2004 | Application of the modified group delay function to speaker identification and discriminationabstractIn this paper, we explore new methods by which speakers can be identified and discriminated, using features derived from the Fourier transform phase. The modified group delay feature (MODGDF) which is a parameterized form of the modified group delay function is used as a front end feature in this study. A Gaussian mixture model (GMM) based speaker identification system is built with the MODGDF as the front end feature. The system is tested on both clean (TIMIT) and noisy telephone (NTIMIT) speech. The results obtained are compared with traditional Mel frequency cepstral coefficients (MFCC) which is derived from the Fourier transform magnitude. When both MFCC and MODGDF were combined, the performance improved by about 4% indicating that both phase and magnitude contain complementary information. In an earlier paper (Murthy et al. (2003)), it was shown that the MODGDF does possess phoneme specific characteristics. In this paper we show that the MODGDF has speaker specific properties. We also make an attempt to understand speaker discriminating characteristics of the MODGDF using the nonlinear mapping technique based on Sammon mapping (Sammon (1969)) and find that the MODGDF empirically demonstrates a certain level of linear separability among speakers. Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
ICASSP (1) | 2 |
| 2004 | Language identification using parallel syllable-like unit recognitionabstractAutomatic spoken language identification (LID) is the task of identifying the language from a short utterance of the speech signal. The most successful approach to LID uses phone recognizers of several languages in parallel. The basic requirement to build a parallel phone recognition (PPR) system is annotated corpora. A novel approach is proposed for the LID task which uses parallel syllable-like unit recognizers, in a framework similar to the PPR approach in the literature. The difference is that unsupervised syllable models are built from the training data. The data is first segmented into syllable-like units. The syllable segments are then clustered using an incremental approach. This results in a set of syllable models for each language. Our initial results on the OGI MLTS corpora show that the performance is 69.5%. We further show that if only a subset of syllable models that are unique (in some sense), are considered, the performance improves to 75.9%. T. Nagarajan 0001, Hema A. Murthy |
ICASSP (1) | 2 |
| 2004 | Cluster and Intrinsic Dimensionality Analysis of the Modified Group Delay Feature for Speaker Classification
Rajesh M. Hegde, Hema A. Murthy |
ICONIP | 2 |
| 2004 | Distributed speaker recognitionabstractSpeech recognition systems are gaining increasing importance with the wide-spread use of mobile and portable devices and other interactive voice response systems. Because of the resource constraints on such devices and the requirements of specific applications, the need to perform speech recognition over a data network becomes inevitable. The requirements of such a system with a human at one end and a machine at the other end are clearly asymmetric. The major focus of this work is to enable speaker recognition for information access over the network. Assuming that at the client end the device is either a Personal Digital Assistant(PDA) or a cellphone, an attempt is made to perform part of computation at the client end, thus conserve bandwidth. Experiments have been performed on both TIMIT data and TIMIT data passed through a speech codec. The results indicate that by performing feature extraction at the client end, the bitrate can be reduced significantly to 13.6kbps with 96% recognition performance. Veena Desai, Hema A. Murthy |
INTERSPEECH | 2 |
| 2004 | Automatic transcription of continuous speech using unsupervised and incremental trainingabstractIn [1], a novel approach is proposed for automatically segmenting and transcribing continuous speech signal without the use of manually annotated speech corpora. In this approach, the continuous speech signal is first automatically segmented into syllable-like units and similar syllable segments are grouped together using an unsupervised and incremental clustering technique. Separate models are generated for each cluster of syllable segments and labels are assigned to them. These syllable models are then used for recognition/transcription. Even though the results in [1] are quite promising, there are some problems in the clustering technique due to (i) the presence of silence segments at the beginning and end of syllable boundaries. (ii) fragmentation of syllables (iii) merging of syllables and (iv) poor initialization of syllable models. In this paper we specifically address these issues, make several refinements to the baseline system, which has resulted in a significant performance improvement of 8% over that of the baseline system described in [1] L. Sarada Ghadiyaram, Hemalatha Nagarajan, T. Nagarajan 0001, Hema A. Murthy |
INTERSPEECH | 4 |
| 2004 | Continuous speech recognition using joint features derived from the modified group delay function and MFCCabstractFeature extraction and selection for continuous speech recognition is a complex task. State of the art speech recognition systems use features that are derived by ignoring the Fourier transform phase. In our earlier studies we have shown the efficacy of The Modified Group Delay Feature (MODGDF) derived from the Fourier transform phase for phoneme, syllable and speaker recognition. In this paper we use the MODGDF and the popular MFCC derived from Fourier transform magnitude to compute joint features for continuous speech recognition of two Indian languages Tamil and Telugu. A novel method of segmentation of the continuous speech signal into syllable like units followed by isolated style recognition using HMMs is used. We further use an innovative technique which transforms the problem of detecting the correct string of syllabic units with maximum likelihood to finding an optimal state sequence locally. The recognition system does not use any language models. The MODGDF gave promising recognition performance for the two languages and compared well with the MFCC. Joint features derived using MODGDF and MFCC gave a 10.6% improvement for both Tamil and Telugu languages. The improvement reinforces the hypothesis that MODGDF captures complementary information to that of the MFCC and can be used along with the MFCC to capture the complete information in the speech signal at functional level and help in avoiding heavy auditory and language models. Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
INTERSPEECH | 2 |
| 2004 | A new prosodic phrasing model for indian language telugu
Nemala Sridhar Krishna, Hema A. Murthy |
INTERSPEECH | 2 |
| 2004 | The modified group delay feature: a new spectral representation of speechabstractAutomatic recognition of speech by machines begins with extraction of meaningful features from the speech signal. Conventional features like the MFCC are derived from the Fourier transform magnitude spectrum, while totally ignoring the phase spectrum. The importance of the Modified group delay feature (MODGDF) derived from the Fourier transform phase spectrum for speaker and phoneme recognition has been presented in our previous efforts. In this paper we try to analyse the feature theoretically and provide justifications in terms of de-correlation, robustness to convolutional and white noise, cluster structures, separability in lower dimensional space, task independence and class separability. The results of speaker identification and continuous speech recognition using the MODGDF as the front end are also presented. Joint features derived from the MODGDF and MFCC gave significant improvements in recognition performance for both speaker and continuous speech recognition tasks. Using the analytical results in the first half of the paper and the results of performance evaluation in the second half, the MODGDF is proposed as an alternative spectral representation of speech. Hema A. Murthy, Rajesh M. Hegde, Venkata Ramana Rao Gadde |
INTERSPEECH | 1 |
| 2004 | Automatic segmentation of continuous speech using minimum phase group delay functions
V. Kamakshi Prasad, T. Nagarajan 0001, Hema A. Murthy |
Speech Commun. | 3 |
| 2003 | The modified group delay function and its application to phoneme recognitionabstractWe explore a new spectral representation of speech signals through group delay functions. The group delay functions by themselves are noisy and difficult to interpret owing to zeroes that are close to the unit circle in the z-domain and these clutter the spectra. A new modified group delay function (Yegnanarayan, B. and Murthy, H.A., IEEE Trans. Sig. Processing, vol.40, p.2281-9, 1992) that reduces the effects of zeroes close to the unit circle is used. Assuming that this new function is minimum phase, the modified group delay spectrum is converted to a sequence of cepstral coefficients. A preliminary phoneme recogniser is built using features derived from these cepstra. Results are compared with those obtained from features derived from the traditional mel frequency cepstral coefficients (MFCC). The baseline MFCC performance is 34.7%, while that of the best modified group delay cepstrum is 39.2%. The performance of the composite MFCC feature, which includes the derivatives and double derivatives, is 60.7%, while that of the composite modified group delay feature is 57.3%. When these two composite features are combined, /spl sim/2% improvement in performance is achieved (62.8%). When this new system is combined with linear frequency cepstra (LFC) (Gadde, V.R.R. et al., The SRI SPINE 2001 Evaluation System. http://elazar.itd.nrl.navy.mil/spine/sri2/presentation/sri2001.html, 2001), the system performance results in another /spl sim/0.8% improvement (63.6%). Hema A. Murthy, Venkata Ramana Rao Gadde |
ICASSP (1) | 1 |
| 2003 | Segmentation of speech into syllable-like unitsabstractIn the development of a syllable-centric ASR system, segmentation of the acoustic signal into syllabic units is an important stage. This paper presents a minimum phase group delay based approach to segment spontaneous speech into syllablelike units. Here, three different minimum phase signals are derived from the short term energy functions of three sub-bands of speech signals, as if it were a magnitude spectrum. The experiments are carried out on Switchboard and OGI-MLTS corpus and the error in segmentation is found to be utmost 40msec for 85% of the syllable segments. T. Nagarajan 0001, Hema A. Murthy, Rajesh M. Hegde |
INTERSPEECH | 2 |
| 2000 | Language identification from short segments of speechabstractAutomatic language identification (LID) from the spoken speech utterance is a challenging problem. In this paper, we present an LID system that works for South Indian languages and Hindi. Each language is modeled using an approach based on Vector Quantisation [1]. The speech is segmented into di erent sounds (CVs) and the performance of the system on each of the segments is studied. Our studies indicate that the presence of some CVs is crucial for each language. We also find that for the same Consonant and Vowel (CV) combination, the quality of the sound is different in di erent languages. We show that once the speech signal is segmented into CVs, it is possible to perform LID on very short segments (100-150ms) of speech itself. Jyotsana Balleda, Hema A. Murthy, T. Nagarajan 0001 |
INTERSPEECH | 2 |
| 2000 | An automatic algorithm for segmenting and labelling a connected digit sequenceabstractGroup delay functions provide an alternative representation of signal information. The main features of group delay functions are the additive and high resolution properties. The Fourier transform (FT) phase is generally featureless due to random polority and wrapping. But the group delay function which is defined as the negative derivative of phase, can be processed to derive significant information such as peaks and valleys in the spectral envelope. In this paper, we show an application of group delay function to solve the segmentation problem in speech. In the proposed method a new signal is generated by symmetrising the short term energy function. The minimum phase group delay function of this signal is computed, the valleys of which correspond to segment boundaries. The proposed technique was tested on manually segmented digit utterances of the TI-DIGITS database. The overall correct segmentation performance is 77.8%. Digitwise recognition performance on the correctly segmented database is 87.1% V. Kamakshi Prasad, Hema A. Murthy |
INTERSPEECH | 2 |
| 1999 | Robust text-independent speaker identification over telephone channelsabstractThis paper addresses the issue of closed-set text-independent speaker identification from samples of speech recorded over the telephone. It focuses on the effects of acoustic mismatches between training and testing data, and concentrates on two approaches: (1) extracting features that are robust against channel variations and (2) transforming the speaker models to compensate for channel effects. First, an experimental study shows that optimizing the front end processing of the speech signal can significantly improve speaker recognition performance. A new filterbank design is introduced to improve the robustness of the speech spectrum computation in the front-end unit. Next, a new feature based on spectral slopes is described. Its ability to discriminate between speakers is shown to be superior to that of the traditional cepstrum. This feature can be used alone or combined with the cepstrum. The second part of the paper presents two model transformation methods that further reduce channel effects. These methods make use of a locally collected stereo database to estimate a speaker-independent variance transformation for each speech feature used by the classifier. The transformations constructed on this stereo database can then be applied to speaker models derived from other databases. Combined, the methods developed in this paper resulted in a 38% relative improvement on the closed-set 30-s training 5-s testing condition of the NIST'95 Evaluation task, after cepstral mean removal. Hema A. Murthy, Françoise Beaufays, Larry Heck, Mitch Weintraub |
IEEE Trans. Speech Audio Process. | 1 |
| 1995 | Transformation of formants for voice conversion using artificial neural networks
M. Narendranath, Hema A. Murthy, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 1994 | Pitch extraction from root cepstrum
Hema A. Murthy |
ICSLP | 1 |
| 1991 | Processing of noisy speech using modified group delay functionsabstractA novel method of processing noisy speech is presented. The method exploits the properties of the negative derivative of the Fourier transform phase spectrum (group delay function) to derive the features of the vocal tract system and the excitation from the speech signal. The key idea used is that the properties of group delay functions for noise and a stable all-pole filter are distinct. Estimation of the spectrum of the vocal tract system and fundamental frequency are treated as problems of spectrum estimation from noisy data. Results of these studies show that intelligible speech can be synthesized from parameters derived from noisy data with an overall signal-to-noise ratio (SNR) as low as 3 dB.> Bayya Yegnanarayana, Hema A. Murthy, V. R. Ramachandran |
ICASSP | 2 |
| 1991 | Speech processing using group delay functions
Hema A. Murthy, Bayya Yegnanarayana |
Signal Process. | 1 |
| 1991 | Formant extraction from group delay function
Hema A. Murthy, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 1990 | Speech enhancement using group delay functions
Bayya Yegnanarayana, Hema A. Murthy, V. R. Ramachandran |
ICSLP | 2 |
| 1989 | A nonparametric method of formant estimation using group delay spectraabstractA novel minimum-phase group delay technique is discussed that is a nonparametric spectral analysis method possessing the ability to demerge closely coupled formants and detect weak formants. It therefore requires no assumptions concerning the underlying nature of the signal, other than that it originates from an LTI filter system. The technique provides a level of performance in formant detection normally associated only with larynx-synchronous techniques. Moreover, it is extremely easy to implement, requiring only three FFT operations per analysis frame.> G. Duncan, Bayya Yegnanarayana, Hema A. Murthy |
ICASSP | 3 |
| 1989 | Formant extraction from Fourier transform phaseabstractA method of extracting formant information from the short-time Fourier transform phase spectrum of speech is proposed. Fourier transform phase has not been used for formant extraction because it appears to be noisy and difficult to interpret. The effects of wrapping of phase (due to zeros close to the unit circle and the linear phase component) make it difficult to derive useful information. The authors develop algorithms to reduce the effects of wrapping.> Hema A. Murthy, K. V. Madhu Murthy, Bayya Yegnanarayana |
ICASSP | 1 |
| 1987 | Reconstruction from Fourier transform phase with applications to speech analysisabstractThis paper addresses the problem of signal reconstruction from Fourier transform phase. In particular, we examine two aspects of this problem. First, we discuss signal reconstruction from the phase spectrum of the short-time Fourier transform(STFT). Next, we examine the problem of signal recovery from partial phase information. We present the results of our studies on reconstruction from partial phase and discuss the application of these results in speech analysis and coding. Bayya Yegnanarayana, S. Tanveer Fathima, Hema A. Murthy |
ICASSP | 3 |