EDBT 2026 Demo / reviewers in the wild / expert
Sriram Ganapathy
dblp:23/4298
· DBLP profile ↗
132ranked-venue papers
22as first author
49since 2021 · last 2026
0000-0002-5779-9066ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 116 · 20 first-author · 39 since 2021Artificial intelligence and machine learning · 87 · 12 first-author · 35 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Mixture-of-Experts model for multimodal emotion recognition in conversations
Soumya Dutta, Smruthi Balaji, Sriram Ganapathy |
Comput. Speech Lang. | 3 |
| 2025 | ULTRAS - Unified Learning of Transformer Representations for Audio and Speech SignalsabstractSelf-supervised learning (SSL) has driven impressive advances in speech processing by adopting time-domain prediction objectives, while audio representation learning frameworks operate on time-frequency spectrograms. Models optimized for one paradigm struggle to transfer to the other, highlighting the need for a joint framework. We propose Unified Learning of Transformer Representations for Audio and Speech (ULTRAS), where the masking and predictive modeling is performed over long patches of the data. The model, based on the transformer architecture, encodes spectral-patches of log-mel spectrogram features. The predictive modeling of masked segments is performed on spectral and temporal targets using a combined loss-function, forcing the representations to encode time and frequency traits. Experiments are performed on a variety of speech and audio tasks, where we illustrate that the ULTRAS framework achieves improved performance over other established baselines. P. E. Ameenudeen, Charumathi Narayanan, Sriram Ganapathy |
ASRU | 3 |
| 2025 | LLM supervised Pre-training for Multimodal Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC) is challenging due to the multimodal nature of the emotion expression. In this paper, we propose to pretrain a text-based recognition model from unsupervised speech transcripts with LLM guidance. These transcriptions are obtained from a raw speech dataset with a pre-trained ASR system. A text LLM model is queried to provide pseudo-labels for these transcripts, and these pseudo-labeled transcripts are subsequently used for learning an utterance level text-based emotion recognition model. We use the utterance level text embeddings for emotion recognition in conversations along with speech embeddings obtained from a recently proposed pre-trained model. A hierarchical way of training the speech-text model is proposed, keeping in mind the conversational nature of the dataset. We perform experiments on three established datasets, namely, IEMOCAP, MELD, and CMU- MOSI, where we illustrate that the proposed model improves over other benchmarks and achieves state-of-the-art results on two out of these three datasets. Soumya Dutta, Sriram Ganapathy |
ICASSP | 2 |
| 2025 | Identifying and Mitigating Mismatched Language Code in Multilingual ASRabstractMultilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can significantly reduce ASR quality. We present a technique to identify and mitigate this issue. We combine off-the-shelf language-ID and language verification models to determine the language code input to the ASR model. The language verification model acts as a gate that decides when to trust the provided language code or use the output of the language-ID model. We compare these approaches with baselines that include vanilla language-ID based and language-independent ASR models. Our experiments on YouTube, SPRING-INX and FLEURS datasets shows the efficacy of the proposed model especially in the mismatched language code setting. Sepand Mavandadi, Kartik Audhkhasi, Shikhar Bharadwaj, Brian Farris, Tongzhou Chen, Bhuvana Ramabhadran, Sriram Ganapathy |
ICASSP | 8 |
| 2025 | Spoken Language Understanding on Unseen Tasks With In-Context LearningabstractSpoken language understanding (SLU) tasks involve diverse skills that probe the information extraction, classification and/or generation capabilities of models. In this setting, task-specific training data may not always be available. While traditional task-specific SLU models are unable to cater to such requirements, the speech-text large language models (LLMs) offer a promising alternative with emergent abilities. However, out of-the-box, our evaluations indicate that the zero/few-shot performance of prominent open-source speech-text LLMs on SLU tasks are not up to the mark. In this paper, we introduce a novel approach to robust task-agnostic fine-tuning using randomized class labels. With this proposed fine-tuning, we illustrate that the performance of the speech-text LLMs on an unseen task is significantly improved over standard approaches. Critically, the proposed approach avoids the requirement of task-specific data annotations for enabling new tasks in speech-text LLMs. Neeraj Agrawal, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2025 | Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
Debarpan Bhattacharya, Apoorva Kulkarni, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2025 | ABHINAYA - A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
Soumya Dutta, Smruthi Balaji, Varada R, Viveka Salinamakki, Sriram Ganapathy |
INTERSPEECH | 5 |
| 2024 | Multimodal Modeling for Spoken Language IdentificationabstractSpoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there is a wealth of other metadata that may be beneficial for this task. In this work, we propose MuSeLI, a Multimodal Spoken Language Identification method, which delves into the use of various metadata sources to enhance language identification. Our study reveals that metadata such as video title, description and geographic location provide substantial information to identify the spoken language of the multimedia recording. We conduct experiments using two diverse public datasets of YouTube videos, and obtain state-of-the-art results on the language identification task. We additionally conduct an ablation study that describes the distinct contribution of each modality for language recognition. Shikhar Bharadwaj, Shikhar Vashishth, Ankur Bapna, Sriram Ganapathy, Vera Axelrod, Siddharth Dalmia, Daan van Esch, Sandy Ritchie, Partha Talukdar, Jason Riesa |
ICASSP | 5 |
| 2024 | Zero Shot Audio To Audio Emotion Transfer With Speaker DisentanglementabstractThe problem of audio-to-audio (A2A) style transfer involves replacing the style features of the source audio with those from the target audio while preserving the content related attributes of the source audio. In this paper, we propose an efficient approach, termed as Zero-shot Emotion Style Transfer (ZEST), that allows the transfer of emotional content present in the given source audio with the one embedded in the target audio while retaining the speaker and speech content from the source. The proposed system builds upon decomposing speech into semantic tokens, speaker representations and emotion embeddings. Using these factors, we propose a framework to reconstruct the pitch contour of the given speech signal and train a decoder that reconstructs the speech signal. The model is trained using a self-supervision based reconstruction loss. During conversion, the emotion embedding is alone derived from the target audio, while rest of the factors are derived from the source audio. In our experiments, we show that, even without using parallel training data or labels from the source or target audio, we illustrate zero shot emotion transfer capabilities of the proposed ZEST model using objective and subjective quality evaluations. Soumya Dutta, Sriram Ganapathy |
ICASSP | 2 |
| 2024 | LLM Augmented LLMs: Expanding Capabilities through CompositionabstractFoundational models with billions of parameters which have been trained on large corpus of data have demonstrated non-trivial skills in a variety of domains. However, due to their monolithic structure, it is challenging and expensive to augment them or impart new skills. On the other hand, due to their adaptation abilities,several new instances of these models are being trained towards new domains and tasks. In this work, we study the problem of efficient and practical composition of existing foundation models with more specific models to enable newer capabilities. To this end, we propose CALM—Composition to Augment Language Models—which introduces cross-attention between models to compose their representations and enable new capabilities. Salient features of CALM are: (i) Scales up LLMs on new tasks by ‘re-using’ existing LLMs along with a few additional parameters and data, (ii) Existing model weights are kept intact, and hence preserves existing capabilities, and (iii) Applies to diverse domains and settings. We illustrate that augmenting PaLM2-S with a smaller model trained on low-resource languages results in an absolute improvement of up to 13% on tasks like translation into English and arithmetic reasoning for low-resource languages. Similarly,when PaLM2-S is augmented with a code-specific model, we see a relative improvement of 40% over the base model for code generation and explanation tasks—on-par with fully fine-tuned counterparts. Rachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta, Sriram Ganapathy, Abhishek Bapna, Partha Talukdar |
ICLR | 5 |
| 2024 | The Second DISPLACE Challenge: DIarization of SPeaker and LAnguage in Conversational Environments
Shareef Babu Kalluri, Prachi Singh, Pratik Roy Chowdhuri, Apoorva Kulkarni, Shikha Baghel, Pradyoth Hegde, Swapnil Sontakke, K. T. Deepak, S. R. Mahadeva Prasanna, Deepu Vijayasenan, Sriram Ganapathy |
INTERSPEECH | 11 |
| 2024 | Improving Self-supervised Pre-training using Accent-Specific Codebooks
Darshan Prabhu, Omkar Nitsure, Preethi Jyothi, Sriram Ganapathy |
INTERSPEECH | 5 |
| 2024 | Summary of the DISPLACE challenge 2023-DIarization of SPeaker and LAnguage in Conversational Environments
Shikha Baghel, Shreyas Ramoji, Somil Jain, Pratik Roy Chowdhuri, Prachi Singh, Deepu Vijayasenan, Sriram Ganapathy |
Speech Commun. | 7 |
| 2024 | Representation Learning With Hidden Unit Clustering for Low Resource Speech ApplicationsabstractIn this paper, we describe an approach to self-supervised representation learning from raw audio using a hidden unit clustering (HUC) framework. The input to the model consists of audio samples that are windowed and processed with 1-D convolutional layers. The learned “time-frequency” representations from the convolutional neural network (CNN) module are further processed with long short term memory (LSTM) layers which generate a contextual vector representation for every windowed segment. The HUC framework, allowing the categorization of the representations into a small number of phoneme-like units, is used to train the model for learning semantically rich speech representations. The targets consist of phoneme-like pseudo labels for each audio segment and these are generated with an iterative$k$-means algorithm. We explore techniques that improve the speaker invariance of the learned representations and illustrate the effectiveness of the proposed approach on two settings, i) completely unsupervised speech applications on the sub-tasks described as part of the ZeroSpeech 2021 challenge and ii) semi-supervised automatic speech recognition (ASR) applications on the TIMIT dataset and on the GramVaani challenge Hindi dataset. In these experiments, we achieve state-of-art results for various Zerospeech tasks. Further, on the ASR experiments, the HUC representations are shown to improve over other established benchmarks based on wav2vec, HuBERT and Best-RQ. Varun Krishna, Tarun Sai, Sriram Ganapathy |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Speech Dereverberation With Frequency Domain Autoregressive ModelingabstractSpeech applications in far-field real world settings often deal with signals that are corrupted by reverberation. The task of dereverberation constitutes an important step to improve the audible quality and to reduce the error rates in applications like automatic speech recognition (ASR). We propose a unified framework of speech dereverberation for improving the speech quality and the ASR performance using the approach of envelope-carrier decomposition provided by an autoregressive (AR) model. The AR model is applied in the frequency domain of the sub-band speech signals to separate the envelope and carrier parts. A novel neural architecture based on dual path long short term memory (DPLSTM) model is proposed, which jointly enhances the sub-band envelope and carrier components. The dereverberated envelope-carrier signals are modulated and the sub-band signals are synthesized to reconstruct the audio signal back. The DPLSTM model for dereverberation of envelope and carrier components also allows the joint learning of the network weights for the down stream ASR task. In the ASR tasks on the REVERB challenge dataset as well as on the VOiCES dataset, we illustrate that the joint learning of speech dereverberation network and the E2E ASR model yields significant performance improvements over the baseline ASR system trained on log-mel spectrogram as well as other benchmarks for dereverberation (average relative improvements of 10-24% over the baseline system). The speech quality improvements, evaluated using subjective listening tests, further highlight the improved quality of the reconstructed audio. Anurenjan Purushothaman, Debottam Dutta, Sriram Ganapathy |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Pseudo-Label Based Supervised Contrastive Loss for Robust Speech RepresentationsabstractThe self supervised learning (SSL) of speech, with discrete tokenization (pseudo-labels), while illustrating performance improvements in low-resource speech recognition, has faced challenges in achieving context invariant and noise robust representations. In this paper, we propose a self-supervised framework based on contrastive loss of the pseudo-labels, obtained from an offline k-means quantizer (tokenizer). We refer to the proposed setting as pseudo-con. The pseudo-con loss, within a batch of training, allows the model to cluster the instances of the same pseudo-label while separating the instances of a different pseudo-label. The proposed pseudo-con loss can also be combined with the cross entropy loss, commonly used in self-supervised learning schemes. We demonstrate the effectiveness of the pseudo-con loss applied for various SSL techniques, like hidden unit bidirectional encoder representations from transformers (HuBERT), best random quantizer (BEST-RQ) and hidden unit clustering (HUC). Our evaluations using the proposed pseudo-con framework achieves state of art results on various sub-tasks of ZeroSpeech 2021 challenge as well as on the context invariance benchmarks. Further, we show significant performance improvements over existing SSL approaches on the TIMIT phoneme recognition task as well as the Librispeech (100h) ASR experiments. Varun Krishna, Sriram Ganapathy |
ASRU | 2 |
| 2023 | MASR: Multi-Label Aware Speech RepresentationabstractIn the recent years, speech representation learning is constructed primarily as a self-supervised learning (SSL) task, using the raw audio signal alone, while ignoring the side-information that is often available for a given speech recording. In this paper, we propose MASR, a Multi-label Aware Speech Representation learning framework, which addresses the aforementioned limitations. MASR enables the inclusion of multiple external knowledge sources to enhance the utilization of meta-data information. The external knowledge sources are incorporated in the form of sample-level pair-wise similarity matrices that are useful in a hard-mining loss. A key advantage of the MASR framework is that it can be combined with any choice of SSL method. Using MASR representations, we perform evaluations on several downstream tasks such as language identification, speech recognition and other non-semantic tasks such as speaker and emotion recognition. In these experiments, we illustrate significant performance improvements for the MASR over other established benchmarks. We perform a detailed analysis on the language identification task to provide insights on how the proposed loss function enables the representations to separate closely related languages. Anjali Raj, Shikhar Bharadwaj, Sriram Ganapathy, Shikhar Vashishth |
ASRU | 3 |
| 2023 | Accented Speech Recognition With Accent-specific CodebooksabstractSpeech accents pose a significant challenge to state-of-the-art automatic speech recognition (ASR) systems.Degradation in performance across underrepresented accents is a severe deterrent to the inclusive adoption of ASR.In this work, we propose a novel accent adaptation approach for end-to-end ASR systems using cross-attention with a trainable set of codebooks.These learnable codebooks capture accent-specific information and are integrated within the ASR encoder layers.The model is trained on accented English speech, while the test data also contained accents which were not seen during training.On the Mozilla Common Voice multi-accented dataset, we show that our proposed approach yields significant performance gains not only on the seen English accents (up to 37% relative improvement in word error rate) but also on the unseen accents (up to 5% relative improvement in WER).Further, we illustrate benefits for a zero-shot transfer setup on the L2Artic dataset.We also compare the performance with other approaches based on accent adversarial training. Fbank Feats CTC Darshan Prabhu, Preethi Jyothi, Sriram Ganapathy, Vinit Unni |
EMNLP | 3 |
| 2023 | Self-Influence Guided Data Reweighting for Language Model Pre-trainingabstractLanguage Models (LMs) pre-trained with selfsupervision on large text corpora have become the default starting point for developing models for various NLP tasks.Once the pre-training corpus has been assembled, all data samples in the corpus are treated with equal importance during LM pre-training.However, due to varying levels of relevance and quality of data, equal importance to all the data samples may not be the optimal choice.While data reweighting has been explored in the context of task-specific supervised learning and LM fine-tuning, model-driven reweighting for pretraining data has not been explored.We fill this important gap and propose PRESENCE, a method for jointly reweighting samples by leveraging self-influence (SI) scores as an indicator of sample importance and pre-training.PRESENCE promotes novelty and stability for model pre-training.Through extensive analysis spanning multiple model sizes, datasets, and tasks, we present PRESENCE as an important first step in the research direction of sample reweighting for pre-training language models. Megh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth, Sarath Chandar, Partha Talukdar |
EMNLP | 3 |
| 2023 | Supervised Hierarchical Clustering Using Graph Neural Networks for Speaker DiarizationabstractConventional methods for speaker diarization involve windowing an audio file into short segments to extract speaker embeddings, followed by an unsupervised clustering of the embeddings. This multistep approach generates speaker assignments for each segment. In this paper, we propose a novel Supervised HierArchical gRaph Clustering algorithm (SHARC) for speaker diarization where we introduce a hierarchical structure using Graph Neural Network (GNN) to perform supervised clustering. The supervision allows the model to update the representations and directly improve the clustering performance, thus enabling a single-step approach for diarization. In the proposed work, the input segment embeddings are treated as nodes of a graph with the edge weights corresponding to the similarity scores between the nodes. We also propose an approach to jointly update the embedding extractor and the GNN model to perform end-to-end speaker diarization (E2E-SHARC). During inference, the hierarchical clustering is performed using node densities and edge existence probabilities to merge the segments until convergence. In the diarization experiments, we illustrate that the proposed E2E-SHARC approach achieves 53% and 44% relative improvements over the baseline systems on benchmark datasets like AMI and Voxconverse, respectively. Prachi Singh, Amrit Kaul, Sriram Ganapathy |
ICASSP | 3 |
| 2023 | The DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments
Shikha Baghel, Shreyas Ramoji, Sidharth, Ranjana H, Prachi Singh, Somil Jain, Pratik Roy Chowdhuri, Kaustubh Kulkarni, Swapnil Padhi, Deepu Vijayasenan, Sriram Ganapathy |
INTERSPEECH | 11 |
| 2023 | Enhancing the EEG Speech Match Mismatch Tasks With Word Boundaries
Akshara Soman, Vidhi Sinha, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2023 | Label Aware Speech Representation Learning For Language Identification
Shikhar Vashishth, Shikhar Bharadwaj, Sriram Ganapathy, Ankur Bapna, Wei Han 0002, Vera Axelrod, Partha Talukdar |
INTERSPEECH | 3 |
| 2022 | Multimodal Transformer with Learnable Frontend and Self Attention for Emotion RecognitionabstractIn this work, we propose a novel approach for multi-modal emotion recognition from conversations using speech and text. The audio representations are learned jointly with a learnable audio front-end (LEAF) model feeding to a CNN based classifier. The text representations are derived from pre-trained bidirectional encoder representations from transformer (BERT) along with a gated recurrent network (GRU). Both the textual and audio representations are separately processed using a bidirectional GRU network with self-attention. Further, the multi-modal information extraction is achieved using a transformer that is input with the textual and audio embeddings at the utterance level. The experiments are performed on the IEMOCAP database, where we show that the proposed framework improves over the current state-of-the-art results under all the common test settings. This is primarily due to the improved emotion recognition performance achieved in the audio domain. Further, we also show that the model is more robust to textual errors caused by an automatic speech recognition (ASR) system. Soumya Dutta, Sriram Ganapathy |
ICASSP | 2 |
| 2022 | Self Supervised Representation Learning with Deep Clustering for Acoustic Unit Discovery from Raw SpeechabstractThe automatic discovery of acoustic sub-word units from raw speech, without any text or labels, is a growing field of research. The key challenge is to derive representations of speech that can be categorized into a small number of phoneme-like units which are speaker invariant and can broadly capture the content variability of speech. In this work, we propose a novel neural network paradigm that uses the deep clustering loss along with the autoregressive contrastive predictive coding (CPC) loss. Both the loss functions, the CPC and the clustering loss, are self-supervised. The clustering cost involves the loss function using the phoneme-like labels generated with an iterative k-means algorithm. The inclusion of this loss ensures that the model representations can be categorized into a small number of automatic speech units. We experiment with several sub-tasks described as part of the Zerospeech 2021 challenge to illustrate the effectiveness of the framework. In these experiments, we show that proposed representation learning approach improves significantly over the previous self-supervision based models as well as the wav2vec family of models on a range of word-level similarity tasks and language modeling tasks. Varun Krishna, Sriram Ganapathy |
ICASSP | 2 |
| 2022 | End-To-End Speech Recognition with Joint Dereverberation of Sub-Band Autoregressive EnvelopesabstractThe end-to-end (E2E) automatic speech recognition (ASR) systems are often required to operate in reverberant conditions, where the long-term sub-band envelopes of the speech are temporally smeared. In this paper, we develop a feature enhancement approach using a neural model operating on sub-band temporal envelopes. The temporal envelopes are modeled using the framework of frequency domain linear prediction (FDLP). The neural enhancement model proposed in this paper performs an envelope gain based enhancement of temporal envelopes. The model architecture consists of a combination of convolutional and long short term memory (LSTM) neural network layers. Further, the envelope dereverberation, feature extraction and acoustic modeling using transformer based E2E ASR can all be jointly optimized for the speech recognition task. We perform E2E speech recognition experiments on the REVERB challenge dataset as well as on the VOiCES dataset. In these experiments, the proposed joint modeling approach yields significant improvements compared to the baseline E2E ASR system (average relative improvements of 21% on the REVERB challenge dataset and about 10% on the VOiCES dataset). Anurenjan Purushothaman, Anirudh Sreeram, Sriram Ganapathy |
ICASSP | 4 |
| 2022 | The Second Dicova Challenge: Dataset and Performance Analysis for Diagnosis of Covid-19 Using AcousticsabstractThe Second Diagnosis of COVID-19 using Acoustics (DiCOVA) Challenge aimed at accelerating the research in acoustics based detection of COVID-19, a topic at the intersection of acoustics, signal processing, machine learning, and healthcare. This paper presents the details of the challenge, which was an open call for researchers to analyze a dataset of audio recordings consisting of breathing, cough and speech signals. This data was collected from individuals with and without COVID-19 infection, and the task in the challenge was a two-class classification. The development set audio recordings were collected from 965 (172 COVID-19 positive) individuals, while the evaluation set contained data from 471 individuals (71 COVID-19 positive). The challenge featured four tracks, one associated with each sound category of cough, speech and breathing, and a fourth fusion track. A baseline system was also released to benchmark the participants. In this paper, we present an overview of the challenge, the rationale for the data collection and the baseline system. Further, a performance analysis for the systems submitted by the 21 participating teams in the leaderboard is also presented. Neeraj Kumar Sharma 0001, Srikanth Raj Chetupalli, Debarpan Bhattacharya, Debottam Dutta, Pravin Mote, Sriram Ganapathy |
ICASSP | 6 |
| 2022 | Leveraging Symmetrical Convolutional Transformer Networks for Speech to Singing Voice Style TransferabstractIn this paper, we propose a model to perform style transfer of speech to singing voice. Contrary to the previous signal processing-based methods, which require high-quality singing templates or phoneme synchronization, we explore a data-driven approach for the problem of converting natural speech to singing voice. We develop a novel neural network architecture, called SymNet, which models the alignment of the input speech with the target melody while preserving the speaker identity and naturalness. The proposed SymNet model is comprised of symmetrical stack of three types of layers - convolutional, transformer, and self-attention layers. The paper also explores novel data augmentation and generative loss annealing methods to facilitate the model training. Experiments are performed on the NUS and NHSS datasets which consist of parallel data of speech and singing voice. In these experiments, we show that the proposed SymNet model improves the objective reconstruction quality significantly over the previously published methods and baseline architectures. Further, a subjective listening test confirms the improved quality of the audio obtained using the proposed approach (absolute improvement of 0.37 in mean opinion score measure over the baseline system). Shrutina Agarwal, Naoya Takahashi, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2022 | Semi-supervised Acoustic and Language Modeling for Hindi ASR
Tarun Sai Bandarupalli, Shakti Rath, Nirmesh J. Shah, Naoyuki Onoe, Sriram Ganapathy |
INTERSPEECH | 5 |
| 2022 | Coswara: A website application enabling COVID-19 screening by analysing respiratory sound samples and health symptoms
Debarpan Bhattacharya, Debottam Dutta, Neeraj Kumar Sharma 0001, Srikanth Raj Chetupalli, Pravin Mote, Sriram Ganapathy, Chandrakiran C, Sahiti Nori, Suhail K. K, Sadhana Gonuguntla, Murali Alagesan |
INTERSPEECH | 6 |
| 2022 | Analyzing the impact of SARS-CoV-2 variants on respiratory sound signalsabstractThe COVID-19 outbreak resulted in multiple waves of infections that have been associated with different SARS-CoV-2 variants.Studies have reported differential impact of the variants on respiratory health of patients.We explore whether acoustic signals, collected from COVID-19 subjects, show computationally distinguishable acoustic patterns suggesting a possibility to predict the underlying virus variant.We analyze the Coswara dataset which is collected from three subject pools, namely, i) healthy, ii) COVID-19 subjects recorded during the delta variant dominant period, and iii) data from COVID-19 subjects recorded during the omicron surge.Our findings suggest that multiple sound categories, such as cough, breathing, and speech, indicate significant acoustic feature differences when comparing COVID-19 subjects with omicron and delta variants.The classification areas-under-the-curve are significantly above chance for differentiating subjects infected by omicron from those infected by delta.Using a score fusion from multiple sound categories, we obtained an area-under-the-curve of 89% and 52.4% sensitivity at 95% specificity.Additionally, a hierarchical three class approach was used to classify the acoustic data into healthy and COVID-19 positive, and further COVID-19 subjects into delta and omicron variants providing high level of 3-class classification accuracy.These results suggest new ways for designing sound based COVID-19 diagnosis approaches. Debarpan Bhattacharya, Debottam Dutta, Neeraj Kumar Sharma 0001, Srikanth Raj Chetupalli, Pravin Mote, Sriram Ganapathy, Chandrakiran C, Sahiti Nori, Suhail K. K, Sadhana Gonuguntla, Murali Alagesan |
INTERSPEECH | 6 |
| 2022 | Speaker conditioned acoustic modeling for multi-speaker conversational ASRabstractIn this paper, we propose a novel approach for the transcription of speech conversations with natural speaker overlap, from single channel speech recordings.The proposed model is a combination of a speaker diarization system and a hybrid automatic speech recognition (ASR) system.The speaker conditioned acoustic model (SCAM) in the ASR system consists of a series of embedding layers which use the speaker activity inputs from the diarization system to derive speaker specific embeddings.The output of the SCAM are speaker specific senones that are used for decoding the transcripts for each speaker in the conversation.In this work, we experiment with the automatic speaker activity decisions generated using an endto-end speaker diarization system.A joint learning approach is also proposed where the diarization model and the ASR acoustic model are jointly optimized.The experiments are performed on the mixed-channel two speaker recordings from the Switchboard corpus of telephone conversations.In these experiments, we show that the proposed acoustic model, incorporating speaker activity decisions and joint optimization, improves significantly over the ASR system with explicit source filtering (relative improvements of 12% in word error rate (WER) over the baseline system). Srikanth Raj Chetupalli, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2022 | Acoustic Representation Learning on Breathing and Speech Signals for COVID-19 Detection
Debottam Dutta, Debarpan Bhattacharya, Sriram Ganapathy, Amir Hossein Poorjam, Deepak Mittal, Maneesh Kumar Singh 0001 |
INTERSPEECH | 3 |
| 2022 | Transformer Networks for Non-Intrusive Speech Quality PredictionabstractThis paper presents the details of our speech quality prediction system submitted to the Conferencing Speech-2022 challenge.The challenge involved the task of non-intrusive speech quality assessment intended for online conferencing applications.We propose two approaches for speech quality prediction in this work.The first approach uses a combination of deep convolutional neural network (CNN) and LSTM neural network with Kullback-Leibler (KL) loss function and cross entropy (CE) loss function for estimating the mean opinion scores (MOS).Our second approach uses transformer based encoder network before applying attention pooling.We observe that our proposed second method gives significant improvements compared to our first method as well as on the baselines provided by the challenge organizers with respect to Pearson Correlation Coefficient (PCC) and Spearman Rank Correlation Coefficient (SRCC) along with reductions in root mean square error (RMSE).The model is also seen to generalize for unseen data resources on the evaluation dataset. M. K. Jayesh, Mukesh Sharma, Praneeth Vonteddu, M. Ali Basha Shaik, Sriram Ganapathy |
INTERSPEECH | 5 |
| 2022 | Dereverberation of autoregressive envelopes for far-field speech recognitionabstractThe task of speech recognition in far-field environments is adversely affected by the reverberant artifacts that elicit as the temporal smearing of the sub-band envelopes. In this paper, we develop a neural model for speech dereverberation using the long-term sub-band envelopes of speech. The sub-band envelopes are derived using frequency domain linear prediction (FDLP) which performs an autoregressive estimation of the Hilbert envelopes. The neural dereverberation model estimates the envelope gain which when applied to reverberant signals suppresses the late reflection components in the far-field signal. The dereverberated envelopes are used for feature extraction in speech recognition. Further, the sequence of steps involved in envelope dereverberation, feature extraction and acoustic modeling for ASR can be implemented as a single neural processing pipeline which allows the joint learning of the dereverberation network and the acoustic model. Several experiments are performed on the REVERB challenge dataset, CHiME-3 dataset and VOiCES dataset. In these experiments, the joint learning of envelope dereverberation and acoustic model yields significant performance improvements over the baseline ASR system based on log-mel spectrogram as well as other past approaches for dereverberation (average relative improvements of 10–24% over the baseline system). A detailed analysis on the choice of hyper-parameters and the cost function involved in envelope dereverberation is also provided. Anurenjan Purushothaman, Anirudh Sreeram, Sriram Ganapathy |
Comput. Speech Lang. | 4 |
| 2022 | PLDA inspired Siamese networks for speaker verification
Shreyas Ramoji, Prashant Krishnan V, Sriram Ganapathy |
Comput. Speech Lang. | 3 |
| 2022 | Towards sound based testing of COVID-19 - Summary of the first Diagnostics of COVID-19 using Acoustics (DiCOVA) Challenge
Neeraj Kumar Sharma 0001, Ananya Muguli, Prashant Krishnan V, Srikanth Raj Chetupalli, Sriram Ganapathy |
Comput. Speech Lang. | 6 |
| 2021 | Self-Supervised Metric Learning With Graph Clustering For Speaker DiarizationabstractIn this paper, we propose a novel algorithm for speaker diarization using metric learning for graph based clustering. The graph clustering algorithms use an adjacency matrix consisting of similarity scores. These scores are computed between speaker embeddings extracted from pairs of audio segments within the given recording. In this paper, we propose an approach that jointly learns the speaker embeddings and the similarity metric using principles of self-supervised learning. The metric learning network implements a neural model of the probabilistic linear discriminant analysis (PLDA). The self-supervision is derived from the pseudo labels obtained from a previous iteration of clustering. The entire model of representation learning and metric learning is trained with a binary cross entropy loss. By combining the self-supervision based metric learning along with the graph-based clustering algorithm, we achieve significant relative improvements of 60% and 7% over the x-vector PLDA agglomerative hierarchical clustering (AHC) approach on AMI and the DIHARD datasets respectively in terms of diarization error rates (DER). Prachi Singh, Sriram Ganapathy |
ASRU | 2 |
| 2021 | Representation Learning for Speech Recognition Using Feedback Based Relevance WeightingabstractIn this work, we propose an acoustic embedding based approach for representation learning in speech recognition. The proposed approach involves two stages comprising of acoustic filterbank learning from raw waveform, followed by modulation filterbank learning. In each stage, a relevance weighting operation is employed that acts as a feature selection module. In particular, the relevance weighting network receives embeddings of the model outputs from the previous time instants as feedback. The proposed relevance weighting scheme allows the respective feature representations to be adaptively selected before propagation to the higher layers. The application of the proposed approach for the task of speech recognition on Aurora-4 and CHiME-3 datasets gives significant performance improvements over baseline systems on raw waveform signal as well as those based on mel representations (average relative improvement of 15% over the mel baseline on Aurora-4 dataset and 7% on CHiME-3 dataset). Purvi Agrawal, Sriram Ganapathy |
ICASSP | 2 |
| 2021 | End-to-End Lyrics Recognition with Voice to Singing Style TransferabstractAutomatic transcription of monophonic/polyphonic music is a challenging task due to the lack of availability of large amounts of transcribed data. In this paper, we propose a data augmentation method that converts natural speech to singing voice based on vocoder based speech synthesizer. This approach, called voice to singing (V2S), performs the voice style conversion by modulating the F0 contour of the natural speech with that of a singing voice. The V2S model based style transfer can generate good quality singing voice thereby enabling the conversion of large corpora of natural speech to singing voice that is useful in building an E2E lyrics transcription system. In our experiments on monophonic singing voice data, the V2S style transfer provides a significant gain (relative improvements of 21 %) for the E2E lyrics transcription system. We also discuss additional components like transfer learning and lyrics based language modeling to improve the performance of the lyrics transcription system. Sakya Basak, Shrutina Agarwal, Sriram Ganapathy, Naoya Takahashi |
ICASSP | 3 |
| 2021 | NISP: A Multi-lingual Multi-accent Dataset for Speaker ProfilingabstractMany commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include physical traits of the speaker like height, age, and gender of the speaker along with the native language of the speaker. Many of the datasets available have only partial information for speaker profiling. In this paper, we attempt to overcome this limitation by developing a new dataset which has speech data from five different Indian languages along with English. The meta-data information for speaker profiling applications like linguistic information, regional information, and physical characteristics of a speaker are also collected. We call this dataset as NITK-IISc Multilingual Multi-accent Speaker Profiling (NISP) dataset. The description of the dataset, potential applications, and baseline results for speaker profiling on this dataset are provided in this paper. Shareef Babu Kalluri, Deepu Vijayasenan, Sriram Ganapathy, Ragesh Rajan M, Prashant Krishnan V |
ICASSP | 3 |
| 2021 | Deep Multiway Canonical Correlation Analysis For Multi-Subject Eeg NormalizationabstractThe normalization of brain recordings from multiple subjects responding to the natural stimuli is one of the key challenges in auditory neuroscience. The objective of this normalization is to transform the brain data in such a way as to remove the inter-subject redundancies and to boost the component related to the stimuli. In this paper, we propose a deep learning framework to improve the correlation of electroencephalography (EEG) data recorded from multiple subjects engaged in an audio listening task. The proposed model extends the linear multi-way canonical correlation analysis (CCA) for audio-EEG analysis using an auto-encoder network with a shared encoder layer. The model is trained to optimize a combined loss involving correlation and reconstruction. The experiments are performed on EEG data collected from subjects listening to natural speech and music. In these experiments, we show that the proposed deep multi-way CCA (DMCCA) based model significantly improves the correlations over the linear multi-way CCA approach with absolute improvements of 0.08 and 0.29 in terms of the Pearson correlation values for speech and music tasks respectively. Jaswanth Reddy Katthi, Sriram Ganapathy |
ICASSP | 2 |
| 2021 | Investigating Feature Selection and Explainability for COVID-19 Diagnostics from Cough SoundsabstractIn this paper, we propose an approach to automatically classify COVID-19 and non-COVID-19 cough samples based on the combination of both feature engineering and deep learning models. In the feature engineering approach, we develop a support vector machine classifier over high dimensional (6373D) space of acoustic features. In the deep learning-based approach, on the other hand, we apply a convolutional neural network trained on the log-mel spectrograms. These two methodologically diverse models are then combined by fusing the probability scores of the models. The proposed system, which ranked 9th on the 2021 Diagnosing COVID-19 using Acoustics (Di- COVA) challenge leaderboard, obtained an area under the receiver operating characteristic curve (AUC) of 0:81 on the blind test data set, which is a 10:9 absolute improvement compared to the baseline. Moreover, we analyze the explainability of the deep learning-based model when detecting COVID-19 from cough signals. Copyright © 2021 ISCA. Flávio Ávila, Amir Hossein Poorjam, Deepak Mittal, Charles Dognin, Ananya Muguli, Srikanth Raj Chetupalli, Sriram Ganapathy, Maneesh Kumar Singh 0001 |
Interspeech | 8 |
| 2021 | Uncovering the Acoustic Cues of COVID-19 Infection
Sriram Ganapathy |
Interspeech | 1 |
| 2021 | DiCOVA Challenge: Dataset, Task, and Baseline System for COVID-19 Diagnosis Using AcousticsabstractThe DiCOVA challenge aims at accelerating research in diagnosing COVID-19 using acoustics (DiCOVA), a topic at the intersection of speech and audio processing, respiratory health diagnosis, and machine learning.This challenge is an open call for researchers to analyze a dataset of sound recordings, collected from COVID-19 infected and non-COVID-19 individuals, for a two-class classification.These recordings were collected via crowdsourcing from multiple countries, through a website application.The challenge features two tracks, one focusing on cough sounds, and the other on using a collection of breath, sustained vowel phonation, and number counting speech recordings.In this paper, we introduce the challenge and provide a detailed description of the task, and present a baseline system for the task. Ananya Muguli, Lancelot Pinto, Nirmala R., Neeraj Kumar Sharma 0001, Prashant Krishnan V, Prasanta Kumar Ghosh, Shrirama Bhat, Srikanth Raj Chetupalli, Sriram Ganapathy, Shreyas Ramoji, Viral Nanda |
Interspeech | 10 |
| 2021 | SRIB-LEAP Submission to Far-Field Multi-Channel Speech Enhancement Challenge for Video ConferencingabstractThis paper presents the details of the SRIB-LEAP submission to the ConferencingSpeech challenge 2021. The challenge involved the task of multi-channel speech enhancement to improve the quality of far field speech from microphone arrays in a video conferencing room. We propose a two stage method involving a beamformer followed by single channel enhancement. For the beamformer, we incorporated self-attention mechanism as inter-channel processing layer in the filter-and-sum network (FaSNet), an end-to-end time-domain beamforming system. The single channel speech enhancement is done in log spectral domain using convolution neural network (CNN)-long short term memory (LSTM) based architecture. We achieved improvements in objective quality metrics - perceptual evaluation of speech quality (PESQ) of 0.5 on the noisy data. On subjective quality evaluation, the proposed approach improved the mean opinion score (MOS) by an absolute measure of 0.9 over the noisy audio. R. G. Prithvi Raj, M. K. Jayesh, Anurenjan Purushothaman, Sriram Ganapathy, M. Ali Basha Shaik |
Interspeech | 5 |
| 2021 | The Third DIHARD Diarization ChallengeabstractDIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise conditions, and conversational domain. Speaker diarization was evaluated under two speech activity conditions (diarization from a reference speech activity vs. diarization from scratch) and 11 diverse domains. The domains span a range of recording conditions and interaction types, including read audio-books, meeting speech, clinical interviews, web videos, and, for the first time, conversational telephone speech. A total of 30 organizations (forming 21teams) from industry and academia submitted 499 valid system outputs. The evaluation results indicate that speaker diarization has improved markedly since DIHARD I, particularly for two-party interactions, but that for many domains (e.g., web video) the problem remains far from solved. Neville Ryant, Prachi Singh, Venkat Krishnamohan, Rajat Varma, Kenneth Church 0001, Christopher Cieri, Jun Du 0002, Sriram Ganapathy, Mark Y. Liberman |
Interspeech | 8 |
| 2021 | LEAP Submission for the Third DIHARD Diarization ChallengeabstractThe LEAP submission for DIHARD-III challenge is described in this paper. The proposed system is composed of a speech bandwidth classifier, and diarization systems fine-tuned for narrowband and wideband speech separately. We use an end-to-end speaker diarization system for the narrowband conversational telephone speech recordings. For the wideband multi-speaker recordings, we use a neural embedding based clustering approach, similar to the baseline system. The embeddings are extracted from a time-delay neural network (called x-vectors) followed by the graph based path integral clustering (PIC) approach. The LEAP system showed 24% and 18% relative improvements for Track-1 and Track-2 respectively over the baseline system provided by the organizers. This paper describes the challenge submission, the post-evaluation analysis and improvements observed on the DIHARD-III dataset. Prachi Singh, Rajat Varma, Venkat Krishnamohan, Srikanth Raj Chetupalli, Sriram Ganapathy |
Interspeech | 5 |
| 2021 | Self-Supervised Representation Learning With Path Integral Clustering for Speaker DiarizationabstractAutomatic speaker diarization techniques typically involve a two-stage processing approach where audio segments of fixed duration are converted to vector representations in the first stage. This is followed by an unsupervised clustering of the representations in the second stage. In most of the prior approaches, these two stages are performed in an isolated manner with independent optimization steps. In this paper, we propose a representation learning and clustering algorithm that can be iteratively performed for improved speaker diarization. The representation learning is based on principles of self-supervised learning while the clustering algorithm is a graph structural method based on path integral clustering (PIC). The representation learning step uses the cluster targets from PIC and the clustering step is performed on embeddings learned from the self-supervised deep model. This iterative approach is referred to as self-supervised clustering (SSC). The diarization experiments are performed on CALLHOME and AMI meeting datasets. In these experiments, we show that the SSC algorithm improves significantly over the baseline system (relative improvements of 13% and 59% on CALLHOME and AMI datasets respectively in terms of diarization error rate (DER)). In addition, the DER results reported in this work improve over several other recent approaches for speaker diarization. Prachi Singh, Sriram Ganapathy |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Unsupervised Neural Mask Estimator for Generalized Eigen-Value Beamforming Based AsrabstractThe state-of-art methods for acoustic beamforming in multi-channel ASR are based on a neural mask estimator that predicts the presence of speech and noise. These models are trained using a paired corpus of clean and noisy recordings (teacher model). In this paper, we attempt to move away from the requirements of having supervised clean recordings for training the mask estimator. The models based on signal enhancement and beamforming using multi-channel linear prediction serve as the required mask estimate. In this way, the model training can also be carried out on real recordings of noisy speech rather than simulated ones alone done in a typical teacher model. Several experiments performed on noisy and reverberant environments in the CHiME-3 corpus as well as the REVERB challenge corpus highlight the effectiveness of the proposed approach. The ASR results for the proposed approach provide performances that are significantly better than a teacher model trained on an out-of-domain dataset and on par with the oracle mask estimators trained on the in-domain dataset. Anirudh Sreeram, Anurenjan Purushothaman, Sriram Ganapathy |
ICASSP | 4 |
| 2020 | 3-D Acoustic Modeling for Far-Field Multi-Channel Speech RecognitionabstractThe conventional approach to automatic speech recognition in multichannel reverberant conditions involves a beamforming based enhancement of the multi-channel speech signal followed by a single channel neural acoustic model. In this paper, we propose to model the multi-channel signal directly using a convolutional neural network (CNN) based architecture which performs the joint acoustic modeling on the three dimensions of time, frequency and channel. The features that are input to the 3-D CNN are extracted by modeling the signal peaks in the spatio-spectral domain using a multivariate autoregressive modeling approach. This AR model is efficient in capturing the channel correlations in the frequency domain of the multi-channel signal. The experiments are conducted on the CHiME-3 and REVERB Challenge dataset using multi-channel reverberant speech. In these experiments, the proposed 3-D feature and acoustic modeling approach provides significant improvements over an ASR system trained with beamformed audio (average relative improvements of 16% and 6% in word error rates for CHiME-3 and REVERB Challenge datasets respectively). Anurenjan Purushothaman, Anirudh Sreeram, Sriram Ganapathy |
ICASSP | 3 |
| 2020 | On The Impact of Language Familiarity in Talker Change DetectionabstractThe ability to detect talker changes when listening to conversational speech is fundamental to perception and understanding of multi-talker speech. In this paper, we propose an experimental paradigm to provide insights on the impact of language familiarity on talker change detection. Two multi-talker speech stimulus sets, one in a language familiar to the listeners (English) and the other unfamiliar (Chinese), are created. A listening test is performed in which listeners indicate the number of talkers in the presented stimuli. Analysis of human performance shows statistically significant results for: (a) lower miss (and a higher false alarm) rate in familiar versus unfamiliar language, and (b) longer response time in familiar versus unfamiliar language. These results signify a link between perception of talker attributes and language proficiency. Subsequently, a machine system is designed to perform the same task. The system makes use of the current state-of-the-art diarization approach with x-vector embeddings. A performance comparison on the same stimulus set indicates that the machine system falls short of human performance by a huge margin, for both languages. Neeraj Kumar Sharma 0001, Venkat Krishnamohan, Sriram Ganapathy, Ahana Gangopadhayay, Lauren K. Fink |
ICASSP | 3 |
| 2020 | Improving Voice Separation by Incorporating End-To-End Speech RecognitionabstractDespite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic nature of speech by taking a transfer learning approach using an end-to-end automatic speech recognition (E2EASR) system. The voice separation is conditioned on deep features extracted from E2EASR to cover the long-term dependence of phonetic aspects. Experimental results on speech separation and enhancement task on the AVSpeech dataset show that the proposed method significantly improves the signal-to-distortion ratio over the baseline model and even outperforms an audio visual model, that utilizes visual information of lip movements. Naoya Takahashi, Mayank Kumar Singh, Sakya Basak, Sudarsanam Parthasaarathy, Sriram Ganapathy, Yuki Mitsufuji |
ICASSP | 5 |
| 2020 | Robust Raw Waveform Speech Recognition Using Relevance Weighted RepresentationsabstractSpeech recognition in noisy and channel distorted scenarios is often challenging as the current acoustic modeling schemes are not adaptive to the changes in the signal distribution in the presence of noise. In this work, we develop a novel acoustic modeling framework for noise robust speech recognition based on relevance weighting mechanism. The relevance weighting is achieved using a sub-network approach that performs feature selection. A relevance sub-network is applied on the output of first layer of a convolutional network model operating on raw speech signals while a second relevance sub-network is applied on the second convolutional layer output. The relevance weights for the first layer correspond to an acoustic filterbank selection while the relevance weights in the second layer perform modulation filter selection. The model is trained for a speech recognition task on noisy and reverberant speech. The speech recognition experiments on multiple datasets (Aurora-4, CHiME-3, VOiCES) reveal that the incorporation of relevance weighting in the neural network architecture improves the speech recognition word error rates significantly (average relative improvements of 10 over the baseline systems). © 2020 ISCA Purvi Agrawal, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2020 | Context Dependent RNNLM for Automatic Transcription of ConversationsabstractConversational speech, while being unstructured at an utterance level, typically has a macro topic which provides larger context spanning multiple utterances. The current language models in speech recognition systems using recurrent neural networks (RNNLM) rely mainly on the local context and exclude the larger context. In order to model the long term dependencies of words across multiple sentences, we propose a novel architecture where the words from prior utterances are converted to an embedding. The relevance of these embeddings for the prediction of next word in the current sentence is found using a gating network. The relevance weighted context embedding vector is combined in the language model to improve the next word prediction, and the entire model including the context embedding and the relevance weighting layers is jointly learned for a conversational language modeling task. Experiments are performed on two conversational datasets - AMI corpus and the Switchboard corpus. In these tasks, we illustrate that the proposed approach yields significant improvements in language model perplexity over the RNNLM baseline. In addition, the use of proposed conversational LM for ASR rescoring results in absolute WER reduction of $1.2$\% on Switchboard dataset and $1.0$\% on AMI dataset over the RNNLM based ASR baseline. Srikanth Raj Chetupalli, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2020 | Audiovisual Correspondence Learning in Humans and MachinesabstractAudiovisual correspondence learning is the task of acquiring the association between images and its corresponding audio. In this paper, we propose a novel experimental paradigm in which unfamiliar pseudo images and pseudowords in audio form are introduced to both humans and machine systems. The task is to learn the association between the pairs of image and audio which is later evaluated with a retrieval task. The machine system used in the study is pretrained with the ImageNet corpus along with the corresponding audio labels. This model is transfer learned for the new image-audio pairs. Using the proposed paradigm, we perform a direct comparison of one-shot, two-shot and three-shot learning performance for humans and machine systems. The human behavioral experiment confirms that the majority of the correspondence learning happens in the first exposure of the audio-visual pair. This paper proposes a machine model which performs on par with the humans in audiovisual correspondence learning. But compared to the machine model, humans exhibited better generalization ability for new input samples with a single exposure. Copyright © 2020 ISCA Venkat Krishnamohan, Akshara Soman, Sriram Ganapathy |
INTERSPEECH | 4 |
| 2020 | Deep Learning Based Dereverberation of Temporal Envelopes for Robust Speech RecognitionabstractAutomatic speech recognition in reverberant conditions is a challenging task as the long-term envelopes of the reverberant speech are temporally smeared. In this paper, we propose a neural model for enhancement of sub-band temporal envelopes for dereverberation of speech. The temporal envelopes are derived using the autoregressive modeling framework of frequency domain linear prediction (FDLP). The neural enhancement model proposed in this paper performs an envelop gain based enhancement of temporal envelopes and it consists of a series of convolutional and recurrent neural network layers. The enhanced sub-band envelopes are used to generate features for automatic speech recognition (ASR). The ASR experiments are performed on the REVERB challenge dataset as well as the CHiME-3 dataset. In these experiments, the proposed neural enhancement approach provides significant improvements over a baseline ASR system with beamformed audio (average relative improvements of 21 on the development set and about 11 on the evaluation set in word error rates for REVERB challenge dataset). © 2020 ISCA Anurenjan Purushothaman, Anirudh Sreeram, Sriram Ganapathy |
INTERSPEECH | 4 |
| 2020 | Neural PLDA Modeling for End-to-End Speaker VerificationabstractWhile deep learning models have made significant advances in supervised classification problems, the application of these models for out-of-set verification tasks like speaker recognition has been limited to deriving feature embeddings. The state-of-the-art x-vector PLDA based speaker verification systems use a generative model based on probabilistic linear discriminant analysis (PLDA) for computing the verification score. Recently, we had proposed a neural network approach for backend modeling in speaker verification called the neural PLDA (NPLDA) where the likelihood ratio score of the generative PLDA model is posed as a discriminative similarity function and the learnable parameters of the score function are optimized using a verification cost. In this paper, we extend this work to achieve joint optimization of the embedding neural network (x-vector network) with the NPLDA network in an end-to-end (E2E) fashion. This proposed end-to-end model is optimized directly from the acoustic features with a verification cost function and during testing, the model directly outputs the likelihood ratio score. With various experiments using the NIST speaker recognition evaluation (SRE) 2018 and 2019 datasets, we show that the proposed E2E model improves significantly over the x-vector PLDA baseline speaker verification system. Shreyas Ramoji, Prashant Krishnan V, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2020 | Coswara - A Database of Breathing, Cough, and Voice Sounds for COVID-19 DiagnosisabstractThe COVID-19 pandemic presents global challenges transcending boundaries of country, race, religion, and economy.The current gold standard method for COVID-19 detection is the reverse transcription polymerase chain reaction (RT-PCR) testing.However, this method is expensive, time-consuming, and violates social distancing.Also, as the pandemic is expected to stay for a while, there is a need for an alternate diagnosis tool which overcomes these limitations, and is deployable at a large scale.The prominent symptoms of COVID-19 include cough and breathing difficulties.We foresee that respiratory sounds, when analyzed using machine learning techniques, can provide useful insights, enabling the design of a diagnostic tool.Towards this, the paper presents an early effort in creating (and analyzing) a database, called Coswara, of respiratory sounds, namely, cough, breath, and voice.The sound samples are collected via worldwide crowdsourcing using a website application.The curated dataset is released as open access.As the pandemic is evolving, the data collection and analysis is a work in progress.We believe that insights from analysis of Coswara can be effective in enabling sound based technology solutions for point-of-care diagnosis of respiratory infection, and in the near future this can help to diagnose COVID-19. Neeraj Kumar Sharma 0001, Prashant Krishnan V, Shreyas Ramoji, Srikanth Raj Chetupalli, Nirmala R., Prasanta Kumar Ghosh, Sriram Ganapathy |
INTERSPEECH | 8 |
| 2020 | Deep Self-Supervised Hierarchical Clustering for Speaker DiarizationabstractThe state-of-the-art speaker diarization systems use agglomerative hierarchical clustering (AHC) which performs the clustering of previously learned neural embeddings. While the clustering approach attempts to identify speaker clusters, the AHC algorithm does not involve any further learning. In this paper, we propose a novel algorithm for hierarchical clustering which combines the speaker clustering along with a representation learning framework. The proposed approach is based on principles of self-supervised learning where the self-supervision is derived from the clustering algorithm. The representation learning network is trained with a regularized triplet loss using the clustering solution at the current step while the clustering algorithm uses the deep embeddings from the representation learning step. By combining the self-supervision based representation learning along with the clustering algorithm, we show that the proposed algorithm improves significantly 29% relative improvement) over the AHC algorithm with cosine similarity for a speaker diarization task on CALLHOME dataset. In addition, the proposed approach also improves over the state-of-the-art system with PLDA affinity matrix with 10% relative improvement in DER. Prachi Singh, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2020 | Supervised I-vector modeling for language and accent recognition
Shreyas Ramoji, Sriram Ganapathy |
Comput. Speech Lang. | 2 |
| 2020 | Automatic speaker profiling from short duration speech data
Shareef Babu Kalluri, Deepu Vijayasenan, Sriram Ganapathy |
Speech Commun. | 3 |
| 2020 | Interpretable Representation Learning for Speech and Audio Signals Based on Relevance WeightingabstractThe learning of interpretable representations from raw data presents significant challenges for time series data like speech. In this work, we propose a relevance weighting scheme that allows the interpretation of the speech representations during the forward propagation of the model itself. The relevance weighting is achieved using a sub-network approach that performs the task of feature selection. A relevance sub-network, applied on the output of first layer of a convolutional neural network model operating on raw speech signals, acts as an acoustic filterbank (FB) layer with relevance weighting. A similar relevance sub-network applied on the second convolutional layer performs modulation filterbank learning with relevance weighting. The full acoustic model consisting of relevance sub-networks, convolutional layers and feed-forward layers is trained for a speech recognition task on noisy and reverberant speech in the Aurora-4, CHiME-3 and VOiCES datasets. The proposed representation learning framework is also applied for the task of sound classification in the UrbanSound8K dataset. A detailed analysis of the relevance weights learned by the model reveals that the relevance weights capture information regarding the underlying speech/audio content. In addition, speech recognition and sound classification experiments reveal that the incorporation of relevance weighting in the neural network architecture improves the performance significantly. Purvi Agrawal, Sriram Ganapathy |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Towards Relevance and Sequence Modeling in Language RecognitionabstractThe task of automatic language identification (LID) involving multiple dialects of the same language family in the presence of noise is a challenging problem. In these scenarios, the identity of the language/dialect may be reliably present only in parts of the temporal sequence of the speech signal. The conventional approaches to LID (and for speaker recognition) ignore the sequence information by extracting long-term statistical summary of the recording assuming an independence of the feature frames. In this paper, we propose a neural network framework utilizing short-sequence information in language recognition. In particular, a new model is proposed for incorporating relevance in language recognition, where parts of speech data are weighted more based on their relevance for the language recognition task. This relevance weighting is achieved using the bidirectional long short-term memory (BLSTM) network with attention modeling. We explore two approaches, the first approach uses segment level i-vector/x-vector representations that are aggregated in the neural model and the second approach where the acoustic features are directly modeled in an end-to-end neural model. Experiments are performed using the language recognition task in NIST LRE 2017 Challenge using clean, noisy and multi-speaker speech data as well as in the RATS language recognition corpus. In these experiments on noisy LRE tasks as well as the RATS dataset, the proposed approach yields significant improvements over the conventional i-vector/x-vector based language recognition approaches as well as with other previous models incorporating sequence information. Bharat Padi, Sriram Ganapathy |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Speaker and Language Aware Training for End-to-End ASRabstractThe end-to-end (E2E) approach to automatic speech recognition (ASR) is a simplified and an elegant approach where a single deep neural network model directly converts the acoustic feature sequence to the text sequence. The current approach to end-to-end ASR uses the neural network model (trained with sequence loss) along with an external character/word based language model (LM) in a decoding pass to output the text sequence. In this work, we propose a new objective function for end-to-end ASR training where the LM score is explicitly introduced in the attention model loss function without any additional training parameters. In this manner, the neural network is made LM aware and this simplifies the model training process. We also propose to incorporate an attention based sequence summary feature in the ASR model which allows the system to be speaker aware. With several E2E ASR experiments on TED-LIUM, WSJ and Librispeech datasets, we show that the proposed speaker and LM aware training improves the ASR performance significantly over the state-of-art E2E approaches. We achieve the best published results reported for WSJ dataset. Shubham Bansal, Karan Malhotra, Sriram Ganapathy |
ASRU | 3 |
| 2019 | Second Language Transfer Learning in Humans and Machines Using Image SupervisionabstractIn the task of language learning, humans exhibit remarkable ability to learn new words from a foreign language with very few instances of image supervision. The question therefore is whether such transfer learning efficiency can be simulated in machines. In this paper, we propose a deep semantic model for transfer learning words from a foreign language (Japanese) using image supervision. The proposed model is a deep audio-visual correspondence network that uses a proxy based triplet loss. The model is trained with large dataset of multi-modal speech/image input in the native language (English). Then, a subset of the model parameters of the audio network are transfer learned to the foreign language words using proxy vectors from the image modality. Using the proxy based learning approach, we show that the proposed machine model achieves transfer learning performance for an image retrieval task which is comparable to the human performance. We also present an analysis that contrasts the errors made by humans and machines in this task. Kiran Praveen, Akshara Soman, Sriram Ganapathy |
ASRU | 4 |
| 2019 | Deep Variational Filter Learning Models for Speech RecognitionabstractWe present a novel approach to derive robust speech representations for automatic speech recognition (ASR) systems. The proposed method uses an unsupervised data-driven modulation filter learning approach that preserves the key modulations of speech signal in spectro-temporal domain. This is achieved by a deep generative modeling framework to learn modulation filters using convolutional variational autoencoder (CVAE). A skip connection based CVAE enables the learning of multiple irredundant modulation filters in the time and frequency modulation domain using temporal and spectral trajectories of input spectrograms. The learnt filters are used to process the spectrogram features for ASR training. The ASR experiments are performed on Aurora-4 (additive noise with channel artifact) and CHiME-3 (additive noise with reverberation) databases. The results show significant improvements for the proposed CVAE model over the baseline features as well as other robust front-ends (average relative improvements of 9% in word error rate over baseline features on Aurora-4 database and 23% on CHiME-3 database). In addition, the performance of the proposed features is highly beneficial for semi-supervised training of ASR when reduced amounts of labeled training data are available (average relative improvements of 29% over baseline features on Aurora-4 database with 30% of the labeled training data). Purvi Agrawal, Sriram Ganapathy |
ICASSP | 2 |
| 2019 | A Deep Neural Network Based End to End Model for Joint Height and Age Estimation from Short Duration SpeechabstractAutomatic height and age prediction of a speaker has a wide variety of applications in speaker profiling, forensics etc. Often in such applications only a few seconds of speech data is available to reliably estimate the speaker parameters. Traditionally, age and height were predicted separately using different estimation algorithms. In this work, we propose a unified DNN architecture to predict both height and age of a speaker for short durations of speech. A novel initialization scheme for the deep neural architecture is introduced, that avoids the requirement for a large training dataset. We evaluate the system on TIMIT dataset where the mean duration of speech segments is around 2.5s. The DNN system is able to improve the age RMSE by at least 0.6 years as compared to a conventional support vector regression system trained on Gaussian Mixture Model mean supervectors. The system achieves an RMSE error of 6.85 and 6.29 cm for male and female height prediction. In case of age estimation, the RMSE errors are 7.60 and 8.63 years for male and female respectively. Analysis of shorter speech segments reveals that even with 1 second speech input the performance degradation is at most 3% compared to the full duration speech files. Shareef Babu Kalluri, Deepu Vijayasenan, Sriram Ganapathy |
ICASSP | 3 |
| 2019 | End-to-end Language Recognition Using Attention Based Hierarchical Gated Recurrent Unit ModelsabstractThe task of automatic language identification (LID) involving multiple dialects of the same language family on short speech recordings is a challenging problem. This can be further complicated for short-duration audio snippets in the presence of noise sources. In these scenarios, the identity of the language/dialect may be reliably present only in parts of the speech embedded in the temporal sequence. The conventional approaches to LID (and for speaker recognition) ignore the sequence information by extracting long-term statistical summary of the recording assuming an independence of the feature frames. In this paper, we propose to develop an end-to-end neural network framework utilizing short-sequence information in language recognition. A hierarchical gated recurrent unit (HGRU) model with attention module is proposed for incorporating relevance in language recognition, where parts of speech data are weighted more based on their relevance for the language recognition task. Experiments are performed using the language recognition task in NIST LRE 2017 Challenge using clean, noisy and multi-speaker speech data. In these experiments, the proposed approach yields significant improvements over the conventional i-vector based language recognition approaches as well as previously proposed approach to language recognition using recurrent networks. Bharat Padi, Sriram Ganapathy |
ICASSP | 3 |
| 2019 | The Leap Speaker Recognition System for NIST SRE 2018 ChallengeabstractThe NIST Speaker Recognition Evaluation (SRE) 2018 challenge comprises an open evaluation of the text independent speaker verification task. This paper summarizes the LEAP speaker verification systems submitted to the NIST SRE 2018. For all the speaker verification approaches, the front-end feature extraction involved the use of neural embeddings from a time delay neural network (TDNN) trained on a speaker discrimination task. These features, called x-vectors, are used in multiple ways for speaker verification task. In the first approach, the x-vectors with pre-processing and dimensionality reduction, are used with probabilistic linear discriminant analysis (PLDA) scoring. The second approach applies a speaker diarization scheme on the test segments containing multiple talkers before speaker verification scoring based on PLDA. The third system uses a local pairwise LDA model for pre-processing the x-vectors which are then scored using a Gaussian back-end. With experiments on the SRE 2018 database, we show that most of the systems achieved noticeable improvements over the NIST baseline in terms of the primary cost metric. Using a system fusion of the various approaches, we obtain significant improvements over the NIST official baseline (average relative improvements of 19.7% and 20.1% for the development and evaluation set respectively). Shreyas Ramoji, Bhargavram Mysore, Anmol Bhatia, Prachi Singh, Harsha Vardhan, Sriram Ganapathy |
ICASSP | 7 |
| 2019 | Analyzing Human Reaction Time for Talker Change DetectionabstractThe ability to detect a change in the input is an essential aspect of perception. In speech communication, we use this ability to identify "talker changes" when listening to conversational speech (such as, audio podcasts). In this paper, we propose to improve our understanding about how fast listeners detect a change in talker, and the acoustic features tracked to identify a voice by designing a novel experimental paradigm. A listening experiment is designed in which listeners indicate the moment of perceived talker change in multi-talker speech utterances. We examine talker change detection performance by probing the human reaction time (RT). A random forest regression is used to model the relationship between RTs and acoustic features. The findings suggest that: (i) RT is less than a second, (ii) RT can be predicted from the difference in acoustic features of segment before and after change, and (iii) there a exists a significant dependence of RT on MFCC-D1 (delta MFCCs) features between segments of speech before and after the change instant. Further, a comparison with a machine system designed for the same task of TCD using speaker diarization principles showed a poor performance relative to the humans. Neeraj Kumar Sharma 0001, Shobhana Ganesh, Sriram Ganapathy, Lori L. Holt |
ICASSP | 3 |
| 2019 | Unsupervised Raw Waveform Representation Learning for ASRabstractThe study aims to investigate the causes, external manifestations and impacts ofderisking phenomenon, increasingly rendered visible in the last years on international financial markets, under the circumstances of the post-crisis global economic developments.The paper reveals that this phenomenon is stilllittle discussed and understood by academics and institutional organizations, consequently, being ignored any theoretical foundation or arguments for the possible guidance, between one or other direction, of policies, norms and regulatory practices in order to monitor and addresstheslippages with adverse effects.In the paper,the stage of actual knowledge of this phenomenon, as concerns the differentforms in which comes out, the main causes and factors of influence, the adverse impact observed and possible counteracting measures is presented.The analysis of the extent to which Romania is affected by de-risking, morevisibleby reducting thenumber of clients and financial transactions, halving the nonperforming loans of the banking system, having also an important social dimension, in the light of the low level of financial inclusion.The paper presents some reference pointsfor decoding this phenomenon, proposing new directions for more in-depth scientific researches,focusing on the identification of specific solutions based on customized analysis, including for Romania"s case,considering the trade-off options, such as stability versus prudentiality in risk taking and risk management, profitability versus social responsibility under uncertainty conditions, banking system interestsversus interests of regulatory and supervisory institutions, monetary policies versus government policy and strategies. Purvi Agrawal, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2019 | A Study of x-Vector Based Speaker Recognition on Short UtterancesabstractThe aim of this work is to gain insights into how the deep neural network (DNN) models should be trained for short utterance evaluation conditions in an x-vector based speaker verification system. The study suggests that the speaker embedding can be extracted with reduced dimensions for short utterance evaluation conditions. When the speaker embedding is extracted from deeper layer which has lower dimension, the x-vector system achieves 14% relative improvement over baseline approach on EER on NIST2010 5sec-5sec truncated conditions. We surmise that since short utterances have less phonetic information speaker discriminative x-vectors can be extracted from a deeper layer of the DNN which captures less phonetic information. Another interesting finding is that the x-vector system achieves 5% relative improvement on NIST2010 5sec-5sec evaluation condition when the back-end PLDA is trained using short utterance development data. The results confirms the intuitive expectation that duration of development utterances and the duration of evaluation utterances should be matched. Finally, for the duration mismatch condition, we propose a variance normalization approach for PLDA training that provides a 4% relative improvement on EER over baseline approach. Ahilan Kanagasundaram, Sridha Sridharan, Sriram Ganapathy, Prachi Singh, Clinton Fookes |
INTERSPEECH | 3 |
| 2019 | Active Learning Methods for Low Resource End-to-End Speech RecognitionabstractRecently developed end-to-end (E2E) automatic speech recognition (ASR) systems demand abundance of transcribed speech data, there are several scenarios where the labeling of speech data is cumbersome and expensive. For a fixed annotation cost, active learning for speech recognition allows to efficiently train the ASR model. In this work, we advance the most common approach for active learning methods which relies on uncertainty sampling technique. In particular, we explore the use of path probability of the decoded sequence as a confidence measure and select the samples with the least confidence for active learning. In order to reduce the sampling bias in active learning, we propose a regularized uncertainty sampling approach that incorporates an i-vector diversity measure. Thus, the active learning in the proposed framework uses a joint score of uncertainty and i-vector diversity. The benefits of the proposed approach are illustrated for an E2E ASR task performed on CSJ and Librispeech datasets. In these experiments, we show that the proposed approach yields considerable improvements over the baseline model using random sampling. Karan Malhotra, Shubham Bansal, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2019 | Attention Based Hybrid i-Vector BLSTM Model for Language RecognitionabstractIn this paper, a hybrid i-vector neural network framework (i-BLSTM) which models the sequence information present in a series of short segment i-vectors for the task of spoken language recognition (LRE) is proposed. A sequence of short segment i-vectors are extracted for every speech utterance and are then modeled using a bidirectional long short-term memory (BLSTM) recurrent neural network (RNN). Attention mechanism inside the neural network relevantly weights segments of the speech utterance and the model learns to give higher weights to parts of speech data which are more helpful to the classification task. The proposed framework performs better in short duration and noisy environments when compared with the conventional i-vector system. Experiments are performed on clean, noisy and multi-speaker speech data from NIST LRE 2017 and RATS language recognition corpus. In these experiments, the proposed approach yields significant improvements (relative improvements of 7.6 - 13 in terms of accuracy for noisy conditions) over the conventional i-vector based language recognition approach and also over an end-to-end LSTM-RNN based approach. Bharat Padi, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2019 | The Second DIHARD Diarization Challenge: Dataset, Task, and BaselinesabstractThis paper introduces the second DIHARD challenge, the second in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions, and conversational domain. The challenge comprises four tracks evaluating diarization performance under two input conditions (single channel vs. multi-channel) and two segmentation conditions (diarization from a reference speech segmentation vs. diarization from scratch). In order to prevent participants from overtuning to a particular combination of recording conditions and conversational domain, recordings are drawn from a variety of sources ranging from read audiobooks to meeting speech, to child language acquisition recordings, to dinner parties, to web video. We describe the task and metrics, challenge design, datasets, and baseline systems for speech enhancement, speech activity detection, and diarization. Neville Ryant, Kenneth Church 0001, Christopher Cieri, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Mark Y. Liberman |
INTERSPEECH | 6 |
| 2019 | LEAP Diarization System for the Second DIHARD ChallengeabstractThis paper presents the LEAP System, developed for the Second DIHARD diarization Challenge. The evaluation data in the challenge is composed of multi-talker speech in restaurants, doctor-patient conversations, child language acquisition recordings in home environments and audio extracted YouTube videos. The LEAP system is developed using two types of embeddings, one based on i-vector representations and the other one based on x-vector representations. The initial diarization output obtained using agglomerative hierarchical clustering (AHC) done on the probabilistic linear discriminant analysis (PLDA) scores is refined using the Variational-Bayes hidden Markov model (VB-HMM) model. We propose a modified VB-HMM model with posterior scaling which provides significant improvements in the final diarization error rate (DER). We also use a domain compensation on the i-vector features to reduce the mis-match between training and evaluation conditions.N(s)TN(s)TN(s)T Using the proposed approaches, we obtain relative improvements in DER of about 7.1% relative for the best individual system over the DIHARD baseline system and about 13.7% relative for the final system combination on evaluation set. An analysis performed using the proposed posterior scaling method shows that scaling results in improved discrimination among the HMM states in the VB-HMM. Prachi Singh, Harsha Vardhan, Sriram Ganapathy, Ahilan Kanagasundaram |
INTERSPEECH | 3 |
| 2018 | 3-D CNN Models for Far-Field Multi-Channel Speech RecognitionabstractAutomatic speech recognition (ASR) in far-field reverberant environments, especially when involving natural conversational multiparty speech conditions, is challenging even with the state-of-the-art recognition methodologies. The two main issues are artifacts in the signal due to reverberation and the presence of multiple speakers. In this paper, we propose a three dimensional (3-D) convolutional neural network (CNN) architecture for multi-channel far-field ASR. This architecture processes time, frequency & channel dimensions of the input spectrogram to learn representations using convolutional layers. Experiments are performed on the REVERB challenge LVCSR task and the augmented multi-party (AMI) LVCSR task using the array microphone recordings. The proposed method shows improvements over the baseline system that uses beamforming of the multi-channel audio along with a 2-D conventional CNN framework (absolute improvements of 1.1 % over the beamformed baseline system on AMI dataset). Sriram Ganapathy, Vijayaditya Peddinti |
ICASSP | 1 |
| 2018 | Enhancement and Analysis of Conversational Speech: JSALT 2017abstractAutomatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinical interviews, extended recordings such as police body cam or child language acquisition data, and any other speech data involving multiple speakers whose voices are not cleanly separated into individual channels. Overlapping speech, environmental noise and suboptimal recording techniques make the problem harder. During the JSALT Summer Workshop at CMU in 2017, an international team of researchers worked on several aspects of this problem, including calibration of the state of the art, detection of overlaps, enhancement of noisy recordings, and classification of shorter speech segments. This paper sketches the workshop's results, and announces plans for a “Diarization Challenge” to encourage further progress. Neville Ryant, Elika Bergelson, Kenneth Church 0001, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Sanjeev Khudanpur, Diana Kowalski, Mahesh Krishnamoorthy, Rajat Kulshreshta, Mark Y. Liberman, Yu-Ding Lu, Matthew Maciejewski, Florian Metze, Ján Profant, Lei Sun 0010, Yu Tsao 0001 |
ICASSP | 6 |
| 2018 | Leveraging LSTM Models for Overlap Detection in Multi-Party MeetingsabstractThe detection of overlapping speech segments is of key importance in speech applications involving analysis of multi-party conversations. The detection problem is challenging because overlapping speech segments are typically captured as short speech utterances far-field microphone recordings. In this paper, we propose detection of overlap segments using a neural network architecture consisting of long-short term memory (LSTM) models. The neural network architecture learns the presence of overlap in speech by identifying the spectrotemporal structure of overlapping speech segments. In order to evaluate the model performance, we perform experiments on simulated overlapped speech generated from the TIMIT database, and natural multi-talker conversational speech in the augmented Multiparty Interaction (AMI) meeting corpus. The proposed approach yields improvements over a Gaussian mixture model based overlap detection system. Furthermore, as an application of overlap detection, integration of overlap detection into speaker diarization task is shown to give improvement in diarization error rate. Neeraj Sajjan, Shobhana Ganesh, Neeraj Kumar Sharma 0001, Sriram Ganapathy, Neville Ryant |
ICASSP | 4 |
| 2018 | Comparison of Unsupervised Modulation Filter Learning Methods for ASRabstractThe widespread deployment of automatic speech recognition (ASR) system in consumer centric applications such as voice interaction and voice search demands the need for noise robustness in such systems. One approach to this problem is to achieve the desired robustness in speech representations used in the ASR. Motivated from studies on robust human speech recognition, we analyse the unsupervised data-driven temporal modulation filter learning for robust feature extraction. In this paper, we compare various unsupervised models for data driven filter learning like convolutional autoencoder (CAE), generative adversarial network (GAN) and convolutional restricted Boltzmann machine (CRBM). The unsupervised models are designed to learn a set of filters from long temporal trajectories of speech sub-band energy. The filters learnt from these models are used for modulation filtering of the input spectrogram before the ASR training. The ASR experiments are performed on Wall Street Journal (WSJ) Aurora-4 database with clean and multi condition training setup. The experimental results obtained from the modulation filtered representations shows considerable robustness to noise, channel distortions and reverberant conditions compared to other feature extraction methods. Among the three approaches compared in this paper, the GAN approach provides the most consistent improvements in ASR accuracy in different training scenarios. Purvi Agrawal, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2018 | Talker Diarization in the Wild: the Case of Child-centered Daylong Audio-recordingsabstractSpeaker diarization (answering 'who spoke when') is a widely researched subject within speech technology. Numerous experiments have been run on datasets built from broadcast news, meeting data, and call centers—the task sometimes appears close to being solved. Much less work has begun to tackle the hardest diarization task of all: spontaneous conversations in real-world settings. Such diarization would be particularly useful for studies of language acquisition, where researchers investigate the speech children produce and hear in their daily lives. In this paper, we study audio gathered with a recorder worn by small children as they went about their normal days. As a result, each child was exposed to different acoustic environments with a multitude of background noises and a varying number of adults and peers. The inconsistency of speech and noise within and across samples poses a challenging task for speaker diarization systems, which we tackled via retraining and data augmentation techniques. We further studied sources of structured variation across raw audio files, including the impact of speaker type distribution, proportion of speech from children, and child age on diarization performance. We discuss the extent to which these findings might generalize to other samples of speech in the wild. Alejandrina Cristià, Shobhana Ganesh, Marisa Casillas, Sriram Ganapathy |
INTERSPEECH | 4 |
| 2018 | Speaker and Language Recognition - From Laboratory Technologies to the Wild
Sriram Ganapathy |
INTERSPEECH | 1 |
| 2018 | Far-Field Speech Recognition Using Multivariate Autoregressive ModelsabstractAutomatic speech recognition (ASR) in far-field reverberant environments is challenging even with the state-of-the-art recognition systems. The main issues are artifacts in the signal due to the long-term reverberation that results in temporal smearing. The autoregressive (AR) modeling approach to speech feature extraction involves representing the high energy regions of the signal which are less susceptible to noise. In this paper, we propose a novel method of speech feature extraction using multivariate AR modeling (MAR) of temporal envelopes. The sub-band discrete cosine transform (DCT) coefficients obtained from multiple speech bands are used in a multivariate linear prediction setting to derive features for speech recognition. For single channel far-field speech recognition, the features are derived using multi-band linear prediction. In the case of multi-channel far-field speech recognition, we use the multi-channel data in the MAR framework. We perform several speech recognition experiments in the REVERB Challenge database for single and multi-microphone settings. In these experiments, the proposed feature extraction method provides significant improvements over baseline methods (average relative improvements of 9.7 % and 3.9 % in single microphone conditions for clean and multi conditions respectively and 6.3 % in multi-microphone conditions). The results with clean training on single microphone conditions further illustrates the effectiveness of the MAR features. Sriram Ganapathy, Madhumita Harish |
INTERSPEECH | 1 |
| 2018 | On Convolutional LSTM Modeling for Joint Wake-Word Detection and Text Dependent Speaker VerificationabstractThe task of personalized keyword detection system which also performs text dependent speaker verification (TDSV) has received substantial interest recently. Conventional approaches to this task involve the development of the TDSV and wakeup -word detection systems separately. In this paper, we show that TDSV and keyword spotting (KWS) can be jointly modeled using the convolutional long short term memory (CLSTM) model architecture, where an initial convolutional feature map is further processed by a LSTM recurrent network. Given a small amount of training data for developing the CLSTM system, we show that the model provides accurate detection of the presence of the keyword in spoken utterance. For the TDSV task, the MTL model can be well regularized using the CLSTM training examples for personalized wake up task. The experiments are performed for KWS wake up detection and TDSV using the combined speech recordings from Wall Street Journal (WSJ) and LibriSpeech corpus. In these experiments with multiple keywords, we illustrate that the proposed approach of MTL significantly improves the performance of previously proposed neural network based text dependent SV systems. We also experimentally illustrate that the CLSTM model provides significant improvements over previously proposed keyword detection systems as well (average relative improvements of 30% over previous approaches). Rajath Kumar, Vaishnavi Yeruva, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2018 | Supervised I-vector Modeling - Theory and ApplicationsabstractOver the last decade, the factor analysis based modeling of a variable length speech utterance into a fixed dimensional vector (termed as i-vector) has been prominently used for many tasks like speaker recognition, language recognition and even in speech recognition. The i-vector model is an unsupervised learning paradigm where the data is initially clustered using a Gaussian Mixture Universal Background Model (GMM-UBM). The adapted means of the Gaussian mixture components are dimensionality reduced using the Total Variability Matrix (TVM) where the latent variables are modeled with a single Gaussian distribution. In this paper, we propose to rework the theory of i-vector modeling using a supervised framework where the speech utterances are associated with a label. Class labels arc introduced in the i-vector model using a mixture Gaussian prior. We show that the proposed model is a generalized i-vector model and the conventional i-vector model turns out to be a special case of this model. This model is applied for a language recognition task using the NIST Language Recognition Evaluation (LRE) 2017 dataset. In these experiments, the supervised i-vector model provides significant improvements over the conventional i-vector model (average relative improvements of 5 % in terms of C-avg). Shreyas Ramoji, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2017 | Deep learning methods for unsupervised acoustic modeling - Leap submission to ZeroSpeech challenge 2017abstractIn this paper, we present our system submission to the ZeroSpeech 2017 Challenge. The track1 of this challenge is intended to develop language independent speech representations that provide the least pairwise ABX distance computed for within speaker and across speaker pairs of spoken words. We investigate two approaches based on deep learning methods for unsupervised modeling. In the first approach, a deep neural network (DNN) is trained on the posteriors of mixture component indices obtained from training a Gaussian mixture model (GMM)-UBM. In the second approach, we develop a similar hidden Markov model (HMM) based DNN model to learn the unsupervised acoustic units provided by HMM state alignments. In addition, we also develop a deep autoencoder which learns language independent embeddings of speech to train the HMM-DNN model. Both the approaches do not use any labeled training data or require any supervision. We perform several experiments using the ZeroSpeech 2017 corpus with the minimal pair ABX error measure. In these experiments, we find that the two proposed approaches significantly improve over the baseline system using MFCC features (average relative improvements of 30-40). Furthermore, the system combination of the two proposed approaches improves the performance over the best individual system. T. K. Ansari, Rajath Kumar, Sonali Singh, Sriram Ganapathy |
ASRU | 4 |
| 2017 | Unsupervised HMM posteriograms for language independent acoustic modeling in zero resource conditionsabstractThe task of language independent acoustic unit modeling in unlabeled raw speech (zero-resource setting) has gained significant interest over the recent years. The main challenge here is the extraction of acoustic representations that elicit good similarity between the same words or linguistic tokens spoken by different speakers and to derive these representations in a language independent manner. In this paper, we explore the use of Hidden Markov Model (HMM) based posteriograms for unsupervised acoustic unit modeling. The states of the HMM (which represent the language independent acoustic units) are initialized using a Gaussian mixture model (GMM) - Universal Background Model (UBM). The trained HMM is subsequently used to generate a temporally contiguous state alignment which are then modeled in a hybrid deep neural network (DNN) model. For the purpose of testing, we use the frame level HMM state posteriors obtained from the DNN as features for the ZeroSpeech challenge task. The minimal pair ABX error rate is measured for both the within and across speaker pairs. With several experiments on multiple languages in the ZeroSpeech corpus, we show that the proposed HMM based posterior features provides significant improvements over the baseline system using MFCC features (average relative improvements of 25% for within speaker pairs and 40% for across speaker pairs). Furthermore, the experiments where the target language is not seen training illustrate the proposed modeling approach is capable of learning global language independent representations. T. K. Ansari, Rajath Kumar, Sonali Singh, Sriram Ganapathy, V. Susheela Devi |
ASRU | 4 |
| 2017 | Leveraging native language speech for accent identification using deep Siamese networksabstractThe problem of automatic accent identification is important for several applications like speaker profiling and recognition as well as for improving speech recognition systems. The accented nature of speech can be primarily attributed to the influence of the speaker's native language on the given speech recording. In this paper, we propose a novel accent identification system whose training exploits speech in native languages along with the accented speech. Specifically, we develop a deep Siamese network based model which learns the association between accented speech recordings and the native language speech recordings. The Siamese networks are trained with i-vector features extracted from the speech recordings using either an unsupervised Gaussian mixture model (GMM) or a supervised deep neural network (DNN) model. We perform several accent identification experiments using the CSLU Foreign Accented English (FAE) corpus. In these experiments, our proposed approach using deep Siamese networks yield significant relative performance improvements of 15.4% on a 10-class accent identification task, over a baseline DNN-based classification system that uses GMM i-vectors. Furthermore, we present a detailed error analysis of the proposed accent identification system. Aditya Siddhant, Preethi Jyothi, Sriram Ganapathy |
ASRU | 3 |
| 2017 | Factor analysis methods for joint speaker verification and spoof detectionabstractThe performance of a speaker verification system is severely degraded by spoofing attacks generated from artificial speech synthesizers. Recently, several approaches have been proposed for classifying natural and synthetic speech (spoof detection) which can be used in conjunction with a speaker verification system. In this paper, we attempt to develop a joint modelling approach which can detect the presence of spoofing attacks while also performing the speaker verification task. We propose a factor modelling approach where the spoof variability subspace and the speaker variability subspace are jointly trained. The lower dimensional projections in these subspaces are used for speaker verification as well as spoof detection tasks. We also investigate the benefits of linear discriminant analysis (LDA), widely used in speaker recognition, for the spoof detection task. Several experiments are performed using the speaker and spoofing (SAS) database. For speaker verification, we compare the performance of the proposed method with a baseline method of fusing a conventional speaker verification system and a spoof detection system. In these experiments, the proposed approach provides substantial improvements for spoof detection (relative improvements of 20% in EER over the baseline) as well as speaker verification under spoofing conditions (relative improvements of 40% in EER over the baseline). Dhanush B. K, Suparna S, Aarthy R, Likhita C, Shashank D, Harish H, Sriram Ganapathy |
ICASSP | 7 |
| 2017 | Speech Representation Learning Using Unsupervised Data-Driven Modulation Filtering for Robust ASRabstractThe performance of an automatic speech recognition (ASR) system degrades severely in noisy and reverberant environments in part due to the lack of robustness in the underlying representations used in the ASR system. On the other hand, the auditory processing studies have shown the importance of modulation filtered spectrogram representations in robust human speech recognition. Inspired by these evidences, we propose a speech representation learning paradigm using data-driven 2- D spectro-temporal modulation filter learning. In particular, multiple representations are derived using the convolutional restricted Boltzmann machine (CRBM) model in an unsupervised manner from the input speech spectrogram. A filter selection criteria based on average number of active hidden units is also employed to select the representations for ASR. The experiments are performed on Wall Street Journal (WSJ) Aurora-4 database with clean and multi condition training setup. In these experiments, the ASR results obtained from the proposed modulation filtering approach shows significant robustness to noise and channel distortions compared to other feature extraction methods (average relative improvements of 19 over baseline features in clean training). Furthermore, the ASR experiments performed on reverberant speech data from the REVERB challenge corpus highlight the benefits of the proposed representation learning scheme for far field speech recognition. Purvi Agrawal, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2017 | IITG-Indigo System for NIST 2016 SRE ChallengeabstractOrientador : Juarez Brandão Lopes Nagendra Kumar 0004, Rohan Kumar Das, Sarfaraz Jelil, Dhanush B. K, H. Kashyap, K. Sri Rama Murty, Sriram Ganapathy, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
INTERSPEECH | 7 |
| 2017 | Increasing the robustness of CNN acoustic models using autoregressive moving average spectrogram features and channel dropout
György Kovács 0001, László Tóth 0001, Dirk Van Compernolle, Sriram Ganapathy |
Pattern Recognit. Lett. | 4 |
| 2017 | Multivariate Autoregressive Spectrogram Modeling for Noisy Speech RecognitionabstractThe performance of an automatic speech recognition (ASR) system is highly degraded in the presence of noise and reverberation. The autoregressive (AR) modeling approach, which preserves the high energy regions of the signal that are less susceptible to noise, first, presents a potential method for robust feature extraction. Second, there are strong correlations in the spectrotemporal domain of the speech signal, which are generally absent in noise. In this letter, we propose a novel method for speech feature extraction, which combines the advantages of AR approach and joint time-frequency processing using the multivariate AR modeling (MAR). Specifically, the subband discrete cosine transform coefficients obtained from multiple speech bands are used in the MAR framework to derive the Riesz temporal envelopes that provide features for ASR. We perform several speech recognition experiments in the Aurora-4 database with clean and multicondition training. In these experiments, the proposed features provide significant improvements over other noise robust feature extraction methods (relative improvements of 24% in clean training and 14% in multicondition training over mel features). Furthermore, the speech recognition experiments in REVERB challenge database illustrates the extension of the MAR modeling method for suppressing reverberant artifacts. Sriram Ganapathy |
IEEE Signal Process. Lett. | 1 |
| 2016 | Speaker age estimation on conversational telephone speech using senone posterior based i-vectorsabstractAutomatic age estimation from speech has a variety of applications including natural human-computer interaction, targeted advertising, customer-agent pairing in call centers, and forensics, to mention a few. Recently, the use of i-vectors has shown promise for automatic age estimation. In this paper, we adopt a phonetically-aware i-vector extractor for the age estimation problem. Such senone i-vector based schemes have demonstrated success in the speaker recognition field. Fixed-length and low-dimensional i-vectors are first conditioned through a linear discriminant analysis (LDA) transform, and then used to train a support vector regression (SVR) model. Additionally, in contrast to previous work, we employ the use of the logarithm of the age as the target in training the SVR to further penalize estimation errors for younger speakers compared with older speakers. The proposed system is evaluated using telephony speech material extracted from the NIST SRE 2008 and 2010 evaluation corpora. Experimental results indicate solid age estimation performance with a mean absolute error (MAE) of 4.7 years for both male and female speakers on the NIST SRE 2010 telephony test set. Seyed Omid Sadjadi, Sriram Ganapathy, Jason W. Pelecanos |
ICASSP | 2 |
| 2016 | An Investigation on the Use of i-Vectors for Robust ASRabstractIn this paper we propose two different i-vector representations that improve the noise robustness of automatic speech recognition (ASR). The first kind of i-vectors is derived from ``noise only'' components of speech provided by an adaptive denoising algorithm, the second variant is extracted from mel filterbank energies containing both speech and noise. The effectiveness of both these representations is shown by combining them with two different kinds of spectral features - the commonly used log-mel filterbank energies and Teager energy spectral coefficients (TESCs). Using two different DNN architectures for acoustic modeling - a standard state-of-the-art sigmoid-based DNN and an advanced architecture using leaky ReLUs, dropout and resealing, we demonstrate the benefit of the proposed representations. On the Aurora-4 multi-condition training task the proposed front-end improves ASR performance by 4%. Dimitrios Dimitriadis, Samuel Thomas 0001, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2016 | The IBM Speaker Recognition System: Recent Advances and Error AnalysisabstractWe present the recent advances along with an error analysis of the IBM speaker recognition system for conversational speech. Some of the key advancements that contribute to our system include: a nearest-neighbor discriminant analysis (NDA) approach (as opposed to LDA) for intersession variability compensation in the i-vector space, the application of speaker and channel-adapted features derived from an automatic speech recognition (ASR) system for speaker recognition, and the use of a DNN acoustic model with a very large number of output units (~10k senones) to compute the frame-level soft alignments required in the i-vector estimation process. We evaluate these techniques on the NIST 2010 SRE extended core conditions (C1-C9), as well as the 10sec-10sec condition. To our knowledge, results achieved by our system represent the best performances published to date on these conditions. For example, on the extended tel-tel condition (C5) the system achieves an EER of 0.59%. To garner further understanding of the remaining errors (on C5), we examine the recordings associated with the low scoring target trials, where various issues are identified for the problematic recordings/trials. Interestingly, it is observed that correcting the pathological recordings not only improves the scores for the target trials but also for the nontarget trials. Seyed Omid Sadjadi, Jason W. Pelecanos, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2015 | Robust speech processing using ARMA spectrogram modelsabstractSpeech applications in noisy and degraded channel conditions continue to be a challenging problem especially when there is a mismatch between the training and test conditions. In this paper, a robust speech feature extraction scheme is developed based on autoregressive moving average (ARMA) modeling that emphasizes high energy regions of the signal with a data driven modulation filter. The peak preserving ability of two dimensional autoregressive (AR) models is used to emphasize the high energy regions in the spectrotemporal domain. The modulation filtering property is achieved by moving average (MA) modeling. The ARMA spectrograms are used to derive features for speech recognition in the Aurora-4 database. In these experiments, the ARMA model features provide significant improvements (relative improvements of 15%) compared to other robust features. Furthermore, the robustness of these features is also verified for language identification (LID) of highly degraded radio channel speech. Here, the ARMA approach achieves relative improvements of up to 20% over the baseline features. Sriram Ganapathy |
ICASSP | 1 |
| 2015 | Nearest neighbor discriminant analysis for language recognitionabstractMany state-of-the-art i-vector based voice biometric systems use linear discriminant analysis (LDA) as a post-processing stage to increase the computational efficiency in the back-end via dimensionality reduction, as well as annihilate the undesired (noisy) directions in the total variability subspace. The traditional approach for computing the LDA transform uses parametric representations for both intra- and inter-class scatter matrices that are based on the Gaussian distribution assumption. However, it is known that the actual distribution of i-vectors may not necessarily be Gaussian, and in particular, in the presence of noise and channel distortions. In addition, the rank of the LDA projection (i.e., the maximum number of available discriminant bases) is limited to the number of classes minus 1. Accordingly, language recognition tasks on noisy data that involve only a few language classes receive limited or no benefit from the LDA post-processing. Motivated by this observation, we present an alternative non-parametric discriminant analysis (NDA) technique that measures both the within- and between-language variation on a local basis using the nearest neighbor rule. The effectiveness of the NDA method is evaluated in the context of noisy language recognition tasks using speech material from the DARPA Robust Automatic Transcription of Speech (RATS) program. Experimental results indicate that NDA is more effective than the traditional parametric LDA for language recognition under noisy and channel degraded conditions. Seyed Omid Sadjadi, Jason W. Pelecanos, Sriram Ganapathy |
ICASSP | 3 |
| 2015 | Investigating factor analysis features for deep neural networks in noisy speech recognitionabstractThe problem of speaker and channel adaptation in deep neural network (DNN) based automatic speech recognition (ASR) sys-tems is of substantial interest in advancing the performance of these systems. Recently, the speaker identity vectors (i-vectors) have shown improvements for ASR systems in matched condi-tions. In this paper, we propose the application of the general factor analysis framework for noisy speech recognition tasks. Several methods for deriving speaker and channel factors are explored including joint factor analysis (JFA) and i-vectors de-rived from DNN posteriors instead of the traditional Universal background model (UBM) approach. We also experiment with the late fusion of i-vector features with bottleneck (BN) fea-tures obtained from a previously trained convolutional neural network (CNN) system. The ASR experiments are performed on the Aspire challenge test data which contains noisy far-field speech while the acoustic models are trained with conversa-tional telephone speech (CTS) data from the Fisher corpus. In these experiments, we show that the factor analysis based meth-ods provide significant improvements in the word error rate (relative improvements of about 11 % compared to the baseline DNN system trained with speaker adapted features). Sriram Ganapathy, Samuel Thomas 0001, Dimitrios Dimitriadis, Steven J. Rennie |
INTERSPEECH | 1 |
| 2014 | Shift-invariant features for speech activity detection in adverse radio-frequency channel conditionsabstractThis work presents a novel approach to speech activity detection for highly degraded radio-frequency channel conditions. In this approach, the audio stream is segmented into short homogeneous segments. Each segment is represented by shift-invariant features. These features provide a coarse histogram-based description of the high-energy trajectories in the time-frequency domain. They are less sensitive to frequency shifting compared to traditional filterbank-based features like Mel-Frequency Cepstral Coefficients (MFCC) and Perceptual Linear Prediction (PLP) coefficients. We evaluate our approach on the speech activity detection task of the Robust Automatic Transcription of Speech (RATS) program. Our experiments show improvements up to 29% relative in the performance in terms of total error on four radio-frequency channels used in RATS compared to the PLP-based baseline system. Mohamed Kamal Omar, Sriram Ganapathy |
ICASSP | 2 |
| 2014 | Analyzing convolutional neural networks for speech activity detection in mismatched acoustic conditionsabstractConvolutional neural networks (CNN) are extensions to deep neural networks (DNN) which are used as alternate acoustic models with state-of-the-art performances for speech recognition. In this paper, CNNs are used as acoustic models for speech activity detection (SAD) on data collected over noisy radio communication channels. When these SAD models are tested on audio recorded from radio channels not seen during training, there is severe performance degradation. We attribute this degradation to mismatches between the two dimensional filters learnt in the initial CNN layers and the novel channel data. Using a small amount of supervised data from the novel channels, the filters can be adapted to provide significant improvements in SAD performance. In mismatched acoustic conditions, the adapted models provide significant improvements (about 10-25%) relative to conventional DNN-based SAD systems. These results illustrate that CNNs have a considerable advantage in fast adaptation for acoustic modeling in these settings. Samuel Thomas 0001, Sriram Ganapathy, George Saon, Hagen Soltau |
ICASSP | 2 |
| 2014 | Robust language identification using convolutional neural network featuresabstractThe language identification (LID) task in the Robust Automatic Transcription of Speech (RATS) program is challenging due to the noisy nature of the audio data collected over highly degraded radio communication channels as well as the use of short duration speech segments for testing. In this paper, we report the recent advances made in the RATS LID task by using bottleneck features from a convolutional neural network (CNN). The CNN, which is trained with labelled data from one of target languages, generates bottleneck features which are used in a Gaussian mixture model (GMM)-ivector LID system. The CNN bottleneck features provide substantial complimentary information to the conventional acoustic features even on languages not seen in its training. Using these bottleneck features in conjunction with acoustic features, we obtain significant improvements (average relative improvements of 25% in terms of equal error rate (EER) compared to the corresponding acoustic system) for the LID task. Furthermore, these improvements are consistent for various choices of acoustic features as well as speech segment durations. Sriram Ganapathy, Kyu Jeong Han, Samuel Thomas 0001, Mohamed Kamal Omar, Maarten Van Segbroeck, Shri Narayanan |
INTERSPEECH | 1 |
| 2014 | Robust Feature Extraction Using Modulation Filtering of Autoregressive ModelsabstractSpeaker and language recognition in noisy and degraded channel conditions continue to be a challenging problem mainly due to the mismatch between clean training and noisy test conditions. In the presence of noise, the most reliable portions of the signal are the high energy regions which can be used for robust feature extraction. In this paper, we propose a front end processing scheme based on autoregressive (AR) models that represent the high energy regions with good accuracy followed by a modulation filtering process. The AR model of the spectrogram is derived using two separable time and frequency AR transforms. The first AR model (temporal AR model) of the sub-band Hilbert envelopes is derived using frequency domain linear prediction (FDLP). This is followed by a spectral AR model applied on the FDLP envelopes. The output 2-D AR model represents a low-pass modulation filtered spectrogram of the speech signal. The band-pass modulation filtered spectrograms can further be derived by dividing two AR models with different model orders (cut-off frequencies). The modulation filtered spectrograms are converted to cepstral coefficients and are used for a speaker recognition task in noisy and reverberant conditions. Various speaker recognition experiments are performed with clean and noisy versions of the NIST-2010 speaker recognition evaluation (SRE) database using the state-of-the-art speaker recognition system. In these experiments, the proposed front-end analysis provides substantial improvements (relative improvements of up to 25%) compared to baseline techniques. Furthermore, we also illustrate the generalizability of the proposed methods using language identification (LID) experiments on highly degraded high-frequency (HF) radio channels and speech recognition experiments on noisy data. Sriram Ganapathy, Sri Harish Reddy Mallidi, Hynek Hermansky |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Unsupervised channel adaptation for language identification using co-trainingabstractLanguage identification (LID) of speech signals in conditions like adverse radio communication channel is a challenging problem. In this paper, we address the scenario of improving the performance of a LID system on mis-matched radio communication channels (not seen in training) given a small amount of speech data without language labels. We develop a co-training procedure using two diverse acoustic LID systems to improve the performance by effectively utilizing the adaptation data. The acoustic LID systems use different features, projection methods and back-end classifiers. Assuming that the classification errors for the diverse LID systems are independent, the co-training procedure improves the classification accuracy of each system. Various LID experiments are performed on the mis-matched channels in a leave-one-out setting for a variety of noise conditions. In these experiments, with small amounts of unsupervised data from the new channel, we show that the proposed co-training procedure provides significant improvement (average relative improvement of 32 %) over the baseline scenario of no-adaptation and noticeable improvements of about 10 % over a self-training framework. Sriram Ganapathy, Mohamed Kamal Omar, Jason W. Pelecanos |
ICASSP | 1 |
| 2013 | TRAP language identification system for RATS phase II evaluationabstractAutomatic language identification or detection of audio data has become an important preprocessing step for speech/speaker recognition and audio data mining. In many surveillance applications, language detection has to be performed on highly degraded audio inputs. In this paper, we present our work on language detection in highly degraded radio channel scenarios. We provide a brief description of the Targeted Robust Audio Processing (TRAP) language detection system builtfor the Phase II Evaluationof the RobustAutomatic Transcription of Speech (RATS) program. This system is a combination of 15 systems with different frontends and speech activity decisions. We also analyze the usefulness of multi-layer perceptron (MLP) based non-linear projection of i-vectors before SVM classification. The proposed backend reduces the Equal Error Rate (EER) by 11%–25% relative compared to the baseline PCA-based feature representation for SVM classification, on the RATS test data consisting of data from eight highfrequency radio communication channels. Index Terms: Language identification (detection), highly degraded radio channel, RATS, i-vector, multi-layer perceptron. Kyu Jeong Han, Sriram Ganapathy, Ming Li 0026, Mohamed Kamal Omar, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Robust speaker recognition using spectro-temporal autoregressive modelsabstractSpeaker recognition in noisy environments is challenging when there is a mis-match in the data used for enrollment and veri-fication. In this paper, we propose a robust feature extraction scheme based on spectro-temporal modulation filtering using two-dimensional (2-D) autoregressive (AR) models. The first step is the AR modeling of the sub-band temporal envelopes by the application of the linear prediction on the sub-band dis-crete cosine transform (DCT) components. These sub-band en-velopes are stacked together and used for a second AR mod-eling step. The spectral envelope across the sub-bands is ap-proximated in this AR model and cepstral features are derived which are used for speaker recognition. The use of AR models emphasizes the focus on the high energy regions which are rel-atively well preserved in the presence of noise. The degree of modulation filtering is controlled using AR model order param-eter. Experiments are performed using noisy versions of NIST 2010 speaker recognition evaluation (SRE) data with a state-of-art speaker recognition system. In these experiments, the proposed features provide significant improvements compared to baseline features (relative improvements of 20 % in terms of equal error rate (EER) and 35 % in terms of miss rate at 10 % false alarm). Sri Harish Reddy Mallidi, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2013 | The IBM speech activity detection system for the DARPA RATS program
George Saon, Samuel Thomas 0001, Hagen Soltau, Sriram Ganapathy, Brian Kingsbury |
INTERSPEECH | 4 |
| 2013 | Enhancing Frequency Shifted Speech Signals in Single Side-Band CommunicationabstractThe spectral quality of speech signals communicated over high-frequency single side band (HF-SSB) radio channels is affected by acoustic artifacts like linear frequency transpositions. In this letter, we propose an approach to automatically estimate and correct for the frequency shift given the degraded signal at the SSB receiver. The proposed method utilizes the harmonic nature of the speech signal in the voiced regions. The fundamental harmonic frequency, obtained from an autoregressive model of the spectrum, is used to estimate the offset value for the current frame. The offset values from the adjacent frames are pooled together to provide the most likely estimate for the received signal. Various experiments are performed on frequency shifted degraded speech signals received from the HF-SSB channel. The enhanced speech signal is also applied for a language identification task (LID) where the models are trained on speech data without any frequency shift. In these experiments, the proposed algorithm provides significant improvements over other baseline offset estimation methods in terms of accuracy of offset estimation as well as the LID system performance (with relative improvements of about 10-25%). Sriram Ganapathy, Jason W. Pelecanos |
IEEE Signal Process. Lett. | 1 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 6 |
| 2012 | Multilingual MLP features for low-resource LVCSR systemsabstractWe introduce a new approach to training multilayer perceptrons (MLPs) for large vocabulary continuous speech recognition (LVCSR) in new languages which have only few hours of annotated in-domain training data (for example, 1 hour of data). In our approach, large amounts of annotated out-of-domain data from multiple languages are used to train multilingual MLP systems without dealing with the different phoneme sets for these languages. Features extracted from these MLP systems are used to train LVCSR systems in the low-resource language similar to the Tandem approach. In our experiments, the proposed features provide a relative improvement of about 30% in an low-resource LVCSR setting with only one hour of training data. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
ICASSP | 2 |
| 2012 | Analysis of Temporal Resolution in Frequency Domain Linear Prediction
Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 1 |
| 2012 | Data-driven Posterior Features for Low Resource Speech Recognition ApplicationsabstractIn low resource settings, with very few hours of training data, state-of-the-art speech recognition systems that require large amounts of task specific training data perform very poorly. We address this issue by building data-driven speech recognition front-ends on significant amounts of task independent data from different languages and genres collected in similar acoustic conditions as the data in the low resource scenario. We show that features derived from these trained front-ends perform significantly better and can alleviate the effect of reduced task specific training data in low resource settings. The proposed features provide a absolute improvement of about 12 % (18 % relative) in an low-resource LVCSR setting with only one hour of training data. We also demonstrate the usefulness of these features for zero-resource speech applications like spoken term discovery, which operate without any transcribed speech to train systems. The proposed features provide significant gains over conventional acoustic features on various information retrieval metrics for this task. Index Terms: Low-resource speech recognition, spoken term discovery, posterior features. Samuel Thomas 0001, Sriram Ganapathy, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 2 |
| 2012 | Noisy channel adaptation in language identificationabstractLanguage identification (LID) of speech data recorded over noisy communication channels is a challenging problem especially when the LID system is tested on speech data from an unseen communication channel (not seen in training). In this paper, we consider the scenario in which a small amount of adaptation data is available from a new communication channel. Various approaches are investigated for efficient utilization of the adaptation data in a supervised as well as unsupervised setting. In a supervised adaptation framework, we show that support vector machines (SVMs) with higher order polynomial kernels (HO-SVM) trained using lower dimensional representations of the the Gaussian mixture model supervectors (GSVs) provide significant performance improvements over the baseline SVM-GSV system. In these LID experiments, we obtain 30% reduction in error-rate with 6 hours of adaptation data for a new channel. For unsupervised adaptation, we develop an iterative procedure for re-labeling the development data using a co-training framework. In these experiments, we obtain considerable improvements(relative improvements of 13 %) over a self-training framework with the HO-SVM models. Sriram Ganapathy, Mohamed Kamal Omar, Jason W. Pelecanos |
SLT | 1 |
| 2011 | Feature normalization for speaker verification in room reverberationabstractThe performance of a typical speaker verification system degrades significantly in reverberant environments. This degradation is partly due to the conventional feature extraction/compensation techniques that use analysis windows which are much shorter than typical room impulse responses. In this paper, we present a feature extraction technique which estimates long-term envelopes of speech in narrow sub-bands using frequency domain linear prediction (FDLP). When speech is corrupted by reverberation, the long-term sub-band envelopes are convolved in time with those of the room impulse response function. In a first order approximation, gain normalization of these envelopes in the FDLP model suppresses the room reverberation artifacts. Experiments are performed on the 8 core conditions of the NIST 2008 speaker recognition evaluation (SRE). In these experiments, the FDLP features provide significant improvements on the interview microphone conditions (relative improvements of 20 30%) over the corresponding baseline system with MFCC features. Sriram Ganapathy, Jason W. Pelecanos, Mohamed Kamal Omar |
ICASSP | 1 |
| 2011 | Modulation Spectrum Analysis for Recognition of Reverberant SpeechabstractRecognition of reverberant speech constitutes a challenging problem for typical speech recognition systems. This is mainly due to the conventional short-term analysis/compensation tech-niques. In this paper, we present a feature extraction technique based on modeling long segments of temporal envelopes of the speech signal in narrow sub-bands using frequency domain lin-ear prediction (FDLP). FDLP provides an all-pole approxima-tion of the Hilbert envelope of the signal by linear prediction on cosine transform of the signal. We show that the FDLP modulation spectrum plays an important role in the robustness of the proposed feature extraction. Automatic speech recogni-tion (ASR) experiments on speech data degraded with a number of room impulse responses (with varying degrees of distortion) show significant performance improvements for the proposed FDLP features when compared to other robust feature extrac-tion techniques (average relative reduction of 40 % in word er-ror rate). Similar improvements are also obtained for far-field data which contain natural reverberation in background noise. Sri Harish Reddy Mallidi, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2010 | Robust spectro-temporal features based on autoregressive models of Hilbert envelopesabstractIn this paper, we present a robust spectro-temporal feature extraction technique using autoregressive models (AR) of sub-band Hilbert envelopes. AR models of Hilbert envelopes are derived using frequency domain linear prediction (FDLP). From the sub-band Hilbert envelopes, spectral features are derived by integrating these envelopes in short-term frames and the temporal features are formed by converting these envelopes into modulation frequency components. The spectral and temporal feature streams are then combined at the phoneme posterior level and are used as the input features for a recognition system. For the proposed features, robustness is achieved by using novel techniques of noise compensation and gain normalization. Phoneme recognition experiments on telephone speech in the HTIMIT database show significant performance improvements for the proposed features when compared to other robust feature techniques (average relative reduction of 10.6 % in phoneme error rate). In addition to the overall phoneme recognition rates, the performance with broad phonetic classes is also reported. Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 1 |
| 2010 | Comparison of modulation features for phoneme recognitionabstractIn this paper, we compare several approaches for the extraction of modulation frequency features from speech signal using a phoneme recognition system. The general framework in these approaches is to decompose the speech signal into a set of sub-bands. Amplitude modulations (AM) in the sub-band signal are used to derive features for automatic speech recognition (ASR). Then, we propose a feature extraction technique which uses autoregressive models (AR) of sub-band Hilbert envelopes in relatively long segments of speech signal. AR models of Hilbert envelopes are derived using frequency domain linear prediction (FDLP). Features are formed by converting the FDLP envelopes into static and dynamic modulation frequency components. In the phoneme recognition experiments using the TIMIT database, the FDLP based modulation frequency features provide significant improvements compared to other techniques (average relative improvement of 7.5% over the base-line features). Furthermore, a detailed analysis is performed to determine the relative contribution of various processing stages in the proposed technique. Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 1 |
| 2010 | Sparse auto-associative neural networks: theory and application to speech recognition
Garimella S. V. S. Sivaram, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2010 | Cross-lingual and multi-stream posterior features for low resource LVCSR systemsabstractWe investigate approaches for large vocabulary continuous speech recognition (LVCSR) system for new languages or new domains using limited amounts of transcribed training data. In these low resource conditions, the performance of conventional LVCSR systems degrade significantly. We propose to train low resource LVCSR system with additional sources of information like annotated data from other languages (German and Spanish) and various acoustic feature streams (short-term and modulation features). We train multilayer perceptrons (MLPs) on these sources of information and use Tandem features derived from the MLPs for low resource LVCSR. In our experiments, the proposed system trained using only one hour of English conversational telephone speech (CTS) provides a relative improvement of 11% over the baseline system. Index Terms: Cross-lingual posterior features, Multi-stream features, Low resource ASR, Tandem features Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2010 | A phoneme recognition framework based on auditory spectro-temporal receptive fieldsabstractWe propose to incorporate features derived using spectrotemporal receptive fields (STRFs) of neurons in the auditory cortex for phoneme recognition. Each of these STRFs is tuned to different auditory frequencies, scales and modulation rates. We select different sets of STRFs which are specific for phonemes in different broad phonetic classes (BPC) of sounds. These STRFs are then used as spectro-temporal filters on spectrograms of speech to extract features for phoneme recognition. For the phoneme recognition task on the TIMIT database, the proposed features show a relative improvement of about 5% over conventional feature extraction techniques. Samuel Thomas 0001, Kailash Patil, Sriram Ganapathy, Nima Mesgarani, Hynek Hermansky |
INTERSPEECH | 3 |
| 2010 | Autoregressive Models of Amplitude Modulations in Audio CompressionabstractWe present a scalable medium bit-rate wide-band audio coding technique based on frequency-domain linear prediction (FDLP). FDLP is an efficient method for representing the long-term amplitude modulations of speech/audio signals using autoregressive models. For the proposed audio codec, relatively long temporal segments (1000 ms) of the input audio signal are decomposed into a set of critically sampled sub-bands using a quadrature mirror filter (QMF) bank. The technique of FDLP is applied on each sub-band to model the sub-band temporal envelopes. The residual of the linear prediction, which represents the frequency modulations in the sub-band signal, are encoded and transmitted along with the envelope parameters. These steps are reversed at the decoder to reconstruct the signal. The proposed codec utilizes a simple signal independent nonadaptive compression mechanism for a wide class of speech and audio signals. The subjective and objective quality evaluations show that the reconstruction signal quality for the proposed FDLP codec compares well with the state-of-the-art audio codecs in the 32-64 kbps range. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Temporal envelope subtraction for robust speech recognition using modulation spectrumabstractIn this paper, we present a new noise compensation technique for modulation frequency features derived from syllable length segments of subband temporal envelopes. The subband temporal envelopes are estimated using frequency domain linear prediction (FDLP). We propose a technique for noise compensation in FDLP where an estimate of the noise envelope is subtracted from the noisy speech envelope. The noise compensated FDLP envelopes are compressed with static (logarithmic) and dynamic (adaptive loops) compression and are transformed into modulation spectral features. Experiments are performed on a phoneme recognition task as well as a connected digit recognition task where the test data is corrupted with variety of noise types at different signal to noise ratios. In these experiments with mismatched train and test conditions, the proposed features provide considerable improvements compared to other state of the art noise robust feature extraction techniques (average relative improvement of 25% and 35% over the baseline PLP features for phoneme and word recognition tasks respectively). Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ASRU | 1 |
| 2009 | Phoneme recognition using spectral envelope and modulation frequency featuresabstractWe present a new feature extraction technique for phoneme recognition that uses short-term spectral envelope and modulation frequency features. These features are derived from sub-band temporal envelopes of speech estimated using frequency domain linear prediction (FDLP). While spectral envelope features are obtained by the short-term integration of the sub-band envelopes, the modulation frequency components are derived from the long-term evolution of the sub-band envelopes. These features are combined at the phoneme posterior level and used as features for a hybrid HMM-ANN phoneme recognizer. For the phoneme recognition task on the TIMIT database, the proposed features show an improvement of 4.7% over the other feature extraction techniques. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
ICASSP | 2 |
| 2009 | Static and dynamic modulation spectrum for speech recognitionabstractWe present a feature extraction technique based on static and dynamic modulation spectrum derived from long-term envelopes in sub-bands. Estimation of the sub-band temporal envelopes is done using Frequency Domain Linear Prediction (FDLP). These sub-band envelopes are compressed with a static (logarithmic) and dynamic (adaptive loops) compression. The compressed sub-band envelopes are transformed into modulation spectral components which are used as features for speech recognition. Experiments are performed on a phoneme recognition task using a hybrid HMM-ANN phoneme recognition system and an ASR task using the TANDEM speech recognition system. The proposed features provide a relative improvements of 3.8 % and 11.5 % in phoneme recognition accuracies for TIMIT and conversation telephone speech (CTS) respectively. Further, these improvements are found to be consistent for ASR tasks on OGI-Digits database (relative improvement of 13.5 %). Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 1 |
| 2009 | Arithmetic coding of sub-band residuals in FDLP speech/audio codecabstractA speech/audio codec based on Frequency Domain Linear Prediction (FDLP) exploits auto-regressive modeling to approximate instantaneous energy in critical frequency sub-bands of relatively long input segments. The current version of the FDLP codec operating at 66 kbps has been shown to provide comparable subjective listening quality results to state-of-the-art codecs on similar bit-rates even without employing standard blocks such as entropy coding or simultaneous masking. This paper describes an experimental work to increase compression efficiency of the FDLP codec by employing entropy coding. Unlike conventional Huffman coding employed in current speech/audio coding systems, we describe an efficient way to exploit arithmetic coding to entropy compress quantized spectral magnitudes of the sub-band FDLP residuals. Such an approach provides 11% (∼ 3 kbps) bit-rate reduction compared to the Huffman coding algorithm (∼ 1 kbps). Petr Motlícek, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2009 | Tandem representations of spectral envelope and modulation frequency features for ASRabstractWe present a feature extraction technique for automatic speech recognition that uses Tandem representation of short-term spectral envelope and modulation frequency features. These features, derived from sub-band temporal envelopes of speech estimated using frequency domain linear prediction, are combined at the phoneme posterior level. Tandem representations derived from these phoneme posteriors are used along with HMM based ASR systems for both small and large vocabulary continuous speech recognition (LVCSR) tasks. For a small vocabulary continuous digit task on the OGI Digits database, the proposed features reduce the word error rate (WER) by 13 % relative to other feature extraction techniques. We obtain a relative reduction of about 14 % in WER for an LVCSR task using the NIST RT05 evaluation data. For phoneme recognition tasks on the TIMIT database these features provide a relative improvement of 13% compared to other techniques. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Temporal masking for bit-rate reduction in audio codec based on Frequency Domain Linear PredictionabstractAudio coding based on frequency domain linear prediction (FDLP) uses auto-regressive model to approximate Hilbert envelopes in frequency sub-bands for relatively long temporal segments. Although the basic technique achieves good quality of the reconstructed signal, there is a need for improving the coding efficiency. In this paper, we present a novel method for the application of temporal masking to reduce the bit-rate in a FDLP based codec. Temporal masking refers to the hearing phenomenon, where the exposure to a sound reduces response to following sounds for a certain period of time (up to 200 ms). In the proposed version of the codec, a first order forward masking model of the human ear is implemented and informal listening experiments using additive white noise are performed to obtain the exact noise masking thresholds. Subsequently, this masking model is employed in encoding the sub- band FDLP carrier signal. Application of the temporal masking in the FDLP codec results in a bit-rate reduction of about 10% without degrading the quality. Performance evaluation is done with perceptual evaluation of audio quality (PEAQ) scores and with subjective listening tests. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky, Harinath Garudadri |
ICASSP | 1 |
| 2008 | Spectral noise shaping: improvements in speech/audio codec based on linear prediction in spectral domainabstractAudio coding based on Frequency Domain Linear Prediction (FDLP) uses auto-regressive models to approximate Hilbert envelopes in frequency sub-bands. Although the basic technique achieves good coding efficiency, there is a need to improve the reconstructed signal quality for tonal signals with impulsive spectral content. For such signals, the quantization noise in the FDLP codec appears as frequency components not present in the input signal. In this paper, we propose a technique of Spectral Noise Shaping (SNS) for improving the quality of tonal signals by applying a Time Domain Linear Prediction (TDLP) filter prior to the FDLP processing. The inverse TDLP filter at the decoder shapes the quantization noise to reduce the artifacts. Application of the SNS technique to the FDLP codec improves the quality of the tonal signals without affecting the bit-rate. Performance evaluation is done with Perceptual Evaluation of Audio Quality (PEAQ) scores and with subjective listening tests. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky, Harinath Garudadri |
INTERSPEECH | 1 |
| 2008 | Front-end for far-field speech recognition based on frequency domain linear predictionabstractAbstract. Automatic Speech Recognition (ASR) systems usually fail when they encounter speech from far-field microphone in reverberant environments. This is due to the application of short-term feature extraction techniques which do not compensate for the artifacts introduced by long room impulse responses. In this paper, we propose a front-end, based on Frequency Domain Linear Prediction (FDLP), that tries to remove reverberation artifacts present in far-field speech. Long temporal segments of far-field speech are analyzed in narrow frequency sub-bands to extract FDLP envelopes and residual signals. Filtering the residual signals with gain normalized inverse FDLP filters result in a set of sub-band signals which are synthesized to reconstruct the signal back. ASR experiments on far-field speech data processed by the proposed front-end show significant improvements (relative reduction of 30 % in word error rate) compared to other robust feature extraction techniques. 2 IDIAP–RR 08-17 1 Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 1 |
| 2008 | Hilbert envelope based spectro-temporal features for phoneme recognition in telephone speechabstractAbstract. In this paper, we present a spectro-temporal feature extraction technique using sub-band Hilbert envelopes of relatively long segments of speech signal. Hilbert envelopes of the sub-bands are estimated using Frequency Domain Linear Prediction (FDLP). Spectral features are derived by integrating the sub-band Hilbert envelopes in short-term frames and the temporal features are formed by converting the FDLP envelopes into modulation frequency components. These are then combined at the phoneme posterior level and are used as the input features for a phoneme recognition system. In order to improve the robustness of the proposed features to telephone speech, the sub-band temporal envelopes are gain normalized prior to feature extraction. Phoneme recognition experiments on telephone speech in the HTIMIT database show significant performance improvements for the proposed features when compared to other robust feature techniques (average relative reduction of 11 % in phoneme error rate).2 IDIAP–RR 08-18 n this paper, we present a spectro-temporal feature extraction technique using sub-band Hilbert envelopes of relatively long segments of speech signal. Hilbert envelopes of the sub-bands are estimated using Frequency Domain Linear Prediction (FDLP). Spectral features are derived by integrating the sub-band Hilbert envelopes in short-term frames and the temporal features are formed by converting Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Recognition of Reverberant Speech Using Frequency Domain Linear PredictionabstractPerformance of a typical automatic speech recognition (ASR) system severely degrades when it encounters speech from reverberant environments. Part of the reason for this degradation is the feature extraction techniques that use analysis windows which are much shorter than typical room impulse responses. We present a feature extraction technique based on modeling temporal envelopes of the speech signal in narrow subbands using frequency domain linear prediction (FDLP). FDLP provides an all-pole approximation of the Hilbert envelope of the signal obtained by linear prediction on cosine transform of the signal. ASR experiments on speech data degraded with a number of room impulse responses (with varying degrees of distortion) show significant performance improvements for the proposed FDLP features when compared to other robust feature extraction techniques (average relative reduction of 24% in word error rate). Similar improvements are also obtained for far-field data which contain natural reverberation in background noise. These results are achieved without any noticeable degradation in performance for clean speech. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
IEEE Signal Process. Lett. | 2 |