EDBT 2026 Demo / reviewers in the wild / expert
Rama Sanand Doddipatla
dblp:47/9231 · also D. Rama Sanand, Doddipatla Rama Sanand, Rama Doddipatla
· DBLP profile ↗
56ranked-venue papers
12as first author
33since 2021 · last 2025
0000-0003-1061-9512ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 11 first-author · 28 since 2021Artificial intelligence and machine learning · 33 · 10 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language ModelsabstractChain-of-Thought (CoT) prompting is a widely used method to improve the reasoning capability of Large Language Models (LLMs). More recently, CoT has been leveraged in Knowledge Distillation (KD) to transfer reasoning capability from a larger LLM to a smaller one. This paper examines the role of CoT in distilling the reasoning capability from larger LLMs to smaller LLMs using white-box KD, analyzing its effectiveness in improving the performance of the distilled models for various natural language reasoning and understanding tasks. We conduct white-box KD experiments using LLMs from the Qwen and Llama2 families, employing CoT data from the CoT-Collection dataset. The distilled models are then evaluated on natural language reasoning and understanding tasks from the BIG-Bench-Hard (BBH) benchmark, which presents complex challenges for smaller LLMs. Experimental results demonstrate the role of CoT in improving white-box KD effectiveness, enabling the distilled models to achieve better average performance in natural language reasoning and understanding tasks from BBH. Cong-Thanh Do, Rama Sanand Doddipatla, Kate M. Knill |
INLG | 2 |
| 2025 | Human ratings of LLM response generation in pair-programming dialogueabstractWe take first steps in exploring whether Large Language Models (LLMs) can be adapted to dialogic learning practices, specifically pair programming — LLMs have primarily been implemented as programming assistants, not fully exploiting their dialogic potential. We used new dialogue data from real pair-programming interactions between students, prompting state-of-the-art LLMs to assume the role of a student, when generating a response that continues the real dialogue. We asked human annotators to rate human and AI responses on the criteria through which we operationalise the LLMs’ suitability for educational dialogue: Coherence, Collaborativeness, and whether they appeared human. Results show model differences, with Llama-generated responses being rated similarly to human answers on all three criteria. Thus, for at least one of the models we investigated, the LLM utterance-level response generation appears to be suitable for pair-programming dialogue. Cecilia Domingo, Paul Piwek, Svetlana Stoyanchev, Michel Wermelinger, Kaustubh Adhikari, Rama Sanand Doddipatla |
INLG | 6 |
| 2024 | Semantic Map-based Generation of Navigation InstructionsabstractWe are interested in the generation of navigation instructions, either in their own right or as training material for robotic navigation task. In this paper, we propose a new approach to navigation instruction generation by framing the problem as an image captioning task using semantic maps as visual input. Conventional approaches employ a sequence of panorama images to generate navigation instructions. Semantic maps abstract away from visual details and fuse the information in multiple panorama images into a single top-down representation, thereby reducing computational complexity to process the input. We present a benchmark dataset for instruction generation using semantic maps, propose an initial model and ask human subjects to manually assess the quality of generated instructions. Our initial investigations show promise in using semantic maps for instruction generation instead of a sequence of panorama images, but there is vast scope for improvement. We release the code for data preparation and model training at https://github.com/chengzu-li/VLGen. Chengzu Li, Simone Teufel, Rama Sanand Doddipatla, Svetlana Stoyanchev |
LREC/COLING | 4 |
| 2024 | Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting ScenariosabstractWe propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions with two active speakers to lie on the shortest path on a sphere between the points given by the d-vectors of each of the active speakers. Using those frame-wise speaker embeddings in clustering-based diarization outperforms segment-level clustering-based diarization systems such as VBx and Spectral Clustering. By extending our approach to a mixture-model-based diarization, the performance can be further improved, approaching the diarization error rates of diarization systems that use a dedicated overlap detection, and outperforming these systems when also employing an additional overlap detection. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2024 | Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
Mohan Li, Simon Keizer, Rama Sanand Doddipatla |
INTERSPEECH | 3 |
| 2024 | WHISMA: A Speech-LLM to Perform Zero-Shot Spoken Language UnderstandingabstractSpeech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken language understanding (SLU) that demonstrates robust performance in various zero-shot settings. WHISMA combines the speech encoder from Whisper with the Llama-3 LLM, and is fine-tuned in a parameter-efficient manner on a comprehensive collection of SLU-related datasets. Our experiments show that WHISMA significantly improves the zero-shot slot filling performance on the SLURP benchmark, achieving a relative gain of 26.6% compared to the current state-of-the-art model. Furthermore, to evaluate WHISMA’s generalisation capabilities to unseen domains, we develop a new task-agnostic benchmark named SLU-GLUE. The evaluation results indicate that WHISMA outperforms an existing speech-LLM (Qwen-Audio) with a relative gain of 33.0%. Mohan Li, Cong-Thanh Do, Simon Keizer, Youmna Farag, Svetlana Stoyanchev, Rama Sanand Doddipatla |
SLT | 6 |
| 2024 | Entity Resolution in Situated Dialog With Unimodal and Multimodal TransformersabstractIn this work we address the entity resolution task for situated multimodal dialog investigating how a unimodal approach, which uses only textual information as input (representing visual attributes as text), compares to a multimodal system, which processes both text and visual information. We analyze two of the top performing models presented in the Tenth Dialog Systems Technology Challenge and propose modifications that enhance their performance on the multimodal coreference resolution task. We evaluate these approaches on in- and out-of-domain settings by training the models on the fashion domain and testing on the furniture domain, and vice-versa, to assess the generalizability of the models. Through systematic analysis, we show that while both systems achieve similar performance on in-domain scenarios, the multimodal system generalizes better to out-of-domain settings. A combination strategy of enhanced unimodal and multimodal systems achieves F1 = 0.80 (5% absolute gain compared to the best performing system). Finally, human performance on the same task is evaluated on a small subset, suggesting that the performance of the current automatic models is on par with people on this task. Alejandro Santorum Varela, Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla, Kate M. Knill |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Robust Recognition of Speaker Emotion With Difference Feature Extraction Using a Few Enrollment UtterancesabstractThis paper presents a novel approach to derive robust representations for speech emotion recognition by extracting speaker-independent features with the help of a speaker-dependent Gaussian mixture model (GMM). Since emotions are subjective and can vary greatly due to speaker behavior, incorporating speaker representations used for speaker verification tasks, such as x-vectors in past studies, have shown to improve the performance of speech emotion recognition, though they are not extracting speaker independent features. In this paper, we propose to derive embeddings that normalize speaker influence in the form of I-vectors, derived using a universal background model (UBM) trained only using the neutral emotion utterances from a single target speaker and an utterance-wise GMM also trained from the same speaker. We show through experiments on three datasets, that the proposed representations outperform methods that employ conventional x-vectors (which are not speaker-independent features) by approx. $3 \%$ absolute on average using as little as 4 enrollment utterances from the target speaker. Daichi Hayakawa, Takehiko Kagoshima, Kenji Iwata, Norbert Braunschweiler, Rama Sanand Doddipatla |
ASRU | 5 |
| 2023 | Towards a Unified End-to-End Language Understanding System for Speech and Text InputsabstractEnd-to-end (E2E) spoken language understanding (SLU) systems facilitate mapping speech inputs directly to semantic outputs, eliminating the need for modular processing of speech-to-text and text-to-semantics sub-tasks using separate models. However, they are now limited to processing speech inputs only, and are not flexible to deal with plain texts. In this paper, we propose an E2E spoken and natural language understanding (SNLU) system that can handle both speech and text within a unified architecture. The system follows the Mask-CTC non-autoregressive approach, and the input flexibility is acquired by partially sharing the decoder between SLU and NLU tasks. Experiments on the SLURP dataset show that the proposed architecture achieves similar performance to using separate E2E SLU and NLU modules, but with relatively 43.7 % less model parameters. We also explore the use of pre-trained speech and language models into the SNLU system, and show that they further improve the performance. Mohan Li, Catalin Zorila, Cong-Thanh Do, Rama Sanand Doddipatla |
ASRU | 4 |
| 2023 | Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting DiarizationabstractUsing a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2023 | Cumulative Attention Based Streaming Transformer ASR with Internal Language Model Joint Training and RescoringabstractThis paper presents an approach to improve the performance of streaming Transformer ASR by introducing an internal language model (ILM) as a part of the decoder layers. In the recently pro- posed cumulative attention (CA) based streaming ASR system, only the last or top few decoder layers are equipped with the CA module. Thus in this work, we propose to train the bottom (non-CA) layers as an ILM using an auxiliary LM loss jointly with the rest of the system. During inference, the outputs of the ILM are interpolated with those of the entire Transformer decoder as done in the conventional external language model (ELM) rescoring. The paper also proposes a refinement to the CA algorithm known as CTC look-ahead, in order to improve the precision of endpoint detection. Experiments conducted on AIShell-1, Aidatatang and Librispeech datasets show that the proposed ILM rescoring method achieves on par or better ASR performance when compared to the ELM rescoring baseline. Also, the CTC look-ahead strategy effectively alleviates the early end-of- speech (EOS) triggering issue suffered by the CA module, without bringing noticeable latency degradation. Mohan Li, Cong-Thanh Do, Rama Sanand Doddipatla |
ICASSP | 3 |
| 2023 | On the Effectiveness of Monoaural Target Source Extraction for Distant end-to-end Automatic Speech RecognitionabstractRecent work on enhancement has shown that frequency domain methods may outperform the time domain approaches, while most of the prior art is focused on reporting objective enhancement metrics on simulated noisy data or use less modern hybrid acoustic models for evaluation. In this paper we investigate the effectiveness of target source extraction for improving the robustness of end-to-end automatic speech recognition in noisy and reverberant conditions. A frequency domain source extraction approach is introduced and compared against a state-of-the-art time domain method using several publicly available simulated and real noisy speech test sets. The results show that the frequency domain method outperforms the time domain one only for simulated conditions, and that it is more stable to window size variations. The experiments also indicate that remixing the unprocessed signal with the enhanced speech (referred to as speaker/source reinforcement) yields similar or better results than by using a matched acoustic model retrained using distortions introduced by enhancement. Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 2 |
| 2023 | A Teacher-Student Approach for Extracting Informative Speaker Embeddings From Speech MixturesabstractWe introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture.To allow for supervised training, a teacherstudent approach is employed: the teacher computes the target embeddings from each speaker's utterance before the utterances are added to form the mixture, and the student embedding extractor is then tasked to reproduce those embeddings from the speech mixture at its input.The system much more reliably verifies the presence or absence of a given speaker in a mixture than a conventional speaker embedding extractor, and even exhibits comparable performance to a multi-channel approach that exploits spatial information for embedding extraction.Further, it is shown that a speaker embedding computed from a mixture can be used to check for the presence of that speaker in another mixture. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2023 | Domain Adaptive Self-supervised Training of Automatic Speech Recognition
Cong-Thanh Do, Rama Sanand Doddipatla, Mohan Li, Thomas Hain |
INTERSPEECH | 2 |
| 2022 | Transformer-Based Streaming ASR with Cumulative AttentionabstractIn this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monotonic chunk-wise attention (MoChA) and head-synchronous decoder-end adaptive computation steps (HS-DACS) algorithms, CA triggers the ASR outputs based on the acoustic information accumulated at each encoding timestep, where the decisions are made using a trainable device, referred to as halting selector. In CA, all the attention heads of the same decoder layer are synchronised to have a unified halting position. This feature effectively alleviates the problem caused by the distinct behaviour of individual heads, which may otherwise give rise to severe latency issues as encountered by MoChA. The ASR experiments conducted on AIShell-1 and Librispeech datasets demonstrate that the proposed CA-based Transformer system can achieve on par or better performance with significant reduction in latency during inference, when compared to other streaming Transformer systems in literature. Mohan Li, Shucong Zhang, Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 4 |
| 2022 | Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech RecognitionabstractImproving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they typically require that the ASR model is retrained to cope with the processing artifacts. In this paper we explore a speaker reinforcement strategy for improving recognition performance without retraining the acoustic model (AM). This is achieved by remixing the enhanced signal with the unprocessed input to alleviate the processing artifacts. We evaluate the proposed approach using a DNN speaker extraction based speech denoiser trained with a perceptually motivated loss function. Results show that (without AM retraining) our method yields about 23% and 25% relative accuracy gains compared with the unprocessed for the monoaural simulated and real CHiME-4 evaluation sets, respectively, and outperforms a state-of-the-art reference method. Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 2 |
| 2022 | Multiple-hypothesis RNN-T Loss for Unsupervised Fine-tuning and Self-training of Neural TransducerabstractThis paper proposes a new approach to perform unsupervised fine-tuning and self-training using unlabeled speech data for recurrent neural network (RNN)-Transducer (RNN-T) end-to-end (E2E) automatic speech recognition (ASR) systems.Conventional systems perform fine-tuning/self-training using ASR hypothesis as the targets when using unlabeled audio data and are susceptible to the ASR performance of the base model.Here in order to alleviate the influence of ASR errors while using unlabeled data, we propose a multiple-hypothesis RNN-T loss that incorporates multiple ASR 1-best hypotheses into the loss function.For the fine-tuning task, ASR experiments on Librispeech show that the multiple-hypothesis approach achieves a relative reduction of 14.2% word error rate (WER) when compared to the single-hypothesis approach, on the test other set.For the self-training task, ASR models are trained using supervised data from Wall Street Journal (WSJ), Aurora-4 along with CHiME-4 real noisy data as unlabeled data.The multiplehypothesis approach yields a relative reduction of 3.3% WER on the CHiME-4's single-channel real noisy evaluation set when compared with the single-hypothesis approach. Cong-Thanh Do, Mohan Li, Rama Sanand Doddipatla |
INTERSPEECH | 3 |
| 2022 | Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition
Mohan Li, Rama Sanand Doddipatla, Catalin Zorila |
INTERSPEECH | 2 |
| 2022 | On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant trainingabstractIn this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired clean speech and real noisy data. It is found that the unpaired clean speech is crucial to improve quality of separated speech from real noisy speech. The proposed method also performs remixing of processed and unprocessed signals to alleviate the processing artifacts. Experiments on the single-channel CHiME-3 real test sets show that the proposed method improves significantly in terms of speech recognition performance over the enhancement system trained either on the mismatched simulated data in a supervised fashion or on the matched real data in an unsupervised fashion. Between 16% and 39% relative WER reduction has been achieved by the proposed system compared to the unprocessed signal using end-to-end and hybrid acoustic models without retraining on distorted data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
INTERSPEECH | 3 |
| 2022 | Combining Structured and Unstructured Knowledge in an Interactive Search Dialogue SystemabstractUsers of interactive search dialogue systems specify their preferences with natural language utterances.However, a schema-driven system is limited to handling the preferences that correspond to the predefined database content.In this work, we present a methodology for extending a schema-driven interactive search dialogue system with the ability to handle unconstrained user preferences.Using unsupervised semantic similarity metrics and text snippets associated with the search items, the system identifies suitable items for the user's unconstrained natural language query.In a crowd-sourced evaluation, the users were asked to chat with our extended restaurant search system.Based on objective metrics and subjective user ratings, we demonstrate the feasibility of using this unsupervised low latency approach to extend a schema-driven search dialogue system to handle unconstrained user preferences. Svetlana Stoyanchev, Suraj Pandey, Simon Keizer, Norbert Braunschweiler, Rama Sanand Doddipatla |
SIGDIAL | 5 |
| 2022 | Non-Autoregressive End-to-End Approaches for Joint Automatic Speech Recognition and Spoken Language UnderstandingabstractThis paper presents the use of non-autoregressive (NAR) approaches for joint automatic speech recognition (ASR) and spoken language understanding (SLU) tasks. The proposed NAR systems employ a Conformer encoder that applies connectionist temporal classification (CTC) to transcribe the speech utterance into raw ASR hypotheses, which are further refined with a bidirectional encoder representations from Transformers (BERT)-like decoder. In the meantime, the intent and slot labels of the utterance are predicted simultaneously using the same decoder. Both Mask-CTC and self-conditioned CTC (SC-CTC) approaches are explored for this study. Experiments conducted on the SLURP dataset show that the proposed SC-Mask-CTC NAR system achieves 3.7% and 3.2% absolute gains in SLU metrics and a competitive level of ASR accuracy, when compared to a Conformer-Transformer based autoregressive (AR) model. Additionally, the NAR systems achieve 6× faster decoding speed than the AR baseline. Mohan Li, Rama Sanand Doddipatla |
SLT | 2 |
| 2022 | Factors in Emotion Recognition With Deep Learning Models Using Speech and Text on Multiple CorporaabstractEmotion recognition performance of deep learning models is influenced by multiple factors such as acoustic condition, textual content, style of emotion expression (e.g. acted, natural), etc. In this paper, multiple factors are analysed by training and evaluating state-of-the-art deep learning models using the input modalities speech, text, and their combination across 6 emotional speech corpora. A novel deep learning model architecture is presented that further improves the state-of-the-art in multimodal emotion recognition with speech and text on the IEMOCAP corpus. Results from models trained on individual corpora show that combining speech and text improves performance only on corpora where the text of utterances varies across different emotions, while it reduced performance on corpora with fixed text expressed in different emotions, where the speech-only models performed better. Further, cross-corpus investigations are presented to understand the robustness to changing acoustic and textual content. Results show that models perform significantly better in matched conditions in particular single corpus models perform better than multi-corpus models, with the latter showing a tendency to be more robust to acoustic variations, while performance still depends on characteristics of both training corpora and test corpus. Norbert Braunschweiler, Rama Sanand Doddipatla, Simon Keizer, Svetlana Stoyanchev |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Study on Cross-Corpus Speech Emotion Recognition and Data AugmentationabstractModels that can handle a wide range of speakers and acoustic conditions are essential in speech emotion recognition (SER). Often, these models tend to show mixed results when presented with speakers or acoustic conditions that were not visible during training. This paper investigates the impact of cross-corpus data complementation and data augmentation on the performance of SER models in matched (test-set from same corpus) and mismatched (test-set from different corpus) conditions. Investigations using six emotional speech corpora that include single and multiple speakers as well as variations in emotion style (acted, elicited, natural) and recording conditions are presented. Observations show that, as expected, models trained on single corpora perform best in matched conditions while performance decreases between 10-40% in mismatched conditions, depending on corpus specific features. Models trained on mixed corpora can be more stable in mismatched contexts, and the performance reductions range from 1 to 8% when compared with single corpus models in matched conditions. Data augmentation yields additional gains up to 4% and seem to benefit mismatched conditions more than matched ones. Norbert Braunschweiler, Rama Sanand Doddipatla, Simon Keizer, Svetlana Stoyanchev |
ASRU | 2 |
| 2021 | Dialogue Strategy Adaptation to New Action Sets Using Multi-Dimensional ModellingabstractA major bottleneck for building statistical spoken dialogue systems for new domains and applications is the need for large amounts of training data. To address this problem, we adopt the multi-dimensional approach to dialogue management and evaluate its potential for transfer learning. Specifically, we exploit pre-trained task-independent policies to speed up training for an extended task-specific action set, in which the single summary action for requesting a slot is replaced by multiple slot-specific request actions. Policy optimisation and evaluation experiments using an agenda-based user simulator show that with limited training data, much better performance levels can be achieved when using the proposed multi-dimensional adaptation method. We confirm this improvement in a crowd-sourced human user evaluation of our spoken dialogue system, comparing partially trained policies. The multi-dimensional system (with adaptation on limited training data in the target scenario) outperforms the one-dimensional baseline (without adaptation on the same amount of training data) by 7% perceived success rate. Simon Keizer, Norbert Braunschweiler, Svetlana Stoyanchev, Rama Sanand Doddipatla |
ASRU | 4 |
| 2021 | Improving HS-DACS Based Streaming Transformer ASR with Deep Reinforcement LearningabstractTransformer-based systems, though have achieved state-of-the-art performance on a wide range of automatic speech recognition (ASR) tasks, are subject to severe latency issues during inference that limit their deployment in real world applications. To enable online decoding, we recently proposed Decoder-end adaptive computation steps (DACS) and the head-synchronous version (HS-DACS) algorithms, which were shown to reduce the computational cost for decoding and close the gap in performance between offline and streaming Transformer ASR. In DACS/HS-DACS based systems, the halting position that triggers the ASR output was determined with an accumulation threshold that was set arbitrarily. In this paper, we propose a deep reinforcement learning approach to improve HS-DACS where an external agent is utilised to optimise the halting position. Experiments on AIShell-1, Tedlium-2 and Librispeech datasets show that the proposed method can further cut down the computation cost in inference with relative gains of 40.0%, 32.8% and 53.4% respectively, and still maintain similar ASR performance when compared with HS-DACS. Mohan Li, Rama Sanand Doddipatla |
ASRU | 2 |
| 2021 | Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech RecognitionabstractThis paper proposes an adaptation method for end-to-end speech recognition. In this method, multiple automatic speech recognition (ASR) 1-best hypotheses are integrated in the computation of the connectionist temporal classification (CTC) loss function. The integration of multiple ASR hypotheses helps alleviating the impact of errors in the ASR hypotheses to the computation of the CTC loss when ASR hypotheses are used. When being applied in semi-supervised adaptation scenarios where part of the adaptation data do not have labels, the CTC loss of the proposed method is computed from different ASR 1-best hypotheses obtained by decoding the unlabeled adaptation data. Experiments are performed in clean and multi-condition training scenarios where the CTC-based end-to-end ASR systems are trained on Wall Street Journal (WSJ) clean training data and CHiME-4 multi-condition training data, respectively, and tested on Aurora-4 test data. The proposed adaptation method yields 6.6% and 5.8% relative word error rate (WER) reductions in clean and multi-condition training scenarios, respectively, compared to a baseline system which is adapted with part of the adaptation data having manual transcriptions using back-propagation fine-tuning. Cong-Thanh Do, Rama Sanand Doddipatla, Thomas Hain |
ICASSP | 2 |
| 2021 | Head-Synchronous Decoding for Transformer-Based Streaming ASRabstractOnline Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder-end Adaptive Computation Steps (DACS) algorithm for online Transformer ASR was shown to achieve state-of-the-art performance and outperform other existing methods. However, like any other online approach, the DACS-based attention heads in each of the Transformer decoder layers operate independently (or asynchronously) and lead to diverged attending positions. Since DACS employs a truncation threshold to determine the halting position, some of the attention weights are cut off untimely and might impact the stability and precision of decoding. To overcome these issues, here we propose a head-synchronous (HS) version of the DACS algorithm, where the boundary of attention is jointly detected by all the DACS heads in each decoder layer. ASR experiments on Wall Street Journal (WSJ), AIShell-1 and Lib- rispeech show that the proposed method consistently outperforms vanilla DACS and achieves state-of-the-art performance. We will also demonstrate that HS-DACS has reduced decoding cost when compared to vanilla DACS. Mohan Li, Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 3 |
| 2021 | Action State Update Approach to Dialogue ManagementabstractUtterance interpretation is one of the main functions of a dialogue manager, which is the key component of a dialogue system. We propose the action state update approach (ASU) for utterance interpretation, featuring a statistically trained binary classifier used to detect dialogue state update actions in the text of a user utterance. Our goal is to interpret referring expressions in user input without a domain-specific natural language understanding component. For training the model, we use active learning to automatically select simulated training examples. With both user-simulated and interactive human evaluations, we show that the ASU approach successfully interprets user utterances in a dialogue system, including those with referring expressions. Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla |
ICASSP | 3 |
| 2021 | Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower LayersabstractAlthough the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in general, freezing the trained feature extractor (the lower layers) and retraining the classifier (the upper layers) on the same dataset leads to worse performance. In this paper, for the first time, we show that the frozen classifier is transferable within the same dataset. We develop a novel top-down training method which can be viewed as an algorithm for searching for high-quality classifiers. We tested this method on automatic speech recognition (ASR) tasks and language modelling tasks. The proposed method consistently improves recurrent neural network ASR models on Wall Street Journal, self-attention ASR models on Switchboard, and AWD-LSTM language models on WikiText-2. Shucong Zhang, Cong-Thanh Do, Rama Sanand Doddipatla, Erfan Loweimi, Peter Bell 0001, Steve Renals |
ICASSP | 3 |
| 2021 | Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning MechanismabstractIn this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs speaker embeddings to identify and extract multiple targets without label permutation ambiguity. To efficiently inform the speaker information to the extraction model, we propose a new speaker conditioning mechanism by designing an additional speaker branch for receiving external speaker embeddings. Experiments on 2-channel WHAMR! data show that the proposed system improves by 9% relative the source separation performance over a strong multi-channel baseline, and it increases the speech recognition accuracy by more than 16% relative over the same baseline. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
ICASSP | 3 |
| 2021 | Teacher-Student MixIT for Unsupervised and Semi-Supervised Speech SeparationabstractIn this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixture invariant training (MixIT) criterion to train a teacher model. The teacher model then estimates separated sources that are used to train a student model with standard permutation invariant training (PIT). The student model can be fine-tuned with supervised data, i.e., paired artificial mixtures and clean speech sources, and further improved via model distillation. Experiments with single and multi channel mixtures show that the teacher-student training resolves the over-separation problem observed in the original MixIT method. Further, the semisupervised performance is comparable to a fully-supervised separation system trained using ten times the amount of supervised data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
Interspeech | 3 |
| 2021 | Transformer-Based Online Speech Recognition with Decoder-end Adaptive Computation StepsabstractTransformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent structures on a number of ASR tasks. However, like other E2E models, Transformer ASR also requires the full input sequence for calculating the attentions on both encoder and decoder, leading to increased latency and posing a challenge for online ASR. The paper proposes Decoder-end Adaptive Computation Steps (DACS) algorithm to address the issue of latency and facilitate online ASR. The proposed algorithm streams the decoding of Transformer ASR by triggering an output after the confidence acquired from the encoder states reaches a certain threshold. Unlike other monotonic attention mechanisms that risk visiting the entire encoder states for each output step, the paper introduces a maximum look-ahead step into the DACS algorithm to prevent from reaching the end of speech too fast. A Chunkwise en-coder is adopted in our system to handle real-time speech inputs. The proposed online Transformer ASR system has been evaluated on Wall Street Journal (WSJ) and AIShell-1 datasets, yielding 5.5% word error rate (WER) and 7.1% character error rate (CER) respectively, with only a minor decay in performance when compared to the offline systems. Mohan Li, Catalin Zorila, Rama Sanand Doddipatla |
SLT | 3 |
| 2021 | An Investigation into the Multi-channel Time Domain Speaker Extraction NetworkabstractThis paper presents an investigation into the effectiveness of spatial features for improving time-domain speaker extraction systems. A two-dimensional Convolutional Neural Network (CNN) based encoder is proposed to capture the spatial information within the multichannel input, which are then combined with the spectral features of a single channel extraction network. Two variants of target speaker extraction methods were tested, one which employs a pre-trained i-vector system to compute a speaker embedding (System A), and one which employs a jointly trained neural network to extract the embeddings directly from time domain enrolment signals (System B). The evaluation was performed on the spatialized WSJ0-2mix dataset using the Signal-to-Distortion Ratio (SDR) metric, and ASR accuracy. In the anechoic condition, more than 10 dB and 7 dB absolute SDR gains were achieved when the 2-D CNN spatial encoder features were included with Systems A and B, respectively. The performance gains in reverberation were lower, however, we have demonstrated that retraining the systems by applying dereverberation preprocessing can significantly boost both the target speaker extraction and ASR performances. Catalin Zorila, Mohan Li, Rama Sanand Doddipatla |
SLT | 3 |
| 2020 | Learning Noise Invariant Features Through Transfer Learning For Robust End-to-End Speech RecognitionabstractEnd-to-end models yield impressive speech recognition results on clean datasets while having inferior performance on noisy datasets. To address this, we propose transfer learning from a clean dataset (WSJ) to a noisy dataset (CHiME4) for connectionist temporal classification models. We argue that the clean classifier (the upper layers of a neural network trained on clean data) can force the feature extractor (the lower layers) to learn the underlying noise invariant patterns in the noisy dataset. While training on the noisy dataset, the clean classifier is either frozen or trained with a small learning rate. The feature extractor is trained with no learning rate re-scaling. The proposed method gives up to 15.5% relative character error rate (CER) reduction compared to models trained only on CHiME-4. Furthermore, we use the test sets of Aurora-4 to perform evaluation on unseen noisy conditions. Our method has significantly lower CERs (11.3% relative on average) on all 14 Aurora-4 test sets compared to the conventional transfer learning method (no learning rate rescale for any layer), indicating our method enables the model to learn noise invariant features. Shucong Zhang, Cong-Thanh Do, Rama Sanand Doddipatla, Steve Renals |
ICASSP | 3 |
| 2020 | On End-to-end Multi-channel Time Domain Speech Separation in Reverberant EnvironmentsabstractThis paper introduces a new method for multi-channel time domain speech separation in reverberant environments. A fully-convolutional neural network structure has been used to directly separate speech from multiple microphone recordings, with no need of conventional spatial feature extraction. To reduce the influence of reverberation on spatial feature extraction, a dereverberation pre-processing method has been applied to further improve the separation performance. A spatialized version of wsj0-2mix dataset has been simulated to evaluate the proposed system. Both source separation and speech recognition performance of the separated signals have been evaluated objectively. Experiments show that the proposed fully-convolutional network improves the source separation metric and the word error rate (WER) by more than 13% and 50% relative, respectively, over a reference system with conventional features. Applying dereverberation as pre-processing to the proposed system can further reduce the WER by 29% relative using an acoustic model trained on clean and reverberated data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
ICASSP | 3 |
| 2019 | An Investigation into the Effectiveness of Enhancement in ASR Training and Test for Chime-5 Dinner Party TranscriptionabstractDespite the strong modeling power of neural network acoustic models, speech enhancement has been shown to deliver additional word error rate improvements if multi-channel data is available. However, there has been a longstanding debate whether enhancement should also be carried out on the ASR training data. In an extensive experimental evaluation on the acoustically very challenging CHiME-5 dinner party data we show that: (i) cleaning up the training data can lead to substantial error rate reductions, and (ii) enhancement in training is advisable as long as enhancement in test is at least as strong as in training. This approach stands in contrast and delivers larger gains than the common strategy reported in the literature to augment the training database with additional artificially degraded speech. Together with an acoustic model topology consisting of initial CNN layers followed by factorized TDNN layers we achieve with 41.6 % and 43.2 % WER on the DEV and EVAL test sets, respectively, a new single-system state-of-the-art result on the CHiME-5 data. This is a 8 % relative improvement compared to the best word error rate published so far for a speech recognizer without system combination. Catalin Zorila, Christoph Böddeker, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ASRU | 3 |
| 2019 | An Unsupervised Learning Approach to Neural-net-supported Wpe DereverberationabstractReverberation degrades signal quality and increases word error rates in automatic speech recognition (ASR). Reverberation suppression is, thus, a key component in listening enhancement devices and ASR front end. The weighted prediction error (WPE) is a prominent and effective method that gained popularity in recent ASR challenges. The need for iterative optimization in WPE leads to high computational cost and instabilities for short signals. Neural net (NN) supported WPE was proposed to alleviate these issues. However, NN training requires parallel data, i.e., reverberant and "clean" (direct sound plus early reflections) speech, which is not available in general. We show that the supporting network can be trained efficiently, without any supervision, using reverberant speech only. Consequently, adaptation to unseen environments is largely simplified. Network training involves the complete de-reverberation system and relies on complex-valued back propagation. The experimental validation confirms that, the proposed approach matches the performance of the method with parallel training data both in terms of perceptual quality and ASR word error rates. Petko Nikolov Petkov, Vasileios Tsiaras, Rama Sanand Doddipatla, Yannis Stylianou |
ICASSP | 3 |
| 2019 | On Reducing the Effect of Speaker Overlap for Chime-5abstractThe CHiME-5 speech separation and recognition challenge was recently shown to pose a difficult task for the current automatic speech recognition systems. Speaker overlap was one of the main difficulties of the challenge. The presence of noise, reverberation and the moving speakers have made the traditional source separation methods ineffective in improving the recognition accuracy. In this paper we have explored several enhancement strategies aimed to reduce the effect of speaker overlap for CHiME-5 without performing source separation. One is based on discarding the overlap segments using the speaker diarisation information from the challenge, another one is a neural network driven automatic gain control enhancement aimed to improve the previous speaker diarisation information, and the last one is based on optimal multi-array data selection. State-of-the-art acoustic models were used to perform the ASR experiments. Results have shown that proposed automatic gain control method yields word error rate (WER) reductions between 2% and 3% absolute on the development set of CHiME-5. Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 2 |
| 2017 | Speaker Adaptation in DNN-Based Speech Synthesis Using d-Vectors
Rama Sanand Doddipatla, Norbert Braunschweiler, Ranniery Maia |
INTERSPEECH | 1 |
| 2016 | Speaker adaptive training in deep neural networks using speaker dependent bottleneck featuresabstractThe paper proposes an approach to perform speaker adaptive training (SAT) in deep neural networks using a two-stage DNN. The first-stage DNN extracts speaker dependent bottleneck (SDBN) features by updating the weights of the BN layer with speaker specific data. Using the SDBN features, a second-stage DNN is trained in the SAT framework. Choosing the BN layer as the speaker dependent layer instead of one of the hidden layers reduces the number of parameters to be tuned using speaker specific data. Experiments are presented on the Aurora4 task, where the input features are normalised with constrained maximum likelihood linear regression (CMLLR) and speaker information is appended in the form of D-vectors. Following an unsupervised adaptation of BN layer, the proposed approach provides a relative gain of 8.6% and 8.9% WER on top of DNNs trained with FBANK features appended with and without D-vectors respectively. A relative gain of 10.3% WER is observed when applied on top of DNNs trained with CMLLR transformed FBANK features, but the gain in performance saturated when combined with D-vectors. It is observed that supervised adaptation with as little as one minute of audio from a specific speaker improved the performance when compared with the baseline. Rama Sanand Doddipatla |
ICASSP | 1 |
| 2015 | Noise-matched training of CRF based sentence end detection models
Madina Hasan, Rama Sanand Doddipatla, Thomas Hain |
INTERSPEECH | 2 |
| 2014 | Speaker dependent bottleneck layer training for speaker adaptation in automatic speech recognitionabstractSpeaker adaptation of deep neural networks (DNN) is difficult, and most commonly performed by changes to the input of the DNNs. Here we propose to learn discriminative feature transformations to obtain speaker normalised bottleneck (BN) features. This is achieved by interpreting the final two hidden layers as speaker specific matrix transformations. The hidden layer weights are updated with data from a specific speaker to learn speaker-dependent discriminative feature transformations. Such simple implementation lends itself to rapid adaptation and flexibility to be used in Speaker Adaptive Training (SAT) frameworks. The performance of this approach is evaluated on a meeting recognition task, using the official NIST RT’07 and RT’09 evaluation test sets. Supervised adaptation of the BN layer shows similar performance to the application of supervised CMLLR as a global transformation, and the combination of these appears to be additive. In unsupervised mode, CMLLR adaptation only yields 3.4% and 2.5% relative word error rate (WER) improvement, on the RT’07 and RT’09 respectively, where the baselines include speaker based cepstral mean and variance normalisation. The combined CMLLR and BN layer speaker adaptation yields a relative WER gain of 4.5% and 4.2% respectively. SAT style BN layer adaptation is attempted and combined with conventional CMLLR SAT, to show that it provides a relative gain of 1.43% and 2.02% on the RT’07 and RT’09 data sets respectively when compared with CMLLR SAT. While the overall gain from BN layer adaptation is small, the results are found to be statistically significant on both the test sets. Index Terms: Deep neural networks, bottleneck features, speaker adaptation, automatic speech recognition. Rama Sanand Doddipatla, Madina Hasan, Thomas Hain |
INTERSPEECH | 1 |
| 2014 | Multi-pass sentence-end detection of lecture speech
Madina Hasan, Rama Sanand Doddipatla, Thomas Hain |
INTERSPEECH | 2 |
| 2013 | Synthetic speaker models using VTLN to improve the performance of children in mismatched speaker conditions for ASR
Rama Sanand Doddipatla, Torbjørn Svendsen |
INTERSPEECH | 1 |
| 2012 | Creating synthetic voices for children by adapting adult average voice using stacked transformations and VTLNabstractThis paper describes experiments in creating personalised children's voices for HMM-based synthesis by adapting either an adult or child average voice. The adult average voice is trained from a large adult speech database, whereas the child average voice is trained using a small database of children's speech. Here we present the idea to use stacked transformations for creating synthetic child voices, where the child average voice is first created from the adult average voice through speaker adaptation using all the pooled speech data from multiple children and then adding child specific speaker adaptation on top of it. VTLN is applied to speech synthesis to see whether it helps the speaker adaptation when only a small amount of adaptation data is available. The listening test results show that the stacked transformations significantly improve speaker adaptation for small amounts of data, but the additional benefit provided by VTLN is not yet clear. Reima Karhila, Rama Sanand Doddipatla, Mikko Kurimo, Peter Smit |
ICASSP | 2 |
| 2012 | VTLN Using Analytically Determined Linear-Transformation on Conventional MFCCabstractIn this paper, we propose a method to analytically obtain a linear-transformation on the conventional Mel frequency cepstral coefficients (MFCC) features that corresponds to conventional vocal tract length normalization (VTLN)-warped MFCC features, thereby simplifying the VTLN processing. There have been many attempts to obtain such a linear-transformation, but all the previously proposed approaches either modify the signal processing (and therefore not conventional MFCC), or the linear-transformation does not correspond to conventional VTLN-warping, or the matrices being estimated and are data dependent. In short, the conventional VTLN part of an automatic speech recognition (ASR) system cannot be simply replaced with any of the previously proposed methods. Umesh proposed the idea to use band-limited interpolation for performing VTLN-warping on MFCC using plain cepstra. Motivated from this work, Panchapagesan and Alwan proposed a linear-transformation to perform VTLN-warping on conventional MFCC. However, in their approach, VTLN warping is specified in the Mel-frequency domain and is not equivalent to conventional VTLN. In this paper, we present an approach which also draws inspiration from the work of Umesh , and which we believe for the first time performs conventional VTLN as a linear-transformation on conventional MFCC using the ideas of band-limited interpolation. Deriving such a linear-transformation to perform VTLN, would allow us to use the VTLN-matrices in transform-based adaptation framework with its associated advantages and yet would require the estimation of a single parameter. Using four different tasks, we show that our proposed approach has almost identical recognition performance to conventional VTLN on both clean and noisy speech data. Rama Sanand Doddipatla, Srinivasan Umesh |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | A Study on Combining VTLN and SAT to Improve the Performance of Automatic Speech Recognition
Rama Sanand Doddipatla, Mikko Kurimo |
INTERSPEECH | 1 |
| 2010 | Revisiting VTLN using linear transformation on conventional MFCCabstractIn this paper, we revisit the linear transformation for VTLN on conventional MFCC proposed by Sanand et al. in [1], using the idea of band-limited interpolation.The filter-bank is modified to include half-filters at zero and nyquist frequencies, as the full symmetric spectrum is required for performing bandlimited interpolation.In this paper, we show that the filter-bank with half-filters does not affect the recognition performance on clean speech (also shown in [1]), but does affect the recognition performance on noisy speech.This motivated us to revisit the linear transformation for VTLN in [1] and propose modifications to undo the affect of half-filters during the feature extraction.We show through recognition experiments that the proposed modifications to the linear transformation have comparable performance as the conventional VTLN approach, still enabling us to perform VTLN using a linear transformation on conventional MFCC. Rama Sanand Doddipatla, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2009 | Improving the performance of VTLN under mismatched speaker conditions and making it approach that of matched speaker conditionsabstractThe performance of conventional VTLN for mis-matched train and test speaker conditions (e.g. adult-train child-test) does not approach the performance of matched speaker conditions (e.g. child-train child-test). In this paper, we investigate this problem and propose methods to reduce this gap in performance. We use our recently proposed linear transformation approach to VTLN, that also enables us to study the effect of Jacobian unlike conventional VTLN. The main advantage of transform-based VTLN over adaptation based approaches (like CMLLR), is that it does not require any matrix estimation. We argue that the degraded VTLN performance under mismatched speaker conditions is due to the significant frequency warping that is necessary for normalization which leads to a mis-match between the correlation in the feature components of the test data and the covariance structure of the trained/normalized model. We show that the use of a global de-correlating transform (MLLT) leads to improved VTLN performance. We finally show that using both Jacobian and MLLT together improves the VTLN performance for mis-matched cases with the performance approaching that of matched speaker conditions. Rama Sanand Doddipatla, Shakti Prasad Rath, Srinivasan Umesh |
ICASSP | 1 |
| 2009 | Characterizing speaker variability using spectral envelopes of vowel sounds
A. N. Harish, Rama Sanand Doddipatla, Srinivasan Umesh |
INTERSPEECH | 2 |
| 2009 | A study on the influence of covariance adaptation on jacobian compensation in vocal tract length normalization
Rama Sanand Doddipatla, Shakti Prasad Rath, Srinivasan Umesh |
INTERSPEECH | 1 |
| 2008 | A computationally efficient approach to warp factor estimation in VTLN using EM algorithm and sufficient statistics
P. T. Akhil, Shakti Prasad Rath, Srinivasan Umesh, Rama Sanand Doddipatla |
INTERSPEECH | 4 |
| 2008 | Use of spectral centre of gravity for generating speaker invariant features for automatic speech recognition
Rama Sanand Doddipatla, Rani R. Sandhya, Srinivasan Umesh |
INTERSPEECH | 1 |
| 2008 | Study of jacobian compensation using linear transformation of conventional MFCC for VTLN
Rama Sanand Doddipatla, Srinivasan Umesh |
INTERSPEECH | 1 |
| 2007 | Speaker-Invariant Features for Automatic Speech Recognition
Srinivasan Umesh, Rama Sanand Doddipatla, G. Praveen |
IJCAI | 2 |
| 2007 | Linear transformation approach to VTLN using dynamic frequency warping
Rama Sanand Doddipatla, D. Dinesh Kumar, Srinivasan Umesh |
INTERSPEECH | 1 |