Katrin Kirchhoff

dblp:21/5031 · DBLP profile ↗
← Back
100ranked-venue papers
20as first author
23since 2021 · last 2025
0000-0002-6645-6030ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 66 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 13 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation
abstract
Existing automatic prompt engineering methods are typically designed for discriminative tasks, where new task prompts are iteratively refined with limited feedback from a single metric reflecting a single aspect. However, these approaches are suboptimal for generative tasks, which require more nuanced guidance beyond a single numeric metric to improve the prompt and optimize multiple aspects of the generated text. To address these challenges, we propose a novel multi-aspect Critique-Suggestion-guided automatic Prompt Optimization (CriSPO) approach. CriSPO introduces a critique-suggestion module as its core component. This module spontaneously discovers aspects, and compares generated and reference texts across these aspects, providing specific suggestions for prompt modification. These clear critiques and actionable suggestions guide a receptive optimizer module to make more substantial changes, exploring a broader and more effective search space. To further improve CriSPO with multi-metric optimization, we introduce an Automatic Suffix Tuning (AST) extension to enhance the performance of task prompts across multiple metrics. We evaluate CriSPO on 4 state-of-the-art Large Language Models (LLMs) across 4 summarization and 5 Question Answering (QA) datasets. Extensive experiments show 3-4% ROUGE score improvement on summarization and substantial improvement of various metrics on QA.
Han He, Qianchu Liu, Chaitanya P. Shivade, Sundararajan Srinivasan, Katrin Kirchhoff
AAAI7
2025 DeAL: Decoding-time Alignment for Large Language Models
abstract
James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchhoff, Dan Roth. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas 0004, Saab Mansour, Katrin Kirchhoff, Dan Roth 0001
ACL (1)8
2025 Zero-resource Speech Translation and Recognition with LLMs
abstract
Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language.
Karel Mundnich, Xing Niu 0001, Prashant Mathur, Srikanth Ronanki, Brady Houston, Veera Raghavendra Elluru, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff
ICASSP13
2024 Revisiting Convolution-free Transformer for Speech Recognition
Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff
INTERSPEECH6
2023 Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale
abstract
Hritik Bansal, Karthik Gopalakrishnan, Saket Dingliwal, Sravan Bodapati, Katrin Kirchhoff, Dan Roth. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hritik Bansal, Karthik Gopalakrishnan 0001, Saket Dingliwal, Sravan Babu Bodapati, Katrin Kirchhoff, Dan Roth 0001
ACL (1)5
2023 Mask the Bias: Improving Domain-Adaptive Generalization of CTC-Based ASR with Internal Language Model Estimation
abstract
End-to-end ASR models trained on large amount of data tend to be implicitly biased towards language semantics of the training data. Internal language model estimation (ILME) has been proposed to mitigate this bias for autoregressive models such as attention-based encoder-decoder and RNN-T. Typically, ILME is performed by modularizing the acoustic and language components of the model architecture, and eliminating the acoustic input to perform log-linear interpolation with the text-only posterior. However, for CTC-based ASR, it is not as straightforward to decouple the model into such acoustic and language components, as CTC log-posteriors are computed in a non-autoregressive manner. In this work, we propose a novel ILME technique for CTC-based ASR models. Our method iteratively masks the audio timesteps to estimate a pseudo log-likelihood of the internal LM by accumulating log-posteriors for only the masked timesteps. Extensive evaluation across multiple out-of-domain datasets reveals that the proposed approach improves WER by up to 9.8% and OOV F1-score by up to 24.6% relative to Shallow Fusion, when only text data from target domain is available. In the case of zero-shot domain adaptation, with no access to any target domain data, we demonstrate that removing the source domain bias with ILME can still outperform Shallow Fusion to improve WER by up to 9.3% relative.
Nilaksh Das, Monica Sunkara, Sravan Babu Bodapati, Jinglun Cai, Devang Kulshreshtha, Jeff Farris, Katrin Kirchhoff
ICASSP7
2023 A Metric-Driven Approach to Conformer Layer Pruning for Efficient ASR Inference
Dhanush Bekal, Karthik Gopalakrishnan 0001, Karel Mundnich, Srikanth Ronanki, Sravan Babu Bodapati, Katrin Kirchhoff
INTERSPEECH6
2023 Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters
Anshu Bhatia, Sanchit Sinha, Saket Dingliwal, Karthik Gopalakrishnan 0001, Sravan Babu Bodapati, Katrin Kirchhoff
INTERSPEECH6
2023 DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer
Goeric Huybrechts, Srikanth Ronanki, Xilai Li, Hadis Nosrati, Sravan Babu Bodapati, Katrin Kirchhoff
INTERSPEECH6
2022 Listen, Know and Spell: Knowledge-Infused Subword Modeling for Improving ASR Performance of OOV Named Entities
abstract
Automatic speech recognition (ASR) is increasingly being used in specialized domains such as medical ASR and news transcription. Owing to the lack of high quality annotated speech data in such domains, off-the-shelf models are commonly employed by fine-tuning on domain-specific data. This poses a significant challenge in transcribing long-tail expressions and out-of-vocabulary (OOV) named entities. On the other hand, readily available knowledge graphs (KGs) provide semantically structured knowledge for such domain-specific named entities. In this work, we propose the Knowledge-Infused Subword Model (KISM), a novel technique for incorporating semantic context from KGs into the ASR pipeline for improving the performance of OOV named entities. Our experiments show that KISM improves OOV recall of an ASR model by 4.58% (absolute) for named entities that were not seen during training.
Nilaksh Das, Polo Chau, Monica Sunkara, Sravan Babu Bodapati, Dhanush Bekal, Katrin Kirchhoff
ICASSP6
2022 Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation
abstract
Audio-visual data allows us to leverage different modalities for downstream tasks. The idea being individual streams can complement each other in the given task, thereby resulting in a model with improved performance. In this work, we present our experimental results on action recognition and video summarization tasks. The proposed modeling approach builds upon the recent advances in contrastive loss based audio-visual representation learning. Temporally cognizant audio-visual discrimination is achieved in a Transformer model by learning with a masked feature reconstruction loss over a fixed time window in addition to learning via contrastive loss. Overall, our results indicate that the addition of temporal information significantly improved the performance of the contrastive loss based framework. We achieve an action classification accuracy of 66.2% versus the next best baseline at 64.7% on the HMDB dataset. For video summarization, we attain an F1 score of 43.5 verses 42.2 on the SumMe dataset.
Chandrashekhar Lavania, Shiva Sundaram, Sundararajan Srinivasan, Katrin Kirchhoff
ICASSP4
2022 Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition
abstract
Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to effectively adapt these representations for speech emotion recognition. Recent public benchmarks show the efficacy of several popular self-supervised speech representations for emotion classification. In this study, we show that the primary difference between the top-performing representations is in predicting valence while the differences in predicting activation and dominance dimensions are less pronounced. However, we show that even the best-performing HuBERT representation underperforms on valence prediction compared to a multimodal model that also incorporates text representation. We address this shortcoming by injecting lexical information into the speech representation using the multimodal model as a teacher. To improve the efficacy of our approach, we propose a novel estimate of the quality of the emotion predictions, to condition teacher-student training. We report new audio-only state-of-the-art concordance correlation coefficient (CCC) values of 0.757, 0.627, 0.671 for activation, valence and dominance predictions, respectively, on the MSP-Podcast corpus, and also state-of-the-art values of 0.667, 0.582, 0.545 on the IEMOCAP corpus.
Sundararajan Srinivasan, Zhaocheng Huang, Katrin Kirchhoff
ICASSP3
2022 Contextual Acoustic Barge-In Classification for Spoken Dialog Systems
Dhanush Bekal, Sundararajan Srinivasan, Srikanth Ronanki, Sravan Babu Bodapati, Katrin Kirchhoff
INTERSPEECH5
2022 Domain Prompts: Towards memory and compute efficient domain adaptation of ASR systems
abstract
Automatic Speech Recognition (ASR) systems have found their use in numerous industrial applications in very diverse domains creating a need to adapt to new domains with small memory and deployment overhead. In this work, we introduce domain prompts, a methodology that involves training a small number of domain embedding parameters to prime a Transformer-based Language Model (LM) to a particular domain. Using this domain-adapted LM for rescoring ASR hypotheses can achieve 7-13% WER reduction for a new domain with just 1000 unlabeled textual domain-specific sentences and a handful of additional parameters. Our method can match or even beat the performance of models fully fine-tuned towards a particular domain with only 0.02% of the parameters. Given the parameter efficiency and negligible deployment overhead, our experiments showcase that such a method is an ideal choice for on-the-fly adaptation of LMs used in ASR systems to progressively scale it to new domains.
Saket Dingliwal, Ashish Shenoy, Sravan Babu Bodapati, Ankur Gandhe, Ravi Gadde, Katrin Kirchhoff
INTERSPEECH6
2022 Directed speech separation for automatic speech recognition of long form conversational speech
abstract
Many of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap.Most of these approaches need an additional stitching step to stitch the separated speech chunks for long form audio. Since most of the approaches involve Permutation Invariant training (PIT), the order of separated speech chunks is nondeterministic and leads to difficulty in accurately stitching homogenous speaker chunks for downstream tasks like Automatic Speech Recognition (ASR).Also, most of these models are trained with synthetic mixtures and do not generalize to real conversational data.In this paper, we propose a speaker conditioned separator trained on speaker embeddings extracted directly from the mixed signal using an over-clustering based approach.This model naturally regulates the order of the separated chunks without the need for an additional stitching step.We also introduce a data sampling strategy with real and synthetic mixtures which generalizes well to real conversation speech.With this model and data sampling technique, we show significant improvements in speaker-attributed word error rate (SA-WER) on Hub5 data.
Rohit Paturi, Sundararajan Srinivasan, Katrin Kirchhoff, Daniel Garcia-Romero
INTERSPEECH3
2022 Personalization of CTC Speech Recognition Models
abstract
End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregressive CTC decoder is often used at inference time due to its speed and simplicity. However, such models are hard to personalize because of their conditional independence assumption that prevents output tokens from previous time steps to influence future predictions. To tackle this, we propose a novel two-way approach that first biases the encoder with attention over a predefined list of rare long-tail and out-of-vocabulary (OOV) words and then uses dynamic boosting and phone alignment network during decoding to further bias the subword pre-dictions. We evaluate our approach on open-source VoxPopuli and in-house medical datasets to showcase a 60% improvement in F1 score on domain-specific rare words over a strong CTC baseline.
Saket Dingliwal, Monica Sunkara, Srikanth Ronanki, Jeff Farris, Katrin Kirchhoff, Sravan Babu Bodapati
SLT5
2022 Exploration of Language-Specific Self-Attention Parameters for Multilingual End-to-End Speech Recognition
abstract
In the last several years, end-to-end (E2E) ASR models have mostly surpassed the performance of hybrid ASR models. E2E is particularly well suited to multilingual approaches because it doesn't require language-specific phone alignments for training. Recent work has improved multilingual E2E modeling over naive data pooling on up to several dozen languages by using both language-specific and language-universal model parameters, as well as providing information about the language being presented to the network. Complementary to previous work we analyze language-specific parameters in the attention mechanism of Conformer-based encoder models. We show that using language-specific parameters in the attention mechanism can improve performance across six languages by up to 12% compared to standard multilingual baselines and up to 36% compared to monolingual baselines, without requiring any additional parameters during monolingual inference nor fine-tuning.
Brady Houston, Katrin Kirchhoff
SLT2
2021 Remember the Context! ASR Slot Error Correction Through Memorization
abstract
Accurate recognition of slot values such as domain specific words or named entities by automatic speech recognition (ASR) systems forms the core of the Goal-oriented Dialogue Systems. Although it is a critical step with direct impact on downstream tasks such as language understanding, many domain agnostic ASR systems tend to perform poorly on domain specific or long tail words. They are often supplemented with slot error correcting systems but it is often hard for any neural model to directly output such rare entity words. To address this problem, we propose$k$-nearest neighbor ($k$-NN) search that outputs domain-specific entities from an explicit datastore. We improve error correction rate by conveniently augmenting a pretrained joint phoneme and text based transformer sequence to sequence model with$k$-NN search during inference. We evaluate our proposed approach on five different domains containing long tail slot entities such as full names, airports, street names, cities, states. Our best performing error correction model shows a relative improvement of 7.4% in word error rate (WER) on rare word entities over the baseline and also achieves a relative WER improvement of 9.8% on an out of vocabulary (OOV) test set.
Dhanush Bekal, Ashish Shenoy, Monica Sunkara, Sravan Babu Bodapati, Katrin Kirchhoff
ASRU5
2021 Transformer-Transducers for Code-Switched Speech Recognition
abstract
We live in a world where 60% of the population can speak two or more languages fluently. Members of these communities constantly switch between languages when having a conversation. As automatic speech recognition (ASR) systems are being deployed to the real-world, there is a need for practical systems that can handle multiple languages both within an utterance or across utterances. In this paper, we present an end-to-end ASR system using a transformer-transducer model architecture for code-switched speech recognition. We propose three modifications over the vanilla model in order to handle various aspects of code-switching. First, we introduce two auxiliary loss functions to handle the low-resource scenario of code-switching. Second, we propose a novel mask-based training strategy with language ID information to improve the label encoder training towards intra-sentential code-switching. Finally, we propose a multi-label/multi-audio encoder structure to leverage the vast monolingual speech corpora towards code-switching. We demonstrate the efficacy of our proposed approaches on the SEAME dataset, a public Mandarin-English code-switching corpus, achieving a mixed error rate of 18.5% and 26.3% on testmanand testsgesets respectively.
Siddharth Dalmia, Yuzong Liu, Srikanth Ronanki, Katrin Kirchhoff
ICASSP4
2021 Neural Inverse Text Normalization
abstract
While there have been several contributions exploring state of the art techniques for text normalization, the problem of inverse text normalization (ITN) remains relatively unexplored. The best known approaches leverage finite state transducer (FST) based models which rely on manually curated rules and are hence not scalable. We propose an efficient and robust neural solution for ITN leveraging transformer based seq2seq models and FST-based text normalization techniques for data preparation. We show that this can be easily extended to other languages without the need for a linguistic expert to manually curate them. We then present a hybrid framework for integrating Neural ITN with an FST to overcome common recoverable errors in production environments. Our empirical evaluations show that the proposed solution minimizes incorrect perturbations (insertions, deletions and substitutions) to ASR output and maintains high quality even on out of domain data. A transformer based model infused with pretraining consistently achieves a lower WER across several datasets and is able to outperform baselines on English, Spanish, German and Italian datasets.
Monica Sunkara, Chaitanya P. Shivade, Sravan Babu Bodapati, Katrin Kirchhoff
ICASSP4
2021 Speaker-Conversation Factorial Designs for Diarization Error Analysis
abstract
Speaker diarization accuracy can be affected by both acoustics and conversation characteristics. Determining the cause of diarization errors is difficult because speaker voice acoustics and conversation structure co-vary, and the interactions between acoustics, conversational structure, and diarization accuracy are complex. This paper proposes a methodology that can distinguish independent marginal effects of acoustic and conversation characteristics on diarization accuracy by remixing conversations in a factorial design. As an illustration, this approach is used to investigate gender-related and language-related accuracy differences with three diarization systems: a baseline system using subsegment x-vector clustering, a variant of it with shorter subsegments, and a third system based on a Bayesian hidden Markov model. Our analysis shows large accuracy disparities for the baseline system primarily due to conversational structure, which are partially mitigated in the other two systems. The illustration thus demonstrates how the methodology can be used to identify and guide diarization model improvements.
Scott Seyfarth, Sundararajan Srinivasan, Katrin Kirchhoff
Interspeech3
2021 Adapting Long Context NLM for ASR Rescoring in Conversational Agents
abstract
Neural Language Models (NLM), when trained and evaluated with context spanning multiple utterances, have been shown to consistently outperform both conventional n-gram language models and NLMs that use limited context. In this paper, we investigate various techniques to incorporate turn based context history into both recurrent (LSTM) and Transformer-XL based NLMs. For recurrent based NLMs, we explore context carry over mechanism and feature based augmentation, where we incorporate other forms of contextual information such as bot response and system dialogue acts as classified by a Natural Language Understanding (NLU) model. To mitigate the sharp nearby, fuzzy far away problem with contextual NLM, we propose the use of attention layer over lexical metadata to improve feature based augmentation. Additionally, we adapt our contextual NLM towards user provided on-the-fly speech patterns by leveraging encodings from a large pre-trained masked language model and performing fusion with a Transformer-XL based NLM. We test our proposed models using N-best rescoring of ASR hypotheses of task-oriented dialogues and also evaluate on downstream NLU tasks such as intent classification and slot labeling. The best performing model shows a relative WER between 1.6% and 9.1% and a slot labeling F1 score improvement of 4% over non-contextual baselines.
Ashish Shenoy, Sravan Babu Bodapati, Monica Sunkara, Srikanth Ronanki, Katrin Kirchhoff
Interspeech5
2021 Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment
abstract
Non-autoregressive encoder-decoder models greatly improve decoding speed over autoregressive models, at the expense of generation quality.To mitigate this, iterative decoding models repeatedly infill or refine the proposal of a non-autoregressive model.However, editing at the level of output sequences limits model flexibility.We instead propose iterative realignment, which by refining latent alignments allows more flexible edits in fewer steps.Our model, Align-Refine, is an end-to-end Transformer which iteratively realigns connectionist temporal classification (CTC) alignments.On the WSJ dataset, Align-Refine matches an autoregressive baseline with a 14× decoding speedup; on LibriSpeech, we reach an LM-free testother WER of 9.0% (19% relative improvement on comparable work) in three iterations.We release our code at https://github.com/ amazon-research/align-refine.
Ethan A. Chi, Julian Salazar, Katrin Kirchhoff
NAACL-HLT3
2020 Masked Language Model Scoring
abstract
Pretrained masked language models (MLMs) require finetuning for most NLP tasks.Instead, we evaluate MLMs out of the box via their pseudo-log-likelihood scores (PLLs), which are computed by masking tokens one by one.We show that PLLs outperform scores from autoregressive language models like GPT-2 in a variety of tasks.By rescoring ASR and NMT hypotheses, RoBERTa reduces an endto-end LibriSpeech model's WER by 30% relative and adds up to +1.7 BLEU on state-of-theart baselines for low-resource translation pairs, with further gains from domain adaptation.We attribute this success to PLL's unsupervised expression of linguistic acceptability without a left-to-right bias, greatly improving on scores from GPT-2 (+10 points on island effects, NPI licensing in BLiMP).One can finetune MLMs to give scores without masking, enabling computation in a single inference pass.In all, PLLs and their associated pseudo-perplexities (PP-PLs) enable plug-and-play use of the growing number of pretrained MLMs; e.g., we use a single cross-lingual model to rescore translations in multiple languages.We release our library for language model scoring at https: //github.com/awslabs/mlm-scoring.
Julian Salazar, Davis Liang, Toan Q. Nguyen, Katrin Kirchhoff
ACL4
2020 Deep Contextualized Acoustic Representations for Semi-Supervised Speech Recognition
abstract
We propose a novel approach to semi-supervised automatic speech recognition (ASR). We first exploit a large amount of unlabeled audio data via representation learning, where we reconstruct a temporal slice of filterbank features from past and future context frames. The resulting deep contextualized acoustic representations (DeCoAR) are then used to train a CTC-based end-to-end ASR system using a smaller amount of labeled audio data. In our experiments, we show that systems trained on DeCoAR consistently outperform ones trained on conventional filterbank features, giving 42% and 19% relative improvement over the baseline on WSJ eval92 and LibriSpeech test-clean, respectively. Our approach can drastically reduce the amount of labeled data required; unsupervised training on LibriSpeech then supervision with 100 hours of labeled data achieves performance on par with training on all 960 hours directly.
Shaoshi Ling, Yuzong Liu, Julian Salazar, Katrin Kirchhoff
ICASSP4
2020 Continual Learning for Multi-Dialect Acoustic Models
Brady Houston, Katrin Kirchhoff
INTERSPEECH2
2020 Multimodal Semi-Supervised Learning Framework for Punctuation Prediction in Conversational Speech
abstract
In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data. Conventional approaches in speech processing typically use forced alignment to encoder per frame acoustic features to word level features and perform multimodal fusion of the resulting acoustic and lexical representations. As an alternative, we explore attention based multimodal fusion and compare its performance with forced alignment based fusion. Experiments conducted on the Fisher corpus show that our proposed approach achieves ~6-9% and ~3-4% absolute improvement (F1 score) over the baseline BLSTM model on reference transcripts and ASR outputs respectively. We further improve the model robustness to ASR errors by performing data augmentation with N-best lists which achieves up to an additional ~2-6% improvement on ASR outputs. We also demonstrate the effectiveness of semi-supervised learning approach by performing ablation study on various sizes of the corpus. When trained on 1 hour of speech and text data, the proposed model achieved ~9-18% absolute improvement over baseline model.
Monica Sunkara, Srikanth Ronanki, Dhanush Bekal, Sravan Babu Bodapati, Katrin Kirchhoff
INTERSPEECH5
2019 Self-attention Networks for Connectionist Temporal Classification in Speech Recognition
abstract
The success of self-attention in NLP has led to recent applications in end-to-end encoder-decoder architectures for speech recognition. Separately, connectionist temporal classification (CTC) has matured as an alignment-free, non-autoregressive approach to sequence transduction, either by itself or in various multitask and decoding frameworks. We propose SAN-CTC, a deep, fully self-attentional network for CTC, and show it is tractable and competitive for end-to-end speech recognition. SAN-CTC trains quickly and outperforms existing CTC models and most encoder-decoder models, with character error rates (CERs) of 4.7% in 1 day on WSJ eval92 and 2.8% in 1 week on LibriSpeech test-clean, with a fixed architecture and one GPU. Similar improvements hold for WERs after LM decoding. We motivate the architecture for speech, evaluate position and down-sampling approaches, and explore how label alphabets (character, phoneme, subword) affect attention heads and performance.
Julian Salazar, Katrin Kirchhoff, Zhiheng Huang
ICASSP2
2019 Speech Audio Super-Resolution for Speech Recognition
Xinyu Li 0003, Venkata Chebiyyam, Katrin Kirchhoff
INTERSPEECH3
2019 Multi-Stream Network with Temporal Attention for Environmental Sound Classification
abstract
Environmental sound classification systems often do not perform robustly across different sound classification tasks and audio signals of varying temporal structures.We introduce a multi-stream convolutional neural network with temporal attention that addresses these problems.The network relies on three input streams consisting of raw audio and spectral features and utilizes a temporal attention function computed from energy changes over time.Training and classification utilizes decision fusion and data augmentation techniques that incorporate uncertainty.We evaluate this network on three commonly used data sets for environmental sound and audio scene classification and achieve new state-of-the-art performance without any changes in network architecture or frontend preprocessing, thus demonstrating better generalizability.
Xinyu Li 0003, Venkata Chebiyyam, Katrin Kirchhoff
INTERSPEECH3
2019 Simple, Fast, Accurate Intent Classification and Slot Labeling for Goal-Oriented Dialogue Systems
abstract
With the advent of conversational assistants like Amazon Alexa, Google Now, etc., dialogue systems are gaining a lot of traction, especially in industrial settings.These systems typically include a Spoken Language understanding component which consists of two tasks: Intent Classification (IC) and Slot Labeling (SL).Generally, these two tasks are modeled together jointly to achieve best performance.However, this joint modeling adds to model obfuscation.In this work, we first design framework for a modularization of joint IC-SL task to enhance architecture transparency.Then, we explore a number of selfattention, convolutional, and recurrent models, contributing a large-scale analysis of modeling paradigms for IC+SL across two datasets.Finally, using this framework, we propose a class of 'label-recurrent' models that are nonrecurrent apart from a 10-dimensional representation of the label history, and show that our proposed systems are highly accurate (achieving over 30% error reduction in SL over the state-of-the-art on the Snips dataset), as well as fast, at 2x the inference and 2/3 to 1/2 the training time of comparable recurrent models, thus giving an edge in critical real-world systems.
Arshit Gupta, John Hewitt, Katrin Kirchhoff
SIGdial3
2018 Development of machine translation technology for assisting health communication: A systematic review
Kristin Dew, Anne M. Turner, Yong K. Choi, Alyssa Bosold, Katrin Kirchhoff
J. Biomed. Informatics5
2017 SVitchboard-II and FiSVer-I: Crafting high quality and low complexity conversational english speech corpora using submodular function optimization
Yuzong Liu, Rishabh Iyer 0001, Katrin Kirchhoff, Jeff A. Bilmes
Comput. Speech Lang.3
2016 Crowdsourced Evaluation of Medical Texts Simplified by Medical Trainees
Yong K. Choi, Anne M. Turner, Katrin Kirchhoff
AMIA3
2016 Novel Front-End Features Based on Neural Graph Embeddings for DNN-HMM and LSTM-CTC Acoustic Modeling
Yuzong Liu, Katrin Kirchhoff
INTERSPEECH2
2016 Graph-Based Semisupervised Learning for Acoustic Modeling in Automatic Speech Recognition
abstract
In this paper, we investigate how to apply graph-based semisupervised learning to acoustic modeling in speech recognition. Graph-based semisupervised learning is a widely used transductive semisupervised learning method in which labeled and unlabeled data are jointly represented as a weighted graph; the resulting graph structure is then used as a constraint during the classification of unlabeled data points. We investigate suitable graph-based learning algorithms for speech data and evaluate two different frameworks for integrating graph-based learning into state-of-the-art, deep neural network (DDN)-based speech recognition systems. The first framework utilizes graph-based learning in parallel with a DNN classifier within a lattice-rescoring framework, whereas the second framework relies on an embedding of graph neighborhood information into continuous space using an autoencoder. We demonstrate significant improvements in framelevel phonetic classification accuracy and consistent reductions in word error rate on large-vocabulary conversational speech recognition tasks.
Yuzong Liu, Katrin Kirchhoff
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Acoustic modeling with neural graph embeddings
abstract
Graph-based learning (GBL) is a form of semi-supervised learning that has been successfully exploited in acoustic modeling in the past. It utilizes manifold information in speech data that is represented as a joint similarity graph over training and test samples. Typically, GBL is used at the output level of an acoustic classifier; however, this setup is difficult to scale to large data sets, and the graph-based learner is not optimized jointly with other components of the speech recognition system. In this paper we explore a different approach where the similarity graph is first embedded into continuous space using a neural autoencoder. Features derived from this encoding are then used at the input level to a standard DNN-based speech recognizer. We demonstrate improved scalability and performance compared to the standard GBL approach as well as significant improvements in word error rate on a medium-vocabulary Switchboard task.
Yuzong Liu, Katrin Kirchhoff
ASRU2
2015 SVitchboard II and fiSVer i: high-quality limited-complexity corpora of conversational English speech
abstract
In this paper, we introduce a set of benchmark corpora of conversational English speech derived from the Switchboard-I and Fisher datasets. Traditional ASR research requires considerable computational resources and has slow experimental turnaround times. Our goal is to introduce these new datasets to researchers in the ASR and machine learning communities (especially in academia), in order to facilitate the development of novel acoustic modeling techniques on smaller but acoustically rich corpora. We select these corpora to maximize an acoustic quality criterion while limiting the vocabulary size (from 10 words up to 10,000 words) with different state-of-the-art submodular function optimization algorithms. We provide baseline word recognition results for both GMM and DNN-based systems and release the corpora definitions and Kaldi training recipes to the public.
Yuzong Liu, Rishabh Iyer 0001, Katrin Kirchhoff, Jeff A. Bilmes
INTERSPEECH3
2015 Morphological Modeling for Machine Translation of English-Iraqi Arabic Spoken Dialogs
abstract
Katrin Kirchhoff, Yik-Cheung Tam, Colleen Richey, Wen Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Katrin Kirchhoff, Yik-Cheung Tam, Colleen Richey, Wen Wang 0001
HLT-NAACL1
2015 Modeling workflow to design machine translation applications for public health practice
Anne M. Turner, Megumu Brownstein, Kate Cole, Hilary Karasz, Katrin Kirchhoff
J. Biomed. Informatics5
2015 Syntactic and Semantic Features For Code-Switching Factored Language Models
abstract
This paper presents our latest investigations on different features for factored language models for Code-Switching speech and their effect on automatic speech recognition (ASR) performance. We focus on syntactic and semantic features which can be extracted from Code-Switching text data and integrate them into factored language models. Different possible factors, such as words, part-of-speech tags, Brown word clusters, open class words and clusters of open class word embeddings are explored. The experimental results reveal that Brown word clusters, part-of-speech tags and open-class words are the most effective at reducing the perplexity of factored language models on the Mandarin-English Code-Switching corpus SEAME. In ASR experiments, the model containing Brown word clusters and part-of-speech tags and the model also including clusters of open class word embeddings yield the best mixed error rate results. In summary, the best language model can significantly reduce the perplexity on the SEAME evaluation set by up to 10.8% relative and the mixed error rate by up to 3.4% relative.
Heike Adel, Ngoc Thang Vu, Katrin Kirchhoff, Dominic Telaar, Tanja Schultz
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Submodularity for Data Selection in Machine Translation
Katrin Kirchhoff, Jeff A. Bilmes
EMNLP1
2014 Unsupervised submodular subset selection for speech data
abstract
We conduct a comparative study on selecting subsets of acoustic data for training phone recognizers. The data selection problem is approached as a constrained submodular optimization problem. Previous applications of this approach required transcriptions or acoustic models trained in a supervised way. In this paper we develop and evaluate a novel and entirely unsupervised approach, and apply it to TIMIT data. Results show that our method consistently outperforms a number of baseline methods while being computationally very efficient and requiring no labeling.
Yuzong Liu, Katrin Kirchhoff, Jeff A. Bilmes
ICASSP3
2014 Submodular subset selection for large-scale speech training data
abstract
We address the problem of subselecting a large set of acoustic data to train automatic speech recognition (ASR) systems. To this end, we apply a novel data selection technique based on constrained submodular function maximization. Though NP-hard, the combinatorial optimization problem can be approximately solved by a simple and scalable greedy algorithm with constant-factor guarantees. We evaluate our approach by subselecting data from 1300 hours of conversational English telephone data to train two types large-vocabulary speech recognizers, one with Gaussian mixture model (GMM) based acoustic models, and another based on deep neural networks (DNNs). We show that training data can be reduced significantly, and that our technique outperforms both random selection and a previously proposed selection method utilizing comparable resources. Notably, using the submodular selection method, the DNN system using only about 5% of the training data is able to achieve performance on par with the GMM system using 100% of the training data - with the baseline subset selection methods, however, the DNN system is unable to accomplish this correspondence.
Yuzong Liu, Katrin Kirchhoff, Chris D. Bartels, Jeff A. Bilmes
ICASSP3
2014 Comparing approaches to convert recurrent neural networks into backoff language models for efficient decoding
abstract
In this paper, we investigate and compare three different possibilities to convert recurrent neural network language models (RNNLMs) into backoff language models (BNLM). While RNNLMs often outperform traditional n-gram approaches in the task of language modeling, their computational demands make them unsuitable for an efficient usage during decoding in an LVCSR system. It is, therefore, of interest to convert them into BNLMs in order to integrate their information into the decoding process. This paper compares three different approaches: a text based conversion, a probability based conversion and an iterative conversion. The resulting language models are evaluated in terms of perplexity and mixed error rate in the context of the Code-Switching data corpus SEAME. Although the best results are obtained by combining the results of all three approaches, the text based conversion approach alone leads to significant improvements on the SEAME corpus as well while offering the highest computational efficiency. In total, the perplexity can be reduced by 11.4% relative on the evaluation set and the mixed error rate by 3.0% relative on the same data set.
Heike Adel, Katrin Kirchhoff, Ngoc Thang Vu, Dominic Telaar, Tanja Schultz
INTERSPEECH2
2014 Combining recurrent neural networks and factored language models during decoding of code-Switching speech
abstract
In this paper, we present our latest investigations of language modeling for Code-Switching. Since there is only little text material for Code-Switching speech available, we integrate syntactic and semantic features into the language modeling process. In particular, we use part-of-speech tags, language identifiers, Brown word clusters and clusters of open class words. We develop factored language models and convert recurrent neural network language models into backoff language models for an efficient usage during decoding. A detailed error analysis reveals the strengths and weaknesses of the different language models. When we interpolate the models linearly, we reduce the perplexity by 15.6% relative on the SEAME evaluation set. This is even slightly better than the result of the unconverted recurrent neural network. We also combine the language models during decoding and obtain a mixed error rate reduction of 4.4% relative on the SEAME evaluation set.
Heike Adel, Dominic Telaar, Ngoc Thang Vu, Katrin Kirchhoff, Tanja Schultz
INTERSPEECH4
2014 Graph-based semi-supervised acoustic modeling in DNN-based speech recognition
abstract
This paper describes the combination of two recent machine learning techniques for acoustic modeling in speech recognition: deep neural networks (DNNs) and graph-based semi-supervised learning (SSL). While DNNs have been shown to be powerful supervised classifiers and have achieved considerable success in speech recognition, graph-based SSL can exploit valuable complementary information derived from the manifold structure of the unlabeled test data. Previous work on graph-based SSL in acoustic modeling has been limited to frame-level classification tasks and has not been compared to, or integrated with, state-of-the-art DNN/HMM recognizers. This paper represents the first integration of graph-based SSL with DNN based speech recognition and analyzes its effect on word recognition performance. The approach is evaluated on two small vocabulary speech recognition tasks and shows a significant improvement in HMM state classification accuracy as well as a consistent reduction in word error rate over a state-of-the-art DNN/HMM baseline.
Yuzong Liu, Katrin Kirchhoff
SLT2
2014 A conjoint analysis framework for evaluating user preferences in machine translation
Katrin Kirchhoff, Daniel Capurro, Anne M. Turner
Mach. Transl.1
2013 Submodular feature selection for high-dimensional acoustic score spaces
abstract
We apply methods for selecting subsets of dimensions from high-dimensional score spaces, and subsets of data for training, using submodular function optimization. Submodular functions provide theoretical performance guarantees while simultaneously retaining extremely fast and scalable optimization via an accelerated greedy algorithm. We evaluate this approach on two applications: data subset selection for phone recognizer training, and semi-supervised learning for phone segment classification. Interestingly, the first application uses submodularity twice: first for score space sub-selection and then for data subset selection. Our approach is computationally efficient but still consistently outperforms a number of baseline methods.
Yuzong Liu, Katrin Kirchhoff, Yisong Song, Jeff A. Bilmes
ICASSP3
2013 Classification of developmental disorders from speech signals using submodular feature selection
abstract
We present our system for the Interspeech 2013 Computational Paralinguistics Autism Sub-challenge. Our contribution focuses on improving classification accuracy of developmental disorders by applying a novel feature selection technique to the rich set of acoustic-prosodic features provided for this purpose. Our feature selection approach is based on submodular function optimization. We demonstrate significant improvements over systems using the full feature set and over a standard feature selection approach. Our final system outperforms the official Challenge baseline system significantly on the development set for both classification tasks, and on the test set for the Typicality task. Finally, we analyze the subselected features and identify the most important ones. Index Terms: classification, feature selection, neural networks, submodular functions
Katrin Kirchhoff, Yuzong Liu, Jeff A. Bilmes
INTERSPEECH1
2013 Graph-based semi-supervised learning for phone and segment classification
abstract
This paper presents several novel contributions to the emerging framework of graph-based semi-supervised learning for speech processing. First, we apply graph-based learning to variable-length segments rather than to the fixed-length vector representations that have been used previously. As part of this work we compare various graph-based learners, and we utilize an efficient feature selection technique for high-dimensional feature spaces that alleviates computational costs and improves the per-formance of graph-based learners. Finally, we present a method to improve regularization during the learning pro-cess. Experimental evaluation on the TIMIT frame and segment classification tasks demonstrates that the graph-based classifiers outperform standard baseline classifiers; furthermore, we find that the best learning algorithms are those that can incorporate prior knowledge. 1.
Yuzong Liu, Katrin Kirchhoff
INTERSPEECH2
2013 Using Document Summarization Techniques for Speech Data Subset Selection
Yuzong Liu, Katrin Kirchhoff, Jeff A. Bilmes
HLT-NAACL3
2012 Evaluating User Preferences in Machine Translation Using Conjoint Analysis
Katrin Kirchhoff, Daniel Capurro, Anne M. Turner
EAMT1
2012 Spectrum Identification using a Dynamic Bayesian Network Model of Tandem Mass Spectra
Ajit P. Singh, John T. Halloran, Jeff A. Bilmes, Katrin Kirchhoff, William Stafford Noble
UAI4
2011 Phonetic Classification Using Controlled Random Walks
abstract
Recently, semi-supervised learning algorithms for phonetic classifiers have been proposed that have obtained promising results. Often, these algorithms attempt to satisfy learning criteria that are not inherent in the standard generative or discriminative training procedures for phonetic classifiers. Graph-based learners in particular utilize an objective function that not only maximizes the classification accuracy on a labeled set but also the global smoothness of the predicted label assignment. In this paper we investigate a novel graph-based semi-supervised learning framework that implements a controlled random walk where different possible moves in the random walk are controlled by probabilities that are dependent on the properties of the graph itself. Experimental results on the TIMIT corpus are presented that demonstrate the effectiveness of this procedure.
Katrin Kirchhoff, Andrei Alexandrescu
INTERSPEECH1
2011 Semi-supervised ranking for document retrieval
Kevin Duh, Katrin Kirchhoff
Comput. Speech Lang.2
2011 Application of statistical machine translation to public health information: a feasibility study
abstract
OBJECTIVE: Accurate, understandable public health information is important for ensuring the health of the nation. The large portion of the US population with Limited English Proficiency is best served by translations of public-health information into other languages. However, a large number of health departments and primary care clinics face significant barriers to fulfilling federal mandates to provide multilingual materials to Limited English Proficiency individuals. This article presents a pilot study on the feasibility of using freely available statistical machine translation technology to translate health promotion materials. DESIGN: The authors gathered health-promotion materials in English from local and national public-health websites. Spanish versions were created by translating the documents using a freely available machine-translation website. Translations were rated for adequacy and fluency, analyzed for errors, manually corrected by a human posteditor, and compared with exclusively manual translations. RESULTS: Machine translation plus postediting took 15-53 min per document, compared to the reported days or even weeks for the standard translation process. A blind comparison of machine-assisted and human translations of six documents revealed overall equivalency between machine-translated and manually translated materials. The analysis of translation errors indicated that the most important errors were word-sense errors. CONCLUSION: The results indicate that machine translation plus postediting may be an effective method of producing multilingual health materials with equivalent quality but lower cost compared to manual translations.
Katrin Kirchhoff, Anne M. Turner, Amittai Axelrod, Francisco Saavedra
J. Am. Medical Informatics Assoc.1
2010 Contextual Modeling for Meeting Translation Using Unsupervised Word Sense Disambiguation
Katrin Kirchhoff
COLING2
2010 Hand Gestures in Disambiguating Types of You Expressions in Multiparty Meetings
Tyler Baldwin, Joyce Y. Chai, Katrin Kirchhoff
SIGDIAL Conference3
2009 Communicative gestures in coreference identification in multiparty meetings
abstract
During multiparty meetings, participants can use non-verbal modalities such as hand gestures to make reference to the shared environment. Therefore, one hypothesis is that incorporating hand gestures can improve coreference identification, a task that automatically identifies what participants refer to with their linguistic expressions. To evaluate this hypothesis, this paper examines the role of hand gestures in coreference identification, in particular, focusing on two questions: (1) what signals can distinguish communicative gestures that can potentially help coreference identification from non-communicative gestures; and (2) in what ways can communicative gestures help coreference identification. Based on the AMI data, our empirical results have shown that the length of gesture production is highly indicative of whether a gesture is communicative and potentially helpful in language understanding. Our experiments on the automated identification of coreferring expressions indicate that while the incorporation of simple gesture features does not improve overall performance, it does show potential on expressions referring to participants, an important and unique component of the meeting domain. A further analysis suggests that communicative gestures provide both redundant and complementary information, but further domain modeling and world knowledge incorporation is required to take full advantage of information that is complementary.
Tyler Baldwin, Joyce Y. Chai, Katrin Kirchhoff
ICMI3
2009 Graph-based Learning for Statistical Machine Translation
Andrei Alexandrescu, Katrin Kirchhoff
HLT-NAACL2
2009 Introduction to the Special Issue on Processing Morphologically Rich Languages
abstract
The 12 papers in this special issue span a variety of speech and language processing applications highlighting the challenges and providing solutions for dealing with the morphological complexity of different languages.
Ruhi Sarikaya, Katrin Kirchhoff, Tanja Schultz, Dilek Hakkani-Tür
IEEE Trans. Speech Audio Process.2
2008 Development of the SRI/nightingale Arabic ASR system
abstract
We describe the large vocabulary automatic speech recognition system developed for Modern Standard Arabic by the SRI/Nightingale team, and used for the 2007 GALE evaluation as part of the speech translation system. We show how system performance is affected by different development choices, ranging from text processing and lexicon to decoding system architecture design. Word error rate results are reported on broadcast news and conversational data from the GALE development and evaluation test sets. Index Terms: speech recognition, large vocabulary, Arabic 1.
Dimitra Vergyri, Arindam Mandal, Wen Wang 0001, Andreas Stolcke, Jing Zheng 0001, Martin Graciarena, David Rybach, Christian Gollan, Ralf Schlüter, Katrin Kirchhoff, Arlo Faria, Nelson Morgan
INTERSPEECH10
2008 Learning to rank with partially-labeled data
abstract
Ranking algorithms, whose goal is to appropriately order a set of objects/documents, are an important component of information retrieval systems. Previous work on ranking algorithms has focused on cases where only labeled data is available for training (i.e. supervised learning). In this paper, we consider the question whether unlabeled (test) data can be exploited to improve ranking performance. We present a framework for transductive learning of ranking functions and show that the answer is affirmative. Our framework is based on generating better features from the test data (via KernelPCA) and incorporating such features via Boosting, thus learning different ranking functions adapted to the individual test queries. We evaluate this method on the LETOR (TREC, OHSUMED) dataset and demonstrate significant improvements.
Kevin Duh, Katrin Kirchhoff
SIGIR2
2007 Graph-based learning for phonetic classification
abstract
We introduce graph-based learning for acoustic-phonetic classification. In graph-based learning, training, and test data points are jointly represented in a wieghted undirected graph characterized by a weight matrix indicating similarties between different samples. Classification of test samples is achieved by label propagation over the entire graph. Although this learning technique is commnly applied in semi- supervised settings, we show how it can also be as a post- processing step to a supervised classifier by imposing additional regularization constraints based on the underlying data manifold. We also present a technique to adapt graph-based learning to large datasets and evaluate our system on a vowel classification task. Our results show that graph-based learning improves significantly over state-of-the-art baselines.
Andrei Alexandrescu, Katrin Kirchhoff
ASRU2
2007 OOV detection by joint word/phone lattice alignment
abstract
We propose a new method for detecting out-of-vocabulary (OOV) words for large vocabulary continuous speech recognition (LVCSR) systems. Our method is based on performing a joint alignment between independently generated word and phone lattices, where the word-lattice is aligned via a recognition lexicon. Based on a similarity measure between phones, we can locate highly mis-aligned regions of time, and then specify those regions as candidate OOVs. This novel approach is implemented using the framework of graphical models (GMs), which enable fast flexible integration of different scores from word lattices, phone lattices, and the similarity measures. We evaluate our method on switchboard data using RT-04 as test set. Experimental results show that our approach provides a promising and scalable new way to detect OOV for LVCSR.
Hui Lin 0001, Jeff A. Bilmes, Dimitra Vergyri, Katrin Kirchhoff
ASRU4
2007 Attention shift decoding for conversational speech recognition
Raghunandan Kumaran, Jeff A. Bilmes, Katrin Kirchhoff
INTERSPEECH3
2007 Semi-automatic error analysis for large-scale statistical machine translation
Katrin Kirchhoff, Owen Rambow, Nizar Habash, Mona T. Diab
MTSummit1
2007 Data-Driven Graph Construction for Semi-Supervised Graph-Based Learning in NLP
Andrei Alexandrescu, Katrin Kirchhoff
HLT-NAACL2
2007 Bridging the gap between human and automatic speech recognition
Louis ten Bosch, Katrin Kirchhoff
Speech Commun.2
2006 Phrase-Based Backoff Models for Machine Translation of Highly Inflected Languages
Katrin Kirchhoff
EACL2
2006 Lexicon Acquisition for Dialectal Arabic Using Transductive Learning
Kevin Duh, Katrin Kirchhoff
EMNLP2
2006 The Vocal Joystick
abstract
The Vocal Joystick is a novel human-computer interface mechanism designed to enable individuals with motor impairments to make use of vocal parameters to control objects on a computer screen (buttons, sliders, etc.) and ultimately electro-mechanical instruments (e.g., robotic arms, wireless home automation devices). We have developed a working prototype of our "VJ-engine" with which individuals can now control computer mouse movement with their voice. The core engine is currently optimized according to a number of criterion. In this paper, we describe the engine system design, engine optimization, and user-interface improvements, and outline some of the signal processing and pattern recognition modules that were successful. Lastly, we present new results comparing the vocal joystick with a state-of-the-art eye tracking pointing device, and show that not only is the Vocal Joystick already competitive, for some tasks it appears to be an improvement.
Jeff A. Bilmes, Jonathan Malkin, Xiao Li 0006, Susumu Harada, Kelley Kilanski, Katrin Kirchhoff, Richard Wright, Amarnag Subramanya, James A. Landay, Patricia Dowden, Howard Jay Chizeck
ICASSP (1)6
2006 Factored Neural Language Models
Andrei Alexandrescu, Katrin Kirchhoff
HLT-NAACL2
2006 Graphical Model Representations of Word Lattices
abstract
We introduce a method for expressing word lattices within a dynamic graphical model. We describe a variety of choices for doing this, including a technique to relax the time information associated with lattice nodes in a way that trades off hypothesis expansion with presumed segmentation boundary accuracy. Our approach uses a set of time-inhomogeneous and algorithmically expressed conditional probability tables to encode the lattice. The approach was implemented as part of the graphical model toolkit, and word error rate improvements on the Switchboard corpus indicate that our technique is a viable means to incorporate large state space speech recognition systems into a graphical model.
Gang Ji, Jeff A. Bilmes, Jeff Michels, Katrin Kirchhoff, Christopher D. Manning
SLT4
2006 Morphology-based language modeling for conversational Arabic speech recognition
Katrin Kirchhoff, Dimitra Vergyri, Jeff A. Bilmes, Kevin Duh, Andreas Stolcke
Comput. Speech Lang.1
2006 Recent innovations in speech-to-text transcription at SRI-ICSI-UW
abstract
We summarize recent progress in automatic speech-to-text transcription at SRI, ICSI, and the University of Washington. The work encompasses all components of speech modeling found in a state-of-the-art recognition system, from acoustic features, to acoustic modeling and adaptation, to language modeling. In the front end, we experimented with nonstandard features, including various measures of voicing, discriminative phone posterior features estimated by multilayer perceptrons, and a novel phone-level macro-averaging for cepstral normalization. Acoustic modeling was improved with combinations of front ends operating at multiple frame rates, as well as by modifications to the standard methods for discriminative Gaussian estimation. We show that acoustic adaptation can be improved by predicting the optimal regression class complexity for a given speaker. Language modeling innovations include the use of a syntax-motivated almost-parsing language model, as well as principled vocabulary-selection techniques. Finally, we address portability issues, such as the use of imperfect training transcripts, and language-specific adjustments required for recognition of Arabic and Mandarin
Andreas Stolcke, Barry Y. Chen, Horacio Franco, Venkata Ramana Rao Gadde, Martin Graciarena, Mei-Yuh Hwang, Katrin Kirchhoff, Arindam Mandal, Nelson Morgan, Tim Ng, Mari Ostendorf, M. Kemal Sönmez, Anand Venkataraman, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001, Qifeng Zhu 0001
IEEE Trans. Speech Audio Process.7
2005 Landmark-Based Speech Recognition: Report of the 2004 Johns Hopkins Summer Workshop
abstract
Three research prototype speech recognition systems are described, all of which use recently developed methods from artificial intelligence (specifically support vector machines, dynamic Bayesian networks, and maximum entropy classification) in order to implement, in the form of an automatic speech recognizer, current theories of human speech perception and phonology (specifically landmark-based speech perception, nonlinear phonology, and articulatory phonology). All three systems begin with a high-dimensional multiframe acoustic-to-distinctive feature transformation, implemented using support vector machines trained to detect and classify acoustic phonetic landmarks. Distinctive feature probabilities estimated by the support vector machines are then integrated using one of three pronunciation models: a dynamic programming algorithm that assumes canonical pronunciation of each word, a dynamic Bayesian network implementation of articulatory phonology, or a discriminative pronunciation model trained using the methods of maximum entropy classification. Log probability scores computed by these models are then combined, using log-linear combination, with other word scores available in the lattice output of a first-pass recognizer, and the resulting combination score is used to compute a second-pass speech recognition output.
Mark Hasegawa-Johnson, James Baker, Sarah Borys, Ken Chen 0001, Emily Coogan, Steven Greenberg, Amit Juneja, Katrin Kirchhoff, Karen Livescu, Srividya Mohan, Jennifer Muller, M. Kemal Sönmez
ICASSP (1)8
2005 Genetic triangulation of graphical models for speech and language processing
abstract
Graphical models are an increasingly popular approach for speech and language processing. As researchers design ever more complex models it becomes crucial to find triangulations that make inference problems tractable. This paper presents a genetic algorithm for triangulation search that is well-suited for speech and language graphical models. It is unique in two ways: First, it can find triangulations appropriate for graphs with a mix of stochastic and deterministic dependencies. Second, the search is guided by optimizing the inference speed (CPU runtime) on real data. We show results on 10 real-world speech and language graphs and demonstrate inference speed-ups over standard triangulation methods. 1.
Chris D. Bartels, Kevin Duh, Jeff A. Bilmes, Katrin Kirchhoff, Simon King 0001
INTERSPEECH4
2005 Development of a conversational telephone speech recognizer for Levantine Arabic
abstract
Many languages, including Arabic, are characterized by a wide variety of different dialects that often differ strongly from each other. When developing speech technology for dialect-rich languages, the portability and reusability of data, algorithms, and system components becomes extremely important. In this paper, we describe the development of a large-vocabulary speech recognition system for Levantine Arabic, which was a new dialectal recognition task for our existing system. We discuss the dialect-specific modeling choices (grapheme vs. phoneme based acoustic models, automatic vowelization techniques, and morphological language models) and investigate to what extent techniques previously tested on other languages are portable to the present task. We present stateof-the-art recognition results on the 2004 Levantine Arabic Rich Transcription evaluation.
Dimitra Vergyri, Katrin Kirchhoff, Venkata Ramana Rao Gadde, Andreas Stolcke, Jing Zheng 0001
INTERSPEECH2
2005 Error-correction detection and response generation in a spoken dialogue system
Ivan Bulyko, Katrin Kirchhoff, Mari Ostendorf, J. Goldberg
Speech Commun.2
2005 Cross-dialectal data sharing for acoustic modeling in Arabic speech recognition
Katrin Kirchhoff, Dimitra Vergyri
Speech Commun.1
2004 Automatic Learning of Language Model Structure
Kevin Duh, Katrin Kirchhoff
COLING2
2004 Cross-dialectal acoustic data sharing for Arabic speech recognition
abstract
The automatic recognition of Arabic dialectal speech is a challenging task since Arabic dialects are essentially spoken varieties, for which only sparse resources (transcriptions and standardized acoustic data) are available to date. In this paper we describe the use of acoustic data from modern standard Arabic (MSA) to improve the recognition of Egyptian conversational Arabic (ECA). The cross-dialectal use of data is complicated by the fact that MSA is written without short vowels and other diacritics and thus has incomplete phonetic information. This problem is addressed by automatically vowelizing MSA data before combining it with ECA data. We described the vowelization procedure as well as speech recognition experiments and show that our technique yields improvements over our baseline system.
Katrin Kirchhoff, Dimitra Vergyri
ICASSP (1)1
2004 Morphology-based language modeling for arabic speech recognition
abstract
Language modeling is a difficult problem for languages with rich morphology. In this paper we investigate the use of morphology-based language models at different stages in a speech recognition system for conversational Arabic. Class-based and single-stream factored language models using mor-phological word representations are applied within an N-best list rescoring framework. In addition, we explore the use of factored language models in first-pass recognition, which is fa-cilitated by two novel procedures: the data-driven optimization of a multi-stream language model structure, and the conversion of a factored language model to a standard word-based model. We evaluate these techniques on a large-vocabulary recognition task and demonstrate that they lead to perplexity and word error rate reductions. 1.
Dimitra Vergyri, Katrin Kirchhoff, Kevin Duh, Andreas Stolcke
INTERSPEECH2
2003 Novel approaches to Arabic speech recognition: report from the 2002 Johns-Hopkins Summer Workshop
abstract
Although Arabic is currently one of the most widely spoken languages in the world, there has been relatively little speech recognition research on Arabic compared to other languages. Moreover, most previous work has concentrated on the recognition of formal rather than dialectal Arabic. This paper reports on our project at the 2002 Johns Hopkins Summer Workshop, which focused on the recognition of dialectal Arabic. Three problems were addressed: (a) the lack of short vowels and other pronunciation information in Arabic texts; (b) the morphological complexity of Arabic; and (c) the discrepancies between dialectal and formal Arabic. We present novel approaches to automatic vowel restoration, morphology-based language modeling and the integration of out-of-corpus language model data, and report significant word error rate improvements on the LDC Arabic CallHome task.
Katrin Kirchhoff, Jeff A. Bilmes, Sourin Das, Nicolae Duta, Melissa Egan, Gang Ji, John Henderson, Daben Liu, Mohammed Noamany, Patrick Schone, Richard M. Schwartz, Dimitra Vergyri
ICASSP (1)1
2003 Multi-stream language identification using data-driven dependency selection
abstract
The most widespread approach to automatic language identification in the past has been the statistical modeling of phone sequences extracted from speech signals. Recently, we have developed an alternative approach to LID based on n-gram modeling of parallel streams of articulatory features, which was shown to have advantages over phone-based systems on short test signals whereas the latter achieved a higher accuracy on longer signals. Additionally, phone and feature streams can be combined to achieve maximum performance. Within this "multi-stream" framework two types of statistical dependencies need to be modeled: (a) dependencies between symbols in individual streams and (b) dependencies between symbols in different streams. The space of possible dependencies is typically too large to be searched exhaustively. We explore the use of genetic algorithms as a method for data-driven dependency selection. The result is a general framework for the discovery and modeling of dependencies between multiple information sources expressed as sequences of symbols, which has implications for other fields beyond language identification, such as speaker identification or language modeling.
Sonia Parandekar, Katrin Kirchhoff
ICASSP (1)2
2003 Factored Language Models and Generalized Parallel Backoff
Jeff A. Bilmes, Katrin Kirchhoff
HLT-NAACL2
2003 Generalized rules for combination and joint training of classifiers
Jeff A. Bilmes, Katrin Kirchhoff
Pattern Anal. Appl.2
2002 Mixed-memory Markov models for Automatic Language Identification
abstract
Automatic language identification (LID) continues to play an integral part in many multilingual speech applications. The most widespread approach to LID is the phonotactic approach, which performs language classification based on the probabilities of phone sequences extracted from the test signal. These probabilities are typically computed using statistical phone n-gram models. In this paper we investigate the approximation of these standard n-gram models by mixed-memory Markov models with application to both a phone-based and an articulatory feature-based LID system. We demonstrate significant improvements in accuracy with a substantially reduced set of parameters on a 10-way language identification task.
Katrin Kirchhoff, Sonia Parandekar, Jeff A. Bilmes
ICASSP1
2002 The 2001 GMTK-based SPINE ASR system
abstract
This paper provides a detailed description of the University of Washington automatic speech recognition (ASR) system for the 2001 DARPA SPeech In Noisy Environments (SPINE) task. Our system makes heavy use of the graphical modeling toolkit (GMTK), a general purpose graphical modeling-based ASR system that allows arbitrary parameter tying, flexible deterministic and stochastic dependencies between variables, and a generalized maximum likelihood parameter estimation algorithm. In our SPINE system, GMTK was used for acoustic model training whereas feature extraction, speaker adaptation, and first-pass decoding were performed by HTK. Our integrated GMTK/HTK system demonstrates the relative merits provided by each tool. Novel aspects of our SPINE system include the capturing of correlations among feature vectors via a globally-shared factored sparse inverse covariance matrix and generalized EM training. 1.
Özgür Çetin, Harriet J. Nock, Katrin Kirchhoff, Jeff A. Bilmes, Mari Ostendorf
INTERSPEECH3
2002 Low-resource noise-robust feature post-processing on Aurora 2.0
Chia-Ping Chen, Jeff A. Bilmes, Katrin Kirchhoff
INTERSPEECH3
2002 Combining acoustic and articulatory feature information for robust speech recognition
Katrin Kirchhoff, Gernot A. Fink, Gerhard Sagerer
Speech Commun.1
2001 Multi-stream statistical n-gram modeling with application to automatic language identification
abstract
Most state-of-the art automatic language identification systems are based on phonotactic information, i.e. languages are identified on the basis of probabilities of phone sequences extracted from the acoustic signal. This approach ignores the potential advantages to be gained from a richer representation of the acoustic signal in terms of parallel streams of subphonemic events. In this paper we develop an alternative approach to language identification which is based on parallel streams of phonetic features and sparse modeling of statistical dependencies between these streams. We present results on the OGI-TS database and show that the feature-based system outperforms a comparable phone-based system significantly while using fewer parameters. Moreover, the feature-based system exhibits a markedly better performance on very short test signals ( 3 seconds). The theoretical approach developed here is of significance not only for language identification but also for related work in pronunciation modeling. 1.
Katrin Kirchhoff, Sonia Parandekar
INTERSPEECH1
2000 Conversational speech recognition using acoustic and articulatory input
abstract
The combination of multiple speech recognizers based on different signal representations is increasingly attracting interest in the speech community. In previous work we presented a hybrid speech recognition system based on the combination of acoustic and articulatory information which achieved significant word error rate reductions under highly noisy conditions on a small-vocabulary numbers recognition task. In this study we extend this approach to large-vocabulary conversational speech recognition using the Gaussian mixture acoustic modeling paradigm. We demonstrate that the articulatory input representation we propose contains information which is complementary to that provided by standard MFCC features, and that their combination can significantly reduce the word error rate on conversational speech. Various combination strategies (feature-level, state-level and word-level combination) are compared and evaluated.
Katrin Kirchhoff, Gernot A. Fink, Gerhard Sagerer
ICASSP1
2000 Directed graphical models of classifier combination: application to phone recognition
abstract
Classifier combination is a technique that often provides appreciable accuracy gains. In this paper, we argue that the underlying statistical model of classifier combination should be made explicit. Using directed graphical models (DGMs), we provide representations of two common combination schemes, the mean and product rules. We also introduce new DGMs that yield novel combination rules. We find that these new DGM-inspired rules can achieve significant accuracy gains on the TIMIT phone-classification task relative to existing combination schemes. 1. INTRODUCTION When multiple independently trained pattern classifiers are combined, the resulting accuracy is often better than any of the individual classifiers. This has been demonstrated for automatic speech recognition (ASR) [7, 10, 18] and for pattern classification [12, 13, 20, 29]. Classifier combination can fuse together different information sources to utilize their complementary information. The sources can be multi-modal, such...
Jeff A. Bilmes, Katrin Kirchhoff
INTERSPEECH2
2000 Speech analysis by rule extraction from trained artificial neural networks
Katrin Kirchhoff
INTERSPEECH1
1999 Dynamic classifier combination in hybrid speech recognition systems using utterance-level confidence values
abstract
A recent development in the hybrid HMM/ANN speech recognition paradigm is the use of several subword classifiers, each of which provides different information about the speech signal. Although the combining methods have obtained promising results, the strategies so far proposed have been relatively simple. In most cases frame-level subword unit probabilities are combined using an unweighted product or sum rule. In this paper, we argue and empirically demonstrate that the classifier combination approach can benefit from a dynamically weighted combination rule, where the weights are derived from higher-than-frame-level confidence values.
Katrin Kirchhoff, Jeff A. Bilmes
ICASSP1
1998 Combining articulatory and acoustic information for speech recognition in noisy and reverberant environments
abstract
Robust speech recognition under varying acoustic conditions may be achieved by exploiting multiple sources of information in the speech signal. In addition to an acoustic signal representation, we use an articulatory representation consisting of pseudoarticulatory features as an additional information source. Hybrid ANN/HMM recognizers using either of these representations are evaluated on a continuous numbers recognition task (OGI Numbers95) under clean, reverberant and noisy conditions. An error analysis of preliminary recognition results shows that the different representations produce qualitatively different errors, which suggests a combination of both representations. We investigate various combination possibilities at the phoneme estimation level and show that significant improvements can been achieved under all three acoustic conditions. 1. INTRODUCTION Whereas most speech recognition systems use a cepstral or spectral representation of the speech signal, there have also been ...
Katrin Kirchhoff
ICSLP1
1996 Syllable-level desynchronisation of phonetic features for speech recognition
Katrin Kirchhoff
ICSLP1