VLDB 2026 Research / reviewers in the wild / expert
Shankar Kumar
dblp:37/3396
· DBLP profile ↗
43ranked-venue papers
8as first author
12since 2021 · last 2025
0009-0001-4307-8102ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Role of Outgoing Connection Heterogeneity in Feedforward Layers of Large Language ModelsabstractWe report on investigations into the characteristics of outgoing connections in feedforward layers of large language models.Our findings show that inner neurons with diverse outgoing connection strengths are more critical to model performance than those with uniform connections.We propose a new fine-tuning loss that takes advantage of this observation by decreasing the outgoing connection entropy in feedforward layers.Using this loss yields gains over standard fine-tuning across two different model families (PaLM-2 and Gemma-2) for downstream tasks in math, coding, and language understanding.To further elucidate the role of outgoing connection heterogeneity, we develop a data-free structured pruning method, which uses entropy to identify and remove neurons.This method is considerably more effective than removing neurons either randomly or based on their magnitude. Felix Stahlberg, Shankar Kumar |
EMNLP | 2 |
| 2025 | Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post EditingabstractLarge Language Models (LLMs) excel at rewriting tasks such as text style transfer and grammatical error correction. While there is considerable overlap between the inputs and outputs in these tasks, the decoding cost still increases with output length, regardless of the amount of overlap. By leveraging the overlap between the input and the output, Kaneko and Okazaki [1] proposed model-agnostic edit span representations to compress the rewrites to save computation. They reported an output length reduction rate of nearly 80% with minimal accuracy impact in four rewriting tasks. In this paper, we propose alternative edit phrase representations inspired by phrase-based statistical machine translation. We systematically compare our phrasal representations with their span representations. We apply the LLM rewriting model to the task of Automatic Speech Recognition (ASR) post editing and show that our target-phrase-only edit representation has the best efficiency-accuracy trade-off. On the LibriSpeech test set, our method closes 50-60% of the WER gap between the edit span model and the full rewrite model while losing only 10-20% of the length reduction rate of the edit span model. Felix Stahlberg, Shankar Kumar |
ICASSP | 3 |
| 2025 | Massive Sound Embedding Benchmark (MSEB)abstractAudio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb. Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001 |
NeurIPS | 6 |
| 2023 | Multi-Output RNN-T Joint Networks for Multi-Task Learning of ASR and Auxiliary TasksabstractWe propose a multi-output joint network architecture for RNN-T transducer, for multi-task modeling of ASR and auxiliary tasks that rely on ASR outputs. Each output of the joint network predicts tar-get labels with disjoint vocabularies for each task, while sharing the same audio features by the encoder and language model features by the prediction network. Each task is trained with an RNN-T loss that marginalizes over all possible paths, and we allow multiple tasks to share the blank logit so that they are synchronized. We demonstrate our method on two auxiliary tasks, namely capitalization and pause prediction, and discuss different considerations for modeling and inference procedures. For capitalization, we successfully distill capitalization labels from a standalone text normalization model, and achieve competitive Uppercase Error Rate (UER) while offering streaming capability and improved inference efficiency. In addition, our model has similar capitalization accuracy compared to a mixed-case ASR model, but obtains improved WERs if integrated with external language models. For pause prediction, we achieve the same performance as the previous two-step approach while providing a simpler training recipe without affecting ASR accuracy. Ding Zhao, Shaojin Ding, Hao Zhang 0010, Shuo-Yiin Chang, David Rybach, Tara N. Sainath, Yanzhang He, Ian McGraw, Shankar Kumar |
ICASSP | 10 |
| 2023 | Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR
W. Ronny Huang, Hao Zhang 0010, Shankar Kumar, Shuo-Yiin Chang, Tara N. Sainath |
INTERSPEECH | 3 |
| 2023 | Fast Text Generation with Text-Editing ModelsabstractText-editing models have recently become a prominent alternative to seq2seq models for monolingual text-generation tasks such as grammatical error correction, simplification, and style transfer. These tasks share a common trait -- they exhibit a large amount of textual overlap between the source and target texts. Text-editing models take advantage of this observation and learn to generate the output by predicting edit operations applied to the source sequence. In contrast, seq2seq models generate outputs word-by-word from scratch thus making them slow at inference time. Text-editing models provide several benefits over seq2seq models including faster inference speed, higher sample efficiency, and better control and explainability of the outputs. This tutorial provides a comprehensive overview of text-editing models and discusses how they can be used to mitigate hallucination and bias, both pressing challenges in the field of text generation. Finally, we discuss how to optimize latency of large language models via distillation to text-editing models and other means. Eric Malmi, Yue Dong 0002, Jonathan Mallinson, Aleksandr Chuklin, Jakub Adámek, Daniil Mirylenka, Felix Stahlberg, Sebastian Krause, Shankar Kumar, Aliaksei Severyn |
KDD | 9 |
| 2023 | Measuring Re-identification RiskabstractCompact user representations (such as embeddings) form the backbone of personalization services. In this work, we present a new theoretical framework to measure re-identification risk in such user representations. Our framework, based on hypothesis testing, formally bounds the probability that an attacker may be able to obtain the identity of a user from their representation. As an application, we show how our framework is general enough to model important real-world applications such as the Chrome's Topics API for interest-based advertising. We complement our theoretical bounds by showing provably good attack algorithms for re-identification that we use to estimate the re-identification risk in the Topics API. We believe this work provides a rigorous and interpretable notion of re-identification risk and a framework to measure it that can be used to inform real-world applications. CJ Carey, Travis Dick, Alessandro Epasto, Adel Javanmard, Josh Karlin, Shankar Kumar, Andrés Muñoz Medina, Vahab S. Mirrokni, Gabriel Henrique Nunes, Sergei Vassilvitskii, Peilin Zhong |
Proc. ACM Manag. Data | 6 |
| 2022 | Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence ModelsabstractIn many natural language processing (NLP) tasks the same input (e.g.source sentence) can have multiple possible outputs (e.g.translations).To analyze how this ambiguity (also known as intrinsic uncertainty) shapes the distribution learned by neural sequence models we measure sentence-level uncertainty by computing the degree of overlap between references in multi-reference test sets from two different NLP tasks: machine translation (MT) and grammatical error correction (GEC).At both the sentence-and the task-level, intrinsic uncertainty has major implications for various aspects of search such as the inductive biases in beam search and the complexity of exact search.In particular, we show that well-known pathologies such as a high number of beam search errors, the inadequacy of the mode, and the drop in system performance with large beam sizes apply to tasks with high level of ambiguity such as MT but not to less uncertain tasks such as GEC.Furthermore, we propose a novel exact n-best search algorithm for neural sequence models, and show that intrinsic uncertainty affects model uncertainty as the model tends to overly spread out the probability mass for uncertain tasks and sentences. Felix Stahlberg, Ilia Kulikov, Shankar Kumar |
ACL (1) | 3 |
| 2022 | Capitalization Normalization for Language Modeling with an Accurate and Efficient Hierarchical RNN ModelabstractCapitalization normalization (truecasing) is the task of restoring the correct case (uppercase or lowercase) of noisy text. We propose a fast, accurate and compact two-level hierarchical word-and-character-based recurrent neural network model. We use the truecaser to normalize user-generated text in a Federated Learning framework for language modeling. A case-aware language model trained on this normalized text achieves the same perplexity as a model trained on text with gold capitalization. In a real user A/B experiment, we demonstrate that the improvement translates to reduced prediction error rates in a virtual keyboard application. Similarly, in an ASR language model fusion experiment, we show reduction in uppercase character error rate and word error rate. Hao Zhang 0010, You-Chi Cheng, Shankar Kumar, W. Ronny Huang, Mingqing Chen, Rajiv Mathews |
ICASSP | 3 |
| 2022 | Sentence-Select: Large-Scale Language Model Data Selection for Rare-Word Speech Recognition
W. Ronny Huang, Cal Peyser, Tara N. Sainath, Ruoming Pang, Trevor Strohman, Shankar Kumar |
INTERSPEECH | 6 |
| 2022 | Jam or Cream First? Modeling Ambiguity in Neural Machine Translation with SCONESabstractThe softmax layer in neural machine translation is designed to model the distribution over mutually exclusive tokens.Machine translation, however, is intrinsically uncertain: the same source sentence can have multiple semantically equivalent translations.Therefore, we propose to replace the softmax activation with a multi-label classification layer that can model ambiguity more effectively.We call our loss function Single-label Contrastive Objective for Non-Exclusive Sequences (SCONES).We show that the multi-label output layer can still be trained on single reference training data using the SCONES loss function.SCONES yields consistent BLEU score gains across six translation directions, particularly for mediumresource language pairs and small beam sizes.By using smaller beam sizes we can speed up inference by a factor of 3.9x and still match or improve the BLEU score obtained using softmax.Furthermore, we demonstrate that SCONES can be used to train NMT models that assign the highest probability to adequate translations, thus mitigating the "beam search curse".Additional experiments on synthetic language pairs with varying levels of uncertainty suggest that the improvements from SCONES can be attributed to better handling of ambiguity. Felix Stahlberg, Shankar Kumar |
NAACL-HLT | 2 |
| 2021 | Lookup-Table Recurrent Language Models for Long Tail Speech RecognitionabstractWe introduce Lookup-Table Language Models (LookupLM), a method for scaling up the size of RNN language models with only a constant increase in the floating point operations, by increasing the expressivity of the embedding table. In particular, we instantiate an (additional) embedding table which embeds the previous n-gram token sequence, rather than a single token. This allows the embedding table to be scaled up arbitrarily -- with a commensurate increase in performance -- without changing the token vocabulary. Since embeddings are sparsely retrieved from the table via a lookup; increasing the size of the table adds neither extra operations to each forward pass nor extra parameters that need to be stored on limited GPU/TPU memory. We explore scaling n-gram embedding tables up to nearly a billion parameters. When trained on a 3-billion sentence corpus, we find that LookupLM improves long tail log perplexity by 2.44 and long tail WER by 23.4% on a downstream speech recognition task over a standard RNN language model baseline, an improvement comparable to a scaling up the baseline by 6.2x the number of floating point operations. W. Ronny Huang, Tara N. Sainath, Cal Peyser, Shankar Kumar, David Rybach, Trevor Strohman |
Interspeech | 4 |
| 2020 | Seq2Edits: Sequence Transduction Using Span-level Edit OperationsabstractWe propose Seq2Edits, an open-vocabulary approach to sequence editing for natural language processing (NLP) tasks with a high degree of overlap between input and output texts.In this approach, each sequence-to-sequence transduction is represented as a sequence of edit operations, where each operation either replaces an entire source span with target tokens or keeps it unchanged.We evaluate our method on five NLP tasks (text normalization, sentence fusion, sentence splitting & rephrasing, text simplification, and grammatical error correction) and report competitive results across the board.For grammatical error correction, our method speeds up inference by up to 5.2x compared to full sequence models because inference time depends on the number of edits rather than the number of target tokens.For text normalization, sentence fusion, and grammatical error correction, our approach improves explainability by associating each edit operation with a human-readable tag. Felix Stahlberg, Shankar Kumar |
EMNLP (1) | 2 |
| 2020 | Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T LossabstractIn this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system. Transformer computation blocks based on self-attention are used to encode both audio and label sequences independently. The activations from both audio and label encoders are combined with a feed-forward layer to compute a probability distribution over the label space for every combination of acoustic frame position and label history. This is similar to the Recurrent Neural Network Transducer (RNN-T) model, which uses RNNs for information encoding instead of Transformer encoders. The model is trained with the RNN-T loss well-suited to streaming decoding. We present results on the LibriSpeech dataset showing that limiting the left context for self-attention in the Transformer layers makes decoding computationally tractable for streaming, with only a slight degradation in accuracy. We also show that the full attention version of our model beats the-state-of-the art accuracy on the LibriSpeech benchmarks. Our results also show that we can bridge the gap between full attention and limited attention versions of our model by attending to a limited number of future frames. Han Lu 0003, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, Shankar Kumar |
ICASSP | 7 |
| 2020 | Improving Tail Performance of a Deliberation E2E ASR Model Using a Large Text CorpusabstractEnd-to-end (E2E) automatic speech recognition (ASR) systems lack the distinct language model (LM) component that characterizes traditional speech systems. While this simplifies the model architecture, it complicates the task of incorporating text-only data into training, which is important to the recognition of tail words that do not occur often in audio-text pairs. While shallow fusion has been proposed as a method for incorporating a pre-trained LM into an E2E model at inference time, it has not yet been explored for very large text corpora, and it has been shown to be very sensitive to hyperparameter settings in the beam search. In this work, we apply shallow fusion to incorporate a very large text corpus into a state-of-the-art E2EASR model. We explore the impact of model size and show that intelligent pruning of the training set can be more effective than increasing the parameter count. Additionally, we show that incorporating the LM in minimum word error rate (MWER) fine tuning makes shallow fusion far less dependent on optimal hyperparameter settings, reducing the difficulty of that tuning problem. Cal Peyser, Sepand Mavandadi, Tara N. Sainath, James Apfel, Ruoming Pang, Shankar Kumar |
INTERSPEECH | 6 |
| 2020 | Data Weighted Training Strategies for Grammatical Error CorrectionabstractRecent progress in the task of Grammatical Error Correction (GEC) has been driven by addressing data sparsity, both through new methods for generating large and noisy pretraining data and through the publication of small and higher-quality finetuning data in the BEA-2019 shared task. Building upon recent work in Neural Machine Translation (NMT), we make use of both kinds of data by deriving example-level scores on our large pretraining data based on a smaller, higher-quality dataset. In this work, we perform an empirical study to discover how to best incorporate delta-log-perplexity, a type of example scoring, into a training schedule for GEC. In doing so, we perform experiments that shed light on the function and applicability of delta-log-perplexity. Models trained on scored data achieve state- of-the-art results on common GEC test sets. Jared Lichtarge, Christopher Alberti, Shankar Kumar |
Trans. Assoc. Comput. Linguistics | 3 |
| 2018 | A Conversational Neural Language Model for Speech Recognition in Digital AssistantsabstractSpeech recognition in digital assistants such as Google Assistant can potentially benefit from the use of conversational context consisting of user queries and responses from the agent. We explore the use of recurrent, Long Short-Term Memory (LSTM), neural language models (LMs) to model the conversations in a digital assistant. Our proposed methods effectively capture the context of previous utterances in a conversation without modifying the underlying LSTM architecture. We demonstrate a 4% relative improvement in recognition performance on Google Assistant queries when using the LSTM LMs to rescore recognition lattices. Eunjoon Cho, Shankar Kumar |
ICASSP | 2 |
| 2018 | RADMM: Recurrent Adaptive Mixture Model with Applications to Domain Robust Language ModelingabstractWe present a new architecture and a training strategy for an adaptive mixture of experts with applications to domain robust language modeling. The proposed model is designed to benefit from the scenario where the training data are available in diverse domains as is the case for YouTube speech recognition. The two core components of our model are an ensemble of parallel long short-term memory (LSTM) expert layers for each domain and another LSTM based network which generates state dependent mixture weights for combining expert LSTM states by linear interpolation. The resulting model is a recurrent adaptive mixture model (RADMM) of domain experts. We train our model on 4.4B words from YouTube speech recognition data. We report results on the YouTube speech recognition test set. Compared with a background LSTM model, we obtain up to 12% relative improvement in perplexity and an improvement in word error rate from 12.3% to 12.1 % while using a lattice rescoring with strong pruning. Kazuki Irie, Shankar Kumar, Michael Nirschl, Hank Liao |
ICASSP | 2 |
| 2018 | Modeling Non-Linguistic Contextual Signals in LSTM Language Models Via Domain AdaptationabstractLanguage Models (LMs) for Automatic Speech Recognition (ASR) can benefit from utilizing non-linguistic contextual signals in modeling. Examples of these signals include the geographical location of the user speaking to the system and/or the identity of the application (app) being spoken to. In practice, the vast majority of input speech queries typically lack annotations of such signals, which poses a challenge to directly train domain-specific LMs. To obtain robust domain LMs, generally an LM which has been pre-trained on general data will be adapted to specific domains. We propose four domain adaptation schemes to improve the domain performance of Long Short-Term Memory (LSTM) LMs, by incorporating app based contextual signals of voice search queries. We show that most of our adaptation strategies are effective, reducing word perplexity up to 21 % relative to a fine-tuned baseline on a held-out domain-specific development set. Initial experiments using a state-of-the-art Italian ASR system show a 3 % relative reduction in WER on top of an unadapted 5-gram LM. In addition, human evaluations show significant improvements on sub-domains from using app signals. Shankar Kumar, Fadi Biadsy, Michael Nirschl, Tomas Vykruta, Pedro J. Moreno 0001 |
ICASSP | 2 |
| 2018 | No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End ModelsabstractFor decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems. This status quo has begun to be challenged recently by end-to-end models which seek to combine acoustic, pronunciation, and language model components into a single neural network. Such systems, which typically predict graphemes or words, simplify the recognition process since they remove the need for a separate expert-curated pronunciation lexicon to map from phoneme-based units to words. However, there has been little previous work comparing phoneme-based versus grapheme-based sub-word units in the end-to-end modeling framework, to determine whether the gains from such approaches are primarily due to the new probabilistic model, or from the joint learning of the various components with grapheme-based units. In this work, we conduct detailed experiments which are aimed at quantifying the value of phoneme-based pronunciation lexica in the context of end-to-end models. We examine phoneme-based end-to-end models, which are contrasted against grapheme-based ones on a large vocabulary English Voice-search task, where we find that graphemes do indeed outperform phonemes. We also compare grapheme and phoneme-based approaches on a multi-dialect English task, which once again confirm the superiority of graphemes, greatly simplifying the system for recognizing multiple dialects. Tara N. Sainath, Rohit Prabhavalkar, Shankar Kumar, Seungji Lee, Anjuli Kannan, David Rybach, Vlad Schogol, Patrick Nguyen, Bo Li 0028, Chung-Cheng Chiu |
ICASSP | 3 |
| 2017 | Lattice rescoring strategies for long short term memory language models in speech recognitionabstractRecurrent neural network (RNN) language models (LMs) and Long Short Term Memory (LSTM) LMs, a variant of RNN LMs, have been shown to outperform traditional N-gram LMs on speech recognition tasks. However, these models are computationally more expensive than N-gram LMs for decoding, and thus, challenging to integrate into speech recognizers. Recent research has proposed the use of lattice-rescoring algorithms using RNNLMs and LSTMLMs as an efficient strategy to integrate these models into a speech recognition system. In this paper, we evaluate existing lattice rescoring algorithms along with new variants on a YouTube speech recognition task. Lattice rescoring using LSTMLMs reduces the word error rate (WER) for this task by 8% relative to the WER obtained using an N-gram LM. Shankar Kumar, Michael Nirschl, Daniel Niels Holtmann-Rice, Hank Liao, Ananda Theertha Suresh, Felix X. Yu |
ASRU | 1 |
| 2017 | Approaches for Neural-Network Language Model Adaptation
Michael Nirschl, Fadi Biadsy, Shankar Kumar |
INTERSPEECH | 4 |
| 2016 | NN-Grams: Unifying Neural Network and n-Gram Language Models for Speech RecognitionabstractWe present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural networks. We report experiments where the model is trained on 26B words. NN-grams are efficient at run-time since they do not include an output soft-max layer. The model is trained using noise contrastive estimation (NCE), an approach that transforms the estimation problem of neural networks into one of binary classification between data samples and noise samples. We present results with noise samples derived from either an n-gram distribution or from speech recognition lattices. NN-grams outperforms an n-gram model on an Italian speech recognition dictation task. Babak Damavandi, Shankar Kumar, Noam Shazeer, Antoine Bruguier |
INTERSPEECH | 2 |
| 2015 | Multilingual Open Relation Extraction Using Cross-lingual ProjectionabstractOpen domain relation extraction systems identify relation and argument phrases in a sentence without relying on any underlying schema. However, current state-of-the-art relation extraction systems are available only for English because of their heavy reliance on linguistic tools such as part-of-speech taggers and dependency parsers. We present a cross-lingual annotation projection method for language independent relation extraction. We evaluate our method on a manually annotated test set and present results on three typologically different languages. We release these manual annotations and extracted relations in ten languages from Wikipedia. Manaal Faruqui, Shankar Kumar |
HLT-NAACL | 2 |
| 2010 | Expected Sequence Similarity Maximization
Cyril Allauzen, Shankar Kumar, Wolfgang Macherey, Mehryar Mohri, Michael Riley 0001 |
HLT-NAACL | 2 |
| 2010 | Model Combination for Machine Translation
John DeNero, Shankar Kumar, Ciprian Chelba, Franz Josef Och |
HLT-NAACL | 2 |
| 2009 | Efficient Minimum Error Rate Training and Minimum Bayes-Risk Decoding for Translation Hypergraphs and Lattices
Shankar Kumar, Wolfgang Macherey, Chris Dyer, Franz Josef Och |
ACL/IJCNLP | 1 |
| 2008 | Lattice Minimum Bayes-Risk Decoding for Statistical Machine Translation
Roy Tromble, Shankar Kumar, Franz Josef Och, Wolfgang Macherey |
EMNLP | 2 |
| 2008 | Video suggestion and discovery for youtube: taking random walks through the view graphabstractThe rapid growth of the number of videos in YouTube provides enormous potential for users to find content of interest to them. Unfortunately, given the difficulty of searching videos, the size of the video repository also makes the discovery of new content a daunting task. In this paper, we present a novel method based upon the analysis of the entire user-video graph to provide personalized video suggestions for users. The resulting algorithm, termed Adsorption, provides a simple method to efficiently propagate preference information through a variety of graphs. We extensively test the results of the recommendations on a three month snapshot of live data from YouTube. Shumeet Baluja, Rohan Seth, Yushi Jing, Jay Yagnik, Shankar Kumar, Deepak Ravichandran, Mohamed Aly 0002 |
WWW | 6 |
| 2007 | Improving Word Alignment with Bridge Languages
Shankar Kumar, Franz Josef Och, Wolfgang Macherey |
EMNLP-CoNLL | 1 |
| 2007 | Segmentation and alignment of parallel text for statistical machine translationabstractWe address the problem of extracting bilingual chunk pairs from parallel text to create training sets for statistical machine translation. We formulate the problem in terms of a stochastic generative process over text translation pairs, and derive two different alignment procedures based on the underlying alignment model. The first procedure is a now-standard dynamic programming alignment model which we use to generate an initial coarse alignment of the parallel text. The second procedure is a divisive clustering parallel text alignment procedure which we use to refine the first-pass alignments. This latter procedure is novel in that it permits the segmentation of the parallel text into sub-sentence units which are allowed to be reordered to improve the chunk alignment. The quality of chunk pairs are measured by the performance of machine translation systems trained from them. We show practical benefits of divisive clustering as well as how system performance can be improved by exploiting portions of the parallel text that otherwise would have to be discarded. We also show that chunk alignment as a first step in word alignment can significantly reduce word alignment error rate. Yonggang Deng, Shankar Kumar, William J. Byrne |
Nat. Lang. Eng. | 2 |
| 2006 | A weighted finite state transducer translation template model for statistical machine translationabstractWe present a Weighted Finite State Transducer Translation Template Model for statistical machine translation. This is a source-channel model of translation inspired by the Alignment Template translation model. The model attempts to overcome the deficiencies of word-to-word translation models by considering phrases rather than words as units of translation. The approach we describe allows us to implement each constituent distribution of the model as a weighted finite state transducer or acceptor. We show that bitext word alignment and translation under the model can be performed with standard finite state machine operations involving these transducers. One of the benefits of using this framework is that it avoids the need to develop specialized search procedures, even for the generation of lattices or N-Best lists of bitext word alignments and translation hypotheses. We report and analyze bitext word alignment and translation performance on the Hansards French-English task and the FBIS Chinese-English task under the Alignment Error Rate, BLEU, NIST and Word Error-Rate metrics. These experiments identify the contribution of each of the model components to different aspects of alignment and translation performance. We finally discuss translation performance with large bitext training sets on the NIST 2004 Chinese-English and Arabic-English MT tasks. Shankar Kumar, Yonggang Deng, William J. Byrne |
Nat. Lang. Eng. | 1 |
| 2006 | Corrections to "Segmental minimum Bayes-risk decoding for automatic speech recognition"abstractThe purpose of this paper is to correct and expand upon the experimental results presented in our recently published paper [1]. In [1, Sec. III-B], we present a risk-based lattice cutting (RLC) procedure to segment ASR word lattices into sequences of smaller sublattices. The purpose of this procedure is to restructure the original lattice to improve the efficiency of minimum Bayes-risk (MBR) and other lattice rescoring procedures. Given that the segmented lattices are to be rescored, it is crucial that no paths from the original lattice be lost in the segmentation process. In the experiments reported in our original publication, some of the original paths were inadvertently discarded from the segmented lattices. This affected the performance of the MBR results presented. In this paper, we briefly review the segmentation algorithm and explain the flaw in our previous experiments. We find consistent minor improvements in word error rate (WER) under the corrected procedure. More importantly, we report experiments confirming that the lattice segmentation procedure does indeed preserve all the paths in the original lattice. Vaibhava Goel, Shankar Kumar, William J. Byrne |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | Minimum Bayes-Risk Decoding for Statistical Machine Translation
Shankar Kumar, William J. Byrne |
HLT-NAACL | 1 |
| 2004 | A Smorgasbord of Features for Statistical Machine Translation
Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alexander Fraser 0001, Shankar Kumar, Libin Shen, Katherine Eng, Viren Jain, Zhen Jin 0007, Dragomir R. Radev |
HLT-NAACL | 7 |
| 2004 | Segmental minimum Bayes-risk decoding for automatic speech recognitionabstractMinimum Bayes-risk (MBR) speech recognizers have been shown to yield improvements over the conventional maximum a-posteriori probability (MAP) decoders through N-best list rescoring and A/sup */ search over word lattices. We present a segmental minimum Bayes-risk decoding (SMBR) framework that simplifies the implementation of MBR recognizers through the segmentation of the N-best lists or lattices over which the recognition is to be performed. This paper presents lattice cutting procedures that underly SMBR decoding. Two of these procedures are based on a risk minimization criterion while a third one is guided by word-level confidence scores. In conjunction with SMBR decoding, these lattice segmentation procedures give consistent improvements in recognition word error rate (WER) on the Switchboard corpus. We also discuss an application of risk-based lattice cutting to multiple-system SMBR decoding and show that it is related to other system combination techniques such as ROVER. This strategy combines lattices produced from multiple ASR systems and is found to give WER improvements in a Switchboard evaluation system. Vaibhava Goel, Shankar Kumar, William J. Byrne |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | A Weighted Finite State Transducer Implementation of the Alignment Template Model for Statistical Machine Translation
Shankar Kumar, William J. Byrne |
HLT-NAACL | 1 |
| 2002 | Minimum Bayes-Risk Word Alignments of Bilingual TextsabstractWe present Minimum Bayes-Risk word alignment for machine translation. This statistical, model-based approach attempts to minimize the expected risk of alignment errors under loss functions that measure alignment quality. We describe various loss functions, including some that incorporate linguistic analysis as can be obtained from parse trees, and show that these approaches can improve alignments of the English-French Hansards. Shankar Kumar, William J. Byrne |
EMNLP | 1 |
| 2002 | Risk based lattice cutting for segmental minimum Bayes-risk decodingabstractMinimum Bayes-Risk (MBR) speech recognizers have been shown to give improvements over the conventional maximum a-posteriori probability (MAP) decoders through N-best list rescoring and search over word lattices. Segmental MBR (SMBR) decoders simplify the implementation of MBR recognizers by segmenting the N-best lists or lattices over which the recognition is performed. We present a lattice cutting procedure that attempts to minimize the total Bayes-Risk of all word strings in the segmented lattice. We provide experimental results on the Switchboard conversational speech corpus showing that this segmentation procedure, in conjunction with SMBR decoding, gives modest but significant improvements over MAP decoders as well as MBR decoders on unsegmented lattices. Shankar Kumar, William J. Byrne |
INTERSPEECH | 1 |
| 2001 | Confidence based lattice segmentation and minimum Bayes-risk decodingabstractMinimum Bayes Risk (MBR) speech recognizers have been shown to yield improvements over the conventional maximum a-posteriori probability (MAP) decoders in the context of N-best list rescoring and search over recognition lattices. Segmental MBR (SMBR) procedures have been developed to simplify implementation of MBR recognizers, by segmenting the N-best list or lattice, to reduce the size of the search space over which MBR recognition is carried out. In this paper we describe lattice cutting as a method to segment recognition word lattices into regions of low confidence and high confidence. We present two SMBR decoding procedures that can be applied on low confidence segment sets. Results obtained on the Switchboard conversational telephone speech corpus show modest but significant improvements relative to MAP decoders. Vaibhava Goel, Shankar Kumar, William J. Byrne |
INTERSPEECH | 2 |
| 2001 | Normalization of non-standard words
Richard Sproat, Alan W. Black, Stanley F. Chen, Shankar Kumar, Mari Ostendorf, Christopher Richards |
Comput. Speech Lang. | 4 |
| 2000 | Segmental minimum Bayes-risk ASR voting strategiesabstractROVER [1] and its successor voting procedures have been shown to be quite effective in reducing the recognition word error rate (WER). The success of these methods has been attributed to their minimum Bayes-risk (MBR) nature: they produce the hypothesis with the least expected word error. In this paper we develop a general procedure within the MBR framework, called segmental MBR recognition, that encompasses current voting techniques and allows further extensions that yield lower expected WER. It also allows incorporation of loss functions other than the WER. We present a derivation of voting procedure of N-best ROVER as an instance of segmental MBR recognition. We then present an extension, called e-ROVER, that alleviates some of the restrictions of N-best ROVER by better approximating the WER. e-ROVER is compared with N-best ROVER on multi-lingual acoustic modeling task and is shown to yield modest yet significant and easily obtained improvements. Vaibhava Goel, Shankar Kumar, William J. Byrne |
INTERSPEECH | 2 |
| 2000 | Unifying HMM and phone-pair segment models
Hsiao-Wuen Hon, Shankar Kumar, Kuansan Wang |
INTERSPEECH | 2 |