Tsendsuren Munkhdalai

dblp:17/11240 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0002-8783-4993ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
abstract
Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang, Shyam Upadhyay. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang 0008, Shyam Upadhyay
ACL (1)6
2024 Improving Speech Recognition for African American English with Audio Classification
abstract
Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to mitigate this is to train or fine-tune models with more representative datasets. But this approach can be hindered by limited in-domain data for training and evaluation. We propose a new way to improve the robustness of a US English short-form speech recognizer using a small amount of out-of-domain (long-form) African American English (AAE) data. We use CORAAL, YouTube and Mozilla Common Voice to train an audio classifier to approximately output whether an utterance is AAE or some other variety including Mainstream American English (MAE). By combining the classifier output with coarse geographic information, we can select a subset of utterances from a large corpus of untranscribed short-form queries for semi-supervised learning at scale. Fine-tuning on this data results in a 38.5% relative word error rate disparity reduction between AAE and MAE without reducing MAE quality.
Shefali Garg, Zhouyuan Huo, Khe Chai Sim, Suzan Schwartz, Mason Chua, Alëna Aksënova, Tsendsuren Munkhdalai, Levi King, Darryl Wright, Zion Mengesha, Dongseong Hwang, Tara N. Sainath, Françoise Beaufays, Pedro J. Moreno 0001
ICASSP7
2024 A Comparison of Parameter-Efficient ASR Domain Adaptation Methods for Universal Speech and Language Models
abstract
A recent paradigm shift in artificial intelligence has seen the rise of foundation models, such as the large language models and the universal speech models. With billions of model parameters and trained with a wide range of data, these foundation models are expected to have a better generalization to different downstream tasks. Efficient adaptation is the key to leveraging these foundation models in a new task or domain. In this paper, we compare several popular parameter-efficient tuning methods, such as vector adaptation, residual adapters, low-rank adapter (LoRA) and prompt-tuning, for automatic speech recognition (ASR) domain adaptation. We use the connectionist temporal classification (CTC) model with Conformer encoder and fused it with a universal language model. We study the effect of adapting either or both of the Conformer encoder and the universal language model. We carry out extensive experiments to study these methods under different hyper-parameter settings and the effect of combining some of these methods. We find that combining vector adaptation and residual adapters with increasing bottleneck dimension achieved the best performance.
Khe Chai Sim, Zhouyuan Huo, Tsendsuren Munkhdalai, Nikhil Siddhartha, Adam Stooke, Zhong Meng, Bo Li 0028, Tara N. Sainath
ICASSP3
2024 Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro J. Moreno 0001
INTERSPEECH4
2023 Re-investigating the Efficient Transfer Learning of Speech Foundation Model using Feature Fusion Methods
Zhouyuan Huo, Khe Chai Sim, Dongseong Hwang, Tsendsuren Munkhdalai, Tara N. Sainath, Pedro J. Moreno 0001
INTERSPEECH4
2023 Dual-Mode NAM: Effective Top-K Context Injection for End-to-End ASR
Zelin Wu, Tsendsuren Munkhdalai, Pat Rondon, Golan Pundak, Khe Chai Sim, Christopher Li
INTERSPEECH2
2022 Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech Recognition
abstract
Fast contextual adaptation has shown to be effective in improving Automatic Speech Recognition (ASR) of rare words and when combined with an on-device personalized training, it can yield an even better recognition result. However, the traditional re-scoring approaches based on an external language model is prone to diverge during the personalized training. In this work, we introduce a model-based end-to-end contextual adaptation approach that is decoder-agnostic and amenable to on-device personalization. Our on-device simulation experiments demonstrate that the proposed approach outperforms the traditional re-scoring technique by 12% relative WER and 15.7% entity mention specific F1-score in a continuous personalization scenario.
Tsendsuren Munkhdalai, Khe Chai Sim, Angad Chandorkar, Mason Chua, Trevor Strohman, Françoise Beaufays
ICASSP1
2022 On-the-fly ASR Corrections with Audio Exemplars
Golan Pundak, Tsendsuren Munkhdalai, Khe Chai Sim
INTERSPEECH2
2022 NAM+: Towards Scalable End-to-End Contextual Biasing for Adaptive ASR
abstract
Attention-based biasing techniques for end-to-end ASR systems are able to achieve large accuracy gains without requiring the inference algorithm adjustments and parameter tuning common to fusion approaches. However, it is challenging to simultaneously scale up attention-based biasing to realistic numbers of biased phrases; maintain in-domain WER gains, while minimizing out-of-domain losses; and run in real time. We present NAM+, an attention-based biasing approach which achieves a 16X inference speedup per acoustic frame over prior work when run with 3,000 biasing entities, as measured on a typical mobile CPU. NAM+ achieves these run-time gains through a combination of Two-Pass Hierarchical Attention and Dilated Context Update. Compared to the adapted baseline, NAM+ further decreases the in-domain WER by up to 12.6% relative, while incurring an out-of-domain WER regression of 20% relative. Compared to the non-adapted baseline, the out-of-domain WER regression is 7.1 % relative.
Tsendsuren Munkhdalai, Zelin Wu, Golan Pundak, Khe Chai Sim, Pat Rondon, Tara N. Sainath
SLT1
2022 Context-Aware Neural Confidence Estimation for Rare Word Speech Recognition
abstract
Confidence estimation for automatic speech recognition (ASR) is important for many downstream tasks. Recently, neural confidence estimation models (CEMs) have been shown to produce accurate confidence scores for predicting word-level errors. These models are built on top of an end-to-end (E2E) ASR and the acoustic embeddings are part of the input features. However, practical E2E ASR systems often incorporate contextual information in the decoder to improve rare word recognition. The CEM is not aware of this and underestimates the confidence of the rare words that have been corrected by the context. In this paper, we propose a context-aware CEM by incorporating context into the encoder using a neural associative memory (NAM) model. It uses attention to detect for presence of the biasing phrases and modify the encoder features. Experiments show that the proposed context-aware CEM using NAM augmented training can improve the AUC-ROC for word error prediction from 0.837 to 0.892.
David Qiu, Tsendsuren Munkhdalai, Yanzhang He, Khe Chai Sim
SLT2
2022 Internal Language Model Personalization of E2E Automatic Speech Recognition Using Random Encoder Features
abstract
End-to-end (E2E) speech-to-text models generally require transcribed audio for training and personalization. We introduce the use of random audio encoder features, rather than speech, to fine-tune the final model layers and acquire new vocabulary from text-only data. This technique can be used for on-device personalization before the user has provided any speech data. We show improvements in the recall of new vocabulary and word error rate (WER) on held-out test sets using simulated user experiments on hybrid autoregressive transducer (HAT) models using conformer-based encoders and simple text embeddings for label processing. We compare this approach to the use of synthetic audio, finding random encoder features to be more beneficial with lower computational cost. Experiments show that the maximum benefit is gained by updating specific network components comprising a subset of those expressing the internal language model.
Adam Stooke, Khe Chai Sim, Mason Chua, Tsendsuren Munkhdalai, Trevor Strohman
SLT4
2021 Diverse Distributions of Self-Supervised Tasks for Meta-Learning in NLP
abstract
Meta-learning considers the problem of learning an efficient learning process that can leverage its past experience to accurately solve new tasks.However, the efficacy of meta-learning crucially depends on the distribution of tasks available for training, and this is often assumed to be known a priori or constructed from limited supervised datasets.In this work, we aim to provide task distributions for meta-learning by considering self-supervised tasks automatically proposed from unlabeled text, to enable large-scale meta-learning in NLP.We design multiple distributions of self-supervised tasks by considering important aspects of task diversity, difficulty, type, domain, and curriculum, and investigate how they affect meta-learning performance.Our analysis shows that all these factors meaningfully alter the task distribution, some inducing significant improvements in downstream few-shot accuracy of the metalearned models.Empirically, results on 20 downstream tasks show significant improvements in few-shot learning -adding up to +4.2% absolute accuracy (on average) to the previous unsupervised meta-learning method, and perform comparably to supervised methods on the FewRel 2.0 benchmark.
Trapit Bansal, Karthick Gunasekaran, Tong Wang 0012, Tsendsuren Munkhdalai, Andrew McCallum
EMNLP (1)4
2021 Learning Associative Inference Using Fast Weight Memory
Imanol Schlag, Tsendsuren Munkhdalai, Jürgen Schmidhuber
ICLR2
2021 Robust Continuous On-Device Personalization for Automatic Speech Recognition
Khe Chai Sim, Angad Chandorkar, Mason Chua, Tsendsuren Munkhdalai, Françoise Beaufays
Interspeech5
2020 Self-Supervised Meta-Learning for Few-Shot Natural Language Classification Tasks
abstract
Self-supervised pre-training of transformer models has revolutionized NLP applications.Such pre-training with language modeling objectives provides a useful initial point for parameters that generalize well to new tasks with fine-tuning.However, fine-tuning is still data inefficient -when there are few labeled examples, accuracy can be low.Data efficiency can be improved by optimizing pre-training directly for future fine-tuning with few examples; this can be treated as a meta-learning problem.However, standard meta-learning techniques require many training tasks in order to generalize; unfortunately, finding a diverse set of such supervised tasks is usually difficult.This paper proposes a self-supervised approach to generate a large, rich, metalearning task distribution from unlabeled text.This is achieved using a cloze-style objective, but creating separate multi-class classification tasks by gathering tokens-to-be blanked from among only a handful of vocabulary terms.This yields as many unique meta-training tasks as the number of subsets of vocabulary terms.We meta-train a transformer model on this distribution of tasks using a recent meta-learning framework.On 17 NLP tasks, we show that this meta-training leads to better few-shot generalization than language-model pre-training followed by finetuning.Furthermore, we show how the self-supervised tasks can be combined with supervised tasks for meta-learning, providing substantial accuracy gains over previous supervised meta-learning.
Trapit Bansal, Rishikesh Jha, Tsendsuren Munkhdalai, Andrew McCallum
EMNLP (1)3
2020 Exploring and Predicting Transferability across NLP Tasks
abstract
Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, Mohit Iyyer. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Tu Vu, Tong Wang 0012, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, Mohit Iyyer
EMNLP (1)3
2020 GEV-NN: A deep neural network architecture for class imbalance problem in binary classification
Lkhagvadorj Munkhdalai, Tsendsuren Munkhdalai, Keun Ho Ryu
Knowl. Based Syst.2
2019 Building Dynamic Knowledge Graphs from Text using Machine Reading Comprehension
Rajarshi Das, Tsendsuren Munkhdalai, Xingdi Yuan, Adam Trischler, Andrew McCallum
ICLR (Poster)2
2019 Metalearned Neural Memory
abstract
We augment recurrent neural networks with an external memory mechanism that builds upon recent progress in metalearning. We conceptualize this memory as a rapidly adaptable function that we parameterize as a deep neural network. Reading from the neural memory function amounts to pushing an input (the key vector) through the function to produce an output (the value vector). Writing to memory means changing the function; specifically, updating the parameters of the neural network to encode desired information. We leverage training and algorithmic techniques from metalearning to update the neural memory function in one shot. The proposed memory-augmented model achieves strong performance on a variety of learning problems, from supervised question answering to reinforcement learning.
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang 0012, Adam Trischler
NeurIPS1
2018 Understanding Deep Learning Performance through an Examination of Test Set Difficulty: A Psychometric Case Study
abstract
Interpreting the performance of deep learning models beyond test set accuracy is challenging. Characteristics of individual data points are often not considered during evaluation, and each data point is treated equally. We examine the impact of a test set question's difficulty to determine if there is a relationship between difficulty and performance. We model difficulty using well-studied psychometric methods on human response patterns. Experiments on Natural Language Inference (NLI) and Sentiment Analysis (SA) show that the likelihood of answering a question correctly is impacted by the question's difficulty. As DNNs are trained with more data, easy examples are learned more quickly than hard examples.
John Lalor, Hao Wu 0055, Tsendsuren Munkhdalai, Hong Yu 0001
EMNLP3
2018 Rapid Adaptation with Conditionally Shifted Neurons
abstract
We describe a mechanism by which artificial neural networks can learn rapid adaptation - the ability to adapt on the fly, with little data, to new tasks - that we call conditionally shifted neurons. We apply this mechanism in the framework of metalearning, where the aim is to replicate some of the flexibility of human learning in machines. Conditionally shifted neurons modify their activation values with task-specific shifts retrieved from a memory module, which is populated rapidly based on limited task experience. On metalearning benchmarks from the vision and language domains, models augmented with conditionally shifted neurons achieve state-of-the-art results.
Tsendsuren Munkhdalai, Xingdi Yuan, Soroush Mehri, Adam Trischler
ICML1
2017 Neural Tree Indexers for Text Understanding
abstract
Recurrent neural networks (RNNs) process input text sequentially and model the conditional transition between word tokens.In contrast, the advantages of recursive networks include that they explicitly model the compositionality and the recursive structure of natural language.However, the current recursive architecture is limited by its dependence on syntactic tree.In this paper, we introduce a robust syntactic parsing-independent tree structured model, Neural Tree Indexers (NTI) that provides a middle ground between the sequential RNNs and the syntactic treebased recursive models.NTI constructs a full n-ary tree by processing the input text with its node function in a bottom-up fashion.Attention mechanism can then be applied to both structure and node function.We implemented and evaluated a binarytree model of NTI, showing the model achieved the state-of-the-art performance on three different NLP tasks: natural language inference, answer sentence selection, and sentence classification, outperforming state-of-the-art recurrent and recursive neural networks 1 .
Tsendsuren Munkhdalai, Hong Yu 0001
EACL (1)1
2017 Neural Semantic Encoders
abstract
We present a memory augmented neural network for natural language understanding: Neural Semantic Encoders.NSE is equipped with a novel memory update rule and has a variable sized encoding memory that evolves over time and maintains the understanding of input sequences through read, compose and write operations.NSE can also access 1 multiple and shared memories.In this paper, we demonstrated the effectiveness and the flexibility of NSE on five different natural language tasks: natural language inference, question answering, sentence classification, document sentiment analysis and machine translation where NSE achieved state-of-the-art performance when evaluated on publically available benchmarks.For example, our shared-memory model showed an encouraging result on neural machine translation, improving an attention-based baseline by approximately 1.0 BLEU.
Tsendsuren Munkhdalai, Hong Yu 0001
EACL (1)1
2017 Reasoning with Memory Augmented Neural Networks for Language Comprehension
Tsendsuren Munkhdalai, Hong Yu 0001
ICLR (Poster)1
2017 Meta Networks
abstract
Neural networks have been successfully applied in applications with a large amount of labeled data. However, the task of rapid generalization on new concepts with small training data while preserving performances on previously learned ones still presents a significant challenge to neural network models. In this work, we introduce a novel meta learning method, Meta Networks (MetaNet), that learns a meta-level knowledge across tasks and shifts its inductive biases via fast parameterization for rapid generalization. When evaluated on Omniglot and Mini-ImageNet benchmarks, our MetaNet models achieve a near human-level performance and outperform the baseline approaches by up to 6\% accuracy. We demonstrate several appealing properties of MetaNet relating to generalization and continual learning.
Tsendsuren Munkhdalai, Hong Yu 0001
ICML1
2015 Self-training in significance space of support vectors for imbalanced biomedical event data
abstract
BACKGROUND: Pairwise relationships extracted from biomedical literature are insufficient in formulating biomolecular interactions. Extraction of complex relations (namely, biomedical events) has become the main focus of the text-mining community. However, there are two critical issues that are seldom dealt with by existing systems. First, an annotated corpus for training a prediction model is highly imbalanced. Second, supervised models trained on only a single annotated corpus can limit system performance. Fortunately, there is a large pool of unlabeled data containing much of the domain background that one can exploit. RESULTS: In this study, we develop a new semi-supervised learning method to address the issues outlined above. The proposed algorithm efficiently exploits the unlabeled data to leverage system performance. We furthermore extend our algorithm to a two-phase learning framework. The first phase balances the training data for initial model induction. The second phase incorporates domain knowledge into the event extraction model. The effectiveness of our method is evaluated on the Genia event extraction corpus and a PubMed document pool. Our method can identify a small subset of the majority class, which is sufficient for building a well-generalized prediction model. It outperforms the traditional self-training algorithm in terms of f-measure. Our model, based on the training data and the unlabeled data pool, achieves comparable performance to the state-of-the-art systems that are trained on a larger annotated set consisting of training and evaluation data.
Tsendsuren Munkhdalai, Oyun-Erdene Namsrai, Keun Ho Ryu
BMC Bioinform.1