Mingbin Xu

dblp:163/1772 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 48% Language models and text generation · 28% Representation and self-supervised learning · 24%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
neural language model
0.312018
Dual Fixed-Size Ordinally Forgetting Encoding (FOFE) for Competitive Neural Language Models · EMNLP 2018
Natural language and speech › Information extraction and text analysis › named entity recognition
mention detection
0.312017
A Local Detection Approach for Named Entity Recognition and Mention Detection · ACL (1) 2017
Natural language and speech › Information extraction and text analysis
named entity recognition
0.312017
A Local Detection Approach for Named Entity Recognition and Mention Detection · ACL (1) 2017
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.312017
Word Embeddings based on Fixed-Size Ordinally Forgetting Encoding · EMNLP 2017

Methods — techniques the papers use, named apart from their topics

fixed-size ordinally forgetting encoding · 0.9dual-FOFE · 0.3truncated SVD · 0.3feedforward neural network · 0.3
YearPublicationVenuePosition
2025 Contextualization of ASR with LLM using phonetic retrieval-based augmentation
abstract
Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a phone book, when the input modality is speech. In this work, we start with a speech recognition task and propose a retrievalbased solution to contextualize the LLM: we first let the LLM detect named entities in speech without any context, then use this named entity as a query to retrieve phonetically similar named entities from a personal database and feed them to the LLM, and finally run context-aware LLM decoding. In a voice assistant task, our solution achieved up to 30.2% relative word error rate reduction and 73.6% relative named entity error rate reduction compared to a baseline system without contextualization. Notably, our solution by design avoids prompting the LLM with the full named entity database, making it highly efficient and applicable to large named entity databases.
Zhihong Lei, Xingyu Na, Mingbin Xu, Ernest Pusateri, Christophe Van Gysel, Shiyi Han, Zhen Huang 0001
ICASSP3
2024 Personalization of CTC-Based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization
abstract
Recent advances in deep learning and automatic speech recognition have improved the accuracy of end-to-end speech recognition systems, but recognition of personal content such as contact names remains a challenge. In this work, we describe our personalization solution for an end-to-end speech recognition system based on connectionist temporal classification. Building on previous work, we present a novel method for generating additional subword tokenizations for personal entities from their pronunciations. We show that using this technique in combination with two established techniques, contextual biasing and wordpiece prior normalization, we are able to achieve personal named entity accuracy on par with a competitive hybrid system.
Zhihong Lei, Ernest Pusateri, Shiyi Han, Leo Liu, Mingbin Xu, Tim Ng, Ruchir Travadi, Youyuan Zhang, Mirko Hannemann, Man-Hung Siu, Zhen Huang 0001
ICASSP5
2024 Enhancing CTC-based speech recognition with diverse modeling units
Shiyi Han, Mingbin Xu, Zhihong Lei, Zhen Huang 0001, Xingyu Na
INTERSPEECH2
2023 Acoustic Model Fusion For End-to-End Speech Recognition
abstract
Recent advances in deep learning and automatic speech recognition (ASR) have enabled the end-to-end (E2E) ASR system and boosted the accuracy to a new level. The E2E systems implicitly model all conventional ASR components, such as the acoustic model (AM) and the language model (LM), in a single network trained on audio-text pairs. Despite this simpler system architecture, fusing a separate LM, trained exclusively on text corpora, into the E2E system has proven to be beneficial. However, the application of LM fusion presents certain drawbacks, such as its inability to address the domain mismatch issue inherent to the internal AM. Drawing inspiration from the concept of LM fusion, we propose the integration of an external AM into the E2E system to better address the domain mismatch. By implementing this novel approach, we have achieved a significant reduction in the word error rate, with an impressive drop of up to 14.3% across varied test sets. We also discovered that this AM fusion approach is particularly beneficial in enhancing named entity recognition.
Zhihong Lei, Mingbin Xu, Shiyi Han, Leo Liu, Zhen Huang 0001, Tim Ng, Ernest Pusateri, Mirko Hannemann, Yaqiao Deng, Man-Hung Siu
ASRU2
2023 Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices
abstract
Federated Learning (FL) is a technique to train models on distributed edge devices with local data samples. Differential Privacy (DP) can be applied with FL to provide a formal privacy guarantee for sensitive data on device. Our goal is to train a large neural network language model (NNLM) on compute-constrained devices while preserving privacy using FL and DP. However, the noise required to guarantee differential privacy increases as the model size grows, which often prevents convergence. We propose Partial Embedding Updates (PEU), a novel technique to reduce the impact of DP-noise by decreasing payload size. Furthermore, we adopt Low Rank Adaptation (LoRA) and Noise Contrastive Estimation (NCE) to reduce the memory demands of large models on compute-constrained devices. We demonstrate in simulation and with real devices that this combination of techniques makes it possible to train large-vocabulary language models while preserving accuracy and privacy.
Mingbin Xu, Congzheng Song, Neha Agrawal, Filip Granqvist, Rogier C. van Dalen, Arturo Argueta, Shiyi Han, Yaqiao Deng, Leo Liu, Anmol Walia, Alex Jin
ICASSP1
2018 Dual Fixed-Size Ordinally Forgetting Encoding (FOFE) for Competitive Neural Language Models
abstract
In this paper, we propose a new approach to employ the fixed-size ordinally-forgetting encoding (FOFE) (Zhang et al., 2015b) in neural languages modelling, called dual-FOFE.The main idea behind dual-FOFE is that it allows to use two different forgetting factors so that it can avoid the trade-off in choosing either small or large values for the single forgetting factor in the original FOFE.In our experiments, we have compared the dual-FOFE based neural network language models (NNLM) against the original FOFE counterparts and various traditional NNLMs.Our results on the challenging Google Billion Words corpus show that both FOFE and dual FOFE yield very strong performance while significantly reducing the computational complexity over other NNLMs.Furthermore, the proposed dual-FOFE method further gives over 10% relative improvement in perplexity over the original FOFE model.
Sedtawut Watcharawittayakul, Mingbin Xu, Hui Jiang 0001
EMNLP2
2017 A Local Detection Approach for Named Entity Recognition and Mention Detection
abstract
In this paper, we study a novel approach for named entity recognition (NER) and mention detection (MD) in natural language processing.Instead of treating NER as a sequence labeling problem, we propose a new local detection approach, which relies on the recent fixed-size ordinally forgetting encoding (FOFE) method to fully encode each sentence fragment and its left/right contexts into a fixedsize representation.Subsequently, a simple feedforward neural network (FFNN) is learned to either reject or predict entity label for each individual text fragment.The proposed method has been evaluated in several popular NER and MD tasks, including CoNLL 2003 NER task and TAC-KBP2015 and TAC-KBP2016 Tri-lingual Entity Discovery and Linking (EDL) tasks.Our method has yielded pretty strong performance in all of these examined tasks.This local detection approach has shown many advantages over the traditional sequence labeling methods.
Mingbin Xu, Hui Jiang 0001, Sedtawut Watcharawittayakul
ACL (1)1
2017 Word Embeddings based on Fixed-Size Ordinally Forgetting Encoding
abstract
In this paper, we propose to learn word embeddings based on the recent fixedsize ordinally forgetting encoding (FOFE) method, which can almost uniquely encode any variable-length sequence into a fixed-size representation.We use FOFE to fully encode the left and right context of each word in a corpus to construct a novel word-context matrix, which is further weighted and factorized using truncated SVD to generate low-dimension word embedding vectors.We have evaluated this alternative method in encoding word-context statistics and show the new FOFE method has a notable effect on the resulting word embeddings.Experimental results on several popular word similarity tasks have demonstrated that the proposed method outperforms many recently popular neural prediction methods as well as the conventional SVD models that use canonical count based techniques to generate word context matrices.
Joseph Sanu, Mingbin Xu, Hui Jiang 0001, Quan Liu 0003
EMNLP2