Minglun Han

dblp:281/7284 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-5120-069XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Integrate-and-Fire Compressor: Learning to Compress Context for LLMs Adaptively
abstract
Large language models (LLMs) face significant challenges in long-context modeling due to increased inference costs, higher latency, and performance degradation caused by information redundancy. Context compression offers a promising solution, but existing methods often rely on fixed strategies that don’t adapt to variations in information density. To address this, we propose the Integrate-and-Fire Compressor (IFC), an adaptive method inspired by the neural integrate-and-fire mechanism. IFC assesses token importance more accurately and dynamically adjusts compression, producing compact and effective context representations for LLMs, even at high compression rates. Experiments on three open-domain QA datasets show that IFC consistently outperforms existing baselines in both compression rate and task performance, highlighting its potential to improve the efficiency and effectiveness of LLMs in handling long contexts.
Xiyun Li, Minglun Han, Bo Xu 0002
ICME5
2024 ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition
abstract
Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues can also benefit in many scenarios. In this paper, we first propose ViLaS (Vision and Language into Automatic Speech Recognition), a novel multimodal ASR model based on the continuous integrate-and-fire (CIF) mechanism, which can integrate visual and textual context simultaneously or separately, to facilitate speech recognition. Next, we introduce an effective training strategy that improves performance in modal-incomplete test scenarios. Then, to explore the effects of integrating vision and language, we create VSDial, a multimodal ASR dataset with multimodal context cues in both Chinese and English versions. Finally, empirical results are reported on the public Flickr8K and self-constructed VSDial datasets. We explore various cross-modal fusion schemes, analyze fine-grained cross-modal alignment on VSDial, and provide insights into the effects of integrating multimodal information on speech recognition.
Ziyi Ni, Minglun Han, Linghui Meng 0001, Jing Shi 0003, Bo Xu 0002
ICASSP2
2023 Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition
abstract
The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with different scales of neuronal dynamics is necessary. Here we introduce four types of neuronal dynamics to post-process the sequential patterns generated from the spiking transformer to get the complex dynamic neuron improved spiking transformer neural network (DyTr-SNN). We found that the DyTr-SNN could handle the non-toy automatic speech recognition task well, representing a lower phoneme error rate, lower computational cost, and higher robustness. These results indicate that the further cooperation of SNNs and neural dynamics at the neuron and network scales might have much in store for the future, especially on the ASR tasks.
Tielin Zhang, Minglun Han, Yi Wang 0077, Duzhen Zhang, Bo Xu 0002
AAAI3
2023 Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding
abstract
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms in medical conversations. In this work, we formalize MSF into a matching problem and propose a Term Semantics Pre-trained Matching Network (TSPMN) that takes both terms and queries as input to model their semantic inter-action. To learn term semantics better, we further design two self-supervised objectives, including Contrastive Term Discrimination (CTD) and Matching-based Mask Term Modeling (MMTM). CTD determines whether it is the masked term in the dialogue for each given term, while MMTM directly predicts the masked ones. Experimental results on two Chinese benchmarks show that TSPMN outperforms strong baselines, especially in few-shot settings1.
Zefa Hu, Xiuyi Chen, Minglun Han, Ziyi Ni, Jing Shi 0003, Bo Xu 0002
ICASSP4
2023 Enhancing Visual Question Answering via Deconstructing Questions and Explicating Answers
abstract
A compositional question refers to a question that involves multiple visual objects, as well as their attributes and relationships, which requires compositional reasoning to answer.Existing VQA models can well answer a compositional question, but few works can give the reasoning process and explain why this answer is given.In this paper, we propose a novel model (DEEX) to enhance visual question answering via DEconstructing questions and EXplicating answers when answering compositional questions.Specifically, DEEX aims to accomplish three sub-tasks: (1) Compositional Question Answering (CQA), (2) Question Deconstructing (QD), and (3) Answer Explicating (AE).We utilize prompt-based multi-task learning to train the proposed DEEX to be able to answer questions and give explanations simultaneously.Experimental results on the GQA dataset demonstrate our method's effectiveness, which can enhance visual question answering by giving corresponding reasoning processes and explanations.
Minglun Han, Jing Shi 0003, Bo Xu 0002
INTERSPEECH2
2023 Knowledge Transfer from Pre-trained Language Models to Cif-based Speech Recognizers via Hierarchical Distillation
Minglun Han, Jing Shi 0003, Bo Xu 0002
INTERSPEECH1
2022 Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection
abstract
Nowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between similar context-specific phrases, which hurts predictions at the token level. In this work, we focus on mitigating confusion problems with fine-grained contextual knowledge selection (FineCoS). In FineCoS, we introduce fine-grained knowledge to reduce the uncertainty of token predictions. Specifically, we first apply phrase selection to narrow the range of phrase candidates, and then conduct token attention on the tokens in the selected phrase candidates. Moreover, we re-normalize the attention weights of most relevant phrases in inference to obtain more focused phrase-level contextual representations, and inject position information to help model better discriminate phrases or tokens. On LibriSpeech and an in-house 160,000-hour dataset, we explore the proposed methods based on an all-neural biasing method, collaborative decoding (ColDec). The proposed methods further bring at most 6.1% relative word error rate reduction on LibriSpeech and 16.4% relative character error rate reduction on the in-house dataset.
Minglun Han, Linhao Dong, Zhenlin Liang, Zejun Ma 0001, Bo Xu 0002
ICASSP1
2021 Cif-Based Collaborative Decoding for End-to-End Contextual Speech Recognition
abstract
End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure and the E2E training hamper injecting context information into them for contextual biasing. Though contextual LAS (CLAS) gives an excellent all-neural solution, the degree of biasing to given contextual information is not explicitly controllable. In this paper, we focus on incorporating contextual information into the continuous integrate-and-fire (CIF) based model that supports contextual biasing in a more controllable fashion. Specifically, an extra context processing network is introduced to extract contextual embeddings, integrate acoustically relevant contextual information and decode the contextual output distribution, thus forming a collaborative decoding with the decoder of the CIF-based model. Evaluated on the named entity rich evaluation sets of HKUST/AISHELL-2, our method brings relative character error rate (CER) reduction of 8.83%/21.13% and relative named entity character error rate (NE-CER) reduction of 40.14%/51.50% when compared with a strong baseline. Besides, it keeps the performance on original evaluation set without degradation.
Minglun Han, Linhao Dong, Bo Xu 0002
ICASSP1