VLDB 2026 Research / reviewers in the wild / expert
Hexin Liu
dblp:125/0337
· DBLP profile ↗
20ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-3998-9229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASRabstractAutomatic Speech Recognition (ASR) aims to convert human speech content into corresponding text. In conversational scenarios, effectively utilizing context can enhance its accuracy. Large Language Models' (LLMs) exceptional long-context understanding and reasoning abilities enable LLM-based ASR (LLM-ASR) to leverage historical context for recognizing conversational speech, which has a high degree of contextual relevance. However, existing conversational LLM-ASR methods use a fixed number of preceding utterances or the entire conversation history as context, resulting in significant ASR confusion and computational costs due to massive irrelevant and redundant information. This paper proposes a multi-modal retrieval-and-selection method named MARS that augments conversational LLM-ASR by enabling it to retrieve and select the most relevant acoustic and textual historical context for the current utterance. Specifically, multi-modal retrieval obtains a set of candidate historical contexts, each exhibiting high acoustic or textual similarity to the current utterance. Multi-modal selection calculates the acoustic and textual similarities for each retrieved candidate historical context and, by employing our proposed near-ideal ranking method to consider both similarities, selects the best historical context. Evaluations on the Interspeech 2025 Multilingual Conversational Speech Language Model Challenge dataset show that the LLM-ASR, when trained on only 1.5K hours of data and equipped with the MARS, outperforms the state-of-the-art top-ranking system trained on 179K hours of data. Bingshen Mu, Hexin Liu, Hongfei Xue, Lei Xie 0001 |
AAAI | 2 |
| 2026 | LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form SpeechabstractForced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts.The multilingual speech understanding and longsequence processing abilities of speech large language models (SLLMs) make them promising for FA in multilingual, crosslingual, and long-form speech settings.However, directly applying the next-token prediction paradigm of SLLMs to FA results in hallucinations and slow inference.To bridge the gap, we propose LLM-ForcedAligner, reformulating FA as a slot-filling paradigm: timestamps are treated as discrete indices, and special timestamp tokens are inserted as slots into the transcript.Conditioned on the speech embeddings and the transcript with slots, the SLLM directly predicts the time indices at slots.During training, causal attention masking with non-shifted input and label sequences allows each slot to predict its own timestamp index based on itself and preceding context, with loss computed only at slot positions.Dynamic slot insertion enables FA at arbitrary positions.Moreover, non-autoregressive inference is supported, avoiding hallucinations and improving speed.Experiments across multilingual, crosslingual, and long-form speech scenarios show that LLM-ForcedAligner achieves a 69%~78% relative reduction in accumulated averaging shift compared with prior methods.Checkpoint and inference code are available at https://huggingface.co/Qwen/ Qwen3-ForcedAligner-0.6B. Bingshen Mu, Xian Shi, Hexin Liu, Jin Xu 0010, Lei Xie 0001 |
ACL (1) | 4 |
| 2026 | Evaluating the Expressive Appropriateness of Speech in Rich ContextsabstractTianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ACL (1) | 13 |
| 2026 | Chronological Thinking in Full-Duplex Spoken Dialogue Language ModelsabstractRecent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user speech streams while generating responses. This simultaneous listening and speaking design enables real-time interaction and the agent can handle dynamic conversational behaviors like user barge-in. However, during the listening phase, existing systems keep the agent idle by repeatedly predicting the silence token, which departs from human behavior: we usually engage in lightweight thinking during conversation rather than remaining absent-minded. Inspired by this, we propose Chronological Thinking, an on-the-fly conversational thinking mechanism that aims to improve response quality in full-duplex SDLMs. Specifically, chronological thinking presents a paradigm shift from conventional LLM thinking approaches, such as Chain-of-Thought, purpose-built for streaming acoustic input. (1) Strictly causal: the agent reasons incrementally while listening, updating internal hypotheses only from past audio with no lookahead. (2) No additional latency: reasoning is amortized during the listening window; once the user stops speaking, the agent halts thinking and begins speaking without further delay. Experiments demonstrate the effectiveness of chronological thinking through both objective metrics and human evaluations show consistent improvements in response quality. Furthermore, chronological thinking robustly handles conversational dynamics and attains competitive performance on full-duplex interaction metrics. Donghang Wu, Chen Chen 0075, Xuerui Yang, Gang Yu 0002, Hexin Liu, Nana Hou, Chng Eng Siong |
SIGDIAL | 8 |
| 2025 | Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware DecodingabstractCode-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both encoder and decoder parts. First, we propose an encoder refiner to enhance the encoder’s capacity of intra-sentence swithching. Second, we propose using two sets of language-aware adapters with different language prompt embeddings to achieve language-specific decoding information in each decoder layer. Then, a fusion module is added to fuse the language-aware decoding. The experimental results using the SEAME dataset show that, compared with the baseline model, the proposed approach achieves a relative MER reduction of 4.1% and 7.2% on the dev_man and dev_sge test sets, respectively, surpassing state-of-the-art methods. Through experiments, we found that the proposed method significantly improves the performance on non-native language in CS speech, indicating that our approach enables Whisper to better distinguish between the two languages. Chenrui Cui, Tianrui Wang, Hexin Liu, Zhaoheng Ni, Lingxuan Ye, Longbiao Wang |
ICASSP | 5 |
| 2025 | GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingabstractSemantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns and contextual cues, to reconstruct incomplete or masked speech signals in noisy environments. However, existing speech enhancement (SE) approaches often overlook the rich semantic information embedded in speech, which is crucial for improving intelligibility, speaker consistency, and overall quality of enhanced speech signals. To enrich the SE model with semantic information, we employ language models as an efficient semantic learner and propose a comprehensive framework tailored for language model-based speech enhancement, called GenSE. Specifically, we approach SE as a conditional language modeling task rather than a continuous signal regression problem defined in existing works. This is achieved by tokenizing speech signals into semantic tokens using a pre-trained self-supervised model and into acoustic tokens using a custom-designed single-quantizer neural codec model. To improve the stability of language model predictions, we propose a hierarchical modeling method that decouples the generation of clean semantic tokens and clean acoustic tokens into two distinct stages. Moreover, we introduce a token chain prompting mechanism during the acoustic token generation stage to ensure timbre consistency throughout the speech enhancement process. Experimental results on benchmark datasets demonstrate that our proposed approach outperforms state-of-the-art SE systems in terms of speech quality and generalization capability. Codes and demos are publicly available at https://anonymous.4open.science/w/gen-se-7F52/. Jixun Yao, Hexin Liu, Chen Chen 0075, Chng Eng Siong, Lei Xie 0001 |
ICLR | 2 |
| 2025 | EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
Jixun Yao, Hexin Liu, Chng Eng Siong, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2025 | Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT ReasoningabstractLarge language models have been extended to the speech domain, leading to the development of speech large language models (SLLMs). While existing SLLMs demonstrate strong performance in speech instruction-following for core languages (e.g., English), they often struggle with non-core languages due to the scarcity of paired speech-text data and limited multilingual semantic reasoning capabilities. To address this, we propose the semi-implicit Cross-lingual Speech Chain-of-Thought (XS-CoT) framework, which integrates speech-to-text translation into the reasoning process of SLLMs. The XS-CoT generates four types of tokens: instruction and response tokens in both core and non-core languages, enabling cross-lingual transfer of reasoning capabilities. To mitigate inference latency in generating target non-core response tokens, we incorporate a semi-implicit CoT scheme into XS-CoT, which progressively compresses the first three types of intermediate reasoning tokens while retaining global reasoning logic during training. By leveraging the robust reasoning capabilities of the core language, XS-CoT improves responses for non-core languages by up to 45% in GPT-4 score when compared to direct supervised fine-tuning on two representative SLLMs, Qwen2-Audio and SALMONN. Moreover, the semi-implicit XS-CoT reduces token delay by more than 50% with a slight drop in GPT-4 scores. Importantly, XS-CoT requires only a small amount of high-quality training data for non-core languages by leveraging the reasoning capabilities of core languages. To support training, we also develop a data pipeline and open-source speech instruction-following datasets in Japanese, German, and French. Hongfei Xue, Yufeng Tang, Hexin Liu, Xuelong Geng, Lei Xie 0001 |
ACM Multimedia | 3 |
| 2024 | Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion ModelabstractXiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, Leibny Paola Garcia Perera, EngSiong Chng, Lina Yao. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xiangyu Zhang 0005, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, L. Paola García-Perera, Chng Eng Siong |
EMNLP | 3 |
| 2024 | When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression DetectionabstractDepression is a critical concern in global mental health, prompting extensive research into AIbased detection methods.Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications.However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities.Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped.In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection.We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks.By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts.This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals.Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines.In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals. Xiangyu Zhang 0005, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps |
EMNLP | 2 |
| 2024 | Unidirectional Brain-Computer Interface: Artificial Neural Network Encoding Natural Images to FMRI Response in the Visual CortexabstractWhile significant advancements in artificial intelligence (AI) have catalyzed progress across various domains, its full potential in understanding visual perception remains underexplored. We propose an artificial neural network dubbed VISION, an acronym for "Visual Interface System for Imaging Output of Neural activity," to mimic the human brain and show how it can foster neuroscientific inquiries. Using visual and contextual inputs, this multimodal model predicts the brain's functional magnetic resonance imaging (fMRI) scan response to natural images. VISION successfully predicts human hemodynamic responses as fMRI voxel values to visual inputs with an accuracy exceeding state-of-the-art performance by 45%. We further probe the trained networks to reveal representational biases in different visual areas, generate experimentally testable hypotheses, and formulate an interpretable metric to associate these hypotheses with cortical functions. With both a model and evaluation metric, the cost and time burdens associated with designing and implementing functional analysis on the visual cortex could be reduced. Our work suggests that the evolution of computational models may shed light on our fundamental understanding of the visual cortex and provide a viable approach toward reliable brain-machine interfaces. Ruixing Liang, Xiangyu Zhang 0005, Hexin Liu, Avisha Kumar, Kelley M. Kempski Leadingham, Joshua Punnoose, L. Paola García-Perera, Amir Manbachi |
ICASSP | 5 |
| 2024 | Enhancing Code-Switching Speech Recognition With Interactive Language BiasesabstractLanguages usually switch within a multilingual speech signal, especially in a bilingual society. This phenomenon is referred to as code-switching (CS), making automatic speech recognition (ASR) challenging under a multilingual scenario. We propose to improve CS-ASR by biasing the hybrid CTC/attention ASR model with multi-level language information comprising frame-and token-level language posteriors. The interaction between various resolutions of language biases is subsequently explored in this work. We conducted experiments on datasets from the ASRU 2019 code-switching challenge. Compared to the baseline, the proposed interactive language biases (ILB) method achieves higher performance and ablation studies highlight the effects of different language biases and their interactions. In addition, the results presented indicate that language bias implicitly enhances internal language modeling, leading to performance degradation after employing an external language model. Hexin Liu, L. Paola García-Perera, Xiangyu Zhang 0005, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 1 |
| 2024 | Bridging Child-Centered Speech Language Identification and Language Diarization via Phonetics
Hexin Liu, L. Paola García-Perera |
INTERSPEECH | 2 |
| 2023 | PQLM - Multilingual Decentralized Portable Quantum Language ModelabstractWith careful manipulation, malicious agents can reverse engineer private information encoded in pre-trained language models. Security concerns motivate the development of quantum pre-training. In this work, we propose a highly portable quantum language model (PQLM) that can easily transmit information to downstream tasks on classical machines. The framework consists of a cloud PQLM built with random Variational Quantum Classifiers (VQC) and local models for downstream applications. We demonstrate the ad hoc portability of the quantum model by extracting only the word embeddings and effectively applying them to downstream tasks on classical machines. Our PQLM exhibits comparable performance to its classical counterpart on both intrinsic evaluation (loss, perplexity) and extrinsic evaluation (multilingual sentiment analysis accuracy) metrics. We also perform ablation studies on the factors affecting PQLM performance to analyze model stability. Our work establishes a theoretical foundation for a portable quantum pre-trained language model that could be trained on private data and made available for public use with privacy protection guarantees. Shuyue Stella Li, Xiangyu Zhang 0005, Hongchao Shu, Ruixing Liang, Hexin Liu, L. Paola García-Perera |
ICASSP | 6 |
| 2023 | Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language DiarizationabstractCode-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and disentangling language information. We incorporate language information within the CS-ASR model by dynamically biasing the model with token-level language posteriors corresponding to outputs of a sequence-to-sequence auxiliary language diarization (LD) module. In contrast, the disentangling process reduces the difference between languages via adversarial training so as to normalize two languages. We conduct experiments on the SEAME dataset. Compared to the baseline model, both the joint optimization with LD and the language posterior bias achieve performance improvement. Comparison of the proposed methods indicates that incorporating language information is more effective than disentangling for reducing language confusion in CS speech. Hexin Liu, Haihua Xu 0001, L. Paola García-Perera, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 1 |
| 2023 | MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarizationabstractTo enhance the reliability and robustness of language identification (LID) and language diarization (LD) systems for heterogeneous populations and scenarios, there is a need for speech processing models to be trained on datasets that feature diverse language registers and speech patterns. We present the MERLIon CCS challenge, featuring a first-of-its-kind Zoom video call dataset of parent-child shared book reading, of over 30 hours with over 300 recordings, annotated by multilingual transcribers using a high-fidelity linguistic transcription protocol. The audio corpus features spontaneous and in-the-wild English-Mandarin code-switching, child-directed speech in non-standard accents with diverse language-mixing patterns recorded in a variety of home environments. This report describes the corpus, as well as LID and LD results for our baseline and several systems submitted to the MERLIon CCS challenge using the corpus. Yi Han Victoria Chua, Hexin Liu, L. Paola García-Perera, Fei Ting Woon, Jinyi Wong, Xiangyu Zhang 0005, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels, Suzy J. Styles |
INTERSPEECH | 2 |
| 2023 | Self-supervised Learning Representation based Accent Recognition with Persistent Accent Memory
Zhiwei Xie 0007, Haihua Xu 0001, Yizhou Peng, Hexin Liu, Hao Huang 0009, Chng Eng Siong |
INTERSPEECH | 5 |
| 2023 | Investigating model performance in language identification: beyond simple error statisticsabstractLanguage development experts need tools that can automatically identify languages from fluent, conversational speech and provide reliable estimates of usage rates at the level of an individual recording. However, LID systems are typically evaluated on metrics such as equal error rate and balanced accuracy, applied at the level of an entire speech corpus. These overview metrics do not provide information about model performance at the level of individual speakers, recordings, or units of speech with different linguistic characteristics. Overview statistics may mask systematic errors in model performance for some subsets of the data, and consequently, have worse performance on data derived from some subsets of human speakers, creating a kind of algorithmic bias. Here, we investigate how well a number of LID systems perform on individual recordings and speech units with different linguistic properties in the MERLIon CCS Challenge featuring accented code-switched child-directed speech. Suzy J. Styles, Yi Han Victoria Chua, Fei Ting Woon, Hexin Liu, L. Paola García-Perera, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels |
INTERSPEECH | 4 |
| 2022 | PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language IdentificationabstractWe propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training.In this model, named PHO-LID, a self-supervised phoneme segmentation task and a LID task share a convolutional neural network (CNN) module, which encodes both language identity and sequential phonemic information in the input speech to generate an intermediate sequence of "phonotactic" embeddings.These embeddings are then fed into transformer encoder layers for utterance-level LID.We call this architecture CNN-Trans.We evaluate it on AP17-OLR data and the MLS14 set of NIST LRE 2017, and show that the PHO-LID model with multitask optimization exhibits the highest LID performance among all models, achieving over 40% relative improvement in terms of average cost on AP17-OLR data compared to a CNN-Trans model optimized only for LID.The visualized confusion matrices imply that our proposed method achieves higher performance on languages of the same cluster in NIST LRE 2017 data than the CNN-Trans model.A comparison between predicted phoneme boundaries and corresponding audio spectrograms illustrates the leveraging of phoneme information for LID. Hexin Liu, L. Paola García-Perera, Andy W. H. Khong, Suzy J. Styles, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2021 | End-to-End Language Diarization for Bilingual Code-Switching SpeechabstractWe propose two end-to-end neural configurations for language diarization on bilingual code-switching speech. The first, a BLSTM-E2E architecture, includes a set of stacked bidirectional LSTMs to compute embeddings and incorporates the deep clustering loss to enforce grouping of languages belonging to the same class. The second, an XSA-E2E architecture, is based on an x-vector model followed by a self-attention encoder. The former encodes frame-level features into segmentlevel embeddings while the latter considers all those embeddings to generate a sequence of segment-level language labels. We evaluated the proposed methods on the dataset obtained from the shared task B in WSTCSMC 2020 and our handcrafted simulated data from the SEAME dataset. Experimental results show that our proposed XSA-E2E architecture achieved a relative improvement of 12.1% in equal error rate and a 7.4% relative improvement on accuracy compared with the baseline algorithm in the WSTCSMC 2020 dataset. Our proposed XSA-E2E architecture achieved an accuracy of 89.84% with a baseline of 85.60% on the simulated data derived from the SEAME dataset. Hexin Liu, L. Paola García-Perera, Justin Dauwels, Andy W. H. Khong, Sanjeev Khudanpur, Suzy J. Styles |
Interspeech | 1 |