Xiao Liu 0029

dblp:82/1364-29 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-8893-366XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
abstract
Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Lei Ji 0001, Xing W, Haizhen Huang, Yeyun Gong, Zhijiang Guo, Xiao Liu 0029, Cheng-Lin Liu 0001
ACL (1)12
2026 EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain
abstract
It is crucial to understand a specific domain by events.Extensive event extraction research has been conducted in many domains such as news, finance, and biology.However, event extraction in scientific domain is still insufficiently supported by comprehensive datasets and tailored methods.Compared with other domains, scientific domain has two characteristics: (1) denser nuggets and events, and (2) more complex information forms.To solve the above problem, considering these two characteristics, we first construct SciEvents, a large-scale multi-event document-level dataset with a schema tailored for scientific domain.It consists of 2,508 documents and 24,381 events under multi-stage manual annotation and quality control.Then, we propose EXCEEDS, an end-to-end scientific event extraction framework by encoding dense nuggets into a grid matrix and simplifying complex event extraction as a nugget-based grid modeling task.Experiments on SciEvents demonstrate state-of-the-art performances of EXCEEDS.Both the SciEvents dataset and the EXCEEDS framework are released publicly to facilitate future research.
Yi-Fan Lu, Xianling Mao, Bo Wang 0134, Xiao Liu 0029, Heyan Huang
ACL (1)4
2026 COSMOS: Connectivity-Oriented Submodular Maximization for Optimal Subgraph Retrieval
abstract
Boci Peng, Xiao Liu, Boren Hu, Yun Zhu, Xuanbo Fan, Yanwei Yue, Chunyu Yang, Yan Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boci Peng, Xiao Liu 0029, Boren Hu, Yun Zhu 0007, Xuanbo Fan, Yanwei Yue, Chunyu Yang 0005, Yan Zhang 0117
ACL (1)2
2026 Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
abstract
Kailai Yang, Xiao Liu, Lei Ji, Hao Li, Xiao Liang, Zhiwei Liu, Yeyun Gong, Peng Cheng, Mao Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kailai Yang, Xiao Liu 0029, Lei Ji 0001, Hao Li 0074, Zhiwei Liu 0003, Yeyun Gong, Peng Cheng 0005, Mao Yang 0004
ACL (1)2
2025 Key-Point-Driven Data Synthesis with Its Enhancement on Mathematical Reasoning
abstract
Large language models have shown great potential in complex reasoning tasks, yet their performance is often hampered by the scarcity of high-quality and reasoning-focused training datasets. Addressing this challenge, we propose Key-PointDriven Data Synthesis (KPDDS), a novel data synthesis framework that synthesizes question-answer pairs by leveraging key points and exemplar practices from authentic data sources. KPDDS ensures the generation of novel questions with rigorous quality control and substantial scalability. As a result, we present KPMath, an extensive synthetic dataset tailored for mathematical reasoning, comprising over 800K questionanswer pairs. Utilizing KPMath and augmenting it with additional reasoning-intensive corpora, we create the comprehensive KPMath-Plus dataset. Our experiments demonstrate that this dataset can enhance the mathematical reasoning performance of models across various architectures and sizes. The Qwen1.5-72B model, fine-tuned on KPMath-Plus, achieves 87.0% accuracy on GSM8K and 58.3% on MATH, surpassing competitors in the 7B to 72B range and best commercial models like GPT-4 across multiple math reasoning datasets.
Xiao Liu 0029, Yeyun Gong, Zhibin Gou, Yelong Shen, Nan Duan 0001, Weizhu Chen
AAAI2
2025 Teaching Your Models to Understand Code via Focal Preference Alignment
abstract
Jie Wu, Haoling Li, Xin Zhang, Xiao Liu, Yangyu Huang, Jianwen Luo, Yizhen Zhang, Zuchao Li, Ruihang Chu, Yujiu Yang, Scarlett Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jie Wu 0001, Haoling Li, Xin Zhang 0099, Xiao Liu 0029, Yangyu Huang, Zuchao Li, Ruihang Chu, Yujiu Yang 0001, Scarlett Li
EMNLP4
2025 Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
abstract
Using large teacher models to guide the training of smaller student models has become the prevailing paradigm for efficient and effective learning. However, vocabulary mismatches between teacher and student language models pose significant challenges in language modeling, resulting in divergent token sequences and output distributions. To overcome these limitations, we propose Vocabulary-agnostic Teacher Guided Language Modeling (VocAgnoLM), a novel approach that bridges the gap caused by vocabulary mismatch through two key methods: (1) Token-level Lexical Alignment, which aligns token sequences across mismatched vocabularies, and (2) Teacher Guided Loss, which leverages the loss of teacher model to guide effective student training. We demonstrate its effectiveness in language modeling with 1B student model using various 7B teacher models with different vocabularies. Notably, with Qwen2.5-Math-Instruct, a teacher model sharing only about 6% of its vocabulary with TinyLlama, VocAgnoLM achieves a 46% performance improvement compared to naive continual pretraining. Furthermore, we demonstrate that VocAgnoLM consistently benefits from stronger teacher models, providing a robust solution to vocabulary mismatches in language modeling.
Haebin Shin, Lei Ji 0001, Xiao Liu 0029, Yeyun Gong
ICML3
2025 Optimizing Large Language Model Training Using FP4 Quantization
abstract
The growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a challenge due to significant quantization errors and limited representational capacity. This work introduces the first FP4 training framework for LLMs, addressing these challenges with two key innovations: a differentiable quantization estimator for precise weight updates and an outlier clamping and compensation strategy to prevent activation collapse. To ensure stability, the framework integrates a mixed-precision training scheme and vector-wise quantization. Experimental results demonstrate that our FP4 framework achieves accuracy comparable to BF16 and FP8, with minimal degradation, scaling effectively to 13B-parameter LLMs trained on up to 100B tokens. With the emergence of next-generation hardware supporting FP4, our framework sets a foundation for efficient ultra-low precision training.
Yeyun Gong, Xiao Liu 0029, Guoshuai Zhao 0001, Ziyue Yang 0002, Baining Guo, Zhengjun Zha, Peng Cheng 0005
ICML3
2025 EpiCoder: Encompassing Diversity and Complexity in Code Generation
abstract
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available.
Yaoxiang Wang, Haoling Li, Xin Zhang 0099, Jie Wu 0001, Xiao Liu 0029, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Yujiu Yang 0001, Jinsong Su, Qi Chen 0009, Scarlett Li
ICML5
2025 Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
abstract
Sungjin Park, Xiao Liu, Yeyun Gong, Edward Choi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xiao Liu 0029, Yeyun Gong, Edward Choi 0003
NAACL (Long Papers)2
2024 Using Left and Right Brains Together: Towards Vision and Language Planning
abstract
Large Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks. However, they inherently operate planning within the language space, lacking the vision and spatial imagination ability. In contrast, humans utilize both left and right hemispheres of the brain for language and visual planning during the thinking process. Therefore, we introduce a novel vision-language planning framework in this work to perform concurrent visual and language planning for tasks with inputs of any form. Our framework incorporates visual planning to capture intricate environmental details, while language planning enhances the logical coherence of the overall system. We evaluate the effectiveness of our framework across vision-language tasks, vision-only tasks, and language-only tasks. The results demonstrate the superior performance of our approach, indicating that the integration of visual and language planning yields better contextually aware task execution.
Jun Cen, Chenfei Wu, Xiao Liu 0029, Shengming Yin, Yixuan Pei, Jinglong Yang, Qifeng Chen 0001, Nan Duan 0001
ICML3
2024 Not All Tokens Are What You Need for Pretraining
abstract
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that ''Not all tokens in a corpus are equally important for language model training''. Our initial analysis examines token-level training dynamics of language model, revealing distinct loss patterns for different tokens. Leveraging these insights, we introduce a new language model called Rho-1. Unlike traditional LMs that learn to predict every next token in a corpus, Rho-1 employs Selective Language Modeling (SLM), which selectively trains on useful tokens that aligned with the desired distribution. This approach involves scoring training tokens using a reference model, and then training the language model with a focused loss on tokens with higher scores. When continual continual pretraining on 15B OpenWebMath corpus, Rho-1 yields an absolute improvement in few-shot accuracy of up to 30% in 9 math tasks. After fine-tuning, Rho-1-1B and 7B achieved state-of-the-art results of 40.6% and 51.8% on MATH dataset, respectively - matching DeepSeekMath with only 3% of the pretraining tokens. Furthermore, when continual pretraining on 80B general tokens, Rho-1 achieves 6.8% average enhancement across 15 diverse tasks, increasing both data efficiency and performance of the language model pre-training.
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 0029, Yelong Shen, Ruochen Xu, Chen Lin 0001, Yujiu Yang 0001, Jian Jiao 0007, Nan Duan 0001, Weizhu Chen
NeurIPS4
2024 LEAD: Liberal Feature-based Distillation for Dense Retrieval
abstract
Knowledge distillation is often used to transfer knowledge from a strong teacher model to a relatively weak student model. Traditional methods include response-based methods and feature-based methods. Response-based methods are widely used but suffer from lower upper limits of performance due to their ignorance of intermediate signals, while feature-based methods have constraints on vocabularies, tokenizers and model architectures. In this paper, we propose a liberal feature-based distillation method (LEAD). LEAD aligns the distribution between the intermediate layers of teacher model and student model, which is effective, extendable, portable and has no requirements on vocabularies, tokenizers, or model architectures. Extensive experiments show the effectiveness of LEAD on widely-used benchmarks, including MS MARCO Passage Ranking, TREC 2019 DL Track, MS MARCO Document Ranking and TREC 2020 DL Track. Our code is available in https://github.com/microsoft/SimXNS/tree/main/LEAD.
Hao Sun 0015, Xiao Liu 0029, Yeyun Gong, Anlei Dong, Jingwen Lu, Yan Zhang 0117, Linjun Yang, Rangan Majumder, Nan Duan 0001
WSDM2
2023 AR-Diffusion: Auto-Regressive Diffusion Model for Text Generation
abstract
Diffusion models have gained significant attention in the realm of image generation due to their exceptional performance. Their success has been recently expanded to text generation via generating all tokens within a sequence concurrently. However, natural language exhibits a far more pronounced sequential dependency in comparison to images, and the majority of existing language models are trained with a left-to-right auto-regressive approach. To account for the inherent sequential characteristic of natural language, we introduce Auto-Regressive Diffusion (AR-Diffusion). AR-Diffusion ensures that the generation of tokens on the right depends on the generated ones on the left, a mechanism achieved through employing a dynamic number of denoising steps that vary based on token position. This results in tokens on the left undergoing fewer denoising steps than those on the right, thereby enabling them to generate earlier and subsequently influence the generation of tokens on the right. In a series of experiments on various text generation tasks, including text summarization, machine translation, and common sense generation, AR-Diffusion clearly demonstrated its superiority over existing diffusion language models and that it can be $100\times\sim600\times$ faster when achieving comparable results. Our code is available at https://github.com/microsoft/ProphetNet/tree/master/AR-diffusion.
Zhihao Fan, Xiao Liu 0029, Hai-Tao Zheng 0002, Yeyun Gong, Yelong Shen, Jian Jiao 0007, Zhongyu Wei, Jian Guo 0016, Nan Duan 0001, Weizhu Chen
NeurIPS3
2023 MASTER: Multi-task Pre-trained Bottlenecked Masked Autoencoders Are Better Dense Retrievers
Kun Zhou 0002, Xiao Liu 0029, Yeyun Gong, Wayne Xin Zhao, Daxin Jiang, Nan Duan 0001, Ji-Rong Wen
ECML/PKDD (2)2
2023 Unsupervised Dense Retrieval Training with Web Anchors
abstract
In this work, we present an unsupervised retrieval method with contrastive learning on web anchors. The anchor text describes the content that is referenced from the linked page. This shows similarities to search queries that aim to retrieve pertinent information from relevant documents. Based on their commonalities, we train an unsupervised dense retriever, Anchor-DR, with a contrastive learning task that matches the anchor text and the linked document. To filter out uninformative anchors (such as "homepage" or other functional anchors), we present a novel filtering technique to only select anchors that contain similar types of information as search queries. Experiments show that Anchor-DR outperforms state-of-the-art methods on unsupervised dense retrieval by a large margin (e.g., by 5.3% NDCG@10 on MSMARCO). The gain of our method is especially significant for search and question answering tasks. Our analysis further reveals that the pattern of anchor-document pairs is similar to that of search query-document pairs. Code available at https://github.com/Veronicium/AnchorDR.
Yiqing Xie, Xiao Liu 0029, Chenyan Xiong
SIGIR2
2023 PROD: Progressive Distillation for Dense Retrieval
abstract
Knowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model. Ideally, we expect the better the teacher is, the better the student performs. However, this expectation does not always come true. It is common that a strong teacher model results in a bad student via distillation due to the nonnegligible gap between teacher and student. To bridge the gap, we propose PROD, a PROgressive Distillation method, for dense retrieval. PROD consists of a teacher progressive distillation and a data progressive distillation to gradually improve the student. To alleviate catastrophic forgetting, we introduce a regularization term in each distillation process. We conduct extensive experiments on seven datasets including five widely-used publicly available benchmarks: MS MARCO Passage, TREC Passage 19, TREC Document 19, MS MARCO Document, and Natural Questions, as well as two industry datasets: Bing-Rel and Bing-Ads. PROD achieves the state-of-the-art in the distillation methods for dense retrieval. Our 6-layer student model even surpasses most of the existing 12-layer models on all five public benchmarks. The code and models are released in https://github.com/microsoft/SimXNS.
Zhenghao Lin, Yeyun Gong, Xiao Liu 0029, Hang Zhang 0029, Chen Lin 0001, Anlei Dong, Jian Jiao 0007, Jingwen Lu, Daxin Jiang, Rangan Majumder, Nan Duan 0001
WWW3
2023 Event Extraction With Dynamic Prefix Tuning and Relevance Retrieval
abstract
We consider event extraction in a generative manner with template-based conditional generation. Although there is a rising trend of casting the task of event extraction as a sequence generation problem with prompts, these generation-based methods have several significant challenges, including using suboptimal prompts, static event type information, and the overwhelming number of irrelevant event types. In this article, we propose a generative template-based method with dynamic prefixes and a relevance retrieval framework for event extraction (GREE) by first integrating context information with type-specific prefixes to learn a context-specific prefix for each context, and then retrieving the relevant event types with an adaptive threshold. Experimental results show that our model achieves competitive results with the state-of-the-art classification-based modelOneIEon ACE 2005 and achieves the best performances on ERE. Additionally, our model is proven to be portable to new types of events effectively.
Heyan Huang, Xiao Liu 0029, Ge Shi 0002, Qian Liu 0012
IEEE Trans. Knowl. Data Eng.2
2022 Dynamic Prefix-Tuning for Generative Template-based Event Extraction
abstract
We consider event extraction in a generative manner with template-based conditional generation.Although there is a rising trend of casting the task of event extraction as a sequence generation problem with prompts, these generation-based methods have two significant challenges, including using suboptimal prompts and static event type information.In this paper, we propose a generative templatebased event extraction method with dynamic prefix (GTEE-DYNPREF) by integrating context information with type-specific prefixes to learn a context-specific prefix for each context.Experimental results show that our model achieves competitive results with the state-ofthe-art classification-based model ONEIE on ACE 2005 and achieves the best performances on ERE.Additionally, our model is proven to be portable to new types of events effectively.
Xiao Liu 0029, Heyan Huang, Ge Shi 0002, Bo Wang 0134
ACL (1)1
2022 BIT-WOW at NLPCC-2022 Task5 Track1: Hierarchical Multi-label Classification via Label-Aware Graph Convolutional Network
Bo Wang 0134, Yi-Fan Lu, Xiaochi Wei, Xiao Liu 0029, Ge Shi 0002, Changsen Yuan, Heyan Huang, Chong Feng 0001, Xianling Mao
NLPCC (2)4
2022 News-driven stock prediction via noisy equity state representation
Heyan Huang, Xiao Liu 0029, Yue Zhang 0004, Chong Feng 0001
Neurocomputing2
2022 End-to-end event factuality prediction using directional labeled graph recurrent network
Xiao Liu 0029, Heyan Huang, Yue Zhang 0004
Inf. Process. Manag.1
2021 BIT-Event at NLPCC-2021 Task 3: Subevent Identification via Adversarial Training
Xiao Liu 0029, Ge Shi 0002, Bo Wang 0134, Changsen Yuan, Heyan Huang, Chong Feng 0001, Lifang Wu
NLPCC (2)1
2020 Dialogue State Induction Using Neural Latent Variable Models
abstract
Dialogue state modules are a useful component in a task-oriented dialogue system. Traditional methods find dialogue states by manually labeling training corpora, upon which neural models are trained. However, the labeling process can be costly, slow, error-prone, and more importantly, cannot cover the vast range of domains in real-world dialogues for customer service. We propose the task of dialogue state induction, building two neural latent variable models that mine dialogue states automatically from unlabeled customer service dialogue records. Results show that the models can effectively find meaningful dialogue states. In addition, equipped with induced dialogue states, a state-of-the-art dialogue system gives better performance compared with not using a dialogue state module.
Qingkai Min, Libo Qin 0001, Zhiyang Teng, Xiao Liu 0029, Yue Zhang 0004
IJCAI4
2019 Distant Supervision for Relation Extraction with Linear Attenuation Simulation and Non-IID Relevance Embedding
abstract
Distant supervision for relation extraction is an efficient method to reduce labor costs and has been widely used to seek novel relational facts in large corpora, which can be identified as a multi-instance multi-label problem. However, existing distant supervision methods suffer from selecting important words in the sentence and extracting valid sentences in the bag. Towards this end, we propose a novel approach to address these problems in this paper. Firstly, we propose a linear attenuation simulation to reflect the importance of words in the sentence with respect to the distances between entities and words. Secondly, we propose a non-independent and identically distributed (non-IID) relevance embedding to capture the relevance of sentences in the bag. Our method can not only capture complex information of words about hidden relations, but also express the mutual information of instances in the bag. Extensive experiments on a benchmark dataset have well-validated the effectiveness of the proposed method.
Changsen Yuan, Heyan Huang, Chong Feng 0001, Xiao Liu 0029, Xiaochi Wei
AAAI4
2019 Open Domain Event Extraction Using Neural Latent Variable Models
abstract
We consider open domain event extraction, the task of extracting unconstraint types of events from news clusters.A novel latent variable neural model is constructed, which is scalable to very large corpus.A dataset is collected and manually annotated, with task-specific evaluation metrics being designed.Results show that the proposed unsupervised model gives better performance compared to the state-of-the-art method for event schema induction.
Xiao Liu 0029, Heyan Huang, Yue Zhang 0004
ACL (1)1
2019 Selective Expression For Event Coreference Resolution on Twitter
abstract
With the growth in popularity and size of social media, there is an urgent need for systems that can recognize the coreference relation between two event mentions in texts from social media. In existing event coreference resolution research, a rich set of linguistic features derived from pre-existing NLP tools and various knowledge bases is often required. This kind of methods restricts domain scalability and leads to the propagation of errors. In this paper, we present a novel selective expression approach based on event trigger to explore the coreferential relationship in high-volume Twitter texts. Firstly, we exploit a bidirectional Long Short Term Memory (Bi-LSTM) to extract the sentence level and mention level features. Then, to selectively express the essential parts of generated features, we apply a gate on sentence level features. Next, to integrate the time information of event mention pairs, we design an auxiliary feature based on triggers and time attributes of the two event mentions. Finally, all these features are concatenated and fed into a classifier to predict the binary coreference relationship between the event mention pair. To evaluate our method, we publish a new dataset EventCoreOnTweet (ECT)1that annotates the coreferential relationship between event mentions and event trigger of each event mention. The experimental results demonstrate that our approach achieves significant performance in the ECT dataset.
Wen-Han Chao, Zhunchen Luo, Xiao Liu 0029, Guobin Sui
IJCNN4
2018 Jointly Multiple Events Extraction via Attention-based Graph Information Aggregation
abstract
Event extraction is of practical utility in natural language processing.In the real world, it is a common phenomenon that multiple events existing in the same sentence, where extracting them are more difficult than extracting a single event.Previous works on modeling the associations between events by sequential modeling methods suffer a lot from the low efficiency in capturing very long-range dependencies.In this paper, we propose a novel Jointly Multiple Events Extraction (JMEE) framework to jointly extract multiple event triggers and arguments by introducing syntactic shortcut arcs to enhance information flow and attention-based graph convolution networks to model graph information.The experiment results demonstrate that our proposed framework achieves competitive results compared with state-of-the-art methods.
Xiao Liu 0029, Zhunchen Luo, Heyan Huang
EMNLP1