Chaojun Xiao

dblp:223/4856 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-6039-0942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
abstract
Yuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001
ACL (1)4
2025 APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
abstract
While long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing long-context queries in a timely manner. To address this, we introduce APB, an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed by reducing compute and enhancing parallelism simultaneously. APB introduces a communication mechanism for essential key-value pairs within a sequence parallelism framework, enabling a faster inference speed while maintaining task performance. We implement APB by incorporating a tailored FlashAttn kernel alongside optimized distribution strategies, supporting diverse models and parallelism configurations. APB achieves speedups of up to 9.2\times, 4.2\times, and 1.6\times compared with FlashAttn, RingAttn, and StarAttn, respectively, without any observable task performance degradation.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Sun Ao, Hao Zhou 0012, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)4
2025 Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
abstract
Activation sparsity denotes the existence of substantial weakly-contributed neurons within feed-forward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1$\times$ speedup compared with its dense version. The codes and checkpoints are available at https://github.com/thunlp/SparsingLaw/.
Yuqi Luo, Xu Han 0007, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu 0001, Maosong Sun 0001
ICML5
2025 Thoroughly Modeling Multi-domain Pre-trained Recommendation as Language
abstract
With the thriving of the pre-trained language model (PLM) widely verified in various NLP tasks, pioneer efforts attempt to explore the possible cooperation of the general textual information in PLM with the personalized behavioral information in user historical behavior sequences to enhance sequential recommendation (SR). However, despite the commonalities of input format and task goal, there are huge gaps between the behavioral and textual information, which obstruct thoroughly modeling SR as language modeling via PLM. To bridge the gap, we propose a novel unified pre-trained language model enhanced sequential recommendation (UPSR) that thoroughly transfers the next item prediction task to a text generation task, aiming to build a unified pre-trained recommendation model for multi-domain recommendation tasks. We formally design five key indicators, namely naturalness, domain consistency, informativeness, noise and ambiguity, and text length, to guide the text \(\rightarrow\) item adaptation (selecting appropriate text to form the item textual representation) and behavior sequence \(\rightarrow\) text sequence adaptation (transferring the sequence of item textual representations into a text sequence) differently for pre-training and fine-tuning stages, which are essential but under-explored by previous works. In experiments, we conduct extensive evaluations on seven datasets with both supervised and zero-shot settings and achieve the overall best performance. Comprehensive model analyses also provide valuable insights for behavior modeling via PLM, shedding light on large pre-trained recommendation models. The source codes will be released in the future.
Zekai Qu, Ruobing Xie, Chaojun Xiao, Yuan Yao 0013, Zhiyuan Liu 0001, Fengzong Lian, Zhanhui Kang, Jie Zhou 0016
ACM Trans. Inf. Syst.3
2024 Robust and Scalable Model Editing for Large Language Models
abstract
Large language models (LLMs) can make predictions using parametric knowledge – knowledge encoded in the model weights – or contextual knowledge – knowledge presented in the context. In many scenarios, a desirable behavior is that LLMs give precedence to contextual knowledge when it conflicts with the parametric knowledge, and fall back to using their parametric knowledge when the context is irrelevant. This enables updating and correcting the model’s knowledge by in-context editing instead of retraining. Previous works have shown that LLMs are inclined to ignore contextual knowledge and fail to reliably fall back to parametric knowledge when presented with irrelevant context. In this work, we discover that, with proper prompting methods, instruction-finetuned LLMs can be highly controllable by contextual knowledge and robust to irrelevant context. Utilizing this feature, we propose EREN (Edit models by REading Notes) to improve the scalability and robustness of LLM editing. To better evaluate the robustness of model editors, we collect a new dataset, that contains irrelevant questions that are more challenging than the ones in existing datasets. Empirical results show that our method outperforms current state-of-the-art methods by a large margin. Unlike existing techniques, it can integrate knowledge from multiple edits, and correctly respond to syntactically similar but semantically unrelated inputs (and vice versa). The source code can be found at https://github.com/thunlp/EREN.
Yingfa Chen, Zhengyan Zhang, Xu Han 0007, Chaojun Xiao, Zhiyuan Liu 0001, Kuai Li, Maosong Sun 0001
LREC/COLING4
2024 Fine-Grained Legal Argument-Pair Extraction via Coarse-Grained Pre-training
abstract
Legal Argument-Pair Extraction (LAE) is dedicated to the identification of interactive arguments targeting the same subject matter within legal complaints and corresponding defenses. This process serves as a foundation for automatically recognizing the focal points of disputes. Current methodologies predominantly conceptualize LAE as a supervised sentence-pair classification problem and usually necessitate extensive manual annotations, thereby constraining their scalability and general applicability. To this end, we present an innovative approach to LAE that focuses on fine-grained alignment of argument pairs, building upon coarse-grained complaint-defense pairs. This strategy stems from two key observations: 1) In general, every argument presented in a legal complaint is likely to be addressed by at least one corresponding argument in the defense. 2) It’s rare for multiple complaint arguments to be addressed by a single defense argument; rather, each complaint argument usually corresponds to a unique defense argument. Motivated by these insights, we develop a specialized pre-training framework. Our model employs pre-training objectives designed to exploit the coarse-grained supervision signals. This enables expressive representations of legal arguments for LAE, even when working with a limited amount of labeled data. To verify the effectiveness of our model, we construct the largest LAE datasets from two representative causes, private lending, and contract dispute. The experimental results demonstrate that our model can effectively capture informative argument knowledge from unlabeled complaint-defense pairs and outperform the unsupervised and supervised baselines by 3.7 and 2.4 points on average respectively. Besides, our model can reach superior accuracy with only half manually annotated data. The datasets and code can be found in https://github.com/thunlp/LAE.
Chaojun Xiao, Yutao Sun, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001
LREC/COLING1
2024 Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
abstract
Legal case retrieval (LCR) aims to provide similar cases as references for a given fact description.This task is crucial for promoting consistent judgments in similar cases, effectively enhancing judicial fairness and improving work efficiency for judges.However, existing works face two main challenges for real-world applications: existing works mainly focus on case-to-case retrieval using lengthy queries, which does not match real-world scenarios; and the limited data scale, with current datasets containing only hundreds of queries, is insufficient to satisfy the training requirements of existing data-hungry neural models.To address these issues, we introduce an automated method to construct synthetic query-candidate pairs and build the largest LCR dataset to date, LEAD, which is hundreds of times larger than existing datasets.This data construction method can provide ample training signals for LCR models.Experimental results demonstrate that model training with our constructed data can achieve state-of-the-art results on two widely-used LCR benchmarks.Besides, the construction method can also be applied to civil cases and achieve promising results.The data and codes can be found in https://github.com/thunlp/LEAD.
Chaojun Xiao, Zhenghao Liu 0001, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP2
2024 Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
abstract
Weilin Zhao, Yuxiang Huang, Xu Han, Wang Xu, Chaojun Xiao, Xinrong Zhang, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu, Maosong Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Weilin Zhao, Yuxiang Huang 0001, Xu Han 0007, Chaojun Xiao, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP5
2024 Exploring the Benefit of Activation Sparsity in Pre-training
abstract
Pre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activation has been explored through post-training methods, its potential in pre-training remains untapped. In this work, we first study how activation properties change during pre-training. Our examination reveals that Transformers exhibit sparse activation throughout the majority of the pre-training process while the activation correlation keeps evolving as training progresses. Leveraging this observation, we propose Switchable Sparse-Dense Learning (SSD). SSD adaptively switches between the Mixtures-of-Experts (MoE) based sparse training and the conventional dense training during the pre-training process, leveraging the efficiency of sparse training and avoiding the static activation correlation of sparse training. Compared to dense training, SSD achieves comparable performance with identical model size and reduces pre-training costs. Moreover, the models trained with SSD can be directly used as MoE models for sparse inference and achieve the same performance as dense models with up to $2\times$ faster inference speed. Codes are available at https://github.com/thunlp/moefication.
Zhengyan Zhang, Chaojun Xiao, Qiujieli Qin, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Ruobing Xie, Maosong Sun 0001, Jie Zhou 0016
ICML2
2024 InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
abstract
Large language models (LLMs) have emerged as a cornerstone in real-world applications with lengthy streaming inputs (e.g., LLM-driven agents). However, existing LLMs, pre-trained on sequences with a restricted maximum length, cannot process longer sequences due to the out-of-domain and distraction issues. Common solutions often involve continual pre-training on longer sequences, which will introduce expensive computational overhead and uncontrollable change in model capabilities. In this paper, we unveil the intrinsic capacity of LLMs for understanding extremely long sequences without any fine-tuning. To this end, we introduce a training-free memory-based method, InfLLM. Specifically, InfLLM stores distant contexts into additional memory units and employs an efficient mechanism to lookup token-relevant units for attention computation. Thereby, InfLLM allows LLMs to efficiently process long sequences with a limited context window and well capture long-distance dependencies. Without any training, InfLLM enables LLMs that are pre-trained on sequences consisting of a few thousand tokens to achieve comparable performance with competitive baselines that continually train these LLMs on long sequences. Even when the sequence length is scaled to 1,024K, InfLLM still effectively captures long-distance dependencies. Our code can be found at https://github.com/thunlp/InfLLM.
Chaojun Xiao, Pengle Zhang, Xu Han 0007, Guangxuan Xiao, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS1
2024 The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential Recommendation
abstract
Sequential recommendation (SR) has seen significant advancements with the help of Pre-trained Language Models (PLMs). Some PLM-based SR models directly use PLM to encode user historical behavior’s text sequences to learn user representations, while there is seldom an in-depth exploration of the capability and suitability of PLM in behavior sequence modeling. In this work, we first conduct extensive model analyses between PLMs and PLM-based SR models, discovering great underutilization and parameter redundancy of PLMs in behavior sequence modeling. Inspired by this, we explore different lightweight usages of PLMs in SR, aiming to maximally stimulate the ability of PLMs for SR while satisfying the efficiency and usability demands of practical systems. We discover that adopting behavior-tuned PLMs for item initializations of conventional ID-based SR models is the most economical framework of PLM-based SR, which would not bring in any additional inference cost but could achieve a dramatic performance boost compared with the original version. Extensive experiments on five datasets show that our simple and universal framework leads to significant improvement compared to classical SR and SOTA PLM-based SR models without additional inference costs. Our code can be found in https://github.com/777pomingzi/Rethinking-PLM-in-RS.
Zekai Qu, Ruobing Xie, Chaojun Xiao, Zhanhui Kang, Xingwu Sun
RecSys3
2023 Plug-and-Play Document Modules for Pre-trained Models
abstract
Chaojun Xiao, Zhengyan Zhang, Xu Han, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Xiangyang Li, Zhonghua Li, Zhao Cao, Maosong Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chaojun Xiao, Zhengyan Zhang, Xu Han 0007, Chi-Min Chan, Yankai Lin 0001, Zhiyuan Liu 0001, Zhao Cao, Maosong Sun 0001
ACL (1)1
2023 Plug-and-Play Knowledge Injection for Pre-trained Language Models
abstract
Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang, Deming Ye, Chaojun Xiao, Xu Han, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhengyan Zhang, Yankai Lin 0001, Deming Ye, Chaojun Xiao, Xu Han 0007, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
ACL (1)6
2023 MUSER: A Multi-View Similar Case Retrieval Dataset
abstract
Similar case retrieval (SCR) is a representative legal AI application that plays a pivotal role in promoting judicial fairness. However, existing SCR datasets only focus on the fact description section when judging the similarity between cases, ignoring other valuable sections (e.g., the court's opinion) that can provide insightful reasoning process behind. Furthermore, the case similarities are typically measured solely by the textual semantics of the fact descriptions, which may fail to capture the full complexity of legal cases from the perspective of legal knowledge. In this work, we present MUSER, a similar case retrieval dataset based on multi-view similarity measurement and comprehensive legal element with sentence-level legal element annotations. Specifically, we select three perspectives (legal fact, dispute focus, and law statutory) and build a comprehensive and structured label schema of legal elements for each of them, to enable accurate and knowledgeable evaluation of case similarities. The constructed dataset originates from Chinese civil cases and contains 100 query cases and 4,024 candidate cases. We implement several text classification algorithms for legal element prediction and various retrieval methods for retrieving similar cases on MUSER. The experimental results indicate that incorporating legal elements can benefit the performance of SCR models, but further efforts are still required to address the remaining challenges posed by MUSER. The source code and dataset are released at https://github.com/THUlawtech/MUSER.
Qingquan Li 0003, Yiran Hu, Chaojun Xiao, Zhiyuan Liu 0001, Maosong Sun 0001, Weixing Shen
CIKM4
2021 Adversarial Language Games for Advanced Natural Language Intelligence
abstract
We study the problem of adversarial language games, in which multiple agents with conflicting goals compete with each other via natural language interactions. While adversarial language games are ubiquitous in human activities, little attention has been devoted to this field in natural language processing. In this work, we propose a challenging adversarial language game called Adversarial Taboo as an example, in which an attacker and a defender compete around a target word. The attacker is tasked with inducing the defender to utter the target word invisible to the defender, while the defender is tasked with detecting the target word before being induced by the attacker. In Adversarial Taboo, a successful attacker and defender need to hide or infer the intention, and induce or defend during conversations. This requires several advanced language abilities, such as adversarial pragmatic reasoning and goal-oriented language interactions in open domain, which will facilitate many downstream NLP tasks. To instantiate the game, we create a game environment and a competition platform. Comprehensive experiments on several baseline attack and defense strategies show promising and interesting results, based on which we discuss some directions for future research.
Yuan Yao 0013, Haoxi Zhong, Zhengyan Zhang, Xu Han 0007, Xiaozhi Wang, Kai Zhang 0033, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001
AAAI7
2020 JEC-QA: A Legal-Domain Question Answering Dataset
abstract
We present JEC-QA, the largest question answering dataset in the legal domain, collected from the National Judicial Examination of China. The examination is a comprehensive evaluation of professional skills for legal practitioners. College students are required to pass the examination to be certified as a lawyer or a judge. The dataset is challenging for existing question answering methods, because both retrieving relevant materials and answering questions require the ability of logic reasoning. Due to the high demand of multiple reasoning abilities to answer legal questions, the state-of-the-art models can only achieve about 28% accuracy on JEC-QA, while skilled humans and unskilled humans can reach 81% and 64% accuracy respectively, which indicates a huge gap between humans and machines on this task. We will release JEC-QA and our baselines to help improve the reasoning ability of machine comprehension models. You can access the dataset from http://jecqa.thunlp.org/.
Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 0001, Zhiyuan Liu 0001, Maosong Sun 0001
AAAI2
2020 How Does NLP Benefit Legal System: A Summary of Legal Artificial Intelligence
abstract
Legal Artificial Intelligence (LegalAI) focuses on applying the technology of artificial intelligence, especially natural language processing, to benefit tasks in the legal domain.In recent years, LegalAI has drawn increasing attention rapidly from both AI researchers and legal professionals, as LegalAI is beneficial to the legal system for liberating legal professionals from a maze of paperwork.Legal professionals often think about how to solve tasks from rulebased and symbol-based methods, while NLP researchers concentrate more on data-driven and embedding methods.In this paper, we describe the history, the current state, and the future directions of research in LegalAI.We illustrate the tasks from the perspectives of legal professionals and NLP researchers and show several representative applications in LegalAI.We conduct experiments and provide an indepth analysis of the advantages and disadvantages of existing works to explore possible future directions.You can find the implementation of our work from https://github. com/thunlp/CLAIM.
Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ACL2
2020 Denoising Relation Extraction from Document-level Distant Supervision
abstract
Distant supervision (DS) has been widely used to generate auto-labeled data for sentencelevel relation extraction (RE), which improves RE performance.However, the existing success of DS cannot be directly transferred to the more challenging document-level relation extraction (DocRE), since the inherent noise in DS may be even multiplied in document level and significantly harm the performance of RE.To address this challenge, we propose a novel pre-trained model for DocRE, which denoises the document-level DS data via multiple pre-training tasks.Experimental results on the large-scale DocRE benchmark show that our model can capture useful information from noisy DS data and achieve promising results.The source code of this paper can be found in https://github.com/thunlp/DSDocRE.
Chaojun Xiao, Yuan Yao 0013, Ruobing Xie, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Fen Lin 0002, Leyu Lin
EMNLP (1)1
2018 Legal Judgment Prediction via Topological Learning
abstract
Legal Judgment Prediction (LJP) aims to predict the judgment result based on the facts of a case and becomes a promising application of artificial intelligence techniques in the legal field.In real-world scenarios, legal judgment usually consists of multiple subtasks, such as the decisions of applicable law articles, charges, fines, and the term of penalty.Moreover, there exist topological dependencies among these subtasks.While most existing works only focus on a specific subtask of judgment prediction and ignore the dependencies among subtasks, we formalize the dependencies among subtasks as a Directed Acyclic Graph (DAG) and propose a topological multi-task learning framework, TOP-JUDGE, which incorporates multiple subtasks and DAG dependencies into judgment prediction.We conduct experiments on several realworld large-scale datasets of criminal cases in the civil law system.Experimental results show that our model achieves consistent and significant improvements over baselines on all judgment prediction tasks.The source code can be obtained from https://github. com/thunlp/TopJudge.
Haoxi Zhong, Zhipeng Guo 0001, Cunchao Tu, Chaojun Xiao, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP4