Xu Han 0007

dblp:19/3011-7 · DBLP profile ↗
← Back
71ranked-venue papers
5as first author
49since 2021 · last 2026
0000-0002-4726-7621ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 68 · 5 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
abstract
Large Language Models (LLMs) with Mixture-of-Experts (MoE) architectures are distinguished by their strong performance scaling with increasing parameters across a wide range of tasks, yet they also suffer from substantial computational and storage overheads. Notably, the performance gains of MoE models do not scale proportionally with the growth in expert parameters. While prior works attempt to reduce parameters via expert-level pruning, merging, or decomposition, they still suffer from challenges in both performance and computational efficiency. In this paper, we address these challenges by introducing micro-expert as a finer-grained compression unit that spans across matrices. We first establish a more fundamental perspective, viewing MoE layers as mixtures of micro-experts, and present CAMERA, a lightweight and training-free framework for identifying micro-expert redundancy. Our analysis uncovers significant variance in micro-expert contributions during decoding. Based on this insight, we further propose CAMERA-P, a structured micro-expert pruning framework, and CAMERA-Q, a mixed-precision quantization idea designed for micro-experts. Extensive experiments on nine downstream tasks show that CAMERA-P consistently outperforms strong baselines under pruning ratios ranging from 20% to 60%. Furthermore, CAMERA-Q achieves superior results under aggressive 2-bit quantization, surpassing existing matrix- and channel-level ideas. Notably, our method enables complete micro-expert analysis of Qwen2-57B-A14B in less than 5 minutes on a single NVIDIA A100-40GB GPU.
Yuzhuang Xu, Xu Han 0007, Yuanchi Zhang, Shiyu Ji, Qingfu Zhu, Wanxiang Che
AAAI2
2026 APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
abstract
Yuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001
ACL (1)3
2026 AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
abstract
Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanle Zhao, Zilin Sang, Qi Shi 0002, Wei-Lun Zhao, Shuo Wang 0013, Duzhen Zhang, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)8
2026 From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
abstract
Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.
Tao Wen 0011, Shuai Shao 0015, Pei Ke, Xu Han 0007, Jie Zou 0001, Tao Tian, Jinjie Qiu, Ke Qin
SIGIR4
2025 APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
abstract
While long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing long-context queries in a timely manner. To address this, we introduce APB, an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed by reducing compute and enhancing parallelism simultaneously. APB introduces a communication mechanism for essential key-value pairs within a sequence parallelism framework, enabling a faster inference speed while maintaining task performance. We implement APB by incorporating a tailored FlashAttn kernel alongside optimized distribution strategies, supporting diverse models and parallelism configurations. APB achieves speedups of up to 9.2\times, 4.2\times, and 1.6\times compared with FlashAttn, RingAttn, and StarAttn, respectively, without any observable task performance degradation.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Sun Ao, Hao Zhou 0012, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)3
2025 FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
abstract
Weilin Zhao, Tengyu Pan, Xu Han, Yudi Zhang, Sun Ao, Yuxiang Huang, Kaihuo Zhang, Weilun Zhao, Yuxuan Li, Jie Zhou, Hao Zhou, Jianyong Wang, Maosong Sun, Zhiyuan Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weilin Zhao, Tengyu Pan, Xu Han 0007, Sun Ao, Yuxiang Huang 0001, Kaihuo Zhang, Wei-Lun Zhao, Jie Zhou 0016, Hao Zhou 0012, Jianyong Wang 0001, Maosong Sun 0001, Zhiyuan Liu 0001
ACL (1)3
2025 LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models
abstract
Zihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Qi Shi, Zhixing Tan, Xu Han, Xiaodong Shi, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuo Wang 0013, Yu Chao, Zhili Li, Qi Shi 0002, Zhixing Tan, Xu Han 0007, Xiaodong Shi, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)10
2025 RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
abstract
Kunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Ruobing Wang, Shuo Wang, Yishan Li, Nan Zhang, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kunlun Zhu, Dingling Xu, Yukun Yan, Zhenghao Liu 0001, Shi Yu 0001, Shuo Wang 0013, Yishan Li, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)11
2025 Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slips
abstract
This study presents a multi-modal multi-granularity tokenizer specifically designed for analyzing ancient Chinese scripts, focusing on the Chu bamboo slip (CBS) script used during the Spring and Autumn and Warring States period (771-256 BCE) in Ancient China. Considering the complex hierarchical structure of ancient Chinese scripts, where a single character may be a combination of multiple sub-characters, our tokenizer first adopts character detection to locate character boundaries. Then it conducts character recognition at both the character and sub-character levels. Moreover, to support the academic community, we assembled the first large-scale dataset of CBSs with over 100K annotated character image scans. On the part-of-speech tagging task built on our dataset, using our tokenizer gives a 5.5% relative improvement in F1-score compared to mainstream sub-word tokenizers. Our work not only aids in further investigations of the specific script but also has the potential to advance research on other forms of ancient Chinese scripts.
Yingfa Chen, Chenlong Hu, Shi Yu 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
COLING6
2025 ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
abstract
Activation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs, serving as a promising paradigm for accelerating model inference. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g., GELU and Swish). Some recent efforts have explored introducing ReLU or its variants as the substitutive activation function to pursue activation sparsity and acceleration, but few can simultaneously obtain high activation sparsity and comparable model performance. This paper introduces a simple and effective method named “ProSparse” to sparsify LLMs while achieving both targets. Specifically, after introducing ReLU activation, ProSparse adopts progressive sparsity regularization with a factor smoothly increasing for multiple stages. This can enhance activation sparsity and mitigate performance degradation by avoiding radical shifts in activation distributions. With ProSparse, we obtain high sparsity of 89.32% for LLaMA2-7B, 88.80% for LLaMA2-13B, and 87.89% for end-size MiniCPM-1B, respectively, with comparable performance to their original Swish-activated versions. These present the most sparsely activated models among open-source LLaMA versions and competitive end-size models. Inference acceleration experiments further demonstrate the significant practical acceleration potential of LLMs with higher activation sparsity, obtaining up to 4.52x inference speedup.
Xu Han 0007, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Zhiyuan Liu 0001, Guangli Li, Maosong Sun 0001
COLING2
2025 Cost-Optimal Grouped-Query Attention for Long-Context Modeling
abstract
Grouped-Query Attention (GQA) is a widely adopted strategy for reducing the computational cost of attention layers in large language models (LLMs).However, current GQA configurations are often suboptimal because they overlook how context length influences inference cost.Since inference cost grows with context length, the most cost-efficient GQA configuration should vary accordingly.In this work, we analyze the relationship among context length, model size, GQA configuration, and model loss, and introduce two innovations:(1) we decouple the total head size from the hidden size, enabling more flexible control over attention FLOPs; and (2) we jointly optimize the model size and the GQA configuration to arrive at a better allocation of inference resources between attention layers and other components.Our analysis reveals that commonly used GQA configurations are highly suboptimal for longcontext scenarios.Moreover, we propose a recipe for deriving cost-optimal GQA configurations.Our results show that for long-context scenarios, one should use fewer attention heads while scaling up the model size.Configurations selected by our recipe can reduce both memory usage and FLOPs by more than 50% compared to Llama-3's GQA, with no degradation in model capabilities.Our findings offer valuable insights for designing efficient longcontext LLMs. 1 Memory (GB)-57.8% -50.8%
Yingfa Chen, Zhen Leng Thai, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP6
2025 VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
abstract
Retrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG systems are solely based on text, rendering it impossible to utilize vision information like layout and images that play crucial roles in real-world multi-modality documents. In this paper, we introduce VisRAG, which tackles this issue by establishing a vision-language model (VLM)-based RAG pipeline. In this pipeline, instead of first parsing the document to obtain text, the document is directly embedded using a VLM as an image and then retrieved to enhance the generation of a VLM. Compared to traditional text-based RAG, VisRAG maximizes the retention and utilization of the data information in the original documents, eliminating the information loss introduced during the parsing process. We collect both open-source and synthetic data to train the retriever in VisRAG and explore a variety of generation methods. Experiments demonstrate that VisRAG outperforms traditional RAG in both the retrieval and generation stages, achieving a 20–40% end-to-end performance gain over traditional text-based RAG pipeline. Further analysis reveals that VisRAG is efficient in utilizing training data and demonstrates strong generalization capability, positioning it as a promising solution for RAG on multi-modality documents. Our code and data are available at https://github.com/openbmb/visrag.
Shi Yu 0001, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu 0001, Shuo Wang 0013, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR9
2025 Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
abstract
Activation sparsity denotes the existence of substantial weakly-contributed neurons within feed-forward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1$\times$ speedup compared with its dense version. The codes and checkpoints are available at https://github.com/thunlp/SparsingLaw/.
Yuqi Luo, Xu Han 0007, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu 0001, Maosong Sun 0001
ICML3
2025 EditEval: Towards Comprehensive and Automatic Evaluation for Text-guided Video Editing
abstract
Recently, video editing task has gained widespread attention due to its practical applications and rapid advancements. However, current automatic evaluation metrics for video editing are mostly poorly aligned with human judgments. Thus, researchers heavily rely on human evaluation, which is not only labor-intensive but also difficult to ensure consistency and objectivity. To address these issues, we propose EditEval, the largest-ever video editing benchmark to comprehensively evaluate the performance of video editing models in three aspects: Textual Faithfulness, Frame Consistency, and Video Fidelity. It includes 200 video clips and 1,010 text prompts, from which 160 instances are sampled to generate 1,280 edited videos using eight open-source video editing models, accompanied by human annotations. Furthermore, we propose EditScore, leveraging the advanced reasoning and comprehension capabilities of Multi-modal Large Language Models (MLLMs) as evaluators to assess edited videos across the aforementioned aspects. Experiments show that the best-performing video editing model only reaches an average score of 3.16 (out of a perfect 5), highlighting the challenge of EditEval. Besides, results from more than 10 MLLMs demonstrate the great potential of utilizing EditScore for automatic evaluation. Notably, for textual faithfulness, EditScore equipped with LLaVA-OneVision-7B achieves a significantly higher Pearson Correlation score compared to previous methods based on CLIP (0.50 vs 0.22). The code and dataset are available at: https://github.com/XMUDeepLIT/EditEval
Bingshuai Liu, Ante Wang, Zijun Min, Chenyang Lyu, Longyue Wang, Xu Han 0007, Peng Li 0030, Jinsong Su
ACM Multimedia7
2025 Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
abstract
Sun Ao, Weilin Zhao, Xu Han, Cheng Yang, Xinrong Zhang, Zhiyuan Liu, Chuan Shi, Maosong Sun. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sun Ao, Weilin Zhao, Xu Han 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Chuan Shi 0001, Maosong Sun 0001
NAACL (Long Papers)3
2025 BurstEngine: An efficient distributed framework for training transformers On extremely Long sequences of over 1M tokens
abstract
Existing methods for training LLMs on long-sequence data, such as Tensor Parallelism and Context Parallelism, exhibit low Model FLOPs Utilization as sequence lengths and number of GPUs increase, especially when sequence lengths exceed 1M tokens. To address these challenges, we propose BurstEngine, an efficient framework designed to train LLMs on long-sequence data. BurstEngine introduces BurstAttention, an optimized distributed attention with lower communication cost than RingAttention. BurstAttention leverages topology-aware ring communication to fully utilize network bandwidth and incorporates fine-grained communication-computation overlap. Furthermore, BurstEngine introduces sequence-level selective checkpointing and fuses the language modeling head with the loss function to reduce memory cost. Additionally, BurstEngine introduces workload balance optimization for various types of attention masking. By integrating these optimizations, BurstEngine achieves a 1.2 × speedup with much lower memory overhead than the state-of-the-art baselines when training LLMs on extremely long sequences of over 1M tokens.
Weilin Zhao, Xu Han 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Chuan Shi 0001, Maosong Sun 0001
SC3
2024 OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
abstract
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Jinyi Hu, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)8
2024 FastFiD: Improve Inference Efficiency of Open Domain Question Answering via Sentence Selection
abstract
Open Domain Question Answering (ODQA) has been advancing rapidly in recent times, driven by significant developments in dense passage retrieval and pretrained language models.Current models typically incorporate the FiD framework, which is composed by a neural retriever alongside an encoder-decoder neural reader.In the answer generation process, the retriever will retrieve numerous passages (around 100 for instance), each of which is then individually encoded by the encoder.Subsequently, the decoder makes predictions based on these encoded passages.Nevertheless, this framework can be relatively time-consuming, particularly due to the extensive length of the gathered passages.To address this, we introduce FastFiD in this paper, a novel approach that executes sentence selection on the encoded passages.This aids in retaining valuable sentences while reducing the context length required for generating answers.Experiments on three commonly used datasets (Natural Questions, TriviaQA and ASQA) demonstrate that our method can enhance the inference speed by 2.3X-5.7X,while simultaneously maintaining the model's performance.Moreover, an in-depth analysis of the model's attention reveals that the selected sentences indeed hold a substantial contribution towards the final answer.The codes are publicly available at https://github.com
Yufei Huang 0008, Xu Han 0007, Maosong Sun 0001
ACL (1)2
2024 MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation
abstract
Xiaozhi Wang, Hao Peng, Yong Guan, Kaisheng Zeng, Jianhui Chen, Lei Hou, Xu Han, Yankai Lin, Zhiyuan Liu, Ruobing Xie, Jie Zhou, Juanzi Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaozhi Wang, Hao Peng 0015, Kaisheng Zeng, Lei Hou 0001, Xu Han 0007, Yankai Lin 0001, Zhiyuan Liu 0001, Ruobing Xie, Jie Zhou 0016, Juan-Zi Li
ACL (1)7
2024 LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks
abstract
LoRA employs lightweight modules to customize large language models (LLMs) for each downstream task or domain, where different learned additional modules represent diverse skills.Combining existing LoRA modules to address new tasks can enhance the reusability of learned LoRA modules, particularly beneficial for tasks with limited annotated data.Most prior works on LoRA combination primarily rely on task-level weights for each involved LoRA, making different examples and tokens share the same LoRA weights.However, in generative tasks, different tokens may necessitate diverse skills to manage.Taking the Chinese math task as an example, understanding the problem description may depend more on the Chinese LoRA, while the calculation part may rely more on the math LoRA.To this end, we propose LoRA-Flow, which utilizes dynamic weights to adjust the impact of different LoRA modules.The weights at each step are determined by a fusion gate with extremely few parameters, which can be learned with only 200 training examples.Experiments across six generative tasks demonstrate that our method consistently outperforms baselines with tasklevel fusion weights.This underscores the necessity of introducing dynamic fusion weights for LoRA combination. 1
Hanqing Wang 0003, Bowen Ping, Shuo Wang 0013, Xu Han 0007, Yun Chen 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)4
2024 UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
abstract
Haoyu Wang, Shuo Wang, Yukun Yan, Xujia Wang, Zhiyu Yang, Yuzhuang Xu, Zhenghao Liu, Liner Yang, Ning Ding, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shuo Wang 0013, Yukun Yan, Xujia Wang, Zhiyu Yang 0001, Yuzhuang Xu, Zhenghao Liu 0001, Liner Yang, Ning Ding 0002, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)10
2024 ınftyBench: Extending Long Context Evaluation Beyond 100K Tokens
abstract
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yingfa Chen, Shengding Hu, Zihang Xu, Moo Khai Hao, Xu Han 0007, Zhen Leng Thai, Shuo Wang 0013, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)7
2024 Robust and Scalable Model Editing for Large Language Models
abstract
Large language models (LLMs) can make predictions using parametric knowledge – knowledge encoded in the model weights – or contextual knowledge – knowledge presented in the context. In many scenarios, a desirable behavior is that LLMs give precedence to contextual knowledge when it conflicts with the parametric knowledge, and fall back to using their parametric knowledge when the context is irrelevant. This enables updating and correcting the model’s knowledge by in-context editing instead of retraining. Previous works have shown that LLMs are inclined to ignore contextual knowledge and fail to reliably fall back to parametric knowledge when presented with irrelevant context. In this work, we discover that, with proper prompting methods, instruction-finetuned LLMs can be highly controllable by contextual knowledge and robust to irrelevant context. Utilizing this feature, we propose EREN (Edit models by REading Notes) to improve the scalability and robustness of LLM editing. To better evaluate the robustness of model editors, we collect a new dataset, that contains irrelevant questions that are more challenging than the ones in existing datasets. Empirical results show that our method outperforms current state-of-the-art methods by a large margin. Unlike existing techniques, it can integrate knowledge from multiple edits, and correctly respond to syntactically similar but semantically unrelated inputs (and vice versa). The source code can be found at https://github.com/thunlp/EREN.
Yingfa Chen, Zhengyan Zhang, Xu Han 0007, Chaojun Xiao, Zhiyuan Liu 0001, Kuai Li, Maosong Sun 0001
LREC/COLING3
2024 Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
abstract
Xinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun, Zhiyuan Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yingfa Chen, Shengding Hu, Xu Han 0007, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun 0001, Zhiyuan Liu 0001
EMNLP4
2024 Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
abstract
Weilin Zhao, Yuxiang Huang, Xu Han, Wang Xu, Chaojun Xiao, Xinrong Zhang, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu, Maosong Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Weilin Zhao, Yuxiang Huang 0001, Xu Han 0007, Chaojun Xiao, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP3
2024 Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
abstract
Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind. Building a competitive counterpart in other languages is highly challenging due to the low-resource nature of non-English multimodal data (i.e., lack of large-scale, high-quality image-text data). In this work, we propose MPM, an effective training paradigm for training large multimodal models in low-resource languages. MPM demonstrates that Multilingual language models can Pivot zero-shot Multimodal learning across languages. Specifically, based on a strong multilingual large language model, multimodal models pretrained on English-only image-text data can well generalize to other languages in a (quasi)-zero-shot manner, even surpassing models trained on image-text data in native languages. Taking Chinese as a practice of MPM, we build large multimodal models VisCPM in image-to-text and text-to-image generation, which achieve state-of-the-art (open-source) performance in Chinese. To facilitate future research, we open-source codes and model weights at https://github.com/OpenBMB/VisCPM.
Jinyi Hu, Yuan Yao 0013, Chongyi Wang, Shan Wang 0015, Yinxu Pan, Tianyu Yu 0002, Hanghao Wu, Haoye Zhang, Xu Han 0007, Yankai Lin 0001, Jiao Xue, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR11
2024 Predicting Emergent Abilities with Infinite Resolution Evaluation
abstract
The scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ''emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy with theoretically infinite resolution, through massive sampling in the decoding phase. With PassUntil, we conduct a quantitative investigation into the scaling law of task performance. The investigation contains two parts. Firstly, a strict task scaling law that is not conventionally known to exist, is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts, which is the first systematic attempt to verify predictable scaling proposed by GPT-4's report. Secondly, underpinned by PassUntil, we are able to study emergent abilities quantitatively. We identify a kind of accelerated emergence whose scaling curve cannot be fitted by standard scaling law function and has a increasing speed. We then examine two hypothesis and imply that the ``multiple circuits hypothesis'' might be responsible for the accelerated emergence.
Shengding Hu, Xin Liu 0086, Xu Han 0007, Chaoqun He, Weilin Zhao, Yankai Lin 0001, Ning Ding 0002, Zebin Ou, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR3
2024 Exploring the Benefit of Activation Sparsity in Pre-training
abstract
Pre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activation has been explored through post-training methods, its potential in pre-training remains untapped. In this work, we first study how activation properties change during pre-training. Our examination reveals that Transformers exhibit sparse activation throughout the majority of the pre-training process while the activation correlation keeps evolving as training progresses. Leveraging this observation, we propose Switchable Sparse-Dense Learning (SSD). SSD adaptively switches between the Mixtures-of-Experts (MoE) based sparse training and the conventional dense training during the pre-training process, leveraging the efficiency of sparse training and avoiding the static activation correlation of sparse training. Compared to dense training, SSD achieves comparable performance with identical model size and reduces pre-training costs. Moreover, the models trained with SSD can be directly used as MoE models for sparse inference and achieve the same performance as dense models with up to $2\times$ faster inference speed. Codes are available at https://github.com/thunlp/moefication.
Zhengyan Zhang, Chaojun Xiao, Qiujieli Qin, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Ruobing Xie, Maosong Sun 0001, Jie Zhou 0016
ICML6
2024 Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
abstract
Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability.
Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS4
2024 InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
abstract
Large language models (LLMs) have emerged as a cornerstone in real-world applications with lengthy streaming inputs (e.g., LLM-driven agents). However, existing LLMs, pre-trained on sequences with a restricted maximum length, cannot process longer sequences due to the out-of-domain and distraction issues. Common solutions often involve continual pre-training on longer sequences, which will introduce expensive computational overhead and uncontrollable change in model capabilities. In this paper, we unveil the intrinsic capacity of LLMs for understanding extremely long sequences without any fine-tuning. To this end, we introduce a training-free memory-based method, InfLLM. Specifically, InfLLM stores distant contexts into additional memory units and employs an efficient mechanism to lookup token-relevant units for attention computation. Thereby, InfLLM allows LLMs to efficiently process long sequences with a limited context window and well capture long-distance dependencies. Without any training, InfLLM enables LLMs that are pre-trained on sequences consisting of a few thousand tokens to achieve comparable performance with competitive baselines that continually train these LLMs on long sequences. Even when the sequence length is scaled to 1,024K, InfLLM still effectively captures long-distance dependencies. Our code can be found at https://github.com/thunlp/InfLLM.
Chaojun Xiao, Pengle Zhang, Xu Han 0007, Guangxuan Xiao, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS3
2024 OneBit: Towards Extremely Low-bit Large Language Models
abstract
Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs. However, current quantization methods suffer severe performance degradation when the bit-width is extremely reduced, and thus focus on utilizing 4-bit or 8-bit values to quantize models. This paper boldly quantizes the weight matrices of LLMs to 1-bit, paving the way for the extremely low bit-width deployment of LLMs. For this target, we introduce a 1-bit model compressing framework named OneBit, including a novel 1-bit parameter representation method to better quantize LLMs as well as an effective parameter initialization method based on matrix decomposition to improve the convergence speed of the quantization framework. Sufficient experimental results indicate that OneBit achieves good performance (at least 81% of the non-quantized performance on LLaMA models) with robust training processes when only using 1-bit weight matrices.
Yuzhuang Xu, Xu Han 0007, Zonghan Yang, Shuo Wang 0013, Qingfu Zhu, Zhiyuan Liu 0001, Wanxiang Che
NeurIPS2
2024 Hyperbolic Pre-Trained Language Model
abstract
In recent years, we have witnessed significant improvements in pre-trained language models (PLM) brought about by the scaling of parameter sizes and data amounts. However, this also brings high computational and storage costs. In this paper, we present a new direction to improve PLMs without scaling parameters and data: adopting a geometric feature space that is more suitable for encoding the intrinsic structured features of text. Although text is generally considered unstructured data, it possesses rich intrinsic structured features that signify syntactic and semantic relationships. Leveraging these structured features is vital for text understanding. Given that structured features are better encoded in hyperbolic spaces than in the Euclidean spaces used by conventional PLMs, we propose that PLMs should operate entirely within hyperbolic spaces. Our experiments demonstrate the superiority of hyperbolic PLMs over Euclidean PLMs across a wide variety of tasks, using the same parameter and data settings. This suggests that altering the geometry of model representation is a promising direction for model enhancement. The code is released athttps://github.com/thunlp/hyperbolic_llm
Weize Chen, Xu Han 0007, Yankai Lin 0001, Kaichen He, Ruobing Xie, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 WebCPM: Interactive Web Search for Chinese Long-form Question Answering
abstract
Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yujia Qin, Zihan Cai, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin 0001, Xu Han 0007, Ning Ding 0002, Ruobing Xie, Fanchao Qi, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0024
ACL (1)8
2023 Plug-and-Play Document Modules for Pre-trained Models
abstract
Chaojun Xiao, Zhengyan Zhang, Xu Han, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Xiangyang Li, Zhonghua Li, Zhao Cao, Maosong Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chaojun Xiao, Zhengyan Zhang, Xu Han 0007, Chi-Min Chan, Yankai Lin 0001, Zhiyuan Liu 0001, Zhao Cao, Maosong Sun 0001
ACL (1)3
2023 Plug-and-Play Knowledge Injection for Pre-trained Language Models
abstract
Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang, Deming Ye, Chaojun Xiao, Xu Han, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhengyan Zhang, Yankai Lin 0001, Deming Ye, Chaojun Xiao, Xu Han 0007, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
ACL (1)7
2023 H3T: Efficient Integration of Memory Optimization and Parallelism for Large-scale Transformer Training
abstract
In recent years, big models based on Transformers have achieved state-of-the-art performance on many artificial intelligence (AI) tasks. Despite the success of these Transformer-based models, their huge parameter size poses a serious challenge to their training, both from the storage and computation perspectives. To this end, memory optimization (e.g., rematerialization and offloading) and parallelism (e.g., data parallelism and model parallelism) are widely explored to make training Transformers more efficient. In this paper, we propose a framework to automatically find an efficient integration of memory optimization and parallelism for High-Throughput Transformer Training (named H3T), which is rarely considered by existing efforts for training big Transformer-based models. Specifically, we design search algorithms to combine appropriate memory optimization strategies and parallelism schemes to achieve a balance between memory overhead and training efficiency. We implement H3T based on an open-source toolkit BMTrain and then use H3T to train the Transformers of different sizes to evaluate the efficiency of H3T. The experimental results show that H3T outperforms the most popular deep learning (DL) toolkit Megatron-DeepSpeed by $1.2\times \sim 4.3\times$ training speed while reducing $34.6\% \sim 80.5\%$ of memory overhead. Moreover, H3T can use only 64 NVIDIA A100 GPUs to train GPT-3-175B, which is very difficult for existing DL toolkits. The source code is available at https://github.com/OpenBMB/BMTrain/tree/h3t.
Xu Han 0007, Weilin Zhao, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS2
2022 Fully Hyperbolic Neural Networks
abstract
Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Weize Chen, Xu Han 0007, Yankai Lin 0001, Hexu Zhao, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
ACL (1)2
2022 PPT: Pre-trained Prompt Tuning for Few-shot Learning
abstract
Prompts for pre-trained language models (PLMs) have shown remarkable performance by bridging the gap between pre-training tasks and various downstream tasks.Among these methods, prompt tuning, which freezes PLMs and only tunes soft prompts, provides an efficient and effective solution for adapting largescale PLMs to downstream tasks.However, prompt tuning is yet to be fully explored.In our pilot experiments, we find that prompt tuning performs comparably with conventional full-model tuning when downstream data are sufficient, whereas it is much worse under fewshot learning settings, which may hinder the application of prompt tuning.We attribute this low performance to the manner of initializing soft prompts.Therefore, in this work, we propose to pre-train prompts by adding soft prompts into the pre-training stage to obtain a better initialization.We name this Pretrained Prompt Tuning framework "PPT".To ensure the generalization of PPT, we formulate similar classification tasks into a unified task form and pre-train soft prompts for this unified task.Extensive experiments show that tuning pre-trained prompts for downstream tasks can reach or even outperform full-model fine-tuning under both full-data and few-shot settings.Our approach is effective and efficient for using large-scale PLMs in practice.The code is publicly available at https:// github.com/thu-coai/PPT.
Yuxian Gu, Xu Han 0007, Zhiyuan Liu 0001, Minlie Huang
ACL (1)2
2022 Cross-Lingual Contrastive Learning for Fine-Grained Entity Typing for Low-Resource Languages
abstract
Xu Han, Yuqi Luo, Weize Chen, Zhiyuan Liu, Maosong Sun, Zhou Botong, Hao Fei, Suncong Zheng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Xu Han 0007, Yuqi Luo, Weize Chen, Zhiyuan Liu 0001, Maosong Sun 0001, Botong Zhou, Suncong Zheng
ACL (1)1
2022 Exploring Mode Connectivity for Pre-trained Language Models
abstract
Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP.From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found.Although plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, little is known about the connection of various minima reached under different adaptation configurations.In this paper, we investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path.We conduct empirical analyses to investigate three questions: (1) how could hyperparameters, specific tuning methods, and training data affect PLM's mode connectivity?(2) How does mode connectivity change during pretraining?(3) How does the PLM's task knowledge change along the path connecting two minima?In general, exploring the mode connectivity of PLMs conduces to understanding the geometric connection of different minima, which may help us fathom the inner workings of PLM downstream adaptation.The codes are publicly available at https://github.com/ thunlp/Mode-Connectivity-PLM.
Yujia Qin, Cheng Qian 0008, Jing Yi, Weize Chen, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
EMNLP6
2022 MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction
abstract
Xiaozhi Wang, Yulin Chen, Ning Ding, Hao Peng, Zimu Wang, Yankai Lin, Xu Han, Lei Hou, Juanzi Li, Zhiyuan Liu, Peng Li, Jie Zhou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xiaozhi Wang, Yulin Chen 0001, Ning Ding 0002, Hao Peng 0015, Yankai Lin 0001, Xu Han 0007, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016
EMNLP7
2022 GACT: Activation Compressed Training for Generic Network Architectures
abstract
Training large neural network (NN) models requires extensive memory resources, and Activation Compression Training (ACT) is a promising approach to reduce training memory footprint. This paper presents GACT, an ACT framework to support a broad range of machine learning tasks for generic NN architectures with limited domain knowledge. By analyzing a linearized version of ACT’s approximate gradient, we prove the convergence of GACT without prior knowledge on operator type or model architecture. To make training stable, we propose an algorithm that decides the compression ratio for each tensor by estimating its impact on the gradient at run time. We implement GACT as a PyTorch library that readily applies to any NN architecture. GACT reduces the activation memory for convolutional NNs, transformers, and graph NNs by up to 8.1x, enabling training with a 4.2x to 24.7x larger batch size, with negligible accuracy loss.
Lianmin Zheng, Dequan Wang, Yukuo Cen, Weize Chen, Xu Han 0007, Jianfei Chen 0001, Zhiyuan Liu 0001, Jie Tang 0001, Joey Gonzalez, Michael W. Mahoney, Alvin Cheung
ICML6
2022 Knowledge Inheritance for Pre-trained Language Models
abstract
Yujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yujia Qin, Yankai Lin 0001, Jing Yi, Xu Han 0007, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
NAACL-HLT5
2021 Adversarial Language Games for Advanced Natural Language Intelligence
abstract
We study the problem of adversarial language games, in which multiple agents with conflicting goals compete with each other via natural language interactions. While adversarial language games are ubiquitous in human activities, little attention has been devoted to this field in natural language processing. In this work, we propose a challenging adversarial language game called Adversarial Taboo as an example, in which an attacker and a defender compete around a target word. The attacker is tasked with inducing the defender to utter the target word invisible to the defender, while the defender is tasked with detecting the target word before being induced by the attacker. In Adversarial Taboo, a successful attacker and defender need to hide or infer the intention, and induce or defend during conversations. This requires several advanced language abilities, such as adversarial pragmatic reasoning and goal-oriented language interactions in open domain, which will facilitate many downstream NLP tasks. To instantiate the game, we create a game environment and a competition platform. Comprehensive experiments on several baseline attack and defense strategies show promising and interesting results, based on which we discuss some directions for future research.
Yuan Yao 0013, Haoxi Zhong, Zhengyan Zhang, Xu Han 0007, Xiaozhi Wang, Kai Zhang 0033, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001
AAAI4
2021 Few-NERD: A Few-shot Named Entity Recognition Dataset
abstract
Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, Zhiyuan Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ning Ding 0002, Yulin Chen 0001, Xiaobin Wang, Xu Han 0007, Pengjun Xie, Hai-Tao Zheng 0002, Zhiyuan Liu 0001
ACL/IJCNLP (1)5
2021 CLEVE: Contrastive Pre-training for Event Extraction
abstract
Ziqi Wang, Xiaozhi Wang, Xu Han, Yankai Lin, Lei Hou, Zhiyuan Liu, Peng Li, Juanzi Li, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ziqi Wang 0003, Xiaozhi Wang, Xu Han 0007, Yankai Lin 0001, Lei Hou 0001, Zhiyuan Liu 0001, Peng Li 0030, Juan-Zi Li, Jie Zhou 0016
ACL/IJCNLP (1)3
2021 Visual Distant Supervision for Scene Graph Generation
abstract
Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive human annotation. In this work, we propose visual distant supervision, a novel paradigm of visual relation learning, which can train scene graph models without any human-labeled data. The intuition is that by aligning commonsense knowledge bases and images, we can automatically create large-scale labeled data to provide distant supervision for visual relation learning. To alleviate the noise in distantly labeled data, we further propose a framework that iteratively estimates the probabilistic relation labels and eliminates the noisy ones. Comprehensive experimental results show that our distantly supervised model outperforms strong weakly supervised and semi-supervised baselines. By further incorporating human-labeled data in a semi-supervised fashion, our model outperforms state-of-the-art fully supervised models by a large margin (e.g., 8.3 micro- and 7.8 macro-recall@50 improvements for predicate classification in Visual Genome evaluation). We make the data and code for this paper publicly available at https://github.com/thunlp/VisualDS.
Yuan Yao 0011, Xu Han 0007, Mengdi Li 0006, Cornelius Weber, Zhiyuan Liu 0001, Stefan Wermter, Maosong Sun 0001
ICCV3
2021 Open Hierarchical Relation Extraction
abstract
Kai Zhang, Yuan Yao, Ruobing Xie, Xu Han, Zhiyuan Liu, Fen Lin, Leyu Lin, Maosong Sun. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Kai Zhang 0033, Yuan Yao 0013, Ruobing Xie, Xu Han 0007, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001
NAACL-HLT4
2021 CSS-LM: A Contrastive Framework for Semi-Supervised Fine-Tuning of Pre-Trained Language Models
abstract
Fine-tuning pre-trained language models (PLMs) has demonstrated its effectiveness on various downstream NLP tasks recently. However, in many scenarios with limited supervised data, the conventional fine-tuning strategies cannot sufficiently capture the important semantic features for downstream tasks. To address this issue, we introduce a novel framework (named ‘`CSS-LM’') to improve the fine-tuning phase of PLMs via contrastive semi-supervised learning. Specifically, given a specific task, we retrieve positive and negative instances from large-scale unlabeled corpora according to their domain-level and class-level semantic relatedness to the task. We then perform contrastive semi-supervised learning on both the retrieved unlabeled instances and original labeled instances to help PLMs capture crucial task-related semantic features. The experimental results show that CSS-LM achieves better results than the conventional fine-tuning strategy on a series of downstream tasks with few-shot settings by up to 7.8%, and outperforms the latest supervised contrastive fine-tuning strategy by up to 7.1%. Our datasets and source code will be available to provide more details.
Yusheng Su, Xu Han 0007, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016, Maosong Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Neural Snowball for Few-Shot Relation Learning
abstract
Knowledge graphs typically undergo open-ended growth of new relations. This cannot be well handled by relation extraction that focuses on pre-defined relations with sufficient training data. To address new relations with few-shot instances, we propose a novel bootstrapping approach, Neural Snowball, to learn new relations by transferring semantic knowledge about existing relations. More specifically, we use Relational Siamese Networks (RSN) to learn the metric of relational similarities between instances based on existing relations and their labeled data. Afterwards, given a new relation and its few-shot instances, we use RSN to accumulate reliable instances from unlabeled corpora; these instances are used to train a relation classifier, which can further identify new facts of the new relation. The process is conducted iteratively like a snowball. Experiments show that our model can gather high-quality instances for better few-shot relation learning and achieves significant improvement compared to baselines. Codes and datasets are released on https://github.com/thunlp/Neural-Snowball.
Tianyu Gao 0001, Xu Han 0007, Ruobing Xie, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001
AAAI2
2020 Continual Relation Learning via Episodic Memory Activation and Reconsolidation
abstract
Continual relation learning aims to continually train a model on new data to learn incessantly emerging novel relations while avoiding catastrophically forgetting old relations.Some pioneering work has proved that storing a handful of historical relation examples in episodic memory and replaying them in subsequent training is an effective solution for such a challenging problem.However, these memorybased methods usually suffer from overfitting the few memorized examples of old relations, which may gradually cause inevitable confusion among existing relations.Inspired by the mechanism in human long-term memory formation, we introduce episodic memory activation and reconsolidation (EMAR) to continual relation learning.Every time neural models are activated to learn both new and memorized data, EMAR utilizes relation prototypes for memory reconsolidation exercise to keep a stable understanding of old relations.The experimental results show that EMAR could get rid of catastrophically forgetting old relations and outperform the state-of-the-art continual learning models.The code and datasets are released on https://github.com/thunlp/ ContinualRE.
Xu Han 0007, Tianyu Gao 0001, Yankai Lin 0001, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
ACL1
2020 Meta-Information Guided Meta-Learning for Few-Shot Relation Classification
abstract
Few-shot classification requires classifiers to adapt to new classes with only a few training instances.State-of-the-art meta-learning approaches such as MAML learn how to initialize and fast adapt parameters from limited instances, which have shown promising results in few-shot classification.However, existing meta-learning models solely rely on implicit instance-based statistics, and thus suffer from instance unreliability and weak interpretability.To solve this problem, we propose a novel meta-information guided meta-learning (MIML) framework, where semantic concepts of classes provide strong guidance for meta-learning in both initialization and adaptation.In effect, our model can establish connections between instance-based information and semantic-based information, which enables more effective initialization and faster adaptation.Comprehensive experimental results on few-shot relation classification demonstrate the effectiveness of the proposed framework.Notably, MIML achieves comparable or superior performance to humans with only one shot on FewRel evaluation.The source code and experiment details of this paper can be obtained from https://github.com/thunlp/MIML.
Bowen Dong 0005, Yuan Yao 0013, Ruobing Xie, Tianyu Gao 0001, Xu Han 0007, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001
COLING5
2020 Dynamic Anticipation and Completion for Multi-Hop Reasoning over Sparse Knowledge Graph
abstract
Multi-hop reasoning has been widely studied in recent years to seek an effective and interpretable method for knowledge graph (KG) completion.Most previous reasoning methods are designed for dense KGs with enough paths between entities, but cannot work well on those sparse KGs that only contain sparse paths for reasoning.On the one hand, sparse KGs contain less information, which makes it difficult for the model to choose correct paths.On the other hand, the lack of evidential paths to target entities also makes the reasoning process difficult.To solve these problems, we propose a multi-hop reasoning model named DacKGR over sparse KGs, by applying novel dynamic anticipation and completion strategies: (1) The anticipation strategy utilizes the latent prediction of embeddingbased models to make our model perform more potential path search over sparse KGs.(2) Based on the anticipation information, the completion strategy dynamically adds edges as additional actions during the path search, which further alleviates the sparseness problem of KGs.The experimental results on five datasets sampled from Freebase, NELL and Wikidata show that our method outperforms state-of-the-art baselines.Our codes and datasets can be obtained from https:// github.com/THU-KEG/DacKGR.
Xu Han 0007, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Wei Zhang 0127, Yichi Zhang 0010, Suhui Wu
EMNLP (1)2
2020 Learning from Context or Names? An Empirical Study on Neural Relation Extraction
abstract
Neural models have achieved remarkable success on relation extraction (RE) benchmarks.However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance of these models.To this end, we empirically study the effect of two main information sources in text: textual context and entity mentions (names).We find that (i) while context is the main source to support the predictions, RE models also heavily rely on the information from entity mentions, most of which is type information, and (ii) existing datasets may leak shallow heuristics via entity mentions and thus contribute to the high performance on RE benchmarks.Based on the analyses, we propose an entity-masked contrastive pre-training framework for RE to gain a deeper understanding on both textual context and type information while avoiding rote memorization of entities or use of superficial cues in mentions.We carry out extensive experiments to support our views, and show that our framework can improve the effectiveness and robustness of neural models in different RE scenarios.All the code and datasets are released at https://github.com/thunlp/
Hao Peng 0015, Tianyu Gao 0001, Xu Han 0007, Yankai Lin 0001, Peng Li 0030, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
EMNLP (1)3
2020 MAVEN: A Massive General Domain Event Detection Dataset
abstract
Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Wangyi Jiang, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Yankai Lin 0001, Jie Zhou 0016
EMNLP (1)3
2020 Denoising Relation Extraction from Document-level Distant Supervision
abstract
Distant supervision (DS) has been widely used to generate auto-labeled data for sentencelevel relation extraction (RE), which improves RE performance.However, the existing success of DS cannot be directly transferred to the more challenging document-level relation extraction (DocRE), since the inherent noise in DS may be even multiplied in document level and significantly harm the performance of RE.To address this challenge, we propose a novel pre-trained model for DocRE, which denoises the document-level DS data via multiple pre-training tasks.Experimental results on the large-scale DocRE benchmark show that our model can capture useful information from noisy DS data and achieve promising results.The source code of this paper can be found in https://github.com/thunlp/DSDocRE.
Chaojun Xiao, Yuan Yao 0013, Ruobing Xie, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Fen Lin 0002, Leyu Lin
EMNLP (1)4
2019 Hybrid Attention-Based Prototypical Networks for Noisy Few-Shot Relation Classification
abstract
The existing methods for relation classification (RC) primarily rely on distant supervision (DS) because large-scale supervised training datasets are not readily available. Although DS automatically annotates adequate amounts of data for model training, the coverage of this data is still quite limited, and meanwhile many long-tail relations still suffer from data sparsity. Intuitively, people can grasp new knowledge by learning few instances. We thus provide a different view on RC by formalizing RC as a few-shot learning (FSL) problem. However, the current FSL models mainly focus on low-noise vision tasks, which makes them hard to directly deal with the diversity and noise of text. In this paper, we propose hybrid attention-based prototypical networks for the problem of noisy few-shot RC. We design instancelevel and feature-level attention schemes based on prototypical networks to highlight the crucial instances and features respectively, which significantly enhances the performance and robustness of RC models in a noisy FSL scenario. Besides, our attention schemes accelerate the convergence speed of RC models. Experimental results demonstrate that our hybrid attention-based models require fewer training iterations and outperform the state-of-the-art baseline models. The code and datasets are released on https://github.com/thunlp/ HATT-Proto.
Tianyu Gao 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
AAAI2
2019 Quantifying Similarity between Relations with Fact Distribution
abstract
We introduce a conceptually simple and effective method to quantify the similarity between relations in knowledge bases.Specifically, our approach is based on the divergence between the conditional probability distributions over entity pairs.In this paper, these distributions are parameterized by a very simple neural network.Although computing the exact similarity is intractable, we provide a sampling-based method to get a good approximation.We empirically show the outputs of our approach significantly correlate with human judgments.By applying our method to various tasks, we also find that (1) our approach could effectively detect redundant relations extracted by open information extraction (Open IE) models, that (2) even the most competitive models for relational classification still make mistakes among very similar relations, and that (3) our approach could be incorporated into negative sampling and softmax classification to alleviate these mistakes.The source code and experiment details of this paper can be obtained from https://github.com/ thunlp/relation-similarity.
Weize Chen, Hao Zhu 0006, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)3
2019 DocRED: A Large-Scale Document-Level Relation Extraction Dataset
abstract
Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yuan Yao 0013, Deming Ye, Peng Li 0030, Xu Han 0007, Yankai Lin 0001, Zhenghao Liu 0001, Zhiyuan Liu 0001, Lixin Huang, Jie Zhou 0016, Maosong Sun 0001
ACL (1)4
2019 ERNIE: Enhanced Language Representation with Informative Entities
abstract
Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks.However, the existing pre-trained language models rarely consider incorporating knowledge graphs (KGs), which can provide rich structured knowledge facts for better language understanding.We argue that informative entities in KGs can enhance language representation with external knowledge.In this paper, we utilize both large-scale textual corpora and KGs to train an enhanced language representation model (ERNIE), which can take full advantage of lexical, syntactic, and knowledge information simultaneously.The experimental results have demonstrated that ERNIE achieves significant improvements on various knowledge-driven tasks, and meanwhile is comparable with the state-of-the-art model BERT on other common NLP tasks.The source code and experiment details of this paper can be obtained from https:// github.com/thunlp/ERNIE.
Zhengyan Zhang, Xu Han 0007, Zhiyuan Liu 0001, Xin Jiang 0002, Maosong Sun 0001, Qun Liu 0001
ACL (1)2
2019 DIAG-NRE: A Neural Pattern Diagnosis Framework for Distantly Supervised Neural Relation Extraction
abstract
Pattern-based labeling methods have achieved promising results in alleviating the inevitable labeling noises of distantly supervised neural relation extraction.However, these methods require significant expert labor to write relation-specific patterns, which makes them too sophisticated to generalize quickly.To ease the labor-intensive workload of pattern writing and enable the quick generalization to new relation types, we propose a neural pattern diagnosis framework, DIAG-NRE, that can automatically summarize and refine highquality relational patterns from noise data with human experts in the loop.To demonstrate the effectiveness of DIAG-NRE, we apply it to two real-world datasets and present both significant and interpretable improvements over state-of-the-art methods.
Shun Zheng 0001, Xu Han 0007, Yankai Lin 0001, Ling Huang 0001, Zhiyuan Liu 0001, Wei Xu 0005
ACL (1)2
2019 GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification
abstract
Fact verification (FV) is a challenging task which requires to retrieve relevant evidence from plain text and use the evidence to verify given claims.Many claims require to simultaneously integrate and reason over several pieces of evidence for verification.However, previous work employs simple models to extract information from evidence without letting evidence communicate with each other, e.g., merely concatenate the evidence for processing.Therefore, these methods are unable to grasp sufficient relational and logical information among the evidence.To alleviate this issue, we propose a graph-based evidence aggregating and reasoning (GEAR) framework which enables information to transfer on a fully-connected evidence graph and then utilizes different aggregators to collect multievidence information.We further employ BERT, an effective pre-trained language representation model, to improve the performance.Experimental results on a large-scale benchmark dataset FEVER have demonstrated that GEAR could leverage multi-evidence information for FV and thus achieves the promising result with a test FEVER score of 67.10%.
Jie Zhou 0024, Xu Han 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)2
2019 FewRel 2.0: Towards More Challenging Few-Shot Relation Classification
abstract
Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tianyu Gao 0001, Xu Han 0007, Hao Zhu 0006, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
EMNLP/IJCNLP (1)2
2019 Adapting Meta Knowledge Graph Information for Multi-Hop Reasoning over Few-Shot Relations
abstract
Xin Lv, Yuxian Gu, Xu Han, Lei Hou, Juanzi Li, Zhiyuan Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yuxian Gu, Xu Han 0007, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001
EMNLP/IJCNLP (1)3
2019 HMEAE: Hierarchical Modular Event Argument Extraction
abstract
Xiaozhi Wang, Ziqi Wang, Xu Han, Zhiyuan Liu, Juanzi Li, Peng Li, Maosong Sun, Jie Zhou, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016, Xiang Ren 0001
EMNLP/IJCNLP (1)3
2019 Open Relation Extraction: Relational Knowledge Transfer from Supervised Data to Unsupervised Data
abstract
Ruidong Wu, Yuan Yao, Xu Han, Ruobing Xie, Zhiyuan Liu, Fen Lin, Leyu Lin, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ruidong Wu, Yuan Yao 0013, Xu Han 0007, Ruobing Xie, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001
EMNLP/IJCNLP (1)3
2018 Neural Knowledge Acquisition via Mutual Attention Between Knowledge Graph and Text
abstract
We propose a general joint representation learning framework for knowledge acquisition (KA) on two tasks, knowledge graph completion (KGC) and relation extraction (RE) from text. In this framework, we learn representations of knowledge graphs (KGs) and text within a unified parameter sharing semantic space. To achieve better fusion, we propose an effective mutual attention between KGs and text. The reciprocal attention mechanism enables us to highlight important features and perform better KGC and RE. Different from conventional joint models, no complicated linguistic analysis or strict alignments between KGs and text are required to train our models. Experiments on relation extraction and entity link prediction show that models trained under our joint framework are significantly improved in comparison with other baselines. Most existing methods for KGC and RE can be easily integrated into our framework due to its flexible architectures. The source code of this paper can be obtained from https://github.com/thunlp/JointNRE.
Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
AAAI1
2018 Adversarial Multi-lingual Neural Relation Extraction
abstract
Multi-lingual relation extraction aims to find unknown relational facts from text in various languages. Existing models cannot well capture the consistency and diversity of relation patterns in different languages. To address these issues, we propose an adversarial multi-lingual neural relation extraction (AMNRE) model, which builds both consistent and individual representations for each sentence to consider the consistency and diversity among languages. Further, we adopt an adversarial training strategy to ensure those consistent sentence representations could effectively extract the language-consistent relation patterns. The experimental results on real-world datasets demonstrate that our AMNRE model significantly outperforms the state-of-the-art models. The source code of this paper can be obtained from https://github.com/thunlp/AMNRE.
Xiaozhi Wang, Xu Han 0007, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
COLING2
2018 Hierarchical Relation Extraction with Coarse-to-Fine Grained Attention
abstract
Distantly supervised relation extraction employs existing knowledge graphs to automatically collect training data.While distant supervision is effective to scale relation extraction up to large-scale corpora, it inevitably suffers from the wrong labeling problem.Many efforts have been devoted to identifying valid instances from noisy data.However, most existing methods handle each relation in isolation, regardless of rich semantic correlations located in relation hierarchies.In this paper, we aim to incorporate the hierarchical information of relations for distantly supervised relation extraction and propose a novel hierarchical attention scheme.The multiple layers of our hierarchical attention scheme provide coarseto-fine granularity to better identify valid instances, which is especially effective for extracting those long-tail relations.The experimental results on a large-scale benchmark dataset demonstrate that our models are capable of modeling the hierarchical information of relations and significantly outperform other baselines.The source code of this paper can be obtained from https://github.com/ thunlp/HNRE.
Xu Han 0007, Pengfei Yu 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Peng Li 0030
EMNLP1
2018 FewRel: A Large-Scale Supervised Few-shot Relation Classification Dataset with State-of-the-Art Evaluation
abstract
We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers.The relation of each sentence is first recognized by distant supervision methods, and then filtered by crowdworkers.We adapt the most recent state-of-the-art few-shot learning methods for relation classification and conduct thorough evaluation of these methods.Empirical results show that even the most competitive few-shot learning models struggle on this task, especially as compared with humans.We also show that a range of different reasoning skills are needed to solve our task.These results indicate that few-shot relation classification remains an open problem and still requires further research.Our detailed analysis points multiple directions for future research.All details and resources about the dataset and baselines are released on http://zhuhao.me/ fewrel.
Xu Han 0007, Hao Zhu 0006, Pengfei Yu 0001, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP1
2018 Put It Back: Entity Typing with Language Model Enhancement
abstract
Entity typing aims to classify semantic types of an entity mention in a specific context.Most existing models obtain training data using distant supervision, and inevitably suffer from the problem of noisy labels.To address this issue, we propose entity typing with language model enhancement.It utilizes a language model to measure the compatibility between context sentences and labels, and thereby automatically focuses more on context-dependent labels.Experiments on benchmark datasets demonstrate that our method is capable of enhancing the entity typing model with information from the language model, and significantly outperforms the stateof-the-art baseline.
Ji Xin, Hao Zhu 0006, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP3