EDBT 2026 Demo / reviewers in the wild / expert
Runxin Xu
dblp:267/5291
· DBLP profile ↗
20ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-3876-2284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3
Liang Chen 0024, Haozhe Zhao, Yinzhen Huang, Tsekai Lin, Weichu Xie, Peiyi Wang, Runxin Xu, Ming Wu 0007, Baobao Chang |
CIKM | 9 |
| 2025 | Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsabstractRecent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities.
However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning. Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang |
ICLR | 10 |
| 2025 | CodeIO: Condensing Reasoning Patterns via Code Input-Output PredictionabstractReasoning is a fundamental capability of Large Language Models. While prior research predominantly focuses on enhancing narrow skills like math or code generation, improving performance on many other reasoning tasks remains challenging due to sparse and fragmented training data. To address this issue, we propose CodeI/O, a novel approach that systematically condenses diverse reasoning patterns inherently embedded in contextually-grounded codes, through transforming the original code into a code input-output prediction format. By training models to predict inputs/outputs given code and test cases entirely in natural language as Chain-of-Thought (CoT) rationales, we expose them to universal reasoning primitives—like logic flow planning, state-space searching, decision tree traversal, and modular decomposition—while decoupling structured reasoning from code-specific syntax and preserving procedural rigor. Experimental results demonstrate CodeI/O leads to consistent improvements across symbolic, scientific, logic, math & numerical, and commonsense reasoning tasks. By matching the existing ground-truth outputs or re-executing the code with predicted inputs, we can verify each prediction and further enhance the CoTs through multi-turn revision, resulting in CodeI/O++ and achieving higher performance. Our data and models will be publicly available. Daya Guo, Dejian Yang, Runxin Xu, Yu Wu 0024, Junxian He |
ICML | 4 |
| 2024 | Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes.However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains.To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension.ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains.Drawing from ArXivCap, we introduce ArXivQA, a questionanswering dataset generated by prompting GPT-4V based on scientific figures.ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4% absolute accuracy gain on a multimodal mathematical reasoning benchmark.Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs.Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domainspecific training yields substantial performance gains.Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements. Lei Li 0039, Yuqi Wang 0003, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu 0049 |
ACL (1) | 3 |
| 2024 | Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsabstractPeiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiyi Wang, Lei Li 0039, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li 0005, Deli Chen, Zhifang Sui |
ACL (1) | 4 |
| 2024 | Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsabstractParameter-efficient fine-tuning (PEFT) is crucial for customizing Large Language Models (LLMs) with constrained resources.Although there have been various PEFT methods for dense-architecture LLMs, PEFT for sparsearchitecture LLMs is still underexplored.In this work, we study the PEFT method for LLMs with the Mixture-of-Experts (MoE) architecture and the contents of this work are mainly threefold: (1) We investigate the dispersion degree of the activated experts in customized tasks, and found that the routing distribution for a specific task tends to be highly concentrated, while the distribution of activated experts varies significantly across different tasks.(2) We propose Expert-Specialized Fine-Tuning, or ESFT, which tunes the experts most relevant to downstream tasks while freezing the other experts and modules; experimental results demonstrate that our method not only improves the tuning efficiency, but also matches or even surpasses the performance of fullparameter fine-tuning.(3) We further analyze the impact of the MoE architecture on expertspecialized fine-tuning.We find that MoE models with finer-grained experts are more advantageous in selecting the combination of experts that are most relevant to downstream tasks, thereby enhancing both the training efficiency and effectiveness.Our code is available at https://github.com/deepseek-ai/ESFT. Zihan Wang 0010, Deli Chen, Damai Dai, Runxin Xu, Zhuoshu Li |
EMNLP | 4 |
| 2023 | An Iteratively Parallel Generation Method with the Pre-Filling Strategy for Document-level Event ExtractionabstractIn document-level event extraction (DEE) tasks, a document typically contains many event records with multiple event roles.Therefore, accurately extracting all event records is a big challenge since the number of event records is not given.Previous works present the entitybased directed acyclic graph (EDAG) generation methods to autoregressively generate event roles, which requires a given generation order.Meanwhile, parallel methods are proposed to generate all event roles simultaneously, but suffer from the inadequate training which manifests zero accuracies on some event roles.In this paper, we propose an Iteratively Parallel Generation method with the Pre-Filling strategy (IPGPF).Event roles in an event record are generated in parallel to avoid order selection, and the event records are iteratively generated to utilize historical results.Experiments on two public datasets show our IPGPF improves 11.7 F1 than previous parallel models and up to 5.1 F1 than auto-regressive models under the control variable settings.Moreover, our enhanced IPGPF outperforms other entityenhanced models and achieves new state-ofthe-art performance 1 .* Work was done when Guanhua was an intern at ByteDance AI Lab.† Corresponding author. 1 Our code is available at https://github.com/ CarlanLark/IPGPF [S6] …, Jinggong Group increased its holdings of the company's stock by 182,038 shares through the secondary market on Dec 15, 2011,… [S7] …, the shares held by Jinggong Group in the company increased from 90,880,020 shares to 91,062,058 shares, … [S9] on Dec 16, 2011, Jinggong Group reduced its holdings of ... 35,000 shares, with an average price of 19.88.[S14] As of the date of this announcement, Jinggong Group holds 91,027,058 shares of the company, … EquityOverweight EquityHolder Jinggong Group Guanhua Huang, Runxin Xu, Jiaze Chen, Zhouwang Yang, Weinan E |
EMNLP | 2 |
| 2023 | TABLEIE: Capturing the Interactions Among Sub-Tasks in Information Extraction via Double TablesabstractInformation Extraction mainly consists of three sub-tasks, Named Entity Recognition, Relation Extraction and Event Extraction. Although these sub-tasks are highly correlated with each other, most previous works simply focus on part of them and ignore the interactions among different sub-tasks. Recently, some graph-based models are proposed to cover all the interactions among different IE sub-tasks. However, the use of Graph Neural Network brings heavy computation burden, damaging the model efficiency. In this paper, we propose a double-table framework, TableIE, to capture the interactions among IE sub-tasks as well as improve the model efficiency. Specifically, TableIE has an entity-relation table and an event table, based on which we propose both within-table and cross-table interaction through a novel table integration technique. Such technique makes use of an information-aware mask to extract more essential information in the table during the integration, which we call discriminative interaction. Our extensive experiments demonstrate that TableIE outperforms the previous state-of-the-art up to 1.4 on the ACE05 dataset. Besides, since TableIE does not involve the time-consuming graph operation, it is also more efficient than the previous graph-based models, with 13x speed-up in the inference stage. Our code is available at https://github.com/PKUnlp-icler/TableIE Jiaxing Lin, Runxin Xu, Baobao Chang |
ICASSP | 2 |
| 2023 | Knowledgeable Salient Span Mask for Enhancing Language Models as Knowledge Base
Cunxiang Wang, Fuli Luo, Yanyang Li, Runxin Xu, Fei Huang 0002, Yue Zhang 0004 |
NLPCC (2) | 4 |
| 2023 | Making Pre-trained Language Models End-to-end Few-shot Learners with Contrastive Prompt TuningabstractPre-trained Language Models (PLMs) have achieved remarkable performance for various language understanding tasks in IR systems, which require the fine-tuning process based on labeled training data. For low-resource scenarios, prompt-based learning for PLMs exploits prompts as task guidance and turns downstream tasks into masked language problems for effective few-shot fine-tuning. In most existing approaches, the high performance of prompt-based learning heavily relies on handcrafted prompts and verbalizers, which may limit the application of such approaches in real-world scenarios. To solve this issue, we present CP-Tuning, an end-to-end Contrastive Prompt Tuning framework for fine-tuning PLMs without any manual engineering of task-specific prompts and verbalizers. It is integrated with the task-invariant continuous prompt encoding technique with fully trainable prompt parameters. We further propose the pair-wise cost-sensitive contrastive learning procedure to optimize the model in order to achieve verbalizer-free class mapping and enhance the task-invariance of prompts. It explicitly learns to distinguish different classes and makes the decision boundary smoother by assigning different costs to easy and hard cases. Experiments over a variety of language understanding tasks and different PLMs show that CP-Tuning outperforms state-of-the-art methods. Ziyun Xu, Chengyu Wang 0001, Minghui Qiu, Fuli Luo, Runxin Xu, Songfang Huang, Jun Huang 0007 |
WSDM | 5 |
| 2022 | From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model CompressionabstractPre-trained Language Models (PLMs) have achieved great success in various Natural Language Processing (NLP) tasks under the pre-training and fine-tuning paradigm. With large quantities of parameters, PLMs are computation-intensive and resource-hungry. Hence, model pruning has been introduced to compress large-scale PLMs. However, most prior approaches only consider task-specific knowledge towards downstream tasks, but ignore the essential task-agnostic knowledge during pruning, which may cause catastrophic forgetting problem and lead to poor generalization ability. To maintain both task-agnostic and task-specific knowledge in our pruned model, we propose ContrAstive Pruning (CAP) under the paradigm of pre-training and fine-tuning. It is designed as a general framework, compatible with both structured and unstructured pruning. Unified in contrastive learn- ing, CAP enables the pruned model to learn from the pre-trained model for task-agnostic knowledge, and fine-tuned model for task-specific knowledge. Besides, to better retain the performance of the pruned model, the snapshots (i.e., the intermediate models at each pruning iteration) also serve as effective supervisions for pruning. Our extensive experiments show that adopting CAP consistently yields significant improvements, especially in extremely high sparsity scenarios. With only 3% model parameters reserved (i.e., 97% sparsity), CAP successfully achieves 99.2% and 96.3% of the original BERT performance in QQP and MNLI tasks. In addition, our probing experiments demonstrate that the model pruned by CAP tends to achieve better generalization ability. Runxin Xu, Fuli Luo, Chengyu Wang 0001, Baobao Chang, Jun Huang 0007, Songfang Huang, Fei Huang 0002 |
AAAI | 1 |
| 2022 | Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and EfficiencyabstractStructured pruning has been extensively studied on monolingual pre-trained language models and is yet to be fully evaluated on their multilingual counterparts.This work investigates three aspects of structured pruning on multilingual pre-trained language models: settings, algorithms, and efficiency.Experiments on nine downstream tasks show several counterintuitive phenomena: for settings, individually pruning for each language does not induce a better result; for algorithms, the simplest method performs the best; for efficiency, a fast model does not imply that it is also small.To facilitate the comparison on all sparsity levels, we present Dynamic Sparsification, a simple approach that allows training the model once and adapting to different model sizes at inference.We hope this work fills the gap in the study of structured pruning on multilingual pre-trained models and sheds light on future research. Yanyang Li, Fuli Luo, Runxin Xu, Songfang Huang, Fei Huang 0002, Liwei Wang 0009 |
ACL (1) | 3 |
| 2022 | An Enhanced Span-based Decomposition Method for Few-Shot Sequence LabelingabstractPeiyi Wang, Runxin Xu, Tianyu Liu, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui |
NAACL-HLT | 2 |
| 2022 | A Two-Stream AMR-enhanced Model for Document-level Event Argument ExtractionabstractRunxin Xu, Peiyi Wang, Tianyu Liu, Shuang Zeng, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Runxin Xu, Peiyi Wang, Tianyu Liu 0001, Shuang Zeng, Baobao Chang, Zhifang Sui |
NAACL-HLT | 1 |
| 2022 | A Double-Graph Based Framework for Frame Semantic ParsingabstractFrame semantic parsing is a fundamental NLP task, which consists of three subtasks: frame identification, argument identification and role classification.Most previous studies tend to neglect relations between different subtasks and arguments and pay little attention to ontological frame knowledge defined in FrameNet.In this paper, we propose a Knowledge-guided Incremental semantic parser with Double-graph (KID).We first introduce Frame Knowledge Graph (FKG), a heterogeneous graph containing both frames and FEs (Frame Elements) built on the frame knowledge so that we can derive knowledgeenhanced representations for frames and FEs.Besides, we propose Frame Semantic Graph (FSG) to represent frame semantic structures extracted from the text with graph structures.In this way, we can transform frame semantic parsing into an incremental graph construction problem to strengthen interactions between subtasks and relations between arguments.Our experiments show that KID outperforms the previous state-of-the-art method by up to 1.7 F1-score on two FrameNet datasets.Our code is availavle at https://github. com/PKUnlp-icler/KID. Runxin Xu, Baobao Chang |
NAACL-HLT | 3 |
| 2021 | ACMo: Angle-Calibrated Moment Methods for Stochastic Optimization
Xunpeng Huang, Runxin Xu, Hao Zhou 0012, Zhengyang Liu 0002, Lei Li 0005 |
AAAI | 2 |
| 2021 | Document-level Event Extraction via Heterogeneous Graph-based Interaction Model with a TrackerabstractRunxin Xu, Tianyu Liu, Lei Li, Baobao Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Runxin Xu, Tianyu Liu 0001, Lei Li 0005, Baobao Chang |
ACL/IJCNLP (1) | 1 |
| 2021 | Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event ClassificationabstractFew-Shot Event Classification (FSEC) aims at developing a model for event prediction, which can generalize to new event types with a limited number of annotated data. Existing FSEC studies have achieved high accuracy on different benchmarks. However, we find they suffer from trigger biases that signify the statistical homogeneity between some trigger words and target event types, which we summarize as trigger overlapping and trigger separability. The biases can result in context-bypassing problem, i.e., correct classifications can be gained by looking at only the trigger words while ignoring the entire context. Therefore, existing models can be weak in generalizing to unseen data in real scenarios. To further uncover the trigger biases and assess the generalization ability of the models, we propose two new sampling methods, Trigger-Uniform Sampling (TUS) and COnfusion Sampling (COS), for the meta tasks construction during evaluation. Besides, to cope with the context-bypassing problem in FSEC models, we introduce adversarial training and trigger reconstruction techniques. Experiments show these techniques help not only improve the performance, but also enhance the generalization ability of models. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Damai Dai, Baobao Chang, Zhifang Sui |
CIKM | 2 |
| 2021 | Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningabstractRecent pretrained language models extend from millions to billions of parameters.Thus the need to fine-tune an extremely large pretrained model with a limited training corpus arises in various downstream tasks.In this paper, we propose a straightforward yet effective fine-tuning technique, CHILD-TUNING, which updates a subset of parameters (called child network) of large pretrained models via strategically masking out the gradients of the non-child network during the backward process.Experiments on various downstream tasks in GLUE benchmark show that CHILD-TUNING consistently outperforms the vanilla fine-tuning by 1.5 ∼ 8.6 average score among four different pretrained models, and surpasses the prior fine-tuning techniques by 0.6 ∼ 1.3 points.Furthermore, empirical results on domain transfer and task transfer show that CHILD-TUNING can obtain better generalization performance by large margins. Runxin Xu, Fuli Luo, Chuanqi Tan, Baobao Chang, Songfang Huang, Fei Huang 0002 |
EMNLP (1) | 1 |
| 2020 | Double Graph Based Reasoning for Document-level Relation ExtractionabstractDocument-level relation extraction aims to extract relations among entities within a document.Different from sentence-level relation extraction, it requires reasoning over multiple sentences across paragraphs.In this paper, we propose Graph Aggregation-and-Inference Network (GAIN), a method to recognize such relations for long paragraphs.GAIN constructs two graphs, a heterogeneous mentionlevel graph (MG) and an entity-level graph (EG).The former captures complex interaction among different mentions and the latter aggregates mentions underlying for the same entities.Based on the graphs we propose a novel path reasoning mechanism to infer relations between entities.Experiments on the public dataset, DocRED, show GAIN achieves a significant performance improvement (2.85 on F1) over the previous state-of-the-art.Our code is available at https://github.com/ PKUnlp-icler/GAIN. Shuang Zeng, Runxin Xu, Baobao Chang, Lei Li 0005 |
EMNLP (1) | 2 |