VLDB 2026 Research / reviewers in the wild / expert
Zhipeng Chen 0001
dblp:22/8437-1
· DBLP profile ↗
16ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0009-4875-5465ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Effective Code-Integrated ReasoningabstractIn this paper, we investigate code-integrated reasoning (CIR), where models generate code when necessary and integrate feedback by executing it through a code interpreter. To acquire this capability, models must learn when and how to use external code tools effectively, which is supported by tool-augmented reinforcement learning (RL). Despite its benefits, tool-augmented RL can still suffer from potential instability in the learning dynamics. In light of this challenge, we present a systematic approach ETIR (Effective TIR) to improving the training effectiveness and stability of tool-augmented RL for code-integrated reasoning. Specifically, we develop enhanced training strategies that balance exploration and stability, progressively building tool-use capabilities while improving reasoning performance. Through extensive experiments on five mainstream mathematical reasoning benchmarks, our model demonstrates significant performance improvements over multiple competitive baselines. Furthermore, we conduct an in-depth analysis of the mechanism of code-integrated reasoning, revealing several key insights, such as the extension of model’s capability boundaries and the simultaneous improvement of reasoning efficiency through code integration. These findings underscore the potential of code-integrated reasoning as a scalable paradigm for advancing robust and efficient language model reasoning. Fei Bai, Yingqian Min, Beichen Zhang 0003, Zhipeng Chen 0001, Wayne Xin Zhao, Zheng Liu 0011, Zhongyuan Wang 0006, Hongteng Xu |
AAAI | 4 |
| 2026 | Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language ModelsabstractThe rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation frameworks.To address this, we introduce OlymMATH, a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions.OlymMATH is the first benchmark to unify dual evaluation paradigms within a single suite: (1) natural language evaluation through OlymMATH-EASY and OlymMATH-HARD, comprising 200 computational problems with numerical answers for objective rule-based assessment, and (2) formal verification through OlymMATH-LEAN, offering 150 problems formalized in Lean 4 for rigorous process-level evaluation.All problems are manually sourced from printed publications to minimize data contamination, verified by experts, and span four core domains.Extensive experiments reveal the benchmark's significant challenge, and our analysis also uncovers consistent performance gaps between languages and identifies cases where models employ heuristic "guessing" rather than rigorous reasoning.To further support community research, we release 582k+ reasoning trajectories, a visualization tool, and expert solutions at https://github.com/RUCAIBox/OlymMATH. Haoxiang Sun, Yingqian Min, Zhipeng Chen 0001, Wayne Xin Zhao, Ji-Rong Wen |
ACL (1) | 3 |
| 2026 | A Survey of Large Language ModelsabstractAbstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities, LLMs necessitate new frameworks for understanding their development, behavior, and societal impact. This survey systematically reviews recent advancements in LLM techniques across four key dimensions: (1) pre-training methodologies, which establish core model capabilities through large-scale self-supervised training, architectural innovations, and data curation strategies; (2) post-training techniques, including supervised fine-tuning and reinforcement learning, which adapt foundational models to downstream tasks and enhance their alignment and safety; (3) utilization strategies, such as in-context learning, prompt engineering, and agentic reasoning, that optimize real-world deployment and enable effective interaction with external environments; and (4) evaluation methods, encompassing benchmarks for key ability dimensions such as core language capabilities, reasoning, and safety, which support comprehensive and reliable assessment of model performance. Additionally, we identify critical research issues, including those concerning theoretical foundations, efficient scaling, alignment, and agentic capability, and highlight the open challenges they present. By synthesizing state-of-the-art insights and emerging trends, this survey aims to provide a systematic and comprehensive framework for understanding the trajectory, current limitations, and future directions of LLM progress. Wayne Xin Zhao, Kun Zhou 0002, Junyi Li 0001, Zican Dong, Yupeng Hou, Beichen Zhang 0003, Yingqian Min, Junjie Zhang 0009, Peiyu Liu 0002, Xiaolei Wang 0005, Yifan Du 0002, Chen Yang 0032, Zhipeng Chen 0001, Jinhao Jiang, Ruiyang Ren, Yifan Li 0009, Xinyu Tang 0004, Zikang Liu 0001, Jian-Yun Nie, Ji-Rong Wen |
Frontiers Comput. Sci. | 15 |
| 2025 | Towards Effective and Efficient Continual Pre-training of Large Language ModelsabstractContinual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly improve the Chinese language ability and scientific reasoning ability of LLMs. To achieve it, we design specific data mixture and curriculum strategies based on existing datasets and synthetic high-quality data. Concretely, we synthesize multidisciplinary scientific QA pairs based on related web pages to guarantee the data quality, and also devise the performance tracking and data mixture adjustment strategy to ensure the training stability. For the detailed designs, we conduct preliminary studies on a relatively small model, and summarize the findings to help optimize our CPT method. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of Llama-3 (8B), including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval). Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE. Jie Chen 0007, Zhipeng Chen 0001, Kun Zhou 0002, Yutao Zhu 0001, Jinhao Jiang, Yingqian Min, Wayne Xin Zhao, Zhicheng Dou, Jiaxin Mao, Yankai Lin 0001, Ruihua Song, Jun Xu 0001, Xu Chen 0017, Rui Yan 0001, Zhewei Wei, Di Hu 0001, Wenbing Huang 0001, Ji-Rong Wen |
ACL (1) | 2 |
| 2025 | Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language ModelsabstractMulti-lingual ability transfer has become increasingly important for the broad application of large language models (LLMs).Existing work highly relies on training with the multi-lingual ability-related data, which may not be available for low-resource languages.To solve it, we propose a Multilingual Abilities Extraction and Combination approach (MAEC), which decomposes and extracts language-agnostic ability-related weights from LLMs, and combines them across different languages by simple addition and subtraction operations without training.Specifically, our MAEC consists of the extraction and combination stages.In the extraction stage, we firstly locate key neurons that are highly related to specific abilities, and then employ them to extract the transferable ability-related weights.In the combination stage, we further select the ability-related tensors that mitigate the linguistic effects, and design a combining strategy based on them and the languagespecific weights, to build the multi-lingual ability-enhanced LLM.To assess the effectiveness of our approach, we conduct extensive experiments on LLaMA-3 8B on mathematical and scientific tasks in both high-resource and low-resource lingual scenarios.Empirical results have shown that MAEC can effectively and efficiently extract and combine the advanced abilities, achieving comparable performance with PaLM.Resources are available at https://github.com/RUCAIBox/MAET. Zhipeng Chen 0001, Kun Zhou 0002, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen |
EMNLP | 1 |
| 2025 | ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming ContestsabstractWith the significant progress of large reasoning models in complex coding and reasoning tasks, existing benchmarks, like LiveCodeBench and CodeElo, are insufficient to evaluate the coding capabilities of large language models (LLMs) in real competition environments. Moreover, current evaluation metrics such as Pass@K fail to capture the reflective abilities of reasoning models. To address these challenges, we propose ICPC-Eval, a top-level competitive coding benchmark designed to probing the frontiers of LLM reasoning. ICPC-Eval includes 118 carefully curated problems from 11 recent ICPC contests held in various regions of the world, offering three key contributions: 1) A challenging realistic ICPC competition scenario, featuring a problem type and difficulty distribution consistent with actual contests. 2) A robust test case generation method and a corresponding local evaluation toolkit, enabling efficient and accurate local evaluation. 3) An effective test-time scaling evaluation metric, Refine@K, which allows iterative repair of solutions based on execution feedback. The results underscore the significant challenge in evaluating complex reasoning abilities: top-tier reasoning models like DeepSeek-R1 often rely on multi-turn code feedback to fully unlock their in-context reasoning potential when compared to non-reasoning counterparts. Furthermore, despite recent advancements in code generation, these models still lag behind top-performing human teams. We release the benchmark at: https://github.com/RUCAIBox/ICPC-Eval Shiyi Xu, Yingqian Min, Zhipeng Chen 0001, Wayne Xin Zhao, Ji-Rong Wen |
NeurIPS | 4 |
| 2024 | Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model AlignmentabstractLarge language models (LLMs) are still struggling in aligning with human preference in complex tasks and scenarios.They are prone to overfit into the unexpected patterns or superficial styles in the training data.We conduct an empirical study that only selects the top-10% most updated parameters in LLMs for alignment training, and see improvements in the convergence process and final performance.It indicates the existence of redundant neurons in LLMs for alignment training.To reduce its influence, we propose a low-redundant alignment method named ALLO, focusing on optimizing the most related neurons with the most useful supervised signals.Concretely, we first identify the neurons that are related to the human preference data by a gradient-based strategy, then identify the alignment-related key tokens by reward models for computing loss.Besides, we also decompose the alignment process into the forgetting and learning stages, where we first forget the tokens with unaligned knowledge and then learn aligned knowledge, by updating different ratios of neurons, respectively.Experimental results on 10 datasets have shown the effectiveness of ALLO.Our code and data are available at https://github.com/RUCAIBox/ALLO. Zhipeng Chen 0001, Kun Zhou 0002, Wayne Xin Zhao, Jingyuan Wang 0001, Ji-Rong Wen |
EMNLP | 1 |
| 2024 | JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis ModelsabstractMathematical reasoning is an important capability of large language models~(LLMs) for real-world applications.
To enhance this capability, existing work either collects large-scale math-related texts for pre-training, or relies on stronger LLMs (\eg GPT-4) to synthesize massive math problems. Both types of work generally lead to large costs in training or synthesis.
To reduce the cost, based on open-source available texts, we propose an efficient way that trains a small LLM for math problem synthesis, to efficiently generate sufficient high-quality pre-training data.
To achieve it, we create a dataset using GPT-4 to distill its data synthesis capability into the small LLM.
Concretely, we craft a set of prompts based on human education stages to guide GPT-4, to synthesize problems covering diverse math knowledge and difficulty levels.
Besides, we adopt the gradient-based influence estimation method to select the most valuable math-related texts.
The both are fed into GPT-4 for creating the knowledge distillation dataset to train the small LLM.
We leverage it to synthesize 6 million math problems for pre-training our JiuZhang3.0 model. The whole process only needs to invoke GPT-4 API 9.3k times and use 4.6B data for training.
Experimental results have shown that JiuZhang3.0 achieves state-of-the-art performance on several mathematical reasoning datasets, under both natural language reasoning and tool manipulation settings.
Our code and data will be publicly released in \url{https://github.com/RUCAIBox/JiuZhang3.0}. Kun Zhou 0002, Beichen Zhang 0003, Zhipeng Chen 0001, Wayne Xin Zhao, Jing Sha, Zhichao Sheng, Shijin Wang 0001, Ji-Rong Wen |
NeurIPS | 4 |
| 2023 | JiuZhang 2.0: A Unified Chinese Pre-trained Language Model for Multi-task Mathematical Problem SolvingabstractAlthough pre-trained language models~(PLMs) have recently advanced the research progress in mathematical reasoning, they are not specially designed as a capable multi-task solver, suffering from high cost for multi-task deployment (e.g. a model copy for a task) and inferior performance on complex mathematical problems in practical applications. To address these issues, we propose JiuZhang 2.0, a unified Chinese PLM specially for multi-task mathematical problem solving. Our idea is to maintain a moderate-sized model and employ the cross-task knowledge sharing to improve the model capacity in a multi-task setting. Specially, we construct a Mixture-of-Experts (MoE) architecture for modeling mathematical text, to capture the common mathematical knowledge across tasks. For optimizing the MoE architecture, we design multi-task continual pre-training and multi-task fine-tuning strategies for multi-task adaptation. These training strategies can effectively decompose the knowledge from the task data and establish the cross-task sharing via expert networks. To further improve the general capacity of solving different complex tasks, we leverage large language models (LLMs) as complementary models to iteratively refine the generated solution by our PLM, via in-context learning. Extensive experiments have demonstrated the effectiveness of our model. Wayne Xin Zhao, Kun Zhou 0002, Beichen Zhang 0003, Zheng Gong 0001, Zhipeng Chen 0001, Yuanhang Zhou, Ji-Rong Wen, Jing Sha, Shijin Wang 0001, Cong Liu 0006 |
KDD | 5 |
| 2022 | ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language ModelsabstractJunyi Li, Tianyi Tang, Zheng Gong, Lixin Yang, Zhuohao Yu, Zhipeng Chen, Jingyuan Wang, Xin Zhao, Ji-Rong Wen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Junyi Li 0001, Zheng Gong 0001, Lixin Yang 0005, Zhuohao Yu 0001, Zhipeng Chen 0001, Jingyuan Wang 0001, Wayne Xin Zhao, Ji-Rong Wen |
NAACL-HLT | 6 |
| 2020 | A Sentence Cloze Dataset for Chinese Machine Reading ComprehensionabstractOwing to the continuous efforts by the Chinese NLP community, more and more Chinese machine reading comprehension datasets become available.To add diversity in this area, in this paper, we propose a new task called Sentence Cloze-style Machine Reading Comprehension (SC-MRC).The proposed task aims to fill the right candidate sentence into the passage that has several blanks.We built a Chinese dataset called CMRC 2019 to evaluate the difficulty of the SC-MRC task.Moreover, to add more difficulties, we also made fake candidates that are similar to the correct ones, which requires the machine to judge their correctness in the context.The proposed dataset contains over 100K blanks (questions) within over 10K passages, which was originated from Chinese narrative stories.To evaluate the dataset, we implement several baseline systems based on the pre-trained models, and the results show that the stateof-the-art model still underperforms human performance by a large margin.We release the dataset and baseline system to further facilitate our community. Yiming Cui 0001, Ting Liu 0001, Ziqing Yang 0001, Zhipeng Chen 0001, Wanxiang Che, Shijin Wang 0001 |
COLING | 4 |
| 2019 | Convolutional Spatial Attention Model for Reading Comprehension with Multiple-Choice QuestionsabstractMachine Reading Comprehension (MRC) with multiplechoice questions requires the machine to read given passage and select the correct answer among several candidates. In this paper, we propose a novel approach called Convolutional Spatial Attention (CSA) model which can better handle the MRC with multiple-choice questions. The proposed model could fully extract the mutual information among the passage, question, and the candidates, to form the enriched representations. Furthermore, to merge various attention results, we propose to use convolutional operation to dynamically summarize the attention values within the different size of regions. Experimental results show that the proposed model could give substantial improvements over various state-of- the-art systems on both RACE and SemEval-2018 Task11 datasets. Zhipeng Chen 0001, Yiming Cui 0001, Shijin Wang 0001 |
AAAI | 1 |
| 2019 | A Span-Extraction Dataset for Chinese Machine Reading ComprehensionabstractYiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, Guoping Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yiming Cui 0001, Ting Liu 0001, Wanxiang Che, Zhipeng Chen 0001, Shijin Wang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Dataset for the First Evaluation on Chinese Machine Reading Comprehension
Yiming Cui 0001, Ting Liu 0001, Zhipeng Chen 0001, Shijin Wang 0001 |
LREC | 3 |
| 2017 | Attention-over-Attention Neural Networks for Reading ComprehensionabstractCloze-style queries are representative problems in reading comprehension. Over the past few months, we have seen much progress that utilizing neural network approach to solve Cloze-style questions. In this paper, we present a novel model called attention-over-attention reader for the Cloze-style reading comprehension task. Our model aims to place another attention mechanism over the document-level attention, and induces "attended attention" for final predictions. Unlike the previous works, our neural network model requires less pre-defined hyper-parameters and uses an elegant architecture for modeling. Experimental results show that the proposed attention-over-attention model significantly outperforms various state-of-the-art systems by a large margin in public datasets, such as CNN and Children's Book Test datasets. Yiming Cui 0001, Zhipeng Chen 0001, Si Wei, Shijin Wang 0001, Ting Liu 0001 |
ACL (1) | 2 |
| 2016 | Consensus Attention-based Neural Networks for Chinese Reading ComprehensionabstractReading comprehension has embraced a booming in recent NLP research. Several institutes have released the Cloze-style reading comprehension data, and these have greatly accelerated the research of machine comprehension. In this work, we firstly present Chinese reading comprehension datasets, which consist of People Daily news dataset and Children’s Fairy Tale (CFT) dataset. Also, we propose a consensus attention-based neural network architecture to tackle the Cloze-style reading comprehension problem, which aims to induce a consensus attention over every words in the query. Experimental results show that the proposed neural network significantly outperforms the state-of-the-art baselines in several public datasets. Furthermore, we setup a baseline for Chinese reading comprehension task, and hopefully this would speed up the process for future research. Yiming Cui 0001, Ting Liu 0001, Zhipeng Chen 0001, Shijin Wang 0001 |
COLING | 3 |