Hao Peng 0015

dblp:69/7742-15 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
19since 2021 · last 2026
0009-0006-7192-5790ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures
abstract
Yi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao, Hao Peng, Xiaobao Wu, Jianhui Chen, Muhan Zhang, Liangming Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zijun Yao 0002, Hao Peng 0015, Xiaobao Wu, Muhan Zhang, Liangming Pan
ACL (1)5
2026 WildReward: Learning Reward Models from In-the-Wild Human Interactions
abstract
Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs.With the widespread deployment of LLMs, in-the-wild interactions have emerged as a rich source of implicit reward signals.This raises the question: Can we develop reward models directly from in-the-wild interactions?In this work, we explore this possibility by adopting WildChat as an interaction source and proposing a pipeline to extract reliable human feedback, yielding 186k high-quality instances for training WILDREWARD via ordinal regression directly on user feedback without preference pairs.Extensive experiments demonstrate that WILDREWARD achieves comparable or even superior performance compared to conventional reward models, with improved calibration and cross-sample consistency.We also observe that WILDREWARD benefits directly from user diversity, where more users yield stronger reward models.Finally, we apply WILDREWARD to online DPO training and observe significant improvements across various tasks.
Hao Peng 0015, Yunjia Qi, Xiaozhi Wang, Zijun Yao 0002, Lei Hou 0001, Juan-Zi Li
ACL (1)1
2026 SimPBL: A Multi-Agent Framework for Project-Based Learning
abstract
Daniel Zhang-Li, Joy Jia Yin Lim, Binglin Liu, Shangqing Tu, Zijun Yao, Hao Peng, Jifan Yu, Haoxuan Li, Zhanxin Hao, Ye He, Zekun Li, Jiangyi Wang, Lei Hou, Bin Xu, Xin Cong, Zhiyuan Liu, Huiqin Liu, Yu Zhang, Juanzi Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Daniel Zhang-Li, Joy Lim Jia Yin, Binglin Liu, Shangqing Tu, Zijun Yao 0002, Hao Peng 0015, Jifan Yu, Haoxuan Li 0003, Zhanxin Hao, Jiangyi Wang, Lei Hou 0001, Bin Xu 0001, Xin Cong, Zhiyuan Liu 0001, Huiqin Liu, Yu Zhang 0186, Juan-Zi Li
ACL (1)6
2026 OpenEP: Open-Ended Future Event Prediction
abstract
Future event prediction (FEP) is a long-standing and crucial task in the world, as understanding the evolution of events enables early risk identification, informed decision-making, and strategic planning. Existing work typically treats event prediction as classification tasks and confines the outcomes of future events to a fixed scope, such as yes/no questions, candidate set, and taxonomy, which is difficult to include all possible outcomes of future events. In this article, we introduce OpenEP (an Open -Ended F EP task), which generates flexible and diverse predictions aligned with real-world scenarios. This is mainly reflected in two aspects: firstly, the predictive questions are diverse, covering different stages of event development and perspectives; secondly, the outcomes are flexible, without constraints on scope or format. To facilitate the study of this task, we construct OpenEPBench , a dynamic dataset where events are annotated on the same day they occur. For question construction, we pose questions from seven perspectives, including time, location, event development, event outcome, event impact, event response, and other, to facilitate an in-depth analysis and understanding of the comprehensive evolution of events. For outcome construction, we collect free-form text containing the outcomes as ground truth to provide semantically complete and detail-enriched outcomes. Furthermore, we propose StkFEP , a stakeholder-enhanced FEP framework, that incorporates event characteristics for open-ended settings. Our method extracts stakeholders involved in events to extend questions to gather diverse information. We also collect both contextually relevant events and historically analogous events to reveal potential evolutionary patterns. Experimental results indicate that accurately predicting future events in open-ended settings is challenging for existing LLMs. In addition, we thoroughly summarize the problems encountered in prediction, hoping to provide insights for future research. Code and data are available at: https://github.com/ncepu-eai/OpenEP .
Hao Peng 0015, Xiaozhi Wang, Lei Hou 0001, Juan-Zi Li
ACM Trans. Inf. Syst.2
2025 Pre-training Distillation for Large Language Models: A Design Space Exploration
abstract
Knowledge distillation (KD) aims to transfer knowledge from a large teacher model to a smaller student model.Previous work applying KD in the field of large language models (LLMs) typically focused on the post-training phase, where the student LLM learns directly from instructions and corresponding responses generated by the teacher model.In this paper, we extend KD to the pre-training phase of LLMs, named pre-training distillation (PD).We first conduct a preliminary experiment using GLM-4-9B as the teacher LLM to distill a 1.9B parameter student LLM, validating the effectiveness of PD.Considering the key impact factors of distillation, we systematically explore the design space of pre-training distillation across four aspects: logits processing, loss selection, scaling law, and offline or online logits.We conduct extensive experiments to explore the design space of pre-training distillation and find better configurations and interesting conclusions, such as larger student LLMs generally benefiting more from pre-training distillation, while a larger teacher LLM does not necessarily guarantee better results.We hope our exploration of the design space will inform future practices in pre-training distillation.
Hao Peng 0015, Yushi Bai, Zijun Yao 0002, Lei Hou 0001, Juan-Zi Li
ACL (1)1
2025 Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
abstract
Reward models (RMs) are crucial for the training and inference-time scaling up of large language models (LLMs). However, existing reward models primarily focus on human preferences, neglecting verifiable correctness signals which have shown strong potential in training LLMs. In this paper, we propose agentic reward modeling, a reward system that combines reward models with verifiable correctness signals from different aspects to provide reliable rewards. We empirically implement a reward agent, named RewardAgent, that combines human preference rewards with two verifiable signals: factuality and instruction following, to provide more reliable rewards. We conduct comprehensive experiments on existing reward model benchmarks and inference-time best-of-n searches on real-world downstream tasks. RewardAgent significantly outperforms vanilla reward models, demonstrating its effectiveness. We further construct training preference pairs using RewardAgent and train an LLM with the DPO objective, achieving superior performance on various NLP benchmarks compared to conventional reward models. Our codes are publicly released to facilitate further research.
Hao Peng 0015, Yunjia Qi, Xiaozhi Wang, Zijun Yao 0002, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
ACL (1)1
2025 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
abstract
Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yushi Bai, Shangqing Tu, Hao Peng 0015, Xiaozhi Wang, Shulin Cao, Jiazheng Xu, Lei Hou 0001, Yuxiao Dong, Jie Tang 0001, Juan-Zi Li
ACL (1)4
2025 Constraint Back-translation Improves Complex Instruction Following of Large Language Models
abstract
Large language models (LLMs) struggle to follow instructions with complex constraints in format, length, etc. Following the conventional instruction-tuning practice, previous works conduct post-training on complex instruction-response pairs generated by feeding complex instructions to advanced LLMs. However, even advanced LLMs cannot follow complex instructions well, thus limiting the quality of generated data. In this work, we find that existing datasets inherently contain implicit complex constraints and propose a novel data generation technique, constraint back-translation. Specifically, we take the high-quality instruction-response pairs in existing datasets and only adopt advanced LLMs to add complex constraints already met by the responses to the instructions, which naturally reduces costs and data noise. In the experiments, we adopt Llama3-70B-Instruct to back-translate constraints and create a high-quality complex instruction-response dataset, named Crab. We present that post-training on Crab improves multiple backbone LLMs' complex instruction-following ability, evaluated on extensive instruction-following benchmarks. We further find that constraint back-translation also serves as a useful auxiliary training objective in post-training. Our code, data, and models are released to facilitate future research.
Yunjia Qi, Hao Peng 0015, Xiaozhi Wang, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
CIKM2
2025 StoryWriter: A Multi-Agent Framework for Long Story Generation
abstract
Long story generation remains a challenge for existing large language models (LLMs), primarily due to two main factors: (1) discourse coherence, which requires plot consistency, logical coherence, and completeness in the long-form generation, and (2) narrative complexity, which requires an interwoven and engaging narrative. In this paper, we present StoryWriter, a modular and open-source multi-agent framework for controllable and scalable long story generation. We conduct both human and automated evaluation, and StoryWriter significantly outperforms existing story generation baselines in both story quality and length. Furthermore, we use StoryWriter to generate a dataset, which contains about 6,000 high-quality long stories, with an average length of 8,000 words. We train the model Llama3.1-8B and GLM4-9B using supervised fine-tuning on LongStory and develop StoryWriterLLAMA and StoryWriterGLM, which demonstrates advanced performance in long story generation. All code, models, and data are made publicly available to encourage further development.
Haotian Xia, Hao Peng 0015, Yunjia Qi, Bin Xu 0001, Juan-Zi Li, Lei Hou 0001, Xiaozhi Wang
CIKM2
2025 MMD-ERE: Multi-Agent Multi-Sided Debate for Event Relation Extraction
abstract
Event relation extraction (ERE) is becoming increasingly important in the era of large language models. An extensive body of research has explored how performance can be further enhanced by the emergence of exciting technologies like chain-of-thought and self-refinement. In this paper, we introduce MMD-ERE, a multi-agent multi-sided debate approach for event relation extraction, which explores the understanding of event relations among different participants before and after debate. Specifically, for organizing the debate, participants are divided into multiple groups, each assigned its own debate topic, and the process effectively integrates both cooperation and confrontation. We also regard the audience as a crucial participant, as their conclusions from an observer’s perspective tend to be more objective. In the end, we explore the understanding of event relations among different participants before and after the debate. Experiments across various ERE tasks and LLMs demonstrate that MMD-ERE outperforms established baselines. Further analysis shows that debates can effectively enhance participants’ understanding of event relations.
Hao Peng 0015, Lei Hou 0001, Juan-Zi Li
COLING2
2025 VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
abstract
Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing large language models (LLMs), with verification engineering playing a central role.However, best practices for RL in instruction following remain underexplored.In this work, we explore the verification challenge in RL for instruction following and propose VERIF, a verification method that combines rule-based code verification with LLM-based verification from a large reasoning model (e.g., QwQ-32B).To support this approach, we construct a highquality instruction-following dataset, VERIN-STRUCT, containing approximately 22,000 instances with associated verification signals.We apply RL training with VERIF to two models, achieving significant improvements across several representative instruction-following benchmarks.The trained models reach state-of-theart performance among models of comparable size and generalize well to unseen constraints.We further observe that their general capabilities remain unaffected, suggesting that RL with VERIF can be integrated into existing RL recipes to enhance overall model performance.We have released our datasets, codes, and models to facilitate future research 1 .
Hao Peng 0015, Yunjia Qi, Xiaozhi Wang, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
EMNLP1
2025 Advancing LLM Reasoning Generalists with Preference Trees
abstract
We introduce EURUS, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B, Llama-3-8B, and Mixtral-8x22B, EURUS models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, EURUX-8X22B outperforms GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 test sets covering five tasks. The strong performance of EURUS can be primarily attributed to ULTRAINTERACT, our newly-curated large-scale, high-quality training data dataset specifically designed for complex reasoning tasks. ULTRAINTERACT can be used in both supervised fine-tuning, preference learning, and reward modeling. It pairs each instruction with a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format, (2) multi-turn interaction trajectories with the environment and the critique, and (3) pairwise positive and negative responses to facilitate preference learning. ULTRAINTERACT allows us to conduct an in-depth exploration of preference learning for reasoning tasks. Our investigation reveals that some well-established preference learning algorithms may be less suitable for reasoning tasks compared to their effectiveness in general conversations. The hypothesis is that in reasoning tasks, the space of correct answers is much smaller than that of incorrect ones, so it is necessary to explicitly increase the reward of chosen data. Therefore, in addition to increasing the reward margin as many preference learning algorithms do, the absolute values of positive responses’ rewards should be positive and may serve as a proxy for performance. Inspired by this, we derive a novel reward modeling objective and empirically that it leads to a stable reward modeling curve and better performance. Together with ULTRAINTERACT, we obtain a strong reward model.
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding 0002, Xingyao Wang 0002, Boji Shan, Zeyuan Liu, Ruobing Xie, Yankai Lin 0001, Zhenghao Liu 0001, Bowen Zhou 0002, Hao Peng 0015, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR14
2025 AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios
abstract
Large Language Models (LLMs) have demonstrated advanced capabilities in real-world agentic applications. Growing research efforts aim to develop LLM-based agents to address practical demands, introducing a new challenge: agentic scenarios often involve lengthy instructions with complex constraints, such as extended system prompts and detailed tool specifications. While adherence to such instructions is crucial for agentic applications, whether LLMs can reliably follow them remains underexplored. In this paper, we introduce AgentIF, the first benchmark for systematically evaluating LLM instruction following ability in agentic scenarios. AgentIF features three key characteristics: (1) Realistic, constructed from $50$ real-world agentic applications. (2) Long, averaging $1,723$ words with a maximum of $15,630$ words. (3) Complex, averaging $11.9$ constraints per instruction, covering diverse constraint types, such as tool specifications and condition constraints.To construct AgentIF, we collect $707$ human-annotated instructions across $50$ agentic tasks from industrial application agents and open-source agentic systems. For each instruction, we annotate the associated constraints and corresponding evaluation metrics, including code-based evaluation, LLM-based evaluation, and hybrid code-LLM evaluation.We use AgentIF to systematically evaluate existing advanced LLMs. We observe that current models generally perform poorly, especially in handling complex constraint structures and tool specifications. We further conduct error analysis and analytical experiments on instruction length and meta constraints, providing some findings about the failure modes of existing LLMs. We have released the code and data to facilitate future research.
Yunjia Qi, Hao Peng 0015, Xiaozhi Wang, Amy Xin, Youfeng Liu, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
NeurIPS2
2024 MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation
abstract
Xiaozhi Wang, Hao Peng, Yong Guan, Kaisheng Zeng, Jianhui Chen, Lei Hou, Xu Han, Yankai Lin, Zhiyuan Liu, Ruobing Xie, Jie Zhou, Juanzi Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaozhi Wang, Hao Peng 0015, Kaisheng Zeng, Lei Hou 0001, Xu Han 0007, Yankai Lin 0001, Zhiyuan Liu 0001, Ruobing Xie, Jie Zhou 0016, Juan-Zi Li
ACL (1)2
2024 ADELIE: Aligning Large Language Models on Information Extraction
abstract
Large language models (LLMs) usually fall short on information extraction (IE) tasks and struggle to follow the complex instructions of IE tasks.This primarily arises from LLMs not being aligned with humans, as mainstream alignment datasets typically do not include IE data.In this paper, we introduce ADELIE (Aligning large language moDELs on Information Extraction), an aligned LLM that effectively solves various IE tasks, including closed IE, open IE, and on-demand IE.We first collect and construct a high-quality alignment corpus IEInstruct for IE.Then we train ADELIE SFT using instruction tuning on IEInstruct.We further train ADELIE SFT with direct preference optimization (DPO) objective, resulting in ADELIE DPO .Extensive experiments on various held-out IE datasets demonstrate that our models (ADELIE SFT and ADELIE DPO ) achieve state-of-the-art (SoTA) performance among open-source models.We further explore the general capabilities of ADELIE, and experimental results reveal that their general capabilities do not exhibit a noticeable decline.We have released the code, data, and models to facilitate further research.
Yunjia Qi, Hao Peng 0015, Xiaozhi Wang, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
EMNLP2
2024 KoLA: Carefully Benchmarking World Knowledge of Large Language Models
abstract
The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems.
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Hao Peng 0015, Zijun Yao 0002, Hanming Li, Zheyuan Zhang 0002, Yushi Bai, Yantao Liu, Amy Xin, Kaifeng Yun, Linlu Gong, Nianyi Lin, Zhi-Li Wu, Yunjia Qi, Weikai Li 0002, Kaisheng Zeng, Ji Qi 0003, Hailong Jin, Jinxin Liu 0002, Yu Gu 0029, Yuan Yao 0011, Ning Ding 0002, Lei Hou 0001, Zhiyuan Liu 0001, Bin Xu 0001, Jie Tang 0001, Juan-Zi Li
ICLR7
2023 Pre-training language model incorporating domain-specific heterogeneous knowledge into a unified representation
Hongyin Zhu, Hao Peng 0015, Zhiheng Lyu, Lei Hou 0001, Juan-Zi Li, Jinghui Xiao
Expert Syst. Appl.2
2022 COPEN: Probing Conceptual Knowledge in Pre-trained Language Models
abstract
Conceptual knowledge is fundamental to human cognition and knowledge bases.However, existing knowledge probing works only focus on evaluating factual knowledge of pre-trained language models (PLMs) and ignore conceptual knowledge.Since conceptual knowledge often appears as implicit commonsense behind texts, designing probes for conceptual knowledge is hard.Inspired by knowledge representation schemata, we comprehensively evaluate conceptual knowledge of PLMs by designing three tasks to probe whether PLMs organize entities by conceptual similarities, learn conceptual properties, and conceptualize entities in contexts, respectively.For the tasks, we collect and annotate 24k data instances covering 393 concepts, which is COPEN, a COnceptual knowledge Probing bENchmark.Extensive experiments on different sizes and types of PLMs show that existing PLMs systematically lack conceptual knowledge and suffer from various spurious correlations.We believe this is a critical bottleneck for realizing human-like cognition in PLMs.COPEN and our codes are publicly released at https: //github.com/THU-KEG/COPEN.
Hao Peng 0015, Xiaozhi Wang, Shengding Hu, Hailong Jin, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Qun Liu 0001
EMNLP1
2022 MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction
abstract
Xiaozhi Wang, Yulin Chen, Ning Ding, Hao Peng, Zimu Wang, Yankai Lin, Xu Han, Lei Hou, Juanzi Li, Zhiyuan Liu, Peng Li, Jie Zhou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xiaozhi Wang, Yulin Chen 0001, Ning Ding 0002, Hao Peng 0015, Yankai Lin 0001, Xu Han 0007, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016
EMNLP4
2020 Learning from Context or Names? An Empirical Study on Neural Relation Extraction
abstract
Neural models have achieved remarkable success on relation extraction (RE) benchmarks.However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance of these models.To this end, we empirically study the effect of two main information sources in text: textual context and entity mentions (names).We find that (i) while context is the main source to support the predictions, RE models also heavily rely on the information from entity mentions, most of which is type information, and (ii) existing datasets may leak shallow heuristics via entity mentions and thus contribute to the high performance on RE benchmarks.Based on the analyses, we propose an entity-masked contrastive pre-training framework for RE to gain a deeper understanding on both textual context and type information while avoiding rote memorization of entities or use of superficial cues in mentions.We carry out extensive experiments to support our views, and show that our framework can improve the effectiveness and robustness of neural models in different RE scenarios.All the code and datasets are released at https://github.com/thunlp/
Hao Peng 0015, Tianyu Gao 0001, Xu Han 0007, Yankai Lin 0001, Peng Li 0030, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
EMNLP (1)1