Xingyao Wang 0002

dblp:264/9892-2 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-3483-8624ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021
YearPublicationVenuePosition
2025 LocAgent: Graph-Guided LLM Agents for Code Localization
abstract
Zhaoling Chen, Robert Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, Xingyao Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhaoling Chen, Robert Tang, Gangda Deng, Fang Wu 0002, Jialong Wu 0010, Viktor Prasanna 0001, Arman Cohan, Xingyao Wang 0002
ACL (1)9
2025 OpenHands: An Open Platform for AI Software Developers as Generalist Agents
abstract
Software is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and effect change in their surrounding environments. In this paper, we introduce OpenHands, a platform for the development of powerful and flexible AI agents that interact with the world in similar ways to a human developer: by writing code, interacting with a command line, and browsing the web. We describe how the platform allows for the implementation of new agents, utilization of various LLMs, safe interaction with sandboxed environments for code execution, and incorporation of evaluation benchmarks. Based on our currently incorporated benchmarks, we perform an evaluation of agents over 13 challenging tasks, including software engineering (e.g., SWE-Bench) and web browsing (e.g., WebArena), amongst others. Released under the permissive MIT license, OpenHands is a community project spanning academia and industry with more than 2K contributions from over 186 contributors in less than six months of development, and will improve going forward.
Xingyao Wang 0002, Boxuan Li, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Yueqi Song, Bowen Li 0002, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang 0002, Binyuan Hui, Junyang Lin
ICLR1
2025 Advancing LLM Reasoning Generalists with Preference Trees
abstract
We introduce EURUS, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B, Llama-3-8B, and Mixtral-8x22B, EURUS models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, EURUX-8X22B outperforms GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 test sets covering five tasks. The strong performance of EURUS can be primarily attributed to ULTRAINTERACT, our newly-curated large-scale, high-quality training data dataset specifically designed for complex reasoning tasks. ULTRAINTERACT can be used in both supervised fine-tuning, preference learning, and reward modeling. It pairs each instruction with a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format, (2) multi-turn interaction trajectories with the environment and the critique, and (3) pairwise positive and negative responses to facilitate preference learning. ULTRAINTERACT allows us to conduct an in-depth exploration of preference learning for reasoning tasks. Our investigation reveals that some well-established preference learning algorithms may be less suitable for reasoning tasks compared to their effectiveness in general conversations. The hypothesis is that in reasoning tasks, the space of correct answers is much smaller than that of incorrect ones, so it is necessary to explicitly increase the reward of chosen data. Therefore, in addition to increasing the reward margin as many preference learning algorithms do, the absolute values of positive responses’ rewards should be positive and may serve as a proxy for performance. Inspired by this, we derive a novel reward modeling objective and empirically that it leads to a stable reward modeling curve and better performance. Together with ULTRAINTERACT, we obtain a strong reward model.
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding 0002, Xingyao Wang 0002, Boji Shan, Zeyuan Liu, Ruobing Xie, Yankai Lin 0001, Zhenghao Liu 0001, Bowen Zhou 0002, Hao Peng 0015, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR5
2025 SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
abstract
Software engineering (SE) is increasingly collaborative, with developers working together on shared complex codebases. Effective collaboration in shared environments requires participants—whether humans or AI agents—to stay on the same page as their environment evolves. When a collaborator’s understanding diverges from the current state—what we term the out-of-sync challenge—the collaborator’s actions may fail, leading to integration issues. In this work, we introduce SyncMind, a framework that systematically defines the out-of-sync problem faced by large language model (LLM) agents in collaborative software engineering (CSE). Based on SyncMind, we create SyncBench, a benchmark featuring 24,332 instances of agent out-of-sync scenarios in real-world CSE derived from 21 popular GitHub repositories with executable verification tests. Experiments on SyncBench uncover critical insights into existing LLM agents’ capabilities and limitations. Besides substantial performance gaps among agents (from Llama-3.1 agents $\leq 3.33%$ to Claude-3.5-Sonnet $\geq 28.18%$), their consistently low collaboration willingness ($\le 4.86%$) suggests fundamental limitations of existing LLM in CSE. However, when collaboration occurs, it positively correlates with out-of-sync recovery success. Minimal performance differences in agents’ resource-aware out-of-sync recoveries further reveal their significant lack of resource awareness and adaptability, shedding light on future development of resource-efficient collaborative systems. Our code and data are openly available on our project website: https://xhguo7.github.io/SyncMind/.
Xuehang Guo, Xingyao Wang 0002, Yangyi Chen, Chi Han, Manling Li, Heng Ji 0001
ICML2
2025 Training Software Engineering Agents and Verifiers with SWE-Gym
abstract
We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories.
Xingyao Wang 0002, Graham Neubig, Navdeep Jaitly, Heng Ji 0001, Alane Suhr, Yizhe Zhang 0002
ICML2
2024 SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
abstract
Large language models (LLMs) often generate inaccurate or fabricated information and generally fail to indicate their confidence, which limits their broader applications.Previous work has elicited confidence from LLMs by direct or self-consistency prompting, or constructing specific datasets for supervised finetuning.The prompting-based approaches have inferior performance, and the training-based approaches are limited to binary or inaccurate group-level confidence estimates.In this work, we present SaySelf, a novel training framework that teaches LLMs to express more fine-grained confidence estimates.In addition, beyond the confidence scores, SaySelf initiates the process of directing LLMs to produce selfreflective rationales that clearly identify gaps in their parametric knowledge and explain their uncertainty.This is achieved by using an LLM to automatically summarize the uncertainties in specific knowledge via natural language.The summarization is based on the analysis of the inconsistency in multiple sampled reasoning chains, and the resulting data is utilized for supervised fine-tuning.Moreover, we utilize reinforcement learning with a meticulously crafted reward function to calibrate the confidence estimates, motivating LLMs to deliver accurate, high-confidence predictions and to penalize overconfidence in erroneous outputs.Experimental results demonstrate the effectiveness of SaySelf in reducing the confidence calibration error and maintaining the task performance.The generated self-reflective rationales are also reasonable and can further contribute to the calibration.The code is made public at https://github.com/xu1868/SaySelf. Direct Prompting / Group-based Calibration Training Self-Consistency Prompting Previous WorkWhat is the name of the younger son of the current President of the United States?Robert Hunter Biden.My overall confidence is 3. Robert Hunter Biden. According to my knowledge, there is a slight possibility that the current President is Trump.My overall confidence is 8.
Shujin Wu, Shizhe Diao, Xiaoze Liu, Xingyao Wang 0002, Yangyi Chen, Jing Gao 0004
EMNLP5
2024 MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
abstract
To solve complex tasks, large language models (LLMs) often require multiple rounds of interactions with the user, sometimes assisted by external tools. However, current evaluation protocols often emphasize benchmark performance with single-turn exchanges, neglecting the nuanced interactions among the user, LLMs, and external tools, while also underestimating the importance of natural language feedback from users. These oversights contribute to discrepancies between research benchmark evaluations and real-world use cases. We introduce MINT, a benchmark that evaluates LLMs' ability to solve tasks with multi-turn interactions by (1) using tools and (2) leveraging natural language feedback. To ensure reproducibility, we provide an evaluation framework where LLMs can access tools by executing Python code and receive users' natural language feedback simulated by GPT-4. We repurpose a diverse set of established evaluation datasets focusing on reasoning, coding, and decision-making and carefully curate them into a compact subset for efficient evaluation. Our analysis of 20 open- and closed-source LLMs offers intriguing findings. (a) LLMs generally benefit from tools and language feedback, with performance gains (absolute, same below) of 1--8% for each turn of tool use and 2--17% with natural language feedback. (b) Better single-turn performance does not guarantee better multi-turn performance. (c) Surprisingly, on the LLMs evaluated, supervised instruction-finetuning (SIFT) and reinforcement learning from human feedback (RLHF) generally hurt multi-turn capabilities. We expect MINT can help measure progress and incentivize research in improving LLMs' capabilities in multi-turn interactions, especially for open-source communities where multi-turn human evaluation can be less accessible compared to commercial LLMs with a larger user base.
Xingyao Wang 0002, Zihan Wang 0010, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng 0009, Heng Ji 0001
ICLR1
2024 CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
abstract
Large language models (LLMs) are often augmented with tools to solve complex tasks. By generating code snippets and executing them through task-specific Application Programming Interfaces (APIs), they can offload certain functions to dedicated external modules, such as image encoding and performing calculations. However, most existing approaches to augment LLMs with tools are constrained by general-purpose APIs and lack the flexibility for tailoring them to specific tasks. In this work, we present CRAFT, a general tool creation and retrieval framework for LLMs. It creates toolsets specifically curated for the tasks and equips LLMs with a component that retrieves tools from these sets to enhance their capability to solve complex tasks. For each task, we collect specific code solutions by prompting GPT-4 to solve the training examples. Following a validation step ensuring the correctness, these solutions are abstracted into code snippets to enhance reusability, and deduplicated for higher quality. At inference time, the language model retrieves snippets from the toolsets and then executes them or generates the output conditioning on the retrieved snippets. Our method is designed to be flexible and offers a plug-and-play approach to adapt off-the-shelf LLMs to unseen domains and modalities, without any finetuning. Experiments on vision-language, tabular processing, and mathematical reasoning tasks show that our approach achieves substantial improvements compared to strong baselines. In addition, our in-depth analysis reveals that: (1) consistent performance improvement can be achieved by scaling up the number of tools and the capability of the backbone models; (2) each component of our approach contributes to the performance gains; (3) the created tools are well-structured and reliable with low complexity and atomicity.
Lifan Yuan, Yangyi Chen, Xingyao Wang 0002, Yi R. Fung 0001, Hao Peng 0009, Heng Ji 0001
ICLR3
2024 Executable Code Actions Elicit Better LLM Agents
abstract
Large Language Model (LLM) agents, capable of performing a broad range of actions, such as invoking tools and controlling robots, show great potential in tackling real-world challenges. LLM agents are typically prompted to produce actions by generating JSON or text in a pre-defined format, which is usually limited by constrained action space (e.g., the scope of pre-defined tools) and restricted flexibility (e.g., inability to compose multiple tools). This work proposes to use executable Python code to consolidate LLM agents’ actions into a unified action space (CodeAct). Integrated with a Python interpreter, CodeAct can execute code actions and dynamically revise prior actions or emit new actions upon new observations through multi-turn interactions. Our extensive analysis of 17 LLMs on API-Bank and a newly curated benchmark shows that CodeAct outperforms widely used alternatives (up to 20% higher success rate). The encouraging performance of CodeAct motivates us to build an open-source LLM agent that interacts with environments by executing interpretable code and collaborates with users using natural language. To this end, we collect an instruction-tuning dataset CodeActInstruct that consists of 7k multi-turn interactions using CodeAct. We show that it can be used with existing data to improve models in agent-oriented tasks without compromising their general capability. CodeActAgent, finetuned from Llama2 and Mistral, is integrated with Python interpreter and uniquely tailored to perform sophisticated tasks (e.g., model training) using existing libraries and autonomously self-debug.
Xingyao Wang 0002, Yangyi Chen, Lifan Yuan, Yizhe Zhang 0002, Yunzhu Li, Hao Peng 0009, Heng Ji 0001
ICML1
2024 R-Tuning: Instructing Large Language Models to Say 'I Don't Know'
abstract
Hanning Zhang, Shizhe Diao, Yong Lin, Yi Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, Tong Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hanning Zhang, Shizhe Diao, Yi R. Fung 0001, Qing Lian, Xingyao Wang 0002, Yangyi Chen, Heng Ji 0001, Tong Zhang 0001
NAACL-HLT6
2023 Code4Struct: Code Generation for Few-Shot Event Structure Prediction
abstract
Large Language Model (LLM) trained on a mixture of text and code has demonstrated impressive capability in translating natural language (NL) into structured code.We observe that semantic structures can be conveniently translated into code and propose CODE4STRUCT to leverage such text-tostructure translation capability to tackle structured prediction tasks.As a case study, we formulate Event Argument Extraction (EAE) as converting text into event-argument structures that can be represented as a class object using code.This alignment between structures and code enables us to take advantage of Programming Language (PL) features such as inheritance 1 and type annotation 2 to introduce external knowledge or add constraints.We show that, with sufficient in-context examples, formulating EAE as a code generation problem is advantageous over using variants of text-based prompts.Despite only using 20 training event instances for each event type, CODE4STRUCT is comparable to supervised models trained on 4,202 instances and outperforms current stateof-the-art (SOTA) trained on 20-shot data by 29.5% absolute F1.By leveraging the inheritance feature of PL, CODE4STRUCT can use 10-shot training data from a sibling event type to predict arguments for zero-resource event types and outperforms the zero-shot baseline by 12% absolute F1. 3 Event Argument Extraction Programming Language (Python) Event / Entity Type Transport, VEH Class definition class Transport, class VEH Hierarchical Event Ontology Movement:Transport Inheritance Inheritance is a way to create a hierarchy of classes in PL.A child class can base upon another class, retaining similar implementation.class Transport(Movement) Event Arguments vehicle Function arguments def function(vehicle=...) Argument ConstraintEach argument can has a list of multiple entities; Argument vehicle should be entities of type VEH. Type Annotation & Argument Default ValueType annotations are used by developers to indicate the data types of variables and input/outputs of functions.If a function is called without the argument, the argument gets its default value (a list in this case). def function( vehicle: List[VEH] = [], … ) Weakly-supervised InformationTransport Event describes someone transporting something in a vehicle from one place to another place. Docstring or Commentsclass Transport(Movement):""" self.agenttransported self.artifact in self.vehiclevehicle from self.origin place to self.destination place."""
Xingyao Wang 0002, Heng Ji 0001
ACL (1)1
2023 ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation
abstract
State-of-the-art vision-language models (VLMs) still have limited performance in structural knowledge extraction, such as relations between objects.In this work, we present ViStruct, a training framework to learn VLMs for effective visual structural knowledge extraction.Two novel designs are incorporated.First, we propose to leverage the inherent structure of programming language to depict visual structural information.This approach enables explicit and consistent representation of visual structural information of multiple granularities, such as concepts, relations, and events, in a well-organized structured format.Second, we introduce curriculum-based learning for VLMs to progressively comprehend visual structures, from fundamental visual concepts to intricate event structures.Our intuition is that lower-level knowledge may contribute to complex visual structure understanding.Furthermore, we compile and release a collection of datasets tailored for visual structural knowledge extraction.We adopt a weakly-supervised approach to directly generate visual event structures from captions for ViStruct training, capitalizing on abundant image-caption pairs from the web.In experiments, we evaluate ViStruct on visual structure prediction tasks, demonstrating its effectiveness in improving the understanding of visual structures.The code is public at https://github.com/ Yangyi-Chen/vi-struct.
Yangyi Chen, Xingyao Wang 0002, Manling Li, Derek Hoiem, Heng Ji 0001
EMNLP2