EDBT 2026 Demo / reviewers in the wild / expert
Shunyu Yao 0006
dblp:156/1038-6
· DBLP profile ↗
21ranked-venue papers
10as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contextual Experience Replay for Self-Improvement of Language AgentsabstractLarge language model (LLM) agents have been applied to sequential decision-making tasks such as web navigation, but without any environment-specific experiences, they often fail in these complex tasks.Moreover, current LLM agents are not designed to continually learn from past experiences during inference time, which could be crucial for them to gain these environment-specific experiences.To address this, we propose Contextual Experience Replay (CER), a training-free framework to enable efficient self-improvement for language agents in their context window.Specifically, CER accumulates and synthesizes past experiences into a dynamic memory buffer.These experiences encompass environment dynamics and common decision-making patterns, allowing the agents to retrieve and augment themselves with relevant knowledge in new tasks, enhancing their adaptability in complex environments.We evaluate CER on the challenging WEBARENA and VISUALWEBARENA benchmarks.On VISUALWEBARENA, CER achieves competitive performance of 31.9%.On WEBARENA, CER also gets a competitive average success rate of 36.7%, relatively improving the success rate of the GPT-4o agent baseline by 51.0%.We also conduct a comprehensive analysis on it to prove its efficiency, validity and understand it better. Yitao Liu, Chenglei Si, Karthik Narasimhan, Shunyu Yao 0006 |
ACL (1) | 4 |
| 2025 | Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case StudyabstractRecent advancements in large language models (LLMs) have significantly enhanced their coding capabilities. However, existing benchmarks predominantly focused on simplified or isolated aspects of coding, such as single-file code generation or repository issue debugging, falling short of measuring the full spectrum of challenges raised by real-world programming activities. In this case study, we explore the performance of LLMs across the entire software development lifecycle with DevEval, encompassing stages including software design, environment setup, implementation, acceptance testing, and unit testing. DevEval features four programming languages, multiple domains, high-quality data collection, and carefully designed and verified metrics for each task. Empirical studies show that current LLMs, including GPT-4, fail to solve the challenges presented within DevEval. Our findings offer actionable insights for the future development of LLMs toward real-world programming applications. Bowen Li 0002, Ziwei Tang, John Yang 0002, Jinyang Li 0003, Shunyu Yao 0006, Chen Qian 0006, Binyuan Hui, Qicheng Zhang, Zhiyin Yu, He Du, Dahua Lin, Chao Peng 0002, Kai Chen 0026 |
COLING | 7 |
| 2025 | {τ}-bench: A Benchmark for \underline{T}ool-\underline{A}gent-\underline{U}ser Interaction in Real-World Domains
Shunyu Yao 0006, Noah Shinn, Pedram Razavi, Karthik Narasimhan |
ICLR | 1 |
| 2025 | When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI CollaborationabstractAs large language models (LLMs) increasingly serve as close collaborators for humans, it is crucial that they express their reasoning in ways that humans can understand and learn from. However, this capability remains relatively less understood and under-evaluated. To address this, we introduce a conceptual framework for such Human-AI knowledge transfer capabilities and conduct the first large-scale user study (N=118) explicitly designed to measure it. In our two-phase setup, humans first ideate with an LLM on problem-solving strategies, then independently implement solutions, isolating the influence of model reasoning on human understanding. Our findings reveal that while model benchmark performance correlates with collaborative outcomes, this relationship is notably inconsistent with significant outliers, highlighting that knowledge transfer is a distinct capability requiring dedicated optimization. Our analysis uncovers behavioral and strategic factors that mediate successful knowledge transfer, and we release our code, dataset, and evaluation framework to support future work on communicatively aligned models. Carlos E. Jimenez, Shunyu Yao 0006, Nick Haber, Diyi Yang, Karthik Narasimhan |
NeurIPS | 3 |
| 2024 | SWE-bench: Can Language Models Resolve Real-world Github Issues?abstractLanguage models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere 1.96% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous. Carlos E. Jimenez, John Yang 0002, Alexander Wettig, Shunyu Yao 0006, Kexin Pei, Ofir Press, Karthik Narasimhan |
ICLR | 4 |
| 2024 | COLLIE: Systematic Construction of Constrained Text Generation TasksabstractText generation under constraints have seen increasing interests in natural language processing, especially with the rapidly improving capabilities of large language models. However, existing benchmarks for constrained generation usually focus on fixed constraint types (e.g. generate a sentence containing certain words) that have proved to be easy for state-of-the-art models like GPT-4. We present COLLIE, a grammar-based framework that allows the specification of rich, compositional constraints with diverse generation levels (word, sentence, paragraph, passage) and modeling challenges (e.g. language understanding, logical reasoning, counting, semantic planning). We also develop tools for automatic extraction of task instances given a constraint structure and a raw text corpus. Using COLLIE, we compile the COLLIE-v1 dataset with 1,132 instances comprising 13 constraint structures. We perform systematic experiments across five state-of-the-art instruction-tuned language models and analyze their performances to reveal shortcomings. COLLIE is designed to be extensible and lightweight, and we hope the community finds it useful to develop more complex constraints and evaluations in the future. Shunyu Yao 0006, Howard Chen 0003, Austin W. Hanjie, Runzhe Yang, Karthik Narasimhan |
ICLR | 1 |
| 2024 | SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringabstractLanguage model agents are increasingly being used to automate complicated tasks in digital environments. Just as humans benefit from powerful software applications, such as integrated development environments, for complex tasks like software engineering, we posit that language model agents represent a new category of end users with their own needs and abilities, and would benefit from specially built interfaces to the software they use. We investigate how the role of interface design affects the performance of language model agents. As a result of this exploration, we introduce SWE-agent: a system that facilitates language model agents to autonomously use computers to solve software engineering tasks. SWE-agent's custom agent-computer interface significantly enhances an agent's ability to create and edit code files, navigate entire repositories, and execute tests and other programs. We evaluate SWE-agent on SWE-bench and HumanEvalFix, achieving state-of-the-art performance on both with a pass@1 rate of 12.5% and 87.7%, respectively, far exceeding the previous state-of-the-art achieved with non-interactive language models. Finally, we provide insight on how the design of the agent-computer interface can impact agents' behavior and performance. John Yang 0002, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao 0006, Karthik Narasimhan, Ofir Press |
NeurIPS | 5 |
| 2023 | EC2: Emergent Communication for Embodied ControlabstractEmbodied control requires agents to leverage multimodal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language instructions support generalization with abstract, symbolic structures. While recent approaches apply contrastive learning to force alignment between the two modalities, we hypothesize better modeling their complementary differences can lead to more holistic representations for downstream adaption. To this end, we propose Emergent Communication for Embodied Control (EC2), a novel scheme to pre-train video-language representations for few-shot embodied control. The key idea is to learn an unsupervised “language” of videos via emergent communication, which bridges the semantics of video details and structures of natural language. We learn embodied representations of video trajectories, emergent language, and natural language using a language model, which is then used to finetune a lightweight policy network for downstream control. Through extensive experiments in Metaworld and Franka Kitchen embodied benchmarks, EC2is shown to consistently outperform previous contrastive learning methods for both videos and texts as task inputs. Further ablations confirm the importance of the emergent language, which is beneficial for both video and language learning, and significantly superior to using pre-trained video captions. We also present a quantitative and qualitative analysis of the emergent language and discuss future directions toward better understanding and leveraging emergent communication in embodied tasks. Yao Mu 0001, Shunyu Yao 0006, Mingyu Ding, Ping Luo 0002, Chuang Gan 0001 |
CVPR | 2 |
| 2023 | ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao 0006, Jeffrey Zhao, Nan Du 0002, Izhak Shafran, Karthik Narasimhan, Yuan Cao 0007 |
ICLR | 1 |
| 2023 | Reflexion: language agents with verbal reinforcement learningabstractLarge language models (LLMs) have been increasingly used to interact with external environments (e.g., games, compilers, APIs) as goal-driven agents. However, it remains challenging for these language agents to quickly and efficiently learn from trial-and-error as traditional reinforcement learning methods require extensive training samples and expensive model fine-tuning. We propose \emph{Reflexion}, a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback. Concretely, Reflexion agents verbally reflect on task feedback signals, then maintain their own reflective text in an episodic memory buffer to induce better decision-making in subsequent trials. Reflexion is flexible enough to incorporate various types (scalar values or free-form language) and sources (external or internally simulated) of feedback signals, and obtains significant improvements over a baseline agent across diverse tasks (sequential decision-making, coding, language reasoning). For example, Reflexion achieves a 91\% pass@1 accuracy on the HumanEval coding benchmark, surpassing the previous state-of-the-art GPT-4 that achieves 80\%. We also conduct ablation and analysis studies using different feedback signals, feedback incorporation methods, and agent types, and provide insights into how they affect performance. We release all code, demos, and datasets at \url{https://github.com/noahshinn024/reflexion}. Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao 0006 |
NeurIPS | 5 |
| 2023 | InterCode: Standardizing and Benchmarking Interactive Coding with Execution FeedbackabstractHumans write code in a fundamentally interactive manner and rely on constant execution feedback to correct errors, resolve ambiguities, and decompose tasks. While LLMs have recently exhibited promising coding capabilities, current coding benchmarks mostly consider a static instruction-to-code sequence transduction process, which has the potential for error propagation and a disconnect between the generated code and its final execution environment. To address this gap, we introduce InterCode, a lightweight, flexible, and easy-to-use framework of interactive coding as a standard reinforcement learning (RL) environment, with code as actions and execution feedback as observations. Our framework is language and platform agnostic, uses self-contained Docker environments to provide safe and reproducible execution, and is compatible out-of-the-box with traditional seq2seq coding methods, while enabling the development of new methods for interactive code generation. We use InterCode to create three interactive code environments with Bash, SQL, and Python as action spaces, leveraging data from the static NL2Bash, Spider, and MBPP datasets. We demonstrate InterCode’s viability as a testbed by evaluating multiple state-of-the-art LLMs configured with different prompting strategies such as ReAct and Plan & Solve. Our results showcase the benefits of interactive code generation and demonstrate that InterCode can serve as a challenging benchmark for advancing code understanding and generation capabilities. InterCode is designed to be easily extensible and can even be used to create new tasks such as Capture the Flag, a popular coding puzzle that is inherently multi-step and involves multiple programming languages. John Yang 0002, Akshara Prabhakar, Karthik Narasimhan, Shunyu Yao 0006 |
NeurIPS | 4 |
| 2023 | Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsabstractLanguage models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices.
Our experiments show that ToT significantly enhances language models’ problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords. For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4\% of tasks, our method achieved a success rate of 74\%. Code repo with all prompts: https://github.com/princeton-nlp/tree-of-thought-llm. Shunyu Yao 0006, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths 0001, Yuan Cao 0007, Karthik Narasimhan |
NeurIPS | 1 |
| 2022 | Multi-Stage Episodic Control for Strategic Exploration in Text Games
Jens Tuyls, Shunyu Yao 0006, Sham M. Kakade, Karthik Narasimhan |
ICLR | 2 |
| 2022 | Linking Emergent and Natural Languages via Corpus Transfer
Shunyu Yao 0006, Mo Yu, Yang Zhang 0001, Karthik Narasimhan, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 1 |
| 2022 | WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsabstractMost existing benchmarks for grounding language in interactive environments either lack realistic linguistic elements, or prove difficult to scale up due to substantial human involvement in the collection of data or feedback signals. We develop WebShop – a simulated e-commerce website environment with 1.18 million real-world products and 12,087 crowd-sourced text instructions. In this environment, an agent needs to navigate multiple types of webpages and issue diverse actions to find, customize, and purchase a product given an instruction. WebShop provides several challenges including understanding compositional instructions, query (re-)formulation, dealing with noisy text in webpages, and performing strategic exploration. We collect over 1,600 human trajectories to first validate the benchmark, then train and evaluate a diverse range of agents using reinforcement learning, imitation learning, and pre-trained image and language models. Our best model achieves a task success rate of 29%, which significantly outperforms rule heuristics but is far lower than expert human performance (59%). We also analyze agent and human trajectories and ablate various model components to provide insights for developing future agents with stronger language understanding and decision making abilities. Finally, we show our agent trained on WebShop exhibits non-trivial sim-to-real transfer when evaluated on amazon.com and ebay.com, indicating the potential value of our benchmark for developing practical web agents that can operate in the wild. Shunyu Yao 0006, Howard Chen 0003, John Yang 0002, Karthik Narasimhan |
NeurIPS | 1 |
| 2021 | Self-Attention Networks Can Process Bounded Hierarchical LanguagesabstractShunyu Yao, Binghui Peng, Christos Papadimitriou, Karthik Narasimhan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shunyu Yao 0006, Binghui Peng, Christos H. Papadimitriou, Karthik Narasimhan |
ACL/IJCNLP (1) | 1 |
| 2021 | Reading and Acting while Blindfolded: The Need for Semantics in Text Game AgentsabstractShunyu Yao, Karthik Narasimhan, Matthew Hausknecht. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shunyu Yao 0006, Karthik Narasimhan, Matthew J. Hausknecht |
NAACL-HLT | 1 |
| 2020 | The fine structure of surprise in intuitive physics: when, why, and how much?
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
CogSci | 3 |
| 2020 | Keep CALM and Explore: Language Models for Action Generation in Text-based GamesabstractText-based games present a unique challenge for autonomous agents to operate in natural language and handle enormous action spaces.In this paper, we propose the Contextual Action Language Model (CALM) to generate a compact set of action candidates at each game state.Our key insight is to train language models on human gameplay, where people demonstrate linguistic priors and a general game sense for promising actions conditioned on game history.We combine CALM with a reinforcement learning agent which re-ranks the generated action candidates to maximize ingame rewards.We evaluate our approach using the Jericho benchmark (Hausknecht et al., 2019a), on games unseen by CALM during training.Our method obtains a 69% relative improvement in average game score over the previous state-of-the-art model.Surprisingly, on half of these games, CALM is competitive with or better than other models that have access to ground truth admissible actions.* * Code and data are available at https://github. com/princeton-nlp/calm-textgame.Observation: You are in the living room.There is a doorway to the east, a wooden door with strange gothic lettering to the west, which appears to be nailed shut, a trophy case, and a large oriental rug in the center of the room.You are carrying: A brass lantern . . . Shunyu Yao 0006, Rohan Rao, Matthew J. Hausknecht, Karthik Narasimhan |
EMNLP (1) | 1 |
| 2019 | Modeling Expectation Violation in Intuitive Physics with Coarse Probabilistic Object RepresentationsabstractFrom infancy, humans have expectations about how objects will move and interact. Even young children expect objects not to move through one another, teleport, or disappear. They are surprised by mismatches between physical expectations and perceptual observations, even in unfamiliar scenes with completely novel objects. A model that exhibits human-like understanding of physics should be similarly surprised, and adjust its beliefs accordingly. We propose ADEPT, a model that uses a coarse (approximate geometry) object-centric representation for dynamic 3D scene understanding. Inference integrates deep recognition networks, extended probabilistic physical simulation, and particle filtering for forming predictions and expectations across occlusion. We also present a new test set for measuring violations of physical expectations, using a range of scenarios derived from developmental psychology. We systematically compare ADEPT, baseline models, and human expectations on this test set. ADEPT outperforms standard network architectures in discriminating physically implausible scenes, and often performs this discrimination at the same level as people. Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
NeurIPS | 3 |
| 2018 | 3D-Aware Scene Manipulation via Inverse GraphicsabstractWe aim to obtain an interpretable, expressive, and disentangled scene representation that contains comprehensive structural and textural information for each object. Previous scene representations learned by neural networks are often uninterpretable, limited to a single object, or lacking 3D knowledge. In this work, we propose 3D scene de-rendering networks (3D-SDN) to address the above issues by integrating disentangled representations for semantics, geometry, and appearance into a deep generative model. Our scene encoder performs inverse graphics, translating a scene into a structured object-wise representation. Our decoder has two components: a differentiable shape renderer and a neural texture generator. The disentanglement of semantics, geometry, and appearance supports 3D-aware scene manipulation, e.g., rotating and moving objects freely while keeping the consistent shape and texture, and changing the object appearance without affecting its shape. Experiments demonstrate that our editing scheme based on 3D-SDN is superior to its 2D counterpart. Shunyu Yao 0006, Tzu-Ming Harry Hsu, Jun-Yan Zhu, Jiajun Wu 0001, Antonio Torralba 0001, William T. Freeman, Josh Tenenbaum |
NeurIPS | 1 |