VLDB 2026 Research / reviewers in the wild / expert
John Yang 0002
dblp:177/0934-2
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Program synthesis and code generation · 40% Software maintenance and evolution · 19% Software testing · 17% | |
| Artificial intelligence
4 papers |
Language models and text generation · 70% Reinforcement learning · 22% Vision and language · 9% | |
| Network and information security
1 paper |
Systems and software security · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 16 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
code generation |
1.0 | 2 | 2024 | SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024 InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023 |
Systems and software security
exploitation |
0.9 | 1 | 2025 | EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities · ICML 2025 |
Systems and software security
vulnerability discovery |
0.9 | 1 | 2025 | EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities · ICML 2025 |
Debugging and program repair
automated program repair |
0.9 | 1 | 2025 | SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025 |
Program synthesis and code generation › code generation with language models
software engineering agents |
0.9 | 1 | 2025 | SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025 |
Empirical software engineering › benchmarking
software engineering benchmarks |
0.9 | 1 | 2025 | SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025 |
Software testing
test generation |
0.9 | 1 | 2025 | SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025 |
Software maintenance and evolution › issue management
issue resolution |
0.8 | 1 | 2024 | SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024 |
Machine learning › Reinforcement learning › reinforcement learning environment
environment design |
0.7 | 1 | 2023 | InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023 |
Program synthesis and code generation › interactive program synthesis
interactive code generation |
0.7 | 1 | 2023 | InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023 |
Natural language and speech › Language models and text generation › LLM agents
web navigation |
0.6 | 1 | 2022 | WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents · NeurIPS 2022 |
Information retrieval
e-commerce search |
0.6 | 1 | 2022 | WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents · NeurIPS 2022 |
Empirical software engineering
mining software repositories |
0.3 | 1 | 2025 | SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025 |
Software testing › test execution
automated test execution |
0.2 | 1 | 2024 | SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering · NeurIPS 2024 |
Software testing › test infrastructure
benchmark construction |
0.2 | 1 | 2024 | SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.5large language model · 2.4language model agent · 1.6execution environment interaction · 1.5prompting · 1.3pre-trained language model · 1.1imitation learning · 1.1large language model agents · 0.9large language model agent · 0.9interactive terminal tooling · 0.9execution environments · 0.9agent-computer interface · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case StudyabstractRecent advancements in large language models (LLMs) have significantly enhanced their coding capabilities. However, existing benchmarks predominantly focused on simplified or isolated aspects of coding, such as single-file code generation or repository issue debugging, falling short of measuring the full spectrum of challenges raised by real-world programming activities. In this case study, we explore the performance of LLMs across the entire software development lifecycle with DevEval, encompassing stages including software design, environment setup, implementation, acceptance testing, and unit testing. DevEval features four programming languages, multiple domains, high-quality data collection, and carefully designed and verified metrics for each task. Empirical studies show that current LLMs, including GPT-4, fail to solve the challenges presented within DevEval. Our findings offer actionable insights for the future development of LLMs toward real-world programming applications. Bowen Li 0002, Ziwei Tang, John Yang 0002, Jinyang Li 0003, Shunyu Yao 0006, Chen Qian 0006, Binyuan Hui, Qicheng Zhang, Zhiyin Yu, He Du, Dahua Lin, Chao Peng 0002, Kai Chen 0026 |
COLING | 5 |
| 2025 | SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?abstractAutonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, with problem statements presented predominantly as text and lacking visual elements such as images. This limited coverage motivates our inquiry into how existing systems might perform on unrepresented software engineering domains
(e.g., front-end, game development, DevOps), which use different programming languages and paradigms. Therefore, we propose SWE-bench Multimodal (SWE-bench M), to evaluate systems on their ability to fix bugs in visual, user-facing JavaScript software. SWE-bench M features 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. Each SWE-bench M task instance contains at least one image in its problem statement or unit tests. Our analysis finds that top-performing SWE-bench systems struggle with SWE-bench M, revealing limitations in visual problem-solving and cross-language generalization. Lastly, we show that SWE-agent’s flexible language-agnostic features enable it to substantially outperform alternatives on SWE-bench M, resolving 12% of task
instances compared to 6% for the next best system. John Yang 0002, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik Narasimhan, Diyi Yang, Sida I. Wang, Ofir Press |
ICLR | 1 |
| 2025 | EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security VulnerabilitiesabstractAlthough language model (LM) agents have demonstrated increased performance in multiple domains, including coding and web-browsing, their success in cybersecurity has been limited. We present *EnIGMA*, an LM agent for autonomously solving Capture The Flag (CTF) challenges. We introduce new tools and interfaces to improve the agent's ability to find and exploit security vulnerabilities, focusing on interactive terminal programs. These novel *Interactive Agent Tools* enable LM agents, for the first time, to run interactive utilities, such as a debugger and a server connection tool, which are essential for solving these challenges.
Empirical analysis on 390 CTF challenges across four benchmarks demonstrate that these new tools and interfaces substantially improve our agent's performance, achieving state-of-the-art results on NYU CTF, Intercode-CTF, and CyBench. Finally, we analyze data leakage, developing new methods to quantify it and identifying a new phenomenon we term *soliloquizing*, where the model self-generates hallucinated observations without interacting with the environment. Talor Abramovich, Meet Udeshi, Minghao Shao, Kilian Lieret, Haoran Xi, Kimberly Milner, Sofija Jancheska, John Yang 0002, Carlos E. Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique 0001, Karthik Narasimhan, Ramesh Karri, Ofir Press |
ICML | 8 |
| 2025 | SWE-smith: Scaling Data for Software Engineering AgentsabstractDespite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point.Existing datasets are small, with at most 1,000s of training instances from 11 or fewer GitHub repositories.The procedures to curate such datasets are often complex, necessitating hundreds of hours of human labor; companion execution environments also take up several terabytes of storage, severely limiting their scalability and usability.To address this pain point, we introduce SWE-smith, a novel pipeline for generating software engineering training data at scale.Given any Python codebase, SWE-smith constructs a corresponding execution environment, then automatically synthesizes 100s to 1,000s of task instances that break existing test(s) in the codebase.Using SWE-smith, we create a dataset of 50k instances sourced from 128 GitHub repositories, an order of magnitude larger than all previous works.We train SWE-agent-LM-32B, achieving 40.2% Pass@1 resolve rate on the SWE-bench Verified benchmark, state of the art among open source models.We open source SWE-smith (collection procedure, task instances, trajectories, models) to lower the barrier of entry for research in LM systems for automated software engineering.All assets available at \url{https://swesmith.com}. John Yang 0002, Kilian Lieret, Carlos E. Jimenez, Alexander Wettig, Kabir Khandpur, Binyuan Hui, Ofir Press, Ludwig Schmidt, Diyi Yang |
NeurIPS | 1 |
| 2024 | Disentangled Prompt Learning for Transferable, Multimodal, Few-Shot Image ClassificationabstractExisting prompting strategies for adapting pretrained vision language models to the downstream task of finegrained attribute classification learn visual variance in a class-specific manner. We present DisPoL, a method for learning disentangled representations that improves the transferability and performance of continuous prompts for downstream classification tasks. Our method decomposes a prompt into separate sub-prompts, then performs late fusion of the corresponding output embeddings in a novel manner. We combine the fixed embedding of a static, context-constraining object sub-prompt and the tunable embedding of a soft sub-prompt for a task-specific attribute using self-attention. By avoiding joint learning of these tokens, the resulting disentangled prompt embeddings are more transferable to unseen objects. We also demonstrate how to use hand-crafted templates to initialize the task-specific soft prompt, improving training efficiency. Through extensive experiments, we show that DisPoL exceeds the performance of existing methods in few-shot settings and highlight its contribution as a parameter-efficient fine-tuning method. John Yang 0002, Alessandro Magnani, Binwei Yang |
IEEE Big Data | 1 |
| 2024 | SWE-bench: Can Language Models Resolve Real-world Github Issues?abstractLanguage models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere 1.96% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous. Carlos E. Jimenez, John Yang 0002, Alexander Wettig, Shunyu Yao 0006, Kexin Pei, Ofir Press, Karthik Narasimhan |
ICLR | 2 |
| 2024 | SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringabstractLanguage model agents are increasingly being used to automate complicated tasks in digital environments. Just as humans benefit from powerful software applications, such as integrated development environments, for complex tasks like software engineering, we posit that language model agents represent a new category of end users with their own needs and abilities, and would benefit from specially built interfaces to the software they use. We investigate how the role of interface design affects the performance of language model agents. As a result of this exploration, we introduce SWE-agent: a system that facilitates language model agents to autonomously use computers to solve software engineering tasks. SWE-agent's custom agent-computer interface significantly enhances an agent's ability to create and edit code files, navigate entire repositories, and execute tests and other programs. We evaluate SWE-agent on SWE-bench and HumanEvalFix, achieving state-of-the-art performance on both with a pass@1 rate of 12.5% and 87.7%, respectively, far exceeding the previous state-of-the-art achieved with non-interactive language models. Finally, we provide insight on how the design of the agent-computer interface can impact agents' behavior and performance. John Yang 0002, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao 0006, Karthik Narasimhan, Ofir Press |
NeurIPS | 1 |
| 2023 | InterCode: Standardizing and Benchmarking Interactive Coding with Execution FeedbackabstractHumans write code in a fundamentally interactive manner and rely on constant execution feedback to correct errors, resolve ambiguities, and decompose tasks. While LLMs have recently exhibited promising coding capabilities, current coding benchmarks mostly consider a static instruction-to-code sequence transduction process, which has the potential for error propagation and a disconnect between the generated code and its final execution environment. To address this gap, we introduce InterCode, a lightweight, flexible, and easy-to-use framework of interactive coding as a standard reinforcement learning (RL) environment, with code as actions and execution feedback as observations. Our framework is language and platform agnostic, uses self-contained Docker environments to provide safe and reproducible execution, and is compatible out-of-the-box with traditional seq2seq coding methods, while enabling the development of new methods for interactive code generation. We use InterCode to create three interactive code environments with Bash, SQL, and Python as action spaces, leveraging data from the static NL2Bash, Spider, and MBPP datasets. We demonstrate InterCode’s viability as a testbed by evaluating multiple state-of-the-art LLMs configured with different prompting strategies such as ReAct and Plan & Solve. Our results showcase the benefits of interactive code generation and demonstrate that InterCode can serve as a challenging benchmark for advancing code understanding and generation capabilities. InterCode is designed to be easily extensible and can even be used to create new tasks such as Capture the Flag, a popular coding puzzle that is inherently multi-step and involves multiple programming languages. John Yang 0002, Akshara Prabhakar, Karthik Narasimhan, Shunyu Yao 0006 |
NeurIPS | 1 |
| 2022 | WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsabstractMost existing benchmarks for grounding language in interactive environments either lack realistic linguistic elements, or prove difficult to scale up due to substantial human involvement in the collection of data or feedback signals. We develop WebShop – a simulated e-commerce website environment with 1.18 million real-world products and 12,087 crowd-sourced text instructions. In this environment, an agent needs to navigate multiple types of webpages and issue diverse actions to find, customize, and purchase a product given an instruction. WebShop provides several challenges including understanding compositional instructions, query (re-)formulation, dealing with noisy text in webpages, and performing strategic exploration. We collect over 1,600 human trajectories to first validate the benchmark, then train and evaluate a diverse range of agents using reinforcement learning, imitation learning, and pre-trained image and language models. Our best model achieves a task success rate of 29%, which significantly outperforms rule heuristics but is far lower than expert human performance (59%). We also analyze agent and human trajectories and ablate various model components to provide insights for developing future agents with stronger language understanding and decision making abilities. Finally, we show our agent trained on WebShop exhibits non-trivial sim-to-real transfer when evaluated on amazon.com and ebay.com, indicating the potential value of our benchmark for developing practical web agents that can operate in the wild. Shunyu Yao 0006, Howard Chen 0003, John Yang 0002, Karthik Narasimhan |
NeurIPS | 3 |