John Yang 0002

dblp:177/0934-2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Program synthesis and code generation · 40% Software maintenance and evolution · 19% Software testing · 17%
Artificial intelligence
4 papers
Language models and text generation · 70% Reinforcement learning · 22% Vision and language · 9%
Network and information security
1 paper
Systems and software security · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
code generation
1.022024
SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023
Systems and software security
exploitation
0.912025
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities · ICML 2025
Systems and software security
vulnerability discovery
0.912025
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities · ICML 2025
Debugging and program repair
automated program repair
0.912025
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025
Program synthesis and code generation
code generation with language models
0.912025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Program synthesis and code generation › code generation with language models
software engineering agents
0.912025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Empirical software engineering › benchmarking
software engineering benchmarks
0.912025
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025
Software testing
test generation
0.912025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Software maintenance and evolution › issue management
issue resolution
0.812024
SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024
Machine learning › Reinforcement learning › reinforcement learning environment
environment design
0.712023
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023
Program synthesis and code generation › interactive program synthesis
interactive code generation
0.712023
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023
Natural language and speech › Language models and text generation › LLM agents
web navigation
0.612022
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents · NeurIPS 2022
Information retrieval
e-commerce search
0.612022
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents · NeurIPS 2022
Empirical software engineering
mining software repositories
0.312025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Software testing › test execution
automated test execution
0.212024
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering · NeurIPS 2024
Software testing › test infrastructure
benchmark construction
0.212024
SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.5large language model · 2.4language model agent · 1.6execution environment interaction · 1.5prompting · 1.3pre-trained language model · 1.1imitation learning · 1.1large language model agents · 0.9large language model agent · 0.9interactive terminal tooling · 0.9execution environments · 0.9agent-computer interface · 0.8
YearPublicationVenuePosition
2025 Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
abstract
Recent advancements in large language models (LLMs) have significantly enhanced their coding capabilities. However, existing benchmarks predominantly focused on simplified or isolated aspects of coding, such as single-file code generation or repository issue debugging, falling short of measuring the full spectrum of challenges raised by real-world programming activities. In this case study, we explore the performance of LLMs across the entire software development lifecycle with DevEval, encompassing stages including software design, environment setup, implementation, acceptance testing, and unit testing. DevEval features four programming languages, multiple domains, high-quality data collection, and carefully designed and verified metrics for each task. Empirical studies show that current LLMs, including GPT-4, fail to solve the challenges presented within DevEval. Our findings offer actionable insights for the future development of LLMs toward real-world programming applications.
Bowen Li 0002, Ziwei Tang, John Yang 0002, Jinyang Li 0003, Shunyu Yao 0006, Chen Qian 0006, Binyuan Hui, Qicheng Zhang, Zhiyin Yu, He Du, Dahua Lin, Chao Peng 0002, Kai Chen 0026
COLING5
2025 SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
abstract
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, with problem statements presented predominantly as text and lacking visual elements such as images. This limited coverage motivates our inquiry into how existing systems might perform on unrepresented software engineering domains (e.g., front-end, game development, DevOps), which use different programming languages and paradigms. Therefore, we propose SWE-bench Multimodal (SWE-bench M), to evaluate systems on their ability to fix bugs in visual, user-facing JavaScript software. SWE-bench M features 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. Each SWE-bench M task instance contains at least one image in its problem statement or unit tests. Our analysis finds that top-performing SWE-bench systems struggle with SWE-bench M, revealing limitations in visual problem-solving and cross-language generalization. Lastly, we show that SWE-agent’s flexible language-agnostic features enable it to substantially outperform alternatives on SWE-bench M, resolving 12% of task instances compared to 6% for the next best system.
John Yang 0002, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik Narasimhan, Diyi Yang, Sida I. Wang, Ofir Press
ICLR1
2025 EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
abstract
Although language model (LM) agents have demonstrated increased performance in multiple domains, including coding and web-browsing, their success in cybersecurity has been limited. We present *EnIGMA*, an LM agent for autonomously solving Capture The Flag (CTF) challenges. We introduce new tools and interfaces to improve the agent's ability to find and exploit security vulnerabilities, focusing on interactive terminal programs. These novel *Interactive Agent Tools* enable LM agents, for the first time, to run interactive utilities, such as a debugger and a server connection tool, which are essential for solving these challenges. Empirical analysis on 390 CTF challenges across four benchmarks demonstrate that these new tools and interfaces substantially improve our agent's performance, achieving state-of-the-art results on NYU CTF, Intercode-CTF, and CyBench. Finally, we analyze data leakage, developing new methods to quantify it and identifying a new phenomenon we term *soliloquizing*, where the model self-generates hallucinated observations without interacting with the environment.
Talor Abramovich, Meet Udeshi, Minghao Shao, Kilian Lieret, Haoran Xi, Kimberly Milner, Sofija Jancheska, John Yang 0002, Carlos E. Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique 0001, Karthik Narasimhan, Ramesh Karri, Ofir Press
ICML8
2025 SWE-smith: Scaling Data for Software Engineering Agents
abstract
Despite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point.Existing datasets are small, with at most 1,000s of training instances from 11 or fewer GitHub repositories.The procedures to curate such datasets are often complex, necessitating hundreds of hours of human labor; companion execution environments also take up several terabytes of storage, severely limiting their scalability and usability.To address this pain point, we introduce SWE-smith, a novel pipeline for generating software engineering training data at scale.Given any Python codebase, SWE-smith constructs a corresponding execution environment, then automatically synthesizes 100s to 1,000s of task instances that break existing test(s) in the codebase.Using SWE-smith, we create a dataset of 50k instances sourced from 128 GitHub repositories, an order of magnitude larger than all previous works.We train SWE-agent-LM-32B, achieving 40.2% Pass@1 resolve rate on the SWE-bench Verified benchmark, state of the art among open source models.We open source SWE-smith (collection procedure, task instances, trajectories, models) to lower the barrier of entry for research in LM systems for automated software engineering.All assets available at \url{https://swesmith.com}.
John Yang 0002, Kilian Lieret, Carlos E. Jimenez, Alexander Wettig, Kabir Khandpur, Binyuan Hui, Ofir Press, Ludwig Schmidt, Diyi Yang
NeurIPS1
2024 Disentangled Prompt Learning for Transferable, Multimodal, Few-Shot Image Classification
abstract
Existing prompting strategies for adapting pretrained vision language models to the downstream task of finegrained attribute classification learn visual variance in a class-specific manner. We present DisPoL, a method for learning disentangled representations that improves the transferability and performance of continuous prompts for downstream classification tasks. Our method decomposes a prompt into separate sub-prompts, then performs late fusion of the corresponding output embeddings in a novel manner. We combine the fixed embedding of a static, context-constraining object sub-prompt and the tunable embedding of a soft sub-prompt for a task-specific attribute using self-attention. By avoiding joint learning of these tokens, the resulting disentangled prompt embeddings are more transferable to unseen objects. We also demonstrate how to use hand-crafted templates to initialize the task-specific soft prompt, improving training efficiency. Through extensive experiments, we show that DisPoL exceeds the performance of existing methods in few-shot settings and highlight its contribution as a parameter-efficient fine-tuning method.
John Yang 0002, Alessandro Magnani, Binwei Yang
IEEE Big Data1
2024 SWE-bench: Can Language Models Resolve Real-world Github Issues?
abstract
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere 1.96% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
Carlos E. Jimenez, John Yang 0002, Alexander Wettig, Shunyu Yao 0006, Kexin Pei, Ofir Press, Karthik Narasimhan
ICLR2
2024 SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
abstract
Language model agents are increasingly being used to automate complicated tasks in digital environments. Just as humans benefit from powerful software applications, such as integrated development environments, for complex tasks like software engineering, we posit that language model agents represent a new category of end users with their own needs and abilities, and would benefit from specially built interfaces to the software they use. We investigate how the role of interface design affects the performance of language model agents. As a result of this exploration, we introduce SWE-agent: a system that facilitates language model agents to autonomously use computers to solve software engineering tasks. SWE-agent's custom agent-computer interface significantly enhances an agent's ability to create and edit code files, navigate entire repositories, and execute tests and other programs. We evaluate SWE-agent on SWE-bench and HumanEvalFix, achieving state-of-the-art performance on both with a pass@1 rate of 12.5% and 87.7%, respectively, far exceeding the previous state-of-the-art achieved with non-interactive language models. Finally, we provide insight on how the design of the agent-computer interface can impact agents' behavior and performance.
John Yang 0002, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao 0006, Karthik Narasimhan, Ofir Press
NeurIPS1
2023 InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
abstract
Humans write code in a fundamentally interactive manner and rely on constant execution feedback to correct errors, resolve ambiguities, and decompose tasks. While LLMs have recently exhibited promising coding capabilities, current coding benchmarks mostly consider a static instruction-to-code sequence transduction process, which has the potential for error propagation and a disconnect between the generated code and its final execution environment. To address this gap, we introduce InterCode, a lightweight, flexible, and easy-to-use framework of interactive coding as a standard reinforcement learning (RL) environment, with code as actions and execution feedback as observations. Our framework is language and platform agnostic, uses self-contained Docker environments to provide safe and reproducible execution, and is compatible out-of-the-box with traditional seq2seq coding methods, while enabling the development of new methods for interactive code generation. We use InterCode to create three interactive code environments with Bash, SQL, and Python as action spaces, leveraging data from the static NL2Bash, Spider, and MBPP datasets. We demonstrate InterCode’s viability as a testbed by evaluating multiple state-of-the-art LLMs configured with different prompting strategies such as ReAct and Plan & Solve. Our results showcase the benefits of interactive code generation and demonstrate that InterCode can serve as a challenging benchmark for advancing code understanding and generation capabilities. InterCode is designed to be easily extensible and can even be used to create new tasks such as Capture the Flag, a popular coding puzzle that is inherently multi-step and involves multiple programming languages.
John Yang 0002, Akshara Prabhakar, Karthik Narasimhan, Shunyu Yao 0006
NeurIPS1
2022 WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
abstract
Most existing benchmarks for grounding language in interactive environments either lack realistic linguistic elements, or prove difficult to scale up due to substantial human involvement in the collection of data or feedback signals. We develop WebShop – a simulated e-commerce website environment with 1.18 million real-world products and 12,087 crowd-sourced text instructions. In this environment, an agent needs to navigate multiple types of webpages and issue diverse actions to find, customize, and purchase a product given an instruction. WebShop provides several challenges including understanding compositional instructions, query (re-)formulation, dealing with noisy text in webpages, and performing strategic exploration. We collect over 1,600 human trajectories to first validate the benchmark, then train and evaluate a diverse range of agents using reinforcement learning, imitation learning, and pre-trained image and language models. Our best model achieves a task success rate of 29%, which significantly outperforms rule heuristics but is far lower than expert human performance (59%). We also analyze agent and human trajectories and ablate various model components to provide insights for developing future agents with stronger language understanding and decision making abilities. Finally, we show our agent trained on WebShop exhibits non-trivial sim-to-real transfer when evaluated on amazon.com and ebay.com, indicating the potential value of our benchmark for developing practical web agents that can operate in the wild.
Shunyu Yao 0006, Howard Chen 0003, John Yang 0002, Karthik Narasimhan
NeurIPS3