Carlos E. Jimenez

dblp:153/0588 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Program synthesis and code generation · 34% Software maintenance and evolution · 21% Software testing · 18%
Artificial intelligence
6 papers
Language models and text generation · 24% Trustworthy machine learning · 22% Vision and language · 16%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 50% Learning and educational technologies · 50%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Human-AI interaction
human-AI collaboration
0.912025
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration · NeurIPS 2025
Learning and educational technologies
knowledge transfer
0.912025
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration · NeurIPS 2025
Systems and software security
exploitation
0.912025
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities · ICML 2025
Systems and software security
vulnerability discovery
0.912025
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities · ICML 2025
Debugging and program repair
automated program repair
0.912025
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025
Program synthesis and code generation
code generation with language models
0.912025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Program synthesis and code generation › code generation with language models
software engineering agents
0.912025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Empirical software engineering › benchmarking
software engineering benchmarks
0.912025
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025
Software testing
test generation
0.912025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Natural language and speech › Language models and text generation
code generation
0.812024
SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024
Software maintenance and evolution › issue management
issue resolution
0.812024
SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity
0.712023
C-STS: Conditional Semantic Textual Similarity · EMNLP 2023
Machine learning › Trustworthy machine learning › robustness
consistency-robustness tradeoff
0.612022
CARETS: A Consistency And Robustness Evaluative Test Suite for VQA · ACL (1) 2022
Machine learning › Efficient and distributed learning
model compression
0.612022
DataMUX: Data Multiplexing for Neural Networks · NeurIPS 2022
Machine learning › Trustworthy machine learning
robustness evaluation
0.612022
CARETS: A Consistency And Robustness Evaluative Test Suite for VQA · ACL (1) 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
DataMUX: Data Multiplexing for Neural Networks · NeurIPS 2022
Computer vision › Vision and language
visual question answering
0.612022
CARETS: A Consistency And Robustness Evaluative Test Suite for VQA · ACL (1) 2022
Natural language and speech › Language models and text generation
large language model
0.312025
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration · NeurIPS 2025
Empirical software engineering
mining software repositories
0.312025
SWE-smith: Scaling Data for Software Engineering Agents · NeurIPS 2025
Software testing › test execution
automated test execution
0.212024
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering · NeurIPS 2024
Software testing › test infrastructure
benchmark construction
0.212024
SWE-bench: Can Language Models Resolve Real-world Github Issues? · ICLR 2024
Natural language and speech › Language models and text generation
natural language understanding
0.212023
C-STS: Conditional Semantic Textual Similarity · EMNLP 2023
Machine learning › Efficient and distributed learning
inference acceleration
0.212022
DataMUX: Data Multiplexing for Neural Networks · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

large language model · 2.4user study · 1.7language model agent · 1.6execution environment interaction · 1.5large language model agents · 0.9large language model agent · 0.9interactive terminal tooling · 0.9execution environments · 0.9agent-computer interface · 0.8dataset construction · 0.7test suite construction · 0.6linear transformation · 0.6demultiplexing · 0.6balanced question generation · 0.6
YearPublicationVenuePosition
2025 SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
abstract
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, with problem statements presented predominantly as text and lacking visual elements such as images. This limited coverage motivates our inquiry into how existing systems might perform on unrepresented software engineering domains (e.g., front-end, game development, DevOps), which use different programming languages and paradigms. Therefore, we propose SWE-bench Multimodal (SWE-bench M), to evaluate systems on their ability to fix bugs in visual, user-facing JavaScript software. SWE-bench M features 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. Each SWE-bench M task instance contains at least one image in its problem statement or unit tests. Our analysis finds that top-performing SWE-bench systems struggle with SWE-bench M, revealing limitations in visual problem-solving and cross-language generalization. Lastly, we show that SWE-agent’s flexible language-agnostic features enable it to substantially outperform alternatives on SWE-bench M, resolving 12% of task instances compared to 6% for the next best system.
John Yang 0002, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik Narasimhan, Diyi Yang, Sida I. Wang, Ofir Press
ICLR2
2025 EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
abstract
Although language model (LM) agents have demonstrated increased performance in multiple domains, including coding and web-browsing, their success in cybersecurity has been limited. We present *EnIGMA*, an LM agent for autonomously solving Capture The Flag (CTF) challenges. We introduce new tools and interfaces to improve the agent's ability to find and exploit security vulnerabilities, focusing on interactive terminal programs. These novel *Interactive Agent Tools* enable LM agents, for the first time, to run interactive utilities, such as a debugger and a server connection tool, which are essential for solving these challenges. Empirical analysis on 390 CTF challenges across four benchmarks demonstrate that these new tools and interfaces substantially improve our agent's performance, achieving state-of-the-art results on NYU CTF, Intercode-CTF, and CyBench. Finally, we analyze data leakage, developing new methods to quantify it and identifying a new phenomenon we term *soliloquizing*, where the model self-generates hallucinated observations without interacting with the environment.
Talor Abramovich, Meet Udeshi, Minghao Shao, Kilian Lieret, Haoran Xi, Kimberly Milner, Sofija Jancheska, John Yang 0002, Carlos E. Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique 0001, Karthik Narasimhan, Ramesh Karri, Ofir Press
ICML9
2025 When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
abstract
As large language models (LLMs) increasingly serve as close collaborators for humans, it is crucial that they express their reasoning in ways that humans can understand and learn from. However, this capability remains relatively less understood and under-evaluated. To address this, we introduce a conceptual framework for such Human-AI knowledge transfer capabilities and conduct the first large-scale user study (N=118) explicitly designed to measure it. In our two-phase setup, humans first ideate with an LLM on problem-solving strategies, then independently implement solutions, isolating the influence of model reasoning on human understanding. Our findings reveal that while model benchmark performance correlates with collaborative outcomes, this relationship is notably inconsistent with significant outliers, highlighting that knowledge transfer is a distinct capability requiring dedicated optimization. Our analysis uncovers behavioral and strategic factors that mediate successful knowledge transfer, and we release our code, dataset, and evaluation framework to support future work on communicatively aligned models.
Carlos E. Jimenez, Shunyu Yao 0006, Nick Haber, Diyi Yang, Karthik Narasimhan
NeurIPS2
2025 SWE-smith: Scaling Data for Software Engineering Agents
abstract
Despite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point.Existing datasets are small, with at most 1,000s of training instances from 11 or fewer GitHub repositories.The procedures to curate such datasets are often complex, necessitating hundreds of hours of human labor; companion execution environments also take up several terabytes of storage, severely limiting their scalability and usability.To address this pain point, we introduce SWE-smith, a novel pipeline for generating software engineering training data at scale.Given any Python codebase, SWE-smith constructs a corresponding execution environment, then automatically synthesizes 100s to 1,000s of task instances that break existing test(s) in the codebase.Using SWE-smith, we create a dataset of 50k instances sourced from 128 GitHub repositories, an order of magnitude larger than all previous works.We train SWE-agent-LM-32B, achieving 40.2% Pass@1 resolve rate on the SWE-bench Verified benchmark, state of the art among open source models.We open source SWE-smith (collection procedure, task instances, trajectories, models) to lower the barrier of entry for research in LM systems for automated software engineering.All assets available at \url{https://swesmith.com}.
John Yang 0002, Kilian Lieret, Carlos E. Jimenez, Alexander Wettig, Kabir Khandpur, Binyuan Hui, Ofir Press, Ludwig Schmidt, Diyi Yang
NeurIPS3
2024 SWE-bench: Can Language Models Resolve Real-world Github Issues?
abstract
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere 1.96% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
Carlos E. Jimenez, John Yang 0002, Alexander Wettig, Shunyu Yao 0006, Kexin Pei, Ofir Press, Karthik Narasimhan
ICLR1
2024 SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
abstract
Language model agents are increasingly being used to automate complicated tasks in digital environments. Just as humans benefit from powerful software applications, such as integrated development environments, for complex tasks like software engineering, we posit that language model agents represent a new category of end users with their own needs and abilities, and would benefit from specially built interfaces to the software they use. We investigate how the role of interface design affects the performance of language model agents. As a result of this exploration, we introduce SWE-agent: a system that facilitates language model agents to autonomously use computers to solve software engineering tasks. SWE-agent's custom agent-computer interface significantly enhances an agent's ability to create and edit code files, navigate entire repositories, and execute tests and other programs. We evaluate SWE-agent on SWE-bench and HumanEvalFix, achieving state-of-the-art performance on both with a pass@1 rate of 12.5% and 87.7%, respectively, far exceeding the previous state-of-the-art achieved with non-interactive language models. Finally, we provide insight on how the design of the agent-computer interface can impact agents' behavior and performance.
John Yang 0002, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao 0006, Karthik Narasimhan, Ofir Press
NeurIPS2
2023 C-STS: Conditional Semantic Textual Similarity
abstract
Ameet Deshpande, Carlos Jimenez, Howard Chen, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen, Karthik Narasimhan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Ameet Deshpande, Carlos E. Jimenez, Howard Chen 0003, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen 0001, Karthik Narasimhan
EMNLP2
2022 CARETS: A Consistency And Robustness Evaluative Test Suite for VQA
abstract
We introduce CARETS, a systematic test suite to measure consistency and robustness of modern VQA models through a series of six fine-grained capability tests.In contrast to existing VQA test sets, CARETS features balanced question generation to create pairs of instances to test models, with each pair focusing on a specific capability such as rephrasing, logical symmetry or image obfuscation.We evaluate six modern VQA systems on CARETS and identify several actionable weaknesses in model comprehension, especially with concepts such as negation, disjunction, or hypernym invariance.Interestingly, even the most sophisticated models are sensitive to aspects such as swapping the order of terms in a conjunction or changing the number of answer choices mentioned in the question.We release CARETS to be used as an extensible tool for evaluating multi-modal model robustness. 1
Carlos E. Jimenez, Olga Russakovsky, Karthik Narasimhan
ACL (1)1
2022 DataMUX: Data Multiplexing for Neural Networks
abstract
In this paper, we introduce \emph{data multiplexing} (DataMUX), a technique that enables deep neural networks to process multiple inputs simultaneously using a single compact representation. DataMUX demonstrates that neural networks are capable of generating accurate predictions over \emph{mixtures} of inputs, resulting in increased inference throughput with minimal extra memory requirements. Our approach uses two key components -- 1) a multiplexing layer that performs a fixed linear transformation to each input before combining them to create a "mixed" representation of the same size as a single input, which is then processed by the base network, and 2) a demultiplexing layer that converts the base network's output back into independent representations before producing predictions for each input. We show the viability of DataMUX for different architectures (Transformers, and to a much lesser extent MLPs and CNNs) across six different tasks spanning sentence classification, named entity recognition and image classification. For instance, DataMUX for Transformers can multiplex up to 20x/40x inputs, achieving up to 11x/18x increase in inference throughput with absolute performance drops of $<2\%$ and $<4\%$ respectively compared to a vanilla Transformer on MNLI, a natural language inference task. We also provide a theoretical construction for multiplexing in self-attention networks and analyze the effect of various design elements in DataMUX.
Vishvak Murahari, Carlos E. Jimenez, Runzhe Yang, Karthik Narasimhan
NeurIPS2