Alex L. Zhang

dblp:389/8697 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Debugging and program repair · 33% Empirical software engineering · 33% Program synthesis and code generation · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 44% Hardware accelerators and domain-specific architectures · 44% Performance modeling and evaluation · 13%
Artificial intelligence
1 paper
Vision and language · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
automated program repair
0.912025
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025
Program synthesis and code generation
code generation with language models
0.912025
KernelBench: Can LLMs Write Efficient GPU Kernels? · ICML 2025
Empirical software engineering › benchmarking
software engineering benchmarks
0.912025
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · ICLR 2025
GPUs and heterogeneous computing
GPU kernel
0.912025
KernelBench: Can LLMs Write Efficient GPU Kernels? · ICML 2025
Hardware accelerators and domain-specific architectures
kernel generation
0.912025
KernelBench: Can LLMs Write Efficient GPU Kernels? · ICML 2025
Performance modeling and evaluation
benchmarking
0.312025
KernelBench: Can LLMs Write Efficient GPU Kernels? · ICML 2025

Methods — techniques the papers use, named apart from their topics

large language model · 1.7iterative refinement · 1.7large language model agents · 0.9large language model agent · 0.9
YearPublicationVenuePosition
2025 SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
abstract
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, with problem statements presented predominantly as text and lacking visual elements such as images. This limited coverage motivates our inquiry into how existing systems might perform on unrepresented software engineering domains (e.g., front-end, game development, DevOps), which use different programming languages and paradigms. Therefore, we propose SWE-bench Multimodal (SWE-bench M), to evaluate systems on their ability to fix bugs in visual, user-facing JavaScript software. SWE-bench M features 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. Each SWE-bench M task instance contains at least one image in its problem statement or unit tests. Our analysis finds that top-performing SWE-bench systems struggle with SWE-bench M, revealing limitations in visual problem-solving and cross-language generalization. Lastly, we show that SWE-agent’s flexible language-agnostic features enable it to substantially outperform alternatives on SWE-bench M, resolving 12% of task instances compared to 6% for the next best system.
John Yang 0002, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik Narasimhan, Diyi Yang, Sida I. Wang, Ofir Press
ICLR3
2025 KernelBench: Can LLMs Write Efficient GPU Kernels?
abstract
Efficient GPU kernels are crucial for building performant machine learning architectures, but writing them is a time-consuming challenge that requires significant expertise; therefore, we explore using language models (LMs) to automate kernel generation. We introduce KernelBench, an open-source framework for evaluating LMs’ ability to write fast and correct kernels on a suite of 250 carefully selected PyTorch ML workloads. KernelBench represents a real-world engineering environment and making progress on the introduced benchmark directly translates to faster practical kernels. We introduce a new evaluation metric $\text{fast}_p$, which measures the percentage of generated kernels that are functionally correct and offer a speedup greater than an adjustable threshold $p$ over baseline. Our experiments across various state-of-the-art models and test-time methods show that frontier reasoning models perform the best out of the box but still fall short overall, matching the PyTorch baseline in less than 20% of the cases. While we show that results can improve by leveraging execution and profiling feedback during iterative refinement, KernelBench remains a challenging benchmark, with its difficulty increasing as we raise speedup threshold $p$.
Anne Ouyang, Simon Guo 0004, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, Azalia Mirhoseini
ICML4