Vijay Kethanaboyina

dblp:408/8488 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Software testing · 44% Program synthesis and code generation · 44% Compilers and program optimization · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
0.912025
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents · NeurIPS 2025
Software testing
performance testing
0.912025
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents · NeurIPS 2025
Compilers and program optimization
software optimization
0.312025
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

large language model · 0.9benchmark evaluation · 0.9
YearPublicationVenuePosition
2025 GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
abstract
Developing high-performance software is a complex task that requires specialized expertise. We introduce GSO, a benchmark for evaluating language models' capabilities in developing high-performance software.We develop an automated pipeline that generates and executes performance tests to analyze repository commit histories to identify 102 challenging optimization tasks across 10 codebases, spanning diverse domains and programming languages.An agent is provided with a codebase and performance test as a precise specification, and tasked to improve the runtime efficiency, which is measured against the expert developer optimization.Our quantitative evaluation reveals that leading SWE-Agents struggle significantly, achieving less than 5% success rate, with limited improvements even with inference-time scaling.Our qualitative analysis identifies key failure modes, including difficulties with low-level languages, practicing lazy optimization strategies, and challenges in accurately localizing bottlenecks.We release the code and artifacts of our benchmark along with agent trajectories to enable future research.
Manish Shetty, Naman Jain, Jinjian Liu, Vijay Kethanaboyina, Koushik Sen, Ion Stoica
NeurIPS4