VLDB 2026 Research / reviewers in the wild / expert
Sophia Kolak
dblp:278/0629
· DBLP profile ↗
2ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Debugging and program repair · 80% Empirical software engineering · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Debugging and program repair
automated program repair |
0.9 | 1 | 2025 | Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models · ICSE 2025 |
Debugging and program repair
fault localization |
0.9 | 1 | 2025 | Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models · ICSE 2025 |
Empirical software engineering › AI for software engineering › machine learning for software engineering
patch classification |
0.9 | 1 | 2025 | Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models · ICSE 2025 |
Debugging and program repair › automated program repair
patch generation |
0.9 | 1 | 2025 | Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models · ICSE 2025 |
Debugging and program repair › fault localization
spectrum-based fault localization |
0.9 | 1 | 2025 | Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models · ICSE 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 0.9entropy analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language ModelsabstractThe problem of software quality has motivated the development of a variety of techniques for Automatic Program Repair (APR). Meanwhile, recent advances in AI and Large Language Models (LLMs) have produced orders of magnitude performance improvements over previous code generation techniques, affording promising opportunities for program repair and its constituent subproblems (e.g., fault localization, patch generation). Because models are trained on large volumes of code in which defects are relatively rare, they tend to both simultaneously perceive faulty code as unlikely (or “unnatural”) and to produce generally correct code (which is more “natural”). This paper comprehensively revisits the idea of (un)naturalness for program repair. We argue that, fundamentally, LLMs can only go so far on their own in reasoning about and fixing buggy code. This motivates the incorporation of traditional tools, which compress useful contextual and analysis information, as a complement to LLMs for repair. We interrogate the role of entropy at every stage of traditional repair, and show that it is indeed usefully complementary to classic techniques. We show that combining measures of naturalness with class Spectrum-Based Fault Localization (SBFL) approaches improves Top-5 scoring by 50 % over SBFL alone. We show that entropy delta, or change in entropy induced by a candidate patch, can improve patch generation efficiency by 24 test suite executions per repair, on average, on our dataset. Finally, we show compelling results that entropy delta for patch classification is highly effective at distinguishing correct from overfitting patches. Overall, our results suggest that LLMs can effectively complement classic techniques for analysis and transformation, producing more efficient and effective automated repair techniques overall. Aidan Z. H. Yang, Sophia Kolak, Vincent J. Hellendoorn, Ruben Martins, Claire Le Goues |
ICSE | 2 |
| 2020 | It Takes a Village to Build a Robot: An Empirical Study of The ROS EcosystemabstractOver the past eleven years, the Robot Operating System (ROS), has grown from a small research project into the most popular framework for robotics development. Composed of packages released on the Rosdistro package manager, ROS aims to simplify development by providing reusable libraries, tools and conventions for building a robot. Still, developing a complete robot is a difficult task that involves bridging many technical disciplines. Experts who create computer vision packages, for instance, may need to rely on software designed by mechanical engineers to implement motor control. As building a robot requires domain expertise in software, mechanical, and electrical engineering, as well as artificial intelligence and robotics, ROS faces knowledge based barriers to collaboration.In this paper, we examine how the necessity of domain specific knowledge impacts the open source collaboration model. We create a comprehensive corpus of package metadata and dependencies over three years in the ROS ecosystem, analyze how collaboration is structured, and study the dependency network evolution. We find that the most widely used ROS packages belong to a small cluster of foundational working groups (FWGs), each organized around a different domain in robotics. We show that the FWGs are growing at a slower rate than the rest of the ecosystem, in terms of their membership and number of packages, yet the number of dependencies on FWGs is increasing at a faster rate. In addition, we mined all ROS packages on GitHub, and showed that 82% rely exclusively on functionality provided by FWGs. Finally, we investigate these highly influential groups and describe the unique model of collaboration they support in ROS. Sophia Kolak, Afsoon Afzal, Claire Le Goues, Michael Hilton 0001, Christopher Steven Timperley |
ICSME | 1 |