VLDB 2026 Research / reviewers in the wild / expert
David O'Brien
dblp:38/6456
· DBLP profile ↗
9ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-9730-9220ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An LLM-Based Agent-Oriented Approach for Automated Code Design Issue LocalizationabstractMaintaining software design quality is crucial for the long-term maintainability and evolution of systems. However, design issues such as poor modularity and excessive complexity often emerge as codebases grow. Developers rely on external tools, such as program analysis techniques, to identify such issues. This work leverages Large Language Models (LLMs) to develop an automated approach for analyzing and localizing design issues. Large language models have demonstrated significant performance on coding tasks, but directly leveraging them for design issue localization is challenging. Large codebases exceed typical LLM context windows, and program analysis tool outputs in non-textual modalities (e.g., graphs or interactive visualizations) are incompatible with LLMs' natural language inputs. To address these challenges, we propose LOCALIZEAGENT, a novel multi-agent framework for effective design issue localization. LOCALIZEAGENT integrates the specialized agents that (1) analyze code to identify potential code design issues, (2) transform program analysis outputs into abstraction-aware LLM-friendly natural language summaries, (3) generate context-aware prompts tailored to specific refactoring types, and (4) leverage LLMs to locate and rank the localized issues based on their relevance. Our evaluation using diverse real-world codebases demonstrates significant improvements over the baseline approaches, with LOCALIZEAGENT achieving$138 \%, 166 \%$, and 206 % relative improvements in exact-match accuracy for localizing information hiding, complexity, and modularity issues, respectively. Fraol Batole, David O'Brien, Tien N. Nguyen, Robert Dyer 0001, Hridesh Rajan |
ICSE | 2 |
| 2024 | Data-Driven Evidence-Based Syntactic Sugar DesignabstractProgramming languages are essential tools for developers, and their evolution plays a crucial role in supporting the activities of developers. One instance of programming language evolution is the introduction of syntactic sugars, which are additional syntax elements that provide alternative, more readable code constructs. However, the process of designing and evolving a programming language has traditionally been guided by anecdotal experiences and intuition. Recent advances in tools and methodologies for mining open-source repositories have enabled developers to make data-driven software engineering decisions. In light of this, this paper proposes an approach for motivating data-driven programming evolution by applying frequent subgraph mining techniques to a large dataset of 166,827,154 open-source Java methods. The dataset is mined by generalizing Java control-flow graphs to capture broad programming language usages and instances of duplication. Frequent subgraphs are then extracted to identify potentially impactful opportunities for new syntactic sugars. Our diverse results demonstrate the benefits of the proposed technique by identifying new syntactic sugars involving a variety of programming constructs that could be implemented in Java, thus simplifying frequent code idioms. This approach can potentially provide valuable insights for Java language designers, and serve as a proof-of-concept for data-driven programming language design and evolution. David O'Brien, Robert Dyer 0001, Tien N. Nguyen, Hridesh Rajan |
ICSE | 1 |
| 2024 | Are Prompt Engineering and TODO Comments Friends or Foes? An Evaluation on GitHub CopilotabstractCode intelligence tools such as GitHub Copilot have begun to bridge the gap between natural language and programming language. A frequent software development task is the management of technical debts, which are suboptimal solutions or unaddressed issues which hinder future software development. Developers have been found to "self-admit" technical debts (SATD) in software artifacts such as source code comments. Thus, is it possible that the information present in these comments can enhance code generative prompts to repay the described SATD? Or, does the inclusion of such comments instead cause code generative tools to reproduce the harmful symptoms of described technical debt? Does the modification of SATD impact this reaction? Despite the heavy maintenance costs caused by technical debt and the recent improvements of code intelligence tools, no prior works have sought to incorporate SATD towards prompt engineering. Inspired by this, this paper contributes and analyzes a dataset consisting of 36,381 TODO comments in the latest available revisions of their respective 102,424 repositories, from which we sample and manually generate 1,140 code bodies using GitHub Copilot. Our experiments show that GitHub Copilot can generate code with the symptoms of SATD, both prompted and unprompted. Moreover, we demonstrate the tool's ability to automatically repay SATD under different circumstances and qualitatively investigate the characteristics of successful and unsuccessful comments. Finally, we discuss gaps in which GitHub Copilot's successors and future researchers can improve upon code intelligence tasks to facilitate AI-assisted software maintenance. David O'Brien, Sumon Biswas, Sayem Mohammad Imtiaz, Rabe Abdalkareem, Emad Shihab, Hridesh Rajan |
ICSE | 1 |
| 2022 | 23 shades of self-admitted technical debt: an empirical study on machine learning softwareabstractIn software development, the term “technical debt” (TD) is used to characterize short-term solutions and workarounds implemented in source code which may incur a long-term cost. Technical debt has a variety of forms and can thus affect multiple qualities of software including but not limited to its legibility, performance, and structure. In this paper, we have conducted a comprehensive study on the technical debts in machine learning (ML) based software. TD can appear differently in ML software by infecting the data that ML models are trained on, thus affecting the functional behavior of ML systems. The growing inclusion of ML components in modern software systems have introduced a new set of TDs. Does ML software have similar TDs to traditional software? If not, what are the new types of ML specific TDs? Which ML pipeline stages do these debts appear? Do these debts differ in ML tools and applications and when they get removed? Currently, we do not know the state of the ML TDs in the wild. To address these questions, we mined 68,820 self-admitted technical debts (SATD) from all the revisions of a curated dataset consisting of 2,641 popular ML repositories from GitHub, along with their introduction and removal. By applying an open-coding scheme and following upon prior works, we provide a comprehensive taxonomy of ML SATDs. Our study analyzes ML SATD type organizations, their frequencies within stages of ML software, the differences between ML SATDs in applications and tools, and quantifies the removal of ML SATDs. The findings discovered suggest implications for ML developers and researchers to create maintainable ML systems. David O'Brien, Sumon Biswas, Sayem Imtiaz, Rabe Abdalkareem, Emad Shihab, Hridesh Rajan |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Fairness and Machine FairnessabstractPrediction-based decisions, which are often made by utilizing the tools of machine learning, influence nearly all facets of modern life. Ethical concerns about this widespread practice have given rise to the field of fair machine learning and a number of fairness measures, mathematically precise definitions of fairness that purport to determine whether a given prediction-based decision system is fair. Following Reuben Binns (2017), we take "fairness" in this context to be a placeholder for a variety of normative egalitarian considerations. We explore a few fairness measures to suss out their egalitarian roots and evaluate them, both as formalizations of egalitarian ideas and as assertions of what fairness demands of predictive systems. We pay special attention to a recent and popular fairness measure, counterfactual fairness, which holds that a prediction about an individual is fair if it is the same in the actual world and any counterfactual world where the individual belongs to a different demographic group (cf. Kusner et al. 2018). Clinton Castro, David O'Brien, Ben Schwan |
AIES | 2 |
| 2004 | Probik: Protein Backbone Motion by Inverse Kinematics
Kimberly Noonan, David O'Brien, Jack Snoeyink |
WAFR | 2 |
| 2002 | Standards-based Sharable Active Guideline Environment (SAGE): A Project to Develop a Universal Framework for Encoding and Disseminating Electronic Clinical Practice Guidelines
Nick Beard, James R. Campbell 0001, Stanley M. Huff, Mauricio Leon, James G. Mansfield, Eric Mays, James C. McClay, David N. Mohr, Mark A. Musen, David O'Brien, Roberto A. Rocha, Anne Saulovich, Sidna M. Tulledge-Scheitel, Samson W. Tu |
AMIA | 10 |
| 2001 | Automatic simplification of particle system dynamicsabstractWe present a novel framework for automatically simplifying the dynamics computation of particle systems to improve simulation speeds. Our approach is based on a physically-based subdivision scheme to generate a hierarchy of approximated motion models or simulation levels of detail (SLOD). At each time step, the SLODs are updated on-the-fly, and the appropriate SLOD is chosen adaptively to reduce computational costs. We have tested a prototype implementation on the simulation of a water fountain and a galaxy system. The preliminary results show a significant performance gain on these scenarios with little loss in the visual appearance of the simulation, indicating the potential to generalize this approach to other dynamical systems. David O'Brien, Susan Fisher, Ming C. Lin |
CA | 1 |
| 1997 | A Strategy and Architecture for The Visualization of Complex Geographical DatasetsabstractThe use of computer visualization as a means to analyze complex geographic datasets is discussed. Visualization is a valuable tool for conducting exploratory data analysis on geographical data; making good use of the human eye's unparalleled ability to recognize structure and relationships that may be inherent within the data. Traditional GIS are extremely poor at visualization, being limited to a very restricted set of visual attributes with which to convey information (position, size, color). The use of a more sophisticated approach is discussed in detail. Specifically, a system to visualise complex environmental datasets is described, which makes use of knowledge concerning the problem domain as well as knowledge concerning human cognition. In the realizations produced, the most salient attributes in the data, for a particular task, are assigned to the most striking visual attributes. Assignments are controlled by heuristics that may be changed to alter system behavior. Results are presented showing the application of this approach on datasets involving several multi-dimensional thematic layers of environmental data, used in mineral exploration. Mark Gahegan, David O'Brien |
Int. J. Pattern Recognit. Artif. Intell. | 2 |