VLDB 2026 Research / reviewers in the wild / expert
John-Paul Ore
dblp:148/2225
· DBLP profile ↗
14ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-8122-3809ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analyzing the dependability of Large Language Models for code clone generationabstractThe ability to generate multiple equivalent versions of the same code segment across different programming languages and within the same language is valuable for code translation, language migration, and code comprehension in education. However, current avenues for generating code clones — through manual creation or specialized software tools — often fail to consistently generate a variety of behaviorally equivalent code clones. Large Language Models (LLMs) offer a promising solution by leveraging their extensive training on diverse codebases to automatically generate code. Unlike traditional methods, LLMs can produce code across a wide variety of programming languages with minimal user effort. Using LLMs for code clone generation could significantly reduce the time and resources needed to create code clones while enhancing their syntactic diversity. In this quantitative empirical study, we investigate the dependability of LLMs as potential generators of code clones. We gathered equivalent code solutions (i.e., behavioral clones) in C++, Java, and Python from thirty-six programming problems from the well-known technical interview practice platform, LeetCode. We query OpenAI’s GPT-3.5, GPT-4, and CodeLlama to generate code clones of the LeetCode solutions. We measure the behavioral equivalence of the LLM-generated clones using a behavioral similarity clustering technique inspired by the code clone detection tool, Simion-based Language Agnostic Code Clones (SLACC). This study reveals that, despite LLMs demonstrating the potential for code generation, their capacity to consistently generate syntactically diverse but behaviorally equivalent code clones is limited. At lower temperature settings, LLMs are more successful in producing behaviorally consistent, syntactically similar code clones within the same language. However, for cross-language cloning tasks and at higher temperature settings and programming difficulties, LLMs introduce greater syntactic diversity and lead to higher rates of compilation and runtime errors, resulting in a decline in behavioral consistency. These findings indicate a need for further quality assurance measures for the use of LLMs for code clone generation. All the data and scripts associated with this paper can be found https://zenodo.org/records/14968618 . Azeeza Eagal, Kathryn T. Stolee, John-Paul Ore |
J. Syst. Softw. | 3 |
| 2024 | Is it a Bug? Understanding Physical Unit Mismatches in Robot SoftwareabstractRobot software is abundant with variables that represent real-world physical units (e.g., meters, seconds). Operations over different units (e.g., adding meters and seconds) may be incorrect and can lead to dangerous system misbehaviors; manually detecting such mistakes is challenging. Current software analysis techniques identify such mismatches using dimensional analysis rules and ROS-specific assumptions to analyze the source code. However, these are ignorant of the fact that physical unit mismatches in robotics code are often intentional (e.g., when operating a differential drive robot), resulting in false positive bug reports that can impede robotics developer trust and productivity. In this work, we study how developers introduce physical unit mismatches by manually inspecting 180 errors detected by the software analysis technique, Phys. We identify three types of physical unit mismatches and present a taxonomy of eight high-level categories of how these errors manifest. We find that developers often make unforced and paradigmatic physical unit mismatches through differential drives, small angle approximations, and controls. We draw insights on current development to inform future research to better detect, categorize, and address meaningful physical unit mismatches. Paulo Canelas, Trenton Tabor, John-Paul Ore, Alcides Fonseca, Claire Le Goues, Christopher Steven Timperley |
ICRA | 3 |
| 2024 | Barriers for Students During Code Change ComprehensionabstractModern code review (MCR) is a key practice for many software engineering organizations, so undergraduate software engineering courses often teach some form of it to prepare students. However, research on MCR describes how many its professional implementations can fail, to say nothing on how these barriers manifest under students' particular contexts. To uncover barriers students face when evaluating code changes during review, we combine interviews and surveys with an observational study. In a junior-level software engineering course, we first interviewed 29 undergraduate students about their experiences in code review. Next, we performed an observational study that presented 44 students from the same course with eight code change comprehension activities. These activities provided students with pull requests of potential refactorings in a familiar code base, collecting feedback on accuracy and challenges. This was followed by a reflection survey. Justin Middleton, John-Paul Ore, Kathryn T. Stolee |
ICSE | 2 |
| 2022 | Understanding Xacro MisunderstandingsabstractThe Xacro XML macro language can be used to augment the Universal Robot Description Format (URDF) and is part of a critical toolchain from geometric representations to simulation, visualization, and system execution. However, mem-bers of the robotics community, especially newcomers, struggle to troubleshoot and understand the interplay between systems and the Xacro preprocessing pipeline. To better understand how system developers struggle with Xacros, we manually examine 712 Xacro-related questions from the question and answer site answers.ros.org and find Xacro misunderstandings fit into eight key categories using a systematic, qualitative approach called Open Coding. By examining the 'tags' applied to questions, we further find that Xacro problems manifest in a befuddlingly broad set of contexts. This hinders onboarding and complicates system developers' understanding of representations and tools in the Robot Operating System. We aim to provide an empirical grounding that identifies and prioritizes impediments to users of open robotics systems, so that tool designers, teachers, and robotics practitioners can devise ways of improving robot software tooling and education. Nicholas Albergo, Vivek Rathi, John-Paul Ore |
ICRA | 3 |
| 2022 | Maktub: Lightweight Robot System Test Creation and AutomationabstractThe rapid expansion of robotics relies on properly configuring and testing hardware and software. Due to the expense and hazard of real-world testing on hardware, robot system testing increasingly utilizes extensive simulation. Creating robot simulation tests requires specialized skills in robot programming and simulation tools. While there are many platforms and tool-kits to create these simulations, they can be cumbersome when combined with automated testing. We present Maktub: a tool for creating tests using Unity and ROS. Maktub leverages the extensive 3D manipulation capabilities of Unity to lower the barrier in creating system tests for robots. A key idea of Maktub is to make tests without needing robotic software development skills. A video demonstration of Maktub can be found here: https://youtu.be/c0Bacy3DlEE, and the source code can be found at https://github.com/RobotCodeLab/Maktub. Amr Moussa, John-Paul Ore |
ASE | 2 |
| 2021 | An Empirical Study on Type Annotations: Accuracy, Speed, and Suggestion EffectivenessabstractType annotations connect variables to domain-specific types. They enable the power of type checking and can detect faults early. In practice, type annotations have a reputation of being burdensome to developers. We lack, however, an empirical understanding of how and why they are burdensome. Hence, we seek to measure the baseline accuracy and speed for developers making type annotations to previously unseen code. We also study the impact of one or more type suggestions. We conduct an empirical study of 97 developers using 20 randomly selected code artifacts from the robotics domain containing physical unit types. We find that subjects select the correct physical type with just 51% accuracy, and a single correct annotation takes about 2 minutes on average. Showing subjects a single suggestion has a strong and significant impact on accuracy both when correct and incorrect, while showing three suggestions retains the significant benefits without the negative effects. We also find that suggestions do not come with a time penalty. We require subjects to explain their annotation choices, and we qualitatively analyze their explanations. We find that identifier names and reasoning about code operations are the primary clues for selecting a type. We also examine two state-of-the-art automated type annotation systems and find opportunities for their improvement. John-Paul Ore, Carrick Detweiler, Sebastian G. Elbaum |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2018 | Assessing the type annotation burdenabstractType annotations provide a link between program variables and domain-specific types. When combined with a type system, these annotations can enable early fault detection. For type annotations to be cost-effective in practice, they need to be both accurate and affordable for developers. We lack, however, an understanding of how burdensome type annotation is for developers. Hence, this work explores three fundamental questions: 1) how accurately do developers make type annotations; 2) how long does a single annotation take; and, 3) if a system could automatically suggest a type annotation, how beneficial to accuracy are correct suggestions and how detrimental are incorrect suggestions? We present results of a study of 71 programmers using 20 random code artifacts that contain variables with physical unit types that must be annotated. Subjects choose a correct type annotation only 51% of the time and take an average of 136 seconds to make a single correct annotation. Our qualitative analysis reveals that variable names and reasoning over mathematical operations are the leading clues for type selection. We find that suggesting the correct type boosts accuracy to 73%, while making a poor suggestion decreases accuracy to 28%. We also explore what state-of-the-art automated type annotation systems can and cannot do to help developers with type annotations, and identify implications for tool developers. John-Paul Ore, Sebastian G. Elbaum, Carrick Detweiler, Lambros Karkazis |
ASE | 1 |
| 2018 | Phys: probabilistic physical unit assignment and inconsistency detectionabstractProgram variables used in robotic and cyber-physical systems often have implicit physical units that cannot be determined from their variable types. Inferring an abstract physical unit type for variables and checking their physical unit type consistency is of particular importance for validating the correctness of such systems. For instance, a variable with the unit of ‘meter’ should not be assigned to another variable with the unit of ‘degree-per-second’. Existing solutions have various limitations such as requiring developers to annotate variables with physical units and only handling variables that are directly or transitively used in popular robotic libraries with known physical unit information. We observe that there are a lot of physical unit hints in these softwares such as variable names and specific forms of expressions. These hints have uncertainty as developers may not respect conventions. We propose to model them with probability distributions and conduct probabilistic inference. At the end, our technique produces a unit distribution for each variable. Unit inconsistencies can then be detected using the highly probable unit assignments. Experimental results on 30 programs show that our technique can infer units for 159.3% more variables compared to the state-of-the-art with more than 88.7% true positives, and inconsistencies detection on 90 programs shows that our technique reports 103.3% more inconsistencies with 85.3% true positives. Sayali Kate, John-Paul Ore, Xiangyu Zhang 0001, Sebastian G. Elbaum, Zhaogui Xu |
ESEC/SIGSOFT FSE | 2 |
| 2017 | Dimensional inconsistencies in code and ROS messages: A study of 5.9M lines of codeabstractThis work presents a study of robot software using the Robot Operating System (ROS), focusing on detecting inconsistencies in physical unit manipulation. We discuss how dimensional analysis, the rules governing how physical quantities are combined, can be used to detect inconsistencies in robot software that are otherwise difficult to detect. Using a corpus of ROS software with 5.9M lines of code, we measure the frequency of these dimensional inconsistencies and find them in 6% (211 / 3,484) of repositories that use ROS. We find that the inconsistency type `Assigning multiple units to a variable' accounts for 75% of inconsistencies in ROS code. We identify the ROS classes and physical units most likely to be involved with dimensional inconsistencies, and find that the ROS Message type geometry_msgs::Twist is involved in over half of all inconsistencies and is used by developers in ways contrary to Twist's intent. We further analyze the frequency of physical units used in ROS programs as a proxy for assessing how developers use ROS, and discuss the practical implications of our results including how to detect and avoid these inconsistencies. John-Paul Ore, Sebastian G. Elbaum, Carrick Detweiler |
IROS | 1 |
| 2017 | Lightweight detection of physical unit inconsistencies without program annotationsabstractSystems interacting with the physical world operate on quantities measured with physical units. When unit operations in a program are inconsistent with the physical units' rules, those systems may suffer. Existing approaches to support unit consistency in programs can impose an unacceptable burden on developers. In this paper, we present a lightweight static analysis approach focused on physical unit inconsistency detection that requires no end-user program annotation, modification, or migration. It does so by capitalizing on existing shared libraries that handle standardized physical units, common in the cyber-physical domain, to link class attributes of shared libraries to physical units. Then, leveraging rules from dimensional analysis, the approach propagates and infers units in programs that use these shared libraries, and detects inconsistent unit usage. We implement and evaluate the approach in a tool, analyzing 213 open-source systems containing +900,000 LOC, finding inconsistencies in 11% of them, with an 87% true positive rate for a class of inconsistencies detected with high confidence. An initial survey of robot system developers finds that the unit inconsistencies detected by our tool are 'problematic', and we investigate how and when these inconsistencies occur. John-Paul Ore, Carrick Detweiler, Sebastian G. Elbaum |
ISSTA | 1 |
| 2017 | Phriky-units: a lightweight, annotation-free physical unit inconsistency detection toolabstractSystems that interact with the physical world use software that represents and manipulates physical quantities. To operate correctly, these systems must obey the rules of how quantities with physical units can be combined, compared, and manipulated. Incorrectly manipulating physical quantities can cause faults that go undetected by the type system, likely manifesting later as incorrect behavior. Existing approaches for inconsistency detection require code annotation, physical unit libraries, or specialized programming languages. We introduce Phriky-Units, a static analysis tool that detects physical unit inconsistencies in robotic software without developer annotations. It does so by capitalizing on existing shared libraries that handle standardized physical units, common in the cyber-physical domain, to link class attributes of shared libraries to physical units. In this work, we describe how Phriky-Units works, provide details of the implementation, and explain how Phriky-Units can be used. Finally we present a summary of an empirical evaluation showing it has an 87% true positive rate for a class of inconsistencies we detect with high-confidence. John-Paul Ore, Carrick Detweiler, Sebastian G. Elbaum |
ISSTA | 1 |
| 2015 | Surface classification for sensor deployment from UAV landingsabstractUsing Unmanned Aerial Vehicles (UAVs) to deploy sensor networks promises an autonomous and useful method of installation in remote or hard to access locations. Some sensors, such as soil moisture sensors, must be physically installed in soft soil, yet UAVs cannot easily determine soil softness with remote sensors. In this paper, we use data from an onboard accelerometer measured during UAV landings to determine the softness of the ground. We collect and analyze over 200 data sets gathered from 8 different materials: foam, carpet, wood, tile, grass, dirt, concrete, and woodchips. Based on this analysis, we examine a number of features from the accelerometer and four classification algorithms: LDA, QDA, SVM, and binary decision trees. The decision tree performs well and is simple to implement onboard the UAV. We implement this in our UAV control system and perform experiments to verify that the UAV can accurately classify the softness of the surface with 90% accuracy. This lays the groundwork for our future work on developing a UAV capable of installing sensors in soft soil. David J. Anthony, Elizabeth Basha, Jared Ostdiek, John-Paul Ore, Carrick Detweiler |
ICRA | 4 |
| 2015 | On air-to-water radio communication between UAVs and water sensor networksabstractOcean monitoring using underwater sensor networks faces communication challenges in retrieving data, communicating large amounts of data between nodes, and covering increasing spatial regions while remaining connected. With underwater sensor networks that are capable of surfacing, unmanned aerial vehicles (UAVs) provide a solution to this by providing radio-based data muling services, but, as this area is still unexplored, the utility of this solution is unclear. In this paper, we examine the theoretical expectations, perform several field experiments, and analyze the communication success rates of 802.15.4 radios near the water surface both communicating between surface nodes as well as between a node and the UAV. These indicate that on the water surface internode radio communication is poor, but node to UAV communication can provide both reasonable ranges and success rates. We additionally measure and analyze the energy aspects of the systems, determining the impacts of parameters such as network size and distance between nodes on the UAV energy. Finally, we consolidate the information into an algorithm outlining how to configure and design hybrid UAV and underwater sensor network systems. Jacob Palmer, Nicholas Yuen, John-Paul Ore, Carrick Detweiler, Elizabeth Basha |
ICRA | 3 |
| 2014 | Controlled sensor network installation with unmanned aerial vehiclesabstractRobots improve wireless sensor network (WSN) deployments by reducing deployment times, deploying nodes to improve coverage, and ferrying data. Utilizing Unmanned Aerial Vehicles (UAVs) to install sensor networks in environmentally sensitive areas is especially valuable, as the UAVs are able to quickly traverse rough and environmentally sensitive terrain. UAV based deployments are challenging, as the UAVs may need to install nodes in a specific orientation or location type, which is difficult to sense from a UAV. We present our work towards resolving these difficulties by first classifying the surface a UAV has landed on, and then conducting a post-deployment analysis of the installation. David J. Anthony, John-Paul Ore, Carrick Detweiler, Elizabeth Basha |
SenSys | 2 |