VLDB 2026 Research / reviewers in the wild / expert
João Castelhano
dblp:171/9502
· DBLP profile ↗
5ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0002-8996-1515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Quality Evaluation of Modern Code Reviews Through Intelligent Biometric Program ComprehensionabstractCode review is an essential practice in software engineering to spot code defects in the early stages of software development. Modern code reviews (e.g., acceptance or rejection of pull requests with Git) have become less formal than classic Fagan's inspections, lightweight, and more reliant on individuals (i.e., reviewers). However, reviewers may encounter mentally demanding challenges during the code review, such as code comprehension difficulties or distractions that might affect the code review quality. This work proposes a novel approach that evaluates the quality of code reviews in terms of bug-finding effectiveness and provides the reviewers with a clear message of whether the review should be repeated, indicating the code regions that may not have been well-reviewed. The proposed approach utilizes biometric information collected from the reviewer during the review process using non-intrusive biofeedback devices (e.g., smartwatches). Biometric measures such as Heart Rate Variability (HRV) and task-evoked pupillary response are captured as a surrogate of the cognitive state of the reviewer (e.g., mental workload) and inexpensive desktop eye-trackers compatible with the software development settings. This work uses Artificial Intelligence techniques to predict the cognitive load from the extracted biomarkers and classify each code region according to a set of features. The final evaluation considers various factors such as code complexity, time of the code review, the experience level of the reviewer, and other factors. Our experimental results show the approach could predict the review quality with 87.77%±4.65 accuracy and a Spearman correlation coefficient of 0.85 (p-value < 0.001) between the predicted and the actual review performance. This evaluation validates the cognitive load measurement using electroencephalography (EEG) signals as ground truth for the HRV and pupil signals. Haytham Hijazi, João Durães, Ricardo Couceiro, João Castelhano, Raul Barbosa, Júlio Medeiros, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira |
IEEE Trans. Software Eng. | 4 |
| 2021 | iReview: an Intelligent Code Review Evaluation Tool using BiofeedbackabstractCode reviews and software inspections are essential for building reliable software. However, current code reviews practice in the software industry (e.g., acceptance or rejection of pull requests with Git) deviates considerably from classic (and expensive) Fagan's inspections. Modern code reviews are lightweight and asynchronous and do not rely on a group of inspectors and inspection meetings any longer. The modern style of code reviews is much more flexible and cost-effective. Still, these advantages come with the price of reducing the quality of code reviews, as a single reviewer generally makes them with all the inherent and the very human limitations of one single look. The reviewer could be distracted, overloaded, under stress, or not even fully understand the code under review. This paper proposes a new tool (iReview) that evaluates the code review quality using biometric measures gathered from code reviewers (often called Biofeedback). Biometric measures such as Heart Rate Variability (HRV) and eye movement dynamics are used to assess the reviewer's comprehension of the code under review. iReview evaluates the quality of each review globally and indicates the code regions that have not been well-reviewed, explaining why those code regions should be reviewed again. The tool uses Artificial Intelligence techniques to classify the code regions into good and bad reviews based on various biometric and non-biometric features. The first results show that iReview can predict the review quality of medium or complex programs with an accuracy ranging from 75% to 87% in detecting bad reviews (i.e., code regions classified as bad reviewed still have undetected bugs). This tool is expected to improve software reliability by ensuring that good reviews have been carried out despite the current lightweight reviewing processes. Haytham Hijazi, José Cruz, João Castelhano, Ricardo Couceiro, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira |
ISSRE | 3 |
| 2019 | Pupillography as Indicator of Programmers' Mental Effort and Cognitive OverloadabstractOur research explores a recent paradigm called Biofeedback Augmented Software Engineering (BASE) that introduces a strong new element in the software development process: the programmers' biofeedback. In this Practical Experience Report we present the results of an experiment to evaluate the possibility of using pupillography to gather biofeedback from the programmers. The idea is to use pupillography to get meta information about the programmers' cognitive and emotional states (stress, attention, mental effort level, cognitive overload,...) during code development to identify conditions that may precipitate programmers making bugs or bugs escaping human attention, and tag the corresponding code locations in the software under development to provide online warnings to the programmer or identify code snippets that will need more intensive testing. The experiments evaluate the use of pupillography as cognitive load predictor, compare the results with the mental effort perceived by programmers using NASATLX, and discuss different possibilities for the use of pupillography as biofeedback sensor in real software development scenarios. Ricardo Couceiro, Gonçalo Duarte, João Durães, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira |
DSN | 4 |
| 2019 | Spotting Problematic Code Lines using Nonintrusive Programmers' BiofeedbackabstractRecent studies have shown that programmers' cognitive load during typical code development activities can be assessed using wearable and low intrusive devices that capture peripheral physiological responses driven by the autonomic nervous system. In particular, measures such as heart rate variability (HRV) and pupillography can be acquired by nonintrusive devices and provide accurate indication of programmers' cognitive load and attention level in code related tasks, which are known elements of human error that potentially lead to software faults. This paper presents an experimental study designed to evaluate the possibility of using HRV and pupillography together with eye tracking to identify and annotate specific code lines (or even finer grain lexical tokens) of the program under development (or under inspection) with information on the cognitive load of the programmer while dealing with such lines of code. The experimental data is discussed in the paper to assess different alternatives for using code annotations representing programmers' cognitive load while producing or reading code. In particular, we propose the use of biofeedback code highlighting techniques to provide online programmer's warnings for potentially problematic code lines that may need a second look at (to remove possible bugs), and biofeedback-driven software testing to optimize testing effort, focusing the tests on code areas with higher bug probability. Ricardo Couceiro, Paulo Carvalho 0001, Miguel Castelo-Branco, Henrique Madeira, Raul Barbosa, João Durães, Gonçalo Duarte, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Nuno Laranjeiro, Júlio Medeiros |
ISSRE | 8 |
| 2016 | WAP: Understanding the Brain at Software DebuggingabstractWe propose that understanding functional patterns of activity in mapped brain regions associated with code comprehension tasks and, more specifically, to the activity of finding bugs in traditional code inspections could reveal useful insights to improve software reliability and to improve the software development process in general. This includes helping to select the best professionals for the debugging effort, improving the conditions for code inspections, and identify new directions to follow for training code reviewers. This paper presents an interdisciplinary study to analyze the brain activity during code inspection tasks using functional magnetic resonance imaging (fMRI), which is a well-established tool in cognitive neuroscience research. We used several programs where realistic bugs representing the most frequent types of software faults found in the field were injected. The code inspectors involved in the research include programmers with different levels of expertise and experience in real code reviews. The goal is to understand brain activity patterns associated with code comprehension tasks and, more specifically, the brain activity when the code reviewer identifies a bug in the code ('eureka' moment), which can be a true positive or a false positive. Our results confirmed that brain areas associated with language processing and mathematics are highly active during code reviewing and shows that there are specific brain activity patterns that can be related to the decision-making moment of suspicion/bug detection. Importantly, the activity at the anterior insula region that we find to play a relevant role in the process of identifying software bugs is positively correlated to the precision of bug detection by the inspectors. This finding provides a new perspective on the role of this region on error awareness and monitoring and of its potential predictive value in predicting the quality of bug removing. João Durães, Henrique Madeira, João Castelhano, Isabel Catarina Duarte, Miguel Castelo-Branco |
ISSRE | 3 |