João Durães

dblp:31/2314 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0002-9697-9991ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 4 first-author · 4 since 2021Security and privacy · 9 · 4 first-authorSystems, architecture and hardware · 5 · 2 first-author
YearPublicationVenuePosition
2026 Complementarity in software code complexity metrics
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira
J. Syst. Softw.4
2025 NRevisit: A Cognitive Behavioral Metric for Code Understandability Assessment
abstract
Measuring code understandability is both highly relevant and exceptionally challenging. This paper proposes a dynamic code understandability assessment method, which estimates a personalized code understandability score from the perspective of the specific programmer handling the code. The method consists of dynamically dividing the code unit under development or review in code regions (invisible to the programmer) and using the number of revisits (NRevisit) to each region as the primary feature for estimating the code understandability score. This approach removes the uncertainty related to the concept of a "typical programmer" assumed by static software code complexity metrics and can be easily implemented using a simple, low-cost, and non-intrusive desktop eye tracker or even a standard computer camera. This metric was evaluated using cognitive load measured through electroencephalography (EEG) in a controlled experiment with 35 programmers. Results show a very high correlation ranging from rs = 0.9067 to rs = 0.9860 (with p nearly 0) between the scores obtained with different alternatives of NRevisit and the ground truth represented by the EEG measurements of programmers’ cognitive load, demonstrating the effectiveness of our approach in reflecting the cognitive effort required for code comprehension. The paper also discusses possible practical applications of NRevisit, including its use in the context of AI-generated code, which is already widely used today.
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira
EASE4
2025 No Vibe Without Comprehension: Measuring Code Understanding in Modern Coding Workflows Using Neurophysiological Signals
abstract
Code comprehension assessment is crucial in modern software engineering contexts, such as the emerging LLM-supported programming paradigm, where evaluating and adjusting LLM-generated code to ensure suitability, correctness, and readability is mandatory. Recent literature offers various code comprehension solutions, ranging from subjective surveys to neurophysiological-based approaches that are more personalized and operational. However, existing proposals often estimate the cognitive load experienced by programmers during code handling, using this measure as a surrogate for code comprehension. This approach has limitations: it is indirect, as other factors influence cognitive load, and a high cognitive load does not necessarily indicate a lack of code understanding. In this paper, we propose a neurophysiological and AI-based solution using a multimodal set of biosensors, including EEG and eyetracking, along with other contextual features to measure the level of code comprehension. Instead of using cognitive load as a surrogate for code comprehension, this work tackles the challenge of assessing code comprehension by employing performance-annotated ground truth. The solution is customizable, allowing adaptation to different industrial requirements, such as stringent safety and reliability needs in mission-critical software or less critical contexts. We analyze various application scenarios to minimize the intrusiveness of the solution while maintaining acceptable performance. Evaluated in a controlled experiment with 50 programmers and 7 code comprehensions tasks, the porposed solution achieved an accuracy of 69% in the prediction of correct code comprehension. This binary modelling achieve an AUC of 75%, demonstrating its viability for measuring code comprehension in modern software development. We believe that such comprehension assessment methods are essential in current “vibe coding” workflows, where AI tools assist programmers interactively, and code understanding levels must be monitored in real-time to ensure effective human-AI collaboration.
Ricardo Saraiva, João Durães, Paulo Carvalho 0001, Henrique Madeira, Haytham Hijazi
ISSRE2
2023 Quality Evaluation of Modern Code Reviews Through Intelligent Biometric Program Comprehension
abstract
Code review is an essential practice in software engineering to spot code defects in the early stages of software development. Modern code reviews (e.g., acceptance or rejection of pull requests with Git) have become less formal than classic Fagan's inspections, lightweight, and more reliant on individuals (i.e., reviewers). However, reviewers may encounter mentally demanding challenges during the code review, such as code comprehension difficulties or distractions that might affect the code review quality. This work proposes a novel approach that evaluates the quality of code reviews in terms of bug-finding effectiveness and provides the reviewers with a clear message of whether the review should be repeated, indicating the code regions that may not have been well-reviewed. The proposed approach utilizes biometric information collected from the reviewer during the review process using non-intrusive biofeedback devices (e.g., smartwatches). Biometric measures such as Heart Rate Variability (HRV) and task-evoked pupillary response are captured as a surrogate of the cognitive state of the reviewer (e.g., mental workload) and inexpensive desktop eye-trackers compatible with the software development settings. This work uses Artificial Intelligence techniques to predict the cognitive load from the extracted biomarkers and classify each code region according to a set of features. The final evaluation considers various factors such as code complexity, time of the code review, the experience level of the reviewer, and other factors. Our experimental results show the approach could predict the review quality with 87.77%±4.65 accuracy and a Spearman correlation coefficient of 0.85 (p-value < 0.001) between the predicted and the actual review performance. This evaluation validates the cognitive load measurement using electroencephalography (EEG) signals as ground truth for the HRV and pupil signals.
Haytham Hijazi, João Durães, Ricardo Couceiro, João Castelhano, Raul Barbosa, Júlio Medeiros, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
IEEE Trans. Software Eng.2
2022 Bi-objective optimization of availability and cost for cloud services
abstract
Cloud-based services are a current approach for developing large-scale applications with advantages such as flexibility, access to on-demand resources, and business agility. The overall application functionality results from complex interactions of many decoupled services, each having its operational specificity. Due to this complexity, the manual configuration of these systems is very arduous, error-prone and likely to impair the quality of service, leading to malfunctioning services, lowering availability and accruing costs. Identifying the optimal solution to simultaneously optimize availability and costs, whilst meeting service level objectives remains a challenge for professionals developing solutions using cloud services. This paper proposes a mathematical formulation of a bi-objective problem to identify the optimal set of solutions for the system configuration. Empirical evaluation of the proposed approach in a case study of a real industrial scenario results in an R-Squared of 0.85, an MSE of 0.021 and an optimization accuracy of 0.928. These methods can help practitioners to keep services at an optimum configuration enabling autonomic service operation, whilst improving availability and cost.
André Bento, João Durães, José Ferreira, Rita Carreira, Filipe Araújo, Raul Barbosa
NCA4
2021 A layered framework for root cause diagnosis of microservices
abstract
Microservice-based architectures feature function-ally independent, well-defined and fine-grained components suit-able for loosely coupled deployments and for building reli-able cloud-native applications. Despite the advantages of this approach, component interactions introduce complexity, thus turning boundary -spanning service operation into a daunting challenge. As systems grow in size, complexity can easily outgrow the cognitive capacity of human operators, who are unable to effectively diagnose faulty microservices. We address this problem by proposing a novel framework to diagnose faulty microservices. Through failure injection and an experimental assessment, our layered diagnosis framework using service response analysis, timing constraints, causality and a ranking algorithm from traces, is able to effectively diagnose faulty microservices. Empirical evaluation of the proposed approach, by examining 130 experi-ments in a representative microservice application in the presence of faults, shows that it can achieve approximately 89% specificity and 77% recall.
André Bento, Jaime Correia, João Durães, Luís Ribeiro, Rita Carreira, Filipe Araújo, Raul Barbosa
NCA3
2019 Pupillography as Indicator of Programmers' Mental Effort and Cognitive Overload
abstract
Our research explores a recent paradigm called Biofeedback Augmented Software Engineering (BASE) that introduces a strong new element in the software development process: the programmers' biofeedback. In this Practical Experience Report we present the results of an experiment to evaluate the possibility of using pupillography to gather biofeedback from the programmers. The idea is to use pupillography to get meta information about the programmers' cognitive and emotional states (stress, attention, mental effort level, cognitive overload,...) during code development to identify conditions that may precipitate programmers making bugs or bugs escaping human attention, and tag the corresponding code locations in the software under development to provide online warnings to the programmer or identify code snippets that will need more intensive testing. The experiments evaluate the use of pupillography as cognitive load predictor, compare the results with the mental effort perceived by programmers using NASATLX, and discuss different possibilities for the use of pupillography as biofeedback sensor in real software development scenarios.
Ricardo Couceiro, Gonçalo Duarte, João Durães, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
DSN3
2019 Spotting Problematic Code Lines using Nonintrusive Programmers' Biofeedback
abstract
Recent studies have shown that programmers' cognitive load during typical code development activities can be assessed using wearable and low intrusive devices that capture peripheral physiological responses driven by the autonomic nervous system. In particular, measures such as heart rate variability (HRV) and pupillography can be acquired by nonintrusive devices and provide accurate indication of programmers' cognitive load and attention level in code related tasks, which are known elements of human error that potentially lead to software faults. This paper presents an experimental study designed to evaluate the possibility of using HRV and pupillography together with eye tracking to identify and annotate specific code lines (or even finer grain lexical tokens) of the program under development (or under inspection) with information on the cognitive load of the programmer while dealing with such lines of code. The experimental data is discussed in the paper to assess different alternatives for using code annotations representing programmers' cognitive load while producing or reading code. In particular, we propose the use of biofeedback code highlighting techniques to provide online programmer's warnings for potentially problematic code lines that may need a second look at (to remove possible bugs), and biofeedback-driven software testing to optimize testing effort, focusing the tests on code areas with higher bug probability.
Ricardo Couceiro, Paulo Carvalho 0001, Miguel Castelo-Branco, Henrique Madeira, Raul Barbosa, João Durães, Gonçalo Duarte, João Castelhano, Isabel Catarina Duarte, César Alexandre Teixeira, Nuno Laranjeiro, Júlio Medeiros
ISSRE6
2016 WAP: Understanding the Brain at Software Debugging
abstract
We propose that understanding functional patterns of activity in mapped brain regions associated with code comprehension tasks and, more specifically, to the activity of finding bugs in traditional code inspections could reveal useful insights to improve software reliability and to improve the software development process in general. This includes helping to select the best professionals for the debugging effort, improving the conditions for code inspections, and identify new directions to follow for training code reviewers. This paper presents an interdisciplinary study to analyze the brain activity during code inspection tasks using functional magnetic resonance imaging (fMRI), which is a well-established tool in cognitive neuroscience research. We used several programs where realistic bugs representing the most frequent types of software faults found in the field were injected. The code inspectors involved in the research include programmers with different levels of expertise and experience in real code reviews. The goal is to understand brain activity patterns associated with code comprehension tasks and, more specifically, the brain activity when the code reviewer identifies a bug in the code ('eureka' moment), which can be a true positive or a false positive. Our results confirmed that brain areas associated with language processing and mathematics are highly active during code reviewing and shows that there are specific brain activity patterns that can be related to the decision-making moment of suspicion/bug detection. Importantly, the activity at the anterior insula region that we find to play a relevant role in the process of identifying software bugs is positively correlated to the precision of bug detection by the inspectors. This finding provides a new perspective on the role of this region on error awareness and monitoring and of its potential predictive value in predicting the quality of bug removing.
João Durães, Henrique Madeira, João Castelhano, Isabel Catarina Duarte, Miguel Castelo-Branco
ISSRE1
2014 Security Benchmarks for Web Serving Systems
abstract
The security of software-based systems is one of the most difficult issues when accessing the suitability of systems to most application scenarios. However, security is very hard to evaluate and quantify, and there are no standard methods to benchmark the security of software systems. This work proposes a novel methodology for benchmarking the security of software-based systems. This methodology uses the notion of risk in a quantifiable way and allows the comparison of functionally-equivalent systems (or different configurations of the same system) to enable users and system integrators to identify and select the most secure one. The benchmark methodology is based on both analytical and experimental steps and can be applicable to any software system. The benchmark procedures and rules guide users on how to instantiate the methodology to specific scenarios and how to execute the benchmark. In this paper we also present an instantiation of the methodology to a case study of web-serving systems and show how to use the results to identify the most secure system under benchmark.
Naaliel Mendes, Henrique Madeira, João Durães
ISSRE3
2013 On Fault Representativeness of Software Fault Injection
abstract
The injection of software faults in software components to assess the impact of these faults on other components or on the system as a whole, allowing the evaluation of fault tolerance, is relatively new compared to decades of research on hardware fault injection. This paper presents an extensive experimental study (more than 3.8 million individual experiments in three real systems) to evaluate the representativeness of faults injected by a state-of-the-art approach (G-SWFIT). Results show that a significant share (up to 72 percent) of injected faults cannot be considered representative of residual software faults as they are consistently detected by regression tests, and that the representativeness of injected faults is affected by the fault location within the system, resulting in different distributions of representative/nonrepresentative faults across files and functions. Therefore, we propose a new approach to refine the faultload by removing faults that are not representative of residual software faults. This filtering is essential to assure meaningful results and to reduce the cost (in terms of number of faults) of software fault injection campaigns in complex software. The proposed approach is based on classification algorithms, is fully automatic, and can be used for improving fault representativeness of existing software fault injection approaches.
Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira
IEEE Trans. Software Eng.3
2010 Representativeness analysis of injected software faults in complex software
abstract
Despite of the existence of several techniques for emulating software faults, there are still open issues regarding representativeness of the faults being injected. An important aspect, not considered by existing techniques, is the non-trivial activation condition (trigger) of real faults, which causes them to elude testing and remain hidden until operation. In this paper, we investigate how the representativeness of injected software faults can be improved regarding the representativeness of triggers, by proposing a set of generic criteria to select representative faults from afaultload. We used the G-SWFIT technique to inject software faults in a DBMS, resulting in over 40 thousands faults and 2 million runs of a real test suite. We analyzed faults with respect to their triggers, and concluded that a non-negligible share (15%) would not realistically elude testing. Our proposed criteria decreased the percentage of non-elusive faults in the faultload, improving its representativeness.
Roberto Natella, Domenico Cotroneo, João Durães, Henrique Madeira
DSN3
2010 Towards Identifying the Best Variables for Failure Prediction Using Injection of Realistic Software Faults
abstract
Predicting failures at runtime is one of the most promising techniques to increase the availability of computer systems. However, failure prediction algorithms are still far from providing satisfactory results. In particular, the identification of the variables that show symptoms of incoming failures is a difficult problem. In this paper we propose an approach for identifying the most adequate variables for failure prediction. Realistic software faults are injected to accelerate the occurrence of system failures and thus generate a large amount of failure related data that is used to select, among hundreds of system variables, a small set that exhibits a clear correlation with failures. The proposed approach was experimentally evaluated using two configurations based on Windows XP. Results show that the proposed approach is quite effective and easy to use and that the injection of software faults is a powerful tool for improving the state of the art on failure prediction.
Ivano Irrera, João Durães, Marco Vieira, Henrique Madeira
PRDC2
2008 Assessing and Comparing Security of Web Servers
abstract
This paper presents an approach to assess security of web servers. This method can be used to compare the security features of different web servers installations and to determine how secure a given web server configuration is. The assessment is done by applying a set of tests designed to check if the system under evaluation fulfils a set of security practices defined by an extensive field study. This work targets the most typical issues related to web servers ranging from classic web servers misconfiguration to the absence of a secure network infrastructure and of well-defined security policies to respond to security incidents. The effectiveness and usefulness of the proposed approach is illustrated through the security assessment and comparison of five different real web servers.
Naaliel Mendes, Afonso Araújo Neto, João Durães, Marco Vieira, Henrique Madeira
PRDC3
2007 Experimental Risk Assessment and Comparison Using Software Fault Injection
abstract
One important question in component-based software development is how to estimate the risk of using COTS components, as the components may have hidden faults and no source code available. This question is particularly relevant in scenarios where it is necessary to choose the most reliable COTS when several alternative components of equivalent functionality are available. This paper proposes a practical approach to assess the risk of using a given software component (COTS or non-COTS). Although we focus on comparing components, the methodology can be useful to assess the risk in individual modules. The proposed approach uses the injection of realistic software faults to assess the impact of possible component failures and uses software complexity metrics to estimate the probability of residual defects in software components. The proposed approach is demonstrated and evaluated in a comparison scenario using two real off-the-shelf components (the RTEMS and the RTLinux real-time operating system) in a realistic application of a satellite data handling application used by the European Space Agency.
Regina Lúcia de Oliveira Moraes, João Durães, Ricardo Barbosa 0003, Eliane Martins, Henrique Madeira
DSN2
2006 Emulation of Software Faults: A Field Data Study and a Practical Approach
abstract
The injection of faults has been widely used to evaluate fault tolerance mechanisms and to assess the impact of faults in computer systems. However, the injection of software faults is not as well understood as other classes of faults (e.g., hardware faults). In this paper, we analyze how software faults can be injected (emulated) in a source-code independent manner. We specifically address important emulation requirements such as fault representativeness and emulation accuracy. We start with the analysis of an extensive collection of real software faults. We observed that a large percentage of faults falls into well-defined classes and can be characterized in a very precise way, allowing accurate emulation of software faults through a small set of emulation operators. A new software fault injection technique (G-SWFIT) based on emulation operators derived from the field study is proposed. This technique consists of finding key programming structures at the machine code-level where high-level software faults can be emulated. The fault-emulation accuracy of this technique is shown. This work also includes a study on the key aspects that may impact the technique accuracy. The portability of the technique is also discussed and it is shown that a high degree of portability can be achieved
João Durães, Henrique Madeira
IEEE Trans. Software Eng.1
2004 Generic Faultloads Based on Software Faults for Dependability Benchmarking
abstract
The most critical component of a dependability benchmark is the faultload, as it should represent a repeatable, portable, representative, and generally accepted set of faults. These properties are essential to achieve the desired standardization level required by a dependability benchmark but, unfortunately, are very hard to achieve. This is particularly true for software faults, which surely accounts for the fact that this important class of faults has never been used in known dependability benchmark proposals. This paper proposes a new methodology for the definition of faultloads based on software faults for dependability benchmarking. Faultload properties such as repeatability, portability and scalability are also analyzed and validated through experimentation using a case study of dependability benchmarking of Web-servers. We concluded that software fault-based faultloads generated using our methodology are appropriate and useful for dependability benchmarking. As our methodology is not tied to any specific software vendor or platform, it can be used to generate faultloads for the evaluation of any software product such as OLTP systems.
João Durães, Henrique Madeira
DSN1
2004 Dependability Benchmarking of Web-Servers
João Durães, Marco Vieira, Henrique Madeira
SAFECOMP1
2003 Definition of Software Fault Emulation Operators: A Field Data Study
abstract
This paper proposes a set of operators for software fault emulation through low-level code mutations. The definition of these operators was based on the analysis of an extensive collection of real software faults. Using the Orthogonal Defect Classification as a starting point, faults were classified in a detailed manner according to the high-level constructs where the faults reside and their effects in the program. We observed that a large percentage of faults fall in well-defined classes and can be characterized in a very precise way, allowing accurate emulation through a small set of mutation operators. The resulting operators closely emulate a broad range of common programmer mistakes. Furthermore, as the mutation is performed directly at the executable code, software faults can be injected in targets for which source code is not available.
João Durães, Henrique Madeira
DSN1
2002 Emulation of Software Faults by Educated Mutations at Machine-Code Level
abstract
This paper proposes a new technique to emulate software faults by educated mutations introduced at the machine-code level and presents an experimental study on the accuracy of the injected faults. The proposed method consists of finding key programming structures at the machine code-level where high-level software faults can be emulated. The main advantage of emulating software faults at the machine-code level is that software faults can be injected even when the source code of the target application is not available, which is very important for the evaluation of COTS components or for the validation of software fault tolerance techniques in COTS based systems. The technique was evaluated using several real programs and different types of faults and, additionally, it includes our study on the key aspects that may impact on the technique accuracy. The portability of the technique is also addressed. The results show that classes of faults such as assignment, checking, interface, and simple algorithm faults can be directly emulated using this technique.
João Durães, Henrique Madeira
ISSRE1
2002 Characterization of Operating Systems Behavior in the Presence of Faulty Drivers through Software Fault Emulation
abstract
This paper proposes a practical way to evaluate the behavior of commercial-off-the-shelf (COTS) operating systems in the presence of faulty device drivers. The proposed method is based on the emulation of software faults in target device drivers and the observation of the behavior of the system and of a workload regarding a comprehensive set of failure modes analyzed according to different dimensions. The emulation of software faults itself is done through the injection at machine-code level of selected mutations that represent the code produced when typical programming errors are made in the high-level language code. An important aspect of the proposed methodology is the use of simple and established practices to evaluate operating systems failure modes, thus allowing its use as a dependability benchmarking technique. The generalization of the methodology to any software system built of discrete and identifiable components is also discussed.
João Durães, Henrique Madeira
PRDC1