Haytham Hijazi

dblp:286/4335 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-4981-3649ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Complementarity in software code complexity metrics
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira
J. Syst. Softw.2
2025 NRevisit: A Cognitive Behavioral Metric for Code Understandability Assessment
abstract
Measuring code understandability is both highly relevant and exceptionally challenging. This paper proposes a dynamic code understandability assessment method, which estimates a personalized code understandability score from the perspective of the specific programmer handling the code. The method consists of dynamically dividing the code unit under development or review in code regions (invisible to the programmer) and using the number of revisits (NRevisit) to each region as the primary feature for estimating the code understandability score. This approach removes the uncertainty related to the concept of a "typical programmer" assumed by static software code complexity metrics and can be easily implemented using a simple, low-cost, and non-intrusive desktop eye tracker or even a standard computer camera. This metric was evaluated using cognitive load measured through electroencephalography (EEG) in a controlled experiment with 35 programmers. Results show a very high correlation ranging from rs = 0.9067 to rs = 0.9860 (with p nearly 0) between the scores obtained with different alternatives of NRevisit and the ground truth represented by the EEG measurements of programmers’ cognitive load, demonstrating the effectiveness of our approach in reflecting the cognitive effort required for code comprehension. The paper also discusses possible practical applications of NRevisit, including its use in the context of AI-generated code, which is already widely used today.
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira
EASE2
2025 No Vibe Without Comprehension: Measuring Code Understanding in Modern Coding Workflows Using Neurophysiological Signals
abstract
Code comprehension assessment is crucial in modern software engineering contexts, such as the emerging LLM-supported programming paradigm, where evaluating and adjusting LLM-generated code to ensure suitability, correctness, and readability is mandatory. Recent literature offers various code comprehension solutions, ranging from subjective surveys to neurophysiological-based approaches that are more personalized and operational. However, existing proposals often estimate the cognitive load experienced by programmers during code handling, using this measure as a surrogate for code comprehension. This approach has limitations: it is indirect, as other factors influence cognitive load, and a high cognitive load does not necessarily indicate a lack of code understanding. In this paper, we propose a neurophysiological and AI-based solution using a multimodal set of biosensors, including EEG and eyetracking, along with other contextual features to measure the level of code comprehension. Instead of using cognitive load as a surrogate for code comprehension, this work tackles the challenge of assessing code comprehension by employing performance-annotated ground truth. The solution is customizable, allowing adaptation to different industrial requirements, such as stringent safety and reliability needs in mission-critical software or less critical contexts. We analyze various application scenarios to minimize the intrusiveness of the solution while maintaining acceptable performance. Evaluated in a controlled experiment with 50 programmers and 7 code comprehensions tasks, the porposed solution achieved an accuracy of 69% in the prediction of correct code comprehension. This binary modelling achieve an AUC of 75%, demonstrating its viability for measuring code comprehension in modern software development. We believe that such comprehension assessment methods are essential in current “vibe coding” workflows, where AI tools assist programmers interactively, and code understanding levels must be monitored in real-time to ensure effective human-AI collaboration.
Ricardo Saraiva, João Durães, Paulo Carvalho 0001, Henrique Madeira, Haytham Hijazi
ISSRE5
2023 Quality Evaluation of Modern Code Reviews Through Intelligent Biometric Program Comprehension
abstract
Code review is an essential practice in software engineering to spot code defects in the early stages of software development. Modern code reviews (e.g., acceptance or rejection of pull requests with Git) have become less formal than classic Fagan's inspections, lightweight, and more reliant on individuals (i.e., reviewers). However, reviewers may encounter mentally demanding challenges during the code review, such as code comprehension difficulties or distractions that might affect the code review quality. This work proposes a novel approach that evaluates the quality of code reviews in terms of bug-finding effectiveness and provides the reviewers with a clear message of whether the review should be repeated, indicating the code regions that may not have been well-reviewed. The proposed approach utilizes biometric information collected from the reviewer during the review process using non-intrusive biofeedback devices (e.g., smartwatches). Biometric measures such as Heart Rate Variability (HRV) and task-evoked pupillary response are captured as a surrogate of the cognitive state of the reviewer (e.g., mental workload) and inexpensive desktop eye-trackers compatible with the software development settings. This work uses Artificial Intelligence techniques to predict the cognitive load from the extracted biomarkers and classify each code region according to a set of features. The final evaluation considers various factors such as code complexity, time of the code review, the experience level of the reviewer, and other factors. Our experimental results show the approach could predict the review quality with 87.77%±4.65 accuracy and a Spearman correlation coefficient of 0.85 (p-value < 0.001) between the predicted and the actual review performance. This evaluation validates the cognitive load measurement using electroencephalography (EEG) signals as ground truth for the HRV and pupil signals.
Haytham Hijazi, João Durães, Ricardo Couceiro, João Castelhano, Raul Barbosa, Júlio Medeiros, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
IEEE Trans. Software Eng.1
2021 iReview: an Intelligent Code Review Evaluation Tool using Biofeedback
abstract
Code reviews and software inspections are essential for building reliable software. However, current code reviews practice in the software industry (e.g., acceptance or rejection of pull requests with Git) deviates considerably from classic (and expensive) Fagan's inspections. Modern code reviews are lightweight and asynchronous and do not rely on a group of inspectors and inspection meetings any longer. The modern style of code reviews is much more flexible and cost-effective. Still, these advantages come with the price of reducing the quality of code reviews, as a single reviewer generally makes them with all the inherent and the very human limitations of one single look. The reviewer could be distracted, overloaded, under stress, or not even fully understand the code under review. This paper proposes a new tool (iReview) that evaluates the code review quality using biometric measures gathered from code reviewers (often called Biofeedback). Biometric measures such as Heart Rate Variability (HRV) and eye movement dynamics are used to assess the reviewer's comprehension of the code under review. iReview evaluates the quality of each review globally and indicates the code regions that have not been well-reviewed, explaining why those code regions should be reviewed again. The tool uses Artificial Intelligence techniques to classify the code regions into good and bad reviews based on various biometric and non-biometric features. The first results show that iReview can predict the review quality of medium or complex programs with an accuracy ranging from 75% to 87% in detecting bad reviews (i.e., code regions classified as bad reviewed still have undetected bugs). This tool is expected to improve software reliability by ensuring that good reviews have been carried out despite the current lightweight reviewing processes.
Haytham Hijazi, José Cruz, João Castelhano, Ricardo Couceiro, Miguel Castelo-Branco, Paulo Carvalho 0001, Henrique Madeira
ISSRE1