VLDB 2026 Research / reviewers in the wild / expert
João R. Campos
dblp:227/2492
· DBLP profile ↗
13ranked-venue papers
5as first author
10since 2021 · last 2027
0000-0002-4623-764XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 first-author · 6 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | PROBE: Benchmarking code generation in large language modelsabstractAbstract Large Language Models (LLMs) are increasingly being used in everyday software engineering tasks, particularly in automated code generation. Despite their widespread adoption, these models remain far from perfect, making systematic and fair evaluation essential to understand their strengths and limitations. In the context of code generation, existing benchmarks are limited: they often target a single programming language and rely primarily on unit test outcomes, while overlooking other critical dimensions such as the overall quality of the generated code and its closeness to a valid solution. To address these gaps, we introduce , an extensible benchmark framework that, unlike prior work, establishes a systematic structure built on diverse and well-defined metrics, representative workloads, varied prompt templates, and a robust experimental procedure. In practice, the code generated by the LLMs is evaluated along three complementary dimensions: functional correctness, proximity to valid solutions, and code quality, enabling a comprehensive assessment of performance. We use to evaluate four open-source and two proprietary models under three prompting strategies across five programming languages. We further complement this analysis with a study of common errors in the code and provide concrete examples, offering clearer insight into where LLMs tend to struggle. Our findings show that, while LLMs achieve promising results, they struggle with harder problems and, in the case of smaller models, with programming languages that have fewer available resources for training, and they often fail due to fundamental and easily avoidable errors that underscore the unreliability of automatically generated code. Rodrigo Pato Nogueira, Marco Vieira, João R. Campos |
Empir. Softw. Eng. | 3 |
| 2026 | Leveraging Large Language Models for Trustworthiness Assessment of Web Applications
Oleksandr Yarotskyi, José D'Abruzzo Pereira, João R. Campos |
ICST | 3 |
| 2025 | Selecting a Data Warehouse Provider: A Daunting Task
Nuno Lourenço 0002, João R. Campos |
DATA | 3 |
| 2025 | Towards the Assessment of Task-based Chatbots: From the TOFU-R Snapshot to the BRASATO Curated DatasetabstractTask-based chatbots are increasingly being used to deliver real services, yet assessing their reliability, security, and robustness remains underexplored, also due to the lack of large-scale, high-quality datasets. The emerging automated quality assessment techniques targeting chatbots often rely on limited pools of subjects, such as custom-made toy examples, or outdated, no longer available, or scarcely popular agents, complicating the evaluation of such techniques. In this paper, we present two datasets and the tool support necessary to create and maintain these datasets. The first dataset is RASA TASK-BASED CHATBOTS FROM GITHUB (TOFU-R), which is a snapshot of the Rasa chatbots available on GitHub, representing the state of the practice in open-source chatbot development with Rasa. The second dataset is BOT RASA COLLECTION (BRASATO), a curated selection of the most relevant chatbots for dialogue complexity, functional complexity, and utility, whose goal is to ease reproducibility and facilitate research on chatbot reliability. Elena Masserini, Diego Clerissi, Daniela Micucci, João R. Campos, Leonardo Mariani |
ISSRE | 4 |
| 2025 | Beyond Functional Correctness: An Empirical Evaluation of Large Language Models for Text-to-Code GenerationabstractLarge Language Models (LLMs) have become increasingly popular for text-to-code generation, a task that involves converting natural language descriptions into code. While they have shown promising results, they are still flawed, especially when dealing with complex problems, making it crucial to have a deep understanding of their strengths and limitations. However, current evaluations are often limited, targeting only a single programming language, considering a small set of related models, and focusing heavily on execution-based metrics such as unit test pass rates. Moreover, they tend to overlook deeper issues like code quality, recurring mistakes, and underlying patterns in the generated code. To address these gaps and properly assess how effective LLMs are at generating quality code, we conduct a comprehensive evaluation of several LLMs across Python and C++ using a large, diverse dataset that spans multiple difficulty levels. Our evaluation combines execution-based and static analysis metrics, enabling us to assess both functional correctness and the structural quality of the code. Unlike prior work, our study also includes an in-depth analysis of the generated code, uncovering common mistakes and failure modes, allowing for a more complete and differentiated view of model capabilities. Results show that, despite their potential, open-source LLMs still struggle with complex tasks and make recurring systematic errors such as missing import statements and incorrect variable scoping. Rodrigo Pato Nogueira, Marco Vieira, João R. Campos |
ISSRE | 3 |
| 2025 | Toward using cyber threat intelligence with machine and deep learning for IoT security: a comprehensive study
Milton de Lima, Carlos Viana, Wellison R. M. Santos, Flávio Neves, João R. Campos, Fernando Aires 0001 |
J. Supercomput. | 5 |
| 2023 | Online Failure Prediction Through Fault Injection and Machine Learning: Methodology and Case StudyabstractOnline Failure Prediction (OFP) is a technique that attempts to predict incoming failures to mitigate their consequences. Machine Learning (ML) has been successfully used to create predictive models for OFP, but failures are rare, and thus failure data are typically not available. Fault injection has been accepted as a viable alternative to generate failure data. However, this raises several challenges, such as how to process the data to create and assess predictive models. The characteristics of fault injection campaigns (e.g., repeated/controlled experiments) and OFP (e.g., autocorrelation) require specific considerations. This work proposes a six-stage methodology for using failure data generated through fault injection to create accurate/representative models for OFP. It overviews the various phases, from generating and processing the data, to creating and assessing the performance of the models up to their deployment, while considering the intrinsic characteristics of the problem. As a case study, we apply the methodology to develop failure predictors for the Linux Operating System (OS). Results show that the proposed methodology led to accurate predictive models that could also generalize to failures that occur under different execution profiles, whilst using traditional techniques resulted in over-optimistic observations. João R. Campos, Ernesto Costa, Marco Vieira |
ISSRE | 1 |
| 2023 | Understanding the Forest: A Visualization Tool to Support Decision Tree AnalysisabstractDecision Trees (DTs) are one of the most widely used supervised Machine Learning algorithms. The algorithm constructs binary tree data structures that partition the data into smaller segments according to different rules. Hence, DTs can be used as a learning process of finding the optimal rules to separate and classify all items of a dataset. Since the algorithm relies on a decision process similar to rule-based decisions, they are easily interpretable. However, DTs can be difficult to analyse when dealing with large datasets and/or with multiple trees, i.e. ensembles. To ease the analysis and validation of these models, we developed a visual tool which includes a set of visualizations that overview and give details of a set of trees. Our tool aims to provide different perspectives over the same data and provide further insights on how decisions are being made. In this article, we overview our design process, present the different visualization models and their iterative validation. We present a use case in the telecommunications domain. In concrete, we use the visual tool to help understand how a model based on DTs decides which is the best channel (i.e., phonecall, e-mail, SMS) to contact a client. Catarina Maçãs, João R. Campos, Nuno Lourenço 0002 |
IV | 2 |
| 2023 | A Machine Learning driven Fault Tolerance Mechanism for UAVs' Flight ControllerabstractUnmanned Aerial Vehicles (UAVs) are susceptible to various hazards (e.g., software or hardware failures, communication failures, or security attacks) that may hinder mission completion or compromise safety by violating the separation minima (i.e., the minimum distance that must be maintained between UAVs in order to ensure safe and efficient operations). To address this issue, this paper proposes a new machine learning-based fault-tolerant mechanism for UAV flight controllers that tolerates GPS-related faults. These faults are of paramount importance (i.e., accidental faults and/or security attacks that eventually cause failures in the GPS function/data), as accurate positioning and tracking are essential to assure safe operation in UAVs. The proposed machine learning models were built using 884,410 data records from 1,985 flight logs collected from the PX4 public repository. The trained models are used to predict the expected position of the UAV during a mission, and separation minima are used as a threshold to detect the GPS hazards by comparing it with the distance between two consecutive position values. When a hazard is detected (i.e., the distance is higher than separation minima), the predicted values by machine learning models are fed into the flight controller’s position estimator (i.e., an Extended Kalman Filter (EKF)). To evaluate the effectiveness of this approach, validation experiments were conducted on several realistically defined missions while being exposed to different types of failure conditions (e.g., GPS signal loss or GPS Spoofing), both with and without using the proposed fault-tolerant mechanism. The results show a remarkable reduction in safety violations (the number of separation minima violations was reduced from 94 to 1). Additionally, the proposed mechanism demonstrated a notable improvement in the distance traveled by UAV and the duration of the flight mission in failure conditions, showing its ability to mitigate faults effectively. These findings support the effectiveness of the proposed fault tolerance mechanism in enhancing UAV safety in the presence of issues caused by GPS. Anamta Khan, João R. Campos, Naghmeh Ramezani Ivaki, Henrique Madeira |
PRDC | 2 |
| 2023 | Online Failure Prediction for Complex Systems: Methodology and Case StudiesabstractOnline Failure Prediction (OFP) allows proactively taking countermeasures before a failure occurs, such as saving data or restarting a system. However, despite its potential contribution to improving dependability, OFP still presents key limitations. Besides the problem of choosing the optimal set of features, assessing predictive models is complex and common procedures for supporting comparison are not available. There is, in fact, little work on developing and assessing failure predictors for complex systems. In this aricle, we present two extensive case studies on distinct Operating Systems (OSs), Linux and Windows, showing that it is possible to create models that can predict different types of incoming failures, highlighting various important considerations such as the operational requirements of the target system. To drive the case studies, we define a well-structured framework for a fair and sound assessment and comparison of alternative predictive solutions. It includes scenarios for choosing the most adequate metrics for the assessment, comparing alternative models, and selecting the best predictor, while considering the need to tolerate perturbations in the data. In practice, we show that, by following a well-defined process, it is possible to develop accurate failure predictors and establish a ranking of the models under evaluation in different scenarios and OSs. João R. Campos, Ernesto Costa, Marco Vieira |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2020 | Fault Injection to Generate Failure Data for Failure Prediction: A Case StudyabstractDue to the complexity of modern software, identifying every fault before deployment is extremely difficult or even not possible. Such residual faults can ultimately lead to failures, often incurring considerable risks or costs. Online Failure Prediction (OFP) is a fault-tolerance technique that attempts to predict the occurrence of failures in the near future and thus prevent/mitigate their consequences. Combined with recent technological developments, Machine Learning (ML) has been successfully used to create predictive models for OFP. However, as failures are rare events, failure data are often not available for building accurate models. Although fault injection has been accepted as a viable solution to generate realistic failure data, fault injectors are difficult to implement/update and thus research on Operating System (OS)-level OFP has become stale, with most works using data from outdated OSs. In this paper, we conduct a comprehensive fault injection campaign on an up-to-date Linux kernel and thoroughly study its behavior in the presence of faults. We then transform the data to explore and assess the predictive performance of various ML techniques for OFP. Finally, we study the influence of different OFP parameters (i.e., lead-time, prediction-window) and compare the results with existing related work. Results suggest that the various failures observed can be grouped into categories that can then be accurately predicted and distinguished by diverse ML models. João R. Campos, Ernesto Costa |
ISSRE | 1 |
| 2020 | On Configuring a Testbed for Dependability Experiments: Guidelines and Fault Injection Case Study
João R. Campos, Ernesto Costa, Marco Vieira |
SAFECOMP | 1 |
| 2019 | Propheticus: Machine Learning Framework for the Development of Predictive Models for Reliable and Secure SoftwareabstractThe growing complexity of software calls for innovative solutions that support the deployment of reliable and secure software. Machine Learning (ML) has shown its applicability to various complex problems and is frequently used in the dependability domain, both for supporting systems design and verification activities. However, using ML is complex and highly dependent on the problem in hand, increasing the probability of mistakes that compromise the results. In this paper, we introduce Propheticus, a ML framework that can be used to create predictive models for reliable and secure software systems. Propheticus attempts to abstract the complexity of ML whilst being easy to use and accommodating the needs of the users. To demonstrate its use, we present two case studies (vulnerability prediction and online failure prediction) that show how it can considerably ease and expedite a thorough ML workflow. João R. Campos, Marco Vieira, Ernesto Costa |
ISSRE | 1 |