VLDB 2026 Research / reviewers in the wild / expert
Bruno Sotto-Mayor
dblp:300/9111
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-4450-6906ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spectrum-based fault diagnosis with partial tracesabstract• Demonstrate the feasibility of software fault diagnosis with partial traces and highlight the advantages of probabilistic reconstruction for fault localization when full execution traces are unavailable. • Propose algorithms, called Rec-Min, Rec-Max, and Rec-Weighted, that build on Spectrum-based Fault Localization (SFL), augmented by a trace reconstruction process that fills the missing parts of the program trace via static code analysis. • Evaluate the algorithms on 109 versions of 10 projects from the Defects4J dataset, showing significant fault localization improvements compared to a baseline approach. • Show that the Rec-Weighted algorithm achieves the best trade-off, reducing wasted effort by 34.4 %. Finding the root cause of observed software bugs, also referred to as software fault diagnosis, is a challenging problem that often must be solved in order to maintain software reliability. Spectrum-Based Fault Localization (SFL) is a powerful technique for automated software fault diagnosis that leverages execution traces. However, having the exact trace of every executed function is not common in real-world software projects, since tracing incurs significant overhead. Instead, many projects maintain log files to track execution, from which only partial execution traces may be extracted. This raises a key challenge: how can we diagnose fault faults (bugs) correctly given only partial execution traces? One way to do so is to run SFL using only the available partial traces. This approach is very limited, as it ignores all components that are not in the partial trace, which may even include the buggy components we are looking for. To overcome this limitation, we propose to use trace reconstruction techniques to synthetically add components to the given partial traces. These techniques analyze possible execution paths, which are extracted via static code analysis. In this work, we consider three trace reconstruction techniques: Rec-Min, Rec-Max, and Rec-Weighted. Rec-Min adds to the partial traces components present in all possible execution paths. Rec-Max adds to the partial trace components that exists in at least one execution path. Rec-Weighted applies a more refined approach. It assigns likelihood scores to the missing components based on the number of execution paths that pass through them. Then, it adapts SFL to intelligently consider these likelihood scores. Empirical evaluation on 109 faulty versions of 10 projects from the Defects4J dataset demonstrates that using our trace reconstruction techniques significantly enhances fault localization compared to the baseline SFL that only uses the partial traces. In some cases, using Rec-Weighted enabled reducing the wasted effort by 34.4 %. Bruno Sotto-Mayor, Roni Stern, Meir Kalech |
J. Syst. Softw. | 1 |
| 2023 | CLEAN++: Code Smells Extraction for C++abstractThe extraction of features is an essential step in the process of mining software repositories. An important feature that has been actively studied in the field of mining software repositories is bad code smells. Bad code smells are patterns in the source code that indicate an underlying issue in the design and implementation of the software. Several tools have been proposed to extract code smells. However, currently, there are no tools that extract a significant number of code smells from software written in C++. Therefore, we propose CLEAN++ (Code smeLls ExtrActioN for C++) [1]. It is an extension of a robust static code analysis tool that implements 35 code smells. To evaluate CLEAN++, we ran it over 44 open-source projects and wrote test cases to validate each code smell. Also, we converted the test cases to Java and used two Java tools to validate the effectiveness of our tool. In the end, we confirmed that the CLEAN++ is successful at detecting code smells.The tool is available at https://github.com/Tomma94/CLEAN-Plus-Plus. Tom Mashiach, Bruno Sotto-Mayor, Gal A. Kaminka, Meir Kalech |
MSR | 2 |
| 2023 | Spectrum-based feature localization for families of systemsabstractIn large code bases, locating the elements that implement concrete features of a system is challenging. This information is paramount for maintenance and evolution tasks, although not always explicitly available. In this work, motivated by the needs of locating features as a first step for feature-based Software Product Line adoption, we propose a solution for improving the performance of existing approaches. For this, relying on an automatic feature localization approach to locate features in single-systems, we propose approaches to deal with feature localization in the context of families of systems, e.g., variants created through opportunistic reuse such as clone-and-own. Our feature localization approaches are built on top of Spectrum-based feature localization (SBFL) techniques, supporting both dynamic feature localization (i.e., using execution traces as input) and static feature localization (i.e., relying on the structural decomposition of the variants’ implementation). Concretely, we provide (i) a characterization of different settings for dynamic SBFL in single systems, (ii) an approach to improve accuracy of dynamic SBFL for families of systems, and (iii) an approach to use SBFL as a static feature localization technique for families of systems. The proposed approaches are evaluated using the consolidated ArgoUML SPL feature localization benchmark. The results suggest that some settings of SBFL favor precision such as using the ranking metrics Wong2, Ochiai2, or Tarantula with high threshold values, while most of the ranking metrics with low thresholds favor recall. The approach to use information from variants increase the precision of dynamic SBFL while maintaining recall even with few number of variants, namely two or three. Finally, the static SBFL approach performs equally in terms of accuracy to other state-of-the-art approaches, such as Formal Concept Analysis and Interdependent Elements. Gabriela Karoline Michelon, Jabier Martinez, Bruno Sotto-Mayor, Aitor Arrieta, Wesley K. G. Assunção, Rui Abreu 0001, Alexander Egyed |
J. Syst. Softw. | 3 |
| 2022 | Exploring Design smells for smell-based defect prediction
Bruno Sotto-Mayor, Amir Elmishali, Meir Kalech, Rui Abreu 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2021 | BEIRUT: Repository Mining for Defect PredictionabstractSoftware Defect Prediction is an important activity used in the Testing Phase of the software development life cycle. Within the research of new defect prediction approaches and the selection of training sets for the classification task, different benchmarks have been analyzed in the literature. They provide several features and defective information over specific software archives. Therefore, they are commonly used in research to evaluate new approaches. However, the current benchmarks contain several limitations, such as lack of project variability, outdated benchmarks, single-version projects, a small number of projects and metrics, unavailable resources, poor usability, and non-extensible tools. Therefore, we introduce a novel tool Bgu rEpository mlning foR bUg predicIion (BEIRUT) for benchmark generation for defect prediction, composed of three main features: Given an open-source repository from GitHub, BEIRUT mines the software repository by (1) selecting the best$k$versions, based on the defective rate of each version, (2) generating training sets and a testing set for defect prediction, composed of a large number of metrics and defective information extracted from each of the selected versions and (3) creating defect prediction models from those extracted metrics. In the end, BEIRUT extracts a diversified catalog of 644 metrics and the defective information from each component of$k$versions, automatically selected based on the rate of defects in each version. They were collected from 512 different projects, starting from 2009. The tool is also supplemented with an easy-to-use web interface that provides a configurable selection of projects and metrics and an interface to manage the defect prediction tasks. Moreover, this tool is adapted to be extended with new projects and new extractors, introducing new metrics to the benchmark. The web service tool can be found at rps.ise.bgu.ac.il/beirut. Amir Elmishali, Bruno Sotto-Mayor, Inbal Roshanski, Amit Sultan, Meir Kalech |
ISSRE | 2 |
| 2021 | Cross-project smell-based defect prediction
Bruno Sotto-Mayor, Meir Kalech |
Soft Comput. | 1 |