EDBT 2026 Demo / reviewers in the wild / expert
Barbara Russo
dblp:r/BarbaraRusso
· DBLP profile ↗
52ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-3737-9264ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 47 · 4 first-author · 18 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How not to get your paper rejected - From the editors' notebook
Miroslaw Staron, Guilherme Horta Travassos, Barbara Russo, Sudipto Ghosh 0001 |
Inf. Softw. Technol. | 3 |
| 2026 | From attack descriptions to vulnerabilities: A sentence transformer-based approachabstractIn the domain of security, vulnerabilities frequently remain undetected even after their exploitation. In this work, vulnerabilities refer to publicly disclosed flaws documented in Common Vulnerabilities and Exposures (CVE) reports. Establishing a connection between attacks and vulnerabilities is essential for enabling timely incident response, as it provides defenders with immediate, actionable insights. However, manually mapping attacks to CVEs is infeasible, thereby motivating the need for automation. This paper evaluates 14 state-of-the-art (SOTA) sentence transformers for automatically identifying vulnerabilities from textual descriptions of attacks. Our results demonstrate that the multi-qa-mpnet-base-dot-v1 (MMPNet) model achieves superior classification performance when using attack Technique descriptions, with an F 1 -score of 89.0, precision of 84.0, and recall of 94.7. Furthermore, it was observed that, on average, 56% of the vulnerabilities identified by the MMPNet model are also represented within the CVE repository in conjunction with an attack, while 61% of the vulnerabilities detected by the model correspond to those cataloged in the CVE repository. A manual inspection of the results revealed the existence of 275 predicted links that were not documented in the MITRE repositories. Consequently, the automation of linking attack techniques to vulnerabilities not only enhances the detection and response capabilities related to software security incidents but also diminishes the duration during which vulnerabilities remain exploitable, thereby contributing to the development of more secure systems. • Automated vulnerability detection using 14 state-of-the-art sentence transformer models. • multi-qa-mpnet-base-dot-v1 achieves an F1 score of 89.0, outperforming other models (e.g., MiniLM, BERT). • Attack technique information yields the highest accuracy (57.3%) for vulnerability prediction. • 275 missing links between attack techniques and vulnerabilities identified through manual validation. • Open-source code and dataset provided to enable reproducible mapping from attacks to CVEs. Refat Othman, Diaeddin Rimawi, Bruno Rossi 0001, Barbara Russo |
J. Syst. Softw. | 4 |
| 2026 | Beyond Balance-Addressing class imbalance in fine-tuning deep learnersabstract• Extends the prior NLBSE tool competition work with an additional model family (SetFit). • Transforms the prototype into a reusable tool ( Beyond Balance ), providing easy-to-use access to the implemented loss-weighting strategies. • Compares the performance impact of loss-weighting across various strategies and model families. • Evaluates the generalizability of loss-weighting across two benchmark datasets from the NLBSE’24 and NLBSE’25 tool competitions. Datasets often contain heavily underrepresented classes. Class imbalance biases models toward frequent classes, reducing performance on rare but important categories; in-process strategies such as loss-weighting remain under-explored for software engineering artefacts. We investigate loss-weighting functions for code comment classification and package our methods into Beyond Balance , a reusable implementation offering multiple weighting strategies for Transformer- and Sentence-Transformer–based models. Loss weighting consistently improves F1 performance across datasets, demonstrating an effective and easily adoptable imbalance-handling technique through Beyond Balance . Moritz Mock, Thomas Borsani, Giuseppe Di Fatta, Barbara Russo |
Sci. Comput. Program. | 4 |
| 2025 | Leveraging Multi-Task Learning to Improve the Detection of SATD and VulnerabilityabstractMulti-task learning is a paradigm that leverages information from related tasks to improve the performance of machine learning. Self-Admitted Technical Debt (SATD) are comments in the code that indicate not-quite-right code introduced for short-term needs, i.e., technical debt (TD). Previous research has provided evidence of a possible relationship between SATD and the existence of vulnerabilities in the code. In this work, we investigate if multi-task learning could leverage the information shared between SATD and vulnerabilities to improve the automatic detection of these issues. To this aim, we implemented VulSATD, a deep learner that detects vulnerable and SATD code based on CodeBERT, a pre-trained transformers model. We evaluated VulSATD on MADE-WIC, a fused dataset of functions annotated for TD (through SATD) and vulnerability. We compared the results using single and multi-task approaches, obtaining no significant differences even after employing a weighted loss. Our findings indicate the need for further investigation into the relationship between these two aspects of low-quality code. Specifically, it is possible that only a subset of technical debt is directly associated with security concerns. Therefore, the relationship between different types of technical debt and software vulnerabilities deserves future exploration and a deeper understanding. Barbara Russo, Jorge Melegati, Moritz Mock |
ICPC | 1 |
| 2024 | Modeling Resilience of Collaborative AI SystemsabstractA Collaborative Artificial Intelligence System (CAIS) performs actions in collaboration with the human to achieve a common goal. CAISs can use a trained AI model to control human-system interaction, or they can use human interaction to dynamically learn from humans in an online fashion. In online learning with human feedback, the AI model evolves by monitoring human interaction through the system sensors in the learning state, and actuates the autonomous components of the CAIS based on the learning in the operational state. Therefore, any disruptive event affecting these sensors may affect the AI model's ability to make accurate decisions and degrade the CAIS performance. Consequently, it is of paramount importance for CAIS managers to be able to automatically track the system performance to understand the resilience of the CAIS upon such disruptive events. In this paper, we provide a new framework to model CAIS performance when the system experiences a disruptive event. With our framework, we introduce a model of performance evolution of CAIS. The model is equipped with a set of measures that aim to support CAIS managers in the decision process to achieve the required resilience of the system. We tested our framework on a real-world case study of a robot collaborating online with the human, when the system is experiencing a disruptive event. The case study shows that our framework can be adopted in CAIS and integrated into the online execution of the CAIS activities. Diaeddin Rimawi, Antonio Liotta, Marco Todescato, Barbara Russo |
CAIN | 4 |
| 2024 | Cybersecurity Defenses: Exploration of CVE Types Through Attack DescriptionsabstractVulnerabilities in software security can remain undiscovered even after being exploited. Linking attacks to vulnerabilities helps experts identify and respond promptly to the incident. This paper introduces VULDAT, a classification tool using a sentence transformer MPNET to identify system vulnerabilities from attack descriptions. Our model was applied to 100 attack techniques from the ATT&CK repository and 685 issues from the CVE repository. Then, we compare the performance of VULDAT against the other eight state-of-the-art classifiers based on sentence transformers. Our findings indicate that our model achieves the best performance with F1 score of 0.85, Precision of 0.86, and Recall of 0.83. Furthermore, we found 56% of CVE reports vulnerabilities associated with an attack were identified by VULDAT, and 61% of identified vulnerabilities were in the CVE repository. Refat Othman, Bruno Rossi 0001, Barbara Russo |
SEAA | 3 |
| 2024 | A Comparison of Vulnerability Feature Extraction Methods from Textual Attack PatternsabstractNowadays, threat reports from cybersecurity vendors incorporate detailed descriptions of attacks within unstructured text. Knowing vulnerabilities that are related to these reports helps cybersecurity researchers and practitioners understand and adjust to evolving attacks and develop mitigation plans. This paper aims to aid cybersecurity researchers and practitioners in choosing attack extraction methods to enhance the monitoring and sharing of threat intelligence. In this work, we examine five feature extraction methods (TF-IDF, LSI, BERT, MiniLM, RoBERTa) and find that Term Frequency-Inverse Document Frequency (TF-IDF) outperforms the other four methods with a precision of 75% and an F1 score of 64%. The findings offer valuable insights to the cybersecurity community, and our research can aid cybersecurity researchers in evaluating and comparing the effectiveness of upcoming extraction methods. Refat Othman, Bruno Rossi 0001, Barbara Russo |
SEAA | 3 |
| 2024 | MADE-WIC: Multiple Annotated Datasets for Exploring Weaknesses In CodeabstractIn this paper, we present MADE-WIC, a large dataset of functions and their comments with multiple annotations for technical debt and code weaknesses leveraging different state-of-the-art approaches. It contains about 860K code functions and more than 2.7M related comments from 12 open-source projects. To the best of our knowledge, no such dataset is publicly available. MADE-WIC aims to provide researchers with a curated dataset on which to test and compare tools designed for the detection of code weaknesses and technical debt. As we have fused existing datasets, researchers have the possibility to evaluate the performance of their tools by also controlling the bias related to the annotation definition and dataset construction. The demonstration video can be retrieved at https://www.youtube.com/watch?v=GaQodPrcb6E. Moritz Mock, Jorge Melegati, Max Kretschmann, Nicolás E. Díaz Ferreyra, Barbara Russo |
ASE | 5 |
| 2023 | CAIS-DMA: A Decision-Making Assistant for Collaborative AI Systems
Diaeddin Rimawi, Antonio Liotta, Marco Todescato, Barbara Russo |
PROFES (1) | 4 |
| 2023 | GResilience: Trading Off Between the Greenness and the Resilience of Collaborative AI Systems
Diaeddin Rimawi, Antonio Liotta, Marco Todescato, Barbara Russo |
ICTSS | 4 |
| 2023 | Actor-Driven Decomposition of Microservices through Multi-level Scalability AssessmentabstractThe microservices architectural style has gained widespread acceptance. However, designing applications according to this style is still challenging. Common difficulties concern finding clear boundaries that guide decomposition while ensuring performance and scalability. With the aim of providing software architects and engineers with a systematic methodology, we introduce a novel actor-driven decomposition strategy to complement the domain-driven design and overcome some of its limitations by reaching a finer modularization yet enforcing performance and scalability improvements. The methodology uses a multi-level scalability assessment framework that supports decision-making over iterative steps. At each iteration, architecture alternatives are quantitatively evaluated at multiple granularity levels. The assessment helps architects to understand the extent to which architecture alternatives increase or decrease performance and scalability. We applied the methodology to drive further decomposition of the core microservices of a real data-intensive smart mobility application and an existing open-source benchmark in the e-commerce domain. The results of an in-depth evaluation show that the approach can effectively support engineers in (i) decomposing monoliths or coarse-grained microservices into more scalable microservices and (ii) comparing among alternative architectures to guide decision-making for their deployment in modern infrastructures that orchestrate lightweight virtualized execution units. Matteo Camilli, Carmine Colarusso, Barbara Russo, Eugenio Zimeo |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Microservices Integrated Performance and Reliability TestingabstractContinuous quality assurance for extra-functional properties of modern software systems is today a big challenge as their complexity is constantly increasing to satisfy market demands. This is the case of microservice systems. They provide high control on the scale of operation by means of fine-grained service decomposition, but this demands careful consideration of the relations between performance of individual microservices and service failures. Matteo Camilli, Antonio Guerriero, Andrea Janes, Barbara Russo, Stefano Russo 0001 |
AST | 4 |
| 2022 | WeakSATD: Detecting Weak Self-admitted Technical DebtabstractSpeeding up development may produce technical debt, i.e., not-quite-right code for which the effort to make it right increases with time as a sort of interest. Developers may be aware of the debt as they admit it in their code comments. Literature reports that such a self-admitted technical debt survives for a long time in a program, but it is not yet clear its impact on the quality of the code in the long term. We argue that self-admitted technical debt contains a number of different weaknesses that may affect the security of a program. Therefore, the longer a debt is not paid back the higher is the risk that the weaknesses can be exploited. To discuss our claim and rise the developers' awareness of the vulnerability of the self-admitted technical debt that is not paid back, we explore the self-admitted technical debt in the Chromium C-code to detect any known weaknesses. In this preliminary study, we first mine the Common Weakness Enumeration repository to define heuristics for the automatic detection and fix of weak code. Then, we parse the C-code to find self-admitted technical debt and the code block it refers to. Finally, we use the heuristics to find weak code snippets associated to self-admitted technical debt and recommend their potential mitigation to developers. Such knowledge can be used to prioritize self-admitted technical debt for repair. A prototype has been developed and applied to the Chromium code. Initial findings report that 55% of self-admitted technical debt code contains weak code of 14 different types. Barbara Russo, Matteo Camilli, Moritz Mock |
MSR | 1 |
| 2022 | Modeling Performance of Microservices Systems with Growth TheoryabstractCONTEXT: The microservices architectural style is gaining momentum in the IT industry. This style does not guarantee that a target system can continuously meet acceptable performance levels. The ability to study the violations of performance requirements and eventually predict them would help practitioners to tune techniques like dynamic load balancing or horizontal scaling to achieve the resilience property. OBJECTIVE: The goal of this work is to study the violations of performance requirements of microservices through time series analysis and provide practical instruments that can detect resilient and non-resilient microservices and possibly predict their performance behavior. METHOD: to model the occurrences of violations of performance requirements as a stochastic process. We applied our method to an in-vitro e-commerce benchmark and an in-production real-world telecommunication system. We interpreted the resulting growth models to characterize the microservices in terms of their transient performance behavior. RESULTS: Our empirical evaluation shows that, in most of the cases, the non-linear S-shaped growth models capture the occurrences of performance violations of resilient microservices with high accuracy. The bounded nature associated with this models tell that the performance degradation is limited and thus the microservice is able to come back to an acceptable performance level even under changes in the nominal number of concurrent users. We also detect cases where linear models represent a better description. These microservices are not resilient and exhibit constant growth and unbounded performance violations over time. The application of our methodology to a real in-production system identified additional resilience profiles that were not observed in the in-vitro experiments. These profiles show the ability of services to react differently to the same solicitation. We found that when a service is resilient it can either decrease the rate of the violations occurrences in a continuous manner or with repeated attempts (periodical or not). CONCLUSIONS: We showed that growth theory can be successfully applied to study the occurences of performance violations of in-vitro and in-production real-world systems. Furthermore, the cost of our model calibration heuristics, based on the mathematical expression of the selected non-linear growth models, is limited. We discussed how the resulting models can shed some light on the trend of performance violations and help engineers to spot problematic microservice operations that exhibit performance issues. Thus, meaningful insights from the application of growth theory have been derived to characterize the behavior of (non) resilient microservices operations. Matteo Camilli, Barbara Russo |
Empir. Softw. Eng. | 2 |
| 2022 | Scalability testing automation using multivariate characterization and detection of software performance antipatterns
Alberto Avritzer, Ricardo Britto 0001, Catia Trubiani, Matteo Camilli, Andrea Janes, Barbara Russo, André van Hoorn, Robert Heinrich, Martina Rapp, Jörg Henß, Ram Kishan Chalawadi |
J. Syst. Softw. | 6 |
| 2022 | Automated test-based learning and verification of performance models for microservices systemsabstractEffective and automated verification techniques able to provide assurances of performance and scalability are highly demanded in the context of microservices systems. In this paper, we introduce a methodology that applies specification-driven load testing to learn the behavior of the target microservices system under multiple deployment configurations. Testing is driven by realistic workload conditions sampled in production. The sampling produces a formal description of the users’ behavior through a Discrete Time Markov Chain. This model drives multiple load testing sessions that query the system under test and feed a Bayesian inference process which incrementally refines the initial model to obtain a complete specification from run-time evidence as a Continuous Time Markov Chain. The complete specification is then used to conduct automated verification by using probabilistic model checking and to compute a configuration score that evaluates alternative deployment options. This paper introduces the methodology, its theoretical foundation, and the toolchain we developed to automate it. Our empirical evaluation shows its applicability, benefits, and costs on a representative microservices system benchmark. We show that the methodology detects performance issues, traces them back to system-level requirements, and, thanks to the configuration score, provides engineers with insights on deployment options. The comparison between our approach and a selected state-of-the-art baseline shows that we are able to reduce the cost up to 73% in terms of number of tests. The verification stage requires negligible execution time and memory consumption. We observed that the verification of 360 system-level requirements took ∼1 minute by consuming at most 34 KB. The computation of the score involved the verification of ∼7k (automatically generated) properties verified in ∼72 seconds using at most ∼50 KB. Matteo Camilli, Andrea Janes, Barbara Russo |
J. Syst. Softw. | 3 |
| 2021 | Risk-Driven Compliance Assurance for Collaborative AI Systems: A Vision Paper
Matteo Camilli, Michael Felderer, Andrea Giusti 0004, Dominik T. Matt, Anna Perini, Barbara Russo, Angelo Susi |
REFSQ | 6 |
| 2021 | A Multivariate Characterization and Detection of Software Performance AntipatternsabstractContext. Software Performance Antipatterns (SPAs) research has focused on algorithms for the characterization, detection, and solution of antipatterns. However, existing algorithms are based on the analysis of runtime behavior to detect trends on several monitored variables (e.g., response time, CPU utilization, and number of threads) using pre-defined thresholds. Objective. In this paper, we introduce a new approach for SPA characterization and detection designed to support continuous integration/delivery/deployment (CI/CDD) pipelines, with the goal of addressing the lack of computationally efficient algorithms. Alberto Avritzer, Ricardo Britto 0001, Catia Trubiani, Barbara Russo, Andrea Janes, Matteo Camilli, André van Hoorn, Robert Heinrich, Martina Rapp, Jörg Henß |
ICPE | 4 |
| 2020 | ODRE Workshop: Probabilistic Dynamic Hard Real-Time Scheduling in HPCabstractIndustry 4.0 is changing fundamentally the way data is collected, stored and analyzed in industrial processes. While this change enables novel application such as flexible manufacturing of highly customized products, the real-time control of these processes, however, has not yet realized its full potential. We believe that modern virtualization techniques, specifically application containers, present a unique opportunity to decouple control functionality from associated hardware. Through it, we can fully realize the potential for highly distributed and transferable industrial processes even with real-time constraints arising from time-critical sub-processes. In this paper, we present a specifically developed orchestration tool to manage the challenges and opportunities of shifting industrial control software from dedicated hardware to bare-metal servers or (edge) cloud computing platforms. Using off-the-shelf technology, the proposed tool can manage the execution of containerized applications on shared resources without compromising hard real-time execution determinism. Through first experimental results, we confirm the viability and analyzed the behavior of resource shared systems with strict real-time requirements. We then describe experiments set out to deliver expected results and gather performance, application scope and limits of the presented approach. Florian Hofer 0001, Martin A. Sehr, Barbara Russo, Alberto L. Sangiovanni-Vincentelli |
ISORC | 3 |
| 2020 | Model-Based Testing Under Parametric Variability of Uncertain Beliefs
Matteo Camilli, Barbara Russo |
SEFM | 2 |
| 2020 | Scalability Assessment of Microservice Architecture Deployment Configurations: A Domain-based Approach Leveraging Operational Profiles and Load TestsabstractMicroservices have emerged as an architectural style for developing distributed applications. Assessing the performance of architecture deployment configurations — e.g., with respect to deployment alternatives — is challenging and must be aligned with the system usage in the production environment. In this paper, we introduce an approach for using operational profiles to generate load tests to automatically assess scalability pass/fail criteria of microservice configuration alternatives. The approach provides a Domain-based metric for each alternative that can, for instance, be applied to make informed decisions about the selection of alternatives and to conduct production monitoring regarding performance-related system properties, e.g., anomaly detection. We have evaluated our approach using extensive experiments in a large bare metal host environment and a virtualized environment. First, the data presented in this paper supports the need to carefully evaluate the impact of increasing the level of computing resources on performance. Specifically, for the experiments presented in this paper, we observed that the evaluated Domain-based metric is a non-increasing function of the number of CPU resources for one of the environments under study. In a subsequent series of experiments, we investigate the application of the approach to assess the impact of security attacks on the performance of architecture deployment configurations. Alberto Avritzer, Vincenzo Ferme, Andrea Janes, Barbara Russo, André van Hoorn, Henning Schulz, Daniel Sadoc Menasché, Vilc Queupe Rufino |
J. Syst. Softw. | 4 |
| 2020 | IEC 61131-3 Software Testing: A Portable Solution for Native ApplicationsabstractProgrammable logic controllers (PLCs) are the most used digital systems in manufacturing industry, but there is little support for test automation of such systems. As net result, testing is mostly done manually or not at all despite the recommendations of the IEC 61131-3 Standard. Attempts to provide an automated testing framework for PLCs have been recently performed with first successful results. The most advanced and promising framework proposes an approach close to object orientation that relies on nonnative language and platform. In this article, we propose a testing library called Advanced Program Organization Unit Testing Framework written in native language, built according to the unit testing paradigm, and supporting automated testing for simple and complex scenarios of IEC 61131-3-compliant PLCs. In this article, we present such library, discuss its performance and advantages, and illustrate its application to a real case study. Florian Hofer 0001, Barbara Russo |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Industrial Control via Application Containers: Migrating from Bare-Metal to IAASabstractWe explore the challenges and opportunities of shifting industrial control software from dedicated hardware to bare-metal servers or cloud computing platforms using off the shelf technologies. In particular, we demonstrate that executing time-critical applications on cloud platforms is viable based on a series of dedicated latency tests targeting relevant real-time configurations. Florian Hofer 0001, Martin A. Sehr, Antonio Iannopollo, Ines Ugalde, Alberto L. Sangiovanni-Vincentelli, Barbara Russo |
CloudCom | 6 |
| 2019 | Log mining to re-construct system behavior: An exploratory study on a large telescope system
Michele Pettinato, Juan Pablo Gil, Patricio Galeas, Barbara Russo |
Inf. Softw. Technol. | 4 |
| 2019 | Automatic Identification and Classification of Software Development Video Tutorial FragmentsabstractSoftware development video tutorials have seen a steep increase in popularity in recent years. Their main advantage is that they thoroughly illustrate how certain technologies, programming languages, etc. are to be used. However, they come with a caveat: there is currently little support for searching and browsing their content. This makes it difficult to quickly find the useful parts in a longer video, as the only options are watching the entire video, leading to wasted time, or fast-forwarding through it, leading to missed information. We present an approach to mine video tutorials found on the web and enable developers to query their contents as opposed to just their metadata. The video tutorials are processed and split into coherent fragments, such that only relevant fragments are returned in response to a query. Moreover, fragments are automatically classified according to their purpose, such as introducing theoretical concepts, explaining code implementation steps, or dealing with errors. This allows developers to set filters in their search to target a specific type of video fragment they are interested in. In addition, the video fragments in CodeTube are complemented with information from other sources, such as Stack Overflow discussions, giving more context and useful information for understanding the concepts. Luca Ponzanelli, Gabriele Bavota, Andrea Mocci, Rocco Oliveto, Massimiliano Di Penta, Sonia Haiduc, Barbara Russo, Michele Lanza 0001 |
IEEE Trans. Software Eng. | 7 |
| 2019 | Listening to the Crowd for the Release Planning of Mobile AppsabstractThe market for mobile apps is getting bigger and bigger, and it is expected to be worth over 100 Billion dollars in 2020. To have a chance to succeed in such a competitive environment, developers need to build and maintain high-quality apps, continuously astonishing their users with the coolest new features. Mobile app marketplaces allow users to release reviews. Despite reviews are aimed at recommending apps among users, they also contain precious information for developers, reporting bugs and suggesting new features. To exploit such a source of information, developers are supposed to manually read user reviews, something not doable when hundreds of them are collected per day. To help developers dealing with such a task, we developed CLAP (Crowd Listener for releAse Planning), a web application able to (i) categorize user reviews based on the information they carry out, (ii) cluster together related reviews, and (iii) prioritize the clusters of reviews to be implemented when planning the subsequent app release. We evaluated all the steps behind CLAP, showing its high accuracy in categorizing and clustering reviews and the meaningfulness of the recommended prioritizations. Also, given the availability of CLAP as a working tool, we assessed its applicability in industrial environments. Simone Scalabrino, Gabriele Bavota, Barbara Russo, Massimiliano Di Penta, Rocco Oliveto |
IEEE Trans. Software Eng. | 3 |
| 2018 | A Quantitative Approach for the Assessment of Microservice Architecture Deployment Alternatives by Automated Performance Testing
Alberto Avritzer, Vincenzo Ferme, Andrea Janes, Barbara Russo, Henning Schulz, André van Hoorn |
ECSA | 4 |
| 2018 | code_call_lens: raising the developer awareness of critical codeabstractAs a developer, it is often complex to foresee the impact of changes in source code on usage, e.g., it is time-consuming to find out all components that will be impacted by a change or estimate the impact on the usability of a failing piece of code. It is therefore hard to decide how much effort in quality assurance is justifiable to obtain the desired business goals. In this paper, to reduce the difficulty for developers to understand the importance of source code, we propose an automated way to provide this information to developers as they are working on a given piece of code. As a proof-of-concept, we developed a plug-in for Microsoft Visual Studio Code that informs about the importance of source code methods based on the frequency of usage by the end-users of the developed software. The plug-in aims to increase the awareness developers have about the importance of source code in an unobtrusive way, helping them to prioritize their effort to quality assurance, technical excellence, and usability. code_call_lens can be downloaded from GitHub at https://github.com/xxMUROxx/vscode.code_call_lens. Andrea Janes, Michael Mairegger, Barbara Russo |
ASE | 3 |
| 2018 | Profiling call changes via motif miningabstractComponents' interactions in software systems evolve over time increasing in complexity and size. Developers might have hard time to master such complexity during their maintenance activities incrementing the risk to make mistakes. Understanding changes of such interactions helps developer plan their re-factoring activities. In this study, we propose a method to study the occurrence of motifs in call graphs and their role in the evolution of a system. In our settings, motifs are patterns of class calls that can arise for many reasons as, for example, by implementing design choices. By mining motifs of the call graph obtained from each system's release, we were able to profile the evolution of 68 releases of five open source systems and show that 1) systems have common motifs that occur non-randomly and persistently over their releases, 2) motifs can be used to describe the evolution of calls, compare systems and eventually reveal releases that underwent major changes, 3) there are no specific motif types that include design patterns in all systems under study, but each system has motifs that likely include them, motifs that do not include them at all, and motifs that include a design pattern and occur only once in every release. Some of the findings resemble the ones for biological / physical systems and, as such, path the way to study the evolution of call graphs as dynamical systems (i.e., as system regulated by analytic functions). Barbara Russo |
MSR | 1 |
| 2017 | What if I Had No Smells?abstractWhat would have happened if I did not have any code smell? This is an interesting question that no previous study, to the best of our knowledge, has tried to answer. In this paper, we present a method for implementing a what-if scenario analysis estimating the number of defective files in the absence of smells. Our industrial case study shows that 20% of the total defective files were likely avoidable by avoiding smells. Such estimation needs to be used with the due care though as it is based on a hypothetical history (i.e., zero number of smells and same process and product change characteristics). Specifically, the number of defective files could even increase for some types of smells. In addition, we note that in some circumstances, accepting code with smells might still be a good option for a company. Davide Falessi, Barbara Russo, Kathleen Mullen |
ESEM | 2 |
| 2017 | Mining Logs to Model the Use of a SystemabstractBackground. Process mining is a technique to build process models from "execution logs" (i.e., events triggered by the execution of a process). State-of-the-art tools can provide process managers with different graphical representations of such models. Managers use these models to compare them with an ideal process model or to support process improvement. They typically select the representation based on their experience and knowledge of the system. Aim. This work studies how to automatically build process models representing the actual intents (or uses) of users while interacting with a software system. Such intents are expressed as a set of actions performed by a user to a system to achieve specific use goals. Method. This work applies the theory of Hidden Markov Models to mine use logs and automatically model the use of a system. Results. Unlike the models generated with process mining tools, the Hidden Markov Models automatically generated in this study provide the intents of a user and can be used to recommend managers with a faithful representation of the use of their systems. Conclusions. The automatic generation of the Hidden Markov Models can achieve a good level of accuracy in representing the actual user's intents provided the log dataset is carefully chosen. In our study, the information contained in one-month set of logs helped automatically build Hidden Markov Models with superior accuracy and similar expressiveness of the models built together with the company's stakeholder. Daniele Gadler, Michael Mairegger, Andrea Janes, Barbara Russo |
ESEM | 4 |
| 2017 | Patterns of developers behaviour: A 1000-hour industrial study
Saulius Astromskis, Gabriele Bavota, Andrea Janes, Barbara Russo, Massimiliano Di Penta |
J. Syst. Softw. | 4 |
| 2016 | Too long; didn't watch!: extracting relevant fragments from software development video tutorialsabstractWhen knowledgeable colleagues are not available, developers resort to offline and online resources, e.g., tutorials, mailing lists, and Q&A websites. These, however, need to be found, read, and understood, which takes its toll in terms of time and mental energy. A more immediate and accessible resource are video tutorials found on the web, which in recent years have seen a steep increase in popularity. Nonetheless, videos are an intrinsically noisy data source, and finding the right piece of information might be even more cumbersome than using the previously mentioned resources. Luca Ponzanelli, Gabriele Bavota, Andrea Mocci, Massimiliano Di Penta, Rocco Oliveto, Barbara Russo, Sonia Haiduc, Michele Lanza 0001 |
ICSE | 7 |
| 2016 | Release planning of mobile apps based on user reviewsabstractDevelopers have to to constantly improve their apps by fixing critical bugs and implementing the most desired features in order to gain shares in the continuously increasing and competitive market of mobile apps. A precious source of information to plan such activities is represented by reviews left by users on the app store. However, in order to exploit such information developers need to manually analyze such reviews. This is something not doable if, as frequently happens, the app receives hundreds of reviews per day. In this paper we introduce CLAP (Crowd Listener for releAse Planning), a thorough solution to (i) categorize user reviews based on the information they carry out (e.g., bug reporting), (ii) cluster together related reviews (e.g., all reviews reporting the same bug), and (iii) automatically prioritize the clusters of reviews to be implemented when planning the subsequent app release. We evaluated all the steps behind CLAP, showing its high accuracy in categorizing and clustering reviews and the meaningfulness of the recommended prioritizations. Also, given the availability of CLAP as a working tool, we assessed its practical applicability in industrial environments. Lorenzo Villarroel, Gabriele Bavota, Barbara Russo, Rocco Oliveto, Massimiliano Di Penta |
ICSE | 3 |
| 2016 | A large-scale empirical study on self-admitted technical debtabstractTechnical debt is a metaphor introduced by Cunningham to indicate "not quite right code which we postpone making it right". Examples of technical debt are code smells and bug hazards. Several techniques have been proposed to detect different types of technical debt. Among those, Potdar and Shihab defined heuristics to detect instances of self-admitted technical debt in code comments, and used them to perform an empirical study on five software systems to investigate the phenomenon. Still, very little is known about the diffusion and evolution of technical debt in software projects. Gabriele Bavota, Barbara Russo |
MSR | 2 |
| 2016 | Using Cohesion and Coupling for Software Remodularization: Is It Enough?abstractRefactoring and, in particular, remodularization operations can be performed to repair the design of a software system and remove the erosion caused by software evolution. Various approaches have been proposed to support developers during the remodularization of a software system. Most of these approaches are based on the underlying assumption that developers pursue an optimal balance between cohesion and coupling when modularizing the classes of their systems. Thus, a remodularization recommender proposes a solution that implicitly provides a (near) optimal balance between such quality attributes. However, there is still no empirical evidence that such a balance is thedesideratumby developers. This article aims at analyzing both objectively and subjectively the aforementioned phenomenon. Specifically, we present the results of (1) a large study analyzing the modularization quality, in terms of package cohesion and coupling, of 100 open-source systems, and (2) a survey conducted with 29 developers aimed at understanding the driving factors they consider when performing modularization tasks. The results achieved have been used to distill a set of lessons learned that might be considered to design more effective remodularization recommenders. Ivan Candela, Gabriele Bavota, Barbara Russo, Rocco Oliveto |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2015 | Four eyes are better than two: On the impact of code reviews on software qualityabstractCode review is advocated as one of the best practices to improve software quality and reduce the likelihood of introducing defects during code change activities. Recent research has shown how code components having a high review coverage (i.e., a high proportion of reviewed changes) tend to be less involved in post-release fixing activities. Yet the relationship between code review and bug introduction or the overall software quality is still largely unexplored. Gabriele Bavota, Barbara Russo |
ICSME | 2 |
| 2015 | Query-based configuration of text retrieval solutions for software engineering tasksabstractText Retrieval (TR) approaches have been used to leverage the textual information contained in software artifacts to address a multitude of software engineering (SE) tasks. However, TR approaches need to be configured properly in order to lead to good results. Current approaches for automatic TR configuration in SE configure a single TR approach and then use it for all possible queries. In this paper, we show that such a configuration strategy leads to suboptimal results, and propose QUEST, the first approach bringing TR configuration selection to the query level. QUEST recommends the best TR configuration for a given query, based on a supervised learning approach that determines the TR configuration that performs the best for each query according to its properties. We evaluated QUEST in the context of feature and bug localization, using a data set with more than 1,000 queries. We found that QUEST is able to recommend one of the top three TR configurations for a query with a 69% accuracy, on average. We compared the results obtained with the configurations recommended by QUEST for every query with those obtained using a single TR configuration for all queries in a system and in the entire data set. We found that using QUEST we obtain better results than with any of the considered TR configurations. Laura Moreno, Gabriele Bavota, Sonia Haiduc, Massimiliano Di Penta, Rocco Oliveto, Barbara Russo, Andrian Marcus |
ESEC/SIGSOFT FSE | 6 |
| 2015 | Mining system logs to learn error predictors: a case study of a telemetry system
Barbara Russo, Giancarlo Succi, Witold Pedrycz |
Empir. Softw. Eng. | 1 |
| 2014 | Evolution of design patterns: a replication studyabstractContext. In 2007, Aversano et al. [2] analysed the evolution of JHotDraw, ArgoUML, and Eclipse JDT between years 2000-2005 to understand the role of frequently changed design patterns. Goal. In this paper, we perform a replication of the study on more recent versions to control for artifactual results. In particular, we investigate whether maturity of software versions can affect the original results. Method. We perform a re-analysis of the original data to learn and correctly deploy the tools used for data collection and analysis and to control instrumental threats that typically affect a replication. Results/Conclusions. Findings confirm that patterns change more frequently when they play a crucial role in the software and when in newer releases they support more advanced features. Bruno Rossi 0001, Barbara Russo |
ESEM | 2 |
| 2012 | Characterizing the roles of classes and their fault-proneness through change metricsabstractMany approaches to determine the fault-proneness of code artifacts rely on historical data of and about these artifacts. These data include the code and how it was changed over time, and information about the changes from version control systems. Each of these can be considered at different levels of granularity. The level of granularity can substantially influence the estimated fault-proneness of a code artifact. Typically, the level of detail oscillates between releases and commits on the one hand, and single lines of code and whole files on the other hand. Not every information may be readily available or feasible to collect at every level, though, nor does more detail necessarily improve the results. Our approach is based on time series of changes in method-level dependencies and churn on a commit-to-commit basis for two systems, Spring and Eclipse. We identify sets of classes with distinct properties of the time series of their change histories. We differentiate between classes based on temporal patterns of change. Based on this differentiation, we show that our measure of structural change in concert with its complement, churn, effectively indicates fault-proneness in classes. We also use windows on time series to select sets of commits and show that changes over short amounts of time do effectively indicate the fault-proneness of classes. Maximilian Steff, Barbara Russo |
ESEM | 2 |
| 2012 | Evolution of features and their dependencies - an explorative study in OSSabstractRelease Planning is the process of decision making about what features are to be implemented (or revised) in which release of a software product. While release planning for proprietary software products is well-studied, little investigation has been performed for open source products. Various types of feature dependencies are known to impact both the planning and the subsequent maintenance process. In this paper, we provide the basic layout of a method to formulate and analyze feature dependencies defined at the code level. Dependencies are defined from evolutionary analysis of the commit graph of OSS code development and syntactical dependencies. We demonstrate our method with an explorative study of an open source project, the Spring Framework. From the analysis of the development cycles of two major releases over forty-one months, we could correlate late, increased feature dependencies with an increased number for subsequent improvements and bug fixes. Maximilian Steff, Barbara Russo, Günther Ruhe |
ESEM | 2 |
| 2012 | Co-evolution of logical couplings and commits for defect estimationabstractLogical couplings between files in the commit history of a software repository are instances of files being changed together. The evolution of couplings over commits' history has been used for the localization and prediction of software defects in software reliability. Couplings have been represented in class graphs and change histories on the class-level have been used to identify defective modules. Our new approach inverts this perspective and constructs graphs of ordered commits coupled by common changed classes. These graphs, thus, represent the co-evolution of commits, structured by the change patterns among classes. We believe that co-evolutionary graphs are a promising new instrument for detecting defective software structures. As a first result, we have been able to correlate the history of logical couplings to the history of defects for every commit in the graph and to identify sub-structures of bug-fixing commits over sub-structures of normal commits. Maximilian Steff, Barbara Russo |
MSR | 2 |
| 2011 | Measuring Architectural Change for Defect Estimation and LocalizationabstractWhile there are many software metrics measuring the architecture of a system and its quality, few are able to assess architectural change qualitatively. Given the sheer size and complexity of current software systems, modifying the architecture of a system can have severe, unintended consequences. We present a method to measure architectural change by way of structural distance and show its strong relationship to defect incidence. We show the validity and potential of the approach in an exploratory analysis of the history and evolution of the Spring Framework. Using other, public datasets, we corroborate the results of our analysis. Maximilian Steff, Barbara Russo |
ESEM | 2 |
| 2011 | Path dependent stochastic models to detect planned and actual technology use: A case study of OpenOffice
Bruno Rossi 0001, Barbara Russo, Giancarlo Succi |
Inf. Softw. Technol. | 2 |
| 2011 | A model of job satisfaction for collaborative development processes
Witold Pedrycz, Barbara Russo, Giancarlo Succi |
J. Syst. Softw. | 2 |
| 2007 | Empirical analysis on the correlation between GCC compiler warnings and revision numbers of source files in five industrial software projects
Raimund Moser, Barbara Russo, Giancarlo Succi |
Empir. Softw. Eng. | 2 |
| 2006 | Identification of defect-prone classes in telecommunication software systems using design metrics
Andrea Janes, Marco Scotto, Witold Pedrycz, Barbara Russo, Milorad Stefanovic, Giancarlo Succi |
Inf. Sci. | 4 |
| 2006 | Early estimation of software size in object-oriented environments a case study in a CMM level 3 software firm
Marco Ronchetti, Giancarlo Succi, Witold Pedrycz, Barbara Russo |
Inf. Sci. | 4 |
| 2005 | An Empirical Exploration of the Distributions of the Chidamber and Kemerer Object-Oriented Metrics Suite
Giancarlo Succi, Witold Pedrycz, Snezana Djokic, Paolo Zuliani, Barbara Russo |
Empir. Softw. Eng. | 5 |
| 2003 | An Empirical Analysis on the Discontinuous Use of Pair Programming
Andrea Janes, Barbara Russo, Paolo Zuliani, Giancarlo Succi |
XP | 2 |
| 2003 | An Investigation on the Occurrence of Service Requests in Commercial Software Applications
Giancarlo Succi, Witold Pedrycz, Milorad Stefanovic, Barbara Russo |
Empir. Softw. Eng. | 4 |