Jesús M. González-Barahona

dblp:66/1148 · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0001-9682-460XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 46 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 A CEFR-Inspired Classification Framework with Fuzzy C-Means to Automate Assessment of Programming Skills in Scratch
Ricardo Hidalgo Aragón, Jesús M. González-Barahona, Gregorio Robles
CSEDU (1)2
2026 Immersion vs. familiarity: A controlled experiment of software evolution visualization in virtual reality
abstract
Abstract Understanding how software systems evolve over time is a key challenge in software engineering. While traditional tools like GitHub offer detailed access to project history, their interface can be fragmented and cognitively demanding. Immersive environments, by contrast, offer new opportunities to visualize software evolution in ways that leverage spatial reasoning and embodied interaction. In this paper, we present a controlled experiment comparing an immersive virtual reality (VR) visualization with a conventional GitHub-based interface for analyzing software evolution across releases. Using the BabiaXR framework and the city metaphor, we built parallel environments where 32 participants explored two real-world Java projects and completed a set of analytical tasks related to structure, activity, and system core identification. Our findings show that participants in VR were not more accurate on average, but engaged more deeply with the data, took more time, and often relied on spatial cues to formulate richer insights. In contrast, participants using the on-screen interface completed tasks faster but occasionally reverted to external tools or skipped complex questions. Qualitative analysis revealed that immersive exploration supported deliberate reasoning, especially in open-ended and structural tasks. Notably, prior experience with VR or software visualization had no significant impact on performance, underscoring the accessibility of immersive tools even for novice users. These results suggest that immersive visualizations may not universally outperform traditional interfaces, but they offer unique cognitive benefits for specific types of software comprehension tasks. We discuss the implications for visualization design and outline future directions involving hybrid analytical environments and extended longitudinal studies.
David Moreno-Lumbreras, Sergio Raúl Montes León, Jesús M. González-Barahona, Gregorio Robles
Empir. Softw. Eng.3
2026 Identification and classification of free, open source software licenses: A systematic literature review
Sergio Raúl Montes León, Gregorio Robles, Jesús M. González-Barahona
J. Syst. Softw.3
2025 Enhancing TCP/IP Architecture Learning through Virtual Reality Technology
abstract
Understanding the TCP/IP architecture in intro-ductory Computer Networks courses presents challenges for undergraduate students due to the abstract nature of concepts such as protocols, layers, services, and encapsulation. While 2D and 3D multimedia animations have been employed to support learning, the educational impact of Virtual Reality (VR) animations remains underexplored. To address this gap, we developed WireXRk, a 3D visualization tool that automatically generates interactive animations of network packets captured from real TCP/IP networks. These an-imations are accessible via VR Head-Mounted Displays (HMDs) or standard web browsers on desktop computers. In a study involving 134 freshmen and sophomores across four Computer Networks engineering courses, students were divided into three groups that utilized either traditional slides, 3D animations on a web browser, or 3D animations with VR HMDs. Among students with higher admission grades, learning outcomes were similar across all groups. However, for students with lower admission grades, significant differences emerged: those using traditional slides or 3D animations with VR HMDs performed better than those using 3D animations on a desktop. This suggests that students with less academic experience may benefit from more dynamic learning environments, reinforcing the need for future pedagogical approaches that integrate VR headsets to enhance engagement and retention.
Eva M. Castro, Pedro de las Heras Quirós, Jesús M. González-Barahona, José Centeno-González, Gregorio Robles
EDUCON3
2024 Towards Identifying Code Proficiency Through the Analysis of Python Textbooks
abstract
Python, one of the most prevalent programming languages today, is widely utilized in various domains, including web development, data science, machine learning, and DevOps. Recent scholarly efforts have proposed a methodology to assess Python competence levels, similar to how proficiency in natural languages is evaluated. This method involves assigning levels of competence to Python constructs—for instance, placing simple ‘print’ statements at the most basic level and abstract base classes at the most advanced. The aim is to gauge the level of proficiency a developer must have to understand a piece of source code. This is particularly crucial for software maintenance and evolution tasks, such as debugging or adding new features. For example, in a code review process, this method could determine the competence level required for reviewers. However, categorizing Python constructs by proficiency levels poses significant challenges. Prior attempts, which relied heavily on expert opinions and developer surveys, have led to considerable discrepancies. In response, this paper presents a new approach to identifying Python competency levels through the systematic analysis of introductory Python programming textbooks. By comparing the sequence in which Python constructs are introduced in these textbooks with the current state of the art, we have uncovered notable discrepancies in the order of introduction of Python constructs. Our study underscores a misalignment in the sequences, demonstrating that pinpointing proficiency levels is not trivial. Insights from the study serve as pivotal steps toward reinforcing the idea that textbooks serve as a valuable source for evaluating developers' proficiency, and particularly in terms of their ability to undertake maintenance and evolution tasks.
Ruksit Rojpaisarnkit, Gregorio Robles, Raula Gaikovina Kula, Dong Wang 0044, Chaiyong Ragkhitwetsagul, Jesús M. González-Barahona, Ken-ichi Matsumoto
ICSME6
2024 Testing the past: can we still run tests in past snapshots for Java projects?
abstract
Abstract Building past snapshots of a software project has shown to be of interest both for researchers and practitioners. However, little attention has been devoted specifically to tests available in those past snapshots, which are fundamental for the maintenance of old versions still in production. The aim of this study is to determine to which extent tests of past snapshots can be executed successfully, which would mean these past snapshots are still testable. Given a software project, we build all its past snapshots from source code, including tests, and then run the tests. When tests do not result in success, we also record the reasons, allowing us to determine factors that make tests fail. We run this method on a total of 86 Java projects. On average, for 52.53% of the project snapshots on which tests can be built, all tests pass. However, on average, 94.14% of tests pass in previous snapshots when we take into account the percentage of tests passing in the snapshots used for building those tests. In real software projects, successfully running tests in past snapshots is not something that we can take for granted: we have found that in a large proportion of the projects we studied this does not happen frequently. We have found that the building from source code is the main limitation when running tests on past snapshots. However, we have found some projects for which tests run successfully in a very large fraction of past snapshots, which allows us to identify good practices. We also provide a framework and metrics to quantify testability (the extent to which we are able to run tests of a snapshot with a success result) of past snapshots from several points of view, which simplifies new analyses on this matter, and could help to measure how any project performs in this respect.
Michel Maes-Bermejo, Micael Gallego, Francisco Gortázar, Gregorio Robles, Jesús M. González-Barahona
Empir. Softw. Eng.5
2024 Hunting bugs: Towards an automated approach to identifying which change caused a bug through regression testing
abstract
Abstract Context Finding code changes that introduced bugs is important both for practitioners and researchers, but doing it precisely is a manual, effort-intensive process. The perfect test method is a theoretical construct aimed at detecting Bug-Introducing Changes (BIC) through a theoretical perfect test. This perfect test always fails if the bug is present, and passes otherwise. Objective To explore a possible automatic operationalization of the perfect test method. Method To use regression tests as substitutes for the perfect test. For this, we transplant the regression tests to past snapshots of the code, and use them to identify the BIC, on a well-known collection of bugs from the Defects4J dataset. Results From 809 bugs in the dataset, when running our operationalization of the perfect test method, for 95 of them the BIC was identified precisely and in the remaining 4 cases, a list of candidates including the BIC was provided. Conclusions We demonstrate that the operationalization of the perfect test method through regression tests is feasible and can be completely automated in practice when tests can be transplanted and run in past snapshots of the code. Given that implementing regression tests when a bug is fixed is considered a good practice, when developers follow it, they can detect effortlessly bug-introducing changes by using our operationalization of the perfect test method.
Michel Maes-Bermejo, Alexander Serebrenik, Micael Gallego, Francisco Gortázar, Gregorio Robles, Jesús M. González-Barahona
Empir. Softw. Eng.6
2024 Software development metrics: to VR or not to VR
abstract
Abstract Context Current data visualization interfaces predominantly rely on 2-D screens. However, the emergence of virtual reality (VR) devices capable of immersive data visualization has sparked interest in exploring their suitability for visualizing software development data. Despite this, there is a lack of detailed investigation into the effectiveness of VR devices specifically for interacting with software development data visualizations. Objective Our objective is to investigate the following question: “How do VR devices compare to traditional screens in visualizing data about software development?” Specifically, we aim to assess the accuracy of conclusions derived from exploring visualizations for understanding the software development process, as well as the time required to reach these conclusions. Method In our controlled experiment, we recruited N=32 volunteers with diverse backgrounds. Participants interacted with similar data visualizations in both VR and traditional screen environments. For the traditional screen setup, we utilized a commercially available set of interactive dashboards based on Kibana, commonly used by Bitergia customers for data insights. In the VR environment, we designed a set of visualizations, tailored to provide an equivalent dataset within a virtual room. Participants answered questions related to software evolution processes, specifically code review and issue tracking, in both VR and traditional screen environments, for two projects. We conducted statistical analyses to compare the correctness of their answers and the time taken for each question. Results Our findings indicate that the correctness of answers in both environments is comparable. Regarding time spent, we observed similar durations, except for complex questions that required examining multiple interconnected visualizations. In such cases, participants in the VR environment were able to answer questions more quickly. Conclusion Based on our results, we conclude that VR immersion can be equally effective as traditional screen setups for understanding software development processes through visualization of relevant metrics in most scenarios. Moreover, VR may offer advantages in comprehending complex tasks that require navigating through multiple interconnected visualizations. However, further experimentation is necessary to validate and reinforce these conclusions.
David Moreno-Lumbreras, Gregorio Robles, Daniel Izquierdo 0001, Jesús M. González-Barahona
Empir. Softw. Eng.4
2024 The influence of the city metaphor and its derivates in software visualization
abstract
The city metaphor is widely used in software visualization to represent complex systems as buildings and structures, providing an intuitive way for developers to understand software components. Various software visualization tools have utilized this approach. Identify the influence of the city metaphor on software visualization research, determine its state-of-the-art status, and identify derived tools and their main characteristics. Conduct a systematic mapping study of 406 publications that reference the first paper on the use of the city metaphor in software visualization and/or the main paper of the CodeCity tool. Analyze the 168 publications from which valuable information could be extracted, and build a complete categoric analysis. The field has grown considerably, with an increasing number of publications since 2001, and a changing research community with evolving interconnections between groups. Researchers have developed more tools that support the city metaphor, but less than 50% of the tools were referenced in their papers. Moreover, 85% of the tools did not use extended reality environments, indicating an opportunity for further exploration. The study demonstrates the active and continually growing presence of the city metaphor in research and its impact on software visualization and its derivatives. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board.
David Moreno-Lumbreras, Jesús M. González-Barahona, Gregorio Robles, Valerio Cosentino
J. Syst. Softw.2
2023 Understanding the NPM Dependencies Ecosystem of a Project Using Virtual Reality
abstract
Modern JavaScript development relies heavily on using Node Package Manager (NPM) modules. These modules are related by dependency relationships, possibly requiring dozens or hundreds of modules to build a complete JavaScript web application. Studying dependencies, in terms of their sustainability, vulnerability, size, defects, etc., is fundamental for the deployment and maintenance of JavaScript web applications. We use a 3D metaphor based on presenting dependencies as an “elevated city”, mapping both dependency relationships and characteristics of interest of each module. We developed a VR (virtual reality) scene representing the dependencies of several web applications using the elevated city metaphor, and exposed industrial experts to it to check its suitability. They explored a medium-sized project, with more than 200 dependencies, sharing their insights. The results highlight different aspects of our approach and how the combination of metrics helps experts to obtain insights from the ecosystem. The feedback shows the usefulness of the visualization to check and explore several aspects of the dependencies of an application, helping to identify problems related to maintainability, license usage, or vulnerabilities, and to design strategies to address them.
David Moreno-Lumbreras, Jesús M. González-Barahona, Michele Lanza 0001
VISSOFT2
2023 The software heritage license dataset (2022 edition)
Jesús M. González-Barahona, Sergio Raúl Montes León, Gregorio Robles, Stefano Zacchiroli
Empir. Softw. Eng.1
2023 Revisiting the reproducibility of empirical software engineering studies based on data retrieved from development repositories
abstract
In 2012, our paper “On the reproducibility of empirical software engineering studies based on data retrieved from development repositories” was published. It proposed a method for assessing the reproducibility of studies based on mining software repositories (MSR studies). Since then, several approaches have happened with respect to the study of the reproducibility of this kind of studies. To revisit the proposals of that paper, analyzing to which extent they remain valid, and how they relate to current initiatives and studies on reproducibility and validation of research results in empirical software engineering. We analyze the most relevant studies affecting assumptions or consequences of the approach of the original paper, and other initiatives related to the evaluation of replicability aspects of empirical software engineering studies. We compare the results of that analysis with the results of the original study, finding similarities and differences. We also run a reproducibility assessment study on current MSR papers. Based on the comparison, and the applicability of the method to current papers, we draw conclusions on the validity of the approach of the original paper. The method proposed in the original paper is still valid, and compares well with other more recent methods. It matches the results of relevant studies on reproducibility, and a systematic comparison with them shows that our approach is aligned with their proposals. Our method has practical use, and complements well the current major initiatives on the review of reproducibility artifacts. As a side result, we learn that the reproducibility of MSR studies has improved during the last decade. We propose to use our approach as a fundamental element of a more profound review of the reproducibility of MSR studies, and of the characterization of validation studies in this realm.
Jesús M. González-Barahona, Gregorio Robles
Inf. Softw. Technol.1
2023 CodeCity: A comparison of on-screen and virtual reality
abstract
Over the past decades, researchers proposed numerous approaches to visualize source code. A popular one is CodeCity, an interactive 3D software visualization representing software system as cities: buildings represent classes (or files) and districts represent packages (or folders). Building dimensions represent values of software metrics, such as number of methods or lines of code. There are many implementations of CodeCity, the vast majority of them running on-screen. Recently, some implementations using virtual reality (VR) have appeared, but the usefulness of CodeCity in VR is still to be proven. Our comparative study aims to answer the question “Is VR well suited for CodeCity, compared to the traditional on-screen implementation?” We performed two experiments with our web-based implementation of CodeCity, which can be used on-screen or in immersive VR. First, we conducted a controlled experiment involving 24 participants from academia and industry. Taking advantage of the obtained feedback, we improved our approach and conducted a second controlled experiment with 26 new participants. Our results show that people using the VR version performed the assigned tasks in much less time, while maintaining a comparable level of correctness. VR is at least equally well-suited as on-screen for visualizing CodeCity, and likely better.
David Moreno-Lumbreras, Roberto Minelli, Andrea Villaverde, Jesús M. González-Barahona, Michele Lanza 0001
Inf. Softw. Technol.4
2022 pycefr: Python competency level through code analysis
abstract
Python is known to be a versatile language, well suited both for beginners and advanced users. Some elements of the language are easier to understand than others: some are found in any kind of code, while some others are used only by experienced programmers. The use of these elements lead to different ways to code, depending on the experience with the language and the knowledge of its elements, the general programming competence and programming skills, etc. In this paper, we present pycefr, a tool that detects the use of the different elements of the Python language, effectively measuring the level of Python proficiency required to comprehend and deal with a fragment of Python code. Following the well-known Common European Framework of Reference for Languages (CEFR), widely used for natural languages, pycefr categorizes Python code in six levels, depending on the proficiency required to create and understand it. We also discuss different use cases for pycefr: identifying code snippets that can be understood by developers with a certain proficiency, labeling code examples in online resources such as Stackoverflow and GitHub to suit them to a certain level of competency, helping in the onboarding process of new developers in Open Source Software projects, etc. A video shows availability and usage of the tool: https://tinyurl.com/ypdt3fwe.
Gregorio Robles, Raula Gaikovina Kula, Chaiyong Ragkhitwetsagul, Tattiya Sakulniwat, Ken-ichi Matsumoto, Jesús M. González-Barahona
ICPC6
2022 Starting the InnerSource Journey: Key Goals and Metrics to Measure Collaboration
Daniel Izquierdo 0001, Jesús Alonso-Gutiérrez, Alberto Pérez García-Plaza, Gregorio Robles, Jesús M. González-Barahona
MSR5
2022 Revisiting the building of past snapshots - a replication and reproduction study
abstract
Abstract Context Building past source code snapshots of a software product is necessary both for research (analyzing the past state of a program) and industry (increasing trustability by reproducibility of past versions, finding bugs by bisecting, backporting bug fixes, among others). A study by Tufano et al. showed in 2016 that many past snapshots cannot be built. Objective We replicate Tufano et al.’s study in 2020, to verify its results and to study what has changed during this time in terms of compilability of a project. Also, we extend it by studying a different set of projects, using additional techniques for building past snapshots, with the aim of extending the validity of its results. Method (i) Replication of the original study, obtaining past snapshots from 79 repositories (with a total of 139,389 commits); and (ii) Reproduction of the original study on a different set of 80 large Java projects, extending the heuristics for building snapshots (300,873 commits). Results We observed degradation of compilability over time, due to vanishing of dependencies and other external artifacts. We validated that the most influential error causing failures in builds are missing external artifacts, and the less influential is compiling errors. We observed some facts that could lead to the effect of the build tool on past compilability. Conclusions We provide details on what aspects have a strong and a shallow influence on past compilability, giving ideas of how to improve it. We could extend previous research on the matter, but could not validate some of the previous results. We offer recommendations on how to make this kind of studies more replicable.
Michel Maes-Bermejo, Micael Gallego, Francisco Gortázar, Gregorio Robles, Jesús M. González-Barahona
Empir. Softw. Eng.5
2022 Development effort estimation in free/open source software from activity in version control systems
abstract
Abstract Effort estimation models are a fundamental tool in software management, and used as a forecast for resources, constraints and costs associated to software development. For Free/Open Source Software (FOSS) projects, effort estimation is especially complex: professional developers work alongside occasional, volunteer developers, so the overall effort (in person-months) becomes non-trivial to determine. The objective of this work it to develop a simple effort estimation model for FOSS projects, based on the historic data of developers’ effort. The model is fed with direct developer feedback to ensure its accuracy. After extracting the personal development profiles of several thousands of developers from 6 large FOSS projects, we asked them to fill in a questionnaire to determine if they should be considered as full-time developers in the project that they work in. Their feedback was used to fine-tune the value of an effort threshold, above which developers can be considered as full-time. With the help of the over 1,000 questionnaires received, we were able to determine, for every project in our sample, the threshold of commits that separates full-time from non-full-time developers. We finally offer guidelines and a tool to apply our model to FOSS projects that use a version control system.
Gregorio Robles, Andrea Capiluppi, Jesús M. González-Barahona, Björn Lundell, Jonas Gamalielsson
Empir. Softw. Eng.3
2021 CodeCity: On-Screen or in Virtual Reality?
abstract
Over the past decades, researchers proposed numerous approaches to visualize source code. A prominent one is CodeCity, an interactive 3D software visualization that leverages the "city metaphor" to represent software system as cities: buildings represent classes (or files) and districts represent packages (or folders). Building dimensions represent values of software metrics, such as the number of methods or the lines of code. There are many implementations of CodeCity, the vast majority of them running on-screen. Recently, some implementations visualizing CodeCity in virtual reality (VR) have appeared. While exciting as a technology, VR’s usefulness remains to be proven.The question we pose is: Is VR well suited to visualize CodeCity, compared to the traditional on-screen implementation?We performed an experiment in our interactive web-based application to visualize CodeCity. Users can fetch data from any git repository and visualize its source code. Our application enables users to navigate CodeCity both on-screen and in an immersive VR environment, using consumer-grade VR headsets like Oculus Quest. Our controlled experiment involved 24 participants from academia and industry. Results show that people using the VR version performed the assigned tasks in much less time, while still maintaining a comparable level of correctness.Therefore, our results show that VR is at least equally well-suited as on-screen for visualizing CodeCity, and likely better.
David Moreno-Lumbreras, Roberto Minelli, Andrea Villaverde, Jesús M. González-Barahona, Michele Lanza 0001
VISSOFT4
2021 A multi-dimensional analysis of technical lag in Debian-based Docker images
Ahmed Zerouali, Tom Mens, Alexandre Decan, Jesús M. González-Barahona, Gregorio Robles
Empir. Softw. Eng.4
2020 How bugs are born: a model to identify how bugs are introduced in software components
abstract
Abstract When identifying the origin of software bugs, many studies assume that “a bug was introduced by the lines of code that were modified to fix it”. However, this assumption does not always hold and at least in some cases, these modified lines are not responsible for introducing the bug. For example, when the bug was caused by a change in an external API. The lack of empirical evidence makes it impossible to assess how important these cases are and therefore, to which extent the assumption is valid. To advance in this direction, and better understand how bugs “are born”, we propose a model for defining criteria to identify the first snapshot of an evolving software system that exhibits a bug. This model, based on the perfect test idea, decides whether a bug is observed after a change to the software. Furthermore, we studied the model’s criteria by carefully analyzing how 116 bugs were introduced in two different open source software projects. The manual analysis helped classify the root cause of those bugs and created manually curated datasets with bug-introducing changes and with bugs that were not introduced by any change in the source code. Finally, we used these datasets to evaluate the performance of four existing SZZ-based algorithms for detecting bug-introducing changes. We found that SZZ-based algorithms are not very accurate, especially when multiple commits are found; the F-Score varies from 0.44 to 0.77, while the percentage of true positives does not exceed 63%. Our results show empirical evidence that the prevalent assumption, “a bug was introduced by the lines of code that were modified to fix it”, is just one case of how bugs are introduced in a software system. Finding what introduced a bug is not trivial: bugs can be introduced by the developers and be in the code, or be created irrespective of the code. Thus, further research towards a better understanding of the origin of bugs in software projects could help to improve design integration tests and to design other procedures to make software development more robust.
Gema Rodríguez-Pérez, Gregorio Robles, Alexander Serebrenik, Andy Zaidman, Daniel M. Germán, Jesús M. González-Barahona
Empir. Softw. Eng.6
2020 Open Source Systems: Enterprise Software and Solutions
Ioannis Stamelos, Iraklis Varlamis, Dimosthenis Anagnostopoulos, Jesús M. González-Barahona
J. Syst. Softw.4
2019 ConPan: a tool to analyze packages in software containers
abstract
Deploying software packages and services into containers is a popular software engineering practice that increases portability and reusability. Docker, the most popular containerization technology, helps DevOps practitioners in their daily activities. Despite being successfully and increasingly employed, containers may include buggy and vulnerable packages that put at risk the environments in which the containers have been deployed. Existing quality and security monitoring tools provide only limited support to analyze Docker containers, thus forcing practitioners to perform additional manual work or develop adhoc scripts when the analysis goes beyond security purposes. This limitation also affects researchers desiring to empirically study the evolution dynamics of Docker containers and their contained packages. To overcome this limitation, we present ConPan, an automated tool to inspect the characteristics of packages in Docker containers, such as their outdatedness and other possible flaws (e.g., bugs and security vulnerabilities). ConPan comes with a CLI and API, and the analysis results can be presented to the user in a variety of formats.
Ahmed Zerouali, Valerio Cosentino, Gregorio Robles, Jesús M. González-Barahona, Tom Mens
MSR4
2019 On the Impact of Outdated and Vulnerable Javascript Packages in Docker Images
abstract
Containerized applications, and in particular Docker images, are becoming a common solution in cloud environments to meet ever-increasing demands in terms of portability, reliability and fast deployment. A Docker image includes all environmental dependencies required to run it, such as specific versions of system and third-party packages. Leveraging on its modularity, an image can be easily embedded in other images, thus simplifying the way of sharing dependencies and building new software. However, the dependencies included in an image may be out of date due to backward compatibility requirements, endangering the environments where the image has been deployed with known vulnerabilities. While previous research efforts have focused on studying the impact of bugs and vulnerabilities of system packages within Docker images, no attention has been given to third-party packages. This paper empirically studies the impact of npm JavaScript package vulnerabilities in Docker images. We based our analysis on 961 images from three official repositories that use Node.js, and 1,099 security reports of packages available on npm, the most popular JavaScript package manager. Our results reveal that the presence of outdated npm packages in Docker images increases the risk of potential security vulnerabilities, suggesting that Docker maintainers should keep their installed JavaScript packages up to date.
Ahmed Zerouali, Valerio Cosentino, Tom Mens, Gregorio Robles, Jesús M. González-Barahona
SANER5
2019 On the Relation between Outdated Docker Containers, Severity Vulnerabilities, and Bugs
abstract
Packaging software into containers is becoming a common practice when deploying services in cloud and other environments. Docker images are one of the most popular container technologies for building and deploying containers. A container image usually includes a collection of software packages, that can have bugs and security vulnerabilities that affect the container health. Our goal is to support container deployers by analysing the relation between outdated containers and vulnerable and buggy packages installed in them. We use the concept of technical lag of a container as the difference between a given container and the most up-to-date container that is possible with the most recent releases of the same collection of packages. For 7,380 official and community Docker images that are based on the Debian Linux distribution, we identify which software packages are installed in them and measure their technical lag in terms of version updates, security vulnerabilities and bugs. We have found, among others, that no release is devoid of vulnerabilities, so deployers cannot avoid vulnerabilities even if they deploy the most recent packages. We offer some lessons learned for container developers in regard to the strategies they can follow to minimize the number of vulnerabilities. We argue that Docker container scan and security management tools should improve their platforms by adding data about other kinds of bugs and include the measurement of technical lag to offer deployers information of when to update.
Ahmed Zerouali, Tom Mens, Gregorio Robles, Jesús M. González-Barahona
SANER4
2019 On the Diversity of Software Package Popularity Metrics: An Empirical Study of npm
abstract
Software systems often leverage on open source software libraries to reuse functionalities. Such libraries are readily available through software package managers like npm for JavaScript. Due to the huge amount of packages available in such package distributions, developers often decide to rely on or contribute to a software package based on its popularity. Moreover, it is a common practice for researchers to depend on popularity metrics for data sampling and choosing the right candidates for their studies. However, the meaning of popularity is relative and can be defined and measured in a diversity of ways, that might produce different outcomes even when considered for the same studies. In this paper, we show evidence of how different is the meaning of popularity in software engineering research. Moreover, we empirically analyse the relationship between different software popularity measures. As a case study, for a large dataset of 175k npm packages, we computed and extracted 9 different popularity metrics from three open source tracking systems: libraries.io, npmjs.com and GitHub. We found that indeed popularity can be measured with different unrelated metrics, each metric can be defined within a specific context. This indicates a need for a generic framework that would use a portfolio of popularity metrics drawing from different concepts.
Ahmed Zerouali, Tom Mens, Gregorio Robles, Jesús M. González-Barahona
SANER4
2019 A formal framework for measuring technical lag in component repositories - and its application to npm
abstract
Abstract Reusable Open Source Software (OSS) components for major programming languages are available in package repositories. Developers rely on package management tools to automate deployments, specifying which package releases satisfy the needs of their applications. However, these specifications may lead to deploying package releases that are outdated, or otherwise undesirable, because they do not include bug fixes, security fixes, or new functionality. In contrast, automatically updating to a more recent release may introduce incompatibility issues. To capture this delicate balance, we formalise a generic model of technical lag, a concept that quantifies to which extent a deployed collection of components is outdated, with respect to the ideal deployment. We operationalise this model for the npm package manager. We empirically analyze the history of package update practices and technical lag for more than 500K packages with about 4M package releases over a seven‐year period. We consider both development and runtime dependencies, and study both direct and transitive dependencies. We also analyze the technical lag of external GitHub applications depending on npm packages. We report our findings, suggesting the need for more awareness of, and integrated tool support for, controlling technical lag in software libraries.
Ahmed Zerouali, Tom Mens, Jesús M. González-Barahona, Alexandre Decan, Eleni Constantinou, Gregorio Robles
J. Softw. Evol. Process.3
2018 What if a bug has a different origin?: making sense of bugs without an explicit bug introducing change
abstract
Background: Many studies in the software research literature on bug fixing are built upon the assumption that "a given bug was introduced by the lines of code that were modified to fix it", or variations of it. Although this assumption seems very reasonable at first glance, there is little empirical evidence supporting it. A careful examination surfaces that there are other possible sources for the introduction of bugs such as modifications to those lines that happened before the last change an changes external to the piece of code being fixed. Goal: We aim at understanding the complex phenomenon of bug introduction and bug fix. Method: We design a preliminary approach distinguishing between bug introducing commits (BIC) and first failing moments (FFM). We apply this approach to Nova and ElasticSearch, two large and well-known open source software projects. Results: In our initial results we obtain that at least 24% bug fixes in Nova and 10% in ElasticSearch have not been caused by a BIC but by co-evolution, compatibility issues or bugs in external API. Merely 26--29% of BICs can be found using the algorithm based on the assumption that "a given bug was introduced by the lines of code that were modified to fix it". Conclusions: The approach allows also for a better framing of the comparison of automatic methods to find bug inducting changes. Our results indicate that more attention should be paid to whether a bug has been introduced and, when it was introduced.
Gema Rodríguez-Pérez, Andy Zaidman, Alexander Serebrenik, Gregorio Robles, Jesús M. González-Barahona
ESEM5
2018 An Empirical Analysis of Technical Lag in npm Package Dependencies
Ahmed Zerouali, Eleni Constantinou, Tom Mens, Gregorio Robles, Jesús M. González-Barahona
ICSR5
2018 [Engineering Paper] Graal: The Quest for Source Code Knowledge
abstract
Source code analysis tools are designed to analyze code artifacts with different intents, which span from improving the quality and security of the software to easing refactoring and reverse engineering activities. However, most tools do not come with features to periodically schedule their analysis or to be executed on a battery of repositories, and lack support to combine their results with other analysis tools. Thus, researchers and practitioners are often forced to develop ad-hoc scripts to meet their needs. This comes at the risk of obtaining wrong results (because of the lack of testing) and of hindering replication by other research teams. In addition, the resulting scripts are often not meant to be customized nor designed for incrementality, scalability and extensibility. In this paper we present Graal, which empowers users with a customizable, scalable and incremental approach to conduct source code analysis and enables relating the obtained results with other software project data. Graal leverages on and extends the functionalities of GrimoireLab, a strong free software tool developed by Bitergia, a company devoted to offer commercial software development analytics, and part of the CHAOSS project of the Linux Foundation.
Valerio Cosentino, Santiago Dueñas, Ahmed Zerouali, Gregorio Robles, Jesús M. González-Barahona
SCAM5
2018 Reproducibility and credibility in empirical software engineering: A case study based on a systematic literature review of the use of the SZZ algorithm
Gema Rodríguez-Pérez, Gregorio Robles, Jesús M. González-Barahona
Inf. Softw. Technol.3
2017 Using Metrics to Track Code Review Performance
abstract
During 2015, some members of the Xen Project Advisory Board became worried about the performance of their code review process. The Xen Project is a free, open source software project developing one of the most popular virtualization platforms in the industry. They use a pre-commit peer review process similar to that in the Linux kernel, based on email messages. They had observed a large increase over time in the number of messages related to code review, and were worried about how this could be a signal of problems with their code review process.
Daniel Izquierdo 0001, Nelson Sekitoleko, Jesús M. González-Barahona, Lars Kurth
EASE3
2016 Characterization of the Xen project code review process: an experience report
abstract
Many software development projects have introduced mandatory code review for every change to the code. This means that the project needs to devote a significant effort to review all proposed changes, and that their merging into the code base may get considerably delayed. Therefore, all those projects need to understand how code review is working, and the delays it is causing in time to merge.
Daniel Izquierdo 0001, Lars Kurth, Jesús M. González-Barahona, Santiago Dueñas, Nelson Sekitoleko
MSR3
2016 Software Engineering Artifact in Software Development Process - Linkage Between Issues and Code Review Processes
abstract
Researchers working with software repositories, often when building performance or quality models, need to recover traceability links between bug reports in issue tracking repositories and reviews in code review systems. However, very often the information stored in bug tracking repositories is not explicitly tagged or linked to the issues reviewing them. Researchers have to adopt various heuristics to tag the data. These includes, for example, identifying if an issue is a bug report or not. In this study we promote a research artifact in software engineering, a reusable unit of research that can be used to support other research endeavors and has acted as a support material that enabled the creation of the results published in a great number of papers until now – linking issues and reviews of the code review software process. We present two state-of-the-practice algorithms on how to link issues and reviews, selecting as our case study the open source cloud computing project of OpenStack. OpenStack enforces strict development guidelines and rules on the quality of the data in its issue tracking repository. We empirically compare the outcome of the two approaches, highlighting the most prominent one.
Dorealda Dalipaj, Jesús M. González-Barahona, Daniel Izquierdo 0001
SoMeT2
2016 Determining the Geographical distribution of a Community by means of a Time-zone Analysis
abstract
Free/libre/open source software projects are usually developed by a geographically distributed community of developers and contributors. In contrast to traditional corporate environments, it is hard to obtain information about how the community is geographically distributed, mainly because participation is open to volunteers and in many cases it is just occasional. During the last years, specially with the increasing implication of institutions, non-profit organizations and companies, there is a growing interest in having information about the geographic location of developers. This is because projects want to be as global as possible, in order to attract new contributors, users and, of course, clients. In this paper we show a methodology to obtain the geographical distribution of a development community by analyzing the source code management system and the mailing lists they use.
Jesús M. González-Barahona, Gregorio Robles, Daniel Izquierdo 0001
OpenSym1
2015 The MetricsGrimoire Database Collection
abstract
The Metrics Grimoire system is composed by a set of tools designed to retrieve data from repositories related to software development. Their aim is to produce organized databases suitable for easy querying with research and industrial purposes. The data in those databases have a similar structure, to easy cross-database studies, and can be enriched with information such as linkage of the multiple identities of actors, or their affiliation. This paper presents the general structure of those databases, and a collection of up-to-date database dumps that are publicly available. They correspond to two well-known projects, Open Stack, and Eclipse, including data from source code management repositories, issue tracking systems, mailing lists, and code review systems.
Jesús M. González-Barahona, Gregorio Robles, Daniel Izquierdo 0001
MSR1
2014 Free/Open Source Software projects as early MOOCs
abstract
This paper presents Free/Libre/Open Source Software (FLOSS) Projects as early Massive Online Open Courses (MOOCs). Being software development a process where learning and collaboration is of major importance, FLOSS projects have in common many characteristics with MOOCs. This is because many FLOSS projects (such as Linux, Apache, GNOME or KDE, among others) are massive, they are open to anyone to participate, and are driven mainly by telematic means. We therefore present the research literature that has studied FLOSS projects from points of view that are close to learning and discuss how the FLOSS community has approached many of the issues related to acquiring knowledge and skills over the Internet and compare them to how currently MOOCs, both xMOOCs and cMOOCs, address these situations.
Gregorio Robles, Hugo Plaza, Jesús M. González-Barahona
EDUCON3
2014 Estimating development effort in Free/Open source software projects by mining software repositories: a case study of OpenStack
abstract
Because of the distributed and collaborative nature of free / open source software (FOSS) projects, the development effort invested in a project is usually unknown, even after the software has been released. However, this information is becoming of major interest, especially ---but not only--- because of the growth in the number of companies for which FOSS has become relevant for their business strategy. In this paper we present a novel approach to estimate effort by considering data from source code management repositories. We apply our model to the OpenStack project, a FOSS project with more than 1,000 authors, in which several tens of companies cooperate. Based on data from its repositories and together with the input from a survey answered by more than 100 developers, we show that the model offers a simple, but sound way of obtaining software development estimations with bounded margins of error.
Gregorio Robles, Jesús M. González-Barahona, Carlos Cervigón, Andrea Capiluppi, Daniel Izquierdo 0001
MSR2
2014 FLOSS 2013: a survey dataset about free software contributors: challenges for curating, sharing, and combining
abstract
In this data paper we describe a data set obtained by means of performing an on-line survey to over 2,000 Free Libre Open Source Software (FLOSS) contributors. The survey includes questions related to personal characteristics (gender, age, civil status, nationality, etc.), education and level of English, professional status, dedication to FLOSS projects, reasons and motivations, involvement and goals. We describe as well the possibilities and challenges of using private information from the survey when linked with other, publicly available data sources. In this regard, an example of data sharing will be presented and legal, ethical and technical issues will be discussed.
Gregorio Robles, Laura Arjona Reina, Alexander Serebrenik, Bogdan Vasilescu, Jesús M. González-Barahona
MSR5
2014 Studying the laws of software evolution in a long-lived FLOSS project
abstract
Some free, open-source software projects have been around for quite a long time, the longest living ones dating from the early 1980s. For some of them, detailed information about their evolution is available in source code management systems tracking all their code changes for periods of more than 15 years. This paper examines in detail the evolution of one of such projects, glibc, with the main aim of understanding how it evolved and how it matched Lehman's laws of software evolution. As a result, we have developed a methodology for studying the evolution of such long-lived projects based on the information in their source code management repository, described in detail several aspects of the history of glibc, including some activity and size metrics, and found how some of the laws of software evolution may not hold in this case. © 2013 The Authors. Journal of Software: Evolution and Process published by John Wiley & Sons Ltd.
Jesús M. González-Barahona, Gregorio Robles, Israel Herraiz, Felipe Ortega
J. Softw. Evol. Process.1
2013 Mining student repositories to gain learning analytics. An experience report
abstract
Engineering students often have to deliver small computer programs in many engineering courses. Instructors have to evaluate these assignments according to the learning goals and their quality, but ensure as well that there is no plagiarism. In this paper, we report the experience of using mining software repositories techniques in a multimedia networks course where students have to submit several software programs. We show how we have proceeded, the tools that we have used and provide some useful links and ideas that other lecturers may use.
Gregorio Robles, Jesús M. González-Barahona
EDUCON2
2013 Intensive metrics for the study of the evolution of open source projects: case studies from apache software foundation projects
abstract
Based on the empirical evidence that the ratio of email messages in public mailing lists to versioning system commits has remained relatively constant along the history of the Apache Software Foundation (ASF), this paper has as goal to study what can be inferred from such a metric for projects of the ASF. We have found that the metric seems to be an intensive metric as it is independent of the size of the project, its activity, or the number of developers, and remains relatively independent of the technology or functional area of the project. Our analysis provides evidence that the metric is related to the technical effervescence and popularity of project, and as such can be a good candidate to measure its healthy evolution. Other, similar metrics -like the ratio of developer messages to commits and the ratio of issue tracker messages to commits- are studied for several projects as well, in order to see if they have similar characteristics.
Santiago Gala-Pérez, Gregorio Robles, Jesús M. González-Barahona, Israel Herraiz
MSR3
2012 Hybrid educational worlds
abstract
The GSyC/LibreSoft research group at the Universidad Rey Juan Carlos has been investigating serious games with smartphones. In our approach, platform developers and educators do not need to create complete virtual worlds, which are in general very time and effort consuming. In the games that have been developed, participants are provided through a mobile phone with the necessary information to imagine a virtual world and interact with other participants and objects in a manner consistent with the learning objectives. In this paper this type of game is presented and the implications of using such hybrid worlds in the creation of serious games for smartphones will be discussed.
Gregorio Robles, Jorge Fernández-González, Jesús M. González-Barahona, Julio Ramiro
EDUCON3
2012 A synchronous on-line competition software to improve and motivate learning
abstract
In this paper we describe a competitive way to learn and review concepts: a championship based on quizzes. Students will compete in “duels” against other students in an instant messaging channel where a program will be launching questions; the student that answers first will get the point. And the first to reach a series of points will win the game. We have designed, implemented and tested a (free) software to automate the entire process of the championship, allowing students to choose the times of their matches, throwing questions to the channel, keeping logs, driving classifications, providing feedback to the lecturer about performance and questions, etc. From several testing championships performed, we have noticed a high involvement and motivation by students. Being the process automated, lecturers do not have to spend much time on this activity even if it takes place outside the classroom (they only have to worry about uploading questions with their answers and keep track of possible revisions students may ask for).
Gregorio Robles, Jesús M. González-Barahona, Arturo Moral
EDUCON2
2012 On the reproducibility of empirical software engineering studies based on data retrieved from development repositories
abstract
Among empirical software engineering studies, those based on data retrieved from development repositories (such as those of source code management, issue tracking or communication systems) are specially suitable for reproduction. However their reproducibility status can vary a lot, from easy to almost impossible to reproduce. This paper explores which elements can be considered to characterize the reproducibility of a study in this area, and how they can be analyzed to better understand the type of reproduction studies they enable or obstruct. One of the main results of this exploration is the need of a systematic approach to asses the reproducibility of a study, due to the complexity of the processes usually involved, and the many details to be taken into account. To address this need, a methodology for assessing the reproducibility of studies is also presented and discussed, as a tool to help to raise awareness about research reproducibility in this field. The application of the methodology in practice has shown how, even for papers aimed to be reproducible, a systematic analysis raises important aspects that render reproduction difficult or impossible. We also show how, by identifying elements and attributes related to reproducibility, it can be better understood which kind of reproduction can be done for a specific study, given the description of datasets, methodologies and parameters it uses.
Jesús M. González-Barahona, Gregorio Robles
Empir. Softw. Eng.1
2011 A Quantitative Examination of the Impact of Featured Articles in Wikipedia
Antonio J. Reinoso, Jesús M. González-Barahona, Rocío Muñoz-Mansilla, Israel Herraiz
ICSOFT (1)2
2010 A Statistical Approach to the Impact of Featured Articles in Wikipedia
Antonio J. Reinoso, Felipe Ortega, Jesús M. González-Barahona, Israel Herraiz
KEOD3
2009 A quantitative approach to the use of the Wikipedia
abstract
This paper presents a quantitative study of the use of the Wikipedia system by its users (both readers and editors), with special focus on the identification of time and kind-of-use patterns, characterization of traffic and workload, and comparative analysis of different language editions. The basis of the study is the filtering and analysis of a large sample of the requests directed to the Wikimedia systems for six weeks, each in a month from November 2007 to April 2008. In particular, we have considered the twenty most frequently visited language editions of the Wikipedia, identifying for each access to any of them the corresponding namespace (sets of resources with uniform semantics), resource name (article names, for example) and action (editions, submissions, history reviews, save operations, etc.). The results found include the identification of weekly and daily patterns, and several correlations between several actions on the articles. In summary, the study shows an overall picture of how the most visited language editions of the Wikipedia are being accessed by their users.
Antonio J. Reinoso, Jesús M. González-Barahona, Gregorio Robles, Felipe Ortega
ISCC2
2009 Evolution of the core team of developers in libre software projects
abstract
In many libre (free, open source) software projects, most of the development is performed by a relatively small number of persons, the ldquocore teamrdquo. The stability and permanence of this group of most active developers is of great importance for the evolution and sustainability of the project. In this position paper we propose a quantitative methodology to study the evolution of core teams by analyzing information from source code management repositories. The most active developers in different periods are identified, and their activity is calculated over time, looking for core team evolution patterns.
Gregorio Robles, Jesús M. González-Barahona, Israel Herraiz
MSR2
2009 Macro-level software evolution: a case study of a large software compilation
abstract
Software evolution studies have traditionally focused on individual products. In this study we scale up the idea of software evolution by considering software compilations composed of a large quantity of independently developed products, engineered to work together. With the success of libre (free, open source) software, these compilations have become common in the form of ‘software distributions’, which group hundreds or thousands of software applications and libraries into an integrated system. We have performed an exploratory case study on one of them, Debian GNU/Linux, finding some significant results. First, Debian has been doubling in size every 2 years, totalling about 300 million lines of code as of 2007. Second, the mean size of packages has remained stable over time. Third, the number of dependencies between packages has been growing quickly. Finally, while C is still by far the most commonly used programming language for applications, use of the C++, Java, and Python languages have all significantly increased. The study helps not only to understand the evolution of Debian, but also yields insights into the evolution of mature libre software systems in general.
Jesús M. González-Barahona, Gregorio Robles, Martin Michlmayr, Juan José Amor, Daniel M. Germán
Empir. Softw. Eng.1
2008 Managing Libre Software Distributions under a Product Line Approach
abstract
Software product lines have already proven to be a successful methodology for building and maintaining a collection of similar software products, based on a common architecture. However, when the base system is heterogeneous and extremely large in size, an extra level of complexity is introduced that should be addressed with appropriate methods and techniques. A good example of this kind of systems is the product family composed by the software distributions composed by libre (free, opensource) software, and based on Linux or BSD kernels. All of them can be considered as a part of a product line, based on a large collection of thousands of packages. One of the main problems faced by these distributions is the increasingly growing number of dependencies among packages, which is already caused problems, with a high risk of rendering the management of such large distributions impossible. In this paper we address some of challenges and main problems of Linux distributions when adopting a product line approach, with special focus to the maintenance and evolution of such systems.
Israel Herraiz, Gregorio Robles, Rafael Capilla, Jesús M. González-Barahona
COMPSAC4
2008 Towards a simplification of the bug report form in eclipse
abstract
We believe that the bug report form of Eclipse contains too many fields, and that for some fields, there are too many options. In this MSR challenge report, we focus in the case of the severity field. That field contains seven different levels of severity. Some of them seem very similar, and it is hard to distinguish among them. Users assign severity, and developers give priority to the reports depending on their severity. However, if users can not distinguish well among the various severity options, they will probably assign different priorities to bugs that require the same priority. We study the mean time to close bugs reported in Eclipse, and how the severity assigned by users affects this time. The results shows that classifying by time to close, there are less clusters of bugs than levels of severity. We therefore conclude that there is a need to make a simpler bug report form.
Israel Herraiz, Daniel M. Germán, Jesús M. González-Barahona, Gregorio Robles
MSR3
2008 Determinism and evolution
abstract
It has been proposed that software evolution follows a Self-Organized Criticality (SOC) dynamics. This fact is supported by the presence of long range correlations in the time series of the number of changes made to the source code over time. Those long range correlations imply that the current state of the project was determined time ago. In other words, the evolution of the software project is governed by a sort of determinism. But this idea seems to contradict intuition. To explore this apparent contradiction, we have performed an empirical study on a sample of 3, 821 libre (free, open source) software projects, finding that their evolution projects is short range correlated. This suggests that the dynamics of software evolution may not be SOC, and therefore that the past of a project does not determine its future except for relatively short periods of time, at least for libre software.
Israel Herraiz, Jesús M. González-Barahona, Gregorio Robles
MSR2
2007 On the prediction of the evolution of libre software projects
abstract
Libre (free / open source) software development is a complex phenomenon. Many actors (core developers, casual contributors, bug reporters, patch submitters, users, etc.), in many cases volunteers, interact in complex patterns without the constrains of formal hierarchical structures or organizational ties. Understanding this complex behavior with enough detail to build explanatory models suitable for prediction is an open challenge, and few results have been published to date in this area. Therefore statistical, non-explanatory models (such as the traditional regression model) have a clear role, and have been used in some evolution studies. Our proposal goes in this direction, but using a model that we have found more useful: time series analysis. Data available from the source code management repository is used to compute the size of the software over its past life, using this information to estimate the future evolution of the project. In this paper we present this methodology and apply it to three large projects, showing how in these cases predictions are more accurate than regression models, and precise enough to estimate with little error their near future evolutions.
Israel Herraiz, Jesús M. González-Barahona, Gregorio Robles, Daniel M. Germán
ICSM2
2006 Towards Community-Driven Development of Educational Materials: The Edukalibre Approach
Jesús M. González-Barahona, Vania Dimitrova, Diego Chaparro, Chris Tebb, Teofilo Romera, Luis Canas, Julika Siemer-Matravers, Styliani Kleanthous
EC-TEL1
2006 Beyond source code: The importance of other artifacts in software development (a case study)
Gregorio Robles, Jesús M. González-Barahona, Juan Julián Merelo Guervós
J. Syst. Softw.2
2000 Libre software environment for robot programming
abstract
When facing the problem of teaching the basis of robot control programming to computer science students, apart from the syllabus of the course, some other requirements have to be considered, such as which is the most appropriate robot and which are the right tools for learning how to control it. In this paper, we describe the tools we have chosen for teaching robotics, focusing on an environment that supports practical assignments. We also analyze the reasons that made us choose each tool, giving special emphasis to the Libre software requirement that we have imposed on every tool we are using. Finally, we present the results and opinions we have obtained from our students and the lessons we have learned by using this Libre software approach.
Vicente Matellán Olivera, Jesús M. González-Barahona, José Centeno-González, Pedro de las Heras Quirós
SMC2