VLDB 2026 Research / reviewers in the wild / expert
Stefano Zacchiroli
dblp:53/2641
· DBLP profile ↗
53ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-4576-136XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 44 · 2 first-author · 17 since 2021Databases, data management, data science and information retrieval · 17 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Theory of computation · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Promises, Perils, and (Timely) Heuristics for Mining Coding Agent ActivityabstractIn 2025, coding agents have seen a very rapid adoption. Coding agents leverage Large Language Models (LLMs) in ways that are markedly different from LLM-based code completion, making their study critical. Moreover, unlike LLM-based completion, coding agents leave visible traces in software repositories, enabling the use of MSR techniques to study their impact on SE practices. This paper documents the promises, perils, and heuristics that we have gathered from studying coding agent activity on GitHub. Romain Robbes, Théo Matricon, Thomas Degueule, Andre Hora 0001, Stefano Zacchiroli |
MSR | 5 |
| 2026 | Determining the intrinsic structure of public software development history: an exploratory study
Antoine Pietri, Guillaume Rousseau, Stefano Zacchiroli |
Empir. Softw. Eng. | 3 |
| 2025 | Automatic Classification of Software Repositories: a Systematic Mapping StudyabstractThe rapid growth of software repositories on development platforms such as GitHub, as well as archives like Software Heritage, prompts the need for better repository classification. Machine learning is increasingly used to automate this classification, but there are no secondary studies analyzing this research landscape. Stefano Balla, Thomas Degueule, Romain Robbes, Jean-Rémy Falleri, Stefano Zacchiroli |
EASE | 5 |
| 2025 | ROSA: Finding Backdoors with FuzzingabstractA code-level backdoor is a hidden access, programmed and concealed within the code of a program. For instance, hard-coded credentials planted in the code of a file server application would enable maliciously logging into all deployed instances of this application. Confirmed software supplychain attacks have led to the injection of backdoors into popular open-source projects, and backdoors have been discovered in various router firmware. Manual code auditing for backdoors is challenging and existing semi-automated approaches can handle only a limited scope of programs and backdoors, while requiring manual reverse-engineering of the audited (binary) program. Graybox fuzzing (automated semi-randomized testing) has grown in popularity due to its success in discovering vulnerabilities and hence stands as a strong candidate for improved backdoor detection. However, current fuzzing knowledge does not offer any means to detect the triggering of a backdoor at runtime. In this work we introduce ROSA, a novel approach (and tool) which combines a state-of-the-art fuzzer (AFL++) with a new metamorphic test oracle, capable of detecting runtime backdoor triggers. To facilitate the evaluation of ROSA, we have created ROSARUM, the first openly available benchmark for assessing the detection of various backdoors in diverse programs. Experimental evaluation shows that ROSA has a level of robustness, speed and automation similar to classical fuzzing. It finds all 17 authentic or synthetic backdooors from ROSARUM in 1 h 30 on average. Compared to existing detection tools, it can handle a diversity of backdoors and programs and it does not rely on manual reverse-engineering of the fuzzed binary code. Dimitri Kokkonis, Michaël Marcozzi, Emilien Decoux, Stefano Zacchiroli |
ICSE | 4 |
| 2025 | Altered Histories in Version Control System Repositories: Evidence from the Trenches
Solal Rapaport, Laurent Pautet, Samuel Tardieu, Stefano Zacchiroli |
ASE | 4 |
| 2025 | Does Functional Package Management Enable Reproducible Builds at Scale? YesabstractReproducible Builds (R-B) guarantee that rebuilding a software package from source leads to bitwise identical artifacts. R-B is a promising approach to increase the integrity of the software supply chain, when installing open source software built by third parties. Unfortunately, despite success stories like high build reproducibility levels in Debian packages, uncertainty remains among field experts on the scalability of R-B to very large package repositories. In this work, we perform the first large-scale study of bitwise reproducibility, in the context of the Nix functional package manager, rebuilding 709816 packages from historical snapshots of the nixpkgs repository, the largest cross-ecosystem open source software distribution, sampled in the period 2017-2023. We obtain very high bitwise reproducibility rates, between 69 and $91 \%$ with an upward trend, and even higher rebuildability rates, over $99 \%$. We investigate unreproducibility causes, showing that about $15 \%$ of failures are due to embedded build dates. We release a novel dataset with all build statuses, logs, as well as full “diffoscopes”: recursive diffs of where unreproducible build artifacts differ. Julien Malka, Stefano Zacchiroli, Théo Zimmermann |
MSR | 2 |
| 2025 | Wild SBOMs: a Large-scale Dataset of Software Bills of Materials from Public CodeabstractDevelopers gain productivity by reusing readily available Free and Open Source Software (FOSS) components. Such practices also bring some difficulties, such as managing licensing, components and related security. One approach to handle those difficulties is to use Software Bill of Materials (SBOMs). While there have been studies on the readiness of practitioners to embrace SBOMs and on the SBOM tools ecosystem, a large scale study on SBOM practices based on SBOM files produced in the wild is still lacking. A starting point for such a study is a large dataset of SBOM files found in the wild. We introduce such a dataset, consisting of over 78 thousand unique SBOM files, deduplicated from those found in over 94 million repositories. We include metadata that contains the standard and format used, quality score generated by the tool sbomqs, number of revisions, filenames and provenance information. Finally, we give suggestions and examples of research that could bring new insights on assessing and improving SBOM real practices. Luis Soeiro, Thomas Robert 0003, Stefano Zacchiroli |
MSR | 3 |
| 2025 | Is This You, LLM? Recognizing AI-written Programs with Multilingual Code StylometryabstractWith the increasing popularity of LLM-based code completers, like GitHub Copilot, the interest in automatically detecting AI-generated code is also increasing-in particular in contexts where the use of LLMs to program is forbidden by policy due to security, intellectual property, or ethical concerns. We introduce a novel technique for AI code stylometry, i.e., the ability to distinguish code generated by LLMs from code written by humans, based on a transformer-based encoder classifier. Differently from previous work, our classifier is capable of detecting AI-written code across 10 different programming languages with a single machine learning model, maintaining high average accuracy across all languages (84.1% ± 3.8%). Together with the classifier we also release H-AIRosettaMP, a novel open dataset for AI code stylometry tasks, consisting of 121 247 code snippets in 10 popular programming languages, labeled as either human-written or AI-generated. The experimental pipeline (dataset, training code, resulting models) is the first fully reproducible one for the AI code stylometry task. Most notably our experiments rely only on open LLMs, rather than on proprietary/closed ones like ChatGPT. Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli |
SANER | 3 |
| 2025 | The impact of the COVID-19 pandemic on women's contribution to public code
Annalí Casanueva Artís, Davide Rossi 0002, Stefano Zacchiroli, Théo Zimmermann |
Empir. Softw. Eng. | 3 |
| 2025 | On the compressibility of large-scale source code datasetsabstractStoring ultra-large amounts of unstructured data (often called objects or blobs) is a fundamental task for several object-based storage engines, data warehouses, data-lake systems, and key–value stores. These systems cannot currently leverage similarities between objects, which could be vital in improving their space and time performance. An important use case in which we can expect the objects to be highly similar is the storage of large-scale versioned source code datasets, such as the Software Heritage Archive (Di Cosmo and Zacchiroli, 2017). This use case is particularly interesting given the extraordinary size (1.5 PiB), the variegated nature, and the high repetitiveness of the at-issue corpus. In this paper we discuss and experiment with content- and context-based compression techniques for source-code collections that tailor known and novel tools to this setting in combination with state-of-the-art general-purpose compressors and the information coming from the Software Heritage Graph. We experiment with our compressors over a random sample of the entire corpus, and four large samples of source code files written in different popular languages: C/C ++ , Java, JavaScript, and Python. We also consider two scenarios of usage for our compressors, called Backup and File-Access scenario, where the latter adds to the former the support for single file retrieval. As a net result, our experiments show (i) how much “compressible” each language is, (ii) which content- or context-based techniques compress better and are faster to (de)compress by possibly supporting individual file access, and (iii) the ultimate compressed size that, according to our estimate, our best solution could achieve in storing all the source code written in these languages and available in the Software Heritage Archive: namely, in 3 TiB (down from their original 78 TiB total size, with an average compression ratio of 4%). Antonio Boffa, Roberto Di Cosmo, Paolo Ferragina, Andrea Guerra, Giovanni Manzini, Giorgio Vinciguerra, Stefano Zacchiroli |
J. Syst. Softw. | 7 |
| 2023 | Assessing the Threat Level of Software Supply Chains with the Log ModelabstractThe use of free and open source software (FOSS) components in all software systems is estimated to be above 90%. With such high usage and because of the heterogeneity of FOSS tools, repositories, developers and ecosystem, the level of complexity of managing software development has also increased. This has amplified both the attack surface for malicious actors and the difficulty of making sure that the software products are free from threats. The rise of security incidents involving high profile attacks is evidence that there is still much to be done to safeguard software products and the FOSS supply chain.Software Composition Analysis (SCA) tools and the study of attack trees help with improving security. However, they still lack the ability to comprehensively address how interactions within the software supply chain may impact security.This work presents a novel approach of assessing threat levels in FOSS supply chains with the log model. This model provides information capture and threat propagation analysis that not only account for security risks that may be caused by attacks and the usage of vulnerable software, but also how they interact with the other elements to affect the threat level for any element in the model. Luis Soeiro, Thomas Robert 0003, Stefano Zacchiroli |
IEEE Big Data | 3 |
| 2023 | The software heritage license dataset (2022 edition)
Jesús M. González-Barahona, Sergio Raúl Montes León, Gregorio Robles, Stefano Zacchiroli |
Empir. Softw. Eng. | 4 |
| 2023 | Using the uniqueness of global identifiers to determine the provenance of Python software source code
Daniel M. Germán, Stefano Zacchiroli |
Empir. Softw. Eng. | 3 |
| 2023 | Robust and scalable content-and-structure indexingabstractAbstract Frequent queries on semi-structured hierarchical data are Content-and-Structure (CAS) queries that filter data items based on their location in the hierarchical structure and their value for some attribute. We propose the Robust and Scalable Content-and-Structure (RSCAS) index to efficiently answer CAS queries on big semi-structured data. To get an index that is robust against queries with varying selectivities, we introduce a novel dynamic interleaving that merges the path and value dimensions of composite keys in a balanced manner. We store interleaved keys in our trie-based RSCAS index, which efficiently supports a wide range of CAS queries, including queries with wildcards and descendant axes. We implement RSCAS as a log-structured merge tree to scale it to data-intensive applications with a high insertion rate. We illustrate RSCAS’s robustness and scalability by indexing data from the Software Heritage (SWH) archive, which is the world’s largest, publicly available source code archive. Kevin Wellenzohn, Michael H. Böhlen, Sven Helmer, Antoine Pietri, Stefano Zacchiroli |
VLDB J. | 5 |
| 2022 | Software Artifact Mining in Software Engineering Conferences: A Meta-AnalysisabstractBackground: Software development results in the production of various types of artifacts: source code, version control system metadata, bug reports, mailing list conversations, test data, etc. Empirical software engineering (ESE) has thrived mining those artifacts to uncover the inner workings of software development and improve its practices. But which artifacts are studied in the field is a moving target, which we study empirically in this paper. Aims: We quantitatively characterize the most frequently mined and co-mined software artifacts in ESE research and the research purposes they support. Method: We conduct a meta-analysis of artifact mining studies published in 11 top conferences in ESE, for a total of 9621 papers. We use natural language processing (NLP) techniques to characterize the types of software artifacts that are most often mined and their evolution over a 16-year period (2004–2020). We analyze the combinations of artifact types that are most often mined together, as well as the relationship between study purposes and mined artifacts. Results: We find that: (1) mining happens in the vast majority of analyzed papers, (2) source code and test data are the most mined artifacts, (3) there is an increasing interest in mining novel artifacts, together with source code, (4) researchers are most interested in the evaluation of software systems and use all possible empirical signals to support that goal. Conclusions: Our study presents a meta analysis of the usage of software artifacts in the field over a period of 16 years using NLP techniques. Zeinab Abou Khalil, Stefano Zacchiroli |
ESEM | 2 |
| 2022 | The General Index of Software Engineering PapersabstractWe introduce the General Index of Software Engineering Papers, a dataset of fulltext-indexed papers from the most prominent scientific venues in the field of Software Engineering. The dataset includes both complete bibliographic information and indexed n-grams (sequence of contiguous words after removal of stopwords and non-words, for a total of 577 276 382 unique n-grams in this release) with length 1 to 5 for 44 581 papers retrieved from 34 venues over the 1971--2020 period. Zeinab Abou Khalil, Stefano Zacchiroli |
MSR | 2 |
| 2022 | Geographic Diversity in Public Code Contributions: An Exploratory Large-Scale Study Over 50 YearsabstractWe conduct an exploratory, large-scale, longitudinal study of 50 years of commits to publicly available version control system repositories, in order to characterize the geographic diversity of contributors to public code and its evolution over time. We analyze in total 2.2 billion commits collected by Software Heritage from 160 million projects and authored by 43 million authors during the 1971--2021 time period. We geolocate developers to 12 world regions derived from the United Nation geoscheme, using as signals email top-level domains, author names compared with names distributions around the world, and UTC offsets mined from commit metadata. Davide Rossi 0002, Stefano Zacchiroli |
MSR | 2 |
| 2022 | A Large-scale Dataset of (Open Source) License Text VariantsabstractWe introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive---the largest publicly available archive of FOSS source code with accompanying development history---all versions of files whose names are commonly used to convey licensing terms to software users and developers. Stefano Zacchiroli |
MSR | 1 |
| 2022 | Efficient Prior Publication Identification for Open Source CodeabstractFree/Open Source Software (FOSS) enables large-scale reuse of preexisting software components. The main drawback is increased complexity in software supply chain management. A common approach to tame such complexity is automated open source compliance, which consists in automating the verification of adherence to various open source management best practices about license obligation fulfillment, vulnerability tracking, software composition analysis, and nearby concerns. Daniele Serafini, Stefano Zacchiroli |
OpenSym | 2 |
| 2020 | Forking Without Clicking: on How to Identify Software Repository ForksabstractThe notion of software "fork" has been shifting over time from the (negative) phenomenon of community disagreements that result in the creation of separate development lines and ultimately software products, to the (positive) practice of using distributed version control system (VCS) repositories to collaboratively improve a single product without stepping on each others toes. In both cases the VCS repositories participating in a fork share parts of a common development history. Antoine Pietri, Guillaume Rousseau, Stefano Zacchiroli |
MSR | 3 |
| 2020 | Determining the Intrinsic Structure of Public Software Development HistoryabstractBackground. Collaborative software development has produced a wealth of version control system (VCS) data that can now be analyzed in full. Little is known about the intrinsic structure of the entire corpus of publicly available VCS as an interconnected graph. Understanding its structure is needed to determine the best approach to analyze it in full and to avoid methodological pitfalls when doing so. Antoine Pietri, Guillaume Rousseau, Stefano Zacchiroli |
MSR | 3 |
| 2020 | The Software Heritage Graph Dataset: Large-scale Analysis of Public Software Development HistoryabstractSoftware Heritage is the largest existing public archive of software source code and accompanying development history. It spans more than five billion unique source code files and one billion unique commits, coming from more than 80 million software projects. These software artifacts were retrieved from major collaborative development platforms (e.g., GitHub, GitLab) and package repositories (e.g., PyPI, Debian, NPM), and stored in a uniform representation linking together source code files, directories, commits, and full snapshots of version control systems (VCS) repositories as observed by Software Heritage during periodic crawls. This dataset is unique in terms of accessibility and scale, and allows to explore a number of research questions on the long tail of public software development, instead of solely focusing on "most starred" repositories as it often happens. Antoine Pietri, Diomidis Spinellis, Stefano Zacchiroli |
MSR | 3 |
| 2020 | Dependency Solving Is Still Hard, but We Are Getting Better at ItabstractDependency solving is a hard (NP-complete) problem in all non-trivial component models due to either mutually incompatible versions of the same packages or explicitly declared package conflicts. As such, software upgrade planning needs to rely on highly specialized dependency solvers, lest falling into pitfalls such as incompleteness—a combination of package versions that satisfy dependency constraints does exist, but the package manager is unable to find it. In this paper we look back at proposals from dependency solving research dating back a few years. Specifically, we review the idea of treating dependency solving as a separate concern in package manager implementations, relying on generic dependency solvers based on tried and tested techniques such as SAT solving, PBO, MILP, etc. By conducting a census of dependency solving capabilities in state-of-the-art package managers we conclude that some proposals are starting to take off (e.g., SAT-based dependency solving) while—with few exceptions—others have not (e.g., outsourcing dependency solving to reusable components). We reflect on why that has been the case and look at novel challenges for dependency solving that have emerged since. Pietro Abate, Roberto Di Cosmo, Georgios Gousios, Stefano Zacchiroli |
SANER | 4 |
| 2020 | Ultra-Large-Scale Repository Analysis via Graph CompressionabstractWe consider the problem of mining the development history—as captured by modern version control systems—of ultra-large-scale software archives (e.g., tens of millions software repositories corresponding). We show that graph compression techniques can be applied to the problem, dramatically reducing the hardware resources needed to mine similarly-sized corpus. As a concrete use case we compress the full Software Heritage archive, consisting of 5 billion unique source code files and 1 billion unique commits, harvested from more than 80 million software projects—encompassing a full mirror of GitHub. The resulting compressed graph fits in less than 100 GB of RAM, corresponding to a hardware cost of less than 300 U.S. dollars. We show that the compressed in-memory representation of the full corpus can be accessed with excellent performances, with edge lookup times close to memory random access. As a sample exploitation experiment we show that the compressed graph can be used to conduct clone detection at this scale, benefiting from main memory access speed. Paolo Boldi, Antoine Pietri, Sebastiano Vigna, Stefano Zacchiroli |
SANER | 4 |
| 2020 | Software provenance tracking at the scale of public source code
Guillaume Rousseau, Roberto Di Cosmo, Stefano Zacchiroli |
Empir. Softw. Eng. | 3 |
| 2019 | The software heritage graph dataset: public software development under one roofabstractSoftware Heritage is the largest existing public archive of software source code and accompanying development history: it currently spans more than five billion unique source code files and one billion unique commits, coming from more than 80 million software projects. This paper introduces the Software Heritage graph dataset: a fully-deduplicated Merkle DAG representation of the Software Heritage archive. The dataset links together file content identifiers, source code directories, Version Control System (VCS) commits tracking evolution over time, up to the full states of VCS repositories as observed by Software Heritage during periodic crawls. The dataset's contents come from major development forges (including GitHub and GitLab), FOSS distributions (e.g., Debian), and language-specific package managers (e.g., PyPI). Crawling information is also included, providing timestamps about when and where all archived source code artifacts have been observed in the wild. The Software Heritage graph dataset is available in multiple formats, including downloadable CSV dumps and Apache Parquet files for local use, as well as a public instance on Amazon Athena interactive query service for ready-to-use powerful analytical processing. Source code file contents are cross-referenced at the graph leaves, and can be retrieved through individual requests using the Software Heritage archive API. Antoine Pietri, Diomidis Spinellis, Stefano Zacchiroli |
MSR | 3 |
| 2018 | Spacetime Characterization of Real-Time Collaborative EditingabstractReal-Time Collaborative Editing (RTCE) is a popular way of instrumenting cooperative work on documents, in particular on the Web. Little is known in the literature yet about RTCE usage patterns in the real world. In this paper we study how a popular RTCE editor (Etherpad) is used in the wild, digging into the edit histories of a large collection of documents (about 14 000 pads), retrieved from one of the most popular public instances of the platform, hosted by the Wikimedia Foundation. The pad analysis is supported by a novel conceptual model that allows to label edit operations as "collaborative" or not depending on their distance-in edit position (space), edit time, or spacetime (both)-from edits made by other authors. The model is applied to classify all edits from the pad corpus. Classification results are further used to characterize the collaboration behavior of pad authors. Findings show that: 1) about half of the pads have a single author and hence witnessed no collaboration; 2) collaboration on common document parts happens often, but it happens asynchronously with authors taking turns in editing; and 3) simultaneous editing of common document parts happens very rarely. These findings help in revisiting early RTCE design decisions (e.g., the granularity of conflict management in RTCE protocols) and give insights on how to address novel needs (e.g., end-to-end encryption and offline editing). Gabriele D'Angelo, Angelo Di Iorio, Stefano Zacchiroli |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2017 | Software Heritage: Scholarly and Educational Synergies with Preserving Our Software Commonsabstractembedded in software source code, that is publicly available and can be freely altered and reused. Free and Open Source Software (FOSS) constitutes the bulk of it. Sadly we seem to be at increasing risk of losing this precious heritage built by the FOSS community over the paste decades: code hosting sites shut down when their popularity decreases, tapes of ancient versions of our toolchain (bit-)rot in basements, etc. The ambitious goal of the Software Heritage project is to contribute to address this risk, by collecting, preserving, and sharing all publicly available software in source code form. Together with its complete development history, as captured by state-of-the art version control systems. Stefano Zacchiroli |
ITiCSE | 1 |
| 2017 | The Debsources Dataset: two decades of free and open source softwareabstractWe present the Debsources Dataset: source code and related metadata spanning two decades of Free and Open Source Software (FOSS) history, seen through the lens of the Debian distribution. The dataset spans more than 3 billion lines of source code as well as metadata about them such as: size metrics (lines of code, disk usage), developer-defined symbols (ctags), file-level checksums (SHA1, SHA256, TLSH), file media types (MIME), release information (which version of which package containing which source code files has been released when), and license information (GPL, BSD, etc). The Debsources Dataset comes as a set of tarballs containing deduplicated unique source code files organized by their SHA1 checksums (the source code), plus a portable PostgreSQL database dump (the metadata). A case study is run to show how the Debsources Dataset can be used to easily and efficiently instrument very long-term analyses of the evolution of Debian from various angles (size, granularity, licensing, etc.), getting a grasp of major FOSS trends of the past two decades. The Debsources Dataset is Open Data, released under the terms of the CC BY-SA 4.0 license, and available for download from Zenodo with DOI reference 10.5281/zenodo.61089. Matthieu Caneill, Daniel M. Germán, Stefano Zacchiroli |
Empir. Softw. Eng. | 3 |
| 2015 | Automatic Application Deployment in the Cloud: from Practice to Theory and Back (Invited Paper)abstractThe problem of deploying a complex software application has been formally investigated in previous work by means of the abstract component model named Aeolus. As the problem turned out to be undecidable, simplified versions of the model were investigated in which decidability was restored by introducing limitations on the ways components are described. In this paper, we take an opposite approach, and investigate the possibility to address a relaxed version of the deployment problem without limiting the expressiveness of the component model. We identify three problems to be solved in sequence: (i) the verification of the existence of a final configuration in which all the constraints imposed by the single components are satisfied, (ii) the generation of a concrete configuration satisfying such constraints, and (iii) the synthesis of a plan to reach such a configuration possibly going through intermediary configurations that violate the non-functional constraints. Roberto Di Cosmo, Michael Lienhardt, Jacopo Mauro, Stefano Zacchiroli, Gianluigi Zavattaro, Jakub Zwolakowski |
CONCUR | 4 |
| 2015 | Automatic Deployment of Services in the Cloud with Aeolus Blender
Roberto Di Cosmo, Antoine Eiche, Jacopo Mauro, Stefano Zacchiroli, Gianluigi Zavattaro, Jakub Zwolakowski |
ICSOC | 4 |
| 2015 | Mining Component Repositories for Installability IssuesabstractComponent repositories play an increasingly relevant role in software life-cycle management, from software distribution to end-user, to deployment and upgrade management. Software components shipped via such repositories are equipped with rich metadata that describe their relationship (e.g., Dependencies and conflicts) with other components. In this practice paper we show how to use a tool, distcheck, that uses component metadata to identify all the components in a repository that cannot be installed (e.g., Due to unsatisfiable dependencies), provides detailed information to help developers understanding the cause of the problem, and fix it in the repository. We report about detailed analyses of several repositories: the Debian distribution, the OPAM package collection, and Drupal modules. In each case, distcheck is able to efficiently identify not installable components and provide valuable explanations of the issues. Our experience provides solid ground for generalizing the use of distcheck to other component repositories. Pietro Abate, Roberto Di Cosmo, Louis Gesbert, Fabrice Le Fessant, Ralf Treinen, Stefano Zacchiroli |
MSR | 6 |
| 2015 | The Debsources Dataset: Two Decades of Debian Source Code MetadataabstractWe present the Debsources Dataset: distribution metadata and source code metrics spanning two decades of Free and Open Source Software (FOSS) history, seen through the lens of the Debian distribution. Debsources is a software platform used to gather, search, and publish on the Web the full source code of the Debian operating system, as well as measures about it. A notable public instance of Debsources is available at http://sources.debian.net, it includes both current and historical releases of Debian. Plugins to compute popular source code metrics (lines of code, defined symbols, disk usage) and other derived data (e.g., Checksums) have been written, integrated, and run on all the source code available on sources.debian.net. The Debsources Dataset is a PostgreSQL database dump of sources.debian.net metadata, as of February 10th, 2015. The dataset contains both Debian-specific metadata -- e.g., which software packages are available in which release, which source code file belong to which package, release dates, etc. -- and source code information gathered by running Debsources plugins. The Debsources Dataset offer a very long-term historical view of the macro-level evolution and constitution of FOSS through the lens of popular, representative FOSS projects of their times. Stefano Zacchiroli |
MSR | 1 |
| 2015 | Editorial
Angelo Di Iorio, Davide Rossi 0002, Stefano Zacchiroli |
J. Web Eng. | 3 |
| 2014 | Debsources: live and historical views on macro-level software evolutionabstractContext. Software evolution has been an active field of research in recent years, but studies on macro-level software evolution---i.e., on the evolution of large software collections over many years---are scarce, despite the increasing popularity of intermediate vendors as a way to deliver software to final users. Matthieu Caneill, Stefano Zacchiroli |
ESEM | 2 |
| 2014 | Automated synthesis and deployment of cloud applicationsabstractComplex networked applications are assembled by connecting software components distributed across multiple machines. Building and deploying such systems is a challenging problem which requires a significant amount of expertise: the system architect must ensure that all component dependencies are satisfied, avoid conflicting components, and add the right amount of component replicas to account for quality of service and fault-tolerance. In a cloud environment, one also needs to minimize the virtual resources provisioned upfront, to reduce the cost of operation. Once the full architecture is designed, it is necessary to correctly orchestrate the deployment phase, to ensure all components are started and connected in the right order. Roberto Di Cosmo, Michael Lienhardt, Ralf Treinen, Stefano Zacchiroli, Jakub Zwolakowski, Antoine Eiche, Alexis Agahi |
ASE | 4 |
| 2014 | Aeolus: A component model for the cloud
Roberto Di Cosmo, Jacopo Mauro, Stefano Zacchiroli, Gianluigi Zavattaro |
Inf. Comput. | 3 |
| 2014 | Learning from the future of component repositories
Pietro Abate, Roberto Di Cosmo, Ralf Treinen, Stefano Zacchiroli |
Sci. Comput. Program. | 4 |
| 2014 | Web Technologies: Selected & extended papers from WT ACM SAC 2012
Angelo Di Iorio, Davide Rossi 0002, Stefano Zacchiroli |
Sci. Comput. Program. | 3 |
| 2013 | Component Reconfiguration in the Presence of Conflicts
Roberto Di Cosmo, Jacopo Mauro, Stefano Zacchiroli, Gianluigi Zavattaro |
ICALP (2) | 3 |
| 2013 | A modular package manager architecture
Pietro Abate, Roberto Di Cosmo, Ralf Treinen, Stefano Zacchiroli |
Inf. Softw. Technol. | 4 |
| 2013 | EditorialabstractThis special issue of Software: Practice and Experience contains extended versions of the best papers accepted for the Web Technologies (WT) track of the 25th ACM Symposium on Applied Computing (SAC), which was held in Taichung, Taiwan in March 2011. In recent years the WT track has attracted an increasing number of high-quality contributions on everything web, by researchers and practitioners from both industry and academia. Because of the nature of the Web, the applied research on these topics has the distinctive potential of leading to tangible changes in our everyday experience in a very short time frame. We are honored to have been able to witness such evolution serving as track chairs for the last 4 years, and we hope to give you a representative glimpse of it with the articles contained in this special issue. In 2011, the WT track has received 31 submissions from 19 countries and accepted ten full papers (for an acceptance rate of about 32%). In addition to that, three submissions have been presented as posters on-site, during the SAC conference. We have selected four among the best papers of the 2011 edition, with a focus on recent evolution of the World Wide Web as an everyday and for everyone platform. Given its focus on the dissemination of tangible software experiences, Software: Practice and Experience is a perfect match for both the research topics and the mixed academic/practitioner approach of the WT track of ACM SAC. We are therefore confident to meet your interests, and we would like to wish you a very good read! Angelo Di Iorio, Davide Rossi 0002, Stefano Zacchiroli |
Softw. Pract. Exp. | 3 |
| 2012 | Why do software packages conflict?abstractDetermining whether two or more packages cannot be installed together is an important issue in the quality assurance process of package-based distributions. Unfortunately, the sheer number of different configurations to test makes this task particularly challenging, and hundreds of such incompatibilities go undetected by the normal testing and distribution process until they are later reported by a user as bugs that we call “conflict defects”. We performed an extensive case study of conflict defects extracted from the bug tracking systems of Debian and Red Hat. According to our results, conflict defects can be grouped into five main categories. We show that with more detailed package meta-data, about 30 % of all conflict defects could be prevented relatively easily, while another 30 % could be found by targeted testing of packages that share common resources or characteristics. These results allow us to make precise suggestions on how to prevent and detect conflict defects in the future. Cyrille Artho, Kuniyasu Suzaki, Roberto Di Cosmo, Ralf Treinen, Stefano Zacchiroli |
MSR | 5 |
| 2012 | Towards a Formal Component Model for the Cloud
Roberto Di Cosmo, Stefano Zacchiroli, Gianluigi Zavattaro |
SEFM | 2 |
| 2012 | Dependency solving: A separate concern in component evolution management
Pietro Abate, Roberto Di Cosmo, Ralf Treinen, Stefano Zacchiroli |
J. Syst. Softw. | 4 |
| 2011 | Supporting software evolution in component-based FOSS systems
Roberto Di Cosmo, Davide Di Ruscio, Patrizio Pelliccione, Alfonso Pierantonio, Stefano Zacchiroli |
Sci. Comput. Program. | 5 |
| 2010 | The Ultimate Debian Database: Consolidating bazaar metadata for Quality Assurance and data miningabstractFLOSS distributions like RedHat and Ubuntu require a lot more complex infrastructures than most other FLOSS projects. In the case of community-driven distributions like Debian, the development of such an infrastructure is often not very organized, leading to new data sources being added in an impromptu manner while hackers set up new services that gain acceptance in the community. Mixing and matching data is then harder than should be, albeit being badly needed for Quality Assurance and data mining. Massive refactoring and integration is not a viable solution either, due to the constraints imposed by the bazaar development model. This paper presents the Ultimate Debian Database (UDD), which is the countermeasure adopted by the Debian project to the above ¿data hell¿. UDD gathers data from various data sources into a single, central SQL database, turning Quality Assurance needs that could not be easily implemented before into simple SQL queries. The paper also discusses the customs that have contributed to the data hell, the lessons learnt while designing UDD, and its applications and potentialities for data mining on FLOSS distributions. Lucas Nussbaum, Stefano Zacchiroli |
MSR | 2 |
| 2010 | Feature Diagrams as Package Dependencies
Roberto Di Cosmo, Stefano Zacchiroli |
SPLC | 2 |
| 2009 | Towards a Model Driven Approach to Upgrade Complex Software Systems
Antonio Cicchetti, Davide Di Ruscio, Patrizio Pelliccione, Alfonso Pierantonio, Stefano Zacchiroli |
ENASE | 5 |
| 2009 | Strong dependencies between software componentsabstractComponent-based systems often describe context requirements in terms of explicit inter-component dependencies. Studying large instances of such systems - such as free and open source software (FOSS) distributions - in terms of declared dependencies between packages is appealing. It is however also misleading when the language to express dependencies is as expressive as Boolean formulae, which is often the case. In such settings, a more appropriate notion of component dependency exists: strong dependency. This paper introduces such notion as a first step towards modeling semantic, rather then syntactic, inter-component relationships. Furthermore, a notion of component sensitivity is derived from strong dependencies, with applications to quality assurance and to the evaluation of upgrade risks. An empirical study of strong dependencies and sensitivity is presented, in the context of one of the largest, freely available, component-based system. Pietro Abate, Roberto Di Cosmo, Jaap Boender, Stefano Zacchiroli |
ESEM | 4 |
| 2008 | Wiki content templatingabstractWiki content templating enables reuse of content structures among wiki pages. In this paper we present a thorough study of this widespread feature, showing how its two state of the art models (functional and creational templating) are sub-optimal. We then propose a third, better, model called lightly constrained (LC) templating and show its implementation in the Moin wiki engine. We also show how LC templating implementations are the appropriate technologies to push forward semantically rich web pages on the lines of (lowercase) semantic web and microformats. Angelo Di Iorio, Fabio Vitali, Stefano Zacchiroli |
WWW | 3 |
| 2007 | User Interaction with the Matita Proof Assistant
Andrea Asperti, Claudio Sacerdoti Coen, Enrico Tassi, Stefano Zacchiroli |
J. Autom. Reason. | 4 |
| 2004 | A Generative Approach to the Implementation of Language Bindings for the Document Object Model
Luca Padovani, Claudio Sacerdoti Coen, Stefano Zacchiroli |
GPCE | 3 |