Marco D'Ambros

dblp:66/2961 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
3since 2021 · last 2025
0009-0008-3765-0871ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 9 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Mining a Century of Swiss Trademark Data
abstract
This paper presents an approach for extracting trademark registration events from the Swiss Official Gazette of Commerce (SOGC), an official daily journal published by the Swiss Confederation since January 1883. Until 2001, the data is only available as scanned documents, which constitute the target dataset of this study. Our approach is composed of a chain of three steps based on state-of-the-art deep learning techniques. We leverage image classification to identify pages containing trademarks (macro segmentation); we apply object detection to identify the portion of the page corresponding to a registration event (micro segmentation); last, we perform information extraction using a document AI technique. We obtain a dataset of ca. 500,000 trademark registration events, extracted from a corpus of 430,000 pages. Each step of our workflow has relatively high accuracy: the macro and micro segmentation steps show precision and recall greater than 95% on a manually constructed dataset. The dataset offers a unique historical perspective on trademark registrations in Switzerland that is not available from any other source. Showcasing what can be achieved with the extracted information, we provide answers to a set of preliminary economics questions.
Daniel Travaglia, Jesper Findahl, Marco D'Ambros, Andrea Mocci, Raphael Parchet
DocEng3
2022 First come first served: the impact of file position on code review
abstract
The most popular code review tools (e.g., Gerrit and GitHub) present the files to review sorted in alphabetical order. Could this choice or, more generally, the relative position in which a file is presented bias the outcome of code reviews? We investigate this hypothesis by triangulating complementary evidence in a two-step study. First, we observe developers’ code review activity. We analyze the review comments pertaining to 219,476 Pull Requests (PRs) from 138 popular Java projects on GitHub. We found files shown earlier in a PR to receive more comments than files shown later, also when controlling for possible confounding factors: e.g., the presence of discussion threads or the lines added in a file. Second, we measure the impact of file position on defect finding in code review. Recruit- ing 106 participants, we conduct an online controlled experiment in which we measure participants’ performance in detecting two unrelated defects seeded into two different files. Participants are assigned to one of two treatments in which the position of the defective files is switched. For one type of defect, participants are not affected by its file’s position; for the other, they have 64% lower odds to identify it when its file is last as opposed to first. Overall, our findings provide evidence that the relative position in which files are presented has an impact on code reviews’ outcome; we discuss these results and implications for tool design and code review.
Enrico Fregnan, Larissa Braz, Marco D'Ambros, Gül Çalikli, Alberto Bacchelli
ESEC/SIGSOFT FSE3
2022 Can Git Repository Visualization Support Educators in Assessing Group Projects?
abstract
In the past years numerous software visualization tools have been introduced to support the analysis of software systems and their evolution as captured in the versioning systems. Usually the target audience of such tools comprises software engineering professionals. In this paper we argue that such tools are also beneficial for educators who need to evaluate the quality of software systems developed by students. However, since the needs of educators are different than those of the software engineering professionals, we discuss several educator needs first. We report several usage examples that we believe are useful for educators when using repository visualization tools. We illustrate them with examples from several student projects from different courses in two universities. We conclude with a series of considerations that should be heeded by both educators and future tool-builders.
Mircea Lungu, Rolf-Helge Pfeiffer, Marco D'Ambros, Michele Lanza 0001, Jesper Findahl
VISSOFT3
2020 Visualizing Interaction Data Inside & Outside the IDE to Characterize Developer Productivity
abstract
Work fragmentation is a common phenomenon in the workspace, and is detrimental to the actual work taking place. To measure and study the impact of work fragmentation in software development, several studies exploited interaction data, i.e., the data generated by the events performed by the developers in the IDE. However, the absence of information on activities performed outside the IDE could lead to a misclassification of development time. In fact, sometimes leaving the IDE is not an interruption of the task at hand, e.g., when consulting API documentation, or when discussing with colleagues in ad-hoc collaboration applications. In this paper, we propose Ferax, a data analytics platform that developers can leverage for retrospection and possibly to improve their productivity. The capabilities of Ferax are twofold: First, it extends Tako, a profiler to record IDE interaction data for Visual Studio Code, with information about which applications were used and which websites were visited. Second, to enable the understanding of productivity and interruptions on developer sessions, Ferax provides interactive visualizations that show the detailed sequence of events inside and outside the IDE, the switches the developer performs by classifying them as productive or possible interruptions, and the time distribution for application usage. As a preliminary evaluation of Ferax we have collected and analyzed real development sessions from a set of master students and two professional developers. We illustrate how a developer can leverage Ferax to characterize her usual habits, to elicit the impact of interruptions, and to better characterize sessions which were only apparently unproductive.
Gabriele Di Rosa, Andrea Mocci, Marco D'Ambros
VISSOFT3
2015 Object-focused environments revisited
Fernando Olivero, Michele Lanza 0001, Marco D'Ambros
Sci. Comput. Program.3
2013 Manhattan: Supporting real-time visual team activity awareness
abstract
Collaboration is essential for the development of complex software systems. An important aspect of collaboration is team awareness: The understanding of the activity of others that provides a context for one's activity. We claim that the current IDE support for awareness is inadequate: The typical setting is to rely on software configuration management systems (SCMs), which are based on an explicit check-out/check-in model. If developers rely only on SCMs information, they become aware of concurrent changes only when they commit their code to the repository. This generates problems such as complex merging and redundant work. Most tools to raise awareness notify developers of emerging conflicts in the form of textual notifications. We propose to improve the notification by using real-time visualization integrated in the IDE to notify developers of team activity. Our approach, implemented in a tool called Manhattan, eases team activity comprehension by relying on a city metaphor. Manhattan depicts a software system as a live city that changes as the underlying system evolves. Within the city, Manhattan renders team activity information, updating developers in real-time about changes implemented by the entire development team. Further, Manhattan provides programmers with immediate feedback about emerging conflicts in which they are involved.
Michele Lanza 0001, Marco D'Ambros, Alberto Bacchelli, Lile Hattori, Francesco Rigotti
ICPC2
2013 Answering software evolution questions: An empirical evaluation
Lile Hattori, Marco D'Ambros, Michele Lanza 0001, Mircea Lungu
Inf. Softw. Technol.2
2012 Ronda: A Fine Grained Collaborative Development Environment
Fernando Olivero, Michele Lanza 0001, Marco D'Ambros
CDVE3
2012 Method-level bug prediction
abstract
Researchers proposed a wide range of approaches to build effective bug prediction models that take into account multiple aspects of the software development process. Such models achieved good prediction performance, guiding developers towards those parts of their system where a large share of bugs can be expected. However, most of those approaches predict bugs on file-level. This often leaves developers with a considerable amount of effort to examine all methods of a file until a bug is located. This particular problem is reinforced by the fact that large files are typically predicted as the most bug-prone. In this paper, we present bug prediction models at the level of individual methods rather than at file-level. This increases the granularity of the prediction and thus reduces manual inspection efforts for developers. The models are based on change metrics and source code metrics that are typically used in bug prediction. Our experiments---performed on 21 Java open-source (sub-)systems---show that our prediction models reach a precision and recall of 84% and 88%, respectively. Furthermore, the results indicate that change metrics significantly outperform source code metrics.
Emanuel Giger, Marco D'Ambros, Martin Pinzger 0001, Harald C. Gall
ESEM2
2012 A Qualitative User Study on Preemptive Conflict Detection
abstract
Preemptive conflict detection is the act of detecting a potential merge conflict at an earlier stage than at check in time, and informing the involved developers about it. Researchers have proposed a number of tools and techniques to detect potential merge conflicts. However, barely any study has been conducted to investigate whether the adoption of such tools and techniques brings benefits to developers. We have conducted a qualitative user study to understand how developers behave when dealing with merging and how this behavior changes when they are exposed to preemptive conflict detection. We report on the analysis of the data collected in the user study, and provide an in-depth discussion on the findings derived from it.
Lile Hattori, Michele Lanza 0001, Marco D'Ambros
ICGSE3
2012 Content classification of development emails
abstract
Emails related to the development of a software system contain information about design choices and issues encountered during the development process. Exploiting the knowledge embedded in emails with automatic tools is challenging, due to the unstructured, noisy, and mixed language nature of this communication medium. Natural language text is often not well-formed and is interleaved with languages with other syntaxes, such as code or stack traces. We present an approach to classify email content at line level. Our technique classifies email lines in five categories (i.e., text, junk, code, patch, and stack trace) to allow one to subsequently apply ad hoc analysis techniques for each category. We evaluated our approach on a statistically significant set of emails gathered from mailing lists of four unrelated open source systems.
Alberto Bacchelli, Tommaso Dal Sasso, Marco D'Ambros, Michele Lanza 0001
ICSE3
2012 Evaluating defect prediction approaches: a benchmark and an extensive comparison
Marco D'Ambros, Michele Lanza 0001, Romain Robbes
Empir. Softw. Eng.1
2011 Miler: a toolset for exploring email data
abstract
Source code is the target and final outcome of software development. By focusing our research and analysis on source code only, we risk forgetting that software is the product of human efforts, where communication plays a pivotal role. One of the most used communications means are emails, which have become vital for any distributed development project. Analyzing email archives is non-trivial, due to the noisy and unstructured nature of emails, the vast amounts of information, the unstandardized storage systems, and the gap with development tools.
Alberto Bacchelli, Michele Lanza 0001, Marco D'Ambros
ICSE3
2011 Effective mining of software repositories
abstract
With the advent of open-source, the Internet, and the consequent widespread adoption of distributed development tools, such as software configuration management and issue tracking systems, a vast amount of valuable information concerning software development and evolution has become available.
Marco D'Ambros, Romain Robbes
ICSM1
2011 Software Evolution Comprehension: Replay to the Rescue
abstract
Developers often need to find answers to questions regarding the evolution of a system when working on its code base. While their information needs require data analysis spanning over different repository types, the source code repository has a pivotal role for program comprehension tasks. However, the coarse-grained nature of the data stored by commit-based software configuration management systems often makes it challenging for a developer to search for an answer. We present Replay, an Eclipse plug-in that allows one to explore the change history of a system by capturing the changes at a finer granularity level than commits, and by replaying the past changes chronologically inside the integrated development environment with the source code at hand. We conducted a controlled experiment to empirically assess whether Replay outperforms a baseline (SVN client in Eclipse) on helping developers to answer common questions related to software evolution. The experiment shows that Replay leads to a decrease in completion time with respect to a set of software evolution comprehension tasks.
Lile Hattori, Marco D'Ambros, Michele Lanza 0001, Mircea Lungu
ICPC2
2011 Enabling program comprehension through a visual object-focused development environment
abstract
Integrated development environments (IDEs) include many tools that provide the means to construct programs. Coincidentally, the very same IDEs are a primary vehicle for program comprehension. We claim that IDEs may be an impediment for program comprehension because they treat software elements as text, which may be counterproductive in the context of program understanding-where abstracting from the source text to the level of structural entities and relationships is the key. We are currently building Gaucho, a visual object-focused environment that allows developers to write programs by creating and manipulating lightweight and intuitive depictions of object-oriented constructs. The research question we investigate here is how such an environment compares with traditional IDEs when it comes to performing program comprehension tasks. To answer our question, we conducted a preliminary controlled experiment with eight subjects, comparing Gaucho against a traditional IDE. We found that Gaucho outperforms the IDE regarding the correctness of the tasks, while it is slower with respect to the completion time. Our preliminary results suggest that alternative-visual-IDEs may be superior to traditional IDEs as program comprehension aids.
Fernando Olivero, Michele Lanza 0001, Marco D'Ambros, Romain Robbes
VL/HCC3
2011 On porting software visualization tools to the web
Marco D'Ambros, Michele Lanza 0001, Mircea Lungu, Romain Robbes
Int. J. Softw. Tools Technol. Transf.1
2010 Are Popular Classes More Defect Prone?
Alberto Bacchelli, Marco D'Ambros, Michele Lanza 0001
FASE2
2010 Commit 2.0: enriching commit comments with visualization
abstract
Software developers use commit comments to document changes and as a mean of communication in their team. However, the support given by IDEs is restricted with this respect, as they limit the users to use only text to document changes.
Marco D'Ambros
ICSE (2)1
2010 Extracting Source Code from E-Mails
abstract
E-mails, used by developers and system users to communicate over a broad range of topics, offer a valuable source of information. If archived, e-mails can be mined to support program comprehension activities and to provide views of a software system that are alternative and complementary to those offered by the source code. However, e-mails are written in natural language, and therefore contain noise that makes it difficult to retrieve the important data. Thus, before conducting an effective system analysis and extracting data for program comprehension, it is necessary to select the relevant messages, and to expose only the meaningful information. In this work we focus both on classifying e-mails that hold fragments of the source code of a system, and on extracting the source code pieces inside the e-mail. We devised and analyzed a number of lightweight techniques to accomplish these tasks. To assess the validity of our techniques, we manually inspected and annotated a statistically significant number of e-mails from five unrelated open source software systems written in Java. With such a benchmark in place, we measured the effectiveness of each technique in terms of precision and recall.
Alberto Bacchelli, Marco D'Ambros, Michele Lanza 0001
ICPC2
2010 An extensive comparison of bug prediction approaches
abstract
Reliably predicting software defects is one of software engineering's holy grails. Researchers have devised and implemented a plethora of bug prediction approaches varying in terms of accuracy, complexity and the input data they require. However, the absence of an established benchmark makes it hard, if not impossible, to compare approaches. We present a benchmark for defect prediction, in the form of a publicly available data set consisting of several software systems, and provide an extensive comparison of the explanative and predictive power of well-known bug prediction approaches, together with novel approaches we devised. Based on the results, we discuss the performance and stability of the approaches with respect to our benchmark and deduce a number of insights on bug prediction models.
Marco D'Ambros, Michele Lanza 0001, Romain Robbes
MSR1
2010 Distributed and Collaborative Software Evolution Analysis with Churrasco
Marco D'Ambros, Michele Lanza 0001
Sci. Comput. Program.1
2009 Visual software evolution reconstruction
abstract
Abstract The analysis of the evolution of large software systems is challenging for many reasons, such as the retrieval and processing of historical information and the large quantity of data that must be dealt with. While recent advances in research have led to the solutions to these problems, a central question remains: How do we deal with this information in a methodical way and where do we start with our analysis? We present a methodology based on interactive visualizations that support the reconstruction of the evolution of software systems. We propose several visualizations which help us to perform software evolution analysis of a system ‘in the large’ and ‘in the small’, and apply them to two large systems. Copyright © 2009 John Wiley & Sons, Ltd.
Marco D'Ambros, Michele Lanza 0001
J. Softw. Maintenance Res. Pract.1
2009 Visualizing Co-Change Information with the Evolution Radar
abstract
Software evolution analysis provides a valuable source of information that can be used both to understand a system's design and predict its future development. While for many program comprehension purposes, it is sufficient to model a single version of a system, there are types of information that can only be recovered when the history of a system is taken into account. Logical coupling, the implicit dependency between software artifacts that have been changed together, is an example of such information. Previous research has dealt with low-level couplings between files, leading to an explosion of the data to be analyzed, or has abstracted the logical couplings to the level of modules, leading to a loss of detailed information. In this paper, we present a visualization-based approach that integrates logical coupling information at different levels of abstraction. This facilitates an in-depth analysis of the logical couplings, and at the same time, leads to a characterization of a system's modules in terms of their logical coupling. The presented approach supports the retrospective analysis of a software system and maintenance activities such as restructuring and redocumentation. We illustrate retrospective analysis on two large open-source software systems.
Marco D'Ambros, Michele Lanza 0001, Mircea Lungu
IEEE Trans. Software Eng.1
2008 Supporting software evolution analysis with historical dependencies and defect information
abstract
More than 90% of the cost of software is due to maintenance and evolution. Understanding the evolution of large software systems is a complex problem, which requires the use of various techniques and the support of tools. Several software evolution approaches put the emphasis on structural entities such as packages, classes and structural relationships. However, software evolution is not only about the history of software artifacts, but it also includes other types of data such as problem reports, mailing list archives etc. We propose an approach which focuses on historical dependencies and defects. We claim that they play an important role in software evolution and they are complementary to techniques based on structural information. We use historical dependencies and defect information to learn about a software system and detect potential problems in the source code. Moreover, based on design flaws detected in the source code, we predict the location of future bugs to focus maintenance activities on the buggy parts of the system. We validated our defect prediction by comparing it with the actual defects reported in the bug tracking system.
Marco D'Ambros
ICSM1