Andrea Mocci

dblp:43/236 · DBLP profile ↗
← Back
42ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-8426-5676ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 36Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
8 papers
Empirical software engineering · 28% Program analysis · 27% Software maintenance and evolution · 20%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › multimedia analysis and retrieval › video retrieval
content-based video retrieval
0.412019
Automatic Identification and Classification of Software Development Video Tutorial Fragments · IEEE Trans. Software Eng. 2019
Empirical software engineering
mining software repositories
0.322015
Use at your own risk: the Java unsafe API in the wild · OOPSLA 2015
Extracting structured data from natural language documents with island parsing · ASE 2011
Empirical software engineering
developer studies
0.322019
Free Hugs - Praising Developers for Their Actions · ICSE (2) 2015
Automatic Identification and Classification of Software Development Video Tutorial Fragments · IEEE Trans. Software Eng. 2019
Program analysis
dynamic analysis
0.222012
Runtime monitoring of component changes with Spy@Runtime · ICSE 2012
Synthesizing intensional behavior models by graph transformation · ICSE 2009
Program analysis › binary analysis
bytecode analysis
0.212015
Use at your own risk: the Java unsafe API in the wild · OOPSLA 2015
Software maintenance and evolution
code review
0.212015
ViDI: The Visual Design Inspector · ICSE (2) 2015
Empirical software engineering › developer studies › developer behavior
developer productivity
0.212015
Free Hugs - Praising Developers for Their Actions · ICSE (2) 2015
Programming languages and type systems
language-based safety
0.212015
Use at your own risk: the Java unsafe API in the wild · OOPSLA 2015
Program analysis › specification mining
behavioral model inference
0.112012
Runtime monitoring of component changes with Spy@Runtime · ICSE 2012
Software testing › software validation
behavioural validation
0.112012
Behavioral validation of JFSL specifications through model synthesis · ICSE 2012
Program verification › dynamic verification › runtime verification
contract validation
0.112012
Behavioral validation of JFSL specifications through model synthesis · ICSE 2012
Requirements engineering and software design › model-driven engineering
model synthesis
0.112012
Behavioral validation of JFSL specifications through model synthesis · ICSE 2012
Program analysis › dynamic analysis
runtime monitoring
0.112012
Runtime monitoring of component changes with Spy@Runtime · ICSE 2012
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
structured information extraction
0.112011
Extracting structured data from natural language documents with island parsing · ASE 2011
Requirements engineering and software design › model-driven engineering › model synthesis
behavior model synthesis
0.112009
Synthesizing intensional behavior models by graph transformation · ICSE 2009
Program analysis
specification mining
0.112009
Synthesizing intensional behavior models by graph transformation · ICSE 2009
Information retrieval › interactive information retrieval
browsing
0.112017
Supporting software developers with a holistic recommender system · ICSE 2017
Information retrieval
multimedia analysis and retrieval
0.112016
Too long; didn't watch!: extracting relevant fragments from software development video tutorials · ICSE 2016
Information retrieval › multimedia analysis and retrieval
video retrieval
0.112016
Too long; didn't watch!: extracting relevant fragments from software development video tutorials · ICSE 2016
Visualization and visual analytics
software visualization
0.112015
ViDI: The Visual Design Inspector · ICSE (2) 2015
Systems and software security
memory safety
0.112015
Use at your own risk: the Java unsafe API in the wild · OOPSLA 2015
Programming languages and type systems › specification language
algebraic specification
0.012012
Behavioral validation of JFSL specifications through model synthesis · ICSE 2012
Software maintenance and evolution
program comprehension
0.012011
Extracting structured data from natural language documents with island parsing · ASE 2011
Graph algorithms and graph theory › graph theory
graph transformation
0.012009
Synthesizing intensional behavior models by graph transformation · ICSE 2009

Methods — techniques the papers use, named apart from their topics

video segmentation · 0.8fragment classification · 0.8static analysis · 0.4repository mining · 0.4black-box dynamic inference · 0.3semantic relationship analysis · 0.3interactive visualization · 0.3video analysis · 0.2text mining · 0.2island parsing · 0.2graph transformation · 0.2finite-state machine inference · 0.2theorem proving · 0.1small-scope synthesis · 0.1
YearPublicationVenuePosition
2025 What Makes a Level Hard in Super Mario Maker 2?
abstract
International audience
Carlo A. Furia, Andrea Mocci
CoG2
2025 Mining a Century of Swiss Trademark Data
abstract
This paper presents an approach for extracting trademark registration events from the Swiss Official Gazette of Commerce (SOGC), an official daily journal published by the Swiss Confederation since January 1883. Until 2001, the data is only available as scanned documents, which constitute the target dataset of this study. Our approach is composed of a chain of three steps based on state-of-the-art deep learning techniques. We leverage image classification to identify pages containing trademarks (macro segmentation); we apply object detection to identify the portion of the page corresponding to a registration event (micro segmentation); last, we perform information extraction using a document AI technique. We obtain a dataset of ca. 500,000 trademark registration events, extracted from a corpus of 430,000 pages. Each step of our workflow has relatively high accuracy: the macro and micro segmentation steps show precision and recall greater than 95% on a manually constructed dataset. The dataset offers a unique historical perspective on trademark registrations in Switzerland that is not available from any other source. Showcasing what can be achieved with the extracted information, we provide answers to a set of preliminary economics questions.
Daniel Travaglia, Jesper Findahl, Marco D'Ambros, Andrea Mocci, Raphael Parchet
DocEng4
2020 Visualizing Interaction Data Inside & Outside the IDE to Characterize Developer Productivity
abstract
Work fragmentation is a common phenomenon in the workspace, and is detrimental to the actual work taking place. To measure and study the impact of work fragmentation in software development, several studies exploited interaction data, i.e., the data generated by the events performed by the developers in the IDE. However, the absence of information on activities performed outside the IDE could lead to a misclassification of development time. In fact, sometimes leaving the IDE is not an interruption of the task at hand, e.g., when consulting API documentation, or when discussing with colleagues in ad-hoc collaboration applications. In this paper, we propose Ferax, a data analytics platform that developers can leverage for retrospection and possibly to improve their productivity. The capabilities of Ferax are twofold: First, it extends Tako, a profiler to record IDE interaction data for Visual Studio Code, with information about which applications were used and which websites were visited. Second, to enable the understanding of productivity and interruptions on developer sessions, Ferax provides interactive visualizations that show the detailed sequence of events inside and outside the IDE, the switches the developer performs by classifying them as productive or possible interruptions, and the time distribution for application usage. As a preliminary evaluation of Ferax we have collected and analyzed real development sessions from a set of master students and two professional developers. We illustrate how a developer can leverage Ferax to characterize her usual habits, to elicit the impact of interruptions, and to better characterize sessions which were only apparently unproductive.
Gabriele Di Rosa, Andrea Mocci, Marco D'Ambros
VISSOFT2
2019 Automatic Identification and Classification of Software Development Video Tutorial Fragments
abstract
Software development video tutorials have seen a steep increase in popularity in recent years. Their main advantage is that they thoroughly illustrate how certain technologies, programming languages, etc. are to be used. However, they come with a caveat: there is currently little support for searching and browsing their content. This makes it difficult to quickly find the useful parts in a longer video, as the only options are watching the entire video, leading to wasted time, or fast-forwarding through it, leading to missed information. We present an approach to mine video tutorials found on the web and enable developers to query their contents as opposed to just their metadata. The video tutorials are processed and split into coherent fragments, such that only relevant fragments are returned in response to a query. Moreover, fragments are automatically classified according to their purpose, such as introducing theoretical concepts, explaining code implementation steps, or dealing with errors. This allows developers to set filters in their search to target a specific type of video fragment they are interested in. In addition, the video fragments in CodeTube are complemented with information from other sources, such as Stack Overflow discussions, giving more context and useful information for understanding the concepts.
Luca Ponzanelli, Gabriele Bavota, Andrea Mocci, Rocco Oliveto, Massimiliano Di Penta, Sonia Haiduc, Barbara Russo, Michele Lanza 0001
IEEE Trans. Software Eng.3
2017 Supporting software developers with a holistic recommender system
abstract
The promise of recommender systems is to provide intelligent support to developers during their programming tasks. Such support ranges from suggesting program entities to taking into account pertinent Q&A pages. However, current recommender systems limit the context analysis to change history and developers' activities in the IDE, without considering what a developer has already consulted or perused, e.g., by performing searches from the Web browser. Given the faceted nature of many programming tasks, and the incompleteness of the information provided by a single artifact, several heterogeneous resources are required to obtain the broader picture needed by a developer to accomplish a task. We present Libra, a holistic recommender system. It supports the process of searching and navigating the information needed by constructing a holistic meta-information model of the resources perused by a developer, analyzing their semantic relationships, and augmenting the web browser with a dedicated interactive navigation chart. The quantitative and qualitative evaluation of Libra provides evidence that a holistic analysis of a developer's information context can indeed offer comprehensive and contextualized support to information navigation and retrieval during software development.
Luca Ponzanelli, Simone Scalabrino, Gabriele Bavota, Andrea Mocci, Rocco Oliveto, Massimiliano Di Penta, Michele Lanza 0001
ICSE4
2017 The code time machine
abstract
Exploring and analyzing the history of changes is an intrinsic part of software evolution comprehension. Existing tools that exploit the data residing in version control repositories provide only limited support for the intuitive navigation of code changes from a historical perspective. We present the Code Time Machine, a lightweight IDE plugin which uses visualization techniques to depict the history of any chosen file augmented with information mined from the underlying versioning system. Inspired by Apple's Time Machine, our tool allows both developers and the system itself to seamlessly move through time. A video of the Code Time Machine can be found at https://youtu.be/meblwFO95oA.
Emad Aghajani, Andrea Mocci, Gabriele Bavota, Michele Lanza 0001
ICPC2
2017 On the uniqueness of code redundancies
abstract
Code redundancy widely occurs in software projects. Researchers have investigated the existence, causes, and impacts of code redundancy, showing that it can be put to good use, for example in the context of code completion. When analyzing source code redundancy, previous studies considered software projects as sequences of tokens, neglecting the role of the syntactic structures enforced by programming languages. However, differences in the redundancy of such structures may jeopardize the performance of applications leveraging code redundancy. We present a study of the redundancy of several types of code constructs in a large-scale dataset of active Java projects mined from GitHub, unveiling that redundancy is not uniform and mainly resides in specific code constructs. We further investigate the implications of the locality of redundancy by analyzing the performance of language models when applied to code completion. Our study discloses the perils of exploiting code redundancy without taking into account its strong locality in specific code constructs.
Bin Lin 0008, Luca Ponzanelli, Andrea Mocci, Gabriele Bavota, Michele Lanza 0001
ICPC3
2017 How developers document pull requests with external references
abstract
Online resources of formal and informal documentation-such as reference manuals, forum discussions and tutorials-have become an asset to software developers, as they allow them to tackle problems and to learn about new tools, libraries, and technologies. This study investigates to what extent and for which purpose developers refer to external online resources when they contribute changes to a repository by raising a pull request. Our study involved (i) a quantitative analysis of over 150k URLs occurring in pull requests posted in GitHub, (ii) a manual coding of the kinds of software evolution activities performed in commits related to a statistically significant sample of 2,130 pull requests referencing external documentation resources, (iii) a survey with 69 participants, who provided feedback on how they use online resources and how they refer to them when filing a pull request. Results of the study indicate that, on the one hand, developers find external resources useful to learn something new or to solve specific problems, and they perceive useful referring such resources to better document changes. On the other hand, both interviews and repository mining suggest that external resources are still rarely referred in document changes.
Fiorella Zampetti, Luca Ponzanelli, Gabriele Bavota, Andrea Mocci, Massimiliano Di Penta, Michele Lanza 0001
ICPC4
2017 Investigating the Use of Code Analysis and NLP to Promote a Consistent Usage of Identifiers
abstract
Meaningless identifiers as well as inconsistent use of identifiers in the source code might hinder code readability and result in increased software maintenance efforts. Over the past years, effort has been devoted to promoting a consistent usage of identifiers across different parts of a system through approaches exploiting static code analysis and Natural Language Processing (NLP). These techniques have been evaluated in small-scale studies, but it is unclear how they compare to each other and how they complement each other. Furthermore, a full-fledged larger empirical evaluation is still missing.,,We aim at bridging this gap. We asked developers of five projects to assess the meaningfulness of the recommendations generated by three techniques, two already existing in the literature (one exploiting static analysis, one using NLP) and a novel one we propose. With a total of 922 rename refactorings evaluated, this is, to the best of our knowledge, the largest empirical study conducted to assess and compare rename refactoring tools promoting a consistent use of identifiers. Our study sheds light on the current state-of-the-art in rename refactoring recommenders, and indicates directions for future work.
Bin Lin 0008, Simone Scalabrino, Andrea Mocci, Rocco Oliveto, Gabriele Bavota, Michele Lanza 0001
SCAM3
2017 How to gamify software engineering
abstract
Software development, like any prolonged and intellectually demanding activity, can negatively affect the motivation of developers. This is especially true in specific areas of software engineering, such as requirements engineering, test-driven development, bug reporting and fixing, where the creative aspects of programming fall short. The developers' engagement might progressively degrade, potentially impacting their work's quality.
Tommaso Dal Sasso, Andrea Mocci, Michele Lanza 0001, Ebrisa Mastrodicasa
SANER2
2017 Mining structured data in natural language artifacts with island parsing
Alberto Bacchelli, Andrea Mocci, Anthony Cleve, Michele Lanza 0001
Sci. Comput. Program.2
2016 Too long; didn't watch!: extracting relevant fragments from software development video tutorials
abstract
When knowledgeable colleagues are not available, developers resort to offline and online resources, e.g., tutorials, mailing lists, and Q&A websites. These, however, need to be found, read, and understood, which takes its toll in terms of time and mental energy. A more immediate and accessible resource are video tutorials found on the web, which in recent years have seen a steep increase in popularity. Nonetheless, videos are an intrinsically noisy data source, and finding the right piece of information might be even more cumbersome than using the previously mentioned resources.
Luca Ponzanelli, Gabriele Bavota, Andrea Mocci, Massimiliano Di Penta, Rocco Oliveto, Barbara Russo, Sonia Haiduc, Michele Lanza 0001
ICSE3
2016 Taming the IDE with fine-grained interaction data
abstract
Integrated Development Environments (IDEs) lack effective support to browse complex relationships between source code elements. As a result, developers are often forced to exploit multiple user interface components at the same time, bringing the IDE into a complex, “chaotic” state. Keeping track of these relationships demands increased source code navigation and cognitive load, leading to productivity deficits documented in observational studies. Beyond small-scale studies, the amount and nature of the chaos experienced by developers in the wild is unclear, and more importantly it is unclear how to tame it. Based on a dataset of fine-grained interaction data, we propose several metrics to characterize and quantify the “level of chaos” of an IDE. Our results suggest that developers spend, on average, more than 30% of their time in a chaotic environment, and that this may affect their productivity. To support developers, we devise and evaluate simple strategies that automatically alter the UI of the IDE. We find that even simple strategies may considerably reduce the level of chaos both in terms of effective space occupancy and time spent in a chaotic environment.
Roberto Minelli, Andrea Mocci, Romain Robbes, Michele Lanza 0001
ICPC2
2016 What Makes a Satisficing Bug Report?
abstract
To ensure quality of software systems, developers use bug reports to track defects. It is in the interest of users and developers that bug reports provide the necessary information to ease the fixing process. Past research found that users do not provide the information that developers deem ideally useful to fix a bug. This raises an interesting question: What is the satisficing information to speed up the bug fixing process? We conducted an observational study on the relation between provided report information and its lifetime, considering more than 650,000 reports from open-source systems using popular bug trackers. We distilled a meta-model for a minimal bug report, establishing a basic layer of core features. We found that few fields influence the resolution time and that customized fields have little impact on it. We performed a survey to investigate what users deem easy to provide in a bug report.
Tommaso Dal Sasso, Andrea Mocci, Michele Lanza 0001
QRS2
2016 Visualizing the Evolution of Working Sets
abstract
As part of their daily work, developers interact with Integrated Development Environments (IDE), generating thousands of events. Together with other aspects of development, this data also captures the modus operandi of the developer, including all the program entities she interacted with during a development session. This "working set" (or context) is leveraged by developers to create and maintain their mental model of the software system at hand. Understanding how developers navigate and interact with source code during a development session is an open question. We present a novel visual approach to understand how working sets evolve during a development session. The visualization incrementally depicts all the program entities involved in a development session, the intensity of the developer activity on them, and the navigation paths that occurred between them. We visualized about a thousand development sessions, and categorized them according to their visual properties.
Roberto Minelli, Andrea Mocci, Michele Lanza 0001
VISSOFT2
2015 Free Hugs - Praising Developers for Their Actions
abstract
Developing software is a complex, intrinsically intellectual, and therefore ephemeral activity, also due to the intangible nature of the end product, the source code. There is a thin red line between a productive development session, where a developer actually does something useful and productive, and a session where the developer essentially produces "fried air", pieces of code whose quality and usefulness are doubtful at best. We believe that well-thought mechanisms of gamification built on fine-grained interaction information mined from the IDE can crystallize and reward good coding behavior. We present our preliminary experience with the design and implementation of a micro-gamification layer built into an object-oriented IDE, which at the end of each development session not only helps the developer to understand what he actually produced, but also praises him in case the development session was productive. Building on this, we envision an environment where the IDE reflects on the deeds of the developers and by providing a historical view also helps to track and reward long-term growth in terms of development skills, not dissimilar from the mechanics of role-playing games.
Roberto Minelli, Andrea Mocci, Michele Lanza 0001
ICSE (2)2
2015 ViDI: The Visual Design Inspector
abstract
We present ViDI (Visual Design Inspector), a novel code review tool which focuses on quality concerns and design inspection as its cornerstones. It leverages visualization techniques to represent the reviewed software and augments the visualization with the results of quality analysis tools. To effectively understand the contribution of a reviewer in terms of the impact of her changes on the overall system quality, ViDI supports the recording and further inspection of reviewing sessions. ViDI is an advanced prototype which we will soon release to the Pharo open-source community.
Yuriy Tymchuk, Andrea Mocci, Michele Lanza 0001
ICSE (2)2
2015 UrbanIt: Visualizing repositories everywhere
abstract
Software evolution is supported by a variety of tools that help developers understand the structure of a software system, analyze its history and support specific classes of analyses. However, the increasingly distributed nature of software development requires basic repository analyses to be always available to developers, even when they cannot access their workstation with full-fledged applications and command-line tools. We present URBANIT, a gesture-based tablet application for the iPad that supports the visualization of software repositories together with useful evolutionary analyses (e.g., version diff) and basic sharing features in a portable and mobile setting. URBANIT is paired with a web application that manages synchronization of multiple repositories.
Andrea Ciani, Roberto Minelli, Andrea Mocci, Michele Lanza 0001
ICSME3
2015 I know what you did last summer: an investigation of how developers spend their time
abstract
Developing software is a complex mental activity, requiring extensive technical knowledge and abstraction capabilities. The tangible part of development is the use of tools to read, inspect, edit, and manipulate source code, usually through an IDE (integrated development environment). Common claims about software development include that program comprehension takes up half of the time of a developer, or that certain UI (user interface) paradigms of IDEs offer insufficient support to developers. Such claims are often based on anecdotal evidence, throwing up the question of whether they can be corroborated on more solid grounds. We present an in-depth analysis of how developers spend their time, based on a fine-grained IDE interaction dataset consisting of ca. 740 development sessions by 18 developers, amounting to 200 hours of development time and 5 million of IDE events. We propose an inference model of development activities to precisely measure the time spent in editing, navigating and searching for artifacts, interacting with the UI of the IDE, and performing corollary activities, such as inspection and debugging. We report several interesting findings which in part confirm and reinforce some common claims, but also disconfirm other beliefs about software development.
Roberto Minelli, Andrea Mocci, Michele Lanza 0001
ICPC2
2015 The plague doctor: a promising cure for the window plague
abstract
Modern Integrated Development Environments (IDEs) are often affected by the "window plague", an overly crowded workspace with many open windows and tabs. The main cause is the lack of navigation support in IDEs, also due to the many -- and not always obvious -- complex relationships that exist between program entities. Researchers have shown that it is possible to mitigate the window plague by exploiting the data obtained by monitoring how developers interact with the user interface of the IDE. However, despite initial results the approach was never fully integrated in an IDE. In our previous work, we implemented DFlow, an automatic interaction profiler that monitors all the fine-grained interactions of the developer with the IDE. Here we present a first prototype of the Plague Doctor, a tool that seamlessly detects the windows that are less likely to be used in the future and automatically closes them. We discuss our long term vision on how to fully exploit the interaction data recorded by DFlow to provide a more effective cure for the window plague.
Roberto Minelli, Andrea Mocci, Michele Lanza 0001
ICPC2
2015 Towards visual reflexion models
abstract
Source code and models of a software system, like architectural views, tend to evolve separately and drift apart over time. Previous research has shown that it is possible to effectively relate them through a reflex ion model, defined as a "summarization of a software system from the viewpoint of a particular high-level model". While effective, the process of constructing and analyzing reflex ion models was supported by text-based tools with limited visual representation. With the original approach, it was relatively hard to understand which parts of the system were represented, and which parts of the system contributed to specific relations in the reflexion model. We present our vision on augmenting the construction and analysis of reflex ion models with visual support, effectively providing the basis for visual reflex ion models. We describe our approach, implemented as a web-based application, and two promising case studies involving two open-source projects.
Marcello Romanelli, Andrea Mocci, Michele Lanza 0001
ICPC2
2015 Summarizing Complex Development Artifacts by Mining Heterogeneous Data
abstract
Summarization is hailed as a promising approach to reduce the amount of information that must be taken in by the person who wants to understand development artifacts, such as pieces of code, bug reports, emails, etc. However, existing approaches treat artifacts as pure textual entities, disregarding the heterogeneous and partially structured nature of most artifacts, which contain intertwined pieces of distinct type, such as source code, diffs, stack traces, human language, etc. We present a novel approach to augment existing summarization techniques (such as LexRank) to deal with the heterogeneous and multidimensional nature of complex artifacts. Our preliminary results on heterogeneous artifacts suggest our approach outperforms the current text-based approaches.
Luca Ponzanelli, Andrea Mocci, Michele Lanza 0001
MSR2
2015 StORMeD: Stack Overflow Ready Made Data
abstract
Stack Overflow is the de facto Question and Answer (Q&A) website for developers, and it has been used in many approaches by software engineering researchers to mine useful data. However, the contents of a Stack Overflow discussion are inherently heterogeneous, mixing natural language, source code, stack traces and configuration files in XML or JSON format. We constructed a full island grammar capable of modeling the set of 700,000 Stack Overflow discussions talking about Java, building a heterogeneous abstract syntax tree (H-AST) of each post (question, answer or comment) in a discussion. The resulting dataset models every Stack Overflow discussion, providing a full H-AST for each type of structured fragment (i.e., JSON, XML, Java, Stack traces), and complementing this information with a set of basic meta-information like term frequency to enable natural language analyses. Our dataset allows the end-user to perform combined analyses of the Stack Overflow by visiting the H-AST of a discussion.
Luca Ponzanelli, Andrea Mocci, Michele Lanza 0001
MSR2
2015 Use at your own risk: the Java unsafe API in the wild
abstract
Java is a safe language. Its runtime environment provides strong safety guarantees that any Java application can rely on. Or so we think. We show that the runtime actually does not provide these guarantees---for a large fraction of today's Java code. Unbeknownst to many application developers, the Java runtime includes a "backdoor" that allows expert library and framework developers to circumvent Java's safety guarantees. This backdoor is there by design, and is well known to experts, as it enables them to write high-performance "systems-level" code in Java. For much the same reasons that safe languages are preferred over unsafe languages, these powerful---but unsafe---capabilities in Java should be restricted. They should be made safe by changing the language, the runtime system, or the libraries. At the very least, their use should be restricted. This paper is a step in that direction. We analyzed 74 GB of compiled Java code, spread over 86,479 Java archives, to determine how Java's unsafe capabilities are used in real-world libraries and applications. We found that 25% of Java bytecode archives depend on unsafe third-party Java code, and thus Java's safety guarantees cannot be trusted. We identify 14 different usage patterns of Java's unsafe capabilities, and we provide supporting evidence for why real-world code needs these capabilities. Our long-term goal is to provide a foundation for the design of new language features to regain safety in Java.
Luis Mastrangelo, Luca Ponzanelli, Andrea Mocci, Michele Lanza 0001, Matthias Hauswirth, Nathaniel Nystrom
OOPSLA3
2015 Blended, not stirred: Multi-concern visualization of large software systems
abstract
While constructing and evolving software systems, developers generate directly and indirectly a large amount of data of diverse nature, such as source code changes, bug tracking information, IDE interactions, stack traces, etc. Often these diverse data sources are processed and visualized in isolation, leading to a partial view of systems. We present a blended approach to visualize several data “ingredients” at once, to give as complete an answer as possible to the question “What happened to the system in the last few days?”. The goal is to enable a quick and comprehensive assessment of what happened to a software system in any given time frame.
Tommaso Dal Sasso, Roberto Minelli, Andrea Mocci, Michele Lanza 0001
VISSOFT3
2015 CEL: Touching software modeling in essence
abstract
Understanding a problem domain is a fundamental prerequisite for good software design. In object-oriented systems design, modeling is the fundamental first phase that focuses on identifying core concepts and their relations. How to properly support modeling is still an open problem, and existing approaches and tools can be very different in nature. On the one hand, lightweight ones, such as pen & paper/whiteboard or CRC cards, are informal and support well the creative aspects of modeling, but produce artifacts that are difficult to store, process and reuse as documentation. On the other hand, more constrained and semi-formal ones, like UML, produce storable and processable structured artifacts with defined semantics, but this comes at the expense of creativity. We believe there exists a middle ground to investigate that maximizes the good of both worlds, that is, by supporting software modeling closer to its essence, with minimal constraints on the developer's creativity and still producing reusable structured artifacts. We also claim that modeling can be best treated by using the emerging technology of touch-based tablets. We present a novel gesture-based modeling approach based on a minimal set of constructs, and CEL, an iPad application, for rapidly creating, manipulating, and storing language agnostic object-oriented software models, which can be exported as skeleton source code in any language of choice. We assess our approach through a controlled qualitative study.
Remo Lemma, Michele Lanza 0001, Andrea Mocci
SANER3
2015 Misery loves company: CrowdStacking traces to aid problem detection
abstract
During software development, exceptions are by no means exceptional: Programmers repeatedly try and test their code to ensure that it works as expected. While doing so, runtime exceptions are raised, pointing out various issues, such as inappropriate usage of an API, convoluted code, as well as defects. Such failures result in stack traces, lists composed of the sequence of method invocations that led to the interruption of the program. Stack traces are useful to debug source code, and if shared also enhance the quality of bug reports. However, they are handled manually and individually, while we argue that they can be leveraged automatically and collectively to enable what we call crowdstacking, the automated collection of stack traces on the scale of a whole development community. We present our crowdstacking approach, supported by Shore-Line Reporter, a tool which seamlessly collects stack traces during program development and execution and stores them on a central repository. We illustrate how thousands of stack traces stemming from the IDEs of several developers can be leveraged to identify common hot spots in the code that are involved in failures, using this knowledge to retrieve relevant and related bug reports and to provide an effective, instant context of the problem to the developer.
Tommaso Dal Sasso, Andrea Mocci, Michele Lanza 0001
SANER2
2015 Code review: Veni, ViDI, vici
abstract
Modern software development sees code review as a crucial part of the process, because not only does it facilitate the sharing of knowledge about the system at hand, but it may also lead to the early detection of defects, ultimately improving the quality of the produced software. Although supported by numerous approaches and tools, code review is still in its infancy, and indeed researchers have pointed out a number of shortcomings in the state of the art. We present a critical analysis of the state of the art of code review tools and techniques, extracting a set of desired features that code review tools should possess. We then present our vision and initial implementation of a novel code review approach named Visual Design Inspection (ViDI), illustrated through a set of usage scenarios. ViDI is based on a combination of visualization techniques, design heuristics, and static code analysis techniques.
Yuriy Tymchuk, Andrea Mocci, Michele Lanza 0001
SANER2
2014 Visual Storytelling of Development Sessions
abstract
Most development activities, like program understanding, source code navigation and editing, are supported by Integrated Development Environments (IDEs). They provide different tools and user interfaces (UI) to interact with the source code, such as browsers, debuggers, and inspectors. It is uncertain how and when programmers use different UI elements of an IDE and to what extent they appropriately support development. Previously we developed DFLOW, a tool that seamlessly records and processes interaction data. Our long-term goal is to assess to what extent the UIs of IDEs support the workflow of developers and whether they can be improved. As a first step we present our approach to analyze development sessions in the form of visual storytelling. We illustrate our initial catalogue of visualizations through two development stories.
Roberto Minelli, Lorenzo Baracchi, Andrea Mocci, Michele Lanza 0001
ICSME3
2014 Improving Low Quality Stack Overflow Post Detection
abstract
Stack Overflow is a popular questions and answers (Q&A) website among software developers. It counts more than two millions of users who actively contribute by asking and answering thousands of questions daily. Identifying and reviewing low quality posts preserves the quality of site's contents and it is crucial to maintain a good user experience. In Stack Overflow the identification of poor quality posts is performed by selected users manually. The system also uses an automated identification system based on textual features. Low quality posts automatically enter a review queue maintained by experienced users. We present an approach to improve the automated system in use at Stack Overflow. It analyzes both the content of a post (e.g., simple textual features and complex readability metrics) and community-related aspects (e.g., popularity of a user in the community). Our approach reduces the size of the review queue effectively and removes misclassified good quality posts.
Luca Ponzanelli, Andrea Mocci, Alberto Bacchelli, Michele Lanza 0001, David Fullerton
ICSME2
2014 Enhanced modeling of DC-DC power converters by means of averaging technique
abstract
An average modeling approach for DC-DC power converters is presented in this paper. It allows the achievement of accurate averaged models of DC-DC converters, which account for dynamic and steady state operations, over both Continuous Conduction Mode (CCM) and Discontinuous Conduction Mode (DCM). In particular, appropriate equivalent switching signals are introduced in order to account for each converter operating state. In addition, a suitable inductor model is introduced in order to improve inductor losses estimation. As a result, the proposed averaged modeling enables an enhanced power losses estimation by accounting for switching and current ripple phenomena. The worth and effectiveness of the proposed modeling approach has been validated through a simulation study, which is performed in the Matlab-Simulink environment and refers to the case of a boost DC-DC converter.
Andrea Mocci, Alessandro Serpi, Ignazio Marongiu, Gianluca Gatto
IECON1
2014 Mining unit tests for code recommendation
abstract
Developers spend a significant portion of their time understanding and learning the correct usage of the APIs of libraries they want to integrate in their projects. However, learning how to effectively use APIs is complex and time consuming. Code recommendation systems play a crucial role facilitating developers in this task by providing to them relevant examples while they code. This paper proposes a novel approach to code recommendation in which code examples are automatically obtained by mining and manipulating unit tests. In this paper we discuss the theoretical and practical implications that underpin this idea. The discussion leads to a series of fascinating research challenges that we organized in a research agenda.
Mohammad Ghafari, Carlo Ghezzi, Andrea Mocci, Giordano Tamburrelli
ICPC3
2014 Collaboration in open-source projects: myth or reality?
abstract
One of the fundamental principles of open-source projects is that they foster collaboration among developers, disregarding their geographical location or personal background. When it comes to software repositories collaboration is a rather ephemeral phenomenon which lacks a clear definition, and it must therefore be mined and modeled. This throws up the question whether what is mined actually maps to reality.
Yuriy Tymchuk, Andrea Mocci, Michele Lanza 0001
MSR2
2014 Visualizing Developer Interactions
abstract
Integrated Development Environments (IDEs) have become the de facto standard vehicle to develop software systems. The user interface (UI) of an IDE offers a staggering amount of facilities to manipulate source code, such as inspectors, debuggers, recommenders, alternative viewers, etc. It is unclear how developers use the UI of an IDE and whether such UIs actually give appropriate support to the developers. We present a visual approach to understand and characterize development sessions from the UI perspective. The tool supporting our approach mines and processes the finest-grained UI-level events making up development sessions and presents them visually. We have collected, visualized, and analyzed hundreds of development sessions and report on our findings.
Roberto Minelli, Andrea Mocci, Michele Lanza 0001, Lorenzo Baracchi
VISSOFT2
2013 A novel continuous-time equivalent circuit for boost DC-DC converters
abstract
A novel continuous-time equivalent circuit suitable for boost DC-DC converters is presented in this paper. It has been developed on the basis of the averaging technique with the aim of achieving a good ripple-free representation of the state variables of the system, whatever the converter operating mode is, i.e. Continuous Conduction Mode (CCM) or Discontinuous Conduction Mode (DCM). This goal can be achieved on condition that an appropriate PWM pattern is introduced, which enables an almost perfect matching between state variables and their corresponding equivalent ones at the start of each sampling time interval. Apart from the inductor and the capacitor, the proposed circuit consists of some equivalent input voltage and output current sources, together with several constant and/or variable resistors, which depend on converter switching signals and circuital parameters. The proposed continuous-time equivalent circuit has been validated through a simulation study, which is performed by means of the software PLECS. Simulation results highlight the worth and the effectiveness of the proposed modelling approach, over both transient and steady state operations.
Gianluca Gatto, Ignazio Marongiu, Andrea Mocci, Alessandro Serpi, Ivan Luigi Spano
IECON3
2013 An improved averaged model for boost DC-DC converters
abstract
An improved averaged model suitable for boost DC-DC converters is presented in this paper. It is developed on the basis of the averaging technique with the aim of appropriately taking into account the effects of switching phenomena. Thus, two appropriate trapezoidal signals have been introduced in order to model the transient evolutions of voltages and currents of the DC-DC converter devices over both turn-on and turn-off of the switch. This enables the synthesis of an appropriate averaged model of boost DC-DC converters, which can be usefully employed in determining average voltage and current evolutions, as well as steady state average powers. The worth and effectiveness of the proposed modelling approach has been validated through a simulation study, which refers to the case of a traditional boost DC-DC converter.
Gianluca Gatto, Ignazio Marongiu, Andrea Mocci, Alessandro Serpi, Ivan Luigi Spano
IECON3
2012 Behavioral validation of JFSL specifications through model synthesis
abstract
Contracts are a popular declarative specification technique to describe the behavior of stateful components in terms of pre/post conditions and invariants. Since each operation is specified separately in terms of an abstract implementation, it may be hard to understand and validate the resulting component behavior from contracts in terms of method interactions. In particular, properties expressed through algebraic axioms, which specify the effect of sequences of operations, require complex theorem proving techniques to be validated. In this paper, we propose an automatic small-scope based approach to synthesize incomplete behavioral abstractions for contracts expressed in the JFSL notation. The proposed abstraction technique enables the possibility to check that the contract behavior is coherent with behavioral properties expressed as axioms of an algebraic specifications. We assess the applicability of our approach by showing how the synthesis methodology can be applied to some classes of contract-based artifacts like specifications of data abstractions and requirement engineering models.
Carlo Ghezzi, Andrea Mocci
ICSE2
2012 Runtime monitoring of component changes with Spy@Runtime
abstract
We present Spy@Runtime, a tool to infer and work with behavior models. Spy@Runtime generates models through a dynamic black box approach and is able to keep them updated with observations coming from actual system execution. We also show how to use models describing the protocol of interaction of a software component to detect and report functional changes as soon as they are discovered. Monitoring functional properties is particularly useful in an open environment in which there is a distributed ownership of a software system. Parts of the system may be changed independently and therefore it becomes necessary to monitor the component's behavior at run time.
Carlo Ghezzi, Andrea Mocci, Mario Sangiorgio
ICSE2
2011 Dynamic synthesis of program invariants using genetic programming
abstract
Symbolic program manipulation plays a key role in program comprehension and verification. Logic formulae are used to represent the program' s state and transformation rules describe the effect of statement executions on the program's state. A well-known problem arises in the case of loops, since the number of iterations is generally unknown. The effect of a loop is therefore abstracted into a loop invariant, whose derivation cannot in general be automated and requires human ingenuity. In this paper, we present a preliminary approach that in tegrates genetic programming into the synthesis of invariant formula that describes the behavior of a loop. We present a specific representation of formulae that works well with loops manipulating arrays. The technique has been validated with a set of relevant examples with increasing complexity. The preliminary results are promising and show the feasibility of our approach.
Luigi Cardamone, Andrea Mocci, Carlo Ghezzi
IEEE Congress on Evolutionary Computation2
2011 Extracting structured data from natural language documents with island parsing
abstract
The design and evolution of a software system leave traces in various kinds of artifacts. In software, produced by humans for humans, many artifacts are written in natural language by people involved in the project. Such entities contain structured information which constitute a valuable source of knowledge for analyzing and comprehending a system's design and evolution. However, the ambiguous and informal nature of narrative is a serious challenge in gathering such information, which is scattered throughout natural language text. We present an approach-based on island parsing-to recognize and enable the parsing of structured information that occur in natural language artifacts. We evaluate our approach by applying it to mailing lists pertaining to three software systems. We show that this approach allows us to extract structured data from emails with high precision and recall.
Alberto Bacchelli, Anthony Cleve, Michele Lanza 0001, Andrea Mocci
ASE4
2010 Automatic Cross Validation of Multiple Specifications: A Case Study
Carlo Ghezzi, Andrea Mocci, Guido Salvaneschi
FASE2
2009 Synthesizing intensional behavior models by graph transformation
abstract
This paper describes an approach (SPY) to recovering the specification of a software component from the observation of its run-time behavior. It focuses on components that behave as data abstractions. Components are assumed to be black boxes that do not allow any implementation inspection. The inferred description may help understand what the component does when no formal specification is available. SPY works in two main stages. First, it builds a deterministic finite-state machine that models the partial behavior of instances of the data abstraction. This is then generalized via graph transformation rules. The rules can generate a possibly infinite number of behavior models, which generalize the description of the data abstraction under an assumption of ldquoregularityrdquo with respect to the observed behavior. The rules can be viewed as a likely specification of the data abstraction. We illustrate how SPY works on relevant examples and we compare it with competing methods.
Carlo Ghezzi, Andrea Mocci, Mattia Monga
ICSE2