Sergio Di Martino

dblp:29/2259 · DBLP profile ↗
← Back
74ranked-venue papers
21as first author
30since 2021 · last 2026
0000-0002-1019-9004ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 40 · 9 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Investigating the adoption and maintenance of web GUI testing: Insights from GitHub repositories
abstract
Web GUI testing is a quality assessment practice aimed at evaluating the functionality of web applications from the perspective of its end users. While prior studies have explored the technical challenges of automated Web GUI testing, fewer works have explored how this practice is applied in real-world web apps. This study aims to investigate the adoption, characteristics, and maintenance of automated web GUI testing practices in open-source web applications, focusing on identifying trends and providing actionable insights for researchers and practitioners. We conducted a large-scale empirical analysis of 472 web applications on the GitHub platform, developed in Java , JavaScript , Python , and TypeScript . These projects use popular browser automation frameworks like Selenium , Playwright , Cypress , and Puppeteer . The study involved examining project characteristics and analyzing the co-evolution and maintenance of automated web GUI tests over time. Our findings empirically document automated web GUI testing adoption patterns in open-source projects, providing insights into the practical drivers behind both initial framework adoption and migration between different testing frameworks. Projects incorporating these tests generally show higher community engagement and consistent maintenance efforts. The analysis reveals that Web GUI tests tend to co-evolve with the underlying applications, reflecting their integration into the development lifecycle. The study provides valuable insights into the prevalence and maintenance of Web GUI testing, highlighting practical implications for improving testing practices. Our findings can guide further research on the matter and support practitioners in enhancing their testing strategies.
Sergio Di Meglio, Luigi L. L. Starace, Valeria Pontillo, Ruben Opdebeeck, Coen De Roover, Sergio Di Martino
Inf. Softw. Technol.6
2026 Web app performance testing in industrial contexts: Supporting workload generation with E2E-Loader++
abstract
Performance testing is essential for ensuring that web applications remain responsive and reliable under varying workloads. Defining meaningful workloads remains a key challenge in performance testing, with existing solutions requiring deployed systems to collect real user behaviors, offering limited support for automated data dependency management, and lacking support for emerging protocols like WebSocket. Our previous work introduced E2E-Loader, a novel approach that supports the definition of performance testing workloads by exploiting existing End-to-End functional test cases, enabling workload creation even before system deployment. While initial results demonstrated technical feasibility, the practical utility and industrial readiness of automated workload generation remained unvalidated in real-world software engineering contexts. In this paper, we present E2E-Load-er++, an enhanced version of our original tool, with significant improvements to dependency detection capabilities, and provide the first rigorous industrial evaluation of our semi-automated performance testing workload generation approach. The comprehensive empirical assessment we conducted involved controlled experiments with professional software engineers working on proprietary industrial applications, aimed at measuring both technical effectiveness and practical impact of E2E-Loader++ on performance testing workflows. Results demonstrate substantial productivity gains: compared to manual approaches, E2E-Loader++ achieved a 62% reduction rate in workload creation time and a 37% reduction rate in the number of required user interactions per minute, while preserving comparable workload quality. Moreover, the tool received strong usability ratings and practitioner acceptance, confirming its potential for real-world adoption in industrial performance testing workflows.
Sergio Di Meglio, Luigi L. L. Starace, Sergio Di Martino
J. Syst. Softw.3
2026 Semi-automated generation of web app performance tests from end-to-end GUI-level tests with E2E-Loader
Sergio Di Meglio, Luigi L. L. Starace, Sergio Di Martino
Sci. Comput. Program.3
2025 Rookie Mistakes: Measuring Software Quality in Student Projects to Guide Educational Enhancement
Sergio Di Martino, Sergio Di Meglio, Anna Rita Fasolino, Luigi L. L. Starace, Porfirio Tramontana
SEAA (3)2
2025 Performance Testing in Open-Source Web Projects: Adoption, Maintenance, and a Change Taxonomy
abstract
Performance testing is crucial to ensuring that web applications meet user expectations under varying workloads. Activities such as stress, load, and smoke testing are designed to simulate different kinds of simultaneous user interactions and assess system behavior. Despite its recognized importance in quality assurance of large-scale web-based systems, witnessed by numerous studies proposing solutions to support these activities, the real-world adoption and evolutionary dynamics of performance tests have received limited attention in the literature. To fill this gap, we analyzed 77 open-source web projects using Apache JMETER and LOCUST. Our study investigates how performance tasks are performed (adoption time, load design, types of tasks), the characteristics of projects that adopt them, and their longterm maintenance. Our findings reveal that performance tests in open-source projects are simple, with a focus on singleuser behaviors and minimal requests, and most tests have low concurrency. Load tests are the most common, followed by smoke and stress tests. Projects with performance tests tend to be larger and more actively maintained. However, tests are mostly long-lived but rarely updated, suggesting potential risks to their relevance and coverage over time. Finally, by creating a taxonomy of performance test changes, we observe recurring patterns of modifications, including workload adjustments, network request changes, and updates to system monitoring.
Sergio Di Meglio, Luigi L. L. Starace, Valeria Pontillo, Ruben Opdebeeck, Coen De Roover, Sergio Di Martino
ICSME6
2025 E2E-Loader: A Tool to Generate Performance Tests from End-to-End GUI-Level Tests
abstract
Performance testing is essential for ensuring that web applications deliver a satisfactory user experience under varying workloads. Crafting meaningful workloads is a key challenge, addressed in previous research by analyzing system logs that reflect real user behaviors. However, these approaches face limitations: they require the system under test to be deployed to collect usage data, offer limited automation for managing data dependencies, and often lack support for modern protocols like Websocket. We present E2E-LOADER, a tool for automating the generation of performance testing workloads for Web Applications. E2E-LOADER leverages existing End-to-End (E2E) GUI-level test cases to create workloads, allowing its use at early stages of development before user data is available. The tool fully supports HTTP and WEBSOCKET-based interactions and includes customizable heuristics to detect data dependencies automatically. E2E-LOADER has been evaluated in previous research in an industrial case study, demonstrating that it produces workloads comparable in quality to those manually designed by practitioners, with significantly less effort and time. The tool and its source code are openly available to support researchers and practitioners in advancing performance testing practices. A screencast showcasing E2E-LOADER in function is available at https://youtu.be/pDWNlllkAhU.
Sergio Di Meglio, Luigi L. L. Starace, Sergio Di Martino
ICST3
2025 E2EGit: A Dataset of End-to-End Web Tests in Open Source Projects
abstract
End-to-end (E2E) testing is a software validation approach that simulates realistic user scenarios throughout the entire workflow of an application. In the context of web applications, E2E testing involves two activities: Graphic User Interface (GUI) testing, which simulates user interactions with the web app’s GUI through web browsers, and performance testing, which evaluates system workload handling. Despite its recognized importance in delivering high-quality web applications, the availability of large-scale datasets featuring real-world E2E web tests remains limited, hindering research in the field.To address this gap, we present E2EGit, a comprehensive dataset of non-trivial open-source web projects collected on GitHub that adopt E2E testing. By analyzing over 5,000 web repositories across popular programming languages (Java, JavaScript, TypeScript and Python), we identified 472 repositories implementing 43,670 automated Web GUI tests with popular browser automation frameworks (Selenium, Playwright, Cypress, Puppeteer), and 84 repositories that featured 271 automated performance tests implemented leveraging the most popular open-source tools (JMeter, LoCust). Among these, 13 repositories implemented both types of testing for a total of 786 Web GUI tests and 61 performance tests. The dataset is available on Zenodo (DOI: 10.5281/zenodo.14234731).
Sergio Di Meglio, Luigi L. L. Starace, Valeria Pontillo, Ruben Opdebeeck, Coen De Roover, Sergio Di Martino
MSR6
2025 Large Language Models in the Travel Domain: An Industrial Experience
abstract
Online property booking platforms are widely used and rely heavily on consistent, up-to-date information about accommodation facilities, often sourced from third-party providers.However, these external data sources are frequently affected by incomplete or inconsistent details, which can frustrate users and result in a loss of market.In response to these challenges, we present an industrial case study involving the integration of Large Language Models (LLMs) into CALEIDOHOTELS, a property reservation platform developed by FERVENTO.We evaluate two well-know LLMs in this context: Mistral 7B, fine-tuned with QLoRA, and Mixtral 8x7B, utilized with a refined system prompt.Both models were assessed based on their ability to generate consistent and homogeneous descriptions while minimizing hallucinations.Mixtral 8x7B outperformed Mistral 7B in terms of completeness (99.6% vs. 93%), precision (98.8% vs. 96%), and hallucination rate (1.2% vs. 4%), producing shorter yet more concise content (249 vs. 277 words on average).However, this came at a significantly higher computational cost: 50GB VRAM and $1.61/hour versus 5GB and $0.16/hour for Mistral 7B.Our findings provide practical insights into the trade-offs between model quality and resource efficiency, offering guidance for deploying LLMs in production environments and demonstrating their effectiveness in enhancing the consistency and reliability of accommodation data.
Sergio Di Meglio, Aniello Somma, Luigi L. L. Starace, Fabio Scippacercola, Giancarlo Sperlì, Sergio Di Martino
SEKE6
2025 Fedflow: a personalized federated learning framework for passenger flow prediction
abstract
Abstract In the Intelligent Public Transportation Systems (IPTS) domain, predicting the number of commuters on-board, entering or leaving a metro train or a bus, i.e. the Passenger Flow (PF), is crucial for optimizing resource allocation and enhancing commuter satisfaction. In urban scenarios, the public transport system is often managed by distinct competing mobility providers. Traditional centralized machine learning models for PF prediction usually require data sharing among such competitors, leading to privacy and economic concerns. To overcome these issues, we propose exploiting Federated Learning (FL) in the PF predictions problem, as only model parameters must be shared among entities. Still, a straightforward application of FL can have some pitfalls. On one hand, it is widely recognized that FL can struggle with data heterogeneity, which is likely in the case of data acquired by distinct companies managing different public mobility services. Moreover, spatio-temporal features are not explicitly handled by classical FL. In this paper, we propose FedFlow: a personalized federated learning framework tailored for PF prediction. The proposed framework encompasses a personalized mechanism meant to refine local models based on client similarities, calculated by only leveraging publicly available domain-dependent information. The proposed framework has been experimentally validated on mobility data collected in a major Italian city, comparing FL predictions obtained by FedFlow against those obtained by LSTM models trained on local data, centralized data, FedAvg, and PerFedAvg. Results show that FedFlow outperforms all the considered adversary techniques. This work demonstrates that our proposal of personalized FL is effective in predicting PF while ensuring data privacy.
Franca Rocco di Torrepadula, Marco Fisichella, Sergio Di Martino, Nicola Mazzocca
Mach. Learn.3
2025 LLM-Based Automation of COSMIC Functional Size Measurement From Use Cases
abstract
COmmon Software Measurement International Consortium (COSMIC) Functional Size Measurement is a method widely used in the software industry to quantify user functionality and measure software size, which is crucial for estimating development effort, cost, and resource allocation. COSMIC measurement is a manual task that requires qualified professionals and effort. To support professionals in COSMIC measurement, we propose an automatic approach, CosMet, that leverages Large Language Models to measure software size starting from use cases specified in natural language. To evaluate the proposed approach, we developed a web tool that implements CosMet using GPT-4 and conducted two studies to assess the approach quantitatively and qualitatively. Initially, we experimented with CosMet on seven software systems, encompassing 123 use cases, and compared the generated results with the ground truth created by two certified professionals. Then, seven professional measurers evaluated the analysis achieved by CosMet and the extent to which the approach reduces the measurement time. The first study's results revealed that CosMet is highly effective in analyzing and measuring use cases. The second study highlighted that CosMet offers a transparent and interpretable analysis, allowing practitioners to understand how the measurement is derived and make necessary adjustments. Additionally, it reduces the manual measurement time by 60-80%.
Gabriele De Vito, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Fabio Palomba
IEEE Trans. Software Eng.2
2024 Can Large Language Models Automatically Generate GIS Reports?
Luigi L. L. Starace, Sergio Di Martino
W2GIS2
2024 Machine Learning for public transportation demand prediction: A Systematic Literature Review
Franca Rocco di Torrepadula, Enea Vincenzo Napolitano, Sergio Di Martino, Nicola Mazzocca
Eng. Appl. Artif. Intell.3
2024 A visual-based toolkit to support mobility data analytics
abstract
The Knowledge Discovery from Data (KDD) process is widely used across various domains to get valuable insights from data. Many platforms, like KNIME or RapidMiner, offer effective tools for KDD analysts, allowing them to perform data analytics tasks in a visual fashion, without writing code. In recent years, the increasing availability of mobility data has led to a surge in KDD-based initiatives from both industry and academia in the Intelligent Transportation Systems (ITS) domain. Still, KDD platforms lack comprehensive support for some typical mobility data manipulation tasks. As a result, mobility data analysis still requires a significant coding phase, with reduced productivity and hindered replicability of results. To address this gap, this paper presents a novel solution aimed at supporting ITS data analysts in defining KDD processes more efficiently. More in detail, we extended the KNIME platform by introducing a collection of new components explicitly tailored to facilitate some peculiar KDD tasks from mobility data. These components encompass critical functionalities such as map coverage analysis, trajectory partitioning and map-matching. To showcase the effectiveness of the proposed solution, we used it to replicate a study published in the ITS data analytics domain. Thanks to our proposal, such replication can be accomplished in a few minutes and with just a few clicks, without any manual coding, resulting in a pipeline that is easier to understand, distribute and re-execute, also for domain experts with no programming experience. Our solution is open-source and freely downloadable from the Knime Hub. In this way, we aim to foster data-driven research and practice in the ITS field, by providing researchers and practitioners with more effective analytics tools to handle mobility data.
Sergio Di Martino, Enrico Landolfi, Nicola Mazzocca, Franca Rocco di Torrepadula, Luigi L. L. Starace
Expert Syst. Appl.1
2024 How to manage massive spatiotemporal dataset from stationary and non-stationary sensors in commercial DBMS?
abstract
Abstract The growing diffusion of the latest information and communication technologies in different contexts allowed the constitution of enormous sensing networks that form the underlying texture of smart environments. The amount and the speed at which these environments produce and consume data are starting to challenge current spatial data management technologies. In this work, we report on our experience handling real-world spatiotemporal datasets: a stationary dataset referring to the parking monitoring system and a non-stationary dataset referring to a train-mounted railway monitoring system. In particular, we present the results of an empirical comparison of the retrieval performances achieved by three different off-the-shelf settings to manage spatiotemporal data, namely the well-established combination of PostgreSQL + PostGIS with standard indexing, a clustered version of the same setup, and then a combination of the basic setup with Timescale, a storage extension specialized in handling temporal data. Since the non-stationary dataset has put much pressure on the configurations above, we furtherly investigated the advantages achievable by combining the TSMS setup with state-of-the-art indexing techniques. Results showed that the standard indexing is by far outperformed by the other solutions, which have different trade-offs. This experience may help researchers and practitioners facing similar problems managing these types of data.
Vincenzo Norman Vitale, Sergio Di Martino, Adriano Peron, Massimiliano Russo, Ermanno Battista
Knowl. Inf. Syst.2
2024 Regression test prioritization leveraging source code similarity with tree kernels
abstract
Abstract Regression test prioritization (RTP) is an active research field, aiming at re‐ordering the tests in a test suite to maximize the rate at which faults are detected. A number of RTP strategies have been proposed, leveraging different factors to reorder tests. Some techniques include an analysis of changed source code, to assign higher priority to tests stressing modified parts of the codebase. Still, most of these change‐based solutions focus on simple text‐level comparisons among versions. We believe that measuring source code changes in a more refined way, capable of discriminating between mere textual changes (e.g., renaming of a local variable) and more structural changes (e.g., changes in the control flow), could lead to significant benefits in RTP, under the assumption that major structural changes are also more likely to introduce faults. To this end, we propose two novel RTP techniques that leverage tree kernels (TK), a class of similarity functions largely used in Natural Language Processing on tree‐structured data. In particular, we apply TKs to abstract syntax trees of source code, to more precisely quantify the extent of structural changes in the source code, and prioritize tests accordingly. We assessed the effectiveness of the proposals by conducting an empirical study on five real‐world Java projects, also used in a number of RTP‐related papers. We automatically generated, for each considered pair of software versions (i.e., old version, new version) in the evolution of the involved projects, 100 variations with artificially injected faults, leading to over 5k different software evolution scenarios overall. We compared the proposed prioritization approaches against well‐known prioritization techniques, evaluating both their effectiveness and their execution times. Our findings show that leveraging more refined code change analysis techniques to quantify the extent of changes in source code can lead to relevant improvements in prioritization effectiveness, while typically introducing negligible overheads due to their execution.
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
J. Softw. Evol. Process.3
2024 GUI testing of Android applications: Investigating the impact of the number of testers on different exploratory testing strategies
abstract
Abstract Graphical user interface (GUI) testing plays a pivotal role in ensuring the quality and functionality of mobile apps. In this context, exploratory testing (ET), a distinctive methodology in which individual testers pursue a creative, and experience‐based approach to test design, is often used as an alternative or in addition to traditional scripted testing. Managing the exploratory testing process is a challenging task that can easily result either in wasteful spending or in inadequate software quality, due to the relative unpredictability of exploratory testing activities, which depend on the skills and abilities of individual testers. A number of works have investigated the diversity of testers' performance when using ET strategies, often in a crowdtesting setting. These works, however, investigated ET effectiveness in detecting bugs, and not in scenarios in which the goal is to generate a re‐executable test suite, as well. Moreover, less work has been conducted on evaluating the impact of adopting different exploratory testing strategies. As a first step toward filling this gap in the literature, in this work, we conduct an empirical evaluation involving four open‐source Android apps and 20 masters students that we believe can be representative of practitioners partaking in exploratory testing activities. The students were asked to generate test suites for the apps using a capture and replay tool and different exploratory testing strategies. We then compare the effectiveness, in terms of aggregate code coverage that different‐sized groups of students using different exploratory testing strategies may achieve. Results provide deeper insights into code coverage dynamics to project managers interested in using exploratory approaches to test simple Android apps, on which they can make more informed decisions.
Sergio Di Martino, Anna Rita Fasolino, Luigi L. L. Starace, Porfirio Tramontana
J. Softw. Evol. Process.1
2023 ECHO: An Approach to Enhance Use Case Quality Exploiting Large Language Models
abstract
UML use cases are commonly used in software engineering to specify the functional requirements of a system since they are an effective tool for interacting with stakeholders thanks to the use of natural languages. However, producing high-quality use cases can be challenging due to the lack of precise guidelines and suitable tools. This can lead to problems, e.g. inaccuracy and incompleteness, in the derived software artifacts and the final product. Recent advancements in Natural Language Processing and Large Language Models (LLMs) can provide the premises for developing tools supporting activities based on natural languages. In this paper, we propose ECHO, a novel approach for supporting software engineers in enhancing the quality of UML use cases using LLMs. Our approach consists of a co-prompt engineering approach and an iterative and interactive process with the LLM to improve the quality of use cases, based on practitioners’ feedback. To prove the feasibility of the proposal, we instantiated the approach using ChatGPT and performed a controlled experiment to assess its effectiveness by involving seven software engineering professionals. Three were part of the experimental group and used ECHO to improve the quality of the use cases. Three others were the control group and enhanced the quality of use cases manually. Finally, the last participant acted as an oracle, blind w.r.t. the groups, and evaluated the quality of the enhanced use cases, both qualitatively by means of a questionnaire, and quantitatively, by means of the Use Case Points metric. Results show that ECHO can effectively support software engineers to improve use cases’ quality thanks to the prompts suitably designed to interact with ChatGPT.
Gabriele De Vito, Fabio Palomba, Carmine Gravino, Sergio Di Martino, Filomena Ferrucci
SEAA4
2023 E2E-Loader: A Framework to Support Performance Testing of Web Applications
abstract
Performance testing is crucial to assess that Web Applications provide a good user experience under different workloads. A workload reproduces the interactions of a number of concurrent users with the system, to observe its actual behavior under stress.Defining meaningful workloads is a key challenge in performance testing, and many solutions have been proposed in the literature to support testers in this task, mostly based on analyzing system logs describing real user behaviors. However, in our industrial and academic experience, we found that these solutions present some limitations, hindering performance testers’ applicability and productivity. In particular, (I) they require the system under test to be actually deployed in order to collect real user behaviors; (II) they offer limited support to automated management of data dependencies; (III) they lack support for emerging protocols, such as WebSocket.In this paper, we present E2E-Loader, a novel approach to automate the design of performance testing workloads for web applications. E2E-Loader generates workloads by exploiting existing End-to-End functional test cases and can be used at an early stage, before the system is deployed and actual user behaviors have been collected. Our solution features full WebSocket support and includes a customizable heuristic to automatically detect data dependencies.We empirically evaluate the proposed approach in an industrial case study. Results are promising and show that the workloads generated with E2E-Loader are generally comparable to those that were manually created by practitioners working with our industrial partner while requiring a fraction of the time to be obtained. Finally, we make E2E-Loader and its source code publicly available for interested practitioners and researchers.
Ermanno Battista, Sergio Di Martino, Sergio Di Meglio, Fabio Scippacercola, Luigi L. L. Starace
ICST2
2023 AI-based Fault-proneness Metrics for Source Code Changes
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
IWSM-Mensura3
2023 Starting a New REST API Project? A Performance Benchmark of Frameworks and Execution Environments
Sergio Di Meglio, Luigi L. L. Starace, Sergio Di Martino
IWSM-Mensura3
2023 Bus Journey Time Prediction with Machine Learning: An Empirical Experience in Two Cities
abstract
With increasing urbanisation, and a growing population, transport within cities has never been more important. Buses are the most widespread form of transport worldwide, often being cheaper and more flexible than rail, but also less reliable. Long term bus journey time predictions are important for advanced journey planning and scheduling of bus services. For this reason, several machine/deep learning techniques have been defined to predict bus journey time. Still, due to the number of involved factors, such as complexity and noise in bus data, road network topology, etc., accurate predictions remain elusive. In this paper we aim at validating some Machine Learning methods recently shown to be effective in the literature, on new bus datasets from Dublin and Genoa. The analysis of the results shows some interesting insights into bus networks, highlighting that the accuracy of the predictions is strongly related to the standard deviation of the whole journey times. It emerges that some bus routes show consistency in the prediction error across methods, and for these routes it makes sense to use methods that are fast and computationally efficient, as there is no benefit to applying more complex algorithms. We use features of the route data distribution to develop an explanatory model for the consistency of the route across methods, with a coefficient of determination ( $$R^2$$ ) of 0.94. Finally, we identify a systematic anomaly in the data in Dublin that alters the performance of the methods.
Laura Dunne, Franca Rocco di Torrepadula, Sergio Di Martino, Gavin McArdle, Davide Nardone
W2GIS3
2023 Mobility Data Analytics with KNOT: The KNime mObility Toolkit
Sergio Di Martino, Nicola Mazzocca, Franca Rocco di Torrepadula, Luigi L. L. Starace
W2GIS1
2022 Change-Aware Regression Test Prioritization using Genetic Algorithms
abstract
Regression testing is a practice aimed at providing confidence that, within software maintenance, the changes in the code base have introduced no faults in previously validated functionalities. With the software industry shifting towards iterative and incremental development with shorter release cycles, the straightforward approach of re-executing the entire test suite on each new version of the software is often unfeasible due to time and resource constraints. In such scenarios, Test Case Prioritization (TCP) strategies aim at providing an effective ordering of the test suite, so that the tests that are more likely to expose faults are executed earlier and fault detection is maximised even when test execution needs to be abruptly terminated due to external constraints. In this work, we propose Genetic-Diff, a TCP strategy based on a genetic algorithm featuring a specifically-designed crossover operator and a novel objective function that combines code coverage metrics with an analysis of changes in the code base. We empirically evaluate the proposed algorithm on several releases of three heterogeneous real-world, open source Java projects, in which we artificially injected faults, and compare the results with other state-of-the-art TCP techniques using fault-detection rate metrics. Findings show that the proposed technique performs generally better than the baselines, especially when there is a limited amount of code changes, which is a common scenario in modern development practices.
Francesco Altiero, Giovanni Colella, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
SEAA4
2022 ReCover: a Curated Dataset for Regression Testing Research
abstract
It is recognized in the literature that finding representative data to conduct regression testing research is non-trivial. In our experience within this field, existing datasets are often affected by issues that limit their applicability. Indeed, these datasets often lack fine-grained coverage information, reference software repositories that are not available anymore, or do not allow researchers to readily build and run the software projects, e.g., to obtain additional information. As a step towards better replicability and data-availability in regression testing research, we introduce ReCover, a dataset of 114 pairs of subsequent versions from 28 open source Java projects from GitHub. In particular, ReCover is intended as a consolidation and enrichment of recent dedicated regression testing datasets proposed in the literature, to overcome some of the above described issues, and to make them ready to use with a broader number of regression testing techniques. To this end, we developed a custom mining tool, that we make available as well, to automatically process two recent, massive regression testing datasets, retaining pairs of software versions for which we were able to (1) retrieve the full source code; (2) build the software in a general-purpose Java/Maven environment (which we provide as a Docker container for ease of replication); and (3) compute fine-grained test coverage metrics. ReCover can be readily employed in regression testing studies, as it bundles in a single package full, buildable source code and detailed coverage reports for all the projects. We envision that its use could foster regression testing research, improving replicability and long-term data availability.
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
MSR3
2022 Bus Passenger Load Prediction: Challenges from an Industrial Experience
Flora Amato, Sergio Di Martino, Nicola Mazzocca, Davide Nardone, Franca Rocco di Torrepadula, Paolo Sannino
W2GIS2
2022 On the Impact of Location-related Terms in Neural Embeddings for Content Similarity Measures in Cultural Heritage Recommender Systems
Antonio Origlia, Sergio Di Martino
W2GIS2
2021 Web Application Testing: Using Tree Kernels to Detect Near-duplicate States in Automated Model Inference
abstract
In the context of End-to-End testing of web applications, automated exploration techniques (a.k.a. crawling) are widely used to infer state-based models of the site under test. These models, in which states represent features of the web application and transitions represent reachability relationships, can be used for several model-based testing tasks, such as test case generation. However, current exploration techniques often lead to models containing many near-duplicate states, i.e., states representing slightly different pages that are in fact instances of the same feature. This has a negative impact on the subsequent model-based testing tasks, adversely affecting, for example, size, running time, and achieved coverage of generated test suites. As a web page can be naturally represented by its tree-structured DOM representation, we propose a novel near-duplicate detection technique to improve the model inference of web applications, based on Tree Kernel (TK) functions. TKs are a class of functions that compute similarity between tree-structured objects, largely investigated and successfully applied in the Natural Language Processing domain. To evaluate the capability of the proposed approach in detecting near-duplicate web pages, we conducted preliminary classification experiments on a freely-available massive dataset of about 100k manually annotated web page pairs. We compared the classification performance of the proposed approach with other state-of-the-art near-duplicate detection techniques. Preliminary results show that our approach performs better than state-of-the-art techniques in the near-duplicate detection classification task. These promising results show that TKs can be applied to near-duplicate detection in the context of web application model inference, and motivate further research in this direction.
Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
ESEM2
2021 Vehicular crowd-sensing: a parametric routing algorithm to increase spatio-temporal road network coverage
abstract
Current vehicles are equipped with a number of environmental sensors to improve safety and quality of life for passengers. Many researchers have shown that these sensors can also be exploited for opportunistic crowd-sensing. Useful new services can be developed on top of these data, like urban surveillance of Smart Cities. The spatio-temporal sensing coverage achievable with Vehicular Crowd-Sensing (VCS), however, is an open issue, since vehicles are not uniformly distributed over the road network, undermining the quality of potential services based on VCS data.In this paper, we present an evolution of the standard A ∗ routing algorithm, meant to increase VCS coverage by selecting a route in a random way among all those satisfying a parametric constraint on the total cost of the path. The proposed solution is based on an edge-computing paradigm, not requiring a central coordination but rather leveraging the computational resources available on-board, significantly reducing the back-end infrastructure costs. The proposed solution has been empirically evaluated on two public datasets of 450,000 real taxi trajectories from two cities, San Francisco and Porto, characterized by a very different road network topology. Results show sensible improvements in terms of achievable spatio-temporal sensing coverage of probe vehicles.
Dario Asprone, Sergio Di Martino, Paola Festa, Luigi L. L. Starace
Int. J. Geogr. Inf. Sci.2
2021 Security-Aware Deployment Optimization of Cloud-Edge Systems in Industrial IoT
abstract
Cloud computing, edge computing, and the Internet of Things are significantly changing from the original architectural models with pure provisioning of virtual resources (and services) to a transparent and adaptive hosting environment, where cloud providers, as well as “on-premise” resources and end nodes, fully realize the “everything-as-a-service” provisioning concept. The optimal design of these architectures, including the selection of optimal services to acquire, is not trivial in the cloud-edge context due to the involvement of a variable number and the type of available resources offerings and to the impact on cost, performance, and other relevant features such as security, almost never considered. This article presents a novel formalization of the cloud-edge allocation problem for the industrial IoT context. The proposed optimization process takes explicitly into account two critical aspects that are often overlooked in similar approaches, namely, the new cloud-edge on-demand service offerings model for the allocation of resources and the impact on the deployed application, in terms of cost, performance, and security policies actually implemented. An efficient yet suboptimal deterministic solver is also presented and compared with a linear programming one. Results are the same in 86% of the cases on the considered data set while our solver is orders of magnitude faster than the linear one.
Valentina Casola, Alessandra De Benedictis, Sergio Di Martino, Nicola Mazzocca, Luigi L. L. Starace
IEEE Internet Things J.3
2021 Comparing the effectiveness of capture and replay against automatic input generation for Android graphical user interface testing
abstract
Summary Exploratory testing and fully automated testing tools represent two viable and cheap alternatives to traditional test‐case‐based approaches for graphical user interface (GUI) testing of Android apps. The former can be executed by capture and replay tools that directly translate execution scenarios registered by testers in test cases, without requiring preliminary test‐case design and advanced programming/testing skills. The latter tools are able to test Android GUIs without tester intervention. Even if these two strategies are widely employed, to the best of our knowledge, no empirical investigation has been performed to compare their performance and obtain useful insights for a project manager to establish an effective testing strategy. In this paper, we present two experiments we carried out to compare the effectiveness of exploratory testing approaches using a capture and replay tool (Robotium Recorder) against three freely available automatic testing tools (AndroidRipper, Sapienz, and Google Robo). The first experiment involved 20 computer engineering students who were asked to record testing executions, under strict temporal limits and no access to the source code. Results were slightly better than those of fully automated tools, but not in a conclusive way. In the second experiment, the same students were asked to improve the achieved testing coverage by exploiting the source code and the coverage obtained in the previous tests, without strict temporal constraints. The results of this second experiment showed that students outperformed the automated tools especially for long/complex execution scenarios. The obtained findings provide useful indications for deciding testing strategies that combine manual exploratory testing and automated testing.
Sergio Di Martino, Anna Rita Fasolino, Luigi L. L. Starace, Porfirio Tramontana
Softw. Test. Verification Reliab.1
2020 An Haptic Interface for Industrial High-Precision Manufacturing Tasks
abstract
Within the Industry 4.0 context, a great number of machineries has been equipped with multiple sensors collecting vast amounts of heterogeneous data, including multimedia ones. In the context of high-precision industrial manufacturing, the output of these sensors can be exploited to leverage on human intelligence for monitoring the quality of the production. Nevertheless, in complex scenarios, the amount of sensed data could lead to a visual and acoustic overload for the Decision Maker. In this poster we propose a multi-modal user interface (UI) we devised to support the Decision Maker in monitoring the outcome of high-precision manufacturing machineries. In particular, to mitigate the acoustic and visual overloads, we propose the use of the haptic channel, both to control the playback of the collected data stream, and to get feedbacks about anomalous situations.
Sergio Di Martino, Vincenzo Norman Vitale
AVI1
2020 Inspecting Code Churns to Prioritize Test Cases
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
ICTSS3
2020 Massive Spatio-Temporal Mobility Data: An Empirical Experience on Data Management Techniques
Sergio Di Martino, Vincenzo Norman Vitale
W2GIS1
2020 Assessing the effectiveness of approximate functional sizing approaches for effort estimation
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro
Inf. Softw. Technol.1
2020 Smart Parking: Using a Crowd of Taxis to Sense On-Street Parking Space Availability
abstract
Monitoring the occupancy of on-street parking spaces on a city-wide scale is still an open issue. Past research demonstrated the viability of parking crowd-sensing by means of the standard on-board sensors of probe vehicles, foreseeing the use of high-mileage vehicles, like taxis. Nevertheless, the achievable spatio-temporal sensing coverage has never been deeply investigated. In this paper, we investigate the suitability of taxi fleets of different sizes to crowd-sense on-street parking availability. We considered 579 road segments in San Francisco (USA), covered both by sensors of the SFpark project and by the GPS traces of 536 taxis. For each of these segments, we computed the taxi transit frequencies, representing the achievable coverage by vehicles equipped with sensors detecting empty parking spots. By combining these frequencies with parking occupancy data coming from SFpark, we estimated the potential quality of crowd-sensed on-street parking information for different fleet sizes. Moreover, we investigated the impact of different misdetection amounts, and Kalman filters to handle them. The results show that a total of 300 taxis can crowd-sense on-street parking availability with an error of up to ±1 stall in 86% of the cases. Moreover, the quality of the sensors is as important as the fleet size (300 taxis with 10% probability of misreadings provide availability information comparable to 486 taxis with 16% probability), while the use of Kalman filters did not lead to statistically significant improvements. In conclusion, the traffic management authorities should consider parking crowd-sensing via probe vehicles as a promising alternative to the expensive deployment of the static parking sensors.
Fabian Bock, Sergio Di Martino, Antonio Origlia
IEEE Trans. Intell. Transp. Syst.2
2019 Comparing Different On-Street Parking Information for Parking Guidance and Information Systems
abstract
Parking search is a highly relevant problem in many cities. Parking Guidance and Information (PGI) systems support drivers by recommending locations and routes with higher chance to find parking. However, the relevance of such systems for on-street parking spaces is barely studied. In this paper, we investigate the consequences of providing the drivers with different levels of parking information to the search. Based on real on-street parking data, we investigated the scenario in which a driver does not find a parking space at the destination and has to decide on the next road to go, given three possible kinds of contextual information: (I) No parking information; (II) static information about the capacity of a road segment and (temporary) parking limitations; (III) real-time information collected from stationary sensors. Clearly the latter has strong implications in terms of deployment and operational costs. These scenarios lead to three different guidance strategies for a PGI system. We conducted empirical experiments on real data from San Francisco and on an artificially altered version of that dataset, to simulate a more competitive parking scenario. Results show that there is a significant reduction of parking search with more informed strategies, and that the use of realtime information offers only a limited improvement over static one. Only in presence of very limited parking availabilities, real-time data becomes more beneficial.
Sergio Di Martino, Vincenzo Norman Vitale, Fabian Bock
IV1
2019 What Is the Impact of On-street Parking Information for Drivers?
Fabian Bock, Sergio Di Martino, Monika Sester
W2GIS2
2019 Predicting the Spatial Impact of Planned Special Events
Sergio Di Martino, Simon Kwoczek, Silvia Rossi 0002
W2GIS1
2019 Industrial Internet of Things: Persistence for Time Series with NoSQL Databases
abstract
With the advent of Internet of Things (IoT) tech-nologies, there is a rapidly growing number of connected devices, producing more and more data, potentially useful for a large number of applications. The streams of data coming from each connected device can be seen as collections of Time Series, which need proper techniques to guarantee their persistence. In particular, these solutions must be able to provide both an effective data ingestion and data retrieval, which are challenging tasks. This problem is particularly sensible in the Industrial IoT (IIoT) context, given the potentially great number of equipment that could be instrumented with sensors generating time series. In this study we present the results of an empirical comparison of three NoSQL Database Management Systems, namely Cassandra, MongoDB and InfluxDB, in maintaining and retrieving gigabytes of real IIoT data, collected from an instrumented dressing machine. Results show that, for our specific Time Series dataset, InfluxDB is able to outperform Cassandra in all the considered tests, and has better overall performance respect to MongoDB.
Sergio Di Martino, Luca Fiadone, Adriano Peron, Alberto Riccabone, Vincenzo Norman Vitale
WETICE1
2018 A Visual Analytics GUI for Multigranular Spatio-Temporal Exploration and Comparison of Open Mobility Data
abstract
Recent technological developments in the fields of positioning and mobile communications gave rise to the availabilityof massive spatio-temporal open datasets about cities. A proper exploitation of these big datasets by decision makers of smart cities could be very useful to analyse and understand mobility patterns, with the final goal of easing many transportation problems, like parking search and traffic. While many research efforts have been aimed at defining powerful visual analytics tools for exploring vehicular trajectory data, to date almost no specifically tailored tools are available to analyse (on-street) parking data and dynamics. To fill this gap, in this paper we present the current state of an on-going research on the development of a visual analytics tool, meant to support decision makers of smart cities in performing multigranular spatio-temporal explorations of mobility open data, like those about parking. Moreover, the proposed GUI offers the possibility to overlay external spatio-temporal datasets as well as to customize the way this data is rendered, to get a better insight on the parking dynamics and its influencing factors.
Camilla Robino, Laura Di Rocco, Sergio Di Martino, Giovanna Guerrini, Michela Bertolotto
IV3
2018 Improving Sensing Coverage of Probe Vehicles with Probabilistic Routing
Dario Asprone, Sergio Di Martino, Paola Festa
W2GIS2
2018 Multigranular Spatio-Temporal Exploration: An Application to On-Street Parking Data
Camilla Robino, Laura Di Rocco, Sergio Di Martino, Giovanna Guerrini, Michela Bertolotto
W2GIS3
2017 Data-Driven Approaches for Smart Parking
Fabian Bock, Sergio Di Martino, Monika Sester
ECML/PKDD (3)2
2017 A comparison of two preference elicitation approaches for museum recommendations
abstract
Summary Recommendation systems based on collaborative filtering methods can be exploited in the context of providing personalized artworks tours within a museum. However, to be effectively used, we have several problems to be addressed: user preferences are not expressed as rating and recommendation systems must provide for new users efficient and simple preferences elicitation processes that do not require much effort and time. In this work, we present and evaluate 2 state‐of‐the‐art approaches that share the aim not to rely on individual item ratings. The first method uses a clustering algorithm to categorize items and provide recommendations, while the second one is inspired by the matrix factorization approach to select a couples of item groups that users have to evaluate to obtain preference profiles. We evaluate the 2 approaches with both an off‐line simulation and a user study with the aim to find the optimal configuration as well as to evaluate the effectiveness of the 2 proposed methods. Results show that the elicitation processes permit to obtain preference profiles in a time substantially less than the baseline method, while the differences in terms of prediction accuracy are minimal.
Silvia Rossi 0002, Francesco Barile, Sergio Di Martino, Davide Improta
Concurr. Comput. Pract. Exp.3
2016 Weighing lexical information for software clustering in the context of architecture recovery
Anna Corazza, Sergio Di Martino, Valerio Maggio, Giuseppe Scanniello
Empir. Softw. Eng.2
2016 Web Effort Estimation: Function Point Analysis vs. COSMIC
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro
Inf. Softw. Technol.1
2015 From Function Points to COSMIC - A Transfer Learning Approach for Effort Estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro
PROFES2
2014 Temporal and Spatial Clustering for a Parking Prediction Service
abstract
It has been estimated that in urban scenarios up to 30% of the traffic is due to vehicles looking for a free parking space. Thanks to recent technological evolutions, it is now possible to have at least a partial coverage of real-time data of parking space availability, and some preliminary mobile services are able to guide drivers towards free parking spaces. Nevertheless, the integration of this data within car navigators is challenging, mainly because (I) current In-Vehicle Telematic systems are not connected, and (II) they have strong limitations in terms of storage capabilities. To overcome these issues, in this paper we present a back-end based approach to learn historical models of parking availability per street. These compact models can then be easily stored on the map in the vehicle. In particular, we investigate the trade-off between the granularity level of the detailed spatial and temporal representation of parking space availability vs. The achievable prediction accuracy, using different spatio-temporal clustering strategies. The proposed solution is evaluated using five months of parking availability data, publicly available from the project Spark, based in San Francisco. Results show that clustering can reduce the needed storage up to 99%, still having an accuracy of around 70% in the predictions.
Felix Richter 0003, Sergio Di Martino, Dirk C. Mattfeld
ICTAI2
2013 Using tabu search to configure support vector regression for effort estimation
abstract
Recent studies have reported that Support Vector Regression (SVR) has the potential as a technique for software development effort estimation. However, its prediction accuracy is heavily influenced by the setting of parameters that needs to be done when employing it. No general guidelines are available to select these parameters, whose choice also depends on the characteristics of the dataset being used. This motivated the work described in (Corazza et al. 2010 ), extended herein. In order to automatically select suitable SVR parameters we proposed an approach based on the use of the meta-heuristics Tabu Search (TS). We designed TS to search for the parameters of both the support vector algorithm and of the employed kernel function, namely RBF. We empirically assessed the effectiveness of the approach using different types of datasets (single and cross-company datasets, Web and not Web projects) from the PROMISE repository and from the Tukutuku database. A total of 21 datasets were employed to perform a 10-fold or a leave-one-out cross-validation, depending on the size of the dataset. Several benchmarks were taken into account to assess both the effectiveness of TS to set SVR parameters and the prediction accuracy of the proposed approach with respect to widely used effort estimation techniques. The use of TS allowed us to automatically obtain suitable parameters’ choices required to run SVR. Moreover, the combination of TS and SVR significantly outperformed all the other techniques. The proposed approach represents a suitable technique for software development effort estimation.
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro, Emilia Mendes
Empir. Softw. Eng.2
2012 LINSEN: An efficient approach to split identifiers and expand abbreviations
abstract
Information Retrieval (IR) techniques are being exploited by an increasing number of tools supporting Software Maintenance activities. Indeed the lexical information embedded in the source code can be valuable for tasks such as concept location, clustering or recovery of traceability links. The application of such IR-based techniques relies on the consistency of the lexicon available in the different artifacts, and their effectiveness can worsen if programmers introduce abbreviations (e.g: rect) and/or do not strictly follow naming conventions such as Camel Case (e.g: UTFtoASCII). In this paper we propose an approach to automatically split identifiers in their composing words, and expand abbreviations. The solution is based on a graph model and performs in linear time with respect to the size of the dictionary, taking advantage of an approximate string matching technique. The proposed technique exploits a number of different dictionaries, referring to increasingly broader contexts, in order to achieve a disambiguation strategy based on the knowledge gathered from the most appropriate domain. The approach has been compared to other splitting and expansion techniques, using freely available oracles for the identifiers extracted from 24 C/C++ and Java open source systems. Results show an improvement in both splitting and expanding performance, in addition to a strong enhancement in the computational efficiency.
Anna Corazza, Sergio Di Martino, Valerio Maggio
ICSM2
2011 Using Web Objects for Development Effort Estimation of Web Applications: A Replicated Study
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro
PROFES1
2011 A Genetic Algorithm to Configure Support Vector Machines for Predicting Fault-Prone Components
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro
PROFES1
2011 A Rich Cloud Application to Improve Sustainable Mobility
Sergio Di Martino, Clemente Giorio, Raffaele Galiero
W2GIS1
2011 Investigating the use of Support Vector Regression for web effort estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
Empir. Softw. Eng.2
2010 A Tree Kernel based approach for clone detection
abstract
Reusing software by copying and pasting is a common practice in software development. This phenomenon is widely known as code cloning. Problems with clones are mainly due to the need of managing each duplication, thus increasing the effort to maintain software systems. Clone detection approaches generally take into account either the syntactic structure (e.g., Abstract Syntax Tree) or lexical elements (e.g., the signature of a function). In this paper we propose an approach to detect code clones, based on syntactic information enriched by lexical elements. To this end, we have defined a Tree Kernel function to compare Abstract Syntax Trees. A preliminary investigation has been also conducted to assess the validity of the proposed approach.
Anna Corazza, Sergio Di Martino, Valerio Maggio, Giuseppe Scanniello
ICSM2
2009 Applying support vector regression for web effort estimation using a cross-company dataset
abstract
Support vector regression (SVR) is a new generation of machine learning algorithms, suitable for predictive data modeling problems. The objective of this paper is to investigate the effectiveness of SVR for Web effort estimation, in particular when dealing with a cross-company dataset. To gain a deeper insight on the method, we carried out an empirical study using four kernels for SVR, namely linear, polynomial, Gaussian, and sigmoid. Moreover, we used two variables' preprocessing strategies (normalization and logarithmic), and two different dependent variables (effort and inverse effort). As a result, SVR was applied using six different configurations for each kernel. As for the dataset, we employed the Tukutuku database, which is widely adopted in Web effort estimation studies. A hold-out approach was adopted to evaluate the prediction accuracy for all the configurations, using two training sets, each containing data on 130 projects randomly selected, and two test sets, each containing the remaining 65 projects. As benchmark, SVR-based predictions were also compared to predictions obtained using manual stepwise regression, case-based reasoning, and Bayesian networks. Our results suggest that SVR performed well, since on the first hold-out, the linear kernel with a logarithmic transformation of variables provided significantly superior prediction accuracy than all the other techniques, while for the second hold-out, the Gaussian kernel achieved significantly superior predictions than all other techniques, except for manual stepwise regression.
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
ESEM2
2009 An Empirical Study on the Use of Web-COBRA and Web Objects to Estimate Web Application Development Effort
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino
ICWE1
2009 Using Support Vector Regression for Web Development Effort Estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
IWSM/Mensura2
2009 Automatic Generation of an Adaptive WebGIS
Sergio Di Martino, Filomena Ferrucci, Gavin McArdle, Giacomo Petillo
W2GIS1
2009 Measures and Techniques for Effort Estimation of Web Applications: an Empirical Study Based on a Single-Company Dataset
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
J. Web Eng.1
2008 Cross-company vs. single-company web effort models using the Tukutuku database: An extended study
Emilia Mendes, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino
J. Syst. Softw.2
2007 Comparing Size Measures for Predicting Web Application Development Effort: A Case Study
abstract
Size represents one of the most important attribute of software products used to predict software development effort. In the past nine years, several measures have been proposed to estimate the size of Web applications, and it is important to determine which one is most effective to predict Web development effort. To this aim in this paper we report on an empirical analysis where, using data from 15 Web projects developed by a software company, we compare four sets of size measures, using two prediction techniques, namely Forward Stepwise Regression (SWR) and Case-Based Reasoning (CBR). All the measures provided good predictions in terms of MMRE, MdMRE, and Pred(0.25) statistics, for both SWR and CBR. Moreover, when using SWR, length measures and Web Objects gave significant better results than Functional measures, however presented similar results to the Tukutuku measures. As for CBR, results did not show any significant differences amongst the four sets of size measures.
Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
ESEM1
2007 Towards the automatic generation of web GIS
abstract
In the present paper, we propose an approach for the development of Web GIS based on WebML, a high-level, formal visual language specifically conceived to design data-intensive Web applications. The proposal is motivated by the observation that Web GIS can be considered as a particular class of data-intensive Web applications. In the paper, we describe the extension of the visual formalism for modeling relevant interaction and navigation operations typical of Web GIS.
Sergio Di Martino, Filomena Ferrucci, Luca Paolino, Monica Sebillo, Genny Tortora, Giuliana Vitiello, Giuseppe Avagliano
GIS1
2007 A WebML-based Visual Language for the Development of Web GIS Applications
abstract
In the present paper, we propose a visual language meant to support the design of Web GIS applications. The proposal is based on the observation that Web GIS can be considered as a particular class of data- intensive Web applications, since they are mainly devoted to handle (spatial) information to and from the user. The success of WebML (Web Modeling Language) for designing traditional data-intensive Web applications suggested us to extend this visual formalism to model relevant interaction and navigation operations typical of Web GIS. The proposed extension consists of a set of content units specifically tailored for GIS concepts and tasks.
Sergio Di Martino, Filomena Ferrucci, Luca Paolino, Monica Sebillo, Giuliana Vitiello, Giuseppe Avagliano
VL/HCC1
2007 A WebML-Based Approach for the Development of Web GIS Applications
Sergio Di Martino, Filomena Ferrucci, Luca Paolino, Monica Sebillo, Giuliana Vitiello, Giuseppe Avagliano
WISE1
2007 A Replicated Study Comparing Web Effort Estimation Techniques
Emilia Mendes, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino
WISE2
2007 Effort estimation: how valuable is it for a web company to use a cross-company data set, compared to using its own single-company data set?
abstract
Previous studies comparing the prediction accuracy of effort models built using Web cross- and single-company data sets have been inconclusive, and as such replicated studies are necessary to determine under what circumstances a company can place reliance on a cross-company effort model. This paper therefore replicates a previous study by investigating how successful a cross-company effort model is: i) to estimate effort for Web projects that belong to a single company and were not used to build the cross-company model; ii) compared to a single-company effort model. Our single-company data set had data on 15 Web projects from a single company and our cross-company data set had data on 68 Web projects from 25 different companies. The effort estimates used in our analysis were obtained by means of two effort estimation techniques, namely forward stepwise regression and case-based reasoning. Our results were similar to those from the replicated study, showing that predictions based on the single-company model were significantly more accurate than those based on the cross-company model.
Emilia Mendes, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino
WWW2
2007 Towards a framework for mining and analysing spatio-temporal datasets
abstract
High‐resolution spatio‐temporal datasets are being collected every day to record the behaviour of several natural phenomena. However, data‐mining techniques are needed to extract relevant patterns from very large repositories and reveal spatial and temporal patterns in the behaviour of these phenomena. To this aim, we propose a system for mining data with spatial and temporal characteristics, and for visualizing and interpreting the results. Within this system, we have developed two complementary 3D visualization environments, one based on Google Earth and one relying on a Java3D graphical user interface. In this paper, we illustrate the main features of the system we have developed, and report on the main results we have obtained by analysing the Hurricane Isabel dataset.
Michela Bertolotto, Sergio Di Martino, Filomena Ferrucci, M. Tahar Kechadi
Int. J. Geogr. Inf. Sci.2
2006 Effort estimation modeling techniques: a case study for web applications
abstract
A reliable effort estimation is crucial for a successful web application development planning. Several approaches exist to address this issue. Among them, the algorithmic approach is one of the most widely used and investigated methods. It is based on suitable effort prediction models which relate the development effort with project characteristics. The size represents one of the most interesting characteristics of software products and several measures can be defined in order to estimate the size of web systems. Moreover, several techniques have been proposed in the literature to build the effort prediction models. Thus, of special interest should be to establish the most effective size measures to be employed in effort prediction models and the most suitable techniques for the model construction. To this aim some empirical studies have been undertaken so far. Since it is widely recognized that several investigations should be performed to verify/confirm empirical results, in the paper we will report on an empirical analysis we have carried out by exploiting data coming from 15 web projects developed by a software company. In particular, for the analysis we have considered two sets of size measures: Length Measures (e.g. number of pages, number of medias, number of client and server side scripts) and Functional Measures (e.g. external input, external output, external query). Moreover, we have employed different techniques, such as Linear Regression, Regression Tree, and Analogy-Based Estimation, in order to determine the one that provides the best prediction.
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello
ICWE2
2006 A COSMIC-FFP Approach to Predict Web Application Development Effort
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello
J. Web Eng.2
2005 A Cosmic-FFP Approach to Estimate WEB Application Development Effort
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Genny Tortora, Giuliana Vitiello
WEBIST2
2004 Handy: A New Interaction Device for Vehicular Information Systems
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Giuseppe Oliviero, Umberto Montemurro, Alessandro Paliotti
Mobile HCI2
2003 An Evaluation of Web3d Technologies from Developer's and End-User's Point of View
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci
SEKE2
2002 An approach for authoring 3D cultural heritage exhibitions on the web
abstract
The development of desktop virtual reality cultural exhibitions on the web is a challenging process, because it requires a collection of skills, ranging from art to 3D Internet technologies, and involves a variety of tasks. The need of suited approaches able to support the development of such exhibitions has motivated the introduction of the approach proposed in the paper. Such an approach is characterized by a strong attention towards the content experts, by a clear identification of the actors involved in the development process, and by a set of visual modeling languages, which support the high-level design of the exhibition and allow a more effective communication between the heterogeneous members of the project. Such modeling languages have been embedded in an authoring system which profitably supports the main figures to carry out their tasks.
Gennaro Costagliola, Sergio Di Martino, Filomena Ferrucci, Fabio Pittarello
SEKE2