Federica Rollo

dblp:194/9921 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-3834-3629ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond traditional models: Foundation models for accurate particulate matter prediction
abstract
Accurate PM 2.5 forecasting is critical for public health, yet traditional deep learning models trained on location-specific data often lack generalizability across geographic contexts. Time Series Foundation Models (TSFMs) offer zero-shot forecasting capabilities through large-scale pre-training, but their efficacy for air quality prediction using low-cost sensor (LCS) data remains unexplored. This study presents an empirical benchmarking study of TSFMs for PM 2.5 prediction using real-world LCS data evaluating Google’s TimesFM and IBM’s Granite against three deep learning architectures (CNN, LSTM, Transformer) across 34 datasets from 10 cities spanning 6 countries. We evaluate two forecasting strategies: direct prediction of reference station-equivalent data and prediction of LCS data with subsequent calibration. Our results show that TSFMs consistently outperform traditional deep learning models, particularly for short-term forecasts, while retaining an stable advantage over longer horizons. Zero-shot TSFMs, applied without any task-specific training, perform competitively across all evaluated sites, providing empirical evidence of transferability to unseen sensor deployments. Fine-tuning offers limited additional benefit over the zero-shot configuration. Direct reference station prediction using fine-tuned Granite emerges as the most effective strategy, achieving the lowest errors and highest R 2 values. No clear performance decline is observed in datasets with substantial missing data compared to those with few missing values when linear interpolation is applied to fill gaps. Input window length has limited impact on forecasting accuracy. These findings establish TSFMs as a promising and practically viable direction for air quality monitoring with LCS networks.
Federica Rollo, Matteo Angelinelli, Martina Casari, Laura Po, Giorgio Pedrazzi, Roberta Turra
Expert Syst. Appl.1
2026 Synthetic dataset generation for theft event extraction in Italian
abstract
Event extraction is the task of automatically identifying and extracting structured information about events from unstructured text. Despite Italian being a well-resourced language, it still lacks annotated datasets specifically designed for fine-grained event extraction. To address this gap, we propose a novel methodology for the generation of synthetic data suitable for fine-grained event extraction tasks. This work is motivated by the high cost and limited scalability of manual annotation. We introduce a controlled synthetic data generation pipeline that strictly adheres to a target annotation schema, providing a scalable alternative to extensive human labeling. The key methodological innovation is a two-phase, document-level generation framework that leverages Large Language Models, ensures structural consistency and mitigates generation biases, enabling the creation of high-quality datasets for complex event extraction scenarios. Using this methodology, we release SYNTH-ITA, the first collection of four medium-scale synthetic datasets for fine-grained Italian event extraction, generated from 10,000 structured crime scenarios each. Experiments conducted on event argument extraction using a QA formulation demonstrate that fine-tuning models on SYNTH-ITA leads to better or comparable performances to models fine-tuned on 200 manually annotated real news articles (+14% improvement with ELECTRA, -0.4% with BERT). Conversely, NER-based models for event argument extraction trained on synthetic data exhibit an 18% performance drop compared to those trained on manually annotated articles. • First medium-scale synthetic datasets for Italian event extraction. • A defined methodology and pipeline to ensure structured annotations and text alignment. • Expert validation and bias mitigation enhance data quality and fairness. • Synthetic data boosts Italian QA performance. • Public release of dataset/code supports Italian NLP tasks and reproducibility.
Giovanni Bonisoli, Federica Rollo, Laura Po
Inf. Process. Manag.2
2025 Scalable Cross-Location Calibration of Low-Cost Air Quality Sensors Using Heterogeneous Data
Martina Casari, Laura Po, Federica Rollo
IEEE Big Data3
2025 Enhancing Low-Cost Air Quality Sensors with AI for Smart Green Routing
abstract
Urban air pollution poses severe health risks, demanding accurate monitoring and exposure-reducing solutions.Low-cost air quality sensors (LCS) provide high spatial resolution but suffer from accuracy limitations that hinder their reliability.This paper presents the AIQS project (AI-enhanced air quality sensor for optimizing green routes), an ongoing initiative that combines artificial intelligence, sensor hardware optimization, and pedestrian routing innovation to address these challenges.AIQS applies machine learning techniques, including Multilayer Perceptrons and fuzzy logic, to correct sensor readings.In parallel, hardware-level optimizations, such as fluid dynamics simulations and pre-treatment modules, are explored to enhance sensor performance.The corrected AQ data is then incorporated into a configurable routing tool capable of estimating pollutant exposure and computing low-exposure pedestrian paths in urban environments.First evaluations, shows that our correction models achieve up to 0.92 R² against reference data across diverse urban environments.The corrected data drives a configurable routing tool that computes paths minimizing cumulative pollution exposure while balancing user preferences (e.g., proximity to green spaces).Preliminary validation in Modena, Italy demonstrates viable "green routes".
Laura Po, Martina Casari, Federica Rollo, Matteo Angelinelli, Giorgio Pedrazzi, Chiara De Pascali, Luca Francioso, Roberta Turra
FedCSIS3
2025 GRAFMOVE: Graph-based Mobility Optimization and Visualization Engine
abstract
Traditional pedestrian routing systems prioritize finding the shortest path, often neglecting user preferences and contextual factors such as green spaces or safety concerns. To address this limitation, we present GRAFMOVE, a customizable routing system that leverages Neo4j and OpenStreetMap data to integrate dynamic, user-specific criteria into pathfinding. GRAFMOVE constructs a footpath graph enriched with contextual data and implements a flexible cost function, enabling users to balance path length with personalized factors. The system includes an interactive dashboard for real-time route visualization and optimization. We demonstrate GRAFMOVE's capabilities through case studies, including Point Of Interest-based routing for tourists and personalized green path recommendations. Our approach advances pedestrian routing by offering a scalable, adaptable solution.
Federica Rollo, Laura Po
SIGSPATIAL/GIS1
2025 MODyPer: Multi-Objective Dynamic Personalized Route Planning for Vulnerable Road Users
abstract
Urban mobility is increasingly shifting towards sustainable modes of transportation, such as walking and cycling, necessitating intelligent route planning systems that cater to dynamic user preferences. This paper introduces MODyPer, a novel graph-based framework designed to optimize route recommendations for vulnerable road users (i.e., pedestrians and cyclists) by integrating multiple objectives, including travel distance, comfort, and environmental criteria. Unlike existing systems that prioritize single objectives (e.g., shortest path), our framework incorporates personalized weighting mechanisms, allowing users to define their preferences dynamically. We applied MODyPer to two real-world urban scenarios of Italian cities, demonstrating its effectiveness in balancing competing objectives while providing timely, user-centric routing.
Federica Rollo, Laura Po
SIGSPATIAL/GIS1
2025 Document-level event extraction from Italian crime news using minimal data
abstract
Event extraction from unstructured text is a critical task in natural language processing, often requiring substantial annotated data. This study presents an approach to document-level event extraction applied to Italian crime news, utilizing large language models (LLMs) with minimal labeled data. Our method leverages zero-shot prompting and in-context learning to effectively extract relevant event information. We address three key challenges: (1) identifying text spans corresponding to event entities, (2) associating related spans dispersed throughout the text with the same entity, and (3) formatting the extracted data into a structured JSON. The findings are promising: LLMs achieve an F1-score of approximately 60% for detecting event-related text spans, demonstrating their potential even in resource-constrained settings. This work represents a significant advancement in utilizing LLMs for tasks traditionally dependent on extensive data, showing that meaningful results are achievable with minimal data annotation. Additionally, the proposed approach outperforms several baselines, confirming its robustness and adaptability to various event extraction scenarios. • Novel use of LLMs for event extraction from Italian news with minimal annotated data. • In-context learning outperforms zero-shot prompting for identifying event spans. • Mixtral achieves top performance in event extraction from Italian crime news. • LLMs outperform QA models, proving robust in data-scarce environments. • In-context learning performance depends on the quality of selected examples.
Giovanni Bonisoli, David Vilares 0001, Federica Rollo, Laura Po
Knowl. Based Syst.3
2023 DICE: a Dataset of Italian Crime Event news
abstract
Extracting events from news stories as the aim of several Natural Language Processing (NLP) applications (e.g., question answering, news recommendation, news summarization) is not a trivial task, due to the complexity of natural language and the fact that news reporting is characterized by journalistic style and norms. Those aspects entail scattering an event description over several sentences within one document (or more documents), applying a mechanism of gradual specification of event-related information. This implies a widespread use of co-reference relations among the textual elements, conveying non-linear temporal information. In addition to this, despite the achievement of state-of-the-art results in several tasks, high-quality training datasets for non-English languages are rarely available.
Giovanni Bonisoli, Maria Pia di Buono, Laura Po, Federica Rollo
SIGIR4
2022 Semi Real-time Data Cleaning of Spatially Correlated Data in Traffic Sensor Networks
abstract
The new Internet of Things (IoT) era is submerging smart cities with data.Various types of sensors are widely used to collect massive amounts of data and to feed several systems such as surveillance, environmental monitoring, and disaster management.In these systems, sensors are deployed to make decisions or to predict an event.However, the accuracy of such decisions or predictions depends upon the reliability of the sensor data.By their nature, sensors are prone to errors, therefore identifying and filtering anomalies is extremely important.This paper proposes an anomaly detection and classification methodology for spatially correlated data of traffic sensors that combines different techniques and is able to distinguish between traffic sensor faults and unusual traffic conditions.The reliability of this methodology has been tested on real-world data.The application on two days affected by car accidents reveals that our approach can detect unusual traffic conditions.Moreover, the data cleaning process could enhance traffic management by ameliorating the traffic model performances. INTRODUCTIONPublic Administrations have begun to capture the large amount of data collected through IoT sensors in order to face the big challenge of sustainable development.Nowadays, many cities are equipped with traffic sensors installed on their road networks.The most diffuse sensor type is the induction loop: static sensors that are embedded under the road surface and provide real-time vehicle count and speed estimation.These data can be used as input to simulate realtime traffic scenarios that can effectively help Public Administration to cope with the mobility challenge -and instantaneously optimizing the transportation flow while sending new instructions to smart city devices like traffic lights.Traffic sensors are of great value for urban traffic modeling.However, they are not free of errors and faults, and the degradation of sensor performance can heavily affect the output of traffic model (Bachechi et al., 2020c).Therefore, detecting faulty traffic sensors is a fundamental step in order to boost the quality of the traffic management system (Bachechi et al., 2020d; Bachechi et al.,
Federica Rollo, Chiara Bachechi, Laura Po
WEBIST1
2021 Using Word Embeddings for Italian Crime News Categorization
abstract
Several studies have shown that the use of embeddings improves outcomes in many Natural Language Processing (NLP) activities, including text categorization.This paper focuses on how word embeddings can be used on newspaper articles related to crimes.The scope is the categorization of the news articles based on the type of crime they report.We compare different Word2Vec models and methods to obtain word embeddings.Then, we exploit both supervised and unsupervised Machine Learning categorization algorithms.Experiments were conducted on an Italian dataset of 15,361 crime news articles showing very promising results.
Giovanni Bonisoli, Federica Rollo, Laura Po
FedCSIS2
2021 ElastiCL: Elastic Quantization for Communication Efficient Collaborative Learning in IoT
abstract
Transmitting updates of high-dimensional models between client IoT devices and the central aggregating server has always been a bottleneck in collaborative learning - especially in uncertain real-world IoT networks where congestion, latency, bandwidth issues are common. In this scenario, gradient quantization is an effective way to reduce bits count when transmitting each model update, but with a trade-off of having an elevated error floor due to higher variance of the stochastic gradients. In this paper, we propose ElastiCL, an elastic quantization strategy that achieves transmission efficiency plus a low error floor by dynamically altering the number of quantization levels during training on distributed IoT devices. Experiments on training ResNet-18, Vanilla CNN shows that ElastiCL can converge in much fewer transmitted bits than fixed quantization level, with little or no compromise on training and test accuracy.
Bharath Sudharsan, Dhruv Sheth, Shailesh Arya, Federica Rollo, Piyush Yadav, Pankesh Patel, John G. Breslin, Muhammad Intizar Ali
SenSys4
2021 Anomaly Detection in Multivariate Spatial Time Series: A Ready-to-Use Implementation
Chiara Bachechi, Federica Rollo, Laura Po, Fabio Quattrini
WEBIST2
2020 Real-Time Data Cleaning in Traffic Sensor Networks
abstract
Through deploying Internet of Things (IoT) technologies, many aspects of the urban environment can be monitored in real-time. Mobility, pollution, parking, waste, lighting can be controlled and managed in an intelligent city thanks to a low-cost sensor network. Such big data streams generated in realtime by sensors need to be handled with appropriate techniques to detect erroneous measurements instantly. In this paper, we implement a fast data cleaning process to remove traffic sensor faults. Then, we present a traffic model that takes advantage of the detection of anomalous data measured by traffic sensors. Experiments on a real case scenario have demonstrated that anomaly detection can further improve the performance of a traffic model in emulating the real urban traffic.
Chiara Bachechi, Federica Rollo, Laura Po
AICCSA2
2020 Crime Event Localization and Deduplication
Federica Rollo, Laura Po
ISWC (2)1