Laura Po

dblp:60/3612 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-3345-176XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Beyond traditional models: Foundation models for accurate particulate matter prediction
abstract
Accurate PM 2.5 forecasting is critical for public health, yet traditional deep learning models trained on location-specific data often lack generalizability across geographic contexts. Time Series Foundation Models (TSFMs) offer zero-shot forecasting capabilities through large-scale pre-training, but their efficacy for air quality prediction using low-cost sensor (LCS) data remains unexplored. This study presents an empirical benchmarking study of TSFMs for PM 2.5 prediction using real-world LCS data evaluating Google’s TimesFM and IBM’s Granite against three deep learning architectures (CNN, LSTM, Transformer) across 34 datasets from 10 cities spanning 6 countries. We evaluate two forecasting strategies: direct prediction of reference station-equivalent data and prediction of LCS data with subsequent calibration. Our results show that TSFMs consistently outperform traditional deep learning models, particularly for short-term forecasts, while retaining an stable advantage over longer horizons. Zero-shot TSFMs, applied without any task-specific training, perform competitively across all evaluated sites, providing empirical evidence of transferability to unseen sensor deployments. Fine-tuning offers limited additional benefit over the zero-shot configuration. Direct reference station prediction using fine-tuned Granite emerges as the most effective strategy, achieving the lowest errors and highest R 2 values. No clear performance decline is observed in datasets with substantial missing data compared to those with few missing values when linear interpolation is applied to fill gaps. Input window length has limited impact on forecasting accuracy. These findings establish TSFMs as a promising and practically viable direction for air quality monitoring with LCS networks.
Federica Rollo, Matteo Angelinelli, Martina Casari, Laura Po, Giorgio Pedrazzi, Roberta Turra
Expert Syst. Appl.4
2026 Synthetic dataset generation for theft event extraction in Italian
abstract
Event extraction is the task of automatically identifying and extracting structured information about events from unstructured text. Despite Italian being a well-resourced language, it still lacks annotated datasets specifically designed for fine-grained event extraction. To address this gap, we propose a novel methodology for the generation of synthetic data suitable for fine-grained event extraction tasks. This work is motivated by the high cost and limited scalability of manual annotation. We introduce a controlled synthetic data generation pipeline that strictly adheres to a target annotation schema, providing a scalable alternative to extensive human labeling. The key methodological innovation is a two-phase, document-level generation framework that leverages Large Language Models, ensures structural consistency and mitigates generation biases, enabling the creation of high-quality datasets for complex event extraction scenarios. Using this methodology, we release SYNTH-ITA, the first collection of four medium-scale synthetic datasets for fine-grained Italian event extraction, generated from 10,000 structured crime scenarios each. Experiments conducted on event argument extraction using a QA formulation demonstrate that fine-tuning models on SYNTH-ITA leads to better or comparable performances to models fine-tuned on 200 manually annotated real news articles (+14% improvement with ELECTRA, -0.4% with BERT). Conversely, NER-based models for event argument extraction trained on synthetic data exhibit an 18% performance drop compared to those trained on manually annotated articles. • First medium-scale synthetic datasets for Italian event extraction. • A defined methodology and pipeline to ensure structured annotations and text alignment. • Expert validation and bias mitigation enhance data quality and fairness. • Synthetic data boosts Italian QA performance. • Public release of dataset/code supports Italian NLP tasks and reproducibility.
Giovanni Bonisoli, Federica Rollo, Laura Po
Inf. Process. Manag.3
2025 Scalable Cross-Location Calibration of Low-Cost Air Quality Sensors Using Heterogeneous Data
Martina Casari, Laura Po, Federica Rollo
IEEE Big Data2
2025 Enhancing Low-Cost Air Quality Sensors with AI for Smart Green Routing
abstract
Urban air pollution poses severe health risks, demanding accurate monitoring and exposure-reducing solutions.Low-cost air quality sensors (LCS) provide high spatial resolution but suffer from accuracy limitations that hinder their reliability.This paper presents the AIQS project (AI-enhanced air quality sensor for optimizing green routes), an ongoing initiative that combines artificial intelligence, sensor hardware optimization, and pedestrian routing innovation to address these challenges.AIQS applies machine learning techniques, including Multilayer Perceptrons and fuzzy logic, to correct sensor readings.In parallel, hardware-level optimizations, such as fluid dynamics simulations and pre-treatment modules, are explored to enhance sensor performance.The corrected AQ data is then incorporated into a configurable routing tool capable of estimating pollutant exposure and computing low-exposure pedestrian paths in urban environments.First evaluations, shows that our correction models achieve up to 0.92 R² against reference data across diverse urban environments.The corrected data drives a configurable routing tool that computes paths minimizing cumulative pollution exposure while balancing user preferences (e.g., proximity to green spaces).Preliminary validation in Modena, Italy demonstrates viable "green routes".
Laura Po, Martina Casari, Federica Rollo, Matteo Angelinelli, Giorgio Pedrazzi, Chiara De Pascali, Luca Francioso, Roberta Turra
FedCSIS1
2025 GRAFMOVE: Graph-based Mobility Optimization and Visualization Engine
abstract
Traditional pedestrian routing systems prioritize finding the shortest path, often neglecting user preferences and contextual factors such as green spaces or safety concerns. To address this limitation, we present GRAFMOVE, a customizable routing system that leverages Neo4j and OpenStreetMap data to integrate dynamic, user-specific criteria into pathfinding. GRAFMOVE constructs a footpath graph enriched with contextual data and implements a flexible cost function, enabling users to balance path length with personalized factors. The system includes an interactive dashboard for real-time route visualization and optimization. We demonstrate GRAFMOVE's capabilities through case studies, including Point Of Interest-based routing for tourists and personalized green path recommendations. Our approach advances pedestrian routing by offering a scalable, adaptable solution.
Federica Rollo, Laura Po
SIGSPATIAL/GIS2
2025 MODyPer: Multi-Objective Dynamic Personalized Route Planning for Vulnerable Road Users
abstract
Urban mobility is increasingly shifting towards sustainable modes of transportation, such as walking and cycling, necessitating intelligent route planning systems that cater to dynamic user preferences. This paper introduces MODyPer, a novel graph-based framework designed to optimize route recommendations for vulnerable road users (i.e., pedestrians and cyclists) by integrating multiple objectives, including travel distance, comfort, and environmental criteria. Unlike existing systems that prioritize single objectives (e.g., shortest path), our framework incorporates personalized weighting mechanisms, allowing users to define their preferences dynamically. We applied MODyPer to two real-world urban scenarios of Italian cities, demonstrating its effectiveness in balancing competing objectives while providing timely, user-centric routing.
Federica Rollo, Laura Po
SIGSPATIAL/GIS2
2025 Document-level event extraction from Italian crime news using minimal data
abstract
Event extraction from unstructured text is a critical task in natural language processing, often requiring substantial annotated data. This study presents an approach to document-level event extraction applied to Italian crime news, utilizing large language models (LLMs) with minimal labeled data. Our method leverages zero-shot prompting and in-context learning to effectively extract relevant event information. We address three key challenges: (1) identifying text spans corresponding to event entities, (2) associating related spans dispersed throughout the text with the same entity, and (3) formatting the extracted data into a structured JSON. The findings are promising: LLMs achieve an F1-score of approximately 60% for detecting event-related text spans, demonstrating their potential even in resource-constrained settings. This work represents a significant advancement in utilizing LLMs for tasks traditionally dependent on extensive data, showing that meaningful results are achievable with minimal data annotation. Additionally, the proposed approach outperforms several baselines, confirming its robustness and adaptability to various event extraction scenarios. • Novel use of LLMs for event extraction from Italian news with minimal annotated data. • In-context learning outperforms zero-shot prompting for identifying event spans. • Mixtral achieves top performance in event extraction from Italian crime news. • LLMs outperform QA models, proving robust in data-scarce environments. • In-context learning performance depends on the quality of selected examples.
Giovanni Bonisoli, David Vilares 0001, Federica Rollo, Laura Po
Knowl. Based Syst.4
2023 DICE: a Dataset of Italian Crime Event news
abstract
Extracting events from news stories as the aim of several Natural Language Processing (NLP) applications (e.g., question answering, news recommendation, news summarization) is not a trivial task, due to the complexity of natural language and the fact that news reporting is characterized by journalistic style and norms. Those aspects entail scattering an event description over several sentences within one document (or more documents), applying a mechanism of gradual specification of event-related information. This implies a widespread use of co-reference relations among the textual elements, conveying non-linear temporal information. In addition to this, despite the achievement of state-of-the-art results in several tasks, high-quality training datasets for non-English languages are rarely available.
Giovanni Bonisoli, Maria Pia di Buono, Laura Po, Federica Rollo
SIGIR3
2022 Road Network Graph Representation for Traffic Analysis and Routing
Chiara Bachechi, Laura Po
ADBIS2
2022 Semi Real-time Data Cleaning of Spatially Correlated Data in Traffic Sensor Networks
abstract
The new Internet of Things (IoT) era is submerging smart cities with data.Various types of sensors are widely used to collect massive amounts of data and to feed several systems such as surveillance, environmental monitoring, and disaster management.In these systems, sensors are deployed to make decisions or to predict an event.However, the accuracy of such decisions or predictions depends upon the reliability of the sensor data.By their nature, sensors are prone to errors, therefore identifying and filtering anomalies is extremely important.This paper proposes an anomaly detection and classification methodology for spatially correlated data of traffic sensors that combines different techniques and is able to distinguish between traffic sensor faults and unusual traffic conditions.The reliability of this methodology has been tested on real-world data.The application on two days affected by car accidents reveals that our approach can detect unusual traffic conditions.Moreover, the data cleaning process could enhance traffic management by ameliorating the traffic model performances. INTRODUCTIONPublic Administrations have begun to capture the large amount of data collected through IoT sensors in order to face the big challenge of sustainable development.Nowadays, many cities are equipped with traffic sensors installed on their road networks.The most diffuse sensor type is the induction loop: static sensors that are embedded under the road surface and provide real-time vehicle count and speed estimation.These data can be used as input to simulate realtime traffic scenarios that can effectively help Public Administration to cope with the mobility challenge -and instantaneously optimizing the transportation flow while sending new instructions to smart city devices like traffic lights.Traffic sensors are of great value for urban traffic modeling.However, they are not free of errors and faults, and the degradation of sensor performance can heavily affect the output of traffic model (Bachechi et al., 2020c).Therefore, detecting faulty traffic sensors is a fundamental step in order to boost the quality of the traffic management system (Bachechi et al., 2020d; Bachechi et al.,
Federica Rollo, Chiara Bachechi, Laura Po
WEBIST3
2021 Using Word Embeddings for Italian Crime News Categorization
abstract
Several studies have shown that the use of embeddings improves outcomes in many Natural Language Processing (NLP) activities, including text categorization.This paper focuses on how word embeddings can be used on newspaper articles related to crimes.The scope is the categorization of the news articles based on the type of crime they report.We compare different Word2Vec models and methods to obtain word embeddings.Then, we exploit both supervised and unsupervised Machine Learning categorization algorithms.Experiments were conducted on an Italian dataset of 15,361 crime news articles showing very promising results.
Giovanni Bonisoli, Federica Rollo, Laura Po
FedCSIS3
2021 Anomaly Detection in Multivariate Spatial Time Series: A Ready-to-Use Implementation
Chiara Bachechi, Federica Rollo, Laura Po, Fabio Quattrini
WEBIST3
2020 Real-Time Data Cleaning in Traffic Sensor Networks
abstract
Through deploying Internet of Things (IoT) technologies, many aspects of the urban environment can be monitored in real-time. Mobility, pollution, parking, waste, lighting can be controlled and managed in an intelligent city thanks to a low-cost sensor network. Such big data streams generated in realtime by sensors need to be handled with appropriate techniques to detect erroneous measurements instantly. In this paper, we implement a fast data cleaning process to remove traffic sensor faults. Then, we present a traffic model that takes advantage of the detection of anomalous data measured by traffic sensors. Experiments on a real case scenario have demonstrated that anomaly detection can further improve the performance of a traffic model in emulating the real urban traffic.
Chiara Bachechi, Federica Rollo, Laura Po
AICCSA3
2020 Visual analytics for spatio-temporal air quality data
abstract
Air pollution is the second biggest environmental concern for Europeans after climate change and the major risk to public health. It is imperative to monitor the spatiotemporal patterns of urban air pollution. The TRAFAIR air quality dashboard is an effective web application to empower decision-makers to be aware of the urban air quality conditions, define new policies, and keep monitoring their effects. The architecture copes with the multidimensionality of data and the real-time visualization challenge of big data streams coming from a network of low-cost sensors. Moreover, it handles the visualization and management of predictive air quality maps series that is produced by an air pollution dispersion model. Air quality data are not only visualized at a limited set of locations at different times but in the continuous space-time domain, thanks to interpolated maps that estimate the pollution at un-sampled locations.
Chiara Bachechi, Federico Desimoni, Laura Po, David Martínez Casas
IV3
2020 Crime Event Localization and Deduplication
Federica Rollo, Laura Po
ISWC (2)2
2020 Empirical evaluation of Linked Data visualization tools
Federico Desimoni, Laura Po
Future Gener. Comput. Syst.2
2019 Implementing an Urban Dynamic Traffic Model
abstract
The world of mobility is constantly evolving and proposing new technologies, such as autonomous driving, electromobility, shared-mobility or even new air transport systems. We do not know how people and things will be moving within cities in 30 years, but for sure we know that road network planning and traffic management will remain critical issues.
Chiara Bachechi, Laura Po
WI2
2015 Open Data for Improving Youth Policies
abstract
The Open Data \textit{philosophy} is based on the idea that certain data should be made ​​available to all citizens, in an open form, without any copyright restrictions, patents or other mechanisms of control. Various government have started to publish open data, first of all USA and UK in 2009, and in 2015, the Open Data Barometer project (www.opendatabarometer.org) states that on 77 diverse states across the world, over 55 percent have developed some form of Open Government Data initiative. We claim Public Administrations, that are the main producers and one of the consumers of Open Data, might effectively extract important information by integrating its own data with open data sources.This paper reports the activities carried on during a one-year research project on Open Data for Youth Policies. The project was mainly devoted to explore the youth situation in the municipalities and provinces of the Emilia Romagna region (Italy), in particular, to examine data on population, education and work.The project goals were: to identify interesting data sources both from the open data community and from the private repositories of local governments of Emilia Romagna region related to the Youth Policies; to integrate them and, to show up the result of the integration by means of a useful navigator tool; in the end, to publish new information on the web as Linked Open Data. This paper also reports the main issues encountered that may seriously affect the entire process of consumption, integration till the publication of open data.
Domenico Beneventano, Sonia Bergamaschi, Luca Gagliardelli, Laura Po
KEOD4
2015 Driving Innovation in Youth Policies with Open Data
Domenico Beneventano, Sonia Bergamaschi, Luca Gagliardelli, Laura Po
IC3K4
2015 Visual Querying LOD sources with LODeX
abstract
The Linked Open Data (LOD) Cloud has more than tripled its sources in just three years (from 295 sources in 2011 to 1014 in 2014). While the LOD data are being produced at a increasing rate, LOD tools lack in producing an high level representation of datasets and in supporting users in the exploration and querying of a source. To overcome the above problems and significantly increase the number of consumers of LOD data, we devised a new method and a tool, called LODeX, that promotes the understanding, navigation and querying of LOD sources both for experts and for beginners. It also provides a standardized and homogeneous summary of LOD sources and supports user in the creation of visual queries on previously unknown datasets.
Fabio Benedetti, Sonia Bergamaschi, Laura Po
K-CAP3
2014 Comparing Topic Models for a Movie Recommendation System
abstract
Recommendation systems have become successful at suggesting content that are likely to be of interest to the user, however their performance greatly suffers when little information about the users preferences are given. In this paper we propose an automated movie recommendation system based on the similarity of movie: given a target movie selected by the user, the goal of the system is to provide a list of those movies that are most similar to the target one, without knowing any user preferences. The Topic Models of Latent Semantic Allocation (LSA) and Latent Dirichlet Allocation (LDA) have been applied and extensively compared on a movie database of two hundred thousand plots. Experiments are an important part of the paper; we examined the topic models behaviour based on standard metrics and on user evaluations, we have conducted performance assessments with 30 users to compare our approach with a commercial system. The outcome was that the performance of LSA was superior to that of LDA in supporting the selection of similar plots. Even if our system does not outperform commercial systems, it does not rely on human effort, thus it can be ported to any domain where natural language descriptions exist. Since it is independent from the number of user ratings, it is able to suggest famous movies as well as old or unheard movies that are still strongly related to the content of the video the user has watched.
Sonia Bergamaschi, Laura Po, Serena Sorrentino
WEBIST (2)2
2013 An iPad Order Management System for Fashion Trade
Ivano Baroni, Sonia Bergamaschi, Laura Po
WEBIST3
2012 A Meta-language for MDX Queries in eLog Business Solution
abstract
The adoption of business intelligence technology in industries is growing rapidly. Business managers are not satisfied with ad hoc and static reports and they ask for more flexible and easy to use data analysis tools. Recently, application interfaces that expand the range of operations available to the user, hiding the underlying complexity, have been developed. The paper presents eLog, a business intelligence solution designed and developed in collaboration between the database group of the University of Modena and Reggio Emilia and eBilling, an Italian SME supplier of solutions for the design, production and automation of documentary processes for top Italian companies. eLog enables business managers to define OLAP reports by means of a web interface and to customize analysis indicators adopting a simple meta-language. The framework translates the user's reports into MDX queries and is able to automatically select the data cube suitable for each query. Over 140 medium and large companies have exploited the technological services of eBilling S.p.A. to manage their documents flows. In particular, eLog services have been used by the major media and telecommunications Italian companies and their foreign annex, such as Sky, Media set, H3G, Tim Brazil etc. The largest customer can provide up to 30 millions mail pieces within 6 months (about 200 GB of data in the relational DBMS). In a period of 18 months, eLog could reach 150 millions mail pieces (1 TB of data) to handle.
Sonia Bergamaschi, Matteo Interlandi, Mario Longo, Laura Po, Maurizio Vincini
ICDE4
2011 Automatic generation of probabilistic relationships for improving schema matching
Laura Po, Serena Sorrentino
Inf. Syst.1
2011 Using semantic techniques to access web data
Raquel Trillo Lado, Laura Po, Sergio Ilarri, Sonia Bergamaschi, Eduardo Mena
Inf. Syst.2
2010 Automatic Lexical Annotation Applied to the SCARLET Ontology Matcher
Laura Po, Sonia Bergamaschi
ACIIDS (2)1
2010 Schema label normalization for improving schema matching
Serena Sorrentino, Sonia Bergamaschi, Maciej Gawinecki, Laura Po
Data Knowl. Eng.4
2009 Schema Normalization for Improving Schema Matching
Serena Sorrentino, Sonia Bergamaschi, Maciej Gawinecki, Laura Po
ER4
2008 Improving Data Integration through Disambiguation Techniques
Laura Po
NLDB1
2007 MELIS - An Incremental Method for the Lexical Annotation of Domain Ontologies
Sonia Bergamaschi, Laura Po, Maurizio Vincini, Paolo Bouquet, Daniel Giacomuzzi, Francesco Guerra 0001
WEBIST (2)2
2007 An Incremental Method for the Lexical Annotation of Domain Ontologies
abstract
In this article, we present MELIS (Meaning Elicitation and Lexical Integration System), a method and a software tool for enabling an incremental process of automatic annotation of local schemas (e.g. relational database schemas, directory trees) with lexical information. The distinguishing and original feature of MELIS is the incremental process: the higher the number of schemas which are processed, the more background/ domain knowledge is cumulated in the system (a portion of domain ontology is learned at every step), the better the performance of the systems on annotating new schemas. MELIS has been tested as a component of the MOMIS-Ontology Builder, a framework able to create a domain ontology representing a set of selected data sources, described with a standard W3C language wherein concepts and attributes are annotated according to the lexical reference database. We describe the MELIS component within the MOMIS-Ontology Builder framework and provide some experimental results of MELIS as a standalone tool and as a component integrated in MOMIS.
Sonia Bergamaschi, Paolo Bouquet, Daniel Giacomuzzi, Francesco Guerra 0001, Laura Po, Maurizio Vincini
Int. J. Semantic Web Inf. Syst.5