VLDB 2026 Research / reviewers in the wild / expert
Fabio Mercorio
dblp:69/7560
· DBLP profile ↗
44ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-6864-2702ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking Distributional Vector Similarity Measures: A SurveyabstractABSTRACT Measuring semantic similarity between words or phrases is central to natural language processing, information retrieval and computational linguistics. Despite their importance, similarity and distance measures are typically chosen by default (e.g., cosine similarity) or in an ad hoc fashion, with little empirical justification. This lack of systematic evaluation creates two gaps: first, the absence of a comprehensive taxonomy of measures that spans set‐based, vector‐based and information‐theoretic approaches; second, the lack of task‐aware benchmarking that quantifies how these measures perform across different models and applications. In this paper, we address these gaps by comparing 15 similarity and distance measures on four NLP tasks (sentence similarity, kNN classification, correlation analysis and visualisation) using multiple benchmark datasets and embedding models. Our results reveal that the effectiveness of similarity measures varies substantially depending on the task and model, challenging the assumption that cosine similarity is universally optimal. These findings highlight the practical risk of relying on default measures and provide a principled basis for selecting similarity functions in NLP, information retrieval and related fields. Erik Cambria, Navid Nobani, Filippo Pallucchini, Fabio Mercorio |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | Synthetic data generation: A tertiary studyabstractSynthetic Data Generation (SDG) is expanding rapidly, yet existing surveys differ widely in scope and methodological quality. This tertiary study systematically searched four major scholarly databases (2015-2025) and, after PRISMA screening and DARE-4 appraisal, 1 identified 17 eligible secondary studies. The evidence reveals a strong concentration in healthcare (58.8% of surveys), limited coverage of non-health domains, and inconsistent reporting of evaluation protocols (e.g., incomplete specification of metrics, data splits, baselines, or evaluation scripts). Fidelity and downstream utility dominate assessment practices, whereas privacy and diversity remain under-examined. Only 4 of 17 surveys provide any reproducibility artefacts. By consolidating these findings, we propose a compact, domain-agnostic evaluation baseline and highlight structural gaps in transparency, domain breadth, and methodological consistency. The study offers actionable guidance for strengthening reproducibility and broadening the evidential foundations of SDG research. Navid Nobani, Giovanni Officioso, Filippo Pallucchini, Giancarlo Sperlì, Fabio Mercorio |
Inf. Process. Manag. | 5 |
| 2026 | VEUCTOR: Training and selecting best vector space models from online job ads for European countriesabstractOver the last decade, word embeddings have enabled machines to represent words and sentences as vectors, enabling researchers to reason on text for tasks like semantic similarity, contextual understanding, machine translation, etc. However, the synthesis of embeddings involves domain-specific parameters that affect semantic accuracy and contextual relevance, often leading to unpredictable biases and inconsistent comparisons. This issue is particularly relevant in labor market analysis, where different embeddings yield varying results, making the selection of the most appropriate model a key element. This paper addresses these challenges by (i) proposing a methodology to train, select, and align vector space models for a target taxonomy, ensuring comparability across dimensions and languages; (ii) applying this approach to 4.5 million job ads in 28 languages, aligning country-specific embeddings using the ESCO taxonomy; (iii) generating over 3000 models over 142 machine days, making the best-performing ones publicly available via VEUCTOR ; and (iv) showing how model choice significantly impacts labor market analysis, revealing substantial variations in occupational skill bundles across embeddings. • We present, formalise, and implement a multilingual methodology to train, select, and align word embedding models using the ESCO taxonomy across 28 European countries. • We generate and evaluate over 3000 embedding models trained on 4.5 million online job advertisements in the frame of an EU Project, using a benchmark-driven approach to optimize semantic alignment. • We release VEUCTOR , a tool that provides access to the best-performing and aligned embeddings, enabling reuse and supporting third-party labor market analyses. • We show that the choice of embedding significantly affects occupational skill bundles and, consequently, labor market analysis outcomes. • We enable reproducible and cross-country labor market intelligence by standardizing model development and alignment across diverse languages and corpora. Emilio Colombo, Simone D'Amico, Fabio Mercorio, Mario Mezzanzanica |
Inf. Sci. | 3 |
| 2025 | Towards the Terminator Economy: Assessing Job Exposure to AI Through LLMsabstractAI and related technologies are reshaping jobs and tasks, either by automating or augmenting human skills in the workplace. Many researchers have been working on estimating if and to what extent jobs and tasks are exposed to the risk of being automatized by AI-related technologies. Our work tackles this issue through a data-driven approach by: (i) developing a reproducible framework that uses cutting-edge open-source large language models to assess the current capabilities of AI and robotics in performing job-related tasks; (ii) formalizing and computing a measure of AI exposure by occupation, the Task Exposure to AI (TEAI) index, and a measure of Task Replacement by AI (TRAI) index, both validated through a human user evaluation and compared with the state-of-the-art. Our results show that the TEAI index is positively correlated with cognitive, problem-solving, and management skills, while it is negatively correlated with social skills. Results also suggest that about one-third of U.S. employment is highly exposed to AI, primarily in high-skill jobs requiring a graduate or postgraduate level of education. We also find that AI exposure is positively associated with employment and wage growth from 2003 to 2023, suggesting that AI has had an overall positive effect on productivity. Considering specifically the TRAI index, we find that even in high-skill occupations, AI exhibits high variability in task substitution, suggesting that AI and humans complement each other within the same occupation, while the allocation of tasks within occupations is likely to change. All results, models, and code are freely available online to allow the community to reproduce our results, compare outcomes, and use our work as a benchmark to monitor AI’s progress over time. Emilio Colombo, Fabio Mercorio, Mario Mezzanzanica, Antonio Serino |
IJCAI | 2 |
| 2025 | ITALIC: An Italian Culture-Aware Natural Language BenchmarkabstractAndrea Seveso, Daniele Potertì, Edoardo Federici, Mario Mezzanzanica, Fabio Mercorio. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Andrea Seveso, Daniele Potertì, Edoardo Federici, Mario Mezzanzanica, Fabio Mercorio |
NAACL (Long Papers) | 5 |
| 2025 | A Benchmark to Evaluate LLMs' Proficiency on Italian Student Competencies
Fabio Mercorio, Mario Mezzanzanica, Daniele Potertì, Antonio Serino, Andrea Seveso |
ECML/PKDD (8) | 1 |
| 2024 | Enriching Skill Taxonomies through Vector Space ModelsabstractHierarchical taxonomies serve as fundamental structures for reasoning with hierarchical concepts across various domains such as healthcare, finance, and economy. However, maintaining their relevance and accuracy is a labor-intensive and error-prone task, demanding experts to identify and revise novel concepts constantly. In this context, distributional semantics techniques offer a promising avenue by suggesting terms likely to be associated with existing concepts. In our study, we propose a method to enhance taxonomies by adding related terms using contextual word embedding as encoders. We introduce VESPATE (VEctor SPAce model for Taxonomy Enrichment), a system designed to automatically expand any given hierarchical taxonomy with new terms using three generative models. Additionally, we integrate VESPATE with human validation to identify and select the most suitable terms for inclusion in the taxonomy. VESPATE was deployed within an EU project to enrich the official European Skill taxonomy, ESCO, with 40K+ digital terms gathered from the Web, aligning ESCO skills with current labor market needs. A total of 924 terms were selected through VESPATE, with 757 new terms subsequently validated by domain experts as correctly matched. Our framework, employing a pool of LLMs as encoders, helped us mitigate the limitations of the generative model, reducing the potential for errors and ensuring precise results in taxonomy enrichment. Additionally, the implementation of VESPATE consistently decreased the human effort required for the project. We evaluated the robustness of our system against a baseline constructed using ESCO’s hierarchy, achieving a 81% Positive Predictive Value (PPV) when combining all three models. Simone D'Amico, Alessia De Santo, Fabio Mercorio, Mario Mezzanzanica |
IEEE Big Data | 3 |
| 2024 | Alignment of Multilingual Embeddings to Estimate Job Similarities in Online Labour MarketabstractIn recent years, word embeddings (WEs) have proven relevant for studying differences and similarities among job professions and skills required by the labour market across countries, providing valuable insights about the labour market dynamics to support policy and decision-making. In such a scenario, aligning WEs constructed across different countries and languages becomes key to allowing experts to reason on the labour market, catching technological and cultural shifts across borders. This paper proposes MEAL, an unsupervised method for aligning monolingual embeddings. Our approach selects a seed lexicon of anchors, i.e. words with the same meaning in both corpora that will be used as pivots in the alignment, without assuming a priori semantic similarities. Indeed, unlike previous literary works, to asses this relationship MEAL takes into account the semantic similarity between the neighbour of the two words in the WE space. Particularly, it chooses optimal anchors that are less susceptible to meaning shift. We deploy MEAL within the research framework of a European H-2020 Project that aims to use AI technologies to predict the future of the European labour market. Specifically, we apply it to the embeddings we train on 7+ millions of Online Job Advertisements (OJAs) collected in 2022. As a main outcome, MEAL allows stakeholders and policymakers (i) to estimate job similarities in Online Labour Markets across Europe, facilitating the assessment of how well these markets align with the taxonomy outlined by the official European Skills and Competences taxonomy, and (ii) to obtain indicators to support a data-driven policy design at a very fine-grained territorial level. Simone D'Amico, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Filippo Pallucchini |
DSAA | 3 |
| 2024 | Model-contrastive explanations through symbolic reasoningabstractExplaining how two machine learning classification models differ in their behaviour is gaining significance in eXplainable AI, given the increasing diffusion of learning-based decision support systems. Human decision-makers deal with more than one machine learning model in several practical situations. Consequently, the importance of understanding how two machine learning models work beyond their prediction performances is key to understanding their behaviour, differences, and likeness. Some attempts have been made to address these problems, for instance, by explaining text classifiers in a time-contrastive fashion. In this paper, we present MERLIN, a novel eXplainable AI approach that provides contrastive explanations of two machine learning models, introducing the concept of model-contrastive explanations. We propose an encoding that allows MERLIN to work with both text and tabular data and with mixed continuous and discrete features. To show the effectiveness of our approach, we evaluate it on an extensive set of benchmark datasets. MERLIN is also implemented as a python-pip package. Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso |
Decis. Support Syst. | 2 |
| 2023 | A survey on XAI and natural language explanations
Erik Cambria, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani |
Inf. Process. Manag. | 3 |
| 2022 | JoTA: Aligning Multilingual Job Taxonomies through Word Embeddings (Student Abstract)abstractWe propose JoTA (Job Taxonomy Alignment), a domain-independent, knowledge-poor method for automatic taxonomy alignment of lexical taxonomies via word embeddings. JoTA associates all the leaf terms of the origin taxonomy to one or many concepts in the destination one, employing a scoring function, which merges the score of a hierarchical method and the score of a classification task. JoTA is developed in the context of an EU Grant aiming at bridging the national taxonomies of EU countries towards the European Skills, Competences, Qualifications and Occupations taxonomy (ESCO) through AI. The method reaches a 0.8 accuracy on recommending top-5 occupations and a wMRR of 0.72. Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica |
AAAI | 3 |
| 2022 | The Good, the Bad, and the Explainer: A Tool for Contrastive Explanations of Text ClassifiersabstractIn the last few years, we have been witnessing the increasing deployment of machine learning-based systems, which act as black boxes whose behaviour is hidden to end-users. As a side-effect, this contributes to increasing the need for explainable methods and tools to support the coordination between humans and ML models towards collaborative decision-making. In this paper, we demonstrate ContrXT, a novel tool that computes the differences in the classification logic of two distinct trained models, reasoning on their symbolic representation through Binary Decision Diagrams. ContrXT is available as a pip package and API. Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani, Andrea Seveso |
IJCAI | 2 |
| 2022 | FFTree: A flexible tree to handle multiple fairness criteria
Alessandro Castelnovo, Andrea Cosentini, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica |
Inf. Process. Manag. | 4 |
| 2022 | XAI for myo-controlled prosthesis: Explaining EMG data for hand gesture classification
Noemi Gozzi, Lorenzo Malandri, Fabio Mercorio, Alessandra Pedrocchi |
Knowl. Based Syst. | 3 |
| 2022 | GraphLMI: A data driven system for exploring labor market information through graph databases
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica |
Multim. Tools Appl. | 3 |
| 2021 | NEO: A System for Identifying New Emerging Occupation from Job AdsabstractWe demonstrate NEO, a tool for automatically enriching the European Occupation and Skill Taxonomy (ESCO) with terms that represents new occupations extracted from million Online Job Advertisements (OJAs). NEO proposes (i) a novel metric that allows one to measure the semantic similarity between words in a taxonomy, and (ii) a set of measures that estimate the adherence of new terms to the most suited taxonomic concept, enabling the user to evaluate the suggestions. To test its effectiveness, NEO has been evaluated over 2M+ 2018 UK job ads, along with a user-study to confirm the usefulness of NEO in the taxonomy enrichment task. Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso |
AAAI | 3 |
| 2021 | A Method for Taxonomy-Aware Embeddings Evaluation (Student Abstract)abstractWhile word embeddings have been showing their effectiveness in capturing semantic and lexical similarities in a large number of domains, in case the corpus used to generate embeddings is associated with a taxonomy (i.e., classification tasks over standard de-jure taxonomies) the common intrinsic and extrinsic evaluation tasks cannot guarantee that the generated embeddings are consistent with the taxonomy. This, as a consequence sharply limits the use of distributional semantics in those domains. To address this issue, we design and implement MEET, which proposes a new measure -HSS- that allows evaluating embeddings from a text corpus preserving the semantic similarity relations of the taxonomy. Navid Nobani, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica |
AAAI | 3 |
| 2021 | Skills2Job: A Recommender System that Encodes Job Offer Embeddings on Graph Databases (Student Abstract)abstractWe propose a recommender system that, starting from a set of users skills, identifies the most suitable jobs as they emerge from a large text of Online Job Vacancies (OJVs). To this aim, we process 2.5M+ OJVs posted in three different countries (United Kingdom, France and Germany), generating several embeddings and performing an intrinsic evaluation of their quality. Besides, we compute a measure of skill importance for each occupation in each country, the Revealed Comparative Advantage (rca). The best vector models, together with the rca, are used to feed a graph database, which will serve as the keystone for the recommender system. Finally, a user study of 10 validates the effectiveness of Skills2Job, both in terms of precision and nDGC. Andrea Seveso, Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica |
AAAI | 4 |
| 2021 | Skills2Graph: Processing million Job Ads to face the Job Skill Mismatch ProblemabstractIn this paper, we present Skills2Graph, a tool that, starting from a set of users’ professional skills, identifies the most suitable jobs as they emerge from a large corpus of 2.5M+ Online Job Vacancies (OJVs) posted in three different countries (the United Kingdom, France, and Germany). To this aim, we rely both on co-occurrence statistics - computing a count-based measure of skill-relevance named Revealed Comparative Advantage (rca) - and distributional semantics - generating several embeddings on the OJVs corpus and performing an intrinsic evaluation of their quality. Results, evaluated through a user study of 10 labor market experts, show a high P@3 for the recommendations provided by Skills2Graph, and a high nDCG (0.985 and 0.984 in a [0,1] range), that indicates a strong correlation between the experts’ scores and the rankings generated by Skills2Graph. Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso |
IJCAI | 3 |
| 2021 | Towards an Explainer-agnostic Conversational XAIabstractExplainable Artificial Intelligence (XAI) is gaining interests in both academia and industry, mainly thanks to the proliferation of darker more complex black-box solutions which are replacing their more transparent ancestors. Believing that the overall performance of an XAI system can be augmented by considering the end-user as a human being, we are studying the ways we can improve the explanations by making them more informative and easier to use from one hand, and interactive and customisable from the other hand. Navid Nobani, Fabio Mercorio, Mario Mezzanzanica |
IJCAI | 2 |
| 2021 | A Human-AI Teaming Approach for Incremental Taxonomy Learning from TextabstractTaxonomies provide a structured representation of semantic relations between lexical terms, acting as the backbone of many applications. The research proposed herein addresses the topic of taxonomy enrichment using an ”human-in-the-loop” semi-supervised approach. I will be investigating possible ways to extend and enrich a taxonomy using corpora of unstructured text data. The objective is to develop a methodological framework potentially applicable to any domain. Andrea Seveso, Fabio Mercorio, Mario Mezzanzanica |
IJCAI | 2 |
| 2021 | TaxoRef: Embeddings Evaluation for AI-driven Taxonomy Refinement
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani |
ECML/PKDD (3) | 2 |
| 2020 | eXDiL: A Tool for Classifying and eXplaining Hospital Discharge Letters
Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso |
CD-MAKE | 1 |
| 2020 | NEO: A Tool for Taxonomy Enrichment with New Emerging Occupations
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso |
ISWC (2) | 3 |
| 2019 | A Tool for Researchers: Querying Big Scholarly Data Through Graph DatabasesabstractWe demonstrate GraphDBLP, a tool to allow researchers for querying the DBLP bibliography as a graph. The DBLP source data were enriched with semantic similarity relationships computed using wordembeddings. A user can interact with the system either via a Web-based GUI or using a shell-interface, both provided with three parametric and pre-defined queries. GraphDBLP would represent a first graph-database instance of the computer scientist network, that can be improved through new relationships and properties on nodes at any time, and this is the main purpose of the tool, that is freely available on Github. To date, GraphDBLP contains 5+ million nodes and 24+ million relationship. Fabio Mercorio, Mario Mezzanzanica, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì |
ECML/PKDD (3) | 1 |
| 2018 | Multimedia story creation on social networks
Flora Amato, Aniello Castiglione, Fabio Mercorio, Mario Mezzanzanica, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì |
Future Gener. Comput. Syst. | 3 |
| 2018 | Classifying online Job Advertisements through Machine Learning
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica |
Future Gener. Comput. Syst. | 3 |
| 2018 | WoLMIS: a labor market intelligence system for classifying web job vacancies
Roberto Boselli, Mirko Cesarini, Stefania Marrara, Fabio Mercorio, Mario Mezzanzanica, Gabriella Pasi, Marco Viviani 0001 |
J. Intell. Inf. Syst. | 4 |
| 2018 | GraphDBLP: a system for analysing networks of computer scientists through graph databases - GraphDBLP
Mario Mezzanzanica, Fabio Mercorio, Mirko Cesarini, Vincenzo Moscato, Antonio Picariello |
Multim. Tools Appl. | 2 |
| 2017 | A Pipeline for Multimedia Twitter Analysis through Graph Databases: Preliminary ResultsabstractTwitter is a microblogging service where users post not only short messages, but also images and other multimedia contents. Twitter can be used for analyzing people public discussions, as a huge amount of messages are continuously broadcasted by users. Analysis have usually focused on the textual part of messages, but the non-negligible number of images exchanged calls for specific attention. In this paper we describe how the tweet multimedia contents can be turned into a knowledge graph and then used for analyzing the messages sent during marketing campaigns. The information extraction and processing pipeline is built on top of off-theshelf APIs and products while the obtained knowledge is modelled through a Graph Database. The resulting knowledge graph was useful to explore and identify similarities among different marketing campaigns carried out using Twitter, providing some preliminary but promising results. Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica, Alessandro Vaccarino |
DATA | 3 |
| 2017 | Using Machine Learning for Labour Market Intelligence
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica |
ECML/PKDD (3) | 3 |
| 2017 | An AI Planning System for Data Cleaning
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica |
ECML/PKDD (3) | 3 |
| 2017 | A language modelling approach for discovering novel labour market occupations from the webabstractThis article presents an approach for the identification of potential new occupations, i.e., professions, not yet codified by the international standard taxonomy ISCO. This work is framed within the research activities of the WoLMIS project, developed by the University of Milano-Bicocca for the CEDEFOP European Agency, which classifies on-line job offers according to the ISCO taxonomy by using machine learning techniques. Stefania Marrara, Gabriella Pasi, Marco Viviani 0001, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica, Marco Pappagallo |
WI | 5 |
| 2016 | Heuristic Planning for Hybrid SystemsabstractPlanning in hybrid systems has been gaining research interest in the Artificial Intelligence community in recent years. Hybrid systems allow for a more accurate representation of real world problems, though solving them is very challenging due to complex system dynamics and a large model feature set. We developed DiNo, a new planner designed to tackle problems set in hybrid domains.DiNo is based on the discretise and validate approach and uses the novel Staged Relaxed Planning Graph+ (SRPG+) heuristic. Wiktor Mateusz Piotrowski, Maria Fox 0001, Derek Long, Daniele Magazzeni, Fabio Mercorio |
AAAI | 5 |
| 2016 | Heuristic Planning for PDDL+ Domains
Wiktor Mateusz Piotrowski, Maria Fox 0001, Derek Long, Daniele Magazzeni, Fabio Mercorio |
IJCAI | 5 |
| 2015 | Applying the AHP to Smart Mobility Services: A Case StudyabstractMaking decision is a far from straightforward process, as it often requires to consider a number of complex criteria whose importance relies on the experiences and the preferences of the decision makers involved. Being able to structure and reproduce this knowledge is a challenging issue in the context of strategic decision making, and also common BI analytics can benefit from the joint use of that knowledge. As a contribution, in this work we describe how a multi criteria decision making technique, i.e., the Analytic Hierarchy Process, has been applied to a smart-mobility context, where the decision goal was to weight the factors that support the innovation of a smart mobility service in the city of Milan. The AHP has been selected as it allows considering both tangible and intangible factors that guide the decision within the model. We employed three distinct kind of stakeholders, namely service providers, over 35, and under 35 users and we synthesised a ranking of criteria on the basis of the preferences they provided. The results shed the light on the different judgments that each group gives to the identified criteria in terms of both ranking and importance. Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica |
DATA | 3 |
| 2015 | A model-based evaluation of data quality activities in KDD
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio |
Inf. Process. Manag. | 4 |
| 2014 | Are the Methodologies for Producing Linked Open Data Feasible for Public Administrations?abstractLinked Open Data (LOD) enable the semantic interoperability of Public Administration (PA) information. Moreover, they allow citizens to reuse public information for creating new services and applications. Although there are many methodologies and guidelines to produce and publish LOD, the PAs still hardly understand and exploit LOD to improve their activities. In this paper we show the use of a set of best practices to support an Italian PA in producing LOD. We show the case of LOD production from existing open datasets related to public services. Together with the production of LOD we present the definition of a reference ontology, the Public Service Ontology, integrated with the datasets. During the application, we highlight and discuss some critical points we found in methodologies and technologies described in the literature, and we identify some potential improvements. Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica |
DATA | 3 |
| 2014 | Improving Data Cleansing Accuracy - A Model-based ApproachabstractAbstract: Research on data quality is growing in importance in both industrial and academic communities, as it aims at deriving knowledge (and then value) from data. Information Systems generate a lot of data useful for studying the dynamics of subjects ’ behaviours or phenomena over time, making the quality of data a crucial aspect for guaranteeing the believability of the overall knowledge discovery process. In such a scenario, data cleansing techniques, i.e., automatic methods to cleanse a dirty dataset, are paramount. However, when multiple cleans-ing alternatives are available a policy is required for choosing between them. The policy design task still relies on the experience of domain-experts, and this makes the automatic identification of accurate policies a signifi-cant issue. This paper extends the Universal Cleaning Process enabling the automatic generation of an accurate cleansing policy derived from the dataset to be analysed. The proposed approach has been implemented and tested on an on-line benchmark dataset, a real-world instance of the Labour Market Domain. Our preliminary results show that our approach would represent a contribution towards the generation of data-driven policy, reducing significantly the domain-experts intervention for policy specification. Finally, the generated results have been made publicly available for downloading. 1 Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio |
DATA | 4 |
| 2013 | Automatic Synthesis of Data Cleansing ActivitiesabstractData cleansing is growing in importance among both public and private organisations, mainly due to the relevant amount of data exploited for supporting decision making processes. This paper is aimed to show how model-based verification algorithms (namely, model checking) can contribute in addressing data cleansing issues, furthermore a new benchmark problem focusing on the labour market dynamic is introduced. The consistent evolution of the data is checked using a model defined on the basis of domain knowledge. Then, we formally introduce the concept of universal cleanser, i.e. an object which summarises the set of all cleansing actions for each feasible data inconsistency (according to a given consistency model), then providing an algorithm which synthesises it. The universal cleanser can be seen as a repository of corrective interventions useful to develop cleansing routines. We applied our approach to a dataset derived from the Italian labour market data, making the whole dataset and outcomes publicly available to the community, so that the results we present can be shared and compared with other techniques Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio |
DATA | 4 |
| 2012 | Data Quality Sensitivity Analysis on Aggregate Indicators
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio |
DATA | 4 |
| 2012 | A universal planning system for hybrid domains
Giuseppe Della Penna, Daniele Magazzeni, Fabio Mercorio |
Appl. Intell. | 3 |
| 2011 | Cost-optimal Strong Planning in Non-deterministic Domains
Giuseppe Della Penna, Fabio Mercorio, Benedetto Intrigila, Daniele Magazzeni, Enrico Tronci |
ICINCO (1) | 2 |
| 2011 | Data Quality through Model Checking Techniques
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio |
IDA | 4 |