Mario Mezzanzanica

dblp:11/6470 · DBLP profile ↗
← Back
21ranked-venue papers in the field
5as first author
7since 2021 · last 2026
0000-0003-0399-2810ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 7 (1 first)Database Systems & Data Management · 6 (3 first)Information Retrieval & Web Search · 3 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 3Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 VEUCTOR: Training and selecting best vector space models from online job ads for European countries
abstract
Over the last decade, word embeddings have enabled machines to represent words and sentences as vectors, enabling researchers to reason on text for tasks like semantic similarity, contextual understanding, machine translation, etc. However, the synthesis of embeddings involves domain-specific parameters that affect semantic accuracy and contextual relevance, often leading to unpredictable biases and inconsistent comparisons. This issue is particularly relevant in labor market analysis, where different embeddings yield varying results, making the selection of the most appropriate model a key element. This paper addresses these challenges by (i) proposing a methodology to train, select, and align vector space models for a target taxonomy, ensuring comparability across dimensions and languages; (ii) applying this approach to 4.5 million job ads in 28 languages, aligning country-specific embeddings using the ESCO taxonomy; (iii) generating over 3000 models over 142 machine days, making the best-performing ones publicly available via VEUCTOR ; and (iv) showing how model choice significantly impacts labor market analysis, revealing substantial variations in occupational skill bundles across embeddings. • We present, formalise, and implement a multilingual methodology to train, select, and align word embedding models using the ESCO taxonomy across 28 European countries. • We generate and evaluate over 3000 embedding models trained on 4.5 million online job advertisements in the frame of an EU Project, using a benchmark-driven approach to optimize semantic alignment. • We release VEUCTOR , a tool that provides access to the best-performing and aligned embeddings, enabling reuse and supporting third-party labor market analyses. • We show that the choice of embedding significantly affects occupational skill bundles and, consequently, labor market analysis outcomes. • We enable reproducible and cross-country labor market intelligence by standardizing model development and alignment across diverse languages and corpora.
Emilio Colombo, Simone D'Amico, Fabio Mercorio, Mario Mezzanzanica
Inf. Sci.4
2025 A Benchmark to Evaluate LLMs' Proficiency on Italian Student Competencies
Fabio Mercorio, Mario Mezzanzanica, Daniele Potertì, Antonio Serino, Andrea Seveso
ECML/PKDD (8)2
2024 Enriching Skill Taxonomies through Vector Space Models
abstract
Hierarchical taxonomies serve as fundamental structures for reasoning with hierarchical concepts across various domains such as healthcare, finance, and economy. However, maintaining their relevance and accuracy is a labor-intensive and error-prone task, demanding experts to identify and revise novel concepts constantly. In this context, distributional semantics techniques offer a promising avenue by suggesting terms likely to be associated with existing concepts. In our study, we propose a method to enhance taxonomies by adding related terms using contextual word embedding as encoders. We introduce VESPATE (VEctor SPAce model for Taxonomy Enrichment), a system designed to automatically expand any given hierarchical taxonomy with new terms using three generative models. Additionally, we integrate VESPATE with human validation to identify and select the most suitable terms for inclusion in the taxonomy. VESPATE was deployed within an EU project to enrich the official European Skill taxonomy, ESCO, with 40K+ digital terms gathered from the Web, aligning ESCO skills with current labor market needs. A total of 924 terms were selected through VESPATE, with 757 new terms subsequently validated by domain experts as correctly matched. Our framework, employing a pool of LLMs as encoders, helped us mitigate the limitations of the generative model, reducing the potential for errors and ensuring precise results in taxonomy enrichment. Additionally, the implementation of VESPATE consistently decreased the human effort required for the project. We evaluated the robustness of our system against a baseline constructed using ESCO’s hierarchy, achieving a 81% Positive Predictive Value (PPV) when combining all three models.
Simone D'Amico, Alessia De Santo, Fabio Mercorio, Mario Mezzanzanica
IEEE Big Data4
2024 Alignment of Multilingual Embeddings to Estimate Job Similarities in Online Labour Market
abstract
In recent years, word embeddings (WEs) have proven relevant for studying differences and similarities among job professions and skills required by the labour market across countries, providing valuable insights about the labour market dynamics to support policy and decision-making. In such a scenario, aligning WEs constructed across different countries and languages becomes key to allowing experts to reason on the labour market, catching technological and cultural shifts across borders. This paper proposes MEAL, an unsupervised method for aligning monolingual embeddings. Our approach selects a seed lexicon of anchors, i.e. words with the same meaning in both corpora that will be used as pivots in the alignment, without assuming a priori semantic similarities. Indeed, unlike previous literary works, to asses this relationship MEAL takes into account the semantic similarity between the neighbour of the two words in the WE space. Particularly, it chooses optimal anchors that are less susceptible to meaning shift. We deploy MEAL within the research framework of a European H-2020 Project that aims to use AI technologies to predict the future of the European labour market. Specifically, we apply it to the embeddings we train on 7+ millions of Online Job Advertisements (OJAs) collected in 2022. As a main outcome, MEAL allows stakeholders and policymakers (i) to estimate job similarities in Online Labour Markets across Europe, facilitating the assessment of how well these markets align with the taxonomy outlined by the official European Skills and Competences taxonomy, and (ii) to obtain indicators to support a data-driven policy design at a very fine-grained territorial level.
Simone D'Amico, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Filippo Pallucchini
DSAA4
2023 A survey on XAI and natural language explanations
Erik Cambria, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani
Inf. Process. Manag.4
2022 FFTree: A flexible tree to handle multiple fairness criteria
Alessandro Castelnovo, Andrea Cosentini, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica
Inf. Process. Manag.5
2021 TaxoRef: Embeddings Evaluation for AI-driven Taxonomy Refinement
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani
ECML/PKDD (3)3
2020 NEO: A Tool for Taxonomy Enrichment with New Emerging Occupations
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso
ISWC (2)4
2019 A Tool for Researchers: Querying Big Scholarly Data Through Graph Databases
abstract
We demonstrate GraphDBLP, a tool to allow researchers for querying the DBLP bibliography as a graph. The DBLP source data were enriched with semantic similarity relationships computed using wordembeddings. A user can interact with the system either via a Web-based GUI or using a shell-interface, both provided with three parametric and pre-defined queries. GraphDBLP would represent a first graph-database instance of the computer scientist network, that can be improved through new relationships and properties on nodes at any time, and this is the main purpose of the tool, that is freely available on Github. To date, GraphDBLP contains 5+ million nodes and 24+ million relationship.
Fabio Mercorio, Mario Mezzanzanica, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì
ECML/PKDD (3)2
2018 WoLMIS: a labor market intelligence system for classifying web job vacancies
Roberto Boselli, Mirko Cesarini, Stefania Marrara, Fabio Mercorio, Mario Mezzanzanica, Gabriella Pasi, Marco Viviani 0001
J. Intell. Inf. Syst.5
2017 A Pipeline for Multimedia Twitter Analysis through Graph Databases: Preliminary Results
abstract
Twitter is a microblogging service where users post not only short messages, but also images and other multimedia contents. Twitter can be used for analyzing people public discussions, as a huge amount of messages are continuously broadcasted by users. Analysis have usually focused on the textual part of messages, but the non-negligible number of images exchanged calls for specific attention. In this paper we describe how the tweet multimedia contents can be turned into a knowledge graph and then used for analyzing the messages sent during marketing campaigns. The information extraction and processing pipeline is built on top of off-theshelf APIs and products while the obtained knowledge is modelled through a Graph Database. The resulting knowledge graph was useful to explore and identify similarities among different marketing campaigns carried out using Twitter, providing some preliminary but promising results.
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica, Alessandro Vaccarino
DATA4
2017 Using Machine Learning for Labour Market Intelligence
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
ECML/PKDD (3)4
2017 An AI Planning System for Data Cleaning
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
ECML/PKDD (3)4
2017 A language modelling approach for discovering novel labour market occupations from the web
abstract
This article presents an approach for the identification of potential new occupations, i.e., professions, not yet codified by the international standard taxonomy ISCO. This work is framed within the research activities of the WoLMIS project, developed by the University of Milano-Bicocca for the CEDEFOP European Agency, which classifies on-line job offers according to the ISCO taxonomy by using machine learning techniques.
Stefania Marrara, Gabriella Pasi, Marco Viviani 0001, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica, Marco Pappagallo
WI6
2015 Applying the AHP to Smart Mobility Services: A Case Study
abstract
Making decision is a far from straightforward process, as it often requires to consider a number of complex criteria whose importance relies on the experiences and the preferences of the decision makers involved. Being able to structure and reproduce this knowledge is a challenging issue in the context of strategic decision making, and also common BI analytics can benefit from the joint use of that knowledge. As a contribution, in this work we describe how a multi criteria decision making technique, i.e., the Analytic Hierarchy Process, has been applied to a smart-mobility context, where the decision goal was to weight the factors that support the innovation of a smart mobility service in the city of Milan. The AHP has been selected as it allows considering both tangible and intangible factors that guide the decision within the model. We employed three distinct kind of stakeholders, namely service providers, over 35, and under 35 users and we synthesised a ranking of criteria on the basis of the preferences they provided. The results shed the light on the different judgments that each group gives to the identified criteria in terms of both ranking and importance.
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
DATA4
2015 A model-based evaluation of data quality activities in KDD
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
Inf. Process. Manag.1
2014 Are the Methodologies for Producing Linked Open Data Feasible for Public Administrations?
abstract
Linked Open Data (LOD) enable the semantic interoperability of Public Administration (PA) information. Moreover, they allow citizens to reuse public information for creating new services and applications. Although there are many methodologies and guidelines to produce and publish LOD, the PAs still hardly understand and exploit LOD to improve their activities. In this paper we show the use of a set of best practices to support an Italian PA in producing LOD. We show the case of LOD production from existing open datasets related to public services. Together with the production of LOD we present the definition of a reference ontology, the Public Service Ontology, integrated with the datasets. During the application, we highlight and discuss some critical points we found in methodologies and technologies described in the literature, and we identify some potential improvements.
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
DATA4
2014 Improving Data Cleansing Accuracy - A Model-based Approach
abstract
Abstract: Research on data quality is growing in importance in both industrial and academic communities, as it aims at deriving knowledge (and then value) from data. Information Systems generate a lot of data useful for studying the dynamics of subjects ’ behaviours or phenomena over time, making the quality of data a crucial aspect for guaranteeing the believability of the overall knowledge discovery process. In such a scenario, data cleansing techniques, i.e., automatic methods to cleanse a dirty dataset, are paramount. However, when multiple cleans-ing alternatives are available a policy is required for choosing between them. The policy design task still relies on the experience of domain-experts, and this makes the automatic identification of accurate policies a signifi-cant issue. This paper extends the Universal Cleaning Process enabling the automatic generation of an accurate cleansing policy derived from the dataset to be analysed. The proposed approach has been implemented and tested on an on-line benchmark dataset, a real-world instance of the Labour Market Domain. Our preliminary results show that our approach would represent a contribution towards the generation of data-driven policy, reducing significantly the domain-experts intervention for policy specification. Finally, the generated results have been made publicly available for downloading. 1
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
DATA1
2013 Automatic Synthesis of Data Cleansing Activities
abstract
Data cleansing is growing in importance among both public and private organisations, mainly due to the relevant amount of data exploited for supporting decision making processes. This paper is aimed to show how model-based verification algorithms (namely, model checking) can contribute in addressing data cleansing issues, furthermore a new benchmark problem focusing on the labour market dynamic is introduced. The consistent evolution of the data is checked using a model defined on the basis of domain knowledge. Then, we formally introduce the concept of universal cleanser, i.e. an object which summarises the set of all cleansing actions for each feasible data inconsistency (according to a given consistency model), then providing an algorithm which synthesises it. The universal cleanser can be seen as a repository of corrective interventions useful to develop cleansing routines. We applied our approach to a dataset derived from the Italian labour market data, making the whole dataset and outcomes publicly available to the community, so that the results we present can be shared and compared with other techniques
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
DATA1
2012 Data Quality Sensitivity Analysis on Aggregate Indicators
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
DATA1
2011 Data Quality through Model Checking Techniques
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
IDA1