Mario Mezzanzanica

dblp:11/6470 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-0399-2810ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 21 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 10 since 2021Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VEUCTOR: Training and selecting best vector space models from online job ads for European countries
abstract
Over the last decade, word embeddings have enabled machines to represent words and sentences as vectors, enabling researchers to reason on text for tasks like semantic similarity, contextual understanding, machine translation, etc. However, the synthesis of embeddings involves domain-specific parameters that affect semantic accuracy and contextual relevance, often leading to unpredictable biases and inconsistent comparisons. This issue is particularly relevant in labor market analysis, where different embeddings yield varying results, making the selection of the most appropriate model a key element. This paper addresses these challenges by (i) proposing a methodology to train, select, and align vector space models for a target taxonomy, ensuring comparability across dimensions and languages; (ii) applying this approach to 4.5 million job ads in 28 languages, aligning country-specific embeddings using the ESCO taxonomy; (iii) generating over 3000 models over 142 machine days, making the best-performing ones publicly available via VEUCTOR ; and (iv) showing how model choice significantly impacts labor market analysis, revealing substantial variations in occupational skill bundles across embeddings. • We present, formalise, and implement a multilingual methodology to train, select, and align word embedding models using the ESCO taxonomy across 28 European countries. • We generate and evaluate over 3000 embedding models trained on 4.5 million online job advertisements in the frame of an EU Project, using a benchmark-driven approach to optimize semantic alignment. • We release VEUCTOR , a tool that provides access to the best-performing and aligned embeddings, enabling reuse and supporting third-party labor market analyses. • We show that the choice of embedding significantly affects occupational skill bundles and, consequently, labor market analysis outcomes. • We enable reproducible and cross-country labor market intelligence by standardizing model development and alignment across diverse languages and corpora.
Emilio Colombo, Simone D'Amico, Fabio Mercorio, Mario Mezzanzanica
Inf. Sci.4
2025 Towards the Terminator Economy: Assessing Job Exposure to AI Through LLMs
abstract
AI and related technologies are reshaping jobs and tasks, either by automating or augmenting human skills in the workplace. Many researchers have been working on estimating if and to what extent jobs and tasks are exposed to the risk of being automatized by AI-related technologies. Our work tackles this issue through a data-driven approach by: (i) developing a reproducible framework that uses cutting-edge open-source large language models to assess the current capabilities of AI and robotics in performing job-related tasks; (ii) formalizing and computing a measure of AI exposure by occupation, the Task Exposure to AI (TEAI) index, and a measure of Task Replacement by AI (TRAI) index, both validated through a human user evaluation and compared with the state-of-the-art. Our results show that the TEAI index is positively correlated with cognitive, problem-solving, and management skills, while it is negatively correlated with social skills. Results also suggest that about one-third of U.S. employment is highly exposed to AI, primarily in high-skill jobs requiring a graduate or postgraduate level of education. We also find that AI exposure is positively associated with employment and wage growth from 2003 to 2023, suggesting that AI has had an overall positive effect on productivity. Considering specifically the TRAI index, we find that even in high-skill occupations, AI exhibits high variability in task substitution, suggesting that AI and humans complement each other within the same occupation, while the allocation of tasks within occupations is likely to change. All results, models, and code are freely available online to allow the community to reproduce our results, compare outcomes, and use our work as a benchmark to monitor AI’s progress over time.
Emilio Colombo, Fabio Mercorio, Mario Mezzanzanica, Antonio Serino
IJCAI3
2025 ITALIC: An Italian Culture-Aware Natural Language Benchmark
abstract
Andrea Seveso, Daniele Potertì, Edoardo Federici, Mario Mezzanzanica, Fabio Mercorio. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Andrea Seveso, Daniele Potertì, Edoardo Federici, Mario Mezzanzanica, Fabio Mercorio
NAACL (Long Papers)4
2025 A Benchmark to Evaluate LLMs' Proficiency on Italian Student Competencies
Fabio Mercorio, Mario Mezzanzanica, Daniele Potertì, Antonio Serino, Andrea Seveso
ECML/PKDD (8)2
2024 Enriching Skill Taxonomies through Vector Space Models
abstract
Hierarchical taxonomies serve as fundamental structures for reasoning with hierarchical concepts across various domains such as healthcare, finance, and economy. However, maintaining their relevance and accuracy is a labor-intensive and error-prone task, demanding experts to identify and revise novel concepts constantly. In this context, distributional semantics techniques offer a promising avenue by suggesting terms likely to be associated with existing concepts. In our study, we propose a method to enhance taxonomies by adding related terms using contextual word embedding as encoders. We introduce VESPATE (VEctor SPAce model for Taxonomy Enrichment), a system designed to automatically expand any given hierarchical taxonomy with new terms using three generative models. Additionally, we integrate VESPATE with human validation to identify and select the most suitable terms for inclusion in the taxonomy. VESPATE was deployed within an EU project to enrich the official European Skill taxonomy, ESCO, with 40K+ digital terms gathered from the Web, aligning ESCO skills with current labor market needs. A total of 924 terms were selected through VESPATE, with 757 new terms subsequently validated by domain experts as correctly matched. Our framework, employing a pool of LLMs as encoders, helped us mitigate the limitations of the generative model, reducing the potential for errors and ensuring precise results in taxonomy enrichment. Additionally, the implementation of VESPATE consistently decreased the human effort required for the project. We evaluated the robustness of our system against a baseline constructed using ESCO’s hierarchy, achieving a 81% Positive Predictive Value (PPV) when combining all three models.
Simone D'Amico, Alessia De Santo, Fabio Mercorio, Mario Mezzanzanica
IEEE Big Data4
2024 Alignment of Multilingual Embeddings to Estimate Job Similarities in Online Labour Market
abstract
In recent years, word embeddings (WEs) have proven relevant for studying differences and similarities among job professions and skills required by the labour market across countries, providing valuable insights about the labour market dynamics to support policy and decision-making. In such a scenario, aligning WEs constructed across different countries and languages becomes key to allowing experts to reason on the labour market, catching technological and cultural shifts across borders. This paper proposes MEAL, an unsupervised method for aligning monolingual embeddings. Our approach selects a seed lexicon of anchors, i.e. words with the same meaning in both corpora that will be used as pivots in the alignment, without assuming a priori semantic similarities. Indeed, unlike previous literary works, to asses this relationship MEAL takes into account the semantic similarity between the neighbour of the two words in the WE space. Particularly, it chooses optimal anchors that are less susceptible to meaning shift. We deploy MEAL within the research framework of a European H-2020 Project that aims to use AI technologies to predict the future of the European labour market. Specifically, we apply it to the embeddings we train on 7+ millions of Online Job Advertisements (OJAs) collected in 2022. As a main outcome, MEAL allows stakeholders and policymakers (i) to estimate job similarities in Online Labour Markets across Europe, facilitating the assessment of how well these markets align with the taxonomy outlined by the official European Skills and Competences taxonomy, and (ii) to obtain indicators to support a data-driven policy design at a very fine-grained territorial level.
Simone D'Amico, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Filippo Pallucchini
DSAA4
2024 Model-contrastive explanations through symbolic reasoning
abstract
Explaining how two machine learning classification models differ in their behaviour is gaining significance in eXplainable AI, given the increasing diffusion of learning-based decision support systems. Human decision-makers deal with more than one machine learning model in several practical situations. Consequently, the importance of understanding how two machine learning models work beyond their prediction performances is key to understanding their behaviour, differences, and likeness. Some attempts have been made to address these problems, for instance, by explaining text classifiers in a time-contrastive fashion. In this paper, we present MERLIN, a novel eXplainable AI approach that provides contrastive explanations of two machine learning models, introducing the concept of model-contrastive explanations. We propose an encoding that allows MERLIN to work with both text and tabular data and with mixed continuous and discrete features. To show the effectiveness of our approach, we evaluate it on an extensive set of benchmark datasets. MERLIN is also implemented as a python-pip package.
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso
Decis. Support Syst.3
2023 A survey on XAI and natural language explanations
Erik Cambria, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani
Inf. Process. Manag.4
2022 JoTA: Aligning Multilingual Job Taxonomies through Word Embeddings (Student Abstract)
abstract
We propose JoTA (Job Taxonomy Alignment), a domain-independent, knowledge-poor method for automatic taxonomy alignment of lexical taxonomies via word embeddings. JoTA associates all the leaf terms of the origin taxonomy to one or many concepts in the destination one, employing a scoring function, which merges the score of a hierarchical method and the score of a classification task. JoTA is developed in the context of an EU Grant aiming at bridging the national taxonomies of EU countries towards the European Skills, Competences, Qualifications and Occupations taxonomy (ESCO) through AI. The method reaches a 0.8 accuracy on recommending top-5 occupations and a wMRR of 0.72.
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica
AAAI4
2022 The Good, the Bad, and the Explainer: A Tool for Contrastive Explanations of Text Classifiers
abstract
In the last few years, we have been witnessing the increasing deployment of machine learning-based systems, which act as black boxes whose behaviour is hidden to end-users. As a side-effect, this contributes to increasing the need for explainable methods and tools to support the coordination between humans and ML models towards collaborative decision-making. In this paper, we demonstrate ContrXT, a novel tool that computes the differences in the classification logic of two distinct trained models, reasoning on their symbolic representation through Binary Decision Diagrams. ContrXT is available as a pip package and API.
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani, Andrea Seveso
IJCAI3
2022 FFTree: A flexible tree to handle multiple fairness criteria
Alessandro Castelnovo, Andrea Cosentini, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica
Inf. Process. Manag.5
2022 GraphLMI: A data driven system for exploring labor market information through graph databases
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica
Multim. Tools Appl.4
2021 NEO: A System for Identifying New Emerging Occupation from Job Ads
abstract
We demonstrate NEO, a tool for automatically enriching the European Occupation and Skill Taxonomy (ESCO) with terms that represents new occupations extracted from million Online Job Advertisements (OJAs). NEO proposes (i) a novel metric that allows one to measure the semantic similarity between words in a taxonomy, and (ii) a set of measures that estimate the adherence of new terms to the most suited taxonomic concept, enabling the user to evaluate the suggestions. To test its effectiveness, NEO has been evaluated over 2M+ 2018 UK job ads, along with a user-study to confirm the usefulness of NEO in the taxonomy enrichment task.
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso
AAAI4
2021 A Method for Taxonomy-Aware Embeddings Evaluation (Student Abstract)
abstract
While word embeddings have been showing their effectiveness in capturing semantic and lexical similarities in a large number of domains, in case the corpus used to generate embeddings is associated with a taxonomy (i.e., classification tasks over standard de-jure taxonomies) the common intrinsic and extrinsic evaluation tasks cannot guarantee that the generated embeddings are consistent with the taxonomy. This, as a consequence sharply limits the use of distributional semantics in those domains. To address this issue, we design and implement MEET, which proposes a new measure -HSS- that allows evaluating embeddings from a text corpus preserving the semantic similarity relations of the taxonomy.
Navid Nobani, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica
AAAI4
2021 Skills2Job: A Recommender System that Encodes Job Offer Embeddings on Graph Databases (Student Abstract)
abstract
We propose a recommender system that, starting from a set of users skills, identifies the most suitable jobs as they emerge from a large text of Online Job Vacancies (OJVs). To this aim, we process 2.5M+ OJVs posted in three different countries (United Kingdom, France and Germany), generating several embeddings and performing an intrinsic evaluation of their quality. Besides, we compute a measure of skill importance for each occupation in each country, the Revealed Comparative Advantage (rca). The best vector models, together with the rca, are used to feed a graph database, which will serve as the keystone for the recommender system. Finally, a user study of 10 validates the effectiveness of Skills2Job, both in terms of precision and nDGC.
Andrea Seveso, Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica
AAAI5
2021 Skills2Graph: Processing million Job Ads to face the Job Skill Mismatch Problem
abstract
In this paper, we present Skills2Graph, a tool that, starting from a set of users’ professional skills, identifies the most suitable jobs as they emerge from a large corpus of 2.5M+ Online Job Vacancies (OJVs) posted in three different countries (the United Kingdom, France, and Germany). To this aim, we rely both on co-occurrence statistics - computing a count-based measure of skill-relevance named Revealed Comparative Advantage (rca) - and distributional semantics - generating several embeddings on the OJVs corpus and performing an intrinsic evaluation of their quality. Results, evaluated through a user study of 10 labor market experts, show a high P@3 for the recommendations provided by Skills2Graph, and a high nDCG (0.985 and 0.984 in a [0,1] range), that indicates a strong correlation between the experts’ scores and the rankings generated by Skills2Graph.
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso
IJCAI4
2021 Towards an Explainer-agnostic Conversational XAI
abstract
Explainable Artificial Intelligence (XAI) is gaining interests in both academia and industry, mainly thanks to the proliferation of darker more complex black-box solutions which are replacing their more transparent ancestors. Believing that the overall performance of an XAI system can be augmented by considering the end-user as a human being, we are studying the ways we can improve the explanations by making them more informative and easier to use from one hand, and interactive and customisable from the other hand.
Navid Nobani, Fabio Mercorio, Mario Mezzanzanica
IJCAI3
2021 A Human-AI Teaming Approach for Incremental Taxonomy Learning from Text
abstract
Taxonomies provide a structured representation of semantic relations between lexical terms, acting as the backbone of many applications. The research proposed herein addresses the topic of taxonomy enrichment using an ”human-in-the-loop” semi-supervised approach. I will be investigating possible ways to extend and enrich a taxonomy using corpora of unstructured text data. The objective is to develop a methodological framework potentially applicable to any domain.
Andrea Seveso, Fabio Mercorio, Mario Mezzanzanica
IJCAI3
2021 TaxoRef: Embeddings Evaluation for AI-driven Taxonomy Refinement
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani
ECML/PKDD (3)3
2020 eXDiL: A Tool for Classifying and eXplaining Hospital Discharge Letters
Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso
CD-MAKE2
2020 NEO: A Tool for Taxonomy Enrichment with New Emerging Occupations
Anna Giabelli, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Andrea Seveso
ISWC (2)4
2020 Fast and effective Big Data exploration by clustering
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Mario Mezzanzanica, Carlo Zaniolo
Future Gener. Comput. Syst.4
2019 A Tool for Researchers: Querying Big Scholarly Data Through Graph Databases
abstract
We demonstrate GraphDBLP, a tool to allow researchers for querying the DBLP bibliography as a graph. The DBLP source data were enriched with semantic similarity relationships computed using wordembeddings. A user can interact with the system either via a Web-based GUI or using a shell-interface, both provided with three parametric and pre-defined queries. GraphDBLP would represent a first graph-database instance of the computer scientist network, that can be improved through new relationships and properties on nodes at any time, and this is the main purpose of the tool, that is freely available on Github. To date, GraphDBLP contains 5+ million nodes and 24+ million relationship.
Fabio Mercorio, Mario Mezzanzanica, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì
ECML/PKDD (3)2
2018 Multimedia story creation on social networks
Flora Amato, Aniello Castiglione, Fabio Mercorio, Mario Mezzanzanica, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì
Future Gener. Comput. Syst.4
2018 Classifying online Job Advertisements through Machine Learning
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
Future Gener. Comput. Syst.4
2018 WoLMIS: a labor market intelligence system for classifying web job vacancies
Roberto Boselli, Mirko Cesarini, Stefania Marrara, Fabio Mercorio, Mario Mezzanzanica, Gabriella Pasi, Marco Viviani 0001
J. Intell. Inf. Syst.5
2018 GraphDBLP: a system for analysing networks of computer scientists through graph databases - GraphDBLP
Mario Mezzanzanica, Fabio Mercorio, Mirko Cesarini, Vincenzo Moscato, Antonio Picariello
Multim. Tools Appl.1
2017 A Pipeline for Multimedia Twitter Analysis through Graph Databases: Preliminary Results
abstract
Twitter is a microblogging service where users post not only short messages, but also images and other multimedia contents. Twitter can be used for analyzing people public discussions, as a huge amount of messages are continuously broadcasted by users. Analysis have usually focused on the textual part of messages, but the non-negligible number of images exchanged calls for specific attention. In this paper we describe how the tweet multimedia contents can be turned into a knowledge graph and then used for analyzing the messages sent during marketing campaigns. The information extraction and processing pipeline is built on top of off-theshelf APIs and products while the obtained knowledge is modelled through a Graph Database. The resulting knowledge graph was useful to explore and identify similarities among different marketing campaigns carried out using Twitter, providing some preliminary but promising results.
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica, Alessandro Vaccarino
DATA4
2017 Using Machine Learning for Labour Market Intelligence
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
ECML/PKDD (3)4
2017 An AI Planning System for Data Cleaning
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
ECML/PKDD (3)4
2017 A language modelling approach for discovering novel labour market occupations from the web
abstract
This article presents an approach for the identification of potential new occupations, i.e., professions, not yet codified by the international standard taxonomy ISCO. This work is framed within the research activities of the WoLMIS project, developed by the University of Milano-Bicocca for the CEDEFOP European Agency, which classifies on-line job offers according to the ISCO taxonomy by using machine learning techniques.
Stefania Marrara, Gabriella Pasi, Marco Viviani 0001, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica, Marco Pappagallo
WI6
2015 Applying the AHP to Smart Mobility Services: A Case Study
abstract
Making decision is a far from straightforward process, as it often requires to consider a number of complex criteria whose importance relies on the experiences and the preferences of the decision makers involved. Being able to structure and reproduce this knowledge is a challenging issue in the context of strategic decision making, and also common BI analytics can benefit from the joint use of that knowledge. As a contribution, in this work we describe how a multi criteria decision making technique, i.e., the Analytic Hierarchy Process, has been applied to a smart-mobility context, where the decision goal was to weight the factors that support the innovation of a smart mobility service in the city of Milan. The AHP has been selected as it allows considering both tangible and intangible factors that guide the decision within the model. We employed three distinct kind of stakeholders, namely service providers, over 35, and under 35 users and we synthesised a ranking of criteria on the basis of the preferences they provided. The results shed the light on the different judgments that each group gives to the identified criteria in terms of both ranking and importance.
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
DATA4
2015 A model-based evaluation of data quality activities in KDD
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
Inf. Process. Manag.1
2014 Are the Methodologies for Producing Linked Open Data Feasible for Public Administrations?
abstract
Linked Open Data (LOD) enable the semantic interoperability of Public Administration (PA) information. Moreover, they allow citizens to reuse public information for creating new services and applications. Although there are many methodologies and guidelines to produce and publish LOD, the PAs still hardly understand and exploit LOD to improve their activities. In this paper we show the use of a set of best practices to support an Italian PA in producing LOD. We show the case of LOD production from existing open datasets related to public services. Together with the production of LOD we present the definition of a reference ontology, the Public Service Ontology, integrated with the datasets. During the application, we highlight and discuss some critical points we found in methodologies and technologies described in the literature, and we identify some potential improvements.
Roberto Boselli, Mirko Cesarini, Fabio Mercorio, Mario Mezzanzanica
DATA4
2014 Improving Data Cleansing Accuracy - A Model-based Approach
abstract
Abstract: Research on data quality is growing in importance in both industrial and academic communities, as it aims at deriving knowledge (and then value) from data. Information Systems generate a lot of data useful for studying the dynamics of subjects ’ behaviours or phenomena over time, making the quality of data a crucial aspect for guaranteeing the believability of the overall knowledge discovery process. In such a scenario, data cleansing techniques, i.e., automatic methods to cleanse a dirty dataset, are paramount. However, when multiple cleans-ing alternatives are available a policy is required for choosing between them. The policy design task still relies on the experience of domain-experts, and this makes the automatic identification of accurate policies a signifi-cant issue. This paper extends the Universal Cleaning Process enabling the automatic generation of an accurate cleansing policy derived from the dataset to be analysed. The proposed approach has been implemented and tested on an on-line benchmark dataset, a real-world instance of the Labour Market Domain. Our preliminary results show that our approach would represent a contribution towards the generation of data-driven policy, reducing significantly the domain-experts intervention for policy specification. Finally, the generated results have been made publicly available for downloading. 1
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
DATA1
2013 Automatic Synthesis of Data Cleansing Activities
abstract
Data cleansing is growing in importance among both public and private organisations, mainly due to the relevant amount of data exploited for supporting decision making processes. This paper is aimed to show how model-based verification algorithms (namely, model checking) can contribute in addressing data cleansing issues, furthermore a new benchmark problem focusing on the labour market dynamic is introduced. The consistent evolution of the data is checked using a model defined on the basis of domain knowledge. Then, we formally introduce the concept of universal cleanser, i.e. an object which summarises the set of all cleansing actions for each feasible data inconsistency (according to a given consistency model), then providing an algorithm which synthesises it. The universal cleanser can be seen as a repository of corrective interventions useful to develop cleansing routines. We applied our approach to a dataset derived from the Italian labour market data, making the whole dataset and outcomes publicly available to the community, so that the results we present can be shared and compared with other techniques
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
DATA1
2012 Data Quality Sensitivity Analysis on Aggregate Indicators
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
DATA1
2011 Data Quality through Model Checking Techniques
Mario Mezzanzanica, Roberto Boselli, Mirko Cesarini, Fabio Mercorio
IDA1