Simone D'Amico

dblp:202/3787 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating the effectiveness of fine-tuning in financial NLP: The case of Social Trading Action Detection
abstract
Financial Natural Language Processing crucially leverages social media for market insights. However, most existing methods for this purpose rely on simple sentiment analysis models, which fail to capture the concrete trading intentions expressed in these discussions. While Large Language Models (LLMs) offer a promising alternative to simplistic sentiment analysis, the actual benefits of fine-tuning across different model families remain unclear in noisy, domain-specific contexts like online forums. To address this gap, we present a comprehensive assessment of the advantages and limitations of fine-tuning for Social Trading Action Detection (STAD), a novel task that aims to classify online posts into actionable categories, namely buy, sell, or other. In addition, we introduce FinReddit-2K, a manually annotated dataset consisting of 2123 Reddit posts, designed to serve as a benchmark for this task. Our experimental analysis goes beyond standard performance metrics and identifies both the types of errors that fine-tuning can successfully mitigate and those that it may inadvertently introduce. Through a systematic evaluation of 57 models, comparing 14 traditional models with 23 zero-shot LLMs and 20 fine-tuned variants, our results show that fine-tuning yields an average F1-score improvement of +15.1%. The best-performing model, a fine-tuned Mistral-7B, achieves an F1-score of 86.0%, although our analysis reveals that fine-tuning fails to produce meaningful performance gains in several scenarios.
Simone D'Amico, Andrea Maurino, Francesco Osborne, Giancarlo Sperlì
Inf. Process. Manag.1
2026 VEUCTOR: Training and selecting best vector space models from online job ads for European countries
abstract
Over the last decade, word embeddings have enabled machines to represent words and sentences as vectors, enabling researchers to reason on text for tasks like semantic similarity, contextual understanding, machine translation, etc. However, the synthesis of embeddings involves domain-specific parameters that affect semantic accuracy and contextual relevance, often leading to unpredictable biases and inconsistent comparisons. This issue is particularly relevant in labor market analysis, where different embeddings yield varying results, making the selection of the most appropriate model a key element. This paper addresses these challenges by (i) proposing a methodology to train, select, and align vector space models for a target taxonomy, ensuring comparability across dimensions and languages; (ii) applying this approach to 4.5 million job ads in 28 languages, aligning country-specific embeddings using the ESCO taxonomy; (iii) generating over 3000 models over 142 machine days, making the best-performing ones publicly available via VEUCTOR ; and (iv) showing how model choice significantly impacts labor market analysis, revealing substantial variations in occupational skill bundles across embeddings. • We present, formalise, and implement a multilingual methodology to train, select, and align word embedding models using the ESCO taxonomy across 28 European countries. • We generate and evaluate over 3000 embedding models trained on 4.5 million online job advertisements in the frame of an EU Project, using a benchmark-driven approach to optimize semantic alignment. • We release VEUCTOR , a tool that provides access to the best-performing and aligned embeddings, enabling reuse and supporting third-party labor market analyses. • We show that the choice of embedding significantly affects occupational skill bundles and, consequently, labor market analysis outcomes. • We enable reproducible and cross-country labor market intelligence by standardizing model development and alignment across diverse languages and corpora.
Emilio Colombo, Simone D'Amico, Fabio Mercorio, Mario Mezzanzanica
Inf. Sci.2
2024 Enriching Skill Taxonomies through Vector Space Models
abstract
Hierarchical taxonomies serve as fundamental structures for reasoning with hierarchical concepts across various domains such as healthcare, finance, and economy. However, maintaining their relevance and accuracy is a labor-intensive and error-prone task, demanding experts to identify and revise novel concepts constantly. In this context, distributional semantics techniques offer a promising avenue by suggesting terms likely to be associated with existing concepts. In our study, we propose a method to enhance taxonomies by adding related terms using contextual word embedding as encoders. We introduce VESPATE (VEctor SPAce model for Taxonomy Enrichment), a system designed to automatically expand any given hierarchical taxonomy with new terms using three generative models. Additionally, we integrate VESPATE with human validation to identify and select the most suitable terms for inclusion in the taxonomy. VESPATE was deployed within an EU project to enrich the official European Skill taxonomy, ESCO, with 40K+ digital terms gathered from the Web, aligning ESCO skills with current labor market needs. A total of 924 terms were selected through VESPATE, with 757 new terms subsequently validated by domain experts as correctly matched. Our framework, employing a pool of LLMs as encoders, helped us mitigate the limitations of the generative model, reducing the potential for errors and ensuring precise results in taxonomy enrichment. Additionally, the implementation of VESPATE consistently decreased the human effort required for the project. We evaluated the robustness of our system against a baseline constructed using ESCO’s hierarchy, achieving a 81% Positive Predictive Value (PPV) when combining all three models.
Simone D'Amico, Alessia De Santo, Fabio Mercorio, Mario Mezzanzanica
IEEE Big Data1
2024 Alignment of Multilingual Embeddings to Estimate Job Similarities in Online Labour Market
abstract
In recent years, word embeddings (WEs) have proven relevant for studying differences and similarities among job professions and skills required by the labour market across countries, providing valuable insights about the labour market dynamics to support policy and decision-making. In such a scenario, aligning WEs constructed across different countries and languages becomes key to allowing experts to reason on the labour market, catching technological and cultural shifts across borders. This paper proposes MEAL, an unsupervised method for aligning monolingual embeddings. Our approach selects a seed lexicon of anchors, i.e. words with the same meaning in both corpora that will be used as pivots in the alignment, without assuming a priori semantic similarities. Indeed, unlike previous literary works, to asses this relationship MEAL takes into account the semantic similarity between the neighbour of the two words in the WE space. Particularly, it chooses optimal anchors that are less susceptible to meaning shift. We deploy MEAL within the research framework of a European H-2020 Project that aims to use AI technologies to predict the future of the European labour market. Specifically, we apply it to the embeddings we train on 7+ millions of Online Job Advertisements (OJAs) collected in 2022. As a main outcome, MEAL allows stakeholders and policymakers (i) to estimate job similarities in Online Labour Markets across Europe, facilitating the assessment of how well these markets align with the taxonomy outlined by the official European Skills and Competences taxonomy, and (ii) to obtain indicators to support a data-driven policy design at a very fine-grained territorial level.
Simone D'Amico, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Filippo Pallucchini
DSAA1
2024 Online Supervised Training of Spaceborne Vision during Proximity Operations using Adaptive Kalman Filtering
abstract
This work presents an Online Supervised Training (OST) method to enable robust vision-based navigation about a non-cooperative spacecraft. Spaceborne Neural Networks (NN) are susceptible to domain gap as they are primarily trained with synthetic images due to the inaccessibility of space. OST aims to close this gap by training a pose estimation NN online using incoming flight images during Rendezvous and Proximity Operations (RPO). The pseudo-labels are provided by an adaptive unscented Kalman filter where the NN is used in the loop as a measurement module. Specifically, the filter tracks the target’s relative orbital and attitude motion, and its accuracy is ensured by robust on-ground training of the NN using only synthetic data. The experiments on real hardware-in-the-loop trajectory images show that OST can improve the NN performance on the target image domain given that OST is performed on images of the target viewed from a diverse set of directions during RPO.
Tae Ha Park, Simone D'Amico
ICRA2
2022 KRAKEN: A Novel Semantic-Based Approach for Keyphrases Extraction
abstract
We propose KRAKEN, a novel approach for the extraction of keyphrases from texts. To this aim, KRAKEN makes use of distributional semantics to identify, as completely as possible, representative portions of documents, i.e. keyphrases. In addition, we define novel metrics to assess a weighted significance to the keyphrases extracted from a document, identifying the most important ones by assessing their semantic similarity with the text of the document they belong to.
Simone D'Amico
IJCAI1