VLDB 2026 Research / reviewers in the wild / expert
Michele Marchetti
dblp:91/8611
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-3692-3600ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Token Reduction in Vision Transformers via Discrete Wavelet Decomposition
Christopher Buratti, Michele Marchetti, Federica Parlapiano, Davide Traini, Domenico Ursino, Luca Virgili |
ICPR (3) | 2 |
| 2026 | Exploiting knowledge graph communities to fine-tune large language modelsabstract• Community-based approach for fine-tuning Large Language Models • Focusing LLM training on Knowledge Graph substructures • Building expert systems by fine-tuning LLMs on domain-specific Knowledge Graph • Outperforming traditional fine-tuning in Knowledge Graph completion metrics Since the introduction of GPT-2, Large Language Models (LLMs) have proven to be able to handle various tasks with impressive performance. However, they sometimes generate incorrect output or even hallucinations. To overcome this problem, many researchers have investigated the possibility of integrating external factual knowledge, such as that encoded in Knowledge Graphs (KGs), into LLMs. Although there are many approaches in the existing literature that integrate KGs and LLMs in different ways, few of them use KGs to fine-tune LLMs, and none of them systematically use KG substructures. In this paper, we propose CoFine (Community-Based Fine-Tuner), an approach to fine-tune an LLM using the communities of a KG. CoFine works as follows: it first divides the KG into communities, each of which contains a homogeneous portion of the knowledge expressed by the KG. It then uses these communities to fine-tune the LLM. This way of proceeding allows LLM fine-tuning to focus on specific homogeneous information contained in the KG expressed by each community. CoFine allows the LLM to achieve a very high accuracy in knowledge completion tasks. This is evidenced by comparisons between CoFine and a baseline LLM fine-tuning approach, which showed that our approach achieves better results for all metrics considered with several KG. Alessia Amelio, Christopher Buratti, Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
Expert Syst. Appl. | 3 |
| 2026 | An ego network-based approach to fine-tune large language models using knowledge graphsabstractIn this paper, we introduce EgoFine, a new approach for fine-tuning Large Language Models (LLMs) based on ego networks extracted from Knowledge Graphs (KGs). EgoFine first identifies the most informative nodes in the KG using degree centrality. It then extracts their ego networks and generates structured training data through random paths within them, thus enabling LLMs to learn domain-specific knowledge. It further constructs negative samples to explicitly model the absence of relationships between entities. We present an experimental campaign involving three KGs (PrimeKG, WN18RR, and YAGO3) and four LLMs (Minerva-350M, Llama3.2-1B, Qwen2-1.5B, and Ministral-3B). This campaign demonstrates that EgoFine outperforms traditional embedding-based methods (e.g., AutoSF, BoxE, NodePiece, PairRE, and TransE), two state-of-the-art approaches integrating KGs and LLMs (e.g., GNN-RAG and KG-Adapter), as well as a baseline approach operating on the same principle as EgoFine but without exploiting the contribution of ego networks. Compared with this last approach, EgoFine improves Hit@1 values up to 47.37%, Mean Reciprocal Rank (MRR) values by up to 46.87%, F1-Score values by up to 36.59%, and Accuracy values by up to 30.91%. The paper also presents an ablation study devoted to evaluating several design choices underlying EgoFine, as well as an analysis of the EgoFine’s behavior when applied on dynamic or noisy KGs. This way of proceeding makes EgoFine particularly beneficial for a variety of real-world applications, including understanding complex biological mechanisms, reasoning about the relationships between legislative sources and court cases, interpreting and explaining complex industrial maintenance and production processes, and understanding the connections between attacks, exploits, and countermeasures in the context of cybersecurity. Alessia Amelio, Christopher Buratti, Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
Inf. Sci. | 3 |
| 2025 | Explaining Vision Transformers Through Similarity-based GraphsabstractVision Transformers (ViTs) have gained recognition in computer vision due to their outstanding performance. Despite their success, the explainability of ViT outputs is still a challenging issue. To address it, we propose a novel explainability method that leverages image patch embeddings from each attention layer of a ViT to construct similarity graphs. The latter are used to generate binary masks by exploring paths starting from specific patches. The masks from all layers are then aggregated into a comprehensive heatmap using the coverage bias formula. We tested our method on two Vision Transformer architectures (ViT-Base and DeiT-Base) and a subset of the ImageNet validation set. Using Insertion and Deletion metrics, we demonstrate the effectiveness of our proposed method compared to similar ones in the literature. Finally, we include a qualitative analysis that shows the capabilities of our method to make ViTs more interpretable. Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
IJCNN | 1 |
| 2025 | Integrating Gradient and Mask-based Approaches for Vision Transformer ExplainabilityabstractVision Transformers (ViTs) have demonstrated outstanding performance across different computer vision tasks thanks to their self-attention mechanism that captures long-range dependencies effectively. However, the inherent complexity of ViTs presents significant challenges in explaining their outputs, which is fundamental in safety-critical domains. To tackle the challenge of explaining ViT outputs, this paper presents Grad-Mask, a novel method that integrates gradients into the mask generation process to create explanation heatmaps. GradMask uses the query, key, and value matrices from each attention layer and computes their gradients with respect to a target class. Afterward, it uses these gradients to generate binary masks, which are then weighted by the corresponding ViT’s confidence scores. Finally, it combines the weighted masks to generate the resulting heatmap. Experimental evaluations on an ImageNet subset with ViT and DeiT (Data-efficient Image Transformer) architectures show that GradMask achieves competitive performance according to standard explainability metrics, such as Insertion, Deletion, and Pointing Game. A hyperparameter analysis confirms the high computational efficiency of GradMask, while an ablation study highlights the importance of combining gradients and masks for the generation of the explanation heatmap. Finally, a qualitative analysis shows the improved explainability of GradMask compared to existing methods, making it a promising approach for understanding ViTs. Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
IJCNN | 1 |
| 2025 | Adaptive patch selection to improve Vision Transformers through Reinforcement LearningabstractAbstract In recent years, Transformers have revolutionized the management of Natural Language Processing tasks, and Vision Transformers (ViTs) promise to do the same for Computer Vision ones. However, the adoption of ViTs is hampered by their computational cost. Indeed, given an image divided into patches, it is necessary to compute for each layer the attention of each patch with respect to all the others. Researchers have proposed many solutions to reduce the computational cost of attention layers by adopting techniques such as quantization, knowledge distillation and manipulation of input images. In this paper, we aim to contribute to the solution of this problem. In particular, we propose a new framework, called AgentViT, which uses Reinforcement Learning to train an agent that selects the most important patches to improve the learning of a ViT. The goal of AgentViT is to reduce the number of patches processed by a ViT, and thus its computational load, while still maintaining competitive performance. We tested AgentViT on CIFAR10, FashionMNIST, and Imagenette $$^+$$ + (which is a subset of ImageNet) in the image classification task and obtained promising performance when compared to baseline ViTs and other related approaches available in the literature. Francesco Cauteruccio, Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
Appl. Intell. | 2 |
| 2025 | Efficient token pruning in Vision Transformers using an attention-based Multilayer NetworkabstractVision Transformers (ViTs), although very successful, have a major limitation to overcome, namely the need for significant computational resources to use them. Several approaches have been proposed to limit the resources required to work with ViTs, aiming at pruning the data provided in input to them. In this paper, we propose Token Reduction via an Attention-based Multilayer network (TRAM), the first approach that achieves this goal using a multilayer network-based representation of the attention matrices. TRAM can work with most ViTs without the need for fine-tuning. It makes several contributions to the literature in this research area; in particular, it is characterized by: (i) a new representation of ViTs based on a multilayer network; (ii) a new approach to evaluate the relevance of tokens based on a new centrality measure computed on the multilayer network; and (iii) an approach to reduce the number of tokens based on this centrality measure. We have validated TRAM by comparing it with several state-of-the-art approaches during an extensive experimental campaign carried out on different image datasets. The results obtained demonstrate not only the efficiency but also the effectiveness of TRAM in reducing the computational load of ViTs while still allowing them to provide accurate results. • TRAM represents tokens using an attention-based multilayer network. • TRAM reduces ViT computational demand without requiring fine-tuning. • TRAM improves FPS and GFlops with near-Vanilla model accuracy. • Visual analysis reveals TRAM’s token selection process. Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
Expert Syst. Appl. | 1 |
| 2025 | Multiplex network-based representation of vision transformers for visual explainabilityabstractAbstract The enormous growth of artificial intelligence (AI), and deep learning (DL) in particular, has led to the widespread use of these systems in a variety of contexts. One DL model capable of addressing complex computer vision tasks is the vision transformer (ViT). Despite its huge success, the reasoning behind the inferences it makes is often unclear, which poses significant challenges in critical scenarios. In this paper, we propose a new approach called MUltiplex Transformer EXplainer (MUTEX), which aims to explain the inferences made by ViTs. MUTEX combines multiplex network-based representations of attention matrices and mask perturbation approaches to provide insight into the inference process of ViTs. By mapping the attention layers of a ViT into a multiplex network, MUTEX is able to analyze the relationships between different parts of the input image and identify the image patches that most influence the inference process. We tested MUTEX on a subset of ImageNet and on BloodMNIST and compared its performance with that of existing visual explainability approaches. In addition, to assess the robustness and adaptability of MUTEX, we conducted a qualitative analysis, along with a hyperparameter and ablation study, which allowed us to further appreciate its potential in visual explainability of ViT. Michele Marchetti, Davide Traini, Domenico Ursino, Luca Virgili |
Neural Comput. Appl. | 1 |
| 2024 | A model-agnostic, network theory-based framework for supporting XAI on classifiersabstractIn recent years, the enormous development of Machine Learning, especially Deep Learning, has led to the widespread adoption of Artificial Intelligence (AI) systems in a large variety of contexts. Many of these systems provide excellent results but act as black-boxes. This can be accepted in various contexts, but there are others (e.g., medical ones) where a result returned by a system cannot be accepted without an explanation on how it was obtained. Explainable AI (XAI) is an area of AI well suited to explain the behavior of AI systems that act as black-boxes. In this paper, we propose a model-agnostic XAI framework to explain the behavior of classifiers. Our framework is based on network theory; thus, it is able to make use of the enormous amount of results that researchers in this area have discovered over time. Being network-based, our framework is completely different from the other model-agnostic XAI approaches. Furthermore, it is parameter-free and is able to handle heterogeneous features that may not even be independent of each other. Finally, it introduces the notion of dyscrasia that allows us to detect not only which features are important in a particular task but also how they interact with each other. Gianluca Bonifazi, Francesco Cauteruccio, Enrico Corradini, Michele Marchetti, Giorgio Terracina, Domenico Ursino, Luca Virgili |
Expert Syst. Appl. | 4 |
| 2024 | A network analysis-based framework to understand the representation dynamics of graph neural networksabstractAbstract In this paper, we propose a framework that uses the theory and techniques of (Social) Network Analysis to investigate the learned representations of a Graph Neural Network (GNN, for short). Our framework receives a graph as input and passes it to the GNN to be investigated, which returns suitable node embeddings. These are used to derive insights on the behavior of the GNN through the application of (Social) Network Analysis theory and techniques. The insights thus obtained are employed to define a new training loss function, which takes into account the differences between the graph received as input by the GNN and the one reconstructed from the node embeddings returned by it. This measure is finally used to improve the performance of the GNN. In addition to describe the framework in detail and compare it with related literature, we present an extensive experimental campaign that we conducted to validate the quality of the results obtained. Gianluca Bonifazi, Francesco Cauteruccio, Enrico Corradini, Michele Marchetti, Domenico Ursino, Luca Virgili |
Neural Comput. Appl. | 4 |
| 2023 | A framework for investigating the dynamics of user and community sentiments in a social platformabstractSocial platforms are the preferred medium for many people to express their opinions on many topics. This has led many professionals from various fields (marketing, politics, research and development, etc.) to demand increasingly advanced approaches capable of analyzing the evolution of user or community sentiments on particular topics. In this paper, we want to make a contribution to addressing this issue. Specifically, we propose a model and a framework to analyze the dynamics of user and community sentiments in a social platform. In particular, our framework currently focuses on three activities, namely: (i) finding users capable of creating and maintaining a community that reflects their sentiment on a topic; (ii) studying how a user or community sentiment on a topic evolves over time; and (iii) investigating the cross-contamination between a user community and its neighborhood. We tested our framework by means of an extensive experimental campaign that we describe in the paper. Our framework is extremely scalable, and further activities can be easily implemented in it in the near future. Gianluca Bonifazi, Francesco Cauteruccio, Enrico Corradini, Michele Marchetti, Giorgio Terracina, Domenico Ursino, Luca Virgili |
Data Knowl. Eng. | 4 |
| 2023 | Representation and compression of Residual Neural Networks through a multilayer network based approach
Alessia Amelio, Gianluca Bonifazi, Francesco Cauteruccio, Enrico Corradini, Michele Marchetti, Domenico Ursino, Luca Virgili |
Expert Syst. Appl. | 5 |
| 2022 | An approach to detect backbones of information diffusers among different communities of a social platform
Gianluca Bonifazi, Francesco Cauteruccio, Enrico Corradini, Michele Marchetti, Alberto Pierini, Giorgio Terracina, Domenico Ursino, Luca Virgili |
Data Knowl. Eng. | 4 |
| 2003 | Towards a Conceptual Framework for UML to Hardware Description Language Mappings
Ian Oliver, Michele Marchetti |
FDL | 2 |