VLDB 2026 Research / reviewers in the wild / expert
Giuseppe Fiameni
dblp:96/10002
· DBLP profile ↗
17ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0001-8687-6609ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synthetic frequency patterns injection for data-agnostic deepfake detectionabstractDeepfake detectors are typically trained on large sets of pristine and generated images, resulting in limited generalization capacity; they excel at identifying deepfakes created through methods encountered during training but struggle with those generated by unknown techniques.This paper introduces a learning approach aimed at significantly enhancing the generalization capabilities of deepfake detectors. Our method takes inspiration from the unique "fingerprints" that image generation processes consistently introduce into the frequency domain. These fingerprints manifest as structured and distinctly recognizable frequency patterns. We propose to train detectors using only pristine images injecting in part of them crafted frequency patterns, simulating the effects of various deepfake generation techniques without being specific to any. These synthetic patterns are based on generic shapes, grids, or auras.We evaluated our approach using diverse architectures across 25 different generation methods. The models trained with our approach were able to perform state-of-the-art deepfake detection, demonstrating also superior generalization capabilities in comparison with previous methods. Indeed, they are untied to any specific generation technique and can effectively identify deepfakes regardless of how they were made.The code to use the proposed approach and reproduce the presented experiments is available at: https://github.com/davide-coccomini/Deepfake-Detection-without-Deepfakes-Generalization-via-Synthetic-Frequency-Patterns-Injection Davide Coccomini, Roberto Caldelli, Claudio Gennaro, Giuseppe Fiameni, Giuseppe Amato 0001, Fabrizio Falchi |
Comput. Vis. Image Underst. | 4 |
| 2026 | Anakin: explainable android malware detection with graph neural networksabstractAbstract Android OS is today the most used Operating System for mobile devices. However, it is susceptible to several malware attacks that may seriously compromise the privacy and security of individuals and organizations. This paper proposes an approach based on a static analysis of decompiled Android PacKages (APKs) to extract critical APIs and detect Android malware. The main contributions lie in the adoption of a graph-based data engineering schema to represent APIs taken from the Function Call Graphs of decompiled APKs and the formulation of a graph-based deep learning approach for explainable malware detection. In particular, the proposed approach, named , implements a Graph Neural Network (GNN) for binary classification (malware versus goodware), and integrates algorithm to disclose how specific API classes and control-flow edges between API calls influence malware alerts. The proposed approach was evaluated by considering 26,527 Android APKs. The results of an extensive and in-depth evaluation show that the presented GNN model achieves higher accuracy than deep neural models trained with traditional API call sequence representations and publicly available related methods. On the other hand, it produces decision explanations that yield interesting insights into the malicious patterns of APKs and support root cause analysis of missed malware alarms. Giuseppina Andresini, Annalisa Appice, Vincenzo Belvedere, Giuseppe Fiameni, Donato Malerba |
Cybersecur. | 4 |
| 2026 | Multimodal predictive process monitoring and its application to explainable clinical pathwaysabstractThis paper presents one of the first contributions in the context of Multimodal Predictive Process Monitoring (MM-PPM) . In recent years, Predictive Process Monitoring (PPM) has evolved at the intersection of process mining, machine learning, and data science, as organizations seek to anticipate the future course of ongoing processes. Traditional PPM mainly relies on structured event log data, but many real-world scenarios generate richer information, including text, images, audio, and video. MM-PPM promises to start addressing this rich data scenario by integrating complementary knowledge from heterogeneous modalities through modality-specific representations and information fusion techniques. The growing digitization of healthcare systems, combined with advances in Artificial Intelligence (AI), has accelerated AI-based PPM for analyzing sequences of clinical events, supporting decision-making, enabling personalized care, and improving clinical facility management. Given these characteristics, clinical pathways represent an ideal domain for experimenting with MM-PPM, as they may naturally involve diverse modalities such as structured records, free-text notes, or medical images. To handle multimodal information available with clinical pathways, we introduce MEDUSA , an MM-PPM approach for outcome prediction, which jointly processes medical image information coupled with the storytelling of structural records and text notes collected during the clinical pathway of a patient until the acquisition of the considered image. The evaluation of MEDUSA is done in a COVID-19 case study, to assess the performance of the proposed approach and explain how specific information within each modality influences the decisions of the predictive model. Vincenzo Pasquadibisceglie, Ivan Donadello, Annalisa Appice, Oswald Lanz, Fabrizio Maria Maggi, Giuseppe Fiameni, Donato Malerba |
Inf. Syst. | 6 |
| 2025 | Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual PromptsabstractFew-shot semantic segmentation aims to segment objects from previously unseen classes using only a limited number of labeled examples. In this paper, we introduce Label Anything, a novel transformer-based architecture designed for multi-prompt, multi-way few-shot semantic segmentation. Our approach leverages diverse visual prompts—points, bounding boxes, and masks—to create a highly flexible and generalizable framework that significantly reduces annotation burden while maintaining high accuracy. Label Anything makes three key contributions: (i) we introduce a new task formulation that relaxes conventional few-shot segmentation constraints by supporting various types of prompts, multi-class classification, and enabling multiple prompts within a single image; (ii) we propose a novel architecture based on transformers and attention mechanisms; and (iii) we design a versatile training procedure allowing our model to operate seamlessly across different N-way K-shot and prompt-type configurations with a single trained model. Our extensive experimental evaluation on the widely used COCO-20i benchmark demonstrates that Label Anything achieves state-of-the-art performance among existing multi-way few-shot segmentation methods, while significantly outperforming leading single-class models when evaluated in multi-class settings. Code and trained models are available at https://github.com/pasqualedem/LabelAnything. Pasquale De Marinis, Nicola Fanelli, Raffaele Scaringi, Emanuele Colonna, Giuseppe Fiameni, Gennaro Vessio, Giovanna Castellano |
ECAI | 5 |
| 2025 | Leveraging a large language model (LLM) to predict hospital admissions of emergency department patientsabstractAn Emergency Department (ED) is a hospital facility that is staffed 24 hours a day, 7 days a week, and provides unscheduled services to patients whose health status requires immediate care. The management of the hospital admissions is one of the most critical processes for an ED. Several literature studies have recently used Artificial Intelligence (AI) methods to predict hospital admissions of ED patients. However, they generally handle ED journey features recording data with “one” multiplicity per journey. They often consider patients’ data collected at the end of ED journeys by rarely paying attention to obtain accurate decisions already during the early stages of ED journeys. Finally, they commonly use AI methods without equipping AI models with explanations of model decisions. Instead, this study illustrates an explainable Predictive Process Monitoring (PPM) method, called LEGOLAS , that uses a Large Language model (LLM) to obtain accurate and explainable decisions regarding the hospital admission of an ED patient. Decisions are already obtained in the early stages of an ED patient journey. A contribution of this study is the use of a story telling to describe ED patient journeys. This story telling accounts for the variety and multiplicity of information commonly recorded in free events logged for ED patients without requiring any complex preprocessing. Another contribution is the exploration of the accuracy performance of several foundation LLMs used with fine-tuning for the ED patient management process. A further contribution is an evaluation study with the real-life event log MIMICEL , to explore the impact and significance of the proposed method in terms of accuracy and earliness of decisions. This evaluation shows that LEGOLAS achieves overall accuracy equal to 0.76 by outperforming related AI methods that stop at an accuracy of 0.66. In addition, LEGOLAS achieves Fscore greater than 0.80 for the two majority classes of the study – Admitted and Home – after only twelve events observed in the ED patient journey. Finally, the evaluation explains which events and which categories of information better contribute to reveal the outcome of running ED patient journeys. Vincenzo Pasquadibisceglie, Annalisa Appice, Donato Malerba, Giuseppe Fiameni |
Expert Syst. Appl. | 4 |
| 2025 | GraphCLIP: Image-graph contrastive learning for multimodal artwork classificationabstractWe present GraphCLIP, a novel contrastive learning framework for multimodal artwork classification that integrates visual and contextual information to improve predictive accuracy and interpretability. Traditional computer vision methods often fall short in visual arts, where context is crucial. GraphCLIP leverages image data and a Knowledge Graph to extract features from both perspectives. Evaluated on the A r t G r a p h dataset, with over 100,000 artworks in 32 styles and 18 genres, GraphCLIP outperforms existing models in single-task (up to + 8 % in F1-score) and multi-task settings (up to + 6 % ), demonstrating robustness even with unseen classes. Additionally, visual and contextual qualitative explanations enhance model transparency. The versatility of GraphCLIP extends beyond art classification: its methodology can be adapted to other domains where integrating diverse data types is essential. (The code is publicly available at: https://github.com/CILAB-ArtGraph/graphclip.git .) • We introduce GraphCLIP, a contrastive learning framework for artwork classification. • GraphCLIP combines visual data with contextual knowledge. • We achieve state-of-the-art performance on the A r t G r a p h dataset. • We demonstrate robustness with unseen classes in distribution shift scenarios. • We provide visual and contextual explanations to enhance model interpretability. Raffaele Scaringi, Giuseppe Fiameni, Gennaro Vessio, Giovanna Castellano |
Knowl. Based Syst. | 2 |
| 2024 | Vito: Vision Transformer Optimization Via Knowledge Distillation On DecodersabstractIn this paper, we propose ViTO, a novel knowledge distillation strategy that aims to convert a CNN model into a transformer-based counterpart that incorporates the advantages of transformers while retaining or improving its inductive bias. Our approach is based on a two-level transformer architecture that includes an inner model for learning visual representations and an outer model that aims to match the teacher’s predictions through autoregression. Specifically, given an image in a batch, the outer model classifies the image by using, in addition to the image’s visual properties, also the predictions it has made on images previously seen within the same batch. The effect of this strategy is to allow the transformer to estimate self- and cross-attention across all input batch images to learn autoregressively intra-class and inter-class correlations.We experimentally validate ViTO on several standard benchmarks obtaining better performance than existing knowledge distillation strategies on transformers. Furthermore, our distilled transformer-based model shows better robustness properties than standard vision transformers, demonstrating the effectiveness of our proposed distillation strategy. Giovanni Bellitto, Renato Sortino, Paolo Spadaro, Simone Palazzo, Federica Proietto Salanitri, Giuseppe Fiameni, Efstratios Gavves, Concetto Spampinato |
ICIP | 6 |
| 2024 | Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Fiameni, Rita Cucchiara |
Int. J. Comput. Vis. | 3 |
| 2023 | Compositional Semantic Mix for Domain Adaptation in Point Cloud SegmentationabstractDeep-learning models for 3D point cloud semantic segmentation exhibit limited generalization capabilities when trained and tested on data captured with different sensors or in varying environments due to domain shift. Domain adaptation methods can be employed to mitigate this domain shift, for instance, by simulating sensor noise, developing domain-agnostic generators, or training point cloud completion networks. Often, these methods are tailored for range view maps or necessitate multi-modal input. In contrast, domain adaptation in the image domain can be executed through sample mixing, which emphasizes input data manipulation rather than employing distinct adaptation modules. In this study, we introduce compositional semantic mixing for point cloud domain adaptation, representing the first unsupervised domain adaptation technique for point cloud segmentation based on semantic and geometric sample mixing. We present a two-branch symmetric network architecture capable of concurrently processing point clouds from a source domain (e.g. synthetic) and point clouds from a target domain (e.g. real-world). Each branch operates within one domain by integrating selected data fragments from the other domain and utilizing semantic information derived from source labels and target (pseudo) labels. Additionally, our method can leverage a limited number of human point-level annotations (semi-supervised) to further enhance performance. We assess our approach in both synthetic-to-real and real-to-real scenarios using LiDAR datasets and demonstrate that it significantly outperforms state-of-the-art methods in both unsupervised and semi-supervised settings. Cristiano Saltori, Fabio Galasso, Giuseppe Fiameni, Nicu Sebe, Fabio Poiesi, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | From Show to Tell: A Survey on Deep Learning-Based Image CaptioningabstractConnecting Vision and Language plays an essential role in Generative Intelligence. For this reason, large research efforts have been devoted to image captioning, i.e. describing images with syntactically and semantically meaningful sentences. Starting from 2015 the task has generally been addressed with pipelines composed of a visual encoder and a language model for text generation. During these years, both components have evolved considerably through the exploitation of object regions, attributes, the introduction of multi-modal connections, fully-attentive approaches, and BERT-like early-fusion strategies. However, regardless of the impressive results, research in image captioning has not reached a conclusive answer yet. This work aims at providing a comprehensive overview of image captioning approaches, from visual encoding and text generation to training strategies, datasets, and evaluation metrics. In this respect, we quantitatively compare many relevant state-of-the-art approaches to identify the most impactful technical innovations in architectures and training strategies. Moreover, many variants of the problem and its open challenges are discussed. The final goal of this work is to serve as a tool for understanding the existing literature and highlighting the future directions for a research area where Computer Vision and Natural Language Processing can find an optimal synergy. Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Silvia Cascianelli, Giuseppe Fiameni, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | CoSMix: Compositional Semantic Mix for Domain Adaptation in 3D LiDAR Segmentation
Cristiano Saltori, Fabio Galasso, Giuseppe Fiameni, Nicu Sebe, Elisa Ricci 0001, Fabio Poiesi |
ECCV (33) | 3 |
| 2022 | GIPSO: Geometrically Informed Propagation for Online Adaptation in 3D LiDAR Segmentation
Cristiano Saltori, Evgeny Krivosheev, Stéphane Lathuilière, Nicu Sebe, Fabio Galasso, Giuseppe Fiameni, Elisa Ricci 0001, Fabio Poiesi |
ECCV (33) | 6 |
| 2022 | Higher-Order Recurrent Network with Space-Time Attention for Video Early Action RecognitionabstractEndowing visual agents with predictive capability is a key step towards video intelligence at scale. Early action recognition aims to predict the action labels before fully observing the complete video frames. Unlike action recognition, the model is asked to forecast the future or the effects by only observing the initial few frames. The strong reasoning ability over the temporal dimension is the key to success. To this end, in this paper, we propose a novel recurrent network with decomposed space-time attention and higher-order design to capture the temporal dependency associated with the specific actions. Our method achieves state-of-the-art performance on Something-Something and EPIC-Kitchens datasets under the early action recognition setting, showing evidence of predictive capability that we attribute to our higher-order recurrent design with space-time attention. Tsung-Ming Tai, Giuseppe Fiameni, Cheng-Kuang Lee, Oswald Lanz |
ICIP | 2 |
| 2022 | Unified Recurrence Modeling for Video Action AnticipationabstractForecasting future events based on evidence of current conditions is an innate skill of human beings, and key for predicting the outcome of any decision making. In artificial vision for example, we would like to predict the next human action before it happens, without observing the future video frames associated to it. Computer vision models for action anticipation are expected to collect the subtle evidence in the preamble of the target actions. In prior studies recurrence modeling often leads to better performance, the strong temporal inference is assumed to be a key element for reasonable prediction. To this end, we propose a unified recurrence modeling for video action anticipation via message passing framework. The information flow in space-time can be described by the interaction between vertices and edges, and the changes of vertices for each incoming frame reflects the underlying dynamics. Our model leverages self-attention as the building blocks for each of the message passing functions. In addition, we introduce different edge learning strategies that can be end-to-end optimized to gain better flexibility for the connectivity between vertices. Our experimental results demonstrate that our proposed method outperforms previous works on the large-scale EPIC-Kitchen dataset. Tsung-Ming Tai, Giuseppe Fiameni, Cheng-Kuang Lee, Simon See, Oswald Lanz |
ICPR | 2 |
| 2022 | A computational approach for progressive architecture shrinkage in action recognitionabstractAbstract Efficiency plays a key role in video understanding modeling, and developing more efficient spatiotemporal deep networks is a key ingredient for enabling their usage in production scenarios. In this work, we propose a methodology for reducing the computational complexity of a video understanding backbone while limiting the drop in accuracy caused by architectural changes. Our approach, named, Progressive Architecture Shrinkage, applies a sequence of reduction operators to the hyperparameters of a network to reduce its computational footprint. The choice of the sequence of operations is automatically optimized in a coordinate‐descent schema, and the approach transfers knowledge from both the initial network and previous stages of the shrinking process by employing a Knowledge Distillation and an adaptive fine‐tuning strategy. As each iteration of the shrinking algorithm requires to train a large‐scale video understanding network, we perform experiments on MARCONI 100—a supercomputer equipped with an IBM Power9 architecture and Volta NVIDIA GPUs. Experimental evaluations are conducted using two backbones and three different action recognition benchmarks. We show that, through our approach, high accuracy levels can be maintained while reducing the number of multiply–adds operations by four times with respect to the original architectures. Code will be made available. Matteo Tomei, Lorenzo Baraldi 0001, Giuseppe Fiameni, Simone Bronzin, Rita Cucchiara |
Softw. Pract. Exp. | 3 |
| 2021 | Deep learning searches for gravitational wave stochastic backgroundsabstractThe background of gravitational waves (GW) has long been studied and remains one of the most exciting aspects in the observation and analysis of gravitational radiation. The paper focuses on the search for the background of gravitational waves using deep neural networks. An astrophysical background due to the presence of many binary black hole coalescences was simulated for Advanced LIGO O3 sensitivity and the Einstein Telescope (ET) design sensitivity. The detection pipeline targets signal data out of the noisy detector background. Its architecture comprises of simulated whitened data as input to three classes of deep neural networks algorithms: a 1D and a 2D convolutional neural network (CNN) and a Long Short Term Memory (LSTM) network. It was found that all three algorithms could distinguish signals from noise with high precision for the ET sensitivity, but the current sensitivity of LIGO is too low to permit the algorithms to learn signal features from the input vectors. Andrei Utina, Francesco Marangio, Filip Morawski, Alberto Iess, Tania Regimbau, Giuseppe Fiameni, Elena Cuoco |
CBMI | 6 |
| 2015 | Interoperability Oriented Architecture: The Approach of EPOS for Solid Earth e-InfrastructuresabstractEPOS is an e-Infrastructure for solid Earh science in Europe. It integrates many heterogeneous Research Infrastructures (RIs) using a novel approach based on the harmonization of existing service and component interfaces. EPOS is designed to provide an architectural framework for new Research Infrastructures in the domain, and to interface with increasing sophistication of existing RIs working with them in co-development from their present state to a future integrated state. The key is the metadata catalogue based on CERIF which provides the virtualization required for EPOS to provide a homogeneous view over the heterogeneity. Architectural concepts together with a plan for integration and collaboration with EPOS nodes in order to interoperate are presented in this paper. Daniele Bailo, Keith G. Jeffery, Alessandro Spinuso, Giuseppe Fiameni |
e-Science | 4 |