Sebastiano Vascon

dblp:133/9513 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-7855-1641ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 An Immersive Virtual Reality Interface for Archaeological Reconstruction
abstract
Archaeological reconstruction of fragmented artifacts presents a complex computational challenge that requires both algorithmic sophistication and expert domain knowledge. We present an immersive Virtual Reality interface that bridges our reconstruction solver with public engagement and expert annotation. Built within the RePAIR project, our system enables users to interact with high-fidelity 3D scans of 2000-year-old fresco fragments from Pompeii’s House of the Painters, attempting manual reconstruction while leveraging AI assistance. Deployed at the Italian Pavilion during EXPO 2025, the system demonstrated both the intrinsic difficulty of archaeological reconstruction (low completion rates) and the potential for human-AI collaboration. Beyond public engagement, our interface serves potentially as a research platform for collecting expert annotations, benchmarking solver performance under varied initial conditions, and generating training data for future reinforcement learning approaches. Our work demonstrates how gamification and immersive technologies can simultaneously democratize cultural heritage and advance research methodologies.
Luca Palmieri 0002, Omidreza Safaei, Marco Ronchese, Marina Khoroshiltseva, Sebastiano Vascon, Marcello Pelillo
AVI5
2026 MMAF: Multimodal Attention Fusion for Molecular Toxicity Prediction
Faiz Ur Rehman, Muhammad Rameez Ur Rahman, Sebastiano Vascon, Marcello Pelillo
ICPR (3)3
2026 Evaluating the robustness of explainable AI in medical image recognition under natural and adversarial data corruption
abstract
Abstract The integration of Explainable AI (XAI) into healthcare promises greater transparency and interpretability of machine learning models, enabling clinicians to understand predictions and make more reliable medical decisions. Yet, the robustness of XAI methods remains uncertain, as small input perturbations can drastically change their explanations, posing critical risks in clinical settings where they may lead to misdiagnoses or inappropriate treatment. Motivated by the central role of XAI in healthcare decision-making, this paper examines its robustness in the presence of data corruption. We systematically evaluate the stability of widely used XAI techniques against both naturally occurring noise (e.g., JPEG compression) and adversarial manipulations that alter explanations without affecting model predictions. To this end, we introduce a set of evaluation metrics that capture complementary aspects of explanation stability, ranging from pixel-level consistency to spatial coherence, and propose a protocol for assessing the resilience of XAI methods across diverse perturbation sources. Our analysis spans three medical imaging datasets, various convolutional and transformer models, and ten post-hoc XAI methods, including Grad-CAM++ for convolutional networks and LibraGrad for vision transformers. We find that current XAI techniques are often unstable, even under imperceptible perturbations. For adversarial noise, a clear set of robust methods emerges, whereas for natural noise, performance varies, with some methods maintaining spatial stability and others preserving pixel-wise consistency. All results together highlight the need for multi-perspective evaluation when selecting XAI techniques in practice.
Sara Repetto, Igor Maljkovic, Michele Lotto, Antonio Emanuele Cinà, Sebastiano Vascon, Fabio Roli
Mach. Learn.5
2026 Multi-view graph pooling via dominant sets for graph classification
abstract
• Dominant Set Multi-View Pooling to integrate topology, coarser graphs, and features. • Develop a dominant set pooling using edge weights to identify all potential clusters. • Design a fusion view attention layer to fuse coarser graphs, topology, and features. • Experimental results show DSMVPool outperforms state-of-the-art on eight benchmarks. Graph pooling is a fundamental operation in Graph Neural Networks (GNNs), designed to simplify graphs by reducing the number of nodes and edges while preserving essential structural information for classification tasks. However, most existing pooling methods tend to overlook edge weights and rely on a single-view pooling strategy that focuses either on local or global topological information, failing to capture the full structural context of the graph. To address these limitations, this study introduces a novel Dominant Set Multi-View Pooling (DSMVPool) method featuring two main contributions. First, we propose a dominant-set cluster pooling approach that analyzes the overall graph architecture and connectivity patterns, identifies potential clusters using edge weight information, and generates a coarser graph view. In addition, we create two complementary pooled views by selecting the most representative nodes based on local topology and node features. Second, we design a fusion-view attention layer that integrates the coarser graph structure with the pooled graph views, enabling our method to simultaneously capture and combine global and local structural information and node features. Extensive experiments on four graph classification benchmarks, covering computer vision, chemical, biological, and social networks, demonstrate that DSMVPool achieves superior performance compared to state-of-the-art methods.
Sebastiano Vascon, Thilo Stadelmann, Marcello Pelillo
Pattern Recognit.2
2025 ECAM: A Contrastive Learning Approach to Avoid Environmental Collision in Trajectory Forecasting
abstract
Human trajectory forecasting is crucial in applications such as autonomous driving, robotics and surveillance. Accurate forecasting requires models to consider various factors, including social interactions, multi-modal predictions, pedestrian intention and environmental context. While existing methods account for these factors, they often overlook the impact of the environment, which leads to collisions with obstacles. This paper introduces ECAM (Environmental Collision Avoidance Module), a contrastive learning-based module to enhance collision avoidance ability with the environment. The proposed module can be integrated into existing trajectory forecasting models, improving their ability to generate collision-free predictions. We evaluate our method on the ETH/UCY dataset and quantitatively and qualitatively demonstrate its collision avoidance capabilities. Our experiments show that state-of-the-art methods significantly reduce (-40/50%) the collision rate when integrated with the proposed module. The code is available at https://github.com/CVML-CFU/ECAM.
Giacomo Rosin, Muhammad Rameez Ur Rahman, Sebastiano Vascon
IJCNN3
2025 Automatizing 3D reconstruction pipelines for speeding-up cultural heritage digitization
Gianluca Bison, Luca Palmieri 0002, Sinem Aslan, Sebastiano Vascon, Marcello Pelillo
Multim. Tools Appl.4
2024 Nash Meets Wertheimer: Using Good Continuation in Jigsaw Puzzles
Marina Khoroshiltseva, Luca Palmieri 0002, Sinem Aslan, Sebastiano Vascon, Marcello Pelillo
ACCV (6)4
2024 Reassembling Broken Objects Using Breaking Curves
Ali Alagrami, Luca Palmieri 0002, Sinem Aslan, Marcello Pelillo, Sebastiano Vascon
ICPR (18)5
2024 Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solving
abstract
This paper proposes the RePAIR dataset that represents a challenging benchmark to test modern computational and data driven methods for puzzle-solving and reassembly tasks. Our dataset has unique properties that are uncommon to current benchmarks for 2D and 3D puzzle solving. The fragments and fractures are realistic, caused by a collapse of a fresco during a World War II bombing at the Pompeii archaeological park. The fragments are also eroded and have missing pieces with irregular shapes and different dimensions, challenging further the reassembly algorithms. The dataset is multi-modal providing high resolution images with characteristic pictorial elements, detailed 3D scans of the fragments and meta-data annotated by the archaeologists. Ground truth has been generated through several years of unceasing fieldwork, including the excavation and cleaning of each fragment, followed by manual puzzle solving by archaeologists of a subset of approx. 1000 pieces among the 16000 available. After digitizing all the fragments in 3D, a benchmark was prepared to challenge current reassembly and puzzle-solving methods that often solve more simplistic synthetic scenarios. The tested baselines show that there clearly exists a gap to fill in solving this computationally complex problem.
Theodore Tsesmelis, Luca Palmieri 0002, Marina Khoroshiltseva, Adeela Islam, Gur Elkin, Ofir Itzhak Shahar, Gianluca Scarpellini, Stefano Fiorini, Yaniv Ohayon, Nadav Alali, Sinem Aslan, Pietro Morerio, Sebastiano Vascon, Elena Gravina, Maria Cristina Napolitano, Giuseppe Scarpati, Gabriel Zuchtriegel, Alexandra Spühler, Michel E. Fuchs, Stuart James, Ohad Ben-Shahar, Marcello Pelillo, Alessio Del Bue
NeurIPS13
2024 Hierarchical Glocal Attention Pooling for Graph Classification
Sebastiano Vascon, Thilo Stadelmann, Marcello Pelillo
Pattern Recognit. Lett.2
2023 The Group Loss++: A Deeper Look Into Group Loss for Deep Metric Learning
abstract
Deep metric learning has yielded impressive results in tasks such as clustering and image retrieval by leveraging neural networks to obtain highly discriminative feature embeddings, which can be used to group samples into different classes. Much research has been devoted to the design of smart loss functions or data mining strategies for training such networks. Most methods consider only pairs or triplets of samples within a mini-batch to compute the loss function, which is commonly based on the distance between embeddings. We propose Group Loss, a loss function based on a differentiable label-propagation method that enforces embedding similarity across all samples of a group while promoting, at the same time, low-density regions amongst data points belonging to different groups. Guided by the smoothness assumption that "similar objects should belong to the same group", the proposed loss trains the neural network for a classification task, enforcing a consistent labelling amongst samples within a class. We design a set of inference strategies tailored towards our algorithm, named Group Loss++ that further improve the results of our model. We show state-of-the-art results on clustering and image retrieval on four retrieval datasets, and present competitive results on two person re-identification datasets, providing a unified framework for retrieval and re-identification.
Ismail Elezi, Jenny Seidenschwarz, Laurin Wagner, Sebastiano Vascon, Alessandro Torcinovich, Marcello Pelillo, Laura Leal-Taixé
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Locality-aware subgraphs for inductive link prediction in knowledge graphs
abstract
Recent methods for inductive reasoning on Knowledge Graphs (KGs) transform the link prediction problem into a graph classification task. They first extract a subgraph around each target link based on the k-hop neighborhood of the target entities, encode the subgraphs using a Graph Neural Network (GNN), then learn a function that maps subgraph structural patterns to link existence. Although these methods have witnessed great successes, increasing k often leads to an exponential expansion of the neighborhood, thereby degrading the GNN expressivity due to oversmoothing. In this paper, we formulate the subgraph extraction as a local clustering procedure that aims at sampling tightly-related subgraphs around the target links, based on a personalized PageRank (PPR) approach. Empirically, on three real-world KGs, we show that reasoning over subgraphs extracted by PPR-based local clustering can lead to a more accurate link prediction model than relying on neighbors within fixed hop distances. Furthermore, we investigate graph properties such as average clustering coefficient and node degree, and show that there is a relation between these and the performance of subgraph-based link prediction.
Hebatallah A. Mohamed Hassan 0001, Diego Pilutti, Stuart James, Alessio Del Bue, Marcello Pelillo, Sebastiano Vascon
Pattern Recognit. Lett.6
2021 The Hammer and the Nut: Is Bilevel Optimization Really Needed to Poison Linear Classifiers?
abstract
One of the most concerning threats for modern AI systems is data poisoning, where the attacker injects maliciously crafted training data to corrupt the system's behavior at test time. Availability poisoning is a particularly worrisome subset of poisoning attacks where the attacker aims to cause a Denial-of-Service (DoS) attack. However, the state-of-the-art algorithms are computationally expensive because they try to solve a complex bi-level optimization problem (the “hammer”). We observed that in particular conditions, namely, where the target model is linear (the “nut”), the usage of computationally costly procedures can be avoided. We propose a counter-intuitive but efficient heuristic that allows contaminating the training set such that the target system's performance is highly compromised. We further suggest a re-parameterization trick to decrease the number of variables to be optimized. Finally, we demonstrate that, under the considered settings, our framework achieves comparable, or even better, performances in terms of the attacker's objective while being significantly more computationally efficient.
Antonio Emanuele Cinà, Sebastiano Vascon, Ambra Demontis, Battista Biggio, Fabio Roli, Marcello Pelillo
IJCNN2
2021 Transductive Visual Verb Sense Disambiguation
abstract
Verb Sense Disambiguation is a well-known task in NLP, the aim is to find the correct sense of a verb in a sentence. Recently, this problem has been extended in a multimodal scenario, by exploiting both textual and visual features of ambiguous verbs leading to a new problem, the Visual Verb Sense Disambiguation (VVSD). Here, the sense of a verb is assigned considering the content of an image paired with it rather than a sentence in which the verb appears. Annotating a dataset for this task is more complex than textual disambiguation, because assigning the correct sense to a pair ofrequires both non-trivial linguistic and visual skills. In this work, differently from the literature, the VVSD task will be performed in a transductive semi-supervised learning (SSL) setting, in which only a small amount of labeled information is required, reducing tremendously the need for annotated data. The disambiguation process is based on a graph-based label propagation method which takes into account mono or multimodal representations forpairs. Experiments have been carried out on the recently published dataset VerSe, the only available dataset for this task. The achieved results outperform the current state-of-the-art by a large margin while using only a small fraction of labeled samples per sense1.
Sebastiano Vascon, Sinem Aslan, Gianluca Bigaglia, Lorenzo Giudice, Marcello Pelillo
WACV1
2020 The Group Loss for Deep Metric Learning
Ismail Elezi, Sebastiano Vascon, Alessandro Torcinovich, Marcello Pelillo, Laura Leal-Taixé
ECCV (7)2
2020 Encoding Brain Networks Through Geodesic Clustering of Functional Connectivity for Multiple Sclerosis Classification
abstract
An important task in brain connectivity research is the classification of patients from healthy subjects. In this work, we present a two-step mathematical framework allowing to discriminate between two groups of people with an application to multiple sclerosis. The proposed approach exploits the properties of the connectivity matrices determined using the covariances between signals of a fixed set of brain areas. These positive semidefinite matrices lay on a Riemannian manifold, allowing to use a geodesic distance defined on this space. In order to generate a vector representation useful for classification purposes, but still preserving the network structure, we encoded the data exploiting the network attractors determined by a geodesic clustering of connectivity matrices. Then clustering centroids were used as a dictionary allowing to encode subject's connectivity matrices as a vector of geodesic distances. A Linear Support Vector Machine was then used to perform classification between subjects. To demonstrate the advantage of using geodesic metrics in this framework, we conducted the same analysis using Euclidean metric. Experimental results validate the fact that employing geodesic metric in this framework leads to a higher classification performance, whereas performance with a Euclidean metric was sub-optimal.
Muhammad Abubakar Yamin, Paola Valsasina, Michael Dayan, Sebastiano Vascon, Jacopo Tessadori, Massimo Filippi, Vittorio Murino, Maria Assunta Rocca, Diego Sona
ICPR4
2020 Biclustering with dominant sets
Matteo Denitto, Manuele Bicego, Alessandro Farinelli, Sebastiano Vascon, Marcello Pelillo
Pattern Recognit.4
2020 Two sides of the same coin: Improved ancient coin classification using Graph Transduction Games
Sinem Aslan, Sebastiano Vascon, Marcello Pelillo
Pattern Recognit. Lett.2
2020 Protein function prediction as a graph-transduction game
Sebastiano Vascon, Marco Frasca 0001, Rocco Tripodi, Giorgio Valentini, Marcello Pelillo
Pattern Recognit. Lett.1
2019 Unsupervised Domain Adaptation using Graph Transduction Games
abstract
Unsupervised domain adaptation (UDA) amounts to assigning class labels to the unlabeled instances of a dataset from a target domain, using labeled instances of a dataset from a related source domain. In this paper we propose to cast this problem in a game-theoretic setting as a non-cooperative game and introduce a fully automatized iterative algorithm for UDA based on graph transduction games (GTG). The main advantages of this approach are its principled foundation, guaranteed termination of the iterative algorithms to a Nash equilibrium (which corresponds to a consistent labeling condition) and soft labels quantifying uncertainty of the label assignment process. We also investigate the beneficial effect of using pseudo-labels from linear classifiers to initialize the iterative process. The performance of the resulting methods is assessed on publicly available object recognition benchmark datasets involving both shallow and deep features. Results of experiments demonstrate the suitability of the proposed game-theoretic approach for solving UDA tasks.
Sebastiano Vascon, Sinem Aslan, Alessandro Torcinovich, Twan van Laarhoven, Elena Marchiori, Marcello Pelillo
IJCNN1
2019 Hypergraph isomorphism using association hypergraphs
Giulia Sandi, Sebastiano Vascon, Marcello Pelillo
Pattern Recognit. Lett.2
2018 Transductive Label Augmentation for Improved Deep Network Learning
abstract
A major impediment to the application of deep learning to real-world problems is the scarcity of labeled data. Small training sets are in fact of no use to deep networks as, due to the large number of trainable parameters, they will very likely be subject to overfitting phenomena. On the other hand, the increment of the training set size through further manual or semi-automatic labellings can be costly, if not possible at times. Thus, the standard techniques to address this issue are transfer learning and data augmentation, which consists of applying some sort of “transformation” to existing labeled instances to let the training set grow in size. Although this approach works well in applications such as image classification, where it is relatively simple to design suitable transformation operators, it is not obvious how to apply it in more structured scenarios. Motivated by the observation that in virtually all application domains it is easy to obtain unlabeled data, in this paper we take a different perspective and propose a label augmentation approach. We start from a small, curated labeled dataset and let the labels propagate through a larger set of unlabeled data using graph transduction techniques. This allows us to naturally use (second-order) similarity information which resides in the data, a source of information which is typically neglected by standard augmentation techniques. In particular, we show that by using known game theoretic transductive processes we can create larger and accurate enough labeled datasets which use results in better trained neural networks. Preliminary experiments are reported which demonstrate a consistent improvement over standard image classification datasets.
Ismail Elezi, Alessandro Torcinovich, Sebastiano Vascon, Marcello Pelillo
ICPR3
2018 Speaker Clustering Using Dominant Sets
abstract
Speaker clustering is the task of forming speaker-specific groups based on a set of utterances. In this paper, we address this task by using Dominant Sets (DS). DS is a graph-based clustering algorithm with interesting properties that fits well to our problem and has never been applied before to speaker clustering. We report on a comprehensive set of experiments on the TIMIT dataset against standard clustering techniques and specific speaker clustering methods. Moreover, we compare performances under different features by using ones learned via deep neural network directly on TIMIT and other ones extracted from a pre-trained VGGVox net. To asses the stability, we perform a sensitivity analysis on the free parameters of our method, showing that performance is stable under parameter changes. The extensive experimentation carried out confirms the validity of the proposed method, reporting state-of-the-art results under three different standard metrics. We also report reference baseline results for speaker clustering on the entire TIMIT dataset for the first time.
Feliks Hibraj, Sebastiano Vascon, Thilo Stadelmann, Marcello Pelillo
ICPR2
2016 Detecting emergent leader in a meeting environment using nonverbal visual features only
abstract
In this paper, we propose an effective method for emergent leader detection in meeting environments which is based on nonverbal visual features. Identifying emergent leader is an important issue for organizations. It is also a well-investigated topic in social psychology while a relatively new problem in social signal processing (SSP). The effectiveness of nonverbal features have been shown by many previous SSP studies. In general, the nonverbal video-based features were not more effective compared to audio-based features although, their fusion generally improved the overall performance. However, in absence of audio sensors, the accurate detection of social interactions is still crucial. Motivating from that, we propose novel, automatically extracted, nonverbal features to identify the emergent leadership. The extracted nonverbal features were based on automatically estimated visual focus of attention which is based on head pose. The evaluation of the proposed method and the defined features were realized using a new dataset which is firstly introduced in this paper including its design, collection and annotation. The effectiveness of the features and the method were also compared with many state of the art features and methods.
Cigdem Beyan, Nicolò Carissimi, Francesca Capozzi, Sebastiano Vascon, Matteo Bustreo, Antonio Pierro, Cristina Becchio, Vittorio Murino
ICMI4
2016 Context aware nonnegative matrix factorization clustering
abstract
In this article we propose a method to refine the clustering results obtained with the nonnegative matrix factorization (NMF) technique, imposing consistency constraints on the final labeling of the data. The research community focused its effort on the initialization and on the optimization part of this method, without paying attention to the final cluster assignments. We propose a game theoretic framework in which each object to be clustered is represented as a player, which has to choose its cluster membership. The information obtained with NMF is used to initialize the strategy space of the players and a weighted graph is used to model the interactions among the players. These interactions allow the players to choose a cluster which is coherent with the clusters chosen by similar players, a property which is not guaranteed by NMF, since it produces a soft clustering of the data. The results on common benchmarks show that our model is able to improve the performances of many NMF formulations.
Rocco Tripodi, Sebastiano Vascon, Marcello Pelillo
ICPR2
2016 Detecting conversational groups in images and sequences: A robust game-theoretic approach
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino
Comput. Vis. Image Underst.1
2014 A Game-Theoretic Probabilistic Approach for Detecting Conversational Groups
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino
ACCV (5)1
2012 A stable graph-based representation for object recognition through high-order matching
Andrea Albarelli, Filippo Bergamasco, Luca Rossi 0004, Sebastiano Vascon, Andrea Torsello
ICPR4