VLDB 2026 Research / reviewers in the wild / expert
Mario Valerio Giuffrida
dblp:179/2301
· DBLP profile ↗
17ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-5232-677XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRUST: Leveraging Text Robustness for Unsupervised Domain AdaptationabstractRecent unsupervised domain adaptation (UDA) methods have shown great success in addressing classical domain shifts (e.g., synthetic-to-real), but they still suffer under complex shifts (e.g. geographical shift), where both the background and object appearances differ significantly across domains. Prior works showed that the language modality can help in the adaptation process, exhibiting more robustness to such complex shifts. In this paper, we introduce TRUST, a novel UDA approach that exploits the robustness of the language modality to guide the adaptation of a vision model. TRUST generates pseudo-labels for target samples from their captions and introduces a novel uncertainty estimation strategy that uses normalised CLIP similarity scores to estimate the uncertainty of the generated pseudo-labels. Such estimated uncertainty is then used to reweight the classification loss, mitigating the adverse effects of wrong pseudo-labels obtained from low-quality captions. To further increase the robustness of the vision model, we propose a multimodal soft-contrastive learning loss that aligns the vision and language feature spaces, by leveraging captions to guide the contrastive training of the vision model on target images. In our contrastive loss, each pair of images acts as both a positive and a negative pair and their feature representations are attracted and repulsed with a strength proportional to the similarity of their captions. This solution avoids the need for hardly determining positive and negative pairs, which is critical in the UDA setting. Our approach outperforms previous methods, setting the new state-of-the-art on classical (DomainNet) and complex (GeoNet) domain shifts. The code is available at https://github.com/MattiaLitrico/TRUST-Leveraging-Text-Robustness-for-Unsupervised-Domain-Adaptation. Mattia Litrico, Mario Valerio Giuffrida, Sebastiano Battiato, Devis Tuia |
AAAI | 2 |
| 2026 | Uncertainty-guided Open-Set Source-Free Unsupervised Domain Adaptation with Target-private Class SegregationabstractStandard Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target, requiring simultaneous access to both source and target data. Moreover, UDA approaches commonly assume that source and target domains share the same labels space. Yet, these two assumptions are hardly satisfied in real-world scenarios. This paper considers the more challenging Source-Free Open-set Domain Adaptation (SF-OSDA) setting, where both assumptions do not hold. We propose a novel approach for SF-OSDA that takes advantage of the granularity of target-private categories by segregating their samples into multiple unknown classes. Starting from an initial clustering-based pseudo-labels initialisation, our method progressively improves the segregation of target-private samples by refining their pseudo-labels with the guide of an uncertainty-based sample selection module. Additionally, we propose a novel contrastive loss, named NL-InfoNCELoss, that, integrating negative learning into self-supervised contrastive learning, enhances the model robustness to noisy pseudo-labels. Extensive experiments on benchmark datasets demonstrate the superiority of our proposed approach over competing methods, establishing new state-of-the-art performance. Notably, additional analyses show that our method is able to learn the underlying semantics of novel classes, opening the possibility to perform novel class discovery. Mattia Litrico, Davide Talon, Sebastiano Battiato, Alessio Del Bue, Mario Valerio Giuffrida, Pietro Morerio |
Int. J. Comput. Vis. | 5 |
| 2026 | Exocentric-to-Egocentric Adaptation for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
Camillo Quattrocchi, Antonino Furnari, Daniele Di Mauro, Mario Valerio Giuffrida, Giovanni Maria Farinella |
Int. J. Comput. Vis. | 4 |
| 2026 | Count2Density: Crowd density estimation without location-level annotationsabstractCrowd density estimation is a well-known computer vision task aimed at estimating the density distribution of people in an image. The primary challenge in this domain is the reliance on fine-grained location-level annotations — i.e., points placed on top of each individual — to train deep networks. Collecting such detailed annotations is both tedious, time-consuming, and poses a significant barrier to scalability for real-world applications. To alleviate this burden, we present Count2Density : a novel pipeline designed to predict meaningful density maps containing quantitative spatial information using only count-level annotations (i.e., the total number of people) during training. To achieve this, Count2Density generates pseudo-density maps leveraging past predictions stored in a Historical Map Bank, thereby reducing confirmation bias. This bank is initialised using an unsupervised saliency estimator to provide an initial spatial prior and is iteratively updated with an Exponential Moving Average of predicted density maps. These pseudo-density maps are obtained by sampling locations from estimated crowd areas using a hypergeometric distribution, with the number of samplings determined by the count-level annotations. To further enhance the spatial awareness of the model and promote robust feature learning, we add a self-supervised contrastive spatial regulariser to encourage similar feature representations within crowded regions while maximising dissimilarity with background regions. Experimental results demonstrate that our approach significantly outperforms cross-domain adaptation methods and achieves better results than recent state-of-the-art approaches in semi-supervised settings across several datasets. Additional analyses validate the effectiveness of each individual component of our pipeline, including the self-supervised contrastive regulariser, confirming the ability of Count2Density to effectively retrieve spatial information from count-level annotations and enabling accurate subregion counting. Mattia Litrico, Michael P. Pound, Sotirios A. Tsaftaris, Sebastiano Battiato, Mario Valerio Giuffrida |
Pattern Recognit. | 6 |
| 2025 | GMT: Guided Mask Transformer for Leaf Instance SegmentationabstractLeaf instance segmentation is a challenging multi-instance segmentation task, aiming to separate and delin-eate each leaf in an image of a plant. Accurate segmentation of each leaf is crucial for plant-related applications such as the fine-grained monitoring of plant growth and crop yield estimation. This task is challenging because of the high similarity (in shape and colour), great size variation, and heavy occlusions among leaf instances. Furthermore, the typically small size of annotated leaf datasets makes it more difficult to learn the distinctive features needed for precise segmentation. We hypothesise that the key to overcoming the these challenges lies in the specific spatial patterns of leaf distri-bution. In this paper, we propose the Guided Mask Trans-former (GMT), which leverages and integrates leaf spa-tial distribution priors into a Transformer-based segmentor. These spatial priors are embedded in a set of guide functions that map leaves at different positions into a more sepa-rable embedding space. Our GMT consistently outperforms the state-of-the-art on three public plant datasets. Our code is available at https://github.com/vios-s/gmt-leaf-ins-seg. Sotirios A. Tsaftaris, Mario Valerio Giuffrida |
WACV | 3 |
| 2025 | Benchmarking computer vision architectures for cloud detection from lidar ceilometer backscatter dataabstractAbstract Cloud detection is fundamental for accurate weather monitoring, often achieved through remote sensing technology, such as satellite imagery or radar. This study explores the use of lidar ceilometer backscatter data, a rich but noisy source of atmospheric information, to enhance cloud detection. Leveraging data acquired from a Lufft CHM 15k ceilometer over three months near Mount Etna, Italy, we gathered a novel dataset comprising time-height plots derived from backscatter profiles. The Weather Research and Forecasting (WRF) model was used for ground-truth data labeling, ensuring reliable model validation. We benchmarked state-of-the-art deep learning architectures, including CNN-based models (e.g., ResNet50, VGG16, InceptionV3, EfficientNet) and the Vision Transformer (ViT), on our collected dataset. Among these, ResNet50 achieved the highest accuracy ( $$89.57\%$$ 89.57 % ), closely followed by ViT ( $$89.36\%$$ 89.36 % ), showcasing the efficacy of residual learning and transformer-based approaches in extracting complex patterns from atmospheric data. Our results highlight the potential of lidar-based systems for accurate cloud detection, complementing other remote sensing technologies. Our work contributes to the field by introducing a publicly available dataset and providing comprehensive benchmarking results that establish a baseline for future research. This study also opens avenues for broader applications of ceilometer data, such as the detection of pollutants and other atmospheric phenomena. Our dataset is publicly available at https://zenodo.org/records/10616434 . Alessio Barbaro Chisari, Luca Guarnera, Alessandro Ortis, Wladimiro Carlo Patatu, Sebastiano Battiato, Mario Valerio Giuffrida |
Vis. Comput. | 6 |
| 2024 | Synchronization Is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
Camillo Quattrocchi, Antonino Furnari, Daniele Di Mauro, Mario Valerio Giuffrida, Giovanni Maria Farinella |
ECCV (72) | 4 |
| 2024 | Federated Learning in a Semi-Supervised Environment for Earth Observation DataabstractWe propose FedRec, a federated learning workflow taking advantage of unlabelled data in a semi-supervised environment to assist in the training of a supervised aggregated model.In our proposed method, an encoder architecture extracting features from unlabelled data is aggregated with the feature extractor of a classification model via weight averaging.The fully connected layers of the supervised models are also averaged in a federated fashion.We show the effectiveness of our approach by comparing it with the state-of-the-art federated algorithm, an isolated and a centralised baseline, on novel cloud detection datasets.Our code is available at https://github.com/CasellaJr/FedRec.* This work has been partly supported by the Spoke "FutureHPC & BigData" of the ICSC -Centro Nazionale di Ricerca in "High Performance Computing, Big Data and Quantum Computing", funded by European Union - Bruno Casella, Alessio Barbaro Chisari, Marco Aldinucci, Sebastiano Battiato, Mario Valerio Giuffrida |
ESANN | 5 |
| 2024 | On the Cloud Detection from Backscattered Images Generated from a Lidar-Based Ceilometer: Current State and OpportunitiesabstractAccurate weather monitoring depends significantly on cloud detection, a crucial process achievable through remote sensing tools such as satellite imagery and radar or through the analysis of data obtained from ceilometers. A ceilometer is a lidar-based device allowing to analyse the atmosphere and detect the presence of particles within clouds. The data retrieved from ceilometers involve analysis of the backscatter of the lidar signal returning to the surface. Given the inherent noise in this data, we leverage deep learning models to detect the presence of clouds in the data. To label the data, we take advantage of a Weather Research & Forecasting (WRF) model, which provided us with ground-truth used for validation purposes. We performed a comparative analysis with current state-of-the-art deep learning architectures on this specialist domain. This comparative analysis shows that the best model is ResNet 50, but also a transformer-based model, such as ViT, achieves great results. These preliminary results pave the scenario for future works aimed at detecting other particles composing the atmosphere, such as polluting agents that can be detected from the ceilometer backscatter data. Alessio Barbaro Chisari, Alessandro Ortis, Luca Guarnera, Wladimiro Carlo Patatu, Rosaria Ausilia Giandolfo, Emanuele Spampinato, Sebastiano Battiato, Mario Valerio Giuffrida |
ICIP | 8 |
| 2024 | TADM: Temporally-Aware Diffusion Model for Neurodegenerative Progression on Brain MRI
Mattia Litrico, Francesco Guarnera, Mario Valerio Giuffrida, Daniele Ravì, Sebastiano Battiato |
MICCAI (2) | 3 |
| 2023 | TouchEnc: a Novel Behavioural Encoding Technique to Enable Computer Vision for Continuous Smartphone User AuthenticationabstractWe are increasingly required to prove our identity when using smartphones through explicit authentication processes such as passwords or physiological biometrics, e.g., authorising online banking transactions or unlocking smartphones. However, these methods are often annoying to input and do not guarantee that the genuine user remains the same. Thus, a modern verification process should differ from traditional authentication. In touch-based biometrics, a new approach must not verify what we draw but how we draw it. Our research proposes TouchEnc, a Deep Learning approach that outperforms conventional methods. Unlike Machine Learning methods, TouchEnc automates the feature extraction from touch gestures. TouchEnc achieves this by transforming and encoding touch behaviour into images, enabling continuous authentication through modern computer vision. Our approach has been tested on a popular and publicly available dataset to demonstrate its effectiveness. Results show that users can authenticate using TouchEnc with a single gesture containing users’ on-screen navigational behaviour, independent of drawing up, down, left, or right. TouchEnc achieves an 8.4% Equal Error Rate and a 96.7% Area Under the Curve using a single gesture. Furthermore, TouchEnc achieves up to 65% better Equal Error Rates when combining gestures compared to the related work. Peter Aaby, Mario Valerio Giuffrida, William J. Buchanan, Zhiyuan Tan 0001 |
TrustCom | 2 |
| 2023 | An omnidirectional approach to touch-based continuous authenticationabstractThis paper focuses on how touch interactions on smartphones can provide a continuous user authentication service through behaviour captured by a touchscreen. While efforts are made to advance touch-based behavioural authentication, researchers often focus on gathering data, tuning classifiers, and enhancing performance by evaluating touch interactions in a sequence rather than independently. However, such systems only work by providing data representing distinct behavioural traits. The typical approach separates behaviour into touch directions and creates multiple user profiles. This work presents an omnidirectional approach which outperforms the traditional method independent of the touch direction - depending on optimal behavioural features and a balanced training set. Thus, we evaluate five behavioural feature sets using the conventional approach against our direction-agnostic method while testing several classifiers, including an Extra-Tree and Gradient Boosting Classifier, which is often overlooked. Results show that in comparison with the traditional, an Extra-Trees classifier and the proposed approach are superior when combining strokes. However, the performance depends on the applied feature set. We find that the TouchAlytics feature set outperforms others when using our approach when combining three or more strokes. Finally, we highlight the importance of reporting the mean area under the curve and equal error rate for single-stroke performance and varying the sequence of strokes separately. Peter Aaby, Mario Valerio Giuffrida, William J. Buchanan, Zhiyuan Tan 0001 |
Comput. Secur. | 2 |
| 2020 | Unsupervised Rotation Factorization in Restricted Boltzmann MachinesabstractFinding suitable image representations for the task at hand is critical in computer vision. Different approaches extending the original Restricted Boltzmann Machine (RBM) model have recently been proposed to offer rotation-invariant feature learning. In this paper, we present an extended novel RBM that learns rotation invariant features by explicitly factorizing for rotation nuisance in 2D image inputs within an unsupervised framework. While the goal is to learn invariant features, our model infers an orientation per input image during training, using information related to the reconstruction error. The training process is regularised by a Kullback-Leibler divergence, offering stability and consistency. We used the γ -score, a measure that calculates the amount of invariance, to mathematically and experimentally demonstrate that our approach indeed learns rotation invariant features. We show that our method outperforms the current state-of-the-art RBM approaches for rotation invariant feature learning on three different benchmark datasets, by measuring the performance with the test accuracy of an SVM classifier. Our implementation is available at https://bitbucket.org/tuttoweb/rotinvrbm. Mario Valerio Giuffrida, Sotirios A. Tsaftaris |
IEEE Trans. Image Process. | 1 |
| 2018 | Root Gap Correction with a Deep Inpainting Model
Hao Chen 0102, Mario Valerio Giuffrida, Sotirios A. Tsaftaris, Peter Doerner |
BMVC | 2 |
| 2018 | Multimodal MR Synthesis via Modality-Invariant Latent RepresentationabstractWe propose a multi-input multi-output fully convolutional neural network model for MRI synthesis. The model is robust to missing data, as it benefits from, but does not require, additional input modalities. The model is trained end-to-end, and learns to embed all input modalities into a shared modality-invariant latent space. These latent representations are then combined into a single fused representation, which is transformed into the target output modality with a learnt decoder. We avoid the need for curriculum learning by exploiting the fact that the various input modalities are highly correlated. We also show that by incorporating information from segmentation masks the model can both decrease its error and generate data with synthetic lesions. We evaluate our model on the ISLES and BRATS data sets and demonstrate statistically significant improvements over state-of-the-art methods for single input tasks. This improvement increases further when multiple input modalities are used, demonstrating the benefits of learning a common latent space, again resulting in a statistically significant improvement over the current best method. Finally, we demonstrate our approach on non skull-stripped brain images, producing a statistically significant improvement over the previous best method. Code is made publicly available at https://github.com/agis85/multimodal_brain_synthesis. Agisilaos Chartsias, Thomas Joyce, Mario Valerio Giuffrida, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Rotation-Invariant Restricted Boltzmann Machine Using Shared Gradient Filters
Mario Valerio Giuffrida, Sotirios A. Tsaftaris |
ICANN (2) | 1 |
| 2015 | On Blind Source Camera Identification
Giovanni Maria Farinella, Mario Valerio Giuffrida, V. Digiacomo, Sebastiano Battiato |
ACIVS | 2 |