VLDB 2026 Research / reviewers in the wild / expert
Roberto Amoroso
dblp:279/5594
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-1033-2485ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perceive. Query & Reason: Enhancing Video QA with Question-Guided Temporal QueriesabstractVideo Question Answering (Video QA) is a challenging video understanding task that requires models to compre-hend entire videos, identify the most relevant information based on contextual cues from a given question, and rea-son accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have trans-formed video QA by leveraging their exceptional common-sense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an ad-ditional space-time alignment poses a considerable chal-lenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA bench-marks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with re-cent advancements in video QA. Roberto Amoroso, Gengyuan Zhang, Rajat Koner, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp |
WACV | 1 |
| 2025 | Learning to mask and permute visual tokens for Vision Transformer pre-trainingabstractThe use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks. In this context, recent approaches have employed the Masked Image Modeling paradigm, which pre-trains a backbone by reconstructing visual tokens associated with randomly masked image patches. This masking approach, however, introduces noise into the input data during pre-training, leading to discrepancies that can impair performance during the fine-tuning phase. Furthermore, input masking neglects the dependencies between corrupted patches, increasing the inconsistencies observed in downstream fine-tuning tasks. To overcome these issues, we propose a new self-supervised pre-training approach, named Masked and Permuted Vision Transformer (MaPeT), that employs autoregressive and permuted predictions to capture intra-patch dependencies. In addition, MaPeT employs auxiliary positional information to reduce the disparity between the pre-training and fine-tuning phases. In our experiments, we employ a fair setting to ensure reliable and meaningful comparisons and conduct investigations on multiple visual tokenizers, including our proposed k -CLIP which directly employs discretized CLIP features. Our results demonstrate that MaPeT achieves competitive performance on ImageNet, compared to baselines and competitors under the same model setting. We release an implementation of our code and models at https://github.com/aimagelab/MaPeT . Lorenzo Baraldi 0002, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Andrea Pilzer, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2025 | Parents and Children: Distinguishing Multimodal Deepfakes from Natural ImagesabstractRecent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively, extracted from CLIP-based models and ResNet or Vision Transformer (ViT)-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2 million images generated from the original COCO image–caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0. Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi 0001, Alberto Del Bimbo, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype GenerationabstractOpen-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Pre-vious works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However, captions provide global information about the semantics of a given image but lack direct localization of individual concepts. Further, training on large-scale datasets inevitably brings significant computational costs. In this paper, we propose FreeDA, a training-free diffusion-augmented method for open-vocabulary semantic segmentation, which leverages the ability of diffusion models to visually localize generated concepts and local-global similarities to match class-agnostic regions with semantic classes. Our approach involves an offline stage in which textual-visual reference embeddings are collected, starting from a large set of captions and leveraging visual and semantic contexts. At test time, these are queried to support the visual matching process, which is carried out by jointly considering class-agnostic regions and global semantic similarities. Extensive analyses demonstrate that FreeDA achieves state-of-the-art performance on five datasets, surpassing previous methods by more than 7.0 average points in terms of mIoU and without requiring any training. Our source code is available at aimagelab.github. io/freeda. Luca Barsellotti, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 2 |
| 2024 | FOSSIL: Free Open-Vocabulary Semantic Segmentation through Synthetic References RetrievalabstractUnsupervised Open-Vocabulary Semantic Segmentation aims to segment an image into regions referring to an arbitrary set of concepts described by text, without relying on dense annotations that are available only for a subset of the categories. Previous works rely on inducing pixel-level alignment in a multi-modal space through contrastive training over vast corpora of image-caption pairs. However, representing a semantic category solely through its textual embedding is insufficient to encompass the wide-ranging variability in the visual appearances of the images associated with that category. In this paper, we propose FOSSIL, a pipeline that enables a self-supervised backbone to perform open-vocabulary segmentation relying only on the visual modality. In particular, we decouple the task into two components: (1) we leverage text-conditioned diffusion models to generate a large collection of visual embeddings, starting from a set of captions. These can be retrieved at inference time to obtain a support set of references for the set of textual concepts. Further, (2) we exploit self-supervised dense features to partition the image into semantically coherent regions. We demonstrate that our approach provides strong performance on different semantic segmentation datasets, without requiring any additional training. Luca Barsellotti, Roberto Amoroso, Lorenzo Baraldi 0001, Rita Cucchiara |
WACV | 2 |
| 2024 | What's Outside the Intersection? Fine-grained Error Analysis for Semantic Segmentation Beyond IoUabstractSemantic segmentation represents a fundamental task in computer vision with various application areas such as autonomous driving, medical imaging, or remote sensing. For evaluating and comparing semantic segmentation models, the mean intersection over union (mIoU) is currently the gold standard. However, while mIoU serves as a valuable benchmark, it does not offer insights into the types of errors incurred by a model. Moreover, different types of errors may have different impacts on downstream applications. To address this issue, we propose an intuitive method for the systematic categorization of errors, thereby enabling a fine-grained analysis of semantic segmentation models. Since we assign each erroneous pixel to precisely one error type, our method seamlessly extends the popular IoU-based evaluation by shedding more light on the false positive and false negative predictions. Our approach is model- and dataset-agnostic, as it does not rely on additional information besides the predicted and ground-truth segmentation masks. In our experiments, we demonstrate that our method accurately assesses model strengths and weaknesses on a quantitative basis, thus reducing the dependence on time-consuming qualitative model inspection. We analyze a variety of state-of-the-art semantic segmentation models, revealing systematic differences across various architectural paradigms. Exploiting the gained insights, we showcase that combining two models with complementary strengths in a straightforward way is sufficient to consistently improve mIoU, even for models setting the current state of the art on ADE20K. We release a toolkit for our evaluation method at https://github.com/mxbh/beyond-iou. Maximilian Bernhard, Roberto Amoroso, Yannic Kindermann, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp, Matthias Schubert |
WACV | 2 |
| 2023 | Superpixel Positional Encoding to Improve ViT-based Semantic Segmentation Models
Roberto Amoroso, Matteo Tomei, Lorenzo Baraldi 0001, Rita Cucchiara |
BMVC | 1 |
| 2023 | A Federated Learning Approach to Traffic Matrix Estimation using Super-resolution TechniquesabstractNetwork measurement and telemetry techniques are central to the management of modern computer networks. Traffic matrix estimation is a popular technique that supports several applications. Existing approaches use statistical methods, which often make invalid assumptions about the structure of the traffic matrix. Data-driven methods, instead, leverage detailed information about the network topology that may be unavailable or impractical to collect. In this work, we propose a super-resolution technique for traffic matrix estimation that can infer fine-grained network traffic. In our experiment, we demonstrate that the proposed approach with high precision outperforms existing data interpolation techniques. We also expand our design by employing a federated learning model to address scalability and improve performance. We find that our model increases the accuracy of the inference with respect to its centralized counterpart. Roberto Amoroso, Lorenzo Pappone, Flavio Esposito |
CCNC | 1 |
| 2021 | Assessing the Role of Boundary-Level Objectives in Indoor Semantic Segmentation
Roberto Amoroso, Lorenzo Baraldi 0001, Rita Cucchiara |
CAIP (1) | 1 |
| 2020 | Estimation of traffic matrices via super-resolution and federated learningabstractNetwork measurement and telemetry techniques are central to the management of today's computer networks. One popular technique with several applications is the estimation of traffic matrices. Existing traffic matrix inference approaches that use statistical methods, often make assumptions on the structure of the matrix that may be invalid. Data-driven methods, instead, often use detailed information about the network topology that may be unavailable or impractical to collect. Roberto Amoroso, Flavio Esposito, Maria Luisa Merani |
CoNEXT | 1 |