Samuele Salti

dblp:31/495 · DBLP profile ↗
← Back
48ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0001-5609-426XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 NVS-HO: A Benchmark for Novel View Synthesis of Handheld Objects
Musawar Ali, Manuel Carranza-García, Nicola Fioraio, Samuele Salti, Luigi Di Stefano
ICPR (2)4
2026 How to Evaluate and Refine Your CAM
Luca Domeniconi, Alessandra Stramiglio, Michele Lombardi 0001, Samuele Salti
ICPR (4)4
2026 The PRISM benchmark: PhotoRealistic Image Synthesis and Manipulation to detect generated images
abstract
The creation of photorealistic synthetic images or the alteration of existing footage is an emerging societal concern that has led researchers to investigate detection of generated content and the release of several benchmarks. However, images generated for the benchmarks are often unrealistic, the reason being twofold: images are generated from scratch using only text conditioning; the generative models used are not the current state-of-the-art. Moreover, resolutions and compression are often different between real and fake images, artificially simplifying the task in the benchmarks. In this paper, we propose PRISM , a new challenging benchmark for generated content detection designed to reflect the complexity of real-world visual data. PRISM includes images produced with recent, high-fidelity generative models and leverages image conditioning to ensure realism. Our benchmark features three levels of image alterations, from subtle manipulations of real images to full generation, while maintaining the same distribution of resolution and compression levels as real data. By enforcing train/test splits where generative models are unseen during training, PRISM provides a grounded and realistic setting for evaluating model generalization and robustness. In addition, we investigate the use of energy-based models and contrastive losses for detecting generated content, and we devise a two-stage training recipe, each stage involving different levels of alterations/generations. This simple method obtains remarkable performance on PRISM and other publicly available benchmarks and proves robust to confounding factors like image compression and resolution. The benchmark and model will be publicly released.
Filippo Bartolucci, Samuele Salti, Giuseppe Lisanti
Comput. Vis. Image Underst.2
2026 Additive decomposition of one-dimensional signals using Transformers
abstract
One-dimensional signal decomposition is a well-established and widely used technique across various scientific fields. It serves as a highly valuable pre-processing step for data analysis. While traditional decomposition techniques often rely on mathematical models, recent research suggests that applying the latest deep learning models to this very ill-posed inverse problem represents an exciting, unexplored area with promising potential. This work presents a novel method for the additive decomposition of one-dimensional signals. We leverage the Transformer architecture to decompose signals into their constituent components: piecewise constant, smooth (trend), highly-oscillatory, and noise components. Our model, trained on synthetic data, achieves excellent accuracy in modeling and decomposing input signals from the same distribution, as demonstrated by the experimental results. • We study additive decomposition of 1D signals with Transformers. • We define a neural architecture for the problem, based on the Transformer encoder. • Our method is more effective and orders of magnitude faster than variational ones. • The proposed method automatically detects the absence of a component.
Samuele Salti, Andrea Pinto, Alessandro Lanza, Serena Morigi
Pattern Recognit. Lett.1
2026 Domain Adaptation for Image Classification of Defects in Semiconductor Manufacturing
abstract
In the semiconductor sector, due to high demand but also strong and increasing competition, time to market and quality are key factors in securing significant market share in various application areas. Thanks to the success of deep learning methods in recent years in the computer vision domain, Industry 4.0 and 5.0 applications, such as defect classification, have achieved remarkable success. In particular, Domain Adaptation (DA) has proven highly effective since it focuses on using the knowledge learned on a (source) domain to adapt and perform effectively on a different but related (target) domain. By improving robustness and scalability, DA minimizes the need for extensive manual re-labeling or retraining of models. This not only reduces computational and resource costs but also allows human experts to focus on high-value tasks. Therefore, we tested the efficacy of DA techniques in semi-supervised and unsupervised settings within the context of the semiconductor field. Moreover, we propose the DBACS approach, a CycleGAN-inspired model enhanced with additional loss terms to improve performance. All the approaches are studied and validated on real-world Electron Microscope images, considering the unsupervised and semi-supervised settings, proving the usefulness of our method in advancing DA techniques for the semiconductor field.
Adrian Poniatowski, Natalie Gentner, Manuel Barusco, Davide Dalle Pezze, Samuele Salti, Gian Antonio Susto
IEEE Trans Autom. Sci. Eng.5
2025 Spatially-aware Weights Tokenization for NeRF-Language Models
abstract
Neural Radiance Fields (NeRFs) are neural networks -- typically multilayer perceptrons (MLPs) -- that represent the geometry and appearance of objects, with applications in vision, graphics, and robotics. Recent works propose understanding NeRFs with natural language using Multimodal Large Language Models (MLLMs) that directly process the weights of a NeRF's MLP. However, these approaches rely on a global representation of the input object, making them unsuitable for spatial reasoning and fine-grained understanding. In contrast, we propose **weights2space**, a self-supervised framework featuring a novel meta-encoder that can compute a sequence of spatial tokens directly from the weights of a NeRF. Leveraging this representation, we build **Spatial LLaNA**, a novel MLLM for NeRFs, capable of understanding details and spatial relationships in objects represented as NeRFs. We evaluate Spatial LLaNA on NeRF captioning and NeRF Q&A tasks, using both existing benchmarks and our novel **Spatial ObjaNeRF** dataset consisting of $100$ manually-curated language annotations for NeRFs. This dataset features 3D models and descriptions that challenge the spatial reasoning capability of MLLMs. Spatial LLaNA outperforms existing approaches across all tasks.
Andrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti, Luigi Di Stefano
NeurIPS4
2025 RendBEV: Semantic Novel View Synthesis for Self-Supervised Bird's Eye View Segmentation
abstract
Bird's Eye View (BEV) semantic maps have recently garnered a lot of attention as a useful representation of the environment to tackle assisted and autonomous driving tasks. However most of the existing work focuses on the fully supervised setting training networks on large annotated datasets. In this work we present RendBEV a new method for the self-supervised training of BEV semantic segmentation networks leveraging differentiable volumetric rendering to receive supervision from semantic perspective views computed by a 2D semantic segmentation model. Our method enables zero-shot BEV semantic segmentation and already delivers competitive results in this challenging setting. When used as pretraining to then fine-tune on labeled BEV ground truth our method significantly boosts performance in low-annotation regimes and sets a new state of the art when fine-tuning on all available labels.
Henrique Piñeiro Monteagudo, Leonardo Taccari, Aurel Pjetri, Francesco Sambo, Samuele Salti
WACV5
2024 Neural Processing of Tri-Plane Hybrid Neural Fields
abstract
Driven by the appealing properties of neural fields for storing and communicating 3D data, the problem of directly processing them to address tasks such as classification and part segmentation has emerged and has been investigated in recent works. Early approaches employ neural fields parameterized by shared networks trained on the whole dataset, achieving good task performance but sacrificing reconstruction quality. To improve the latter, later methods focus on individual neural fields parameterized as large Multi-Layer Perceptrons (MLPs), which are, however, challenging to process due to the high dimensionality of the weight space, intrinsic weight space symmetries, and sensitivity to random initialization. Hence, results turn out significantly inferior to those achieved by processing explicit representations, e.g., point clouds or meshes. In the meantime, hybrid representations, in particular based on tri-planes, have emerged as a more effective and efficient alternative to realize neural fields, but their direct processing has not been investigated yet. In this paper, we show that the tri-plane discrete data structure encodes rich information, which can be effectively processed by standard deep-learning machinery. We define an extensive benchmark covering a diverse set of fields such as occupancy, signed/unsigned distance, and, for the first time, radiance fields. While processing a field with the same reconstruction quality, we achieve task performance far superior to frameworks that process large MLPs and, for the first time, almost on par with architectures handling explicit representations.
Adriano Cardace, Pierluigi Zama Ramirez, Francesco Ballerini, Allan Zhou, Samuele Salti, Luigi Di Stefano
ICLR5
2024 LLaNA: Large Language and NeRF Assistant
abstract
Multimodal Large Language Models (MLLMs) have demonstrated an excellent understanding of images and 3D data. However, both modalities have shortcomings in holistically capturing the appearance and geometry of objects. Meanwhile, Neural Radiance Fields (NeRFs), which encode information within the weights of a simple Multi-Layer Perceptron (MLP), have emerged as an increasingly widespread modality that simultaneously encodes the geometry and photorealistic appearance of objects. This paper investigates the feasibility and effectiveness of ingesting NeRF into MLLM. We create LLaNA, the first general-purpose NeRF-language assistant capable of performing new tasks such as NeRF captioning and Q&A. Notably, our method directly processes the weights of the NeRF’s MLP to extract information about the represented objects without the need to render images or materialize 3D data structures. Moreover, we build a dataset of NeRFs with text annotations for various NeRF-language tasks with no human intervention. Based on this dataset, we develop a benchmark to evaluate the NeRF understanding capability of our method. Results show that processing NeRF weights performs favourably against extracting 2D or 3D representations from NeRFs.
Andrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti, Luigi Di Stefano
NeurIPS4
2024 Booster: A Benchmark for Depth From Images of Specular and Transparent Surfaces
abstract
Estimating depth from images nowadays yields outstanding results, both in terms of in-domain accuracy and generalization. However, we identify two main challenges that remain open in this field: dealing with non-Lambertian materials and effectively processing high-resolution images. Purposely, we propose a novel dataset that includes accurate and dense ground-truth labels at high resolution, featuring scenes containing several specular and transparent surfaces. Our acquisition pipeline leverages a novel deep space-time stereo framework, enabling easy and accurate labeling with sub-pixel precision. The dataset is composed of 606 samples collected in 85 different scenes, each sample includes both a high-resolution pair (12 Mpx) as well as an unbalanced stereo pair (Left: 12 Mpx, Right: 1.1 Mpx), typical of modern mobile devices that mount sensors with different resolutions. Additionally, we provide manually annotated material segmentation masks and 15 K unlabeled samples. The dataset is composed of a train set and two test sets, the latter devoted to the evaluation of stereo and monocular depth estimation networks. Our experiments highlight the open challenges and future research directions in this field.
Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Deep Learning on Object-Centric 3D Neural Fields
abstract
In recent years, Neural Fields (NFs) have emerged as an effective tool for encoding diverse continuous signals such as images, videos, audio, and 3D shapes. When applied to 3D data,NFs offer a solution to the fragmentation and limitations associated with prevalent discrete representations. However, given thatNFs are essentially neural networks, it remains unclear whether and how they can be seamlessly integrated into deep learning pipelines for solving downstream tasks. This paper addresses this research problem and introducesnf2vec, a framework capable of generating a compact latent representation for an inputNFin a single inference pass. We demonstrate thatnf2veceffectively embeds 3D objects represented by the inputNFs and showcase how the resulting embeddings can be employed in deep learning pipelines to successfully address various tasks, all while processing exclusivelyNFs. We test this framework on severalNFs used to represent 3D surfaces, such as unsigned/signed distance and occupancy fields. Moreover, we demonstrate the effectiveness of our approach with more complexNFs that encompass both geometry and appearance of 3D objects such as neural radiance fields.
Pierluigi Zama Ramirez, Luca De Luigi, Daniele Sirocchi, Adriano Cardace, Riccardo Spezialetti, Francesco Ballerini, Samuele Salti, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Neural Disparity Refinement
abstract
We propose a framework that combines traditional, hand-crafted algorithms and recent advances in deep learning to obtain high-quality, high-resolution disparity maps from stereo images. By casting the refinement process as a continuous feature sampling strategy, our neural disparity refinement network can estimate an enhanced disparity map at any output resolution. Our solution can process any disparity map produced by classical stereo algorithms, as well as those predicted by modern stereo networks or even different depth-from-images approaches, such as the COLMAP structure-from-motion pipeline. Nonetheless, when deployed in the former configuration, our framework performs at its best in terms of zero-shot generalization from synthetic to real images. Moreover, its continuous formulation allows for easily handling the unbalanced stereo setup very diffused in mobile phones.
Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Dynamic Bird's Eye View Reconstruction of Driving Accidents
abstract
The consequences of vehicle crashes are extremely costly, especially in industrial contexts, where the loss of income due to the vehicle unavailability while the incident is investigated adds to the damage produced by the event. The ongoing shift toward more connected vehicles, featuring sensors and cameras, offers the opportunity to alleviate such losses by speeding up the resolution of disputes. In this paper, we show how data routinely collected by connected vehicles can be fused to attain automatic reconstruction of the crash dynamic, a key element that has to be provided by drivers to submit a First Notification of Loss. We build upon state-of-the-art methods in areas such as SLAM, depth estimation and object detection to create a reconstruction of the scene with the vehicles involved localized both in space and time, which we present in an animated bird’s eye view. Our pipeline is evaluated on a challenging benchmark of real world videos and it is shown to create reliable reconstructions of the moment of the impact in more than 50% of scenes and overall good reconstructions in about 37% of them.
Marco Boschi, Luca De Luigi, Samuele Salti, Francesco Sambo, Douglas Coimbra de Andrade, Leonardo Taccari, Alex Quintero Garcia
IEEE Trans. Intell. Transp. Syst.3
2023 ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects
abstract
In this paper, we focus on the problem of rendering novel views from a Neural Radiance Field (NeRF) under unobserved light conditions. To this end, we introduce a novel dataset, dubbed ReNe (Relighting NeRF), framing real world objects under one-light-at-time (OLAT) conditions, annotated with accurate ground-truth camera and light poses. Our acquisition pipeline leverages two robotic arms holding, respectively, a camera and an omni-directional point-wise light source. We release a total of 20 scenes depicting a variety of objects with complex geometry and challenging materials. Each scene includes 2000 images, acquired from 50 different points of views under 40 different OLAT conditions. By leveraging the dataset, we perform an ablation study on the relighting capability of variants of the vanilla NeRF architecture and identify a lightweight architecture that can render novel views of an object under novel light conditions, which we use to establish a non-trivial baseline for the dataset. Dataset and benchmark are available at https://eyecan-ai.
Marco Toschi, Riccardo De Matteo, Riccardo Spezialetti, Daniele De Gregorio, Luigi Di Stefano, Samuele Salti
CVPR6
2023 Deep Learning on Implicit Neural Representations of Shapes
Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
ICLR5
2023 Depth Self-Supervision for Single Image Novel View Synthesis
abstract
In this paper, we tackle the problem of generating a novel image from an arbitrary viewpoint given a single frame as input. While existing methods operating in this setup aim at predicting the target view depth map to guide the synthesis, without explicit supervision over such a task, we jointly optimize our framework for both novel view synthesis and depth estimation to unleash the synergy between the two at its best. Specifically, a shared depth decoder is trained in a self-supervised manner to predict depth maps that are consistent across the source and target views. Our results demonstrate the effectiveness of our approach in addressing the challenges of both tasks allowing for higher-quality generated images, as well as more accurate depth for the target viewpoint.
Giovanni Minelli, Matteo Poggi, Samuele Salti
IROS3
2023 Self-Distillation for Unsupervised 3D Domain Adaptation
abstract
Point cloud classification is a popular task in 3D vision. However, previous works, usually assume that point clouds at test time are obtained with the same procedure or sensor as those at training time. Unsupervised Domain Adaptation (UDA) instead, breaks this assumption and tries to solve the task on an unlabeled target domain, leveraging only on a supervised source domain. For point cloud classification, recent UDA methods try to align features across domains via auxiliary tasks such as point cloud reconstruction, which however do not optimize the discriminative power in the target domain in feature space. In contrast, in this work, we focus on obtaining a discriminative feature space for the target domain enforcing consistency between a point cloud and its augmented version. We then propose a novel iterative self-training methodology that exploits Graph Neural Networks in the UDA context to refine pseudo-labels. We perform extensive experiments and set the new state-of-the art in standard UDA benchmarks for point cloud classification. Finally, we show how our approach can be extended to more complex tasks such as part segmentation.
Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
WACV4
2023 Learning Good Features to Transfer Across Tasks and Domains
abstract
Availability of labelled data is the major obstacle to the deployment of deep learning algorithms for computer vision tasks in new domains. The fact that many frameworks adopted to solve different tasks share the same architecture suggests that there should be a way of reusing the knowledge learned in a specific setting to solve novel tasks with limited or no additional supervision. In this work, we first show that such knowledge can be shared across tasks by learning a mapping between task-specific deep features in a given domain. Then, we show that this mapping function, implemented by a neural network, is able to generalize to novel unseen domains. Besides, we propose a set of strategies to constrain the learned feature spaces, to ease learning and increase the generalization capability of the mapping network, thereby considerably improving the final performance of our framework. Our proposal obtains compelling results in challenging synthetic-to-real adaptation scenarios by transferring knowledge between monocular depth estimation and semantic segmentation tasks.
Pierluigi Zama Ramirez, Adriano Cardace, Luca De Luigi, Alessio Tonioni, Samuele Salti, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Lightweight and Effective Convolutional Neural Networks for Vehicle Viewpoint Estimation From Monocular Images
abstract
Vehicle viewpoint estimation from monocular images is a crucial component for autonomous driving vehicles and for fleet management applications. In this paper, we make several contributions to advance the state-of-the-art on this problem. We show the effectiveness of applying a smoothing filter to the output neurons of a Convolutional Neural Network (CNN) when estimating vehicle viewpoint. We point out the overlooked fact that, under the same viewpoint, the appearance of a vehicle is strongly influenced by its position in the image plane, which renders viewpoint estimation from appearance an ill-posed problem. We show how, by inserting in the model a CoordConv layer to provide the coordinates of the vehicle, we are able to solve such ambiguity and greatly increase performance. Finally, we introduce a new data augmentation technique that improves viewpoint estimation on vehicles that are closer to the camera or partially occluded. All these improvements let a lightweight CNN reach optimal results while keeping inference time low. An extensive evaluation on a viewpoint estimation benchmark (Pascal3D+) and on actual vehicle camera data (nuScenes) shows that our method significantly outperforms the state-of-the-art in vehicle viewpoint estimation, both in terms of accuracy and memory footprint.
Simone Magistri, Marco Boschi, Francesco Sambo, Douglas Coimbra de Andrade, Matteo Simoncini, Luca Kubin, Leonardo Taccari, Luca De Luigi, Samuele Salti
IEEE Trans. Intell. Transp. Syst.9
2022 Cross-Spectral Neural Radiance Fields
abstract
We propose X-NeRF, a novel method to learn a Cross-Spectral scene representation given images captured from cameras with different light spectrum sensitivity, based on the Neural Radiance Fields formulation. X-NeRF optimizes camera poses across spectra during training and exploits Normalized Cross-Device Coordinates (NXDC) to render images of different modalities from arbitrary viewpoints, which are aligned and at the same resolution. Experiments on 16 forward-facing scenes, featuring color, multi-spectral and infrared images, confirm the effectiveness of X-NeRF at modeling Cross-Spectral scene representations.
Matteo Poggi, Pierluigi Zama Ramirez, Fabio Tosi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
3DV4
2022 Open Challenges in Deep Stereo: the Booster Dataset
abstract
We present a novel high-resolution and challenging stereo dataset framing indoor scenes annotated with dense and accurate ground-truth disparities. Peculiar to our dataset is the presence of several specular and transparent surfaces, i.e. the main causes of failures for state-of-the-art stereo networks. Our acquisition pipeline leverages a novel deep space-time stereo framework which allows for easy and accurate labeling with sub-pixel precision. We re-lease a total of 419 samples collected in 64 different scenes and annotated with dense ground-truth disparities. Each sample include a high-resolution pair (12 Mpx) as well as an unbalanced pair (Left: 12 Mpx, Right: 1.1 Mpx). Additionally, we provide manually annotated material segmentation masks and 15K unlabeled samples. We evaluate state-of-the-art deep networks based on our dataset, highlighting their limitations in addressing the open challenges in stereo and drawing hints for future research.
Pierluigi Zama Ramirez, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
CVPR4
2022 RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation
abstract
We address the problem of registering synchronized color (RGB) and multi-spectral (MS) images featuring very different resolution by solving stereo matching correspondences. Purposely, we introduce a novel RGB-MS dataset framing 13 different scenes in indoor environments and providing a total of 34 image pairs annotated with semi-dense, high-resolution ground-truth labels in the form of disparity maps. To tackle the task, we propose a deep learning architecture trained in a self-supervised manner by exploiting a further RGB camera, required only during training data acquisition. In this setup, we can conveniently learn cross-modal matching in the absence of ground-truth labels by distilling knowledge from an easier RGB-RGB matching task based on a collection of about 11K unlabeled image triplets. Experiments show that the proposed pipeline sets a good performance bar (1.16 pixels average registration error) for future research on this novel, challenging task.
Fabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
CVPR4
2022 Learning the Space of Deep Models
abstract
Embedding of large but redundant data, such as images or text, in a hierarchy of lower-dimensional spaces is one of the key features of representation learning approaches, which nowadays provide state-of-the-art solutions to problems once believed hard or impossible to solve. In this work1, in a plot twist with a strong meta aftertaste, we show how trained deep models are as redundant as the data they are optimized to process, and how it is therefore possible to use deep learning models to embed deep learning models. In particular, we show that it is possible to use representation learning to learn a fixed-size, low-dimensional embedding space of trained deep models and that such space can be explored by interpolation or optimization to attain ready-to-use models. We find that it is possible to learn an embedding space of multiple instances of the same architecture and of multiple architectures. We address image classification and neural representation of signals, showing how our embedding space can be learnt so as to capture the notions of performance and 3D shape, respectively. In the Multi-Architecture setting we also show how an embedding trained only on a subset of architectures can learn to generate already-trained instances of architectures it never sees instantiated at training time.
Gianluca Berardi, Luca De Luigi, Samuele Salti, Luigi Di Stefano
ICPR3
2022 Plugging Self-Supervised Monocular Depth into Unsupervised Domain Adaptation for Semantic Segmentation
abstract
Although recent semantic segmentation methods have made remarkable progress, they still rely on large amounts of annotated training data, which are often infeasible to collect in the autonomous driving scenario. Previous works usually tackle this issue with Unsupervised Domain Adaptation (UDA), which entails training a network on synthetic images and applying the model to real ones while minimizing the discrepancy between the two domains. Yet, these techniques do not consider additional information that may be obtained from other tasks. Differently, we propose to exploit self-supervised monocular depth estimation to improve UDA for semantic segmentation. On one hand, we deploy depth to realize a plug-in component which can inject complementary geometric cues into any existing UDA method. We further rely on depth to generate a large and varied set of samples to Self-Train the final model. Our whole proposal allows for achieving state-of-the-art performance (58.8 mIoU) in the GTA5 → CS benchmark. Code is available at https://github.com/CVLAB-Unibo/d4-dbst.
Adriano Cardace, Luca De Luigi, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
WACV4
2022 Shallow Features Guide Unsupervised Domain Adaptation for Semantic Segmentation at Class Boundaries
abstract
Although deep neural networks have achieved remarkable results for the task of semantic segmentation, they usually fail to generalize towards new domains, especially when performing synthetic-to-real adaptation. Such domain shift is particularly noticeable along class boundaries, invalidating one of the main goals of semantic segmentation that consists in obtaining sharp segmentation masks.In this work, we specifically address this core problem in the context of Unsupervised Domain Adaptation and present a novel low-level adaptation strategy that allows us to obtain sharp predictions. Moreover, inspired by recent self-training techniques, we introduce an effective data augmentation that alleviates the noise typically present at semantic boundaries when employing pseudo-labels for self-training. Our contributions can be easily integrated into other popular adaptation frameworks, and extensive experiments show that they effectively improve performance along class boundaries.
Adriano Cardace, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
WACV3
2022 Unsupervised Learning of Local Equivariant Descriptors for Point Clouds
abstract
Correspondences between 3D keypoints generated by matching local descriptors are a key step in 3D computer vision and graphic applications. Learned descriptors are rapidly evolving and outperforming the classical handcrafted approaches in the field. Yet, to learn effective representations they require supervision through labeled data, which are cumbersome and time-consuming to obtain. Unsupervised alternatives exist, but they lag in performance. Moreover, invariance to viewpoint changes is attained either by relying on data augmentation, which is prone to degrading upon generalization on unseen datasets, or by learning from handcrafted representations of the input which are already rotation invariant but whose effectiveness at training time may significantly affect the learned descriptor. We show how learning an equivariant 3D local descriptor instead of an invariant one can overcome both issues. LEAD (Local EquivAriant Descriptor) combines Spherical CNNs to learn an equivariant representation together with plane-folding decoders to learn without supervision. Through extensive experiments on standard surface registration datasets, we show how our proposal outperforms existing unsupervised methods by a large margin and achieves competitive results against the supervised approaches, especially in the practically very relevant scenario of transfer learning.
Marlon Marcon, Riccardo Spezialetti, Samuele Salti, Luciano Silva, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Unsafe Maneuver Classification From Dashcam Video and GPS/IMU Sensors Using Spatio-Temporal Attention Selector
abstract
In this paper, we propose a novel deep learning architecture to classify unsafe driving maneuvers from dashcam and IMU data. Such architecture processes the output of an object detection algorithm in combination with raw video frames and GPS/IMU data. At the core of the architecture there is a novel Spatio-Temporal Attention Selector (STAS) module, which (1) extracts features describing the evolution of each object in the scene over time and (2) leverages multi-head dot product attention to select the relevant ones,i.e., the dangerous ones or the ones in danger, to perform classification. We also introduce a simple but effective methodology to increase the benefit of fine-tuning the backbone network. Our method is shown to achieve higher performance than other approaches in the literature applying attention over single frames.
Matteo Simoncini, Douglas Coimbra de Andrade, Leonardo Taccari, Samuele Salti, Luca Kubin, Fabio Schoen, Francesco Sambo
IEEE Trans. Intell. Transp. Syst.4
2021 Neural Disparity Refinement for Arbitrary Resolution Stereo
abstract
We introduce a novel architecture for neural disparity refinement aimed at facilitating deployment of 3D computer vision on cheap and widespread consumer devices, such as mobile phones. Our approach relies on a continuous formulation that enables to estimate a refined disparity map at any arbitrary output resolution. Thereby, it can handle effectively the unbalanced camera setup typical of nowadays mobile phones, which feature both high and low resolution RGB sensors within the same device. Moreover, our neural network can process seamlessly the output of a variety of stereo methods and, by refining the disparity maps computed by a traditional matching algorithm like SGM, it can achieve unpaired zero-shot generalization performance compared to state-of-the-art end-to-end stereo models.
Filippo Aleotti, Fabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
3DV5
2021 RefRec: Pseudo-labels Refinement via Shape Reconstruction for Unsupervised 3D Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) for point cloud classification is an emerging research problem with relevant practical motivations. Reliance on multi-task learning to align features across domains has been the standard way to tackle it. In this paper, we take a different path and propose RefRec, the first approach to investigate pseudo-labels and self-training in UDA for point clouds. We present two main innovations to make self-training effective on 3D data: i) refinement of noisy pseudo-labels by matching shape descriptors that are learned by the unsupervised task of shape reconstruction on both domains; ii) a novel self-training protocol that learns domain-specific decision boundaries and reduces the negative impact of mislabelled target samples and in-domain intra-class variability. RefRec sets the new state of the art in both standard benchmarks used to test UDA for point cloud classification, showcasing the effectiveness of self-training for this important problem.
Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
3DV4
2020 Distilled Semantics for Comprehensive Scene Understanding from Videos
abstract
Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explicit supervision from ground truth annotations, particularly hard to source for these two tasks. In this paper, we take an additional step toward holistic scene understanding with monocular cameras by learning depth and motion alongside with semantics, with supervision for the latter provided by a pre-trained network distilling proxy ground truth images. We address the three tasks jointly by a) a novel training protocol based on knowledge distillation and self-supervision and b) a compact network architecture which enables efficient scene understanding on both power hungry GPUs and low-power embedded platforms. We thoroughly assess the performance of our framework and show that it yields state-of-the-art results for monocular depth estimation, optical flow and motion segmentation.
Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Luigi Di Stefano, Stefano Mattoccia
CVPR5
2020 Learning to Orient Surfaces by Self-supervised Spherical CNNs
abstract
Defining and reliably finding a canonical orientation for 3D surfaces is key to many Computer Vision and Robotics applications. This task is commonly addressed by handcrafted algorithms exploiting geometric cues deemed as distinctive and robust by the designer. Yet, one might conjecture that humans learn the notion of the inherent orientation of 3D objects from experience and that machines may do so alike. In this work, we show the feasibility of learning a robust canonical orientation for surfaces represented as point clouds. Based on the observation that the quintessential property of a canonical orientation is equivariance to 3D rotations, we propose to employ Spherical CNNs, a recently introduced machinery that can learn equivariant representations defined on the Special Ortoghonal group SO(3). Specifically, spherical correlations compute feature maps whose elements define 3D rotations. Our method learns such feature maps from raw data by a self-supervised training procedure and robustly selects a rotation to transform the input point cloud into a learned canonical orientation. Thereby, we realize the first end-to-end learning approach to define and extract the canonical orientation of 3D shapes, which we aptly dub Compass. Experiments on several public datasets prove its effectiveness at orienting local surface patches as well as whole objects.
Riccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva, Samuele Salti, Luigi Di Stefano
NeurIPS5
2019 Learning Across Tasks and Domains
abstract
Recent works have proven that many relevant visual tasks are closely related one to another. Yet, this connection is seldom deployed in practice due to the lack of practical methodologies to transfer learned concepts across different training processes. In this work, we introduce a novel adaptation framework that can operate across both task and domains. Our framework learns to transfer knowledge across tasks in a fully supervised domain (e.g., synthetic data) and use this knowledge on a different domain where we have only partial supervision (e.g., real data). Our proposal is complementary to existing domain adaptation techniques and extends them to cross tasks scenarios providing additional performance gains. We prove the effectiveness of our framework across two challenging tasks (i.e., monocular depth estimation and semantic segmentation) and four different domains (Synthia, Carla, Kitti, and Cityscapes).
Pierluigi Zama Ramirez, Alessio Tonioni, Samuele Salti, Luigi Di Stefano
ICCV3
2019 Learning an Effective Equivariant 3D Descriptor Without Supervision
abstract
Establishing correspondences between 3D shapes is a fundamental task in 3D Computer Vision, typically ad- dressed by matching local descriptors. Recently, a few at- tempts at applying the deep learning paradigm to the task have shown promising results. Yet, the only explored way to learn rotation invariant descriptors has been to feed neural networks with highly engineered and invariant representations provided by existing hand-crafted descriptors, a path that goes in the opposite direction of end-to-end learning from raw data so successfully deployed for 2D images. In this paper, we explore the benefits of taking a step back in the direction of end-to-end learning of 3D descriptors by disentangling the creation of a robust and distinctive rotation equivariant representation, which can be learned from unoriented input data, and the definition of a good canonical orientation, required only at test time to obtain an invariant descriptor. To this end, we leverage two re- cent innovations: spherical convolutional neural networks to learn an equivariant descriptor and plane folding de- coders to learn without supervision. The effectiveness of the proposed approach is experimentally validated by out- performing hand-crafted and learned descriptors on a standard benchmark.
Riccardo Spezialetti, Samuele Salti, Luigi Di Stefano
ICCV2
2018 Learning to Detect Good 3D Keypoints
Alessio Tonioni, Samuele Salti, Federico Tombari, Riccardo Spezialetti, Luigi Di Stefano
Int. J. Comput. Vis.2
2015 Learning a Descriptor-Specific 3D Keypoint Detector
abstract
Keypoint detection represents the first stage in the majority of modern computer vision pipelines based on automatically established correspondences between local descriptors. However, no standard solution has emerged yet in the case of 3D data such as point clouds or meshes, which exhibit high variability in level of detail and noise. More importantly, existing proposals for 3D keypoint detection rely on geometric saliency functions that attempt to maximize repeatability rather than distinctiveness of the selected regions, which may lead to sub-optimal performance of the overall pipeline. To overcome these shortcomings, we cast 3D keypoint detection as a binary classification between points whose support can be correctly matched by a predefined 3D descriptor or not, thereby learning a descriptor-specific detector that adapts seamlessly to different scenarios. Through experiments on several public datasets, we show that this novel approach to the design of a keypoint detector represents a flexible solution that, nonetheless, can provide state-of-the-art descriptor matching performance.
Samuele Salti, Federico Tombari, Riccardo Spezialetti, Luigi Di Stefano
ICCV1
2015 Traffic sign detection via interest region extraction
Samuele Salti, Alioscia Petrelli, Federico Tombari, Nicola Fioraio, Luigi Di Stefano
Pattern Recognit.1
2015 Synergistic Change Detection and Tracking
abstract
Visual tracking in image streams acquired by static cameras is usually based on change detection and recursive Bayesian estimation, such an approach laying at the core of many practical applications. Yet, the interaction between the change detector and the Bayesian filter is typically designed heuristically. Differently, this paper develops a sound framework to model and implement a bidirectional communication flow between the two processes. In our Bayesian loop, change detection provides well-defined observation likelihood to the recursive filter and the filter prediction provides an informative prior to the change detector, which deploys Bayesian reasoning alike. The loop is developed for the two major variants of Bayesian filters used in tracking, namely the Kalman filter and the particle filter. Experiments on publicly available videos and a novel challenging data set show that the proposed interaction scheme outperforms several state-of-the-art trackers.
Samuele Salti, Alessandro Lanza, Luigi Di Stefano
IEEE Trans. Circuits Syst. Video Technol.1
2014 Automatic detection of pole-like structures in 3D urban environments
abstract
This work aims at automatic detection of man-made pole-like structures in scans of urban environments acquired by a 3D sensor mounted on top a moving vehicle. Pole-like structures, such as e.g. road signs and streetlights, are widespread in these environments, and their reliable detection is relevant to applications dealing with autonomous navigation, facility damage detection, city planning and maintenance. Yet, due to the characteristic thin shape, detection of man-made pole-like structures is significantly prone to both noise as well as occlusions and clutter, the latter being pervasive nuisances when scanning urban environments. Our approach is based on a “local” stage, whereby local features are classified and clustered together, followed by a “global” stage aimed at further classification of candidate entities. The proposed pipeline turns out effective in experiments on a standard publicly available dataset as well as on a challenging dataset acquired during the project for validation purposes.
Federico Tombari, Nicola Fioraio, Tommaso Cavallari, Samuele Salti, Alioscia Petrelli, Luigi Di Stefano
IROS4
2014 SHOT: Unique signatures of histograms for surface and texture description
Samuele Salti, Federico Tombari, Luigi Di Stefano
Comput. Vis. Image Underst.1
2013 Keypoints from Symmetries by Wave Propagation
abstract
The paper conjectures and demonstrates that repeatable keypoints based on salient symmetries at different scales can be detected by a novel analysis grounded on the wave equation rather than the heat equation underlying tradi-tional Gaussian scale–space theory. While the image struc-tures found by most state-of-the-art detectors, such as blobs and corners, occur typically on planar highly textured sur-faces, salient symmetries are widespread in diverse kinds of images, including those related to untextured objects, which are hardly dealt with by current feature-based recog-nition pipelines. We provide experimental results on stan-dard datasets and also contribute with a new dataset fo-cused on untextured objects. Based on the positive exper-imental results, we hope to foster further research on the promising topic of scale invariant analysis through the wave equation. 1.
Samuele Salti, Alessandro Lanza, Luigi Di Stefano
CVPR1
2013 A traffic sign detection pipeline based on interest region extraction
abstract
In this paper we present a pipeline for automatic detection of traffic signs in images. The proposed system can deal with high appearance variations, which typically occur in traffic sign recognition applications, especially with strong illumination changes and dramatic scale changes. Unlike most existing systems, our pipeline is based on interest regions extraction rather than a sliding window detection scheme. The proposed approach has been specialized and tested in three variants, each aimed at detecting one of the three categories of Mandatory, Prohibitory and Danger traffic signs. Our proposal has been evaluated experimentally within the German Traffic Sign Detection Benchmark competition.
Samuele Salti, Alioscia Petrelli, Federico Tombari, Nicola Fioraio, Luigi Di Stefano
IJCNN1
2013 Performance Evaluation of 3D Keypoint Detectors
Federico Tombari, Samuele Salti, Luigi Di Stefano
Int. J. Comput. Vis.2
2013 On-line Support Vector Regression of the transition model for the Kalman filter
Samuele Salti, Luigi Di Stefano
Image Vis. Comput.1
2012 Adaptive Appearance Modeling for Video Tracking: Survey and Evaluation
abstract
Long-term video tracking is of great importance for many applications in real-world scenarios. A key component for achieving long-term tracking is the tracker's capability of updating its internal representation of targets (the appearance model) to changing conditions. Given the rapid but fragmented development of this research area, we propose a unified conceptual framework for appearance model adaptation that enables a principled comparison of different approaches. Moreover, we introduce a novel evaluation methodology that enables simultaneous analysis of tracking accuracy and tracking success, without the need of setting application-dependent thresholds. Based on the proposed framework and this novel evaluation methodology, we conduct an extensive experimental comparison of trackers that perform appearance model adaptation. Theoretical and experimental analyses allow us to identify the most effective approaches as well as to highlight design choices that favor resilience to errors during the update process. We conclude the paper with a list of key open research challenges that have been singled out by means of our experimental comparison.
Samuele Salti, Andrea Cavallaro, Luigi Di Stefano
IEEE Trans. Image Process.1
2011 Background subtraction by non-parametric probabilistic clustering
abstract
We present a background subtraction approach aimed at efficiency and robustness to common source of disturbance such as gradual and sudden illumination changes, camera gain and exposure variations, noise. At each new frame, a non-parametric mixture-based probabilistic clustering is performed to segment the image into changed and unchanged pixels with respect to a fixed background. A two-components mixture, a two-dimensional discrete feature space, a non-parametric model for the components likelihood and a proper initial guess are the key ingredients of this novel algorithm that, besides dealing effectively with the discrimination of photometric and semantic changes, exhibits very high computational efficiency. Experiments are presented, proving the achieved state-of-the-art robustness-efficiency trade-off.
Alessandro Lanza, Samuele Salti, Luigi Di Stefano
AVSS2
2011 A combined texture-shape descriptor for enhanced 3D feature matching
abstract
Motivated by the increasing availability of 3D sensors capable of delivering both shape and texture information, this paper presents a novel descriptor for feature matching in 3D data enriched with texture. The proposed approach stems from the theory of a recently proposed descriptor for 3D data which relies on shape only, and represents its generalization to the case of multiple cues associated with a 3D mesh. The proposed descriptor, dubbed CSHOT, is demonstrated to notably improve the accuracy of feature matching in challenging object recognition scenarios characterized by the presence of clutter and occlusions.
Federico Tombari, Samuele Salti, Luigi Di Stefano
ICIP2
2010 On the Use of Implicit Shape Models for Recognition of Object Categories in 3D Data
Samuele Salti, Federico Tombari, Luigi Di Stefano
ACCV (3)1
2010 Unique Signatures of Histograms for Local Surface Description
Federico Tombari, Samuele Salti, Luigi Di Stefano
ECCV (3)2