Luigi Di Stefano

dblp:00/2029 · DBLP profile ↗
← Back
119ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0001-6014-6421ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 78 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 75 · 5 first-author · 19 since 2021Systems, architecture and hardware · 10Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 NVS-HO: A Benchmark for Novel View Synthesis of Handheld Objects
Musawar Ali, Manuel Carranza-García, Nicola Fioraio, Samuele Salti, Luigi Di Stefano
ICPR (2)5
2025 SiM3D: Single-Instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark
abstract
We propose SiM3D, the first benchmark considering the integration of multiview and multimodal information for comprehensive 3D anomaly detection and segmentation (ADS), where the task is to produce a voxel-based Anomaly Volume. Moreover, SiM3D focuses on a scenario of high interest in manufacturing: single-instance anomaly detection, where only one object, either real or synthetic, is available for training. In this respect, SiM3D stands out as the first ADS benchmark that addresses the challenge of generalising from synthetic training data to real test data. SiM3D includes a novel multimodal multiview dataset acquired using top-tier industrial sensors and robots. The dataset features multiview high-resolution images (12 Mpx) and point clouds (7M points) for 333 instances of eight types of objects, alongside a CAD model for each type. We also provide manually annotated 3D segmentation GTs for anomalous test samples. To establish reference baselines for the proposed multiview 3D ADS task, we adapt prominent singleview methods and assess their performance using novel metrics that operate on Anomaly Volumes.
Alex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia, Alessandro Oliva, Giuseppe Lisanti, Luigi Di Stefano
ICCV7
2025 Spatially-aware Weights Tokenization for NeRF-Language Models
abstract
Neural Radiance Fields (NeRFs) are neural networks -- typically multilayer perceptrons (MLPs) -- that represent the geometry and appearance of objects, with applications in vision, graphics, and robotics. Recent works propose understanding NeRFs with natural language using Multimodal Large Language Models (MLLMs) that directly process the weights of a NeRF's MLP. However, these approaches rely on a global representation of the input object, making them unsuitable for spatial reasoning and fine-grained understanding. In contrast, we propose **weights2space**, a self-supervised framework featuring a novel meta-encoder that can compute a sequence of spatial tokens directly from the weights of a NeRF. Leveraging this representation, we build **Spatial LLaNA**, a novel MLLM for NeRFs, capable of understanding details and spatial relationships in objects represented as NeRFs. We evaluate Spatial LLaNA on NeRF captioning and NeRF Q&A tasks, using both existing benchmarks and our novel **Spatial ObjaNeRF** dataset consisting of $100$ manually-curated language annotations for NeRFs. This dataset features 3D models and descriptions that challenge the spatial reasoning capability of MLLMs. Spatial LLaNA outperforms existing approaches across all tasks.
Andrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti, Luigi Di Stefano
NeurIPS5
2024 Multimodal Industrial Anomaly Detection by Crossmodal Feature Mapping
abstract
Recent advancements have shown the potential of leveraging both point clouds and images to localize anomalies. Nevertheless, their applicability in industrial manufacturing is often constrained by significant drawbacks, such as the use of memory banks, which lead to a substantial increase in terms of memory footprint and inference time. We propose a novel light and fast framework that learns to map features from one modality to the other on nominal samples and detect anomalies by pinpointing inconsistencies between observed and mapped features. Extensive experiments show that our approach achieves state-of-the-art detection and segmentation performance, in both the standard and few-shot settings, on the MVTec 3D-AD dataset while achieving faster inference and occupying less memory than previous multimodal AD methods. Furthermore, we propose a layer pruning technique to improve memory and time efficiency with a marginal sacrifice in performance.
Alex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti, Luigi Di Stefano
CVPR4
2024 Neural Processing of Tri-Plane Hybrid Neural Fields
abstract
Driven by the appealing properties of neural fields for storing and communicating 3D data, the problem of directly processing them to address tasks such as classification and part segmentation has emerged and has been investigated in recent works. Early approaches employ neural fields parameterized by shared networks trained on the whole dataset, achieving good task performance but sacrificing reconstruction quality. To improve the latter, later methods focus on individual neural fields parameterized as large Multi-Layer Perceptrons (MLPs), which are, however, challenging to process due to the high dimensionality of the weight space, intrinsic weight space symmetries, and sensitivity to random initialization. Hence, results turn out significantly inferior to those achieved by processing explicit representations, e.g., point clouds or meshes. In the meantime, hybrid representations, in particular based on tri-planes, have emerged as a more effective and efficient alternative to realize neural fields, but their direct processing has not been investigated yet. In this paper, we show that the tri-plane discrete data structure encodes rich information, which can be effectively processed by standard deep-learning machinery. We define an extensive benchmark covering a diverse set of fields such as occupancy, signed/unsigned distance, and, for the first time, radiance fields. While processing a field with the same reconstruction quality, we achieve task performance far superior to frameworks that process large MLPs and, for the first time, almost on par with architectures handling explicit representations.
Adriano Cardace, Pierluigi Zama Ramirez, Francesco Ballerini, Allan Zhou, Samuele Salti, Luigi Di Stefano
ICLR6
2024 LLaNA: Large Language and NeRF Assistant
abstract
Multimodal Large Language Models (MLLMs) have demonstrated an excellent understanding of images and 3D data. However, both modalities have shortcomings in holistically capturing the appearance and geometry of objects. Meanwhile, Neural Radiance Fields (NeRFs), which encode information within the weights of a simple Multi-Layer Perceptron (MLP), have emerged as an increasingly widespread modality that simultaneously encodes the geometry and photorealistic appearance of objects. This paper investigates the feasibility and effectiveness of ingesting NeRF into MLLM. We create LLaNA, the first general-purpose NeRF-language assistant capable of performing new tasks such as NeRF captioning and Q&A. Notably, our method directly processes the weights of the NeRF’s MLP to extract information about the represented objects without the need to render images or materialize 3D data structures. Moreover, we build a dataset of NeRFs with text annotations for various NeRF-language tasks with no human intervention. Based on this dataset, we develop a benchmark to evaluate the NeRF understanding capability of our method. Results show that processing NeRF weights performs favourably against extracting 2D or 3D representations from NeRFs.
Andrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti, Luigi Di Stefano
NeurIPS5
2024 Booster: A Benchmark for Depth From Images of Specular and Transparent Surfaces
abstract
Estimating depth from images nowadays yields outstanding results, both in terms of in-domain accuracy and generalization. However, we identify two main challenges that remain open in this field: dealing with non-Lambertian materials and effectively processing high-resolution images. Purposely, we propose a novel dataset that includes accurate and dense ground-truth labels at high resolution, featuring scenes containing several specular and transparent surfaces. Our acquisition pipeline leverages a novel deep space-time stereo framework, enabling easy and accurate labeling with sub-pixel precision. The dataset is composed of 606 samples collected in 85 different scenes, each sample includes both a high-resolution pair (12 Mpx) as well as an unbalanced stereo pair (Left: 12 Mpx, Right: 1.1 Mpx), typical of modern mobile devices that mount sensors with different resolutions. Additionally, we provide manually annotated material segmentation masks and 15 K unlabeled samples. The dataset is composed of a train set and two test sets, the latter devoted to the evaluation of stereo and monocular depth estimation networks. Our experiments highlight the open challenges and future research directions in this field.
Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Deep Learning on Object-Centric 3D Neural Fields
abstract
In recent years, Neural Fields (NFs) have emerged as an effective tool for encoding diverse continuous signals such as images, videos, audio, and 3D shapes. When applied to 3D data,NFs offer a solution to the fragmentation and limitations associated with prevalent discrete representations. However, given thatNFs are essentially neural networks, it remains unclear whether and how they can be seamlessly integrated into deep learning pipelines for solving downstream tasks. This paper addresses this research problem and introducesnf2vec, a framework capable of generating a compact latent representation for an inputNFin a single inference pass. We demonstrate thatnf2veceffectively embeds 3D objects represented by the inputNFs and showcase how the resulting embeddings can be employed in deep learning pipelines to successfully address various tasks, all while processing exclusivelyNFs. We test this framework on severalNFs used to represent 3D surfaces, such as unsigned/signed distance and occupancy fields. Moreover, we demonstrate the effectiveness of our approach with more complexNFs that encompass both geometry and appearance of 3D objects such as neural radiance fields.
Pierluigi Zama Ramirez, Luca De Luigi, Daniele Sirocchi, Adriano Cardace, Riccardo Spezialetti, Francesco Ballerini, Samuele Salti, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 Neural Disparity Refinement
abstract
We propose a framework that combines traditional, hand-crafted algorithms and recent advances in deep learning to obtain high-quality, high-resolution disparity maps from stereo images. By casting the refinement process as a continuous feature sampling strategy, our neural disparity refinement network can estimate an enhanced disparity map at any output resolution. Our solution can process any disparity map produced by classical stereo algorithms, as well as those predicted by modern stereo networks or even different depth-from-images approaches, such as the COLMAP structure-from-motion pipeline. Nonetheless, when deployed in the former configuration, our framework performs at its best in terms of zero-shot generalization from synthetic to real images. Moreover, its continuous formulation allows for easily handling the unbalanced stereo setup very diffused in mobile phones.
Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects
abstract
In this paper, we focus on the problem of rendering novel views from a Neural Radiance Field (NeRF) under unobserved light conditions. To this end, we introduce a novel dataset, dubbed ReNe (Relighting NeRF), framing real world objects under one-light-at-time (OLAT) conditions, annotated with accurate ground-truth camera and light poses. Our acquisition pipeline leverages two robotic arms holding, respectively, a camera and an omni-directional point-wise light source. We release a total of 20 scenes depicting a variety of objects with complex geometry and challenging materials. Each scene includes 2000 images, acquired from 50 different points of views under 40 different OLAT conditions. By leveraging the dataset, we perform an ablation study on the relighting capability of variants of the vanilla NeRF architecture and identify a lightweight architecture that can render novel views of an object under novel light conditions, which we use to establish a non-trivial baseline for the dataset. Dataset and benchmark are available at https://eyecan-ai.
Marco Toschi, Riccardo De Matteo, Riccardo Spezialetti, Daniele De Gregorio, Luigi Di Stefano, Samuele Salti
CVPR5
2023 Learning Depth Estimation for Transparent and Mirror Surfaces
abstract
Inferring the depth of transparent or mirror (ToM) surfaces represents a hard challenge for either sensors, algorithms, or deep networks. We propose a simple pipeline for learning to estimate depth properly for such surfaces with neural networks, without requiring any ground-truth annotation. We unveil how to obtain reliable pseudo labels by in-painting ToM objects in images and processing them with a monocular depth estimation model. These labels can be used to fine-tune existing monocular or stereo networks, to let them learn how to deal with ToM surfaces. Experimental results on the Booster dataset show the dramatic improvements enabled by our remarkably simple proposal.
Alex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, Luigi Di Stefano
ICCV6
2023 Deep Learning on Implicit Neural Representations of Shapes
Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
ICLR6
2023 Self-Distillation for Unsupervised 3D Domain Adaptation
abstract
Point cloud classification is a popular task in 3D vision. However, previous works, usually assume that point clouds at test time are obtained with the same procedure or sensor as those at training time. Unsupervised Domain Adaptation (UDA) instead, breaks this assumption and tries to solve the task on an unlabeled target domain, leveraging only on a supervised source domain. For point cloud classification, recent UDA methods try to align features across domains via auxiliary tasks such as point cloud reconstruction, which however do not optimize the discriminative power in the target domain in feature space. In contrast, in this work, we focus on obtaining a discriminative feature space for the target domain enforcing consistency between a point cloud and its augmented version. We then propose a novel iterative self-training methodology that exploits Graph Neural Networks in the UDA context to refine pseudo-labels. We perform extensive experiments and set the new state-of-the art in standard UDA benchmarks for point cloud classification. Finally, we show how our approach can be extended to more complex tasks such as part segmentation.
Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
WACV5
2023 ScanNeRF: a Scalable Benchmark for Neural Radiance Fields
abstract
In this paper, we propose the first-ever real benchmark thought for evaluating Neural Radiance Fields (NeRFs) and, in general, Neural Rendering (NR) frameworks. We design and implement an effective pipeline for scanning real objects in quantity and effortlessly. Our scan station is built with less than 500$ hardware budget and can collect roughly 4000 images of a scanned object in just 5 minutes. Such a platform is used to build ScanNeRF, a dataset characterized by several train/val/test splits aimed at benchmarking the performance of modern NeRF methods under different conditions. Accordingly, we evaluate three cuttingedge NeRF variants on it to highlight their strengths and weaknesses. The dataset is available on our project page, together with an online benchmark to foster the development of better and better NeRFs.
Luca De Luigi, Damiano Bolognini, Federico Domeniconi, Daniele De Gregorio, Matteo Poggi, Luigi Di Stefano
WACV6
2023 Learning Good Features to Transfer Across Tasks and Domains
abstract
Availability of labelled data is the major obstacle to the deployment of deep learning algorithms for computer vision tasks in new domains. The fact that many frameworks adopted to solve different tasks share the same architecture suggests that there should be a way of reusing the knowledge learned in a specific setting to solve novel tasks with limited or no additional supervision. In this work, we first show that such knowledge can be shared across tasks by learning a mapping between task-specific deep features in a given domain. Then, we show that this mapping function, implemented by a neural network, is able to generalize to novel unseen domains. Besides, we propose a set of strategies to constrain the learned feature spaces, to ease learning and increase the generalization capability of the mapping network, thereby considerably improving the final performance of our framework. Our proposal obtains compelling results in challenging synthetic-to-real adaptation scenarios by transferring knowledge between monocular depth estimation and semantic segmentation tasks.
Pierluigi Zama Ramirez, Adriano Cardace, Luca De Luigi, Alessio Tonioni, Samuele Salti, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Cross-Spectral Neural Radiance Fields
abstract
We propose X-NeRF, a novel method to learn a Cross-Spectral scene representation given images captured from cameras with different light spectrum sensitivity, based on the Neural Radiance Fields formulation. X-NeRF optimizes camera poses across spectra during training and exploits Normalized Cross-Device Coordinates (NXDC) to render images of different modalities from arbitrary viewpoints, which are aligned and at the same resolution. Experiments on 16 forward-facing scenes, featuring color, multi-spectral and infrared images, confirm the effectiveness of X-NeRF at modeling Cross-Spectral scene representations.
Matteo Poggi, Pierluigi Zama Ramirez, Fabio Tosi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
3DV6
2022 Open Challenges in Deep Stereo: the Booster Dataset
abstract
We present a novel high-resolution and challenging stereo dataset framing indoor scenes annotated with dense and accurate ground-truth disparities. Peculiar to our dataset is the presence of several specular and transparent surfaces, i.e. the main causes of failures for state-of-the-art stereo networks. Our acquisition pipeline leverages a novel deep space-time stereo framework which allows for easy and accurate labeling with sub-pixel precision. We re-lease a total of 419 samples collected in 64 different scenes and annotated with dense ground-truth disparities. Each sample include a high-resolution pair (12 Mpx) as well as an unbalanced pair (Left: 12 Mpx, Right: 1.1 Mpx). Additionally, we provide manually annotated material segmentation masks and 15K unlabeled samples. We evaluate state-of-the-art deep networks based on our dataset, highlighting their limitations in addressing the open challenges in stereo and drawing hints for future research.
Pierluigi Zama Ramirez, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
CVPR6
2022 RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation
abstract
We address the problem of registering synchronized color (RGB) and multi-spectral (MS) images featuring very different resolution by solving stereo matching correspondences. Purposely, we introduce a novel RGB-MS dataset framing 13 different scenes in indoor environments and providing a total of 34 image pairs annotated with semi-dense, high-resolution ground-truth labels in the form of disparity maps. To tackle the task, we propose a deep learning architecture trained in a self-supervised manner by exploiting a further RGB camera, required only during training data acquisition. In this setup, we can conveniently learn cross-modal matching in the absence of ground-truth labels by distilling knowledge from an easier RGB-RGB matching task based on a collection of about 11K unlabeled image triplets. Experiments show that the proposed pipeline sets a good performance bar (1.16 pixels average registration error) for future research on this novel, challenging task.
Fabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
CVPR6
2022 Learning the Space of Deep Models
abstract
Embedding of large but redundant data, such as images or text, in a hierarchy of lower-dimensional spaces is one of the key features of representation learning approaches, which nowadays provide state-of-the-art solutions to problems once believed hard or impossible to solve. In this work1, in a plot twist with a strong meta aftertaste, we show how trained deep models are as redundant as the data they are optimized to process, and how it is therefore possible to use deep learning models to embed deep learning models. In particular, we show that it is possible to use representation learning to learn a fixed-size, low-dimensional embedding space of trained deep models and that such space can be explored by interpolation or optimization to attain ready-to-use models. We find that it is possible to learn an embedding space of multiple instances of the same architecture and of multiple architectures. We address image classification and neural representation of signals, showing how our embedding space can be learnt so as to capture the notions of performance and 3D shape, respectively. In the Multi-Architecture setting we also show how an embedding trained only on a subset of architectures can learn to generate already-trained instances of architectures it never sees instantiated at training time.
Gianluca Berardi, Luca De Luigi, Samuele Salti, Luigi Di Stefano
ICPR4
2022 Plugging Self-Supervised Monocular Depth into Unsupervised Domain Adaptation for Semantic Segmentation
abstract
Although recent semantic segmentation methods have made remarkable progress, they still rely on large amounts of annotated training data, which are often infeasible to collect in the autonomous driving scenario. Previous works usually tackle this issue with Unsupervised Domain Adaptation (UDA), which entails training a network on synthetic images and applying the model to real ones while minimizing the discrepancy between the two domains. Yet, these techniques do not consider additional information that may be obtained from other tasks. Differently, we propose to exploit self-supervised monocular depth estimation to improve UDA for semantic segmentation. On one hand, we deploy depth to realize a plug-in component which can inject complementary geometric cues into any existing UDA method. We further rely on depth to generate a large and varied set of samples to Self-Train the final model. Our whole proposal allows for achieving state-of-the-art performance (58.8 mIoU) in the GTA5 → CS benchmark. Code is available at https://github.com/CVLAB-Unibo/d4-dbst.
Adriano Cardace, Luca De Luigi, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
WACV5
2022 Shallow Features Guide Unsupervised Domain Adaptation for Semantic Segmentation at Class Boundaries
abstract
Although deep neural networks have achieved remarkable results for the task of semantic segmentation, they usually fail to generalize towards new domains, especially when performing synthetic-to-real adaptation. Such domain shift is particularly noticeable along class boundaries, invalidating one of the main goals of semantic segmentation that consists in obtaining sharp segmentation masks.In this work, we specifically address this core problem in the context of Unsupervised Domain Adaptation and present a novel low-level adaptation strategy that allows us to obtain sharp predictions. Moreover, inspired by recent self-training techniques, we introduce an effective data augmentation that alleviates the noise typically present at semantic boundaries when employing pseudo-labels for self-training. Our contributions can be easily integrated into other popular adaptation frameworks, and extensive experiments show that they effectively improve performance along class boundaries.
Adriano Cardace, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
WACV4
2022 Feature disentangling and reciprocal learning with label-guided similarity for multi-label image retrieval
Luigi Di Stefano
Neurocomputing4
2022 Unsupervised Learning of Local Equivariant Descriptors for Point Clouds
abstract
Correspondences between 3D keypoints generated by matching local descriptors are a key step in 3D computer vision and graphic applications. Learned descriptors are rapidly evolving and outperforming the classical handcrafted approaches in the field. Yet, to learn effective representations they require supervision through labeled data, which are cumbersome and time-consuming to obtain. Unsupervised alternatives exist, but they lag in performance. Moreover, invariance to viewpoint changes is attained either by relying on data augmentation, which is prone to degrading upon generalization on unseen datasets, or by learning from handcrafted representations of the input which are already rotation invariant but whose effectiveness at training time may significantly affect the learned descriptor. We show how learning an equivariant 3D local descriptor instead of an invariant one can overcome both issues. LEAD (Local EquivAriant Descriptor) combines Spherical CNNs to learn an equivariant representation together with plane-folding decoders to learn without supervision. Through extensive experiments on standard surface registration datasets, we show how our proposal outperforms existing unsupervised methods by a large margin and achieves competitive results against the supervised approaches, especially in the practically very relevant scenario of transfer learning.
Marlon Marcon, Riccardo Spezialetti, Samuele Salti, Luciano Silva, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Continual Adaptation for Deep Stereo
Matteo Poggi, Alessio Tonioni, Fabio Tosi, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Neural Disparity Refinement for Arbitrary Resolution Stereo
abstract
We introduce a novel architecture for neural disparity refinement aimed at facilitating deployment of 3D computer vision on cheap and widespread consumer devices, such as mobile phones. Our approach relies on a continuous formulation that enables to estimate a refined disparity map at any arbitrary output resolution. Thereby, it can handle effectively the unbalanced camera setup typical of nowadays mobile phones, which feature both high and low resolution RGB sensors within the same device. Moreover, our neural network can process seamlessly the output of a variety of stereo methods and, by refining the disparity maps computed by a traditional matching algorithm like SGM, it can achieve unpaired zero-shot generalization performance compared to state-of-the-art end-to-end stereo models.
Filippo Aleotti, Fabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano
3DV7
2021 RefRec: Pseudo-labels Refinement via Shape Reconstruction for Unsupervised 3D Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) for point cloud classification is an emerging research problem with relevant practical motivations. Reliance on multi-task learning to align features across domains has been the standard way to tackle it. In this paper, we take a different path and propose RefRec, the first approach to investigate pseudo-labels and self-training in UDA for point clouds. We present two main innovations to make self-training effective on 3D data: i) refinement of noisy pseudo-labels by matching shape descriptors that are learned by the unsupervised task of shape reconstruction on both domains; ii) a novel self-training protocol that learns domain-specific decision boundaries and reduces the negative impact of mislabelled target samples and in-domain intra-class variability. RefRec sets the new state of the art in both standard benchmarks used to test UDA for point cloud classification, showcasing the effectiveness of self-training for this important problem.
Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano
3DV5
2020 Distilled Semantics for Comprehensive Scene Understanding from Videos
abstract
Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explicit supervision from ground truth annotations, particularly hard to source for these two tasks. In this paper, we take an additional step toward holistic scene understanding with monocular cameras by learning depth and motion alongside with semantics, with supervision for the latter provided by a pre-trained network distilling proxy ground truth images. We address the three tasks jointly by a) a novel training protocol based on knowledge distillation and self-supervision and b) a compact network architecture which enables efficient scene understanding on both power hungry GPUs and low-power embedded platforms. We thoroughly assess the performance of our framework and show that it yields state-of-the-art results for monocular depth estimation, optical flow and motion segmentation.
Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Luigi Di Stefano, Stefano Mattoccia
CVPR6
2020 Effective Deployment of CNNs for 3DoF Pose Estimation and Grasping in Industrial Settings
abstract
In this paper we investigate how to effectively deploy deep learning in practical industrial settings, such as robotic grasping applications. When a deep-learning based solution is proposed, usually lacks of any simple method to generate the training data. In the industrial field, where automation is the main goal, not bridging this gap is one of the main reasons why deep learning is not as widespread as it is in the academic world. For this reason, in this work we developed a system composed by a 3-DoF Pose Estimator based on Convolutional Neural Networks (CNNs) and an effective procedure to gather massive amounts of training images in the field with minimal human intervention. By automating the labeling stage, we also obtain very robust systems suitable for production-level usage. An open source implementation of our solution is provided, alongside with the dataset used for the experimental evaluation.
Daniele De Gregorio, Riccardo Zanella, Gianluca Palli, Luigi Di Stefano
ICPR4
2020 Learning to Orient Surfaces by Self-supervised Spherical CNNs
abstract
Defining and reliably finding a canonical orientation for 3D surfaces is key to many Computer Vision and Robotics applications. This task is commonly addressed by handcrafted algorithms exploiting geometric cues deemed as distinctive and robust by the designer. Yet, one might conjecture that humans learn the notion of the inherent orientation of 3D objects from experience and that machines may do so alike. In this work, we show the feasibility of learning a robust canonical orientation for surfaces represented as point clouds. Based on the observation that the quintessential property of a canonical orientation is equivariance to 3D rotations, we propose to employ Spherical CNNs, a recently introduced machinery that can learn equivariant representations defined on the Special Ortoghonal group SO(3). Specifically, spherical correlations compute feature maps whose elements define 3D rotations. Our method learns such feature maps from raw data by a self-supervised training procedure and robustly selects a rotation to transform the input point cloud into a learned canonical orientation. Thereby, we realize the first end-to-end learning approach to define and extract the canonical orientation of 3D shapes, which we aptly dub Compass. Experiments on several public datasets prove its effectiveness at orienting local surface patches as well as whole objects.
Riccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva, Samuele Salti, Luigi Di Stefano
NeurIPS6
2020 Real-Time RGB-D Camera Pose Estimation in Novel Scenes Using a Relocalisation Cascade
abstract
Camera pose estimation is an important problem in computer vision, with applications as diverse as simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques match the current image against keyframes with known poses coming from a tracker, directly regress the pose, or establish correspondences between keypoints in the current image and points in the scene in order to estimate the pose. In recent years, regression forests have become a popular alternative to establish such correspondences. They achieve accurate results, but have traditionally needed to be trained offline on the target scene, preventing relocalisation in new environments. Recently, we showed how to circumvent this limitation by adapting a pre-trained forest to a new scene on the fly. The adapted forests achieved relocalisation performance that was on par with that of offline forests, and our approach was able to estimate the camera pose in close to real time, which made it desirable for systems that require online relocalisation. In this paper, we present an extension of this work that achieves significantly better relocalisation performance whilst running fully in real time. To achieve this, we make several changes to the original approach: (i) instead of simply accepting the camera pose hypothesis produced by RANSAC without question, we make it possible to score the final few hypotheses it considers using a geometric approach and select the most promising one; (ii) we chain several instantiations of our relocaliser (with different parameter settings) together in a cascade, allowing us to try faster but less accurate relocalisation first, only falling back to slower, more accurate relocalisation as necessary; and (iii) we tune the parameters of our cascade, and the individual relocalisers it contains, to achieve effective overall performance. Taken together, these changes allow us to significantly improve upon the performance our original state-of-the-art method was able to achieve on the well-known 7-Scenes and Stanford 4 Scenes benchmarks. As additional contributions, we present a novel way of visualising the internal behaviour of our forests, and use the insights gleaned from this to show how to entirely circumvent the need to pre-train a forest on a generic scene.
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien P. C. Valentin, Victor Adrian Prisacariu, Luigi Di Stefano, Philip Torr 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 Unsupervised Domain Adaptation for Depth Prediction from Images
abstract
State-of-the-art approaches to infer dense depth measurements from images rely on CNNs trained end-to-end on a vast amount of data. However, these approaches suffer a drastic drop in accuracy when dealing with environments much different in appearance and/or context from those observed at training time. This domain shift issue is usually addressed by fine-tuning on smaller sets of images from the target domain annotated with depth labels. Unfortunately, relying on such supervised labeling is seldom feasible in most practical settings. Therefore, we propose an unsupervised domain adaptation technique which does not require groundtruth labels. Our method relies only on image pairs and leverages on classical stereo algorithms to produce disparity measurements alongside with confidence estimators to assess upon their reliability. We propose to fine-tune both depth-from-stereo as well as depth-from-mono architectures by a novel confidence-guided loss function that handles the measured disparities as noisy labels weighted according to the estimated confidence. Extensive experimental results based on standard datasets and evaluation protocols prove that our technique can address effectively the domain shift issue with both stereo and monocular depth prediction architectures and outperforms other state-of-the-art unsupervised loss functions that may be alternatively deployed to pursue domain adaptation.
Alessio Tonioni, Matteo Poggi, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Semiautomatic Labeling for Deep Learning in Robotics
abstract
In this article, we propose an augmented reality semiautomatic labeling (ARS), a semiautomatic method which leverages on moving a 2-D camera by means of a robot, proving precise camera tracking, and an augmented reality pen (ARP) to define initial object bounding box, to create large labeled data sets with minimal human intervention. By removing the burden of generating annotated data from humans, we make the deep learning technique applied to computer vision, which typically requires very large data sets, truly automated and reliable. With the ARS pipeline, we created two novel data sets effortlessly, one on electromechanical components (industrial scenario) and other on fruits (daily-living scenario) and trained two state-of-the-art object detectors robustly, based on convolutional neural networks, such as you only look once (YOLO) and single shot detector (SSD). With respect to conventional manual annotation of 1000 frames that takes us slightly more than 10 h, the proposed approach based on ARS allows to annotate 9 sequences of about 35 000 frames in less than 1 h, with a gain factor of about 450. Moreover, both the precision and recall of object detection is increased by about 15% with respect to manual labeling. All our software is available as a robot operating system (ROS) package in a public repository alongside with the novel annotated data sets. Note to Practitioners-This article was motivated by the lack of a simple and effective solution for the generation of data sets usable to train a data-driven model, such as a modern deep neural network, so as to make them accessible in an industrial environment. Specifically, a deep learning robot guidance vision system would require such a large amount of manually labeled images that it would be too expensive and impractical for a real use case, where system reconfigurability is a fundamental requirement. With our system, on the other hand, especially in the field of industrial robotics, the cost of image labeling can be reduced, for the first time, to nearly zero, thus paving the way for self-reconfiguring systems with very high performance (as demonstrated by our experimental results). One of the limitations of this approach is the need to use a manual method for the detection of objects of interest in the preliminary stages of the pipeline (ARP or graphical interface). A feasible extension, related to the field of collaborative robotics, could be used to exploit the robot itself, manually moved by the user, even for this preliminary stage, so as to eliminate any source of inaccuracy.
Daniele De Gregorio, Alessio Tonioni, Gianluca Palli, Luigi Di Stefano
IEEE Trans Autom. Sci. Eng.4
2019 GFrames: Gradient-Based Local Reference Frame for 3D Shape Matching
abstract
We introduce GFrames, a novel local reference frame (LRF) construction for 3D meshes and point clouds. GFrames are based on the computation of the intrinsic gradient of a scalar field defined on top of the input shape. The resulting tangent vector field defines a repeatable tangent direction of the local frame at each point; importantly, it directly inherits the properties and invariance classes of the underlying scalar function, making it remarkably robust under strong sampling artifacts, vertex noise, as well as non-rigid deformations. Existing local descriptors can directly benefit from our repeatable frames, as we showcase in a selection of 3D vision and shape analysis applications where we demonstrate state-of-the-art performance in a variety of challenging settings.
Simone Melzi, Riccardo Spezialetti, Federico Tombari, Michael M. Bronstein, Luigi Di Stefano, Emanuele Rodolà
CVPR5
2019 Learning to Adapt for Stereo
abstract
Real world applications of stereo depth estimation require models that are robust to dynamic variations in the environment. Even though deep learning based stereo methods are successful, they often fail to generalize to unseen variations in the environment, making them less suitable for practical applications such as autonomous driving. In this work, we introduce a ``learning-to-adapt'' framework that enables deep stereo methods to continuously adapt to new target domains in an unsupervised manner. Specifically, our approach incorporates the adaptation procedure into the learning objective to obtain a base set of parameters that are better suited for unsupervised online adaptation. To further improve the quality of the adaptation, we learn a confidence measure that effectively masks the errors introduced during the unsupervised adaptation. We evaluate our method on synthetic and real-world stereo datasets and our experiments evidence that learning-to-adapt is, indeed beneficial for online adaptation on vastly different domains.
Alessio Tonioni, Oscar Rahnama, Thomas Joy, Luigi Di Stefano, Thalaiyasingam Ajanthan, Philip Torr 0001
CVPR4
2019 Real-Time Self-Adaptive Deep Stereo
abstract
Deep convolutional neural networks trained end-to-end are the state-of-the-art methods to regress dense disparity maps from stereo pairs. These models, however, suffer from a notable decrease in accuracy when exposed to scenarios significantly different from the training set (e.g., real vs synthetic images, etc.). We argue that it is extremely unlikely to gather enough samples to achieve effective training/tuning in any target domain, thus making this setup impractical for many applications. Instead, we propose to perform unsupervised and continuous online adaptation of a deep stereo network, which allows for preserving its accuracy in any environment. However, this strategy is extremely computationally demanding and thus prevents real-time inference. We address this issue introducing a new lightweight, yet effective, deep stereo architecture, Modularly ADaptive Network(MADNet), and developing a Modular ADaptation (MAD) algorithm, which independently trains sub-portions of the network. By deploying MADNet together with MAD we introduce the first real-time self-adaptive deep stereo system enabling competitive performance on heterogeneous datasets. Our code is publicly available at https://github.com/CVLAB-Unibo/Real-time-self-adaptive-deep-stereo.
Alessio Tonioni, Fabio Tosi, Matteo Poggi, Stefano Mattoccia, Luigi Di Stefano
CVPR5
2019 Learning Across Tasks and Domains
abstract
Recent works have proven that many relevant visual tasks are closely related one to another. Yet, this connection is seldom deployed in practice due to the lack of practical methodologies to transfer learned concepts across different training processes. In this work, we introduce a novel adaptation framework that can operate across both task and domains. Our framework learns to transfer knowledge across tasks in a fully supervised domain (e.g., synthetic data) and use this knowledge on a different domain where we have only partial supervision (e.g., real data). Our proposal is complementary to existing domain adaptation techniques and extends them to cross tasks scenarios providing additional performance gains. We prove the effectiveness of our framework across two challenging tasks (i.e., monocular depth estimation and semantic segmentation) and four different domains (Synthia, Carla, Kitti, and Cityscapes).
Pierluigi Zama Ramirez, Alessio Tonioni, Samuele Salti, Luigi Di Stefano
ICCV4
2019 Learning an Effective Equivariant 3D Descriptor Without Supervision
abstract
Establishing correspondences between 3D shapes is a fundamental task in 3D Computer Vision, typically ad- dressed by matching local descriptors. Recently, a few at- tempts at applying the deep learning paradigm to the task have shown promising results. Yet, the only explored way to learn rotation invariant descriptors has been to feed neural networks with highly engineered and invariant representations provided by existing hand-crafted descriptors, a path that goes in the opposite direction of end-to-end learning from raw data so successfully deployed for 2D images. In this paper, we explore the benefits of taking a step back in the direction of end-to-end learning of 3D descriptors by disentangling the creation of a robust and distinctive rotation equivariant representation, which can be learned from unoriented input data, and the definition of a good canonical orientation, required only at test time to obtain an invariant descriptor. To this end, we leverage two re- cent innovations: spherical convolutional neural networks to learn an equivariant descriptor and plane folding de- coders to learn without supervision. The effectiveness of the proposed approach is experimentally validated by out- performing hand-crafted and learned descriptors on a standard benchmark.
Riccardo Spezialetti, Samuele Salti, Luigi Di Stefano
ICCV3
2019 Domain invariant hierarchical embedding for grocery products recognition
Alessio Tonioni, Luigi Di Stefano
Comput. Vis. Image Underst.2
2018 Let's Take a Walk on Superpixels Graphs: Deformable Linear Objects Segmentation and Model Estimation
Daniele De Gregorio, Gianluca Palli, Luigi Di Stefano
ACCV (2)3
2018 Geometry Meets Semantics for Semi-supervised Monocular Depth Estimation
Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, Luigi Di Stefano
ACCV (3)5
2018 Exploiting semantics in adversarial training for image-level domain adaptation
abstract
Performance achievable by modern deep learning approaches are directly related to the amount of data used at training time. Unfortunately, the annotation process is notoriously tedious and expensive, especially for pixel-wise tasks like semantic segmentation. Recent works have proposed to rely on synthetically generated imagery to ease the training set creation. However, models trained on these kind of data usually under-perform on real images due to the well known issue of domain shift. We address this problem by learning a domain-to-domain image translation GAN to shrink the gap between real and synthetic images. Peculiarly to our method, we introduce semantic constraints into the generation process to both avoid artifacts and guide the synthesis. To prove the effectiveness of our proposal, we show how a semantic segmentation CNN trained on images from the synthetic GTA dataset adapted by our method can improve performance by more than 16% mIoU with respect to the same model trained on synthetic images.
Pierluigi Zama Ramirez, Alessio Tonioni, Luigi Di Stefano
IPAS3
2018 A deep learning pipeline for product recognition on store shelves
abstract
Recognition of grocery products in store shelves poses peculiar challenges. Firstly, the task mandates the recognition of an extremely high number of different items, in the order of several thousands for medium-small shops, with many of them featuring small inter and intra class variability. Then, available product databases usually include just one or a few studio-quality images per product (referred to herein as reference images), whilst at test time recognition is performed on pictures displaying a portion of a shelf containing several products and taken in the store by cheap cameras (referred to herein as query images). Moreover, as the items on sale in a store as well as their appearance change frequently overtime, a practical recognition system should handle seamlessly new products/packages. We developed a deep learning based pipeline to solve this task. First we deploy state of the art object detectors to obtain an initial product-agnostic item detection, then, we pursue product recognition through a similarity search between global descriptors computed on reference and cropped query images. To maximize performance, we learn an ad-hoc global descriptor by a CNN trained on reference images based on an image embedding loss. We have tested our pipeline on the standard grocery product [1] dataset and improved the currents state of the art. While computationally expensive at training time our system turn out not only accurate but also quite fast at test time.
Alessio Tonioni, Eugenio Serra, Luigi Di Stefano
IPAS3
2018 Learning to Detect Good 3D Keypoints
Alessio Tonioni, Samuele Salti, Federico Tombari, Riccardo Spezialetti, Luigi Di Stefano
Int. J. Comput. Vis.5
2017 Learning confidence measures in the wild
Fabio Tosi, Matteo Poggi, Stefano Mattoccia, Alessio Tonioni, Luigi Di Stefano
BMVC5
2017 On-the-Fly Adaptation of Regression Forests for Online Camera Relocalisation
abstract
Camera relocalisation is an important problem in computer vision, with applications in simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques either match the current image against keyframes with known poses coming from a tracker, or establish 2D-to-3D correspondences between keypoints in the current image and points in the scene in order to estimate the camera pose. Recently, regression forests have become a popular alternative to establish such correspondences. They achieve accurate results, but must be trained offline on the target scene, preventing relocalisation in new environments. In this paper, we show how to circumvent this limitation by adapting a pre-trained forest to a new scene on the fly. Our adapted forests achieve relocalisation performance that is on par with that of offline forests, and our approach runs in under 150ms, making it desirable for real-time systems that require online relocalisation.
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien P. C. Valentin, Luigi Di Stefano, Philip Torr 0001
CVPR5
2017 Unsupervised Adaptation for Deep Stereo
abstract
Recent ground-breaking works have shown that deep neural networks can be trained end-to-end to regress dense disparity maps directly from image pairs. Computer generated imagery is deployed to gather the large data corpus required to train such networks, an additional fine-tuning allowing to adapt the model to work well also on real and possibly diverse environments. Yet, besides a few public datasets such as Kitti, the ground-truth needed to adapt the network to a new scenario is hardly available in practice. In this paper we propose a novel unsupervised adaptation approach that enables to fine-tune a deep learning stereo model without any ground-truth information. We rely on off-the-shelf stereo algorithms together with state-of-the-art confidence measures, the latter able to ascertain upon correctness of the measurements yielded by former. Thus, we train the network based on a novel loss-function that penalizes predictions disagreeing with the highly confident disparities provided by the algorithm and enforces a smoothness constraint. Experiments on popular datasets (KITTI 2012, KITTI 2015 and Middlebury 2014) and other challenging test images demonstrate the effectiveness of our proposal.
Alessio Tonioni, Matteo Poggi, Stefano Mattoccia, Luigi Di Stefano
ICCV4
2017 SkiMap: An efficient mapping framework for robot navigation
abstract
We present a novel mapping framework for robot navigation which features a multi-level querying system capable to obtain rapidly representations as diverse as a 3D voxel grid, a 2.5D height map and a 2D occupancy grid. These are inherently embedded into a memory and time efficient core data structure organized as a Tree of SkipLists. Compared to the well-known Octree representation, our approach exhibits a better time efficiency, thanks to its simple and highly parallelizable computational structure, and a similar memory footprint when mapping large workspaces. Peculiarly within the realm of mapping for robot navigation, our framework supports real-time erosion and re-integration of measurements upon reception of optimized poses from the sensor tracker, so as to improve continuously the accuracy of the map.
Daniele De Gregorio, Luigi Di Stefano
ICRA2
2016 Pairwise Registration by Local Orientation Cues
abstract
Abstract Inspired by recent work on robust and fast computation of 3D Local Reference Frames (LRFs), we propose a novel pipeline for coarse registration of 3D point clouds. Key to the method are: (i) the observation that any two corresponding points endowed with an LRF provide a hypothesis on the rigid motion between two views, (ii) the intuition that feature points can be matched based solely on cues directly derived from the computation of the LRF, (iii) a feature detection approach relying on a saliency criterion which captures the ability to establish an LRF repeatably. Unlike related work in literature, we also propose a comprehensive experimental evaluation based on diverse kinds of data (such as those acquired by laser scanners, Kinect and stereo cameras) as well as on quantitative comparison with respect to other methods. We also address the issue of setting the many parameters that characterize coarse registration pipelines fairly and realistically. The experimental evaluation vouches that our method can handle effectively data acquired by different sensors and is remarkably fast.
Alioscia Petrelli, Luigi Di Stefano
Comput. Graph. Forum2
2016 A Global Hypothesis Verification Framework for 3D Object Recognition in Clutter
abstract
Pipelines to recognize 3D objects despite clutter and occlusions usually end up with a final verification stage whereby recognition hypotheses are validated or dismissed based on how well they explain sensor measurements. Unlike previous work, we propose a Global Hypothesis Verification (GHV) approach which regards all hypotheses jointly so as to account for mutual interactions. GHV provides a principled framework to tackle the complexity of our visual world by leveraging on a plurality of recognition paradigms and cues. Accordingly, we present a 3D object recognition pipeline deploying both global and local 3D features as well as shape and color. Thereby, and facilitated by the robustness of the verification process, diverse object hypotheses can be gathered and weak hypotheses need not be suppressed too early to trade sensitivity for specificity. Experiments demonstrate the effectiveness of our proposal, which significantly improves over the state-of-art and attains ideal performance (no false negatives, no false positives) on three out of the six most relevant and challenging benchmark datasets.
Aitor Aldoma, Federico Tombari, Luigi Di Stefano, Markus Vincze
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 RGB-D Visual Search with Compact Binary Codes
abstract
As integration of depth sensing into mobile devices is likely forthcoming, we investigate on merging appearance and shape information for mobile visual search. Accordingly, we propose an RGB-D search engine architecture that can attain high recognition rates with peculiarly moderate bandwidth requirements. Our experiments include a comparison to the CDVS (Compact Descriptors for Visual Search) pipeline, candidate to become part of the MPEG-7 standard, and contribute to elucidate on the merits and limitations of joint deployment of depth and color in mobile visual search.
Alioscia Petrelli, Danilo Pau, Emanuele Plebani, Luigi Di Stefano
3DV4
2015 Large-scale and drift-free surface reconstruction using online subvolume registration
abstract
Depth cameras have helped commoditize 3D digitization of the real-world. It is now feasible to use a single Kinect-like camera to scan in an entire building or other large-scale scenes. At large scales, however, there is an inherent chal-lenge of dealing with distortions and drift due to accumu-lated pose estimation errors. Existing techniques suffer from one or more of the following: a) requiring an expensive offline global optimization step taking hours to compute; b) needing a full second pass over the input depth frames to correct for accumulated errors; c) relying on RGB data alongside depth data to optimize poses; or d) requiring the user to create explicit loop closures to allow gross alignment errors to be resolved. In this paper, we present a method that addresses all of these issues. Our method supports online model correction, without needing to reprocess or store any input depth data. Even while performing global correction of a large 3D model, our method takes only minutes rather than hours to compute. Our model does not require any explicit loop closures to be detected and, finally, relies on depth data alone, allowing operation in low-lighting conditions. We show qualitative results on many large scale scenes, high-lighting the lack of error and drift in our reconstructions. We compare to state of the art techniques and demonstrate large-scale dense surface reconstruction “in the dark”, a capability not offered by RGB-D techniques. 1.
Nicola Fioraio, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Luigi Di Stefano, Shahram Izadi
CVPR4
2015 Learning a Descriptor-Specific 3D Keypoint Detector
abstract
Keypoint detection represents the first stage in the majority of modern computer vision pipelines based on automatically established correspondences between local descriptors. However, no standard solution has emerged yet in the case of 3D data such as point clouds or meshes, which exhibit high variability in level of detail and noise. More importantly, existing proposals for 3D keypoint detection rely on geometric saliency functions that attempt to maximize repeatability rather than distinctiveness of the selected regions, which may lead to sub-optimal performance of the overall pipeline. To overcome these shortcomings, we cast 3D keypoint detection as a binary classification between points whose support can be correctly matched by a predefined 3D descriptor or not, thereby learning a descriptor-specific detector that adapts seamlessly to different scenarios. Through experiments on several public datasets, we show that this novel approach to the design of a keypoint detector represents a flexible solution that, nonetheless, can provide state-of-the-art descriptor matching performance.
Samuele Salti, Federico Tombari, Riccardo Spezialetti, Luigi Di Stefano
ICCV4
2015 Volume-Based Semantic Labeling with Signed Distance Functions
Tommaso Cavallari, Luigi Di Stefano
PSIVT2
2015 Traffic sign detection via interest region extraction
Samuele Salti, Alioscia Petrelli, Federico Tombari, Nicola Fioraio, Luigi Di Stefano
Pattern Recognit.5
2015 Synergistic Change Detection and Tracking
abstract
Visual tracking in image streams acquired by static cameras is usually based on change detection and recursive Bayesian estimation, such an approach laying at the core of many practical applications. Yet, the interaction between the change detector and the Bayesian filter is typically designed heuristically. Differently, this paper develops a sound framework to model and implement a bidirectional communication flow between the two processes. In our Bayesian loop, change detection provides well-defined observation likelihood to the recursive filter and the filter prediction provides an informative prior to the change detector, which deploys Bayesian reasoning alike. The loop is developed for the two major variants of Bayesian filters used in tracking, namely the Kalman filter and the particle filter. Experiments on publicly available videos and a novel challenging data set show that the proposed interaction scheme outperforms several state-of-the-art trackers.
Samuele Salti, Alessandro Lanza, Luigi Di Stefano
IEEE Trans. Circuits Syst. Video Technol.3
2014 Interest Points via Maximal Self-Dissimilarities
Federico Tombari, Luigi Di Stefano
ACCV (2)2
2014 Automatic detection of pole-like structures in 3D urban environments
abstract
This work aims at automatic detection of man-made pole-like structures in scans of urban environments acquired by a 3D sensor mounted on top a moving vehicle. Pole-like structures, such as e.g. road signs and streetlights, are widespread in these environments, and their reliable detection is relevant to applications dealing with autonomous navigation, facility damage detection, city planning and maintenance. Yet, due to the characteristic thin shape, detection of man-made pole-like structures is significantly prone to both noise as well as occlusions and clutter, the latter being pervasive nuisances when scanning urban environments. Our approach is based on a “local” stage, whereby local features are classified and clustered together, followed by a “global” stage aimed at further classification of candidate entities. The proposed pipeline turns out effective in experiments on a standard publicly available dataset as well as on a challenging dataset acquired during the project for validation purposes.
Federico Tombari, Nicola Fioraio, Tommaso Cavallari, Samuele Salti, Alioscia Petrelli, Luigi Di Stefano
IROS6
2014 SHOT: Unique signatures of histograms for surface and texture description
Samuele Salti, Federico Tombari, Luigi Di Stefano
Comput. Vis. Image Underst.3
2013 Joint Detection, Tracking and Mapping by Semantic Bundle Adjustment
abstract
In this paper we propose a novel Semantic Bundle Adjustment framework whereby known rigid stationary objects are detected while tracking the camera and mapping the environment. The system builds on established tracking and mapping techniques to exploit incremental 3D reconstruction in order to validate hypotheses on the presence and pose of sought objects. Then, detected objects are explicitly taken into account for a global semantic optimization of both camera and object poses. Thus, unlike all systems proposed so far, our approach allows for solving jointly the detection and SLAM problems, so as to achieve object detection together with improved SLAM accuracy.
Nicola Fioraio, Luigi Di Stefano
CVPR2
2013 Keypoints from Symmetries by Wave Propagation
abstract
The paper conjectures and demonstrates that repeatable keypoints based on salient symmetries at different scales can be detected by a novel analysis grounded on the wave equation rather than the heat equation underlying tradi-tional Gaussian scale–space theory. While the image struc-tures found by most state-of-the-art detectors, such as blobs and corners, occur typically on planar highly textured sur-faces, salient symmetries are widespread in diverse kinds of images, including those related to untextured objects, which are hardly dealt with by current feature-based recog-nition pipelines. We provide experimental results on stan-dard datasets and also contribute with a new dataset fo-cused on untextured objects. Based on the positive exper-imental results, we hope to foster further research on the promising topic of scale invariant analysis through the wave equation. 1.
Samuele Salti, Alessandro Lanza, Luigi Di Stefano
CVPR3
2013 BOLD Features to Detect Texture-less Objects
abstract
Object detection in images withstanding significant clutter and occlusion is still a challenging task whenever the object surface is characterized by poor informative content. We propose to tackle this problem by a compact and distinctive representation of groups of neighboring line segments aggregated over limited spatial supports and invariant to rotation, translation and scale changes. Peculiarly, our proposal allows for leveraging on the inherent strengths of descriptor-based approaches, i.e. robustness to occlusion and clutter and scalability with respect to the size of the model library, also when dealing with scarcely textured objects.
Federico Tombari, Alessandro Franchi, Luigi Di Stefano
ICCV3
2013 Multimodal cue integration through Hypotheses Verification for RGB-D object recognition and 6DOF pose estimation
abstract
This paper proposes an effective algorithm for recognizing objects and accurately estimating their 6DOF pose in scenes acquired by a RGB-D sensor. The proposed method is based on a combination of different recognition pipelines, each exploiting the data in a diverse manner and generating object hypotheses that are ultimately fused together in an Hypothesis Verification stage that globally enforces geometrical consistency between model hypotheses and the scene. Such a scheme boosts the overall recognition performance as it enhances the strength of the different recognition pipelines while diminishing the impact of their specific weaknesses. The proposed method outperforms the state-of-the-art on two challenging benchmark datasets for object recognition comprising 35 object models and, respectively, 176 and 353 scenes.
Aitor Aldoma, Federico Tombari, Johann Prankl, Andreas Richtsfeld, Luigi Di Stefano, Markus Vincze
ICRA5
2013 A traffic sign detection pipeline based on interest region extraction
abstract
In this paper we present a pipeline for automatic detection of traffic signs in images. The proposed system can deal with high appearance variations, which typically occur in traffic sign recognition applications, especially with strong illumination changes and dramatic scale changes. Unlike most existing systems, our pipeline is based on interest regions extraction rather than a sliding window detection scheme. The proposed approach has been specialized and tested in three variants, each aimed at detecting one of the three categories of Mandatory, Prohibitory and Danger traffic signs. Our proposal has been evaluated experimentally within the German Traffic Sign Detection Benchmark competition.
Samuele Salti, Alioscia Petrelli, Federico Tombari, Nicola Fioraio, Luigi Di Stefano
IJCNN5
2013 Performance Evaluation of 3D Keypoint Detectors
Federico Tombari, Samuele Salti, Luigi Di Stefano
Int. J. Comput. Vis.3
2013 On-line Support Vector Regression of the transition model for the Kalman filter
Samuele Salti, Luigi Di Stefano
Image Vis. Comput.2
2012 A Global Hypotheses Verification Method for 3D Object Recognition
Aitor Aldoma, Federico Tombari, Luigi Di Stefano, Markus Vincze
ECCV (3)3
2012 Performance Evaluation of Full Search Equivalent Pattern Matching Algorithms
abstract
Pattern matching is widely used in signal processing, computer vision, and image and video processing. Full search equivalent algorithms accelerate the pattern matching process and, in the meantime, yield exactly the same result as the full search. This paper proposes an analysis and comparison of state-of-the-art algorithms for full search equivalent pattern matching. Our intention is that the data sets and tests used in our evaluation will be a benchmark for testing future pattern matching algorithms, and that the analysis concerning state-of-the-art algorithms could inspire new fast algorithms. We also propose extensions of the evaluated algorithms and show that they outperform the original formulations.
Wanli Ouyang, Federico Tombari, Stefano Mattoccia, Luigi Di Stefano, Wai-kuen Cham
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Adaptive Appearance Modeling for Video Tracking: Survey and Evaluation
abstract
Long-term video tracking is of great importance for many applications in real-world scenarios. A key component for achieving long-term tracking is the tracker's capability of updating its internal representation of targets (the appearance model) to changing conditions. Given the rapid but fragmented development of this research area, we propose a unified conceptual framework for appearance model adaptation that enables a principled comparison of different approaches. Moreover, we introduce a novel evaluation methodology that enables simultaneous analysis of tracking accuracy and tracking success, without the need of setting application-dependent thresholds. Based on the proposed framework and this novel evaluation methodology, we conduct an extensive experimental comparison of trackers that perform appearance model adaptation. Theoretical and experimental analyses allow us to identify the most effective approaches as well as to highlight design choices that favor resilience to errors during the update process. We conclude the paper with a list of key open research challenges that have been singled out by means of our experimental comparison.
Samuele Salti, Andrea Cavallaro, Luigi Di Stefano
IEEE Trans. Image Process.3
2011 Energy-aware objects abandon / removal detection
abstract
A major issue for video surveillance embedded systems is the need to continuously perform a number of highly demanding operations even when the analyzed scene does not show peculiar or interesting features, so that power consumption is a critical issue. In this paper we present a low-power multimodal embedded video surveillance system aimed at detecting objects abandoned/removed in/from a static monitored scene. Energy-awareness is achieved by means of an efficient and scalable objects abandon/removal detection algorithm, a Linux governor that controls CPU frequency and operating mode so as to establish an optimal trade-off between fulfilling the application efficiency-accuracy requirements and maximizing battery life and, finally, a pyroelectric infrared sensor that allows to wake up the CPU only when video processing is actually needed.
Alessandro Lanza, Michele Magno, Davide Brunelli, Luigi Di Stefano, Luca Benini
AVSS4
2011 Background subtraction by non-parametric probabilistic clustering
abstract
We present a background subtraction approach aimed at efficiency and robustness to common source of disturbance such as gradual and sudden illumination changes, camera gain and exposure variations, noise. At each new frame, a non-parametric mixture-based probabilistic clustering is performed to segment the image into changed and unchanged pixels with respect to a fixed background. A two-components mixture, a two-dimensional discrete feature space, a non-parametric model for the components likelihood and a proper initial guess are the key ingredients of this novel algorithm that, besides dealing effectively with the discrimination of photometric and semantic changes, exhibits very high computational efficiency. Experiments are presented, proving the achieved state-of-the-art robustness-efficiency trade-off.
Alessandro Lanza, Samuele Salti, Luigi Di Stefano
AVSS3
2011 On the repeatability of the local reference frame for partial shape matching
abstract
We investigate on local reference frames (LRF) deployed with 3D descriptors to achieve invariance to objects' pose. We address the task of matching together partial views of surfaces and propose an experimental study on a large corpus of real data which allows for clearly ranking existing LRF proposals based on their repeatability. Then, drawing inspiration from analysis of the experimental findings, we formulate a new proposal which, in particular, peculiarly includes a procedure aimed at estimating a repeatable LRF also at border features, which is very important when matching partial views of surfaces. Experiments show that the new proposal neatly outperforms existing methods in terms of repeatability, is computationally very efficient and provide relevant benefits in practical applications
Alioscia Petrelli, Luigi Di Stefano
ICCV2
2011 A combined texture-shape descriptor for enhanced 3D feature matching
abstract
Motivated by the increasing availability of 3D sensors capable of delivering both shape and texture information, this paper presents a novel descriptor for feature matching in 3D data enriched with texture. The proposed approach stems from the theory of a recently proposed descriptor for 3D data which relies on shape only, and represents its generalization to the case of multiple cues associated with a 3D mesh. The proposed descriptor, dubbed CSHOT, is demonstrated to notably improve the accuracy of feature matching in challenging object recognition scenarios characterized by the presence of clutter and occlusions.
Federico Tombari, Samuele Salti, Luigi Di Stefano
ICIP3
2011 Online learning for automatic segmentation of 3D data
abstract
We propose a method to perform automatic segmentation of 3D scenes based on a standard classifier, whose learning model is continuously improved by means of new samples, and a grouping stage, that enforces local consistency among classified labels. The new samples are automatically delivered to the system by a feedback loop based on a feature selection approach that exploits the outcome of the grouping stage. By experimental results on several datasets we demonstrate that the proposed online learning paradigm is effective in increasing the accuracy of the whole 3D segmentation thanks to the improvement of the learning model of the classifier by means of newly acquired, unsupervised data.
Federico Tombari, Luigi Di Stefano, Simone Giardino
IROS2
2011 Statistical Change Detection by the Pool Adjacent Violators Algorithm
abstract
In this paper, we present a statistical change detection approach aimed at being robust with respect to the main disturbance factors acting in real-world applications such as illumination changes, camera gain and exposure variations, noise. We rely on modeling the effects of disturbance factors on images as locally order-preserving transformations of pixel intensities plus additive noise. This allows us to identify within the space of all of the possible image change patterns the subspace corresponding to disturbance factors effects. Hence, scene changes can be detected by a-contrario testing the hypothesis that the measured pattern is due to disturbance factors, that is, by computing a distance between the pattern and the subspace. By assuming additive Gaussian noise, the distance can be computed within a maximum likelihood nonparametric isotonic regression framework. In particular, the projection of the pattern onto the subspace is computed by an O(N) iterative procedure known as Pool Adjacent Violators algorithm.
Alessandro Lanza, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Efficient template matching for multi-channel images
abstract
Template matching is a computationally intensive problem aimed at locating a template within a image. When dealing with images having more than one channel, the computational burden becomes even more dramatic. For this reason, in this paper we investigate on a methodology to speed-up template matching on multi-channel images without deteriorating the outcome of the search. In particular, we propose a fast, exhaustive technique based on the Zero-mean Normalized Cross-Correlation (ZNCC) inspired from previous work related to grayscale images. Experimental testing performed over thousands of template matching instances demonstrates the efficiency of our proposal.
Stefano Mattoccia, Federico Tombari, Luigi Di Stefano
Pattern Recognit. Lett.3
2011 Adaptive Low Resolution Pruning for fast Full Search-equivalent pattern matching
Federico Tombari, Wanli Ouyang, Luigi Di Stefano, Wai-kuen Cham
Pattern Recognit. Lett.3
2010 On the Use of Implicit Shape Models for Recognition of Object Categories in 3D Data
Samuele Salti, Federico Tombari, Luigi Di Stefano
ACCV (3)3
2010 Accurate and Efficient Background Subtraction by Monotonic Second-Degree Polynomial Fitting
abstract
We present a background subtraction approach aimed at efficiency and accuracy also in presence of common sources of disturbance such as illumination changes, camera gain and exposure variations, noise. The novelty of the proposal relies on a-priori modeling the local effect of disturbs on small neighborhoods of pixel intensities as a monotonic, homogeneous, second-degree polynomial transformation plus additive Gaussian noise. This allows for classifying pixels as changed or unchanged by an efficient inequality-constrained least-squares fitting procedure. Experiments prove that the approach is state-of-the-art in terms of efficiency-accuracy tradeoff on challenging sequences characterized by disturbs yielding sudden and strong variations of the background appearance.
Alessandro Lanza, Federico Tombari, Luigi Di Stefano
AVSS3
2010 Unique Signatures of Histograms for Local Surface Description
Federico Tombari, Samuele Salti, Luigi Di Stefano
ECCV (3)3
2010 Stereo for robots: Quantitative evaluation of efficient and low-memory dense stereo algorithms
abstract
Despite the significant number of stereo vision algorithms proposed in literature in the last decade, most proposals are notably computationally demanding and/or memory hungry so that it is unfeasible to employ them in application scenarios requiring real-time or near real-time processing on platforms with limited resources such as embedded devices. In this paper, we have selected the subset of proposals that appears more suited to the above requirements and, since literature lacks a proper comparison between these methods, we propose a quantitative experimental evaluation aimed at highlighting the best performing approach under the two criteria of accuracy and efficiency. The evaluation is performed on a standard benchmark dataset as well as on a novel dataset, acquired by means of an active technique, characterized by realistic working conditions.
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano
ICARCV3
2010 A 3D reconstruction system based on improved spacetime stereo
abstract
Spacetime stereo is a promising technique for accurate 3D reconstruction based on randomly varying illumination and temporal integration of the stereo matching cost. In this paper we show that the standard spacetime stereo approach can be improved in terms of accuracy of disparity estimation and convergence speed by adoption of suitable matching algorithms based on adaptive support windows. We also present a practical and cost-effective 3D reconstruction system that deploys the proposed improved spacetime method together with cheap commercial off-the-shelf hardware (a PC, a stereo camera and a projector). Experimental results show that the proposed system can yield rapidly accurate 3D reconstruction of various types of objects and faces.
Federico Tombari, Luigi Di Stefano, Stefano Mattoccia, Andrea Mainetti
ICARCV2
2010 Robust and efficient background subtraction by quadratic polynomial fitting
abstract
We present a background subtraction algorithm aimed at efficiency and robustness to common sources of disturbance such as illumination changes, camera gain and exposure variations, noise. The approach relies on modeling the local effect of disturbance factors on a neighborhood of pixel intensities as a second-degree polynomial transformation plus additive Gaussian noise. This allows for classifying pixels as changed or unchanged by a simple least-squares polynomial fitting procedure. Experimental results prove that the approach is state-of-the-art in challenging sequences characterized by sources of disturbance yielding sudden and strong background appearance changes.
Alessandro Lanza, Federico Tombari, Luigi Di Stefano
ICIP3
2010 Mobile Visual Search using Smart-M3
abstract
Mobile Visual Search is a new technology that aims at linking physical objects with digital information by taking pictures of them using a mobile device with on-board camera. In this paper we present an easily extendible and adaptable framework for Mobile Visual Search applications that takes advantage of the Smart-M3 interoperability platform. After introducing Smart-M3, we describe the proposed architecture and illustrate its use in two novel visual search scenarios.
Alessandro Franchi, Luigi Di Stefano, Tullio Salmon Cinotti
ISCC2
2010 Energy aware multimodal embedded video surveillance
abstract
One of the major challenges in embedded system is reduction of power consumption. So far most of the microprocessors provide power saving mechanisms by changing the power operating mode as well as the frequency and core voltage at runtime. One of best methods is the management of available resources and operating systems like Linux define subsystems for power consumption management which require coordination and cooperation of hardware, kernel, and user-space applications, offering power savings options when the CPU is active as well as when it is inactive. In this paper we present a multimodal embedded visual surveillance system for the detection of abandoned/removed objects which exploits a Linux governor to control CPU frequency and operating mode to establish an optimal trade-off between fulfilling the application response time and accuracy requirements and maximizing battery life. To adopt an aggressive power management and to keep the whole system in sleep mode, a pyroelctric infrared sensor is also used to wake up the CPU only when video processing is actually needed.
Michele Magno, Alessandro Lanza, Davide Brunelli, Luigi Di Stefano, Luca Benini
VLSI-SoC4
2009 Enhanced Low-Resolution Pruning for Fast Full-Search Template Matching
Stefano Mattoccia, Federico Tombari, Luigi Di Stefano
ACIVS3
2009 A Template Analysis Methodology to Improve the Efficiency of Fast Matching Algorithms
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano, Fabio Regoli, Riccardo Viti
ACIVS3
2009 Bayesian Order-Consistency Testing with Class Priors Derivation for Robust Change Detection
abstract
In this paper we propose a formalization of change detection as a Bayesian order-consistency test, based on the assumption that disturbance factors such as illumination changes and variations of camera parameters do not change the ordering between noiseless intensities within a neighborhood of pixels. The assumption of additive, zero-mean, i.i.d. gaussian noise allows for testing the composite order-consistency hypothesis by efficient computation of the marginal likelihood. Moreover, since the above formalization enables to incorporate changed/unchanged class priors seamlessly, we also propose a simple method to derive informative priors based on the calculation of marginal likelihoods at reduced resolution. Experimental results on challenging test sequences characterized by sudden and strong illumination changes prove the effectiveness of the proposed approach.
Alessandro Lanza, Luigi Di Stefano, Luca Soffritti
AVSS2
2009 Multimodal Abandoned/Removed Object Detection for Low Power Video Surveillance Systems
abstract
Low-cost and low-power video surveillance systems based on networks of wireless video sensors will enter soon the marketplace with the promise of flexibility, quick deployment and providing accurate and real-time visual data. Energy autonomy and efficiency of the implemented algorithms are undoubtedly the primary design challenges to be addressed on systems subject to low computational capabilities and memory constraints. In this paper we present a low-power video sensor node designed for low-cost video surveillance which is able to detect abandoned and removed objects. The system exploits multi-modal sensor integration which saves on-board power consumption. In particular a pyroelectric infrared (PIR) sensor is exploited to optimize the use of the camera, grabbing images only when required in order to obtain the maximum efficiency from event recognition. Our fixed-point ARM-based approach is characterized in terms of runtime execution and power consumption, while efficiency is demonstrated by experimental results and compared with floating point implementations.
Michele Magno, Federico Tombari, Davide Brunelli, Luigi Di Stefano, Luca Benini
AVSS4
2009 Full-Search-Equivalent Pattern Matching with Incremental Dissimilarity Approximations
abstract
This paper proposes a novel method for fast pattern matching based on dissimilarity functions derived from the Lp norm, such as the Sum of Squared Differences (SSD) and the Sum of Absolute Differences (SAD). The proposed method is full-search equivalent, i.e. it yields the same results as the Full Search (FS) algorithm. In order to pursue computational savings the method deploys a succession of increasingly tighter lower bounds of the adopted Lp norm-based dissimilarity function. Such bounding functions allow for establishing a hierarchy of pruning conditions aimed at skipping rapidly those candidates that cannot satisfy the matching criterion. The paper includes an experimental comparison between the proposed method and other full-search equivalent approaches known in literature, which proves the remarkable computational efficiency of our proposal.
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 An Evaluation Methodology for Image Mosaicing Algorithms
Pietro Azzari, Luigi Di Stefano, Stefano Mattoccia
ACIVS2
2008 Graffiti Detection Using a Time-Of-Flight Camera
Federico Tombari, Luigi Di Stefano, Stefano Mattoccia, Andrea Zanetti
ACIVS2
2008 Multi-view Access Monitoring and Singularization in Interlocks
abstract
We present a method aimed at monitoring access to interlocks and secured entrance areas, which deploys two views in order to robustly perform intrusion detection and singularization. The main contributions are represented by an original approach to perform background subtraction, which is particularly robust against sudden illumination changes, shadows and photometric distortions, and by the use of a feature extraction and classification approach which allows to reliably determine an estimation of the number of people currently occupying the monitored area. Our system is designed to operate in very small interlocks and can work in a substantially unstructured environment.
Luigi Di Stefano, Federico Tombari, Stefano Mattoccia, Matteo Balasso
AVSS1
2008 Classification and evaluation of cost aggregation methods for stereo correspondence
abstract
In the last decades several cost aggregation methods aimed at improving the robustness of stereo correspondence within local and global algorithms have been proposed. Given the recent developments and the lack of an appropriate comparison, in this paper we survey, classify and compare experimentally on a standard data set the main cost aggregation approaches proposed in literature. The experimental evaluation addresses both accuracy and computational requirements, so as to outline the best performing methods under these two criteria.
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano, Elisa Addimanda
CVPR3
2008 Reliable rejection of mismatching candidates for efficient ZNCC template matching
abstract
This paper presents a method that reduces the computational cost of template matching based on the zero-mean normalized cross-correlation (ZNCC) without compromising the accuracy of the results. A very effective condition is determined at a small and fixed cost that allow to rapidly detect a large number of mismatching candidates with no need to compute the ZNCC score. Then, thanks to the use of an additional set of conditions, the computation of the whole ZNCC function is typically required only for a very small number of candidates. Experimental results demonstrate the effectiveness of our approach.
Stefano Mattoccia, Federico Tombari, Luigi Di Stefano
ICIP3
2008 Markerless Augmented Reality Using Image Mosaics
Pietro Azzari, Luigi Di Stefano, Federico Tombari, Stefano Mattoccia
ICISP2
2008 Near real-time stereo based on effective cost aggregation
abstract
Recent research activity on stereo matching has proved the efficacy of local approaches based on advanced cost aggregation strategies in accurately retrieving 3D information. However, accuracy is typically achieved at expense of computational efficiency, with best methods being far from meeting real-time requirements. On the other side, basic real-time local algorithms relying on a rectangular correlation window suffer from significant ambiguity along depth borders and untextured areas. This work proposes a novel local approach aimed at maximizing the speed-accuracy trade-off by means of an efficient segmentation-based cost aggregation strategy.
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano, Elisa Addimanda
ICPR3
2008 Fast Full-Search Equivalent Template Matching by Enhanced Bounded Correlation
abstract
We propose a novel algorithm, referred to as enhanced bounded correlation (EBC), that significantly reduces the number of computations required to carry out template matching based on normalized cross correlation (NCC) and yields exactly the same result as the full search algorithm. The algorithm relies on the concept of bounding the matching function: finding an efficiently computable upper bound of the NCC rapidly prunes those candidates that cannot provide a better NCC score with respect to the current best match. In this framework, we apply a succession of increasingly tighter upper bounding functions based on Cauchy-Schwarz inequality. Moreover, by including an online parameter prediction step into EBC, we obtain a parameter free algorithm that, in most cases, affords computational advantages very similar to those attainable by optimal offline parameter tuning. Experimental results show that the proposed algorithm can significantly accelerate a full-search equivalent template matching process and outperforms state-of-the-art methods.
Stefano Mattoccia, Federico Tombari, Luigi Di Stefano
IEEE Trans. Image Process.3
2007 Stereo Vision Enabling Precise Border Localization Within a Scanline Optimization Framework
Stefano Mattoccia, Federico Tombari, Luigi Di Stefano
ACCV (2)3
2007 Robust Multi-View Change Detection
abstract
We present a multi-view change detection approach aimed at being robust \nwith respect to common “disturbance factors” yielding image changes in realworld \napplications. Disturbance factors causing “slow” or “fast-and-global” \nimage variations, such as light changes and dynamic adjustments of camera \nparameters (e.g. auto-exposure and auto-gain control), are dealt with by a \nproper single-view change detector run independently on each view. The \ncomputed change masks are then fused into a “synergy mask” defined into a \ncommon virtual top-view, so as to detect and filter-out “fast-and-local” image \nchanges due to physical points lying on the ground surface (e.g. shadows cast \nby moving objects and light spots hitting the ground surface).
Alessandro Lanza, Luigi Di Stefano, Jérôme Berclaz, François Fleuret, Pascal Fua
BMVC2
2007 Segmentation-Based Adaptive Support for Accurate Stereo Correspondence
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano
PSIVT3
2006 People Tracking Using a Time-of-Flight Depth Sensor
abstract
Visually track several moving persons engaged in close interactions is known to be a very hard problem, though 3-D approaches based on stereo vision and plan-view maps offer much promise for dealing effectively with major issues such as occlusions and quick changes in body pose and appearance. However, in case of untextured scenes due to homogeneous objects or poor illumination, stereo-based tracking systems rapidly drop their performance. In this work, we present a real time people tracking system able to work even under severe low-lighting conditions. The system relies on a novel active sensor that provides brightness and depth images based on a Time of Flight (TOF) technology. The tracking algorithm is simple yet efficient, being based on geometrical constraints and invariants. Experiments accomplished under changing lighting conditions and involving multiple people closely interacting with each other have proved the reliability of the system.
Alessandro Bevilacqua, Luigi Di Stefano, Pietro Azzari
AVSS2
2006 Detecting Changes in Grey Level Sequences by ML Isotonic Regression
abstract
We present a robust and efficient change detection algorithm for grey-level sequences. A deep investigation of the effects of disturbance factors (illumination changes and automatic or manual adjustments of the camera transfer function, such as AGC, AE and \gamma-correction) on image brightness allows to assume locally an order-preservation of pixel intensities. By a simple statistical modelling of camera noise, an ML isotonic regression procedure can thus be applied to perform change detection. Although the proposed approach may be used as a stand-alone pixel-level change detector, here we apply it to reduced-resolution images. In fact, we aim at using the algorithm as the coarse-level of a coarse-to-fine change detector we presented in [2].
Alessandro Lanza, Luigi Di Stefano
AVSS2
2006 Template Matching Based on the L_p Norm Using Sufficient Conditions with Incremental Approximations
abstract
This paper proposes a novel algorithm aimed at speeding-up template matching based on the L_p norm. The algorithm is exhaustive, i.e. it yields the same results as a Full Search (FS) template matching process, and is based on the deployment of tight lower bounds that can be derived by using together the triangular inequality and partial evaluations of the L_p norm. In order to deploy this, template and image subwindows are properly partitioned. The experimental results prove that the proposed algorithm allows speeding-up the FS process and also (when applied to the L_2 norm) the exhaustive approach based on the Fast Fourier Transform.
Federico Tombari, Stefano Mattoccia, Luigi Di Stefano
AVSS3
2005 An effective real-time mosaicing algorithm apt to detect motion through background subtraction using a PTZ camera
abstract
Nowadays, many visual surveillance systems exploit pan/tilt/zoom (PTZ) cameras to increase the field of view of a surveyed area. The background subtraction technique is widespread to detect moving objects with a high accuracy using one stationary camera. Extending such algorithms to work with moving cameras requires to have a background mosaic at one's disposal. Many solutions using mosaic background subtraction have been proposed, which offer real time capabilities or high quality of the detected objects. However, most of them rely on prior assumptions which limit the camera motion or the algorithm to work with a depth field of view only. In this work we propose some innovative solutions to achieve a real time mosaic background apt to work with existing background subtraction algorithms to yield excellent foreground object masks. Extensive experiments accomplished on challenging indoor and outdoor scenes permit to assess the quality of the mosaic as well as of the detected moving masks.
Pietro Azzari, Luigi Di Stefano, Alessandro Bevilacqua
AVSS2
2005 Coarse-to-fine strategy for robust and efficient change detectors
abstract
We present a novel approach to change detection based on a coarse-to-fine strategy. An efficient coarse-level detection is proposed that filters out most of the possible false changes, thus attaining reliable and tight coarse-grain super-masks of the truly changed areas. The subsequent fine-level detection can thus "focus the attention" just on the "interesting" parts of the frame and perform a robust selective background updating procedure by considering the complement of these masks. Besides, the analysis of a strip of pixels surrounding each coarse-grain blob allows to infer information on light changes possibly occurring in the blob's vicinity. Although any algorithm can be used as the final fine-level detection, here we show how the approach applies to a particular algorithm we devised, based on a non-parametric statistical modelling of the camera noise.
Alessandro Bevilacqua, Luigi Di Stefano, Alessandro Lanza
AVSS2
2005 An effective multi-stage background generation algorithm
abstract
In this paper we present a new background generation algorithm and show experimental results aimed at assessing its performance comparatively with respect to two other representative approaches. The algorithm is able to extract a stationary background from a short bootstrap sequence in which moving objects can also be present. The method works with pixel-wise temporal statistics and consists of three subsequent stages. The first two try to isolate for each pixel the stationary background process from the possible foreground processes due to the moving objects covering the pixel, thus voting the background process temporal median as the good background value. The third stage completes the background generation by means of a non-parametric statistical model of the temporal camera noise, inferred from the statistics computed in the previous stage.
Alessandro Bevilacqua, Luigi Di Stefano, Alessandro Lanza
AVSS2
2005 Using local and global object's information to track vehicles in urban scenes
abstract
In intelligent transportation systems (ITS's) vehicle tracking is necessary to permit high-level analysis, such as vehicle counting or classification. Nowadays, the need for a precise vehicle behavior analysis is growing mainly in urban intersections. Typical urban traffic scenes contain high-cluttered areas where static and dynamic occlusions take place and objects are missed. Tracking systems relying on monocular cameras are widespread. However, often they are misled by the complicated object interactions occurring in those areas, thus yielding errors in the higher level modules. The real-time tracking system we have conceived relies on an algorithm exploiting local and global information from corner points and whole object's features that allows us to keep track of many different objects in challenging urban scenarios. We assess our results through extensive on-field testing by manually extracting the ground truth from different sequences taken by real world traffic monitoring systems.
Alessandro Bevilacqua, Luigi Di Stefano, Stefano Vaccari
AVSS2
2005 A novel approach to change detection based on a coarse-to-fine strategy
abstract
We present a novel approach to the change detection problem based on a coarse-to-fine strategy. The basic idea consists in assigning to an efficient preliminary coarse-level detection the task to filter out the well known possible false changes (e.g., those due to camera noise and small displacements, or to scene illumination changes). This provides the subsequent fine-level detection with reliable supermasks of the true changed areas in the scene. In this way, the fine-level detection can "focus the attention" on limited parts of the frames, thus yielding remarkable advantages in terms of computational efficiency. Here, just a coarse-level detection algorithm based on background subtraction and on the concept of structure is presented, to stress that any pixel-level algorithm can be used afterwards and benefit in terms of robustness as well as of computational efficiency.
Alessandro Bevilacqua, Luigi Di Stefano, Alessandro Lanza, Gianpaolo Capelli
ICIP (2)2
2005 ZNCC-based template matching using bounded partial correlation
Luigi Di Stefano, Stefano Mattoccia, Federico Tombari
Pattern Recognit. Lett.1
2004 An efficient change detection algorithm based on a statistical non-parametric camera noise model
abstract
In this paper we present a change detection algorithm for grey level sequences based on the background subtraction technique, which achieves a good trade-off between time performance and detection quality. The basic idea consists in separating the background process into a deterministic background process and a stochastic camera noise process. The assumption that statistics of the camera noise for a pixel only depends on its current grey level allows to infer a nonparametric statistical camera noise model once and for all arising from a short bootstrap sequence. Hence, 256 couples of lower and upper deterministic thresholds are extracted, to be used in the background subtraction step. While the deterministic nature of the background model as well as of the thresholds lead to an efficient algorithm, utilising 256 couples of different thresholds results in a very sensitive detection. Experimental results allow to assess both the efficiency and the effectiveness of the method we devised.
Alessandro Bevilacqua, Luigi Di Stefano, Alessandro Lanza
ICIP2
2004 A fast area-based stereo matching algorithm
Luigi Di Stefano, Massimiliano Marchionni, Stefano Mattoccia
Image Vis. Comput.1
2003 A Change-Detection Algorithm Based on Structure and Colour
abstract
The paper proposes a novel change-detection algorithm for automated video surveillance applications. The algorithm is based on the idea of incorporating into the background model a set of simple low-level features capable of capturing effectively "structural" (i.e. robust with respect to illumination variations) information. Thanks to this approach, and unlike most conventional change-detection algorithms, the proposed algorithm is capable of handling correctly still and slow objects as well as of working properly throughout very long time spans. Moreover, the algorithm can naturally interact with the higher-level processing modules found in advanced video-based surveillance systems in order to allow for flexible and intelligent background maintenance.
Luigi Di Stefano, Stefano Mattoccia, Martino Mola
AVSS1
2003 A sufficient condition based on the Cauchy-Schwarz inequality for efficient template matching
abstract
The paper proposes a technique aimed at reducing the number of calculations required to carry out an exhaustive template matching process based on the normalized cross correlation (NCC). The technique deploys an effective sufficient condition, relying on the recently introduced concept of bounded partial correlation that allows rapid elimination of the points that cannot provide a better cross-correlation score with respect to the current best candidate. In this paper we devise a novel sufficient condition based on the Cauchy-Schwarz inequality and compare the experimental results with those attained using the standard NCC-based template matching algorithm and the already known sufficient condition based on the Jensen inequality.
Luigi Di Stefano, Stefano Mattoccia
ICIP (1)1
2003 Fast template matching using bounded partial correlation
Luigi Di Stefano, Stefano Mattoccia
Mach. Vis. Appl.1
2002 Quantitative evaluation of area-based stereo matching
abstract
This paper presents a quantitative evaluation of two area-based stereo matching approaches aimed at real-time applications. The former approach follows a standard scheme based on a double matching phase (DMP) and aimed at rejecting unreliable disparities by enforcing the left-right consistency constraint. This approach is representative of several implementations known in literature. The latter approach, proposed recently, relies on a single matching phase (SMF) and rejects unreliable matches by detecting violations of the uniqueness constraints. We review first the two approaches and analyse their differences. Successively, we provide quantitative results, using standard stereo pairs with ground truth, aimed at evaluating and comparing the two approaches in terms of matching reliability and speed.
Luigi Di Stefano, Massimiliano Marchionni, Stefano Mattoccia, Giovanni Neri
ICARCV1
2000 Perception of Depth Information by Means of a Wire-Actuated Haptic Interface
abstract
The VIDET project is aimed at investigating the possibility of developing a wearable robotic system for helping the mobility of visually impaired persons. The basic idea involves the conversion of real-time depth data gathered through stereo-vision into a virtual, "bas-relief" model perceivable by means of a haptic interface. In this paper we describe the real-time stereo system, review the basic principles of the two main haptic devices developed so far, and present new experimental results concerning extraction of depth data by the stereo system and haptic perception of the virtual model recovered from stereo-data.
Paolo Arcara, Luigi Di Stefano, Stefano Mattoccia, Claudio Melchiorri, Gabriele Vassura
ICRA2
1997 A new phase extraction algorithm for phase profilometry
Luigi Di Stefano, Francis M. Boland
Mach. Vis. Appl.1
1995 Detection of Circular Objects by Wave Propagation on a Mesh-Connected Computer
Rita Cucchiara, Luigi Di Stefano, Massimo Piccardi
J. Parallel Distributed Comput.2
1993 Processing of variable size images on a cellular array: Performance analysis with the Abingdon Cross Benchmark
abstract
Handling a continuous flow of variable size images is a requirement for real time computer vision machines. A modular system based on a small size SIMD cellular array of 1-bit processing elements has been developed with this goal in mind and it is now evaluated against the Abingdon Cross Benchmark specifications. The benchmark tests the combination of algorithms and architecture and generates a quality factor expressed as the ratio of the image lateral size and the processing time. The examined machine supports an efficient means to automatically partition, process and reconstruct images larger than the array size. The authors briefly describe the system, discuss the selected algorithms and present performance results and estimates for several system configurations.>
Massimo Piccardi, Luigi Di Stefano, Rita Cucchiara, Tullio Salmon Cinotti
ASAP2