Massimo Bertozzi

dblp:89/1570 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0003-1463-5384ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorSystems, architecture and hardware · 3 · 2 first-authorComputer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CalibBEV: LiDAR-Camera Calibration via BEV Alignment
abstract
We present CalibBEV, a novel Bird’s Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shared 3D spatial representation, enabling accurate and robust cross-modal calibration. CalibBEV extracts sensor-wise BEV features from each modality using domain-specific architectures and estimates the calibration matrix through a two-step alignment process. First, we perform an implicit alignment by regressing a coarse calibration matrix directly from the BEV features. To ease this alignment, we enforce semantic consistency between BEV representations across modalities using a contrastive loss inspired by CLIP, guiding both networks toward a unified feature space. In the second step, we leverage our BEV formulation to explicitly align the features of one modality with the other, refining the initial coarse estimate into a final, more accurate calibration matrix. CalibBEV significantly outperforms prior point-to-pixel matching methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 51% and 68%, and the Relative Translation Error (RTE) by 80% and 91%, respectively, compared to previous methods.
Filippo D'Addeo, Lorenzo Cipelli, Adriano Cardace, Emanuele Ghelfi, Andrea Zinelli, Massimo Bertozzi
WACV6
2026 End-to-End LiDAR-Camera Calibration via Multi-Modal Correspondences Estimation and Explicit BEV Alignment
abstract
Abstract In this work, we present a Bird’s Eye View (BEV) Alignment approach for the LiDAR-Camera calibration task. Building upon previous BEV-based work, we extract sensor-wise BEV features from each input modality using domain-specific architectures. Then, we employ a CNN-based encoder to align the two BEVs and estimate the calibration matrix. However, corresponding 2D and 3D features may be spatially distant in BEV space, and as a consequence the encoder alone might struggle to learn the height dimension and estimate the correct registration matrix. To address this, we introduce an implicit alignment step to cross-attend the downsampled 3D features with those from RGB for computing point-to-pixel correspondences and estimating a coarse calibration matrix. To improve the implicit alignment, we also enforce the prediction of correct point-to-pixel correspondences by direct supervision of the similarity matrix computed into the cross attention module. Then, the coarsely aligned 3D features and the RGB features are fed to the BEV Alignment step, in which the CNN-based encoder refines the coarse estimate into a final, more accurate calibration matrix. Notably, both the steps are optimized in an end-to-end fashion. Our method significantly outperforms previous point-to-pixel matching methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 74% and 79%, and the Relative Translation Error (RTE) by 90% and 95%, respectively, compared to previous methods.
Lorenzo Cipelli, Filippo D'Addeo, Emanuele Ghelfi, Marcello Ceresini, Andrea Bertogalli, Federico Pirazzoli, Massimo Bertozzi
Int. J. Comput. Vis.7
2025 Improving 3D Multi-View Object Detection via Explicit Query Supervision
abstract
Perception is a crucial aspect of an autonomous driving system. One essential task is represented by multi-camera 3D object detection, which allows an intelligent vehicle to detect the surrounding obstacles using a camera-only setup. Currently, there are a lot of different approaches trying to solve this task, with many of them being transformer-based. Specifically, most of these make use of object queries instead of a Bird's Eye View plane to directly represent the set of possible detections and avoid any post-processing operation, like non-maxima suppression. However, the ambiguous supervision caused by the bipartite matching loss typically leads to training instability. To overcome this limitation, we propose an additional module able to “push” the object queries toward the locations that more likely contain obstacles, providing both better insights into their position to the detection module and stabilizing the bipartite matching during training. We evaluate our proposal against different objet queries-based baselines both on the nuScenes dataset test and validation sets. Specifically, compared to the lightweight PETR architecture, we highlight an increase of 1.6% both in NDS and mAP under the same configuration settings.
Filippo D'Addeo, Andrea Zinelli, Massimo Bertozzi
IV3
2025 Mamba-ST: State Space Model for Efficient Style Transfer
abstract
The goal of style transfer is, given a content image and a style source, generating a new image preserving the content but with the artistic representation of the style source. Most of the state-of-the-art architectures use transformers or diffusion-based models to perform this task, despite the heavy computational burden that they require. In particular, transformers use self- and cross-attention layers which have large memory footprint, while diffusion models require high inference time. To overcome the above, this paper explores a novel design of Mamba, an emergent State-Space Model (SSM), called Mamba-ST, to perform style transfer. To do so, we adapt Mamba linear equation to simulate the behavior of cross-attention layers, which are able to combine two separate embeddings into a single output, but drastically reducing memory usage and time complexity. We modified the Mamba's inner equations so to accept inputs from, and combine, two separate data streams. To the best of our knowledge, this is the first attempt to adapt the equations of SSMs to a vision task like style transfer without requiring any other module like cross-attention or custom normalization layers. An extensive set of experiments demonstrates the superiority and efficiency of our method in performing style transfer compared to transformers and diffusion models. Results show improved quality in terms of both ArtFID and FID metrics. Code is available at https://github.com/FilippoBotti/MambaST.
Filippo Botti, Alex Ergasti, Leonardo Rossi, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001
WACV6
2025 Swin2-MoSE: A new single image supersolution model for remote sensing
abstract
Abstract Due to the limitations of current optical and sensor technologies and the high cost of updating them, the spectral and spatial resolution of satellites may not always meet desired requirements. For these reasons, Remote‐Sensing Single‐Image Super‐Resolution (RS‐SISR) techniques have gained significant interest. In this paper, Swin2‐MoSE model is proposed, an enhanced version of Swin2SR. The model introduces MoE‐SM, an enhanced Mixture‐of‐Experts (MoE) to replace the Feed‐Forward inside all Transformer block. MoE‐SM is designed with Smart‐Merger, and new layer for merging the output of individual experts, and with a new way to split the work between experts, defining a new per‐example strategy instead of the commonly used per‐token one. Furthermore, it is analyzed how positional encodings interact with each other, demonstrating that per‐channel bias and per‐head bias can positively cooperate. Finally, the authors propose to use a combination of Normalized‐Cross‐Correlation (NCC) and Structural Similarity Index Measure (SSIM) losses, to avoid typical MSE loss limitations. Experimental results demonstrate that Swin2‐MoSE outperforms any Swin derived models by up to 0.377–0.958 dB (PSNR) on task of , and resolution‐upscaling ( and OLI2MSI datasets). It also outperforms SOTA models by a good margin, proving to be competitive and with excellent potential, especially for complex tasks. Additionally, an analysis of computational costs is also performed. Finally, the efficacy of Swin2‐MoSE is shown, applying it to a semantic segmentation task (SeasoNet dataset). Code and pretrained are available on https://github.com/IMPLabUniPr/swin2‐mose/tree/official_code
Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001
IET Image Process.4
2025 MARS: Paying More Attention to Visual Attributes for Text-Based Person Search
abstract
Text-Based Person Search (TBPS) is a problem that gained significant interest within the research community. The task is that of retrieving one or more images of a specific individual based on a textual description. The multi-modal nature of the task requires learning representations that bridge text and image data within a shared latent space. Existing TBPS systems face two major challenges. One is defined as inter-identity noise that is due to the inherent vagueness and imprecision of text descriptions, and it indicates how descriptions of visual attributes can be generally associated to different people; the other is the intra-identity variations, which are all those nuisances, e.g., pose, illumination, that can alter the visual appearance of the same textual attributes for a given subject. To address these issues, this article presents a novel TBPS architecture named Mae-Attribute-Relation-Sensitive (MARS), which enhances current state-of-the-art models by introducing two key components: a Visual Reconstruction Loss and an Attribute Loss. The former employs a Masked AutoEncoder trained to reconstruct randomly masked image patches with the aid of the textual description. In doing so the model is encouraged to learn more expressive representations and textual–visual relations in the latent space. The attribute loss, instead, balances the contribution of different types of attributes, defined as adjective–noun chunks of text. This loss ensures that every attribute is taken into consideration in the person retrieval process. Extensive experiments on three commonly used datasets, namely CUHK-PEDES, ICFG-PEDES, and RSTPReid, report performance improvements, with significant gains in the Mean Average Precision (mAP) metric w.r.t. the current state of the art. Code will be available at https://github.com/ErgastiAlex/MARS .
Alex Ergasti, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Memory-Augmented Online Video Anomaly Detection
abstract
The ability to understand the surrounding scene is of paramount importance for Autonomous Vehicles (AVs). This paper presents a system capable to work in an online fashion, giving an immediate response to the arise of anomalies surrounding the AV, exploiting only the videos captured by a dash-mounted camera. Our architecture, called MOVAD, relies on two main modules: a Short-Term Memory Module to extract information related to the ongoing action, implemented by a Video Swin Transformer (VST), and a Long-Term Memory Module injected inside the classifier that considers also remote past information and action context thanks to the use of a Long-Short Term Memory (LSTM) network. The strengths of MOVAD are not only linked to its excellent performance, but also to its straightforward and modular architecture, trained in a end-to-end fashion with only RGB frames with as less assumptions as possible, which makes it easy to implement and play with. We evaluated the performance of our method on Detection of Traffic Anomaly (DoTA) dataset, a challenging collection of dash-mounted camera videos of accidents. After an extensive ablation study, MOVAD is able to reach an AUC score of 82.17%, surpassing the current state-of-the-art by +2.87 AUC. Our code and pretrained are available online on https://github.com/IMPLabUniPr/movad/tree/movad_vad
Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001
ICASSP4
2024 CFTS-GAN: Continual Few-Shot Teacher Student for Generative Adversarial Networks
Munsif Ali, Leonardo Rossi, Massimo Bertozzi
ICPR (25)3
2023 FrankenMask: Manipulating semantic masks with transformers for face parts editing
abstract
In this paper, we propose FrankenMask, a novel framework that allows swapping and rearranging face parts in semantic masks for automatic editing of shape-related facial attributes. This is a novel yet challenging task as substituting face parts in a semantic mask requires to account for possible spatial misalignment and the adaptation of surrounding regions. We obtain such a feature by combining a Transformer encoder to learn the spatial relationships of facial parts, with an encoder–decoder architecture, which reconstructs a complete mask from the composition of local parts. Reconstruction and attribute classification results demonstrate the effective synthesis of facial images, while showing the generation of accurate and plausible facial attributes. Code is available at https://github.com/TFonta/FrankenMask_semantic.
Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Leonardo Galteri, Stefano Berretti, Massimo Bertozzi, Andrea Prati 0001
Pattern Recognit. Lett.6
2023 Unsupervised Discovery and Manipulation of Continuous Disentangled Factors of Variation
abstract
Learning a disentangled representation of a distribution in a completely unsupervised way is a challenging task that has drawn attention recently. In particular, much focus has been put in separating factors of variation (i.e., attributes) within the latent code of a Generative Adversarial Network (GAN). Achieving that permits control of the presence or absence of those factors in the generated samples by simply editing a small portion of the latent code. Nevertheless, existing methods that perform very well in a noise-to-image setting often fail when dealing with a real data distribution, i.e., when the discovered attributes need to be applied to real images. However, some methods are able to extract and apply a style to a sample but struggle to maintain its content and identity, while others are not able to locally apply attributes and end up achieving only a global manipulation of the original image. In this article, we propose a completely (i.e., truly ) unsupervised method that is able to extract a disentangled set of attributes from a data distribution and apply them to new samples from the same distribution by preserving their content. This is achieved by using an image-to-image GAN that maps an image and a random set of continuous attributes to a new image that includes those attributes. Indeed, these attributes are initially unknown and they are discovered during training by maximizing the mutual information between the generated samples and the attributes’ vector. Finally, the obtained disentangled set of continuous attributes can be used to freely manipulate the input samples. We prove the effectiveness of our method over a series of datasets and show its application on various tasks, such as attribute editing, data augmentation, and style transfer.
Tomaso Fontanini, Luca Donati, Massimo Bertozzi, Andrea Prati 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Arbitrary Point Cloud Upsampling with Spherical Mixture of Gaussians
abstract
Generating dense point clouds from sparse raw data benefits downstream 3D understanding tasks, but existing models are limited to a fixed upsampling ratio or to a short range of integer values. In this paper, we present APU-SMOG, a Transformer-based model for Arbitrary Point cloud Upsampling (APU). The sparse input is firstly mapped to a Spherical Mixture of Gaussians (SMOG) distribution, from which an arbitrary number of points can be sampled. Then, these samples are fed as queries to the Transformer decoder, which maps them back to the target surface. Extensive qualitative and quantitative evaluations show that APU-SMOG outperforms state-of-the-art fixed-ratio methods, while effectively enabling upsampling with any scaling factor, including non-integer values, with a single trained model. The code will be made available.
Anthony Dell'Eva, Marco Orsingher, Massimo Bertozzi
3DV3
2022 Revisiting PatchMatch Multi-View Stereo for Urban 3D Reconstruction
abstract
In this paper, a complete pipeline for image-based 3D reconstruction of urban scenarios is proposed, based on PatchMatch Multi-View Stereo (MVS). Input images are firstly fed into an off-the-shelf visual SLAM system to extract camera poses and sparse keypoints, which are used to initialize PatchMatch optimization. Then, pixelwise depths and normals are iteratively computed in a multi-scale framework with a novel depth-normal consistency loss term and a global refinement algorithm to balance the inherently local nature of PatchMatch. Finally, a large-scale point cloud is generated by back-projecting multi-view consistent estimates in 3D. The proposed approach is carefully evaluated against both classical MVS algorithms and monocular depth networks on the KITTI dataset, showing state of the art performances.
Marco Orsingher, Paolo Zani, Paolo Medici, Massimo Bertozzi
IV4
2018 Introduction to the Special Issue on Applications of Mechatronic and Embedded Systems (MESA) in ITS
abstract
Embedded systems result from the integration between mechanical and electronic components (hardware) and the information-driven functions (software). Embedded systems play a key role in the development of mechatronic systems, which involves finding an optimal balance between the basic mechanical structure, sensor and actuators, automatic digital information processing and control.
Massimo Bertozzi, Primo Zingaretti
IEEE Trans. Intell. Transp. Syst.1
2018 Self-Localization Based on Visual Lane Marking Maps: An Accurate Low-Cost Approach for Autonomous Driving
abstract
Autonomous driving in public roads requires precise localization within the range of few centimeters. Even the best localization systems based on GNSS cannot always reach this level of precision, especially in an urban environment, where the signal is disturbed by surrounding buildings and artifacts. Recent works have shown the advantage of using maps as a precise, robust, and reliable way of localization. Typical approaches use the set of current readings from the vehicle sensors to estimate its position on the map. The approach presented in this paper exploits a short-range visual lane marking detector and a dead reckoning system to construct a registry of the detected back lane markings corresponding to the last 240 m driven. This information is used to search in the map the most similar section, to determine the vehicle localization in the map reference. Additional filtering is used to obtain a more robust estimation for the localization. The accuracy obtained is sufficiently high to allow autonomous driving in a narrow road. The system uses a low-cost architecture of sensors and the algorithm is light enough to run on low-power embedded architecture.
Rafael P. Vivacqua, Massimo Bertozzi, Pietro Cerri, Felipe N. Martins, Raquel Frizera Vassallo
IEEE Trans. Intell. Transp. Syst.2
2015 360° Detection and tracking algorithm of both pedestrian and vehicle using fisheye images
abstract
All-around view is a mandatory element for autonomous vehicles. The European V-Charge project seeks to develop an autonomous vehicle using only low-cost sensors. This paper presents a detection and tracking algorithm that covers all the area around the vehicle using 4 fisheye cameras only. The algorithm is able to detect pedestrians and vehicles and track them, using cylindrical images. This paper presents the whole pipeline, from the image un-warping to the classification and the tracking algorithms, together with some results.
Massimo Bertozzi, Luca Castangia, Stefano Cattani, Antonio Prioletti, Pietro Versari
Intelligent Vehicles Symposium1
2015 Introduction to the Special Issue on Mechatronic and Embedded Systems and Applications in ITS
abstract
The papers in this special issue were presented at the The Mechatronic and Embedded Technologies in Intelligent Transportation Systems Symposium which includes contributions in technologies, methodologies and, in particular, applications of mechatronic and embedded systems in any aspect of ITSs, from automatic vehicle localization and monitoring to autonomous vehicles and environmental perception.
Massimo Bertozzi, Yanqing Gao, Primo Zingaretti
IEEE Trans. Intell. Transp. Syst.1
2011 VIAC: An out of ordinary experiment
abstract
This paper presents the preliminary results of VIAC, the VisLab Intercontinental Autonomous Challenge, a test of autonomous driving along an unknown route from Italy to China. It took 3 months to run the entire test; all data have been logged, including all data generated by the sensors, vehicle data, and GPS info. This huge amount of information has been packed during the trip, compressed, and transferred back to Parma for further processing. This data is now ready for a deep analysis of the various systems performance, with the aim of virtually running the whole trip multiple times with improved versions of the software. This paper discusses some preliminary figures obtained by the analysis of the data collected during the test. More information will be generated by a deeper analysis, which will take additional time, being the data about 40 terabyte in size.
Massimo Bertozzi, Luca Bombini, Alberto Broggi, Michele Buzzoni, Elena Cardarelli, Stefano Cattani, Pietro Cerri, Alessandro Coati, Stefano Debattisti, Andrea Falzoni, Rean Isabella Fedriga, Mirko Felisa, Luca Gatti, Alessandro Giacomazzo, Paolo Grisleri, Maria Chiara Laghi, Luca Mazzei, Paolo Medici, Matteo Panciroli, Pier Paolo Porta, Paolo Zani, Pietro Versari
Intelligent Vehicles Symposium1
2007 Pedestrian detection by means of far-infrared stereo vision
Massimo Bertozzi, Alberto Broggi, Claudio Caraffi, Michael S. Del Rose, Mirko Felisa, G. Vezzoni
Comput. Vis. Image Underst.1
2006 Vehicle detection by means of stereo vision-based obstacles features extraction and monocular pattern analysis
abstract
This paper presents a stereo vision system for the detection and distance computation of a preceding vehicle. It is divided in two major steps. Initially, a stereo vision-based algorithm is used to extract relevant three-dimensional (3-D) features in the scene, these features are investigated further in order to select the ones that belong to vertical objects only and not to the road or background. These 3-D vertical features are then used as a starting point for preceding vehicle detection; by using a symmetry operator, a match against a simplified model of a rear vehicle's shape is performed using a monocular vision-based approach that allows the identification of a preceding vehicle. In addition, using the 3-D information previously extracted, an accurate distance computation is performed.
Gwenaëlle Toulminet, Massimo Bertozzi, Stéphane Mousset, Abdelaziz Bensrhair, Alberto Broggi
IEEE Trans. Image Process.2
2002 Artificial vision in road vehicles
abstract
The last few decades have witnessed the birth and growth of a new sensibility to transportation efficiency. In particular the need for efficient and improved people and goods mobility has pushed researchers to address the problem of intelligent transportation systems. This paper surveys the most advanced approaches to (partial) customization of the road following task, using on-board systems based on artificial vision. The functionalities of lane detection, obstacle detection and pedestrian detection are described and classified, and their possible application in future road vehicles is discussed.
Massimo Bertozzi, Alberto Broggi, Massimo Cellario, Alessandra Fascioli, Paolo Lombardi, Marco Porta
Proc. IEEE1
2002 Quintic G2-splines for the iterative steering ofvision-based autonomous vehicles
abstract
This paper presents a new motion planning primitive to be used for the iterative steering of vision-based autonomous vehicles. This primitive is a parameterized quintic spline, denoted as /spl eta/-spline, that allows interpolating an arbitrary sequence of points with overall second-order geometric (G/sup 2/-) continuity. Issues such as completeness, minimality, regularity, symmetry, and flexibility of these G/sup 2/-splines are addressed in the exposition. The development of the new primitive is tightly connected to the inversion control of nonholonomic car-like vehicles. The paper also exposes a supervisory strategy for iterative steering that integrates feedback vision data processing with the feedforward inversion control.
Aurelio Piazzi, Corrado Guarino Lo Bianco, Massimo Bertozzi, Alessandra Fascioli, Alberto Broggi
IEEE Trans. Intell. Transp. Syst.3
2001 Self-Calibration of a Stereo Vision System for Automotive Applications
abstract
In this paper a calibration method for on-board cameras used on the ARGO autonomous vehicle is presented. A number of markers have been placed on the vehicle's hood, which is framed by the vision system. Due to the knowledge of the markers' position it is possible to compute the camera position and orientation with respect to the vehicle. By using fast computations, this procedure, of basic importance when the camera head has pan-tilt capabilities, can be performed during autonomous driving, without slowing down normal operations.
Alberto Broggi, Massimo Bertozzi, Alessandra Fascioli
ICRA2
2000 Visual perception of obstacles and vehicles for platooning
abstract
Presents the methods for sensing obstacles and vehicles implemented on the University of Parma experimental vehicle (ARGO). The ARGO project is briefly described along with its main objectives; the prototype vehicle and its functionalities are presented. The perception of the environment is performed through the processing of images acquired from the vehicle. Details about the stereo vision-based detection of generic obstacles are given, along with a measurement of the performance of the method; then a new approach for leading vehicles detection is described, relying on symmetry detection in monocular images. The paper concludes with a description of the current implementation of the control system, based on a gain scheduled controller, which allows the vehicle to follow the road or other vehicles.
Alberto Broggi, Massimo Bertozzi, Alessandra Fascioli, Corrado Guarino Lo Bianco, Aurelio Piazzi
IEEE Trans. Intell. Transp. Syst.2
1999 Tools for code optimization and system evaluation of the image processing system PAPRICA-3
Massimo Bertozzi, Alberto Broggi
J. Syst. Archit.1
1998 Stereo inverse perspective mapping: theory and applications
Massimo Bertozzi, Alberto Broggi, Alessandra Fascioli
Image Vis. Comput.1
1998 GOLD: a parallel real-time stereo vision system for generic obstacle and lane detection
abstract
This paper describes the generic obstacle and lane detection system (GOLD), a stereo vision-based hardware and software architecture to be used on moving vehicles to increment road safety. Based on a full-custom massively parallel hardware, it allows to detect both generic obstacles (without constraints on symmetry or shape) and the lane position in a structured environment (with painted lane markings) at a rate of 10 Hz. Thanks to a geometrical transform supported by a specific hardware module, the perspective effect is removed from both left and right stereo images; the left is used to detect lane markings with a series of morphological filters, while both remapped stereo images are used for the detection of free-space in front of the vehicle. The output of the processing is displayed on both an on-board monitor and a control-panel to give visual feedbacks to the driver. The system was tested on the mobile laboratory (MOB-LAB) experimental land vehicle, which was driven for more than 3000 km along extra-urban roads and freeways at speeds up to 80 km/h, and demonstrated its robustness with respect to shadows and changing illumination conditions, different road textures, and vehicle movement.
Massimo Bertozzi, Alberto Broggi
IEEE Trans. Image Process.1
1997 A real-time oriented system for vehicle detection
Massimo Bertozzi, Alberto Broggi, Stefano Castelluccio
J. Syst. Archit.1
1996 A stereo vision system for real-time automotive obstacle detection
abstract
This work presents a system for obstacle detection in a pair of images acquired by a stereo vision device installed on a moving vehicle. The whole system is structured in a pipeline of two different computational engines: a massively parallel architecture, PAPRICA, devoted to low-level image processing and a traditional serial architecture running medium-level tasks. A geometrical transformation, based on the assumption of a flat road in front of the vehicle, is performed to remove the perspective effect from both images. The difference between the results is used for the detection of free-space in front of the vehicle, thus allowing to avoid the high computational tasks involved in traditional stereo vision approaches; the geometrical transformation is performed by a specific hardware device integrated in PAPRICA architecture. The system was tested on the MOB-LAB experimental land vehicle, which was driven for more than 3000 km along extra-urban roads and freeways at speeds up to 80 km/h, and demonstrated its robustness with respect to shadows and changing illumination conditions, different road textures, and vehicle movement.
Massimo Bertozzi, Alberto Broggi, Alessandra Fascioli
ICIP (2)1