Leonardo Taccari

dblp:99/10257 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0800-4893ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Color Is Not Enough: Dataset and Method for Identifying Relevant Traffic Lights in Driving Scenes
abstract
Accurate localization and classification of traffic lights in driving scenes are crucial for enhancing road scene understanding in various intelligent vehicles applications. However, determining which traffic lights are relevant for the ego-vehicle remains an under-explored challenge. In this paper, we address both thelocaltask of identifying the state and relevance of each traffic light in an image and the strictly relatedglobaltask of recommending the correct course of action for the ego-vehicle (should it stop?). We propose a novel architecture, which not only localizes each traffic light and identifies its relevance with respect to the ego-vehicle, but also generates a global recommendation. To address the scarcity of datasets with these types of annotations, we introduce the Verizon Connect Traffic Light Dataset (VZC-TLD), the first U.S. dataset that provides 3,000 images annotated with traffic light boxes, states, and relevance. Experimental results on both VZC-TLD and the DriveU Traffic Light Dataset (DTLD) show that our unified approach is indeed effective, and leads to significant improvements over approaches that do not exploit the synergies between the local and global tasks. Dataset is available at:https://leotac.github.io/vzc-tld
Tomaso Trinci, Simone Magistri, Tommaso Bianconcini, Leonardo Taccari, Leonardo Sarti, Francesco Sambo
IEEE Trans. Intell. Transp. Syst.4
2025 ViewpointDepth: A New Dataset for Monocular Depth Estimation Under Viewpoint Shifts
abstract
Monocular depth estimation is a critical task for autonomous driving and many other computer vision applications. While significant progress has been made in this field, the effects of viewpoint shifts on depth estimation models remain largely underexplored. This paper introduces a novel dataset and evaluation methodology to quantify the impact of different camera positions and orientations on monocular depth estimation performance. We propose a ground truth strategy based on homography estimation and object detection, eliminating the need for expensive LIDAR sensors. We collect a diverse dataset of road scenes from multiple viewpoints and use it to assess the robustness of a modern depth estimation model to geometric shifts. After assessing the validity of our strategy on a public dataset, we provide valuable insights into the limitations of current models and highlight the importance of considering viewpoint variations in real-world applications.
Aurel Pjetri, Stefano Caprasecca, Leonardo Taccari, Matteo Simoncini, Henrique Piñeiro Monteagudo, Wallace Walter, Douglas Coimbra de Andrade, Francesco Sambo, Andrew D. Bagdanov
IV3
2025 RendBEV: Semantic Novel View Synthesis for Self-Supervised Bird's Eye View Segmentation
abstract
Bird's Eye View (BEV) semantic maps have recently garnered a lot of attention as a useful representation of the environment to tackle assisted and autonomous driving tasks. However most of the existing work focuses on the fully supervised setting training networks on large annotated datasets. In this work we present RendBEV a new method for the self-supervised training of BEV semantic segmentation networks leveraging differentiable volumetric rendering to receive supervision from semantic perspective views computed by a 2D semantic segmentation model. Our method enables zero-shot BEV semantic segmentation and already delivers competitive results in this challenging setting. When used as pretraining to then fine-tune on labeled BEV ground truth our method significantly boosts performance in low-annotation regimes and sets a new state of the art when fine-tuning on all available labels.
Henrique Piñeiro Monteagudo, Leonardo Taccari, Aurel Pjetri, Francesco Sambo, Samuele Salti
WACV2
2024 Dynamic Bird's Eye View Reconstruction of Driving Accidents
abstract
The consequences of vehicle crashes are extremely costly, especially in industrial contexts, where the loss of income due to the vehicle unavailability while the incident is investigated adds to the damage produced by the event. The ongoing shift toward more connected vehicles, featuring sensors and cameras, offers the opportunity to alleviate such losses by speeding up the resolution of disputes. In this paper, we show how data routinely collected by connected vehicles can be fused to attain automatic reconstruction of the crash dynamic, a key element that has to be provided by drivers to submit a First Notification of Loss. We build upon state-of-the-art methods in areas such as SLAM, depth estimation and object detection to create a reconstruction of the scene with the vehicles involved localized both in space and time, which we present in an animated bird’s eye view. Our pipeline is evaluated on a challenging benchmark of real world videos and it is shown to create reliable reconstructions of the moment of the impact in more than 50% of scenes and overall good reconstructions in about 37% of them.
Marco Boschi, Luca De Luigi, Samuele Salti, Francesco Sambo, Douglas Coimbra de Andrade, Leonardo Taccari, Alex Quintero Garcia
IEEE Trans. Intell. Transp. Syst.6
2023 Lightweight and Effective Convolutional Neural Networks for Vehicle Viewpoint Estimation From Monocular Images
abstract
Vehicle viewpoint estimation from monocular images is a crucial component for autonomous driving vehicles and for fleet management applications. In this paper, we make several contributions to advance the state-of-the-art on this problem. We show the effectiveness of applying a smoothing filter to the output neurons of a Convolutional Neural Network (CNN) when estimating vehicle viewpoint. We point out the overlooked fact that, under the same viewpoint, the appearance of a vehicle is strongly influenced by its position in the image plane, which renders viewpoint estimation from appearance an ill-posed problem. We show how, by inserting in the model a CoordConv layer to provide the coordinates of the vehicle, we are able to solve such ambiguity and greatly increase performance. Finally, we introduce a new data augmentation technique that improves viewpoint estimation on vehicles that are closer to the camera or partially occluded. All these improvements let a lightweight CNN reach optimal results while keeping inference time low. An extensive evaluation on a viewpoint estimation benchmark (Pascal3D+) and on actual vehicle camera data (nuScenes) shows that our method significantly outperforms the state-of-the-art in vehicle viewpoint estimation, both in terms of accuracy and memory footprint.
Simone Magistri, Marco Boschi, Francesco Sambo, Douglas Coimbra de Andrade, Matteo Simoncini, Luca Kubin, Leonardo Taccari, Luca De Luigi, Samuele Salti
IEEE Trans. Intell. Transp. Syst.7
2022 Detection of Stop Sign Violations From Dashcam Data
abstract
In this article we present a novel machine learning pipeline for automatic detection of stop sign violations from dashcam videos, Inertial Measurement Units (IMU) and Global Positioning System (GPS) data. We developed a two-step approach, including a detector (Stop Sign Detector) capable of identifying stop signs presence, position, and size within video frames, followed by a classifier (Stop Violation Classifier) that assesses the presence of violations along with a severity score. The Stop Sign Detector is a deep convolutional neural network (CNN) for image classification, which leverages the information contained in its deeper layer feature maps in order to extract estimates of position and size of the detected stop signs. The Stop Violation Classifier fuses the information provided by the Stop Sign Detector with IMU/GPS data to assess the presence and severity of a stop sign violation. The proposed approach has been tested on several thousands of real-world videos, recorded from US vehicles, in all kinds of weather conditions, times of the day and environments. Our method achieves an area under the precision-recall curve of 94% with a required computational time of 2.4 seconds to process a 16-second video entirely on CPU.
Luca Bravi, Luca Kubin, Stefano Caprasecca, Douglas Coimbra de Andrade, Matteo Simoncini, Leonardo Taccari, Francesco Sambo
IEEE Trans. Intell. Transp. Syst.6
2022 Deep Crash Detection From Vehicular Sensor Data With Multimodal Self-Supervision
abstract
The ability to detect vehicle accidents from on-board sensor data is of the utmost importance to provide prompt assistance to prevent injuries and fatalities. In this article, we present a novel deep learning method capable of analyzing time series recorded from Inertial Measurement Units (IMU) and GPS devices to recognize the presence of an accident along with its severity. We propose a neural architecture capable of exploiting the different sensor streams (i.e., acceleration, gyroscope, and GPS speed), a multimodal contrastive self-supervised training procedure, and an ad-hoc stack of data augmentation techniques, specifically designed to counteract the extreme class imbalance and to improve the generalization capabilities of the whole pipeline. The proposed method has been validated against several state-of-the-art methods on a large and highly imbalanced dataset, composed of more than 200 thousand time series collected from US vehicles, with different vehicle sizes and traveling on different types of road. Our method achieves an average-precision score (AP) of 0.9 in the detection of crashes and 0.76 in the detection of severe crashes, significantly outperforming all the other approaches, and has small footprint and latency, so that it can easily be deployed on embedded devices.
Luca Kubin, Tommaso Bianconcini, Douglas Coimbra de Andrade, Matteo Simoncini, Leonardo Taccari, Francesco Sambo
IEEE Trans. Intell. Transp. Syst.5
2022 Unsafe Maneuver Classification From Dashcam Video and GPS/IMU Sensors Using Spatio-Temporal Attention Selector
abstract
In this paper, we propose a novel deep learning architecture to classify unsafe driving maneuvers from dashcam and IMU data. Such architecture processes the output of an object detection algorithm in combination with raw video frames and GPS/IMU data. At the core of the architecture there is a novel Spatio-Temporal Attention Selector (STAS) module, which (1) extracts features describing the evolution of each object in the scene over time and (2) leverages multi-head dot product attention to select the relevant ones,i.e., the dangerous ones or the ones in danger, to perform classification. We also introduce a simple but effective methodology to increase the benefit of fine-tuning the backbone network. Our method is shown to achieve higher performance than other approaches in the literature applying attention over single frames.
Matteo Simoncini, Douglas Coimbra de Andrade, Leonardo Taccari, Samuele Salti, Luca Kubin, Fabio Schoen, Francesco Sambo
IEEE Trans. Intell. Transp. Syst.3
2014 Maximum Throughput Network Routing Subject to Fair Flow Allocation
Edoardo Amaldi, Stefano Coniglio, Leonardo Taccari
ISCO3