VLDB 2026 Research / reviewers in the wild / expert
Filippo Ghilotti
dblp:415/2763
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
3D vision · 46% Autonomous driving · 30% Robot navigation and mapping · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › multimodal perception
LiDAR-camera fusion |
0.9 | 1 | 2025 | Self-Supervised Sparse Sensor Fusion for Long Range Perception · ICCV 2025 |
Robotics › Autonomous driving
perception |
0.9 | 1 | 2025 | Self-Supervised Sparse Sensor Fusion for Long Range Perception · ICCV 2025 |
Computer vision › 3D vision › 3d scene modeling
scene representation |
0.9 | 1 | 2025 | Self-Supervised Sparse Sensor Fusion for Long Range Perception · ICCV 2025 |
Robotics › Robot navigation and mapping
sensor fusion |
0.9 | 1 | 2025 | Self-Supervised Sparse Sensor Fusion for Long Range Perception · ICCV 2025 |
Robotics › Autonomous driving
trajectory prediction |
0.3 | 1 | 2025 | Self-Supervised Sparse Sensor Fusion for Long Range Perception · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
self-supervised pretraining · 0.9bird's-eye-view representation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniLiPs: Unified LiDAR Pseudo-Labeling with Geometry-Grounded Dynamic Scene DecompositionabstractUnlabeled LiDAR logs, in autonomous driving applications, are inherently a gold mine of dense 3D geometry hiding in plain sight - yet they are almost useless without human labels, highlighting a dominant cost barrier for autonomous-perception research. In this work we tackle this bottleneck by leveraging temporal-geometric consistency across LiDAR sweeps to lift and fuse cues from text and 2 Dvision foundation models directly into 3D, without any manual input. We introduce an unsupervised multimodal pseudo-labeling method relying on strong geometric priors learned from temporally accumulated LiDAR maps, alongside with a novel iterative update rule that enforces joint geometric-semantic consistency, and vice-versa detecting moving objects from inconsistencies. Our method simultaneously produces 3D semantic labels, 3D bounding boxes, and dense LiDAR scans, demonstrating robust generalization across three datasets. We experimentally validate that our method compares favorably to existing semantic segmentation and object detection pseudo-labeling methods, which often require additional manual supervision. We confirm that even a small fraction of our geometrically consistent, densified LiDAR improves depth prediction by 51.5 % and 22.0 % MAE in the 80-150 and 150-250 meters range, respectively. Filippo Ghilotti, Samuel Brucker, Nahku Saidy, Matteo Matteucci, Mario Bijelic, Felix Heide |
3DV | 1 |
| 2026 | Too Tiny to See: Hazardous Obstacle Detection Dataset and EvaluationabstractWe introduce a novel dataset and evaluation approach for long-range depth prediction of small objects that enables consistent comparison across direct time-of-flight (ToF) sensors and learned depth estimation methods. In autonomous driving, accurate depth perception is essential for identifying and locating surrounding elements and determining safe driving paths. Traditional depth metrics focus on distance accuracy but fail to evaluate a key factor at long ranges: distinguishing small, slightly elevated structures from the ground - crucial for anticipating obstacles and making safe driving decisions. At far distances, imagebased systems suffer from resolution limitations that tend to oversmooth the ground plane, causing elevated objects to be mistaken as texture patterns on the surface. Conversely, scanning LiDAR systems may return only a single point from an elevated object due to steep incident angles and sparse returns, preventing accurate differentiation from the ground. This hampers a fair comparison of object presence and shape. To address this, we propose a framework that evaluates how well the estimated point clouds preserve semantic content relative to ground-truth data. We leverage neural network-based feature extraction to assess structural similarity, enabling a modality-agnostic evaluation of object-level fidelity. Our method also supports analysis of the trade-off between resolution and accuracy, investigating performances across sensor types - such as highresolution cameras versus LiDAR - and conditions, including day and night scenarios. This enables a more comprehensive understanding of the capabilities and limitations of current depth prediction approaches in real-world settings. Topi Miekkala, Samuel Brucker, Stefanie Walz, Filippo Ghilotti, Andrea Ramazzina, Dominik Scheuble, Pasy Pyykonen, Mario Bijelic, Felix Heide |
3DV | 4 |
| 2025 | Self-Supervised Sparse Sensor Fusion for Long Range PerceptionabstractOutside of urban hubs, autonomous cars and trucks have to master driving on intercity highways. Safe, long-distance highway travel at speeds exceeding 100 km/h demands perception distances of at least 250 m, which is about five times the 50-100m typically addressed in city driving, to allow sufficient planning and braking margins. Increasing the perception ranges also allows to extend autonomy from light two-ton passenger vehicles to large-scale forty-ton trucks, which need a longer planning horizon due to their high inertia. However, most existing perception approaches focus on shorter ranges and rely on Bird's Eye View (BEV) representations, which incur quadratic increases in memory and compute costs as distance grows. To overcome this limitation, we built on top of a sparse representation and introduced an efficient 3D encoding of multi-modal and temporal features, along with a novel self-supervised pre-training scheme that enables large-scale learning from unlabeled camera-LiDAR data. Our approach extends perception distances to 250 meters and achieves an 26.6% improvement in mAP in object detection and a decrease of 30.5% in Chamfer Distance in LiDAR forecasting compared to existing methods, reaching distances up to 250 meters. Project Page: https://light.princeton.edu/lrs4fusion/ Edoardo Palladin, Samuel Brucker, Filippo Ghilotti, Praveen Narayanan, Mario Bijelic, Felix Heide |
ICCV | 3 |