EDBT 2026 Demo / reviewers in the wild / expert
Vitor Campagnolo Guizilini
dblp:81/7230 · also Vitor Campanholo Guizilini, Vitor Guizilini
· DBLP profile ↗
54ranked-venue papers
20as first author
30since 2021 · last 2026
0000-0002-8715-8307ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 19 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 12 first-author · 21 since 2021Systems, architecture and hardware · 15 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastMap: Revisiting Structure from Motion Through First-Order OptimizationabstractWe propose FastMap, a new global structure from motion method focused on speed and simplicity. Previous methods like COLMAP and GLOMAP are able to estimate high-precision camera poses, but suffer from poor scalability when the number of matched keypoint pairs becomes large, mainly due to the time-consuming process of secondorder Gauss-Newton optimization. Instead, we design our method solely based on first-order optimizers. To obtain maximal speedup, we identify and eliminate two key performance bottlenecks: computational complexity and the kernel implementation of each optimization step. Through extensive experiments, we show that FastMap is up to 10 × faster than COLMAP and GLOMAP with GPU acceleration and achieves comparable pose accuracy. Project webpage: https://jiahao.ai/fastmap. Muhammad Zubair Irshad, Igor Vasiljevic, Matthew R. Walter, Vitor Campagnolo Guizilini, Gregory Shakhnarovich |
3DV | 6 |
| 2025 | Incorporating Dense Metric Depth into Neural 3D Representations for View Synthesis and RelightingabstractCapturing photo realistic appearance and geometry of scenes is a fundamental problem in computer vision and graphics with a set of mature tools and solutions for content creation [12], [67], large scale scene mapping [5], augmented reality and cinematography [6], [80], [97]. Enthusiast level 3D photogrammetry, especially for small or tabletop scenes, has been supercharged by more capable smartphone cameras and new toolboxes like RealityCapture and NeRF-Studio. A subset of these solutions are geared towards view synthesis where the focus is on photo-realistic view interpolation rather than recovery of accurate scene geometry. These solutions take the “shape-radiance ambiguity”[58] into stride by decoupling the scene transmissivity (related to geometry) from the scene appearance prediction. But without diverse training views, several neural scene representations (e.g. [39], [73], [76]) are prone to poor shape reconstructions while estimating accurate appearance. Arkadeep Narayan Chaudhury, Igor Vasiljevic, Sergey Zakharov, Vitor Campagnolo Guizilini, Rares Ambrus, Srinivasa G. Narasimhan, Christopher G. Atkeson |
3DV | 4 |
| 2025 | GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusionabstract3D reconstruction from a single image is a long-standing problem in computer vision. Learning-based methods address its inherent scale ambiguity by leveraging increasingly large labeled and unlabeled datasets, to produce geometric priors capable of generating accurate predictions across domains. As a result, state of the art approaches show impressive performance in zero-shot relative and metric depth estimation. Recently, diffusion models have exhibited remarkable scalability and generalizable properties in their learned representations. However, because these models repurpose tools originally designed for image generation, they can only operate on dense ground-truth, which is not available for most depth labels, especially in real-world settings. In this paper we present GRIN, an efficient diffusion model designed to ingest sparse unstructured training data. We use image features with 3D geometric positional encodings to condition the diffusion process both globally and locally, generating depth predictions at a pixel-level. With comprehensive experiments across eight indoor and outdoor datasets, we show that GRIN establishes a new state of the art in zero-shot metric monocular depth estimation even when trained from scratch. Vitor Campagnolo Guizilini, Pavel Tokmakov, Achal Dave, Rares Ambrus |
3DV | 1 |
| 2025 | Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric DiffusionabstractCurrent methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation. Vitor Campagnolo Guizilini, Muhammad Zubair Irshad, Dian Chen 0005, Gregory Shakhnarovich, Rares Ambrus |
CVPR | 1 |
| 2025 | ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic GraspingabstractRobotic grasping is a cornerstone capability of embodied systems. Many methods directly output grasps from partial information without modeling the geometry of the scene, leading to suboptimal motion and even collisions. To address these issues, we introduce ZeroGrasp, a novel framework that simultaneously performs 3D reconstruction and grasp pose prediction in near real-time. A key insight of our method is that occlusion reasoning and modeling the spatial relationships between objects is beneficial for both accurate reconstruction and grasping. We couple our method with a novel large-scale synthetic dataset, which comprises 1M photo-realistic images, high-resolution 3D reconstructions and 11.3B physically-valid grasp pose annotations for 12K objects from the Objaverse-LVIS dataset. We evaluate Zero-Grasp on the GraspNet-1B benchmark as well as through real-world robot experiments. ZeroGrasp achieves state-of-the-art performance and generalizes to novel real-world objects by leveraging synthetic data. https://sh8.io/#/zerograsp Shun Iwase, Muhammad Zubair Irshad, Katherine Liu, Vitor Campagnolo Guizilini, Robert Lee, Takuya Ikeda, Ayako Amma, Koichi Nishiwaki, Kris Makoto Kitani, Rares Ambrus, Sergey Zakharov |
CVPR | 4 |
| 2025 | Learning Temporally Consistent Video Depth from Video Diffusion PriorsabstractThis work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: xdimlab.github.io/ChronoDepth. Jiahao Shao, Youmin Zhang 0008, Yujun Shen, Vitor Campagnolo Guizilini, Yue Wang 0041, Matteo Poggi, Yiyi Liao |
CVPR | 6 |
| 2025 | SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian SplattingabstractReconstructing articulated objects prevalent in daily environments is crucial for applications in augmented/virtual reality and robotics. However, existing methods face scalability limitations (requiring 3D supervision or costly annotations), robustness issues (being susceptible to local optima), and rendering shortcomings (lacking speed or photorealism). We introduce SplArt, a self-supervised, category-agnostic framework that leverages 3D Gaussian Splatting (3DGS) to reconstruct articulated objects and infer kinematics from two sets of posed RGB images captured at different articulation states, enabling real-time photorealistic rendering for novel viewpoints and articulations. SplArt augments 3DGS with a differentiable mobility parameter per Gaussian, achieving refined part segmentation. A multi-stage optimization strategy is employed to progressively handle reconstruction, part segmentation, and articulation estimation, significantly enhancing robustness and accuracy. SplArt exploits geometric self-supervision, effectively addressing challenging scenarios without requiring 3D annotations or category-specific priors. Evaluations on established and newly proposed benchmarks, along with applications to real-world scenarios using a handheld RGB camera, demonstrate SplArt's state-of-the-art performance and real-world practicality. Code is publicly available at https://github.com/ripl/splart. Shengjie Lin, Jiading Fang, Muhammad Zubair Irshad, Vitor Campagnolo Guizilini, Rares Ambrus, Gregory Shakhnarovich, Matthew R. Walter |
ICCV | 4 |
| 2025 | PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World UnderstandingabstractUnderstanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in reasoning and task planning for embodied agents, their ability to comprehend physical phenomena remains extremely limited.
To close this gap, we introduce PhysBench, a comprehensive benchmark designed to evaluate VLMs' physical world understanding capability across a diverse set of tasks.
PhysBench contains 10,002 entries of interleaved video-image-text data, categorized into four major domains: physical object properties, physical object relationships, physical scene understanding, and physics-based dynamics, further divided into 19 subclasses and 8 distinct capability dimensions.
Our extensive experiments, conducted on 75 representative VLMs, reveal that while these models excel in common-sense reasoning, they struggle with understanding the physical world---likely due to the absence of physical knowledge in their training data and the lack of embedded physical priors.
To tackle the shortfall, we introduce PhysAgent, a novel framework that combines the generalization strengths of VLMs with the specialized expertise of vision models, significantly enhancing VLMs' physical understanding across a variety of tasks, including an 18.4\% improvement on GPT-4o.
Furthermore, our results demonstrate that enhancing VLMs' physical world understanding capabilities can help embodied agents such as MOKA.
We believe that PhysBench and PhysAgent offer valuable insights and contribute to bridging the gap between VLMs and physical world understanding. [Project Page is here](https://physbench.github.io/) Wei Chow, Jiageng Mao, Boyi Li 0001, Daniel Seita, Vitor Campagnolo Guizilini, Yue Wang 0041 |
ICLR | 5 |
| 2025 | Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual OdometryabstractMonocular visual odometry is a key technology in various autonomous systems. Traditional feature-based methods suffer from failures due to poor lighting, insufficient texture, and large motions. In contrast, recent learning-based dense SLAM methods exploit iterative dense bundle adjustment to address such failure cases, and achieve robust and accurate localization in a wide variety of real environments, without depending on domain-specific supervision. However, despite its potential, the methods still struggle with scenarios involving large motion and object dynamics. In this study, we diagnose key weaknesses in a popular learning-based dense SLAM model (DROID-SLAM) by analyzing major failure cases on outdoor benchmarks and exposing various shortcomings of its optimization process. We then propose the use of self-supervised priors leveraging a frozen large-scale pre-trained monocular depth estimator to initialize the dense bundle adjustment process, leading to robust visual odometry without the need to fine-tune the SLAM backbone. Despite its simplicity, the proposed method demonstrates significant improvements on KITTI odometry, as well as the challenging DDAD benchmark. The project page: https://toyotafrc.github.io/SGInit-Proj/ Takayuki Kanai, Igor Vasiljevic, Vitor Campagnolo Guizilini, Kazuhiro Shintani |
IROS | 3 |
| 2024 | Towards Realistic Scene Generation with LiDAR Diffusion ModelsabstractDiffusion models (DMs) excel in photorealistic image synthesis, but their adaptation to LiDAR scene generation poses a substantial hurdle. This is primarily because DMs operating in the point space struggle to preserve the curve-like patterns and 3D geometry of LiDAR scenes, which consumes much of their representation power. In this paper, we propose LiDAR Diffusion Models (LiDMs) to generate LiDAR-realistic scenes from a latent space tailored to capture the realism of LiDAR scenes by incorporating geometric priors into the learning pipeline. Our method targets three major desiderata: pattern realism, geometry realism, and object realism. Specifically, we introduce curve-wise compression to simulate real-world LiDAR patterns, point-wise coordinate supervision to learn scene geometry, and patch-wise encoding for a full 3D object context. With these three core designs, our method achieves competitive performance on unconditional LiDAR generation in 64-beam scenario and state of the art on conditional LiDAR generation, while maintaining high efficiency compared to point-based DMs (up to 107× faster). Further-more, by compressing LiDAR scenes into a latent space, we enable the controllability of DMs with various conditions such as semantic maps, camera views, and text prompts. Our code and pretrained weights are available at htt ps: //github.com/hancyran/LiDAR-Diffusion. Haoxi Ran, Vitor Campagnolo Guizilini, Yue Wang 0041 |
CVPR | 2 |
| 2024 | NeRF-MAE: Masked AutoEncoders for Self-supervised 3D Representation Learning for Neural Radiance Fields
Muhammad Zubair Irshad, Sergey Zakharov, Vitor Campagnolo Guizilini, Adrien Gaidon, Zsolt Kira, Rares Ambrus |
ECCV (88) | 3 |
| 2024 | Zero-Shot Multi-object Scene Completion
Shun Iwase, Katherine Liu, Vitor Campagnolo Guizilini, Adrien Gaidon, Kris Makoto Kitani, Rares Ambrus, Sergey Zakharov |
ECCV (3) | 3 |
| 2024 | Transcrib3D: 3D Referring Expression Resolution through Large Language ModelsabstractIf robots are to work effectively alongside people, they must be able to interpret natural language references to objects in their 3D environment. Understanding 3D referring expressions is challenging—it requires the ability to both parse the 3D structure of the scene and correctly ground free-form language in the presence of distraction and clutter. We introduce Transcrib3D, an approach that brings together 3D detection methods and the emergent reasoning capabilities of large language models (LLMs). Transcrib3D uses text as the unifying medium, which allows us to sidestep the need to learn shared representations connecting multi-modal inputs, which would require massive amounts of annotated 3D data. As a demonstration of its effectiveness, Transcrib3D achieves state-of-the-art results on 3D reference resolution benchmarks, with a great leap in performance from previous multi-modality baselines. To improve upon zero-shot performance and facilitate local deployment on edge computers and robots, we propose self-correction for fine-tuning that trains smaller models, resulting in performance close to that of large models. We show that our method enables a real robot to perform pick-and-place tasks given queries that contain challenging referring expressions. Code will be available at https://ripl.github.io/Transcrib3D. Jiading Fang, Xiangshan Tan, Shengjie Lin, Igor Vasiljevic, Vitor Campagnolo Guizilini, Hongyuan Mei, Rares Ambrus, Gregory Shakhnarovich, Matthew R. Walter |
IROS | 5 |
| 2024 | Score Distillation via Reparametrized DDIMabstractWhile 2D diffusion models generate realistic, high-detail images, 3D shape generation methods like Score Distillation Sampling (SDS) built on these 2D diffusion models produce cartoon-like, over-smoothed shapes. To help explain this discrepancy, we show that the image guidance used in Score Distillation can be understood as the velocity field of a 2D denoising generative process, up to the choice of a noise term. In particular, after a change of variables, SDS resembles a high-variance version of Denoising Diffusion Implicit Models (DDIM) with a differently-sampled noise term: SDS introduces noise i.i.d. randomly at each step, while DDIM infers it from the previous noise predictions. This excessive variance can lead to over-smoothing and unrealistic outputs. We show that a better noise approximation can be recovered by inverting DDIM in each SDS update step. This modification makes SDS's generative process for 2D images almost identical to DDIM. In 3D, it removes over-smoothing, preserves higher-frequency detail, and brings the generation quality closer to that of 2D samplers. Experimentally, our method achieves better or similar 3D generation quality compared to other state-of-the-art Score Distillation methods, all without training additional neural networks or multi-view supervision, and providing useful insights into relationship between 2D and 3D asset generation with diffusion models. Artem Lukoianov, Haitz Sáez de Ocáriz Borde, Kristjan Greenewald, Vitor Campagnolo Guizilini, Timur M. Bagautdinov, Vincent Sitzmann, Justin Solomon 0001 |
NeurIPS | 4 |
| 2024 | $SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth EstimationabstractIncorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior, aiding in the generation of robust multi-view features for 3D scene understanding. In this paper, we explore the application of equivariant multi-view learning to depth estimation, not only recognizing its significance for computer vision and robotics but also addressing the limitations of previous research. Most prior studies have either overlooked equivariance in this setting or achieved only approximate equivariance through data augmentation, which often leads to inconsistencies across different reference frames. To address this issue, we propose to embed $SE(3)$ equivariance into the Perceiver IO architecture. We employ Spherical Harmonics for positional encoding to ensure 3D rotation equivariance, and develop a specialized equivariant encoder and decoder within the Perceiver IO architecture. To validate our model, we applied it to the task of stereo depth estimation, achieving state of the art results on real-world datasets without explicit geometric constraints or extensive data augmentation. Yinshuang Xu, Dian Chen 0005, Katherine Liu, Sergey Zakharov, Rares Ambrus, Kostas Daniilidis, Vitor Campagnolo Guizilini |
NeurIPS | 7 |
| 2023 | Viewpoint Equivariance for Multi-View 3D Object Detectionabstract3D object detection from visual sensors is a corner-stone capability of robotic systems. State-of-the-art methods focus on reasoning and decoding object bounding boxes from multi-view camera input. In this work we gain intuition from the integral role of multi-view consistency in 3D scene understanding and geometric learning. To this end, we introduce VEDet, a novel 3D object detection framework that exploits 3D multi-view geometry to improve localization through viewpoint awareness and equivariance. VEDet leverages a query-based transformer architecture and encodes the 3D scene by augmenting image features with positional encodings from their 3D perspective geometry. We design view-conditioned queries at the output level, which enables the generation of multiple virtual frames during training to learn viewpoint equivariance by enforcing multi-view consistency. The multi-view geometry injected at the input level as positional encodings and regularized at the loss level provides rich geometric cues for 3D object detection, leading to state-of-the-art performance on the nuScenes benchmark. The code and model are made available at https://github.com/TRI-ML/VEDet. Dian Chen 0005, Jie Li 0031, Vitor Campagnolo Guizilini, Rares Ambrus, Adrien Gaidon |
CVPR | 3 |
| 2023 | Towards Zero-Shot Scale-Aware Monocular Depth EstimationabstractMonocular depth estimation is scale-ambiguous, and thus requires scale supervision to produce metric predictions. Even so, the resulting models will be geometry-specific, with learned scales that cannot be directly transferred across domains. Because of that, recent works focus instead on relative depth, eschewing scale in favor of improved up-to-scale zero-shot transfer. In this work we introduce ZeroDepth, a novel monocular depth estimation framework capable of predicting metric scale for arbitrary test images from different domains and camera parameters. This is achieved by (i) the use of input-level geometric embeddings that enable the network to learn a scale prior over objects; and (ii) decoupling the encoder and decoder stages, via a variational latent representation that is conditioned on single frame information. We evaluated ZeroDepth targeting both outdoor (KITTI, DDAD, nuScenes) and indoor (NYUv2) benchmarks, and achieved a new state-of-the-art in both settings using the same pre-trained model, outperforming methods that train on in-domain data and require test-time scaling to produce metric estimates. Project page: https://sites.google.com/view/tri-zerodepth. Vitor Campagnolo Guizilini, Igor Vasiljevic, Dian Chen 0005, Rares Ambrus, Adrien Gaidon |
ICCV | 1 |
| 2023 | DeLiRa: Self-Supervised Depth, Light, and Radiance FieldsabstractDifferentiable volumetric rendering is a powerful paradigm for 3D reconstruction and novel view synthesis. However, standard volume rendering approaches struggle with degenerate geometries in the case of limited viewpoint diversity, a common scenario in robotics applications. In this work, we propose to use the multi-view photometric objective from the self-supervised depth estimation literature as a geometric regularizer for volumetric rendering, significantly improving novel view synthesis without requiring additional information. Building upon this insight, we explore the explicit modeling of scene geometry using a generalist Transformer, jointly learning a radiance field as well as depth and light fields with a set of shared latent codes. We demonstrate that sharing geometric information across tasks is mutually beneficial, leading to improvements over single-task learning without an increase in network complexity. Our DeLiRa architecture achieves state-of-the-art results on the ScanNet benchmark, enabling high quality volumetric rendering as well as real-time novel view and depth synthesis in the limited viewpoint diversity setting. Our project page is https://sites.google.com/view/tri-delira. Vitor Campagnolo Guizilini, Igor Vasiljevic, Jiading Fang, Rares Ambrus, Sergey Zakharov, Vincent Sitzmann, Adrien Gaidon |
ICCV | 1 |
| 2023 | NeO 360: Neural Fields for Sparse View Synthesis of Outdoor ScenesabstractRecent implicit neural representations have shown great results for novel view synthesis. However, existing methods require expensive per-scene optimization from many views hence limiting their application to real-world unbounded urban settings where the objects of interest or backgrounds are observed from very few views. To mitigate this challenge, we introduce a new approach called NeO 360, Neural fields for sparse view synthesis of outdoor scenes. NeO 360 is a generalizable method that reconstructs 360° scenes from a single or a few posed RGB images. The essence of our approach is in capturing the distribution of complex real-world outdoor 3D scenes and using a hybrid image-conditional triplanar representation that can be queried from any world point. Our representation combines the best of both voxel-based and bird’s-eye-view (BEV) representations and is more effective and expressive than each. NeO 360’s representation allows us to learn from a large collection of unbounded 3D scenes while offering generalizability to new views and novel scenes from as few as a single image during inference. We demonstrate our approach on the pro posed challenging 360° unbounded dataset, called NeRDS 360, and show that NeO 360 outperforms state-of-the-art generalizable methods for novel view synthesis while also offering editing and composition capabilities. Project page: zubair-irshad.github.io/projects/neo360.html Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Vitor Campagnolo Guizilini, Thomas Kollar, Adrien Gaidon, Zsolt Kira, Rares Ambrus |
ICCV | 4 |
| 2023 | Depth Is All You Need for Monocular 3D DetectionabstractA key contributor to recent progress in 3D detection from single images is monocular depth estimation. Existing methods focus on how to leverage depth explicitly, by generating pseudo-pointclouds or providing attention cues for image features. More recent works leverage depth prediction as a pretraining task and fine-tune the depth representation while training it for 3D detection. However, the adaptation is limited in scale by manual labels. In this work, we propose further aligning the depth representation with the target domain in an unsupervised fashion. Our methods leverage commonly available LiDAR or RGB videos during training time to fine-tune the depth representation, which leads to improved 3D detectors. Especially when using RGB videos, we show that our two-stage training by first generating depth pseudo-labels is critical, because of the inconsistency in loss distribution between the two tasks. With either type of reference data, our multi-task learning approach improves over the state of the art on both KITTI and NuScenes, while matching the test-time complexity of its single-task sub-network. Source code and pretrained models are available on https://github.com/TRI-ML/DD3D. Dennis Park, Jie Li 0031, Dian Chen 0005, Vitor Campagnolo Guizilini, Adrien Gaidon |
ICRA | 4 |
| 2023 | Robust Self-Supervised Extrinsic Self-CalibrationabstractAutonomous vehicles and robots need to operate over a wide variety of scenarios in order to complete tasks efficiently and safely. Multi-camera self-supervised monocular depth estimation from videos is a promising way to reason about the environment, as it generates metrically scaled geometric predictions from visual data without requiring additional sensors. However, most works assume well-calibrated extrinsics to fully leverage this multi-camera setup, even though accurate and efficient calibration is still a challenging problem. In this work, we introduce a novel method for extrinsic calibration that builds upon the principles of self-supervised monocular depth and ego-motion learning. Our proposed curriculum learning strategy uses monocular depth and pose estimators with velocity supervision to estimate extrinsics, and then jointly learns extrinsic calibration along with depth and pose for a set of overlapping cameras rigidly attached to a moving vehicle. Experiments on a benchmark multi-camera dataset (DDAD) demonstrate that our method enables self-calibration in various scenes robustly and efficiently compared to a traditional vision-based pose estimation pipeline. Furthermore, we demonstrate the benefits of extrinsics self-calibration as a way to improve depth prediction via joint optimization. The project page: https://sites.google.com/tri.global/tri-sesc Takayuki Kanai, Igor Vasiljevic, Vitor Campagnolo Guizilini, Adrien Gaidon, Rares Ambrus |
IROS | 3 |
| 2022 | Multi-Frame Self-Supervised Depth with TransformersabstractMulti-frame depth estimation improves over single-frame approaches by also leveraging geometric relationships between images via feature matching, in addition to learning appearance-based features. In this paper we revisit feature matching for self-supervised monocular depth estimation, and propose a novel transformer architecture for cost volume generation. We use depth-discretized epipolar sampling to select matching candidates, and refine predictions through a series of self- and cross-attention layers. These layers sharpen the matching probability between pixel features, improving over standard similarity metrics prone to ambiguities and local minima. The refined cost volume is decoded into depth estimates, and the whole pipeline is trained end-to-end from videos using only a photometric objective. Experiments on the KITTI and DDAD datasets show that our DepthFormer architecture establishes a new state of the art in self-supervised monocular depth estimation, and is even competitive with highly specialized supervised single-frame architectures. We also show that our learned cross-attention network yields representations transferable across datasets, increasing the effectiveness of pre-training strategies. Project page: https://sites.google.com/tri.global/depthformer. Vitor Campagnolo Guizilini, Rares Ambrus, Dian Chen 0005, Sergey Zakharov, Adrien Gaidon |
CVPR | 1 |
| 2022 | Depth Field Networks For Generalizable Multi-view Scene Representation
Vitor Campagnolo Guizilini, Igor Vasiljevic, Jiading Fang, Rare Ambru, Gregory Shakhnarovich, Matthew R. Walter, Adrien Gaidon |
ECCV (32) | 1 |
| 2022 | SpOT: Spatiotemporal Modeling for 3D Object Tracking
Colton Stearns, Davis Rempe, Jie Li 0031, Rares Ambrus, Sergey Zakharov, Vitor Campagnolo Guizilini, Yanchao Yang 0001, Leonidas J. Guibas |
ECCV (38) | 6 |
| 2022 | Photo-realistic Neural Domain Randomization
Sergey Zakharov, Rares Ambrus, Vitor Campagnolo Guizilini, Wadim Kehl, Adrien Gaidon |
ECCV (25) | 3 |
| 2022 | Self-Supervised Camera Self-Calibration from VideoabstractCamera calibration is integral to robotics and computer vision algorithms that seek to infer geometric properties of the scene from visual input streams. In practice, calibration is a laborious procedure requiring specialized data collection and careful tuning. This process must be repeated whenever the parameters of the camera change, which can be a frequent occurrence for mobile robots and autonomous vehicles. In contrast, self-supervised depth and ego-motion estimation approaches can bypass explicit calibration by in-ferring per-frame projection models that optimize a view-synthesis objective. In this paper, we extend this approach to explicitly calibrate a wide range of cameras from raw videos in the wild. We propose a learning algorithm to regress per-sequence calibration parameters using an efficient family of general camera models. Our procedure achieves self-calibration results with sub-pixel reprojection error, outperforming other learning-based methods. We validate our approach on a wide variety of camera geometries, including perspective, fisheye, and catadioptric. Finally, we show that our approach leads to improvements in the downstream task of depth estimation, achieving state-of-the-art results on the EuRoC dataset with greater computational efficiency than contemporary methods. The project page: https://sites.google.com/ttic.edu/self-sup-self-calib Jiading Fang, Igor Vasiljevic, Vitor Campagnolo Guizilini, Rares Ambrus, Gregory Shakhnarovich, Adrien Gaidon, Matthew R. Walter |
ICRA | 3 |
| 2021 | Sparse Auxiliary Networks for Unified Monocular Depth Prediction and CompletionabstractEstimating scene geometry from data obtained with cost-effective sensors is key for robots and self-driving cars. In this paper, we study the problem of predicting dense depth from a single RGB image (monodepth) with optional sparse measurements from low-cost active depth sensors. We introduce Sparse Auxiliary Networks (SANs), a new module enabling monodepth networks to perform both the tasks of depth prediction and completion, depending on whether only RGB images or also sparse point clouds are available at inference time. First, we decouple the image and depth map encoding stages using sparse convolutions to process only the valid depth map pixels. Second, we inject this information, when available, into the skip connections of the depth prediction network, augmenting its features. Through extensive experimental analysis on one indoor (NYUv2) and two outdoor (KITTI and DDAD) benchmarks, we demonstrate that our proposed SAN architecture is able to simultaneously learn both tasks, while achieving a new state of the art in depth prediction by a significant margin. Vitor Campagnolo Guizilini, Rares Ambrus, Wolfram Burgard, Adrien Gaidon |
CVPR | 1 |
| 2021 | Geometric Unsupervised Domain Adaptation for Semantic SegmentationabstractSimulators can efficiently generate large amounts of labeled synthetic data with perfect supervision for hard-to-label tasks like semantic segmentation. However, they introduce a domain gap that severely hurts real-world performance. We propose to use self-supervised monocular depth estimation as a proxy task to bridge this gap and improve sim-to-real unsupervised domain adaptation (UDA). Our Geometric Unsupervised Domain Adaptation method (GUDA)1learns a domain-invariant representation via a multi-task objective combining synthetic semantic supervision with real-world geometric constraints on videos. GUDA establishes a new state of the art in UDA for semantic segmentation on three benchmarks, outperforming methods that use domain adversarial learning, self-training, or other self-supervised proxy tasks. Furthermore, we show that our method scales well with the quality and quantity of synthetic data while also improving depth prediction. Vitor Campagnolo Guizilini, Jie Li 0031, Rares Ambrus, Adrien Gaidon |
ICCV | 1 |
| 2021 | Is Pseudo-Lidar needed for Monocular 3D Object detection?abstractRecent progress in 3D object detection from single images leverages monocular depth estimation as a way to produce 3D pointclouds, turning cameras into pseudo-lidar sensors. These two-stage detectors improve with the accuracy of the intermediate depth estimation network, which can itself be improved without manual labels via large-scale self-supervised learning. However, they tend to suffer from overfitting more than end-to-end methods, are more complex, and the gap with similar lidar-based detectors remains significant. In this work, we propose an end-to-end, single stage, monocular 3D object detector, DD3D, that can benefit from depth pre-training like pseudo-lidar methods, but without their limitations. Our architecture is designed for effective information transfer between depth estimation and 3D detection, allowing us to scale with the amount of unlabeled pre-training data. Our method achieves state-of-the-art results on two challenging benchmarks, with 16.34% and 9.28% AP for Cars and Pedestrians (respectively) on the KITTI-3D benchmark, and 41.5% mAP on NuScenes. Dennis Park, Rares Ambrus, Vitor Campagnolo Guizilini, Jie Li 0031, Adrien Gaidon |
ICCV | 3 |
| 2021 | MarioNette: Self-Supervised Sprite LearningabstractArtists and video game designers often construct 2D animations using libraries of sprites---textured patches of objects and characters. We propose a deep learning approach that decomposes sprite-based video animations into a disentangled representation of recurring graphic elements in a self-supervised manner. By jointly learning a dictionary of possibly transparent patches and training a network that places them onto a canvas, we deconstruct sprite-based content into a sparse, consistent, and explicit representation that can be easily used in downstream tasks, like editing or analysis. Our framework offers a promising approach for discovering recurring visual patterns in image collections without supervision. Dmitriy Smirnov 0001, Michaël Gharbi, Matthew Fisher, Vitor Campagnolo Guizilini, Alexei A. Efros, Justin Solomon 0001 |
NeurIPS | 4 |
| 2020 | Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motionabstractSelf-supervised learning has emerged as a powerful tool for depth and ego-motion estimation, leading to state-of-the-art results on benchmark datasets. However, one significant limitation shared by current methods is the assumption of a known parametric camera model - usually the standard pinhole geometry - leading to failure when applied to imaging systems that deviate significantly from this assumption (e.g., catadioptric cameras or underwater imaging). In this work, we show that self-supervision can be used to learn accurate depth and ego-motion estimation without prior knowledge of the camera model. Inspired by the geometric model of Grossberg and Nayar, we introduce Neural Ray Surfaces (NRS), convolutional networks that represent pixel-wise projection rays, approximating a wide range of cameras. NRS are fully differentiable and can be learned end-to-end from unlabeled raw videos. We demonstrate the use of NRS for self-supervised learning of visual odometry and depth estimation from raw videos obtained using a wide variety of camera systems, including pinhole, fisheye, and catadioptric. Igor Vasiljevic, Vitor Campagnolo Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Gregory Shakhnarovich, Adrien Gaidon |
3DV | 2 |
| 2020 | 3D Packing for Self-Supervised Monocular Depth EstimationabstractAlthough cameras are ubiquitous, robotic platforms typically rely on active sensors like LiDAR for direct 3D perception. In this work, we propose a novel self-supervised monocular depth estimation method combining geometry with a new deep network, PackNet, learned only from unlabeled monocular videos. Our architecture leverages novel symmetrical packing and unpacking blocks to jointly learn to compress and decompress detail-preserving representations using 3D convolutions. Although self-supervised, our method outperforms other self, semi, and fully supervised methods on the KITTI benchmark. The 3D inductive bias in PackNet enables it to scale with input resolution and number of parameters without overfitting, generalizing better on out-of-domain data such as the NuScenes dataset. Furthermore, it does not require large-scale supervised pretraining on ImageNet and can run in real-time. Finally, we release DDAD (Dense Depth for Automated Driving), a new urban driving dataset with more challenging and accurate depth evaluation, thanks to longer-range and denser ground-truth depth generated from high-density LiDARs mounted on a fleet of self-driving cars operating world-wide. Vitor Campagnolo Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raventos, Adrien Gaidon |
CVPR | 1 |
| 2020 | Real-Time Panoptic Segmentation From Dense DetectionsabstractPanoptic segmentation is a complex full scene parsing task requiring simultaneous instance and semantic segmentation at high resolution. Current state-of-the-art approaches cannot run in real-time, and simplifying these architectures to improve efficiency severely degrades their accuracy. In this paper, we propose a new single-shot panoptic segmentation network that leverages dense detections and a global self-attention mechanism to operate in real-time with performance approaching the state of the art. We introduce a novel parameter-free mask construction method that substantially reduces computational complexity by efficiently reusing information from the object detection and semantic segmentation sub-tasks. The resulting network has a simple data flow that requires no feature map re-sampling, enabling significant hardware acceleration. Our experiments on the Cityscapes and COCO benchmarks show that our network works at 30 FPS on 1024x2048 resolution, trading a 3% relative performance degradation from the current state of the art for up to 440% faster inference. Rui Hou 0007, Jie Li 0031, Arjun Bhargava, Allan Raventos, Vitor Campagnolo Guizilini, Jerome P. Lynch, Adrien Gaidon |
CVPR | 5 |
| 2020 | Semantically-Guided Representation Learning for Self-Supervised Monocular Depth
Vitor Campagnolo Guizilini, Rui Hou 0007, Jie Li 0031, Rares Ambrus, Adrien Gaidon |
ICLR | 1 |
| 2020 | Neural Outlier Rejection for Self-Supervised Keypoint Learning
Jiexiong Tang, Hanme Kim, Vitor Campagnolo Guizilini, Sudeep Pillai, Rares Ambrus |
ICLR | 3 |
| 2019 | Dynamic Hilbert Maps: Real-Time Occupancy Predictions in Changing EnvironmentsabstractThis paper addresses the problem of learning instantaneous occupancy levels of dynamic environments and predicting future occupancy levels. Due to the complexity of most real environments, such as urban streets or crowded areas, the efficient and robust incorporation of temporal dependencies into otherwise static occupancy models remains a challenge. We propose a method to capture the uncertainty of moving objects and incorporate this uncertainty information into a continuous occupancy map represented in a rich high-dimensional feature space. This data-efficient model not only allows us to learn the occupancy states incrementally, but also makes predictions about what the future occupancy states will be. Experiments performed using 2D and 3D laser data collected from crowded unstructured outdoor environments show that the proposed methodology can accurately predict occupancy states for areas of around 1000 m2at 10 Hz, making the proposed framework ideal for online applications under real-time constraints. Vitor Campagnolo Guizilini, Ransalu Senanayake, Fabio Ramos 0001 |
ICRA | 1 |
| 2019 | Segmenting and Detecting Nematode in Coffee Crops Using Aerial Images
Alexandre J. Oliveira, Gleice A. de Assis, Vitor Campagnolo Guizilini, Elaine Ribeiro de Faria, Jefferson R. Souza |
ICVS | 3 |
| 2018 | Iterative Continuous Convolution for 3D Template Matching and Global LocalizationabstractThis paper introduces a novel methodology for 3D template matching that is scalable to higher-dimensional spaces and larger kernel sizes. It uses the Hilbert Maps framework to model raw pointcloud information as a continuous occupancy function, and we derive a closed-form solution to the convolution operation that takes place directly in the Reproducing Kernel Hilbert Space defining these functions. The result is a third function modeling activation values, that can be queried at arbitrary resolutions with logarithmic complexity, and by iteratively searching for high similarity areas we can determine matching candidates. Experimental results show substantial speed gains over standard discrete convolution techniques, such as sliding window and fast Fourier transform, along with a significant decrease in memory requirements, without accuracy loss. This efficiency allows the proposed methodology to be used in areas where discrete convolution is currently infeasible. As a practical example we explore the key problem in robotics of global localization, in which a vehicle must be positioned on a map using only its current sensor information, and provide comparisons with other state-of-the-art techniques in terms of computational speed and accuracy. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
AAAI | 1 |
| 2018 | Learning to Race Through Coordinate Descent Bayesian OptimisationabstractIn the automation of many kinds of processes, the observable outcome can often be described as the combined effect of an entire sequence of actions, or controls, applied throughout the process execution. In these cases, strategies to optimise control policies for individual stages of the process are not applicable, and instead the whole policy needs to be optimised at once. On the other hand, the cost to evaluate the policy's performance might also be high, being desirable that a solution can be found with as few interactions as possible with the real system. We consider the problem of optimising control policies to allow a robot to complete a given race track within a minimum amount of time. We assume that the robot has no prior information about the track or its own dynamical model, just an initial valid driving example. Localisation is only applied to monitor the robot and to provide an indication of its position along the track's centre axis. With that in mind, we propose a method for finding a policy that minimises the time per lap while keeping the vehicle on the track using a Bayesian optimisation (BO) approach over a reproducing kernel Hilbert space. We apply an algorithm to search more efficiently over high-dimensional policy-parameter spaces with BO, by iterating over each dimension individually, in a sequential coordinate descent-like scheme. Experiments demonstrate the performance of the algorithm against other methods in a simulated car racing environment. Rafael Oliveira 0001, Fernando H. M. Rocha, Lionel Ott, Vitor Campagnolo Guizilini, Fabio Ramos 0001, Valdir Grassi Jr. |
ICRA | 4 |
| 2018 | Failure Detection in Row Crops From UAV Images Using Morphological OperatorsabstractThe detection of failures (DF) in coffee crops is fundamental in evaluating product quality and the optimal occupation of planted areas. The use of unmanned aerial vehicles (UAVs) in precision agriculture has great potential as a tool to analyze critical parameters in cultivation, among them the detection of planting failures. This letter presents a novel methodology for DF from aerial images, obtained using a UAV capable of collecting high-resolution RGB images. The proposed approach uses mathematical morphology operators to detect failures over planted areas and returns both the individual positions of these failures and total failure length (sum of empty spaces between plants), thus facilitating decision making for further actions. Results show that the proposed DF method is reliable for accurately identifying failures over rows of planted coffee crops. Henrique Candido de Oliveira, Vitor Campagnolo Guizilini, Israel P. Nunes, Jefferson R. Souza |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Unsupervised Feature Learning for 3D Scene Reconstruction with Occupancy MapsabstractThis paper addresses the task of unsupervised feature learning for three-dimensional occupancy mapping, as a way to segment higher-level structures based on raw unorganized point cloud data. In particular, we focus on detecting planar surfaces, which are common in most structured or semi-structured environments. This segmentation is then used to minimize the amount of parameters necessary to properly create a 3D occupancy model of the surveyed space, thus increasing computational speed and decreasing memory requirements. As the 3D modeling tool, an extension to Hilbert Maps was selected, since it naturally uses a feature-based representation of the environment to achieve real-time performance. Experiments conducted in simulated and real large-scale datasets show a substantial gain in performance, while decreasing the amount of stored information by orders of magnitude without sacrificing accuracy. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
AAAI | 1 |
| 2017 | Markovian jump linear systems-based filtering for visual and GPS aided inertial navigation systemabstractVisual-Inertial SLAM methods have become a very important technology for several applications in robotics. This kind of approach usually is composed by sensors as rate gyros, accelerometers and monocular cameras. Magnetometers and GPS modules generally used for outdoors are absent in the SLAM system observation, since the magnetometer measurements deteriorate in the presence of ferromagnetic materials and the GPS module signals are unavailable indoors or in urban environments. In order to make use of all these sensors, we propose Markovian jump linear systems (MJLS) to model the modes of operation of the navigation system based on available sensors and their reliability. An extended Kalman filter for MJLS fuses the sensor data and estimates the motion using the best mode of operation for each particular time instant. Experimental results are presented to show the effectiveness of the proposed method, in situations that would pose a challenge for standard data fusion techniques. Roberto S. Inoue, Vitor Campagnolo Guizilini, Marco H. Terra, Fabio Ramos 0001 |
IROS | 2 |
| 2017 | Variational Hilbert Regression with Applications to Terrain Modeling
Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
ISRR | 1 |
| 2017 | Bayesian Optimisation for Safe Navigation Under Localisation Uncertainty
Rafael Oliveira 0001, Lionel Ott, Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
ISRR | 3 |
| 2016 | Route planning for active classification with UAVsabstractThe mapping of agricultural crops by capturing images obtained with UAVs enables fast environmental monitoring and diagnosis in large areas. Airborne monitoring in agriculture can a substantially impacts on the identification of diseases and produce accurate information on affected areas. The problem can be formulated as a classification task on aerial images with significant opportunities to impact other fields. This paper presents an active learning method through route planning for improvements in the knowledge on visited areas and minimization uncertainties about the classification of diseases in crops. Binary Logistic Regression and Gaussian Process were used for the detection of pathologies and map interpolation, respectively. A Bayesian optimization strategy is also proposed for the planning of an informative trajectory, which resulted in a maximized search for affected areas in an initially unknown environment. Kelen Cristiane Teixeira Vivaldini, Vitor Campagnolo Guizilini, Matheus D. Croce, Thiago H. Martinelli, Denis F. Wolf, Fabio Ramos 0001 |
ICRA | 2 |
| 2016 | Large-scale 3D scene reconstruction with Hilbert Mapsabstract3D scene reconstruction involves the volumetric modeling of space, and it is a fundamental step in a wide variety of robotic applications, including grasping, obstacle avoidance, path planning, mapping and many others. Nowadays, sensors are able to quickly collect vast amounts of data, and the challenge has become one of storing and processing all this information in a timely manner, especially if real-time performance is required. Recently, a novel technique for the stochastic learning of discriminative models through continuous occupancy maps was proposed: Hilbert Maps [18], that is able to represent the input space at an arbitrary resolution while capturing statistical relationships between measurements. The original framework was proposed for 2D environments, and here we extend it to higher-dimensional spaces, addressing some of the challenges brought by the curse of dimensionality. Namely, we propose a method for the automatic selection of feature coordinate locations, and introduce the concept of localized automatic relevance determination (LARD) to the Hilbert Maps framework, in which different dimensions in the projected Hilbert space operate within independent length-scale values. The proposed technique was tested against other state-of-the-art 3D scene reconstruction tools in three different datasets: a simulated indoors environment, RIEGL laser scans and dense LSD-SLAM pointclouds. The results testify to the proposed framework's ability to model complex structures and correctly interpolate over unobserved areas of the input space while achieving real-time training and querying performances. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
IROS | 1 |
| 2015 | A Nonparametric Online Model for Air Quality PredictionabstractWe introduce a novel method for the continuous online prediction of particulate matter in the air (more specifically, PM10 and PM2.5) given sparse sensor information. A nonparametric model is developed using Gaussian Processes, which eschews the need for an explicit formulation of internal -- and usually very complex -- dependencies between meteorological variables. Instead, it uses historical data to extrapolate pollutant values both spatially (in areas with no sensor information) and temporally (the near future). Each prediction also contains a respective variance, indicating its uncertainty level and thus allowing a probabilistic treatment of results. A novel training methodology (Structural Cross-Validation) is presented, which preserves the spatio-temporal structure of available data during the hyperparameter optimization process. Tests were conducted using a real-time feed from a sensor network in an area of roughly 50x80 km, alongside comparisons with other techniques for air pollution prediction. The promising results motivated the development of a smartphone applicative and a website, currently in use to increase the efficiency of air quality monitoring and control in the area. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
AAAI | 1 |
| 2015 | Automatic detection of Ceratocystis wilt in Eucalyptus crops from aerial imagesabstractOne of the challenges in precision agriculture is the detection of diseased crops in agricultural environments. This paper presents a methodology to detect the Ceratocystis wilt disease in Eucalyptus crops. An unmanned aerial vehicle is used to obtain high-resolution RGB images of a predefined area. The methodology enables the extraction of visual features from image regions and uses several supervised machine learning (ML) techniques to classify regions into three classes: ground, healthy and diseased plants. Several learning techniques were compared using data obtained from a commercial Eucalyptus plantation. Experimental results show that the GP learning model is more reliable than the other learning methods for accurately identifying diseased trees. Jefferson R. Souza, Caio C. T. Mendes, Vitor Campagnolo Guizilini, Kelen Cristiane Teixeira Vivaldini, Adimara Colturato, Fabio Ramos 0001, Denis F. Wolf |
ICRA | 3 |
| 2014 | Dense motion segmentation for first-person activity recognitionabstractIn this paper, we propose a dense motion segmentation method for human daily activity recognition from a wearable device - "Smart Glasses". The glasses are embedded with a camera, which allows the system to automatically recognise the wearer's activities from a first-person perspective. This application can be broadly applied to patients, elderly, safety workers, e-health monitoring, or anyone requiring cognitive assistance or guidance on their activities of daily living (ADLs). We validate our system in challenging real-world scenarios, and compare two feature extraction approaches: averaged optical flow and a combined dense motion segmentation approach. We classify them using LogitBoost (on Decision Stumps) and Support Vector Machine (SVM). We also suggest the optimal settings of the classifiers through cross-validation over our ADLs database. The results show that the optical flow with average pooling has a good performance when classifying general locomotive activities. The results also indicate the benefits that dense motion segmentation features can have on reliably classify activities involving a moving object, such as hands. We achieve an overall accuracy of up to 69.76% on 12 ADLs using local classifiers, and with a Hidden Markov Model (HMM) process this accuracy improves to up to 89.59%. Kai Zhan, Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
ICARCV | 2 |
| 2014 | Online self-supervised multi-instance segmentation of dynamic objectsabstractThis paper presents a method for the continuous segmentation of dynamic objects using only a vehicle mounted monocular camera without any prior knowledge of the object's appearance. Prior work in online static/dynamic segmentation [1] is extended to identify multiple instances of dynamic objects by introducing an unsupervised motion clustering step. These clusters are then used to update a multi-class classifier within a self-supervised framework. In contrast to many tracking-by-detection based methods, our system is able to detect dynamic objects without any prior knowledge of their visual appearance shape or location. Furthermore, the classifier is used to propagate labels of the same object in previous frames, which facilitates the continuous tracking of individual objects based on motion. The proposed system is evaluated using recall and false alarm metrics in addition to a new multi-instance labelled dataset to measure the performance of segmenting multiple instances of objects. Alex Bewley, Vitor Campagnolo Guizilini, Fabio Ramos 0001, Ben Upcroft |
ICRA | 2 |
| 2013 | Online self-supervised segmentation of dynamic objectsabstractWe address the problem of automatically segmenting dynamic objects in an urban environment from a moving camera without manual labelling, in an online, self-supervised learning manner. We use input images obtained from a single uncalibrated camera placed on top of a moving vehicle, extracting and matching pairs of sparse features that represent the optical flow information between frames. This optical flow information is initially divided into two classes, static or dynamic, where the static class represents features that comply to the constraints provided by the camera motion and the dynamic class represents the ones that do not. This initial classification is used to incrementally train a Gaussian Process (GP) classifier to segment dynamic objects in new images. The hyperparameters of the GP covariance function are optimized online during navigation, and the available self-supervised dataset is updated as new relevant data is added and redundant data is removed, resulting in a near-constant computing time even after long periods of navigation. The output is a vector containing the probability that each pixel in the image belongs to either the static or dynamic class (ranging from 0 to 1), along with the corresponding uncertainty estimate of the classification. Experiments conducted in an urban environment, with cars and pedestrians as dynamic objects and no prior knowledge or additional sensors, show promising results even when the vehicle is moving at considerable speeds (up to 50 km/h). This scenario produces a large quantity of featureless regions and false matches that is very challenging for conventional approaches. Results obtained using a portable camera device also testify to our algorithm's ability to generalize over different environments and configurations without any fine-tuning of parameters. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
ICRA | 1 |
| 2012 | Semi-parametric models for visual odometryabstractThis paper introduces a novel framework for estimating the motion of a robotic car from image information, a scenario widely known as visual odometry. Most current monocular visual odometry algorithms rely on a calibrated camera model and recover relative rotation and translation by tracking image features and applying geometrical constraints. This approach has some drawbacks: translation is recovered up to a scale, it requires camera calibration which can be tricky under certain conditions, and uncertainty estimates are not directly obtained. We propose an alternative approach that involves the use of semi-parametric statistical models as means to recover scale, infer camera parameters and provide uncertainty estimates given a training dataset. As opposed to conventional non-parametric machine learning procedures, where standard models for egomotion would be neglected, we present a novel framework in which the existing parametric models and powerful non-parametric Bayesian learning procedures are combined. We devise a multiple output Gaussian Process (GP) procedure, named Coupled GP, that uses a parametric model as the mean function and a non-stationary covariance function to map image features directly into vehicle motion. Additionally, this procedure is also able to infer joint uncertainty estimates (full covariance matrices) for rotation and translation. Experiments performed using data collected from a single camera under challenging conditions show that this technique outperforms traditional methods in trajectories of several kilometers. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
ICRA | 1 |
| 2011 | Visual odometry learning for unmanned aerial vehiclesabstractThis paper addresses the problem of using visual information to estimate vehicle motion (a.k.a. visual odometry) from a machine learning perspective. The vast majority of current visual odometry algorithms are heavily based on geometry, using a calibrated camera model to recover relative translation (up to scale) and rotation by tracking image features over time. Our method eliminates the need for a parametric model by jointly learning how image structure and vehicle dynamics affect camera motion. This is achieved with a Gaussian Process extension, called Coupled GP, which is trained in a supervised manner to infer the underlying function mapping optical flow to relative translation and rotation. Matched image features parameters are used as inputs and linear and angular velocities are the outputs in our non-linear multi-task regression problem. We show here that it is possible, using a single uncalibrated camera and establishing a first-order temporal dependency between frames, to jointly estimate not only a full 6 DoF motion (along with a full covariance matrix) but also relative scale, a non-trivial problem in monocular configurations. Experiments were performed with imagery collected with an unmanned aerial vehicle (UAV) flying over a deserted area at speeds of 100-120 km/h and altitudes of 80-100 m, a scenario that constitutes a challenge for traditional visual odometry estimators. Vitor Campagnolo Guizilini, Fabio Ramos 0001 |
ICRA | 1 |
| 2008 | Solving the Online SLAM Problem with an Omnidirectional Vision System
Vitor Campagnolo Guizilini, Jun Okamoto |
ICONIP (1) | 1 |