Steven Lake Waslander

dblp:18/7142 · also Steven L. Waslander · DBLP profile ↗
← Back
71ranked-venue papers
2as first author
31since 2021 · last 2026
0000-0003-4217-4415ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 1 first-author · 26 since 2021Systems, architecture and hardware · 32 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 SCATR: Mitigating New Instance Suppression in LiDAR-based Tracking-by-Attention via Second Chance Assignment and Track Query Dropout
abstract
LiDAR-based tracking-by-attention (TBA) frameworks inherently suffer from high false negative errors, leading to a significant performance gap compared to traditional LiDAR-based tracking-by-detection (TBD) methods. This paper introduces SCATR, a novel LiDAR-based TBA model designed to address this fundamental challenge systematically. SCATR leverages recent progress in vision-based tracking and incorporates targeted training strategies specifically adapted for LiDAR. Our work’s core innovations are two architecture-agnostic training strategies for TBA methods: Second Chance Assignment and Track Query Dropout. Second Chance Assignment is a novel ground truth assignment that concatenates unassigned track queries to the proposal queries before bipartite matching, giving these track queries a second chance to be assigned to a ground truth object and effectively mitigating the conflict between detection and tracking tasks inherent in tracking-by-attention. Track Query Dropout is a training method that diversifies supervised object query configurations to efficiently train the decoder to handle different track query sets, enhancing robustness to missing or newborn tracks. Experiments on the nuScenes tracking benchmark demonstrate that SCATR achieves state-of-the-art performance among LiDAR-based TBA methods, outperforming previous works by 7.6% AMOTA and successfully bridging the long-standing performance gap between LiDAR-based TBA and TBD methods. Ablation studies further validate the effectiveness and generalization of Second Chance Assignment and Track Query Dropout. Code can be found at the following link: https://github.com/TRAILab/SCATR
Brian Cheong, Sandro Papais, Steven Lake Waslander
WACV4
2026 FlowCLAS: Enhancing Normalizing Flow-Based Anomaly Segmentation Via Contrastive Learning
abstract
Anomaly segmentation is an essential capability for safety-critical robotics applications that must be aware of unexpected events. Normalizing flows (NFs), a class of generative models, are a promising approach for this task due to their ability to model the inlier data distribution efficiently. However, their performance falters in dynamic scenes, where complex, multi-modal data distributions cause them to struggle with identifying out-of-distribution samples, leaving a performance gap to leading discriminative methods. To address this limitation, we introduce FlowCLAS, a hybrid framework that enhances the traditional maximum likelihood objective of NFs with a discriminative, contrastive loss. Leveraging Outlier Exposure, this objective explicitly enforces a separation between normal and anomalous features in the latent space, retaining the probabilistic foundation of NFs while embedding the discriminative power they lack. The strength of this approach is demonstrated by FlowCLAS establishing new state-of-the-art (SOTA) performance across multiple challenging anomaly segmentation benchmarks for robotics, including Fishyscapes Lost & Found, Road Anomaly, SegmentMeIfYouCan-ObstacleTrack, and ALLO. Our experiments also show that this contrastive approach is more effective than other outlier-based training strategies for NFs, successfully bridging the performance gap to leading discriminative methods. Project page: https://trailab.github.io/FlowCLAS
Selina Leveugle, Paul Grouchy, Chris Langley, Svetlana Stolpner, Jonathan Kelly, Steven Lake Waslander
WACV7
2025 Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection
abstract
The current state-of-the-art methods in domain adaptive object detection (DAOD) use Mean Teacher self-labelling, where a teacher model, directly derived as an exponential moving average of the student model, is used to generate labels on the target domain which are then used to improve both models in a positive loop. This couples learning and generating labels on the target domain, and other recent works also leverage the generated labels to add additional domain alignment losses. We believe this coupling is brittle and excessively constrained: there is no guarantee that a student trained only on source data can generate accurate target domain labels and initiate the positive feedback loop, and much better target domain labels can likely be generated by using a large pretrained network that has been exposed to much more data. Vision foundational models are exactly such models, and they have shown impressive task generalization capabilities even when frozen. We want to leverage these models for DAOD and introduce DINO Teacher, which consists of two components. First, we train a new labeller on source data only using a large frozen DINOv2 backbone and show it generates more accurate labels than Mean Teacher. Next, we align the student’s source and target image patch features with those from a DINO encoder, driving source and target representations closer to the generalizable DINO representation. We obtain state-of-the-art performance on multiple DAOD datasets. Code available at https://github.com/TRAILab/DINO_Teacher.
Marc-Antoine Lavoie, Anas Mahmoud 0002, Steven Lake Waslander
CVPR3
2025 PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point Clouds
abstract
Self-supervised learning (SSL) on 3D point clouds has the potential to learn feature representations that can transfer to diverse sensors and multiple downstream perception tasks. However, recent SSL approaches fail to define pretext tasks that retain geometric information such as object pose and scale, which can be detrimental to the performance of downstream localization and geometry-sensitive 3D scene understanding tasks, such as 3D semantic segmentation and 3D object detection. We propose PSA-SSL, a novel extension to point cloud SSL that learns object pose and size-aware (PSA) features. Our approach defines a self-supervised bounding box regression pretext task, which retains object pose and size information. Furthermore, we incorporate LiDAR beam pattern augmentation on input point clouds, which encourages learning sensor-agnostic features. Our experiments demonstrate that with a single pretrained model, our light-weight yet effective extensions achieve significant improvements on 3D semantic segmentation with limited labels across popular autonomous driving datasets (Waymo, nuScenes, SemanticKITTI). Moreover, our approach outperforms other state-of-the-art SSL methods on 3D semantic segmentation (using up to 10 times less labels), as well as on 3D object detection. Our code will be released on https://github.com/TRAILab/PSA-SSL.
Barza Nisar, Steven Lake Waslander
CVPR2
2025 ForeSight: Multi-View Streaming Joint Object Detection and Trajectory Forecasting
abstract
We introduce ForeSight, a novel joint detection and forecasting framework for vision-based 3D perception in autonomous vehicles. Traditional approaches treat detection and forecasting as separate sequential tasks, limiting their ability to leverage temporal cues. ForeSight addresses this limitation with a multi-task streaming and bidirectional learning approach, allowing detection and forecasting to share query memory and propagate information seamlessly. The forecast-aware detection transformer enhances spatial reasoning by integrating trajectory predictions from a multiple hypothesis forecast memory queue, while the streaming forecast transformer improves temporal consistency using past forecasts and refined detections. Unlike tracking-based methods, ForeSight eliminates the need for explicit object association, reducing error propagation with a tracking-free model that efficiently scales across multi-frame sequences. Experiments on the nuScenes dataset show that ForeSight achieves state-of-the-art performance, achieving an EPA of 54.9%, surpassing previous methods by 9.3%, while also attaining the best mAP and minADE among multi-view detection and forecasting models.
Sandro Papais, Brian Cheong, Steven Lake Waslander
ICCV4
2025 SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion Prediction
abstract
Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dynamic, human-robot-mixed environments. However, the scarcity of large-scale driving datasets has hindered the development of robust and generalizable motion prediction models, limiting their ability to capture complex interactions and road geometries. Inspired by recent advances in natural language processing (NLP) and computer vision (CV), self-supervised learning (SSL) has gained significant attention in the motion prediction community for learning rich and transferable scene representations. Nonetheless, existing pre-training methods for motion prediction have largely focused on specific model architectures and single dataset, limiting their scalability and generalizability. To address these challenges, we propose SmartPretrain, a general and scalable SSL framework for motion prediction that is both model-agnostic and dataset-agnostic. Our approach integrates contrastive and reconstructive SSL, leveraging the strengths of both generative and discriminative paradigms to effectively represent spatiotemporal evolution and interactions without imposing architectural constraints. Additionally, SmartPretrain employs a dataset-agnostic scenario sampling strategy that integrates multiple datasets, enhancing data volume, diversity, and robustness. Extensive experiments on multiple datasets demonstrate that SmartPretrain consistently improves the performance of state-of-the-art prediction models across datasets, data splits and main metrics. For instance, SmartPretrain significantly reduces the MissRate of Forecast-MAE by 10.6\%. These results highlight SmartPretrain's effectiveness as a unified, scalable solution for motion prediction, breaking free from the limitations of the small-data regime.
Yang Zhou 0054, Hao Shao, Steven Lake Waslander, Hongsheng Li 0001, Yu Liu 0015
ICLR4
2025 OpenNav: Open-World Navigation with Multimodal Large Language Models
abstract
Pre-trained large language models (LLMs) have demonstrated strong common-sense reasoning abilities, making them promising for robotic navigation and planning tasks. However, despite recent progress, bridging the gap between language descriptions and actual robot actions in the open-world, beyond merely invoking limited predefined motion primitives, remains an open challenge. In this work, we aim to enable robots to interpret and decompose complex language instructions, ultimately synthesizing a sequence of trajectory points to complete diverse navigation tasks given open-set instructions and open-set objects. We observe that multi-modal large language models (MLLMs) exhibit strong cross-modal understanding when processing free-form language instructions, demonstrating robust scene comprehension. More importantly, leveraging their code-generation capability, MLLMs can interact with vision-language perception models to generate compositional 2D bird-eye-view value maps, effectively integrating semantic knowledge from MLLMs with spatial information from maps to reinforce the robot’s spatial understanding. To further validate our approach, we effectively leverage large-scale autonomous vehicle datasets (AVDs) to validate our proposed zero-shot vision-language navigation framework in outdoor navigation tasks, demonstrating its capability to execute a diverse range of free-form natural language navigation instructions while maintaining robustness against object detection errors and linguistic ambiguities. Furthermore, we validate our system on a Husky robot in both indoor and outdoor scenes, demonstrating its real-world robustness and applicability. Supplementary videos are available at https://trailab.github.io/OpenNav-website/
Mingfeng Yuan, Steven Lake Waslander
IROS3
2024 Multiple View Geometry Transformers for 3D Human Pose Estimation
abstract
In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer designs, which struggle to resolve geometric information accurately, particularly during occlusion. In-stead, we propose a novel hybrid model, MVGFormer, which has a series of geometric and appearance modules organized in an iterative manner. The geometry modules are learning-free and handle all viewpoint-dependent 3D tasks geometrically which notably improves the model's gener-alization ability. The appearance modules are learnable and are dedicated to estimating 2D poses from image signals end-to-end which enables them to achieve accurate es-timates even when occlusion occurs, leading to a model that is both accurate and generalizable to new cameras and geometries. We evaluate our approach for both in-domain and out-of-domain settings, where our model consistently outperforms state-of-the-art methods, and especially does so by a significant margin in the out-of-domain setting. We will release the code and models: https://github.com/XunshanMan/MVGFormer.
Ziwei Liao, Chunyu Wang 0001, Han Hu 0001, Steven Lake Waslander
CVPR5
2024 LMDrive: Closed-Loop End-to-End Driving with Large Language Models
abstract
Despite significant recent progress in the field of autonomous driving, modern methods still struggle and can incur serious accidents when encountering long-tail unfore-seen events and challenging urban scenarios. On the one hand, large language models (LLM) have shown impres-sive reasoning capabilities that approach “Artificial Gen-eral Intelligence”. On the other hand, previous autonomous driving methods tend to rely on limited-format inputs (e.g., sensor data and navigation waypoints), restricting the vehi-cle's ability to understand language information and inter-act with humans. To this end, this paper introduces LM-Drive, a novel language-guided, end-to-end, closed-loop autonomous driving framework. LMDrive uniquely processes and integrates multimodal sensor data with naturallanguage instructions, enabling interaction with humans and navigation software in realistic instructional settings. To facilitate research in language-based closed-loop autonomous driving, we also publicly release the corresponding dataset which includes approximately 64K instruction-following data clips, and the LangAuto benchmark that tests the system's ability to handle complex instructions and challenging driving scenarios. Extensive closed-loop experiments are conducted to demonstrate LMDrive's effectiveness. To the best of our knowledge, we're the very first work to leverage LLMs for closed-loop end-to-end autonomous driving. Code is available on our webpage.
Hao Shao, Guanglu Song, Steven Lake Waslander, Yu Liu 0015, Hongsheng Li 0001
CVPR5
2024 SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction
abstract
Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dy-namic, human-robot-mixed environments. Context information, such as road maps and surrounding agents' states, provides crucial geometric and semantic information for motion behavior prediction. To this end, recent works explore two-stage prediction frameworks where coarse trajectories are first proposed, and then used to select critical context information for trajectory refinement. However, they either incur a large amount of computation or bring limited improvement, if not both. In this paper, we introduce a novel scenario-adaptive refinement strategy, named SmartRefine, to refine prediction with minimal additional computation. Specifically, SmartRefine can comprehensively adapt refinement configurations based on each scenario's properties, and smartly chooses the number of refinement iterations by introducing a quality score to measure the prediction quality and remaining refinement potential of each scenario. SmartRefine is designed as a generic and flexible approach that can be seamlessly integrated into most state-of-the-art motion prediction models. Experiments on Argoverse (1 & 2) show that our method consistently improves the prediction accuracy of multiple state-of-the-art prediction models. Specifically, by adding SmartRefine to QCNet, we outper-form all published ensemble-free works on the Argoverse 2 leaderboard (single agent track) at submission11November 2023.: Compre-hensive studies are also conducted to ablate design choices and explore the mechanism behind multi-iteration refinement. Codes are available at our webpage.
Yang Zhou 0054, Hao Shao, Steven Lake Waslander, Hongsheng Li 0001, Yu Liu 0015
CVPR4
2024 JDT3D: Addressing the Gaps in LiDAR-Based Tracking-by-Attention
Brian Cheong, Jiachen Zhou 0002, Steven Lake Waslander
ECCV (66)3
2024 Image-to-Lidar Relational Distillation for Autonomous Driving Data
Anas Mahmoud 0002, Ali Harakeh, Steven Lake Waslander
ECCV (62)3
2024 UncertaintyTrack: Exploiting Detection and Localization Uncertainty in Multi-Object Tracking
abstract
Multi-object tracking (MOT) methods have seen a significant boost in performance recently, due to strong interest from the research community and steadily improving object detection methods. The majority of tracking methods, which follow the tracking-by-detection (TBD) paradigm, blindly trust the incoming detections with no sense of their associated localization uncertainty. This lack of uncertainty awareness poses a problem in safety-critical tasks such as autonomous driving where passengers could be put at risk due to erroneous detections that have propagated to downstream tasks, including MOT. While there are existing works in probabilistic object detection that predict the localization uncertainty around the boxes, no work in 2D MOT for autonomous driving has studied whether these estimates are meaningful enough to be leveraged effectively in object tracking. We introduce UncertaintyTrack, a collection of extensions that can be applied to multiple TBD trackers to account for localization uncertainty estimates from probabilistic object detectors. Experiments on the Berkeley Deep Drive MOT dataset show that the combination of our method and informative uncertainty estimates reduces the number of ID switches by around 19% and improves mMOTA by 2-3%. The source code is available at https://github.com/TRAILab/UncertaintyTrack
Steven Lake Waslander
ICRA2
2024 Uncertainty-aware 3D Object-Level Mapping with Deep Shape Priors
abstract
3D object-level mapping is a fundamental problem in robotics, which is especially challenging when object CAD models are unavailable during inference. We propose a framework that can reconstruct high-quality object-level maps for unknown objects. Our approach takes multiple RGB-D images as input and outputs dense 3D shapes and 9-DoF poses (including 3 scale parameters) for detected objects. The core idea is to leverage a learnt generative model for a category of object shapes as priors and to formulate a probabilistic, uncertainty-aware optimization framework for 3D reconstruction. We derive a probabilistic formulation that propagates shape and pose uncertainty through two novel loss functions. Unlike current state-of-the-art approaches, we explicitly model the uncertainty of the object shapes and poses during our optimization, resulting in a high-quality object-level mapping system. Moreover, the estimated shape and pose uncertainties, which we demonstrate can accurately reflect the true errors of our object maps, can be useful for downstream robotics tasks such as active vision. We perform extensive evaluations on indoor and outdoor real-world datasets, achieving substantial improvements over state-of-the-art methods. Our code is available at https://github.com/TRAILab/UncertainShapePose.
Ziwei Liao, Jun Yang 0053, Jingxing Qian, Angela P. Schoellig, Steven Lake Waslander
ICRA5
2024 SWTrack: Multiple Hypothesis Sliding Window 3D Multi-Object Tracking
abstract
Modern robotic systems are required to operate in dense dynamic environments, requiring highly accurate real-time track identification and estimation. For 3D multi-object tracking, recent approaches process a single measurement frame recursively with greedy association and are prone to errors in ambiguous association decisions. Our method, Sliding Window Tracker (SWTrack), yields more accurate association and state estimation by batch processing many frames of sensor data while being capable of running online in real-time. The most probable track associations are identified by evaluating all possible track hypotheses across the temporal sliding window. A novel graph optimization approach is formulated to solve the multidimensional assignment problem with lifted graph edges introduced to account for missed detections and graph sparsity enforced to retain real-time efficiency. We evaluate our SWTrack implementation on the NuScenes autonomous driving dataset to demonstrate improved tracking performance.
Sandro Papais, Robert Ren, Steven Lake Waslander
ICRA3
2024 Active Pose Refinement for Textureless Shiny Objects using the Structured Light Camera
abstract
6D pose estimation of textureless shiny objects has become an essential problem in many robotic applications. Many pose estimators require high-quality depth data, often measured by structured light cameras. However, when objects have shiny surfaces (e.g., metal parts), these cameras fail to sense complete depths from a single viewpoint due to the specular reflection, resulting in a significant drop in the final pose accuracy. To mitigate this issue, we present a complete active vision framework for 6D object pose refinement and next-best-view prediction. Specifically, we first develop an optimization-based pose refinement module for the structured light camera. Our system then selects the next best camera viewpoint to collect depth measurements by minimizing the predicted uncertainty of the object pose. Compared to previous approaches, we additionally predict measurement uncertainties of future viewpoints by online rendering, which significantly improves the next-best-view prediction performance. We test our method on the real-world ROBI dataset. The results show that our pose refinement module outperforms the traditional ICP-based approach when given the same input depth data, and our next-best-view strategy can achieve high object pose accuracy with significantly fewer viewpoints than the heuristic-based policies.
Jun Yang 0053, Steven Lake Waslander
IROS3
2024 DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
abstract
We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that predicts a rich neural scene representation from sparse, single-frame multi-view camera inputs with limited view overlap, and is trained self-supervised with differentiable rendering to reconstruct RGB, depth, or feature images. Our first insight is to exploit per-scene optimized Neural Radiance Fields (NeRFs) by generating dense depth and virtual camera targets from them, which helps our model to learn enhanced 3D geometry from sparse non-overlapping image inputs. Second, to learn a semantically rich 3D representation, we propose distilling features from pre-trained 2D foundation models, such as CLIP or DINOv2, thereby enabling various downstream tasks without the need for costly 3D human annotations. To leverage these two insights, we introduce a novel model architecture with a two-stage lift-splat-shoot encoder and a parameterized sparse hierarchical voxel representation. Experimental results on the NuScenes and Waymo NOTR datasets demonstrate that DistillNeRF significantly outperforms existing comparable state-of-the-art self-supervised methods for scene reconstruction, novel view synthesis, and depth estimation; and it allows for competitive zero-shot 3D semantic occupancy prediction, as well as open-world scene understanding through distilled foundation model features. Demos and code will be available at https://distillnerf.github.io/.
Seung Wook Kim 0001, Jiawei Yang 0002, Cunjun Yu, Boris Ivanovic, Steven Lake Waslander, Yue Wang 0041, Sanja Fidler, Marco Pavone 0001, Péter Karkus
NeurIPS6
2024 Multi-view 3D Object Reconstruction and Uncertainty Modelling with Neural Shape Prior
abstract
3D object reconstruction is important for semantic scene understanding. It is challenging to reconstruct detailed 3D shapes from monocular images directly due to a lack of depth information, occlusion and noise. Most current methods generate deterministic object models without any awareness of the uncertainty of the reconstruction. We tackle this problem by leveraging a neural object representation which learns an object shape distribution from large dataset of 3d object models and maps it into a latent space. We propose a method to model uncertainty as part of the representation and define an uncertainty-aware encoder which generates latent codes with uncertainty directly from individual input images. Further, we propose a method to propagate the uncertainty in the latent code to SDF values and generate a 3d object mesh with local uncertainty for each mesh component. Finally, we propose an incremental fusion method under a Bayesian framework to fuse the latent codes from multi-view observations. We evaluate the system in both synthetic and real datasets to demonstrate the effectiveness of uncertainty-based fusion to improve 3D object reconstruction accuracy.
Ziwei Liao, Steven Lake Waslander
WACV2
2024 Bayesian Embeddings for Few-Shot Open World Recognition
abstract
As autonomous decision-making agents move from narrow operating environments to unstructured worlds, learning systems must move from a closed-world formulation to an open-world and few-shot setting in which agents continuously learn new classes from small amounts of information. This stands in stark contrast to modern machine learning systems that are typically designed with a known set of classes and a large number of examples for each class. In this work we extend embedding-based few-shot learning algorithms to the open-world recognition setting. We combine Bayesian non-parametric class priors with an embedding-based pre-training scheme to yield a highly flexible framework which we refer to as few-shot learning for open world recognition (FLOWR). We benchmark our framework on open-world extensions of the common MiniImageNet and TieredImageNet few-shot learning datasets. Our results show, compared to prior methods, strong classification accuracy performance and up to a 12% improvement in H-measure (a measure of novel class detection) from our non-parametric open-world few-shot learning scheme.
John Willes, James Harrison, Ali Harakeh, Chelsea Finn, Marco Pavone 0001, Steven Lake Waslander
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Estimating Regression Predictive Distributions with Sample Networks
abstract
Estimating the uncertainty in deep neural network predictions is crucial for many real-world applications. A common approach to model uncertainty is to choose a parametric distribution and fit the data to it using maximum likelihood estimation. The chosen parametric form can be a poor fit to the data-generating distribution, resulting in unreliable uncertainty estimates. In this work, we propose SampleNet, a flexible and scalable architecture for modeling uncertainty that avoids specifying a parametric form on the output distribution. SampleNets do so by defining an empirical distribution using samples that are learned with the Energy Score and regularized with the Sinkhorn Divergence. SampleNets are shown to be able to well-fit a wide range of distributions and to outperform baselines on large-scale real-world regression tasks.
Ali Harakeh, Jordan S. K. Hu, Naiqing Guan, Steven Lake Waslander, Liam Paull
AAAI4
2023 Self-Supervised Image-to-Point Distillation via Semantically Tolerant Contrastive Loss
abstract
An effective framework for learning 3D representations for perception tasks is distilling rich self-supervised image features via contrastive learning. However, image-to-point representation learning for autonomous driving datasets faces two main challenges: 1) the abundance of self-similarity, which results in the contrastive losses pushing away semantically similar point and image regions and thus disturbing the local semantic structure of the learned representations, and 2) severe class imbalance as pretraining gets dominated by over-represented classes. We propose to alleviate the self-similarity problem through a novel semantically tolerant image-to-point contrastive loss that takes into consideration the semantic distance between positive and negative image regions to minimize contrasting semantically similar point and image regions. Additionally, we address class imbalance by designing a class-agnostic balanced loss that approximates the degree of class imbalance through an aggregate sample-to-samples semantic similarity measure. We demonstrate that our semantically-tolerant contrastive loss with class balancing improves state-of-the-art 2D-to-3D representation learning in all evaluation settings on 3D semantic segmentation. Our method consistently outperforms state-of-the-art 2D-to-3D representation learning frameworks across a wide range of 2D self-supervised pretrained models.
Anas Mahmoud 0002, Jordan S. K. Hu, Tianshu Kuai, Ali Harakeh, Liam Paull, Steven Lake Waslander
CVPR6
2023 ReasonNet: End-to-End Driving with Temporal and Global Reasoning
abstract
The large-scale deployment of autonomous vehicles is yet to come, and one of the major remaining challenges lies in urban dense traffic scenarios. In such cases, it remains challenging to predict the future evolution of the scene and future behaviors of objects, and to deal with rare adverse events such as the sudden appearance of occluded objects. In this paper, we present ReasonNet, a novel end-to-end driving framework that extensively exploits both temporal and global information of the driving scene. By reasoning on the temporal behavior of objects, our method can effectively process the interactions and relationships among features in different frames. Reasoning about the global information of the scene can also improve overall perception performance and benefit the detection of adverse events, especially the anticipation of potential danger from occluded objects. For comprehensive evaluation on occlusion events, we also release publicly a driving simulation benchmark DriveOcclusionSim consisting of diverse occlusion events. We conduct extensive experiments on multiple CARLA benchmarks, where our model outperforms all prior methods, ranking first on the sensor track of the public CARLA Leaderboard [53].
Hao Shao, Ruobing Chen 0005, Steven Lake Waslander, Hongsheng Li 0001, Yu Liu 0015
CVPR4
2023 6D Pose Estimation for Textureless Objects on RGB Frames using Multi-View Optimization
abstract
6D pose estimation of textureless objects is a valuable but challenging task for many robotic applications. In this work, we propose a framework to address this challenge using only RGB images acquired from multiple viewpoints. The core idea of our approach is to decouple 6D pose estimation into a sequential two-step process, first estimating the 3D translation and then the 3D rotation of each object. This decoupled formulation first resolves the scale and depth ambiguities in single RGB images, and uses these estimates to accurately identify the object orientation in the second stage, which is greatly simplified with an accurate scale estimate. Moreover, to accommodate the multi-modal distribution present in rotation space, we develop an optimization scheme that explicitly handles object symmetries and counteracts measurement uncertainties. In comparison to the state-of-the-art multi-view approach, we demonstrate that the proposed approach achieves substantial improvements on a challenging 6D pose estimation dataset for textureless objects.
Jun Yang 0053, Wenjie Xue, Sahar Ghavidel, Steven Lake Waslander
ICRA4
2023 Dense Voxel Fusion for 3D Object Detection
abstract
Camera and LiDAR sensor modalities provide complementary appearance and geometric information useful for detecting 3D objects for autonomous vehicle applications. However, current end-to-end fusion methods are challenging to train and underperform state-of-the-art LiDAR-only detectors. Sequential fusion methods suffer from a limited number of pixel and point correspondences due to point cloud sparsity, or their performance is strictly capped by the detections of one of the modalities. Our proposed solution, Dense Voxel Fusion (DVF) is a sequential fusion method that generates multi-scale dense voxel feature representations, improving expressiveness in low point density regions. To enhance multi-modal learning, we train directly with projected ground truth 3D bounding box labels, avoiding noisy, detector-specific 2D predictions. Both DVF and the multi-modal training approach can be applied to any voxel-based LiDAR backbone. DVF ranks 3rdamong published fusion methods on KITTI’s 3D car detection benchmark without introducing additional trainable parameters, nor requiring stereo images or dense depth labels. In addition, DVF significantly improves 3D vehicle detection performance of voxel-based methods on the Waymo Open Dataset.
Anas Mahmoud 0002, Jordan S. K. Hu, Steven Lake Waslander
WACV3
2022 Point Density-Aware Voxels for LiDAR 3D Object Detection
abstract
LiDAR has become one of the primary 3D object detection sensors in autonomous driving. However, LiDAR's diverging point pattern with increasing distance results in a non-uniform sampled point cloud ill-suited to discretized volumetric feature extraction. Current methods either rely on voxelized point clouds or use inefficient farthest point sampling to mitigate detrimental effects caused by density variation but largely ignore point density as a feature and its predictable relationship with distance from the LiDAR sensor. Our proposed solution, Point Density-Aware Voxel network (PDV), is an end-to-end two stage LiDAR 3D object detection architecture that is designed to account for these point density variations. PDV efficiently localizes voxel features from the 3D sparse convolution backbone through voxel point centroids. The spatially localized voxel features are then aggregated through a density-aware RoI grid pooling module using kernel density estimation (KDE) and self attention with point density positional encoding. Finally, we exploit LiDAR's point density to distance relationship to refine our final bounding box confidences. PDV outperforms all state-of-the-art methods on the Waymo Open Dataset and achieves competitive results on the KITTI dataset.
Jordan S. K. Hu, Tianshu Kuai, Steven Lake Waslander
CVPR3
2022 Next-Best-View Prediction for Active Stereo Cameras and Highly Reflective Objects
abstract
Depth acquisition with the active stereo camera is a challenging task for highly reflective objects. When setup permits, multi-view fusion can provide increased levels of depth completion. However, due to the slow acquisition speed of high-end active stereo cameras, collecting a large number of viewpoints for a single scene is generally not practical. In this work, we propose a next-best-view framework to strategically select camera viewpoints for completing depth data on reflective objects. In particular, we explicitly model the specular reflection of reflective surfaces based on the Phong reflection model and a photometric response function. Given the object CAD model and grayscale image, we employ an RGB-based pose estimator to obtain current pose predictions from the existing data, which is used to form predicted surface normal and depth hypotheses, and allows us to then assess the information gain from a subsequent frame for any candidate viewpoint. Using this formulation, we implement an active perception pipeline which is evaluated on a challenging real-world dataset. The evaluation results demonstrate that our active depth acquisition method outperforms two strong baselines for both depth completion and object pose estimation performance.
Jun Yang 0053, Steven Lake Waslander
ICRA2
2022 LiDAR-MIMO: Efficient Uncertainty Estimation for LiDAR-based 3D Object Detection
abstract
The estimation of uncertainty in robotic vision, such as 3D object detection, is an essential component in developing safe autonomous systems aware of their own performance. However, the deployment of current uncertainty estimation methods in 3D object detection remains challenging due to timing and computational constraints. To tackle this issue, we propose LiDAR-MIMO, an adaptation of the multi-input multi-output (MIMO) uncertainty estimation method to the LiDAR-based 3D object detection task. Our method modifies the original MIMO by performing multi-input at the feature level to ensure the detection, uncertainty estimation, and runtime performance benefits are retained despite the limited capacity of the underlying detector and the large computational costs of point cloud processing. We compare LiDAR-MIMO with MC dropout and ensembles as baselines and show comparable uncertainty estimation results with only a small number of output heads. Further, LiDAR-MIMO can be configured to be twice as fast as MC dropout and ensembles, while achieving higher mAP than MC dropout and approaching that of ensembles.
Matthew Pitropov, Chengjie Huang, Vahdat Abdelzad, Krzysztof Czarnecki 0001, Steven Lake Waslander
IV5
2022 A Review and Comparative Study on Probabilistic Object Detection in Autonomous Driving
abstract
Capturing uncertainty in object detection is indispensable for safe autonomous driving. In recent years, deep learning has become the de-facto approach for object detection, and many probabilistic object detectors have been proposed. However, there is no summary on uncertainty estimation in deep object detection, and existing methods are either built with different network architectures and uncertainty estimation methods, or evaluated on different datasets with a wide range of evaluation metrics. As a result, a comparison among methods remains challenging, as does the selection of a model that best suits a particular application. This paper aims to alleviate this problem by providing a review and comparative study on existing probabilistic object detection methods for autonomous driving applications. First, we provide an overview of practical uncertainty estimation methods in deep learning, and then systematically survey existing methods and evaluation metrics for probabilistic object detection. Next, we present a strict comparative study for probabilistic object detection based on an image detector and three public autonomous driving datasets. Finally, we present a discussion of the remaining challenges and future works. Code has been made available athttps://github.com/asharakeh/pod_compare.git.
Di Feng, Ali Harakeh, Steven Lake Waslander, Klaus Dietmayer
IEEE Trans. Intell. Transp. Syst.3
2021 Categorical Depth Distribution Network for Monocular 3D Object Detection
abstract
Monocular 3D object detection is a key problem for autonomous vehicles, as it provides a solution with simple configuration compared to typical multi-sensor systems. The main challenge in monocular 3D detection lies in accurately predicting object depth, which must be inferred from object and scene cues due to the lack of direct range measurement. Many methods attempt to directly estimate depth to assist in 3D detection, but show limited performance as a result of depth inaccuracy. Our proposed solution, Categorical Depth Distribution Network (CaDDN), uses a predicted categorical depth distribution for each pixel to project rich contextual feature information to the appropriate depth interval in 3D space. We then use the computationally efficient bird’s-eye-view projection and single-stage detector to produce the final output detections. We design CaDDN as a fully differentiable end-to-end approach for joint depth estimation and object detection. We validate our approach on the KITTI 3D object detection benchmark, where we rank 1stamong published monocular methods. We also provide the first monocular 3D detection results on the newly released Waymo Open Dataset. We provide a code release for CaDDN which is made available.
Cody Reading, Ali Harakeh, Julia Chae, Steven Lake Waslander
CVPR4
2021 Estimating and Evaluating Regression Predictive Uncertainty in Deep Object Detectors
Ali Harakeh, Steven Lake Waslander
ICLR2
2021 ROBI: A Multi-View Dataset for Reflective Objects in Robotic Bin-Picking
abstract
In robotic bin-picking applications, the perception of texture-less, highly reflective parts is a valuable but challenging task. The high glossiness can introduce fake edges in RGB images and inaccurate depth measurements, especially in heavily cluttered bin scenarios. In this paper, we present the ROBI (Reflective Objects in BIns) dataset, a public dataset for 6D object pose estimation and multi-view depth fusion in robotic bin-picking scenarios. The ROBI dataset includes a total of 63 bin-picking scenes captured with two active stereo cameras: a high-cost Ensenso sensor and a low-cost RealSense sensor. For each scene, the monochrome/RGB images and depth maps are captured from sampled view spheres around the scene, and are annotated with accurate 6D poses of visible objects and an associated visibility score. For evaluating the performance of depth fusion, we captured the ground truth depth maps by high-cost Ensenso camera with objects coated in anti-reflective scanning spray. To show the utility of the dataset, we evaluated the representative algorithms of 6D object pose estimation and multi-view depth fusion on the full dataset. Evaluation results demonstrate the difficulty of highly reflective objects, especially in difficult cases due to the degradation of depth data quality, severe occlusions, and cluttered scenes. The ROBI dataset is available online at https://www.trailab.utias.utoronto.ca/robi.
Jun Yang 0053, Yizhou Gao, Steven Lake Waslander
IROS4
2020 BayesOD: A Bayesian Approach for Uncertainty Estimation in Deep Object Detectors
abstract
When incorporating deep neural networks into robotic systems, a major challenge is the lack of uncertainty measures associated with their output predictions. Methods for uncertainty estimation in the output of deep object detectors (DNNs) have been proposed in recent works, but have had limited success due to 1) information loss at the detectors nonmaximum suppression (NMS) stage, and 2) failure to take into account the multitask, many-to-one nature of anchor-based object detection. To that end, we introduce BayesOD, an uncertainty estimation approach that reformulates the standard object detector inference and Non-Maximum suppression components from a Bayesian perspective. Experiments performed on four common object detection datasets show that BayesOD provides uncertainty estimates that are better correlated with the accuracy of detections, manifesting as a significant reduction of 9.77%-13.13% on the minimum Gaussian uncertainty error metric and a reduction of 1.63%-5.23% on the minimum Categorical uncertainty error metric. Code will be released at https://github.com/asharakeh/bayes-od-rc.
Ali Harakeh, Michael Smart, Steven Lake Waslander
ICRA3
2020 Object-Centric Stereo Matching for 3D Object Detection
abstract
Safe autonomous driving requires reliable 3D object detection-determining the 6 DoF pose and dimensions of objects of interest. Using stereo cameras to solve this task is a cost-effective alternative to the widely used LiDAR sensor. The current state-of-the-art for stereo 3D object detection takes the existing PSMNet stereo matching network, with no modifications, and converts the estimated disparities into a 3D point cloud, and feeds this point cloud into a LiDAR-based 3D object detector. The issue with existing stereo matching networks is that they are designed for disparity estimation, not 3D object detection; the shape and accuracy of object point clouds are not the focus. Stereo matching networks commonly suffer from inaccurate depth estimates at object boundaries, which we define as streaking, because background and foreground points are jointly estimated. Existing networks also penalize disparity instead of the estimated position of object point clouds in their loss functions. We propose a novel 2D box association and object-centric stereo matching method that only estimates the disparities of the objects of interest to address these two issues. Our method achieves state-of-the-art results on the KITTI 3D and BEV benchmarks.
Alex D. Pon, Jason Ku 0001, Steven Lake Waslander
ICRA4
2020 AC/DCC : Accurate Calibration of Dynamic Camera Clusters for Visual SLAM
abstract
In order to relate information across cameras in a Dynamic Camera Cluster (DCC), an accurate time-varying set of extrinsic calibration transformations need to be determined. Previous calibration approaches rely solely on collecting measurements from a known fiducial target which limits calibration accuracy as insufficient excitation of the gimbal is achieved. In this paper, we improve DCC calibration accuracy by collecting measurements over the entire configuration space of the gimbal and achieve a 10X improvement in pixel re-projection error. We perform a joint optimization over the calibration parameters between any number of cameras and unknown joint angles using a pose-loop error optimization approach, thereby avoiding the need for overlapping fields-of-view. We test our method in simulation and provide a calibration sensitivity analysis for different levels of camera intrinsic and joint angle noise. In addition, we provide a novel analysis of the degenerate parameters in the calibration when joint angle values are unknown, which avoids situations in which the calibration cannot be uniquely recovered. The calibration code will be made available at https://github.com/TRAILab/AC-DCC.
Jason Rebello, Angus Fung, Steven Lake Waslander
ICRA3
2020 Confidence Guided Stereo 3D Object Detection with Split Depth Estimation
abstract
Accurate and reliable 3D object detection is vital to safe autonomous driving. Despite recent developments, the performance gap between stereo-based methods and LiDAR-based methods is still considerable. Accurate depth estimation is crucial to the performance of stereo-based 3D object detection methods, particularly for those pixels associated with objects in the foreground. Moreover, stereo-based methods suffer from high variance in the depth estimation accuracy, which is often not considered in the object detection pipeline. To tackle these two issues, we propose CG-Stereo, a confidence-guided stereo 3D object detection pipeline that uses separate decoders for foreground and background pixels during depth estimation, and leverages the confidence estimation from the depth estimation network as a soft attention mechanism in the 3D object detector. Our approach outperforms all state-of-the-art stereo-based 3D detectors on the KITTI benchmark.
Jason Ku 0001, Steven Lake Waslander
IROS3
2020 TruPercept: Trust Modelling for Autonomous Vehicle Cooperative Perception from Synthetic Data
abstract
Inter-vehicle communication for autonomous vehicles (AVs) stands to provide significant benefits in terms of perception robustness. We propose a novel approach for AVs to communicate perceptual observations, tempered by trust modelling of peers providing reports. Based on the accuracy of reported object detections as verified locally, communicated messages can be fused to augment perception performance beyond line of sight and at great distance from the ego vehicle. Also presented is a new synthetic dataset which can be used to test cooperative perception. The TruPercept dataset includes unreliable and malicious behaviour scenarios to experiment with some challenges cooperative perception introduces. The TruPercept runtime and evaluation framework allows modular component replacement to facilitate ablation studies as well as the creation of new trust scenarios we are able to show.
Braden Hurl, Robin Cohen, Krzysztof Czarnecki 0001, Steven Lake Waslander
IV4
2019 Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction
abstract
We present MonoPSR, a monocular 3D object detection method that leverages proposals and shape reconstruction. First, using the fundamental relations of a pinhole camera model, detections from a mature 2D object detector are used to generate a 3D proposal per object in a scene. The 3D location of these proposals prove to be quite accurate, which greatly reduces the difficulty of regressing the final 3D bounding box detection. Simultaneously, a point cloud is predicted in an object centered coordinate system to learn local scale and shape information. However, the key challenge is how to exploit shape information to guide 3D localization. As such, we devise aggregate losses, including a novel projection alignment loss, to jointly optimize these tasks in the neural network to improve 3D localization accuracy. We validate our method on the KITTI benchmark where we set new state-of-the-art results among published monocular methods, including the harder pedestrian and cyclist classes, while maintaining efficient run-time.
Jason Ku 0001, Alex D. Pon, Steven Lake Waslander
CVPR3
2019 Improving 3D Object Detection for Pedestrians with Virtual Multi-View Synthesis Orientation Estimation
abstract
Accurately estimating the orientation of pedestrians is an important and challenging task for autonomous driving because this information is essential for tracking and predicting pedestrian behavior. This paper presents a flexible Virtual Multi-View Synthesis module that can be adopted into 3D object detection methods to improve orientation estimation. The module uses a multi-step process to acquire the fine-grained semantic information required for accurate orientation estimation. First, the scene's point cloud is densified using a structure preserving depth completion algorithm and each point is colorized using its corresponding RGB pixel. Next, virtual cameras are placed around each object in the densified point cloud to generate novel viewpoints, which preserve the object's appearance. We show that this module greatly improves the orientation estimation on the challenging pedestrian class on the KITTI benchmark. When used with the open-source 3D detector AVOD-FPN, we outperform all other published methods on the pedestrian Orientation, 3D, and Bird's Eye View benchmarks.
Jason Ku 0001, Alex D. Pon, Sean Walsh 0003, Steven Lake Waslander
IROS4
2019 Precise Synthetic Image and LiDAR (PreSIL) Dataset for Autonomous Vehicle Perception
abstract
We introduce the Precise Synthetic Image and LiDAR (PreSIL) dataset for autonomous vehicle perception. Grand Theft Auto V (GTA V), a commercial video game, has a large detailed world with realistic graphics, which provides a diverse data collection environment. Existing works creating synthetic LiDAR data for autonomous driving with GTA V have not released their datasets, rely on an in-game raycasting function which represents people as cylinders, and can fail to capture vehicles past 30 metres. Our work creates a precise LiDAR simulator within GTA V which collides with detailed models for all entities no matter the type or position. The PreSIL dataset consists of over 50,000 frames and includes high-definition images with full resolution depth information, semantic segmentation (images), point-wise segmentation (point clouds), and detailed annotations for all vehicles and people. Collecting additional data with our framework is entirely automatic and requires no human annotation of any kind. We demonstrate the effectiveness of our dataset by showing an improvement of up to 5% average precision on the KITTI 3D Object Detection benchmark challenge when state-of-the-art 3D object detection networks are pre-trained with our data. The data and code are available at https://tinyurl.com/y3tb9sxy.
Braden Hurl, Krzysztof Czarnecki 0001, Steven Lake Waslander
IV3
2019 Multistep Prediction of Dynamic Systems With Recurrent Neural Networks
abstract
In this paper, we address the state initialization problem in recurrent neural networks (RNNs), which seeks proper values for the RNN initial states at the beginning of a prediction interval. The proposed methods employ various forms of neural networks (NNs) to generate proper initial state values for RNNs. A variety of RNNs are trained using the proposed NN initialization schemes for modeling two aerial vehicles, a helicopter and a quadrotor, from experimental data. It is shown that the RNN initialized by the NN-based initialization method outperforms the washout method which is commonly used to initialize RNNs. Furthermore, a comprehensive study of RNNs trained for multistep prediction of the two aerial vehicles is presented. The multistep prediction of the quadrotor is enhanced using a hybrid model, which combines a simplified physics-based motion model of the vehicle with RNNs. While the maximum translational and rotational velocities in the Quadrotor data set are about 4 m/s and 3.8 rad/s, respectively, the hybrid model produces predictions, over 1.9 s, which remain within 9 cm/s and 0.12 rad/s of the measured translational and rotational velocities, with 99% confidence on the test data set.
Nima Mohajerin, Steven Lake Waslander
IEEE Trans. Neural Networks Learn. Syst.2
2018 Encoderless Gimbal Calibration of Dynamic Multi-Camera Clusters
abstract
Dynamic Camera Clusters (DCCs) are multi-camera systems where one or more cameras are mounted on actuated mechanisms such as a gimbal. Existing methods for DCC calibration rely on joint angle measurements to resolve the time-varying transformation between the dynamic and static camera. This information is usually provided by motor encoders, however, joint angle measurements are not always readily available on off-the-shelf mechanisms. In this paper, we present an encoderless approach for DCC calibration which simultaneously estimates the kinematic parameters of the transformation chain as well as the unknown joint angles. We also demonstrate the integration of an encoderless gimbal mechanism with a state-of-the art VIO algorithm, and show the extensions required in order to perform simultaneous online estimation of the joint angles and vehicle localization state. The proposed calibration approach is validated both in simulation and on a physical DCC composed of a 2-DOF gimbal mounted on a UAV. Finally, we show the experimental results of the calibrated mechanism integrated into the OKVIS VIO package, and demonstrate successful online joint angle estimation while maintaining localization accuracy that is comparable to a standard static multi-camera configuration.
Christopher L. Choi, Jason Rebello, Leonid Koppel, Pranav Ganti, Arun Das 0005, Steven Lake Waslander
ICRA6
2018 Deep Learning a Quadrotor Dynamic Model for Multi-Step Prediction
abstract
We develop a multi-step motion prediction modeling method for dynamic systems over long horizons using deep learning. Building on previous work, we propose a novel hybrid network architecture, by combining deep recurrent neural networks with a quadrotor motion model created using classic system identification methods. The proposed model takes only the initial system state and motor speeds over the prediction horizon as inputs and returns robust state predictions for up to two seconds of motion at 100 Hz. We employ recurrent neural network state initialization during training, to exploit real-world dataset collected from quadrotor vehicle flights in an indoor flight arena. Our experiments demonstrate that the proposed hybrid network model consistently outperforms both black box and rigid body dynamics predictions over single and multi-step prediction scenarios, with an order of magnitude improvements in velocity estimates in particular.
Nima Mohajerin, Melissa Mozifian, Steven Lake Waslander
ICRA3
2018 Joint 3D Proposal Generation and Object Detection from View Aggregation
abstract
We present AVOD, an Aggregate View Object Detection network for autonomous driving scenarios. The proposed neural network architecture uses LIDAR point clouds and RGB images to generate features that are shared by two subnetworks: a region proposal network (RPN) and a second stage detector network. The proposed RPN uses a novel architecture capable of performing multimodal feature fusion on high resolution feature maps to generate reliable 3D object proposals for multiple object classes in road scenes. Using these proposals, the second stage detection network performs accurate oriented 3D bounding box regression and category classification to predict the extents, orientation, and classification of objects in 3D space. Our proposed architecture is shown to produce state of the art results on the KITTI 3D object detection benchmark [1] while running in real time with a low memory footprint, making it a suitable candidate for deployment on autonomous vehicles. Code is available at: https://github.com/kujason/avod.
Jason Ku 0001, Melissa Mozifian, Jungwook Lee, Ali Harakeh, Steven Lake Waslander
IROS5
2017 State initialization for recurrent neural network modeling of time-series data
abstract
To use a Recurrent Neural Network (RNN) for time series modeling, it is essential to properly initialize the network, that is, to set the hidden neuron outputs properly at the initial time. Normally, an RNN is initialized with zero state values or at steady state. In the context of dynamic system identification, such initializations imply the system to be modelled is in steady state, i.e., capturing transient behaviour of the system is difficult if the network states are not properly initialized. If the network initial states are not calculable from the training data, then a method to infer them, both throughout the training and validation phases, is needed. In this paper, we use a feed forward neural network to initialize a structurally deep recurrent neural network in learning and multi-step prediction of the altitude of a real quadrotor vehicle. To the best of our knowledge, this is the first time a neural network has outperformed a physics based model for multi-step time series prediction from recorded quadrotor flight data.
Nima Mohajerin, Steven Lake Waslander
IJCNN2
2017 Autonomous active calibration of a dynamic camera cluster using next-best-view
abstract
Dynamic camera cluster (DCC) calibration determines a time-varying set of extrinsic calibration transformations between cameras in a multi-camera cluster, where one or more cameras are mounted on actuated mechanisms. In this paper, we present a novel active vision approach for DCC calibration, which directly reduces the parameter uncertainty by selecting calibration measurements using an information theoretic next-best-view policy. Our system automatically selects the next best measurement for the calibration by determining the optimal actuator inputs which minimize the predicted covariance of the extrinsic parameters. We show that our method is able to successfully estimate calibration parameters up to a user specified accuracy with no manual excitation. We test our method in simulation on a variety of actuated mechanisms and validate the results on a real 3 axis gimbal, and demonstrate our approach is able to achieve accurate calibrations using fewer measurement sets when compared to existing approaches.
Jason Rebello, Arun Das 0005, Steven Lake Waslander
IROS3
2016 The constriction decomposition method for coverage path planning
abstract
The task of coverage path planning in 2D indoor and outdoor environments is classified as a NP-hard problem, and has been an active research topic for over 30 years. We derive a novel, exact cellular decomposition method called the Constriction Decomposition Method (CDM) and apply it to complex indoor environments. The CDM rapidly identifies constriction points in the environment and decomposes the environment into easily traversable cells by exploiting the geometric information contained in the environment's straight skeleton. Several heuristic path planning methods that find paths that completely cover each cell are explored. We apply our method on a complex indoor office-like environment and compare our results to existing methods.
Stanley Brown, Steven Lake Waslander
IROS2
2016 Calibration of a dynamic camera cluster for multi-camera visual SLAM
abstract
Multi-camera clusters used for visual SLAM assume a fixed calibration between the cameras, which places many limitations on its performance, and directly excludes all configurations where a camera in the cluster is mounted to a moving component. In this work, we present a calibration method for dynamic multi-camera clusters, where one or more of the cluster cameras is mounted to an actuated mechanism, such as a gimbal or robotic manipulator. Our calibration approach parametrizes the actuated mechanism using the Denavit-Hartenberg convention, then determines the calibration parameters which allow for the estimation of the time varying extrinsic transformations between camera frames. We validate our calibration approach using a dynamic camera cluster consisting of a static camera and a camera mounted to a pan-tilt unit, and demonstrate that the dynamic camera cluster can provide accurate tracking when used to perform SLAM.
Arun Das 0005, Steven Lake Waslander
IROS2
2016 Degenerate motions in multicamera cluster SLAM with non-overlapping fields of view
Michael J. Tribou, David Wang 0001, Steven Lake Waslander
Image Vis. Comput.3
2015 Entropy based keyframe selection for Multi-Camera Visual SLAM
abstract
Although many state-of-the-art visual SLAM algorithms use keyframes to help alleviate the computational requirements of performing online bundle adjustment, little consideration is taken for specific keyframe selection. In this work, we propose two entropy based methods which aim to insert keyframes that will directly improve the system's ability to localize. The first approach inserts keyframes based on the cumulative point entropy reduction in the existing map, while the second approach uses the predicted point flow discrepancy to select keyframes which best initializes new features for the camera to track against in the future. We implement the proposed methods within the Multi-Camera Parallel Mapping and Tracking framework, and demonstrate the effectiveness of our methods using ground truth data collected using an indoor positioning system.
Arun Das 0005, Steven Lake Waslander
IROS2
2015 Modelling a Quadrotor Vehicle Using a Modular Deep Recurrent Neural Network
abstract
In this paper, the Modular Deep Recurrent Neural Network (MODERNN) framework is studied for learning a Multi-Input-Multi-Output (MIMO) model of a quad rotor. Comparing a Single-Input-Single-Output (SISO) system, a MIMO system is much harder to model because of the intercoupling of the system variables as well as the multi-dimensionality of the input and output spaces. In this paper it is shown that the MODERNN framework is capable of modelling complex MIMO dynamical mappings, such as a simulated MIMO model (4-by-4) of a quad rotor vehicle in the presence of noise and ground effect.
Nima Mohajerin, Steven Lake Waslander
SMC2
2015 Planning Paths for Package Delivery in Heterogeneous Multirobot Teams
abstract
This paper addresses the task scheduling and path planning problem for a team of cooperating vehicles performing autonomous deliveries in urban environments. The cooperating team comprises two vehicles with complementary capabilities, a truck restricted to travel along a street network, and a quadrotor micro-aerial vehicle of capacity one that can be deployed from the truck to perform deliveries. The problem is formulated as an optimal path planning problem on a graph and the goal is to find the shortest cooperative route enabling the quadrotor to deliver items at all requested locations. The problem is shown to be NP-hard. A solution is then proposed using a novel reduction to the Generalized Traveling Salesman Problem, for which well-established heuristic solvers exist. The heterogeneous delivery problem contains as a special case the problem of scheduling deliveries from multiple static warehouses. We propose two additional algorithms, based on enumeration and a reduction to the traveling salesman problem, for this special case. Simulation results compare the performance of the presented algorithms and demonstrate examples of delivery route computations over real urban street maps.
Neil Mathew, Stephen L. Smith 0001, Steven Lake Waslander
IEEE Trans Autom. Sci. Eng.3
2015 Path Following Using Dynamic Transverse Feedback Linearization for Car-Like Robots
abstract
This paper presents an approach for designing path-following controllers for the kinematic model of car-like mobile robots using transverse feedback linearization with dynamic extension. This approach is applicable to a large class of paths and its effectiveness is experimentally demonstrated on a Chameleon R100 Ackermann steering robot. Transverse feedback linearization makes the desired path attractive and invariant, while the dynamic extension allows the closed-loop system to achieve the desired motion along the path.
Adeel Akhtar, Christopher Nielsen, Steven Lake Waslander
IEEE Trans. Robotics3
2015 Multirobot Rendezvous Planning for Recharging in Persistent Tasks
abstract
This paper addresses a multirobot scheduling problem in which autonomous unmanned aerial vehicles (UAVs) must be recharged during a long-term mission. The proposal is to introduce a separate team of dedicated charging robots that the UAVs can dock with in order to recharge. The goal is to schedule and plan minimum cost paths for charging robots such that they rendezvous with and replenish the UAVs, as needed, during the mission. The approach is to discretize the 3-D UAV flight trajectories into sets of projected charging points on the ground, thus allowing the problem to be abstracted onto a partitioned graph. Solutions consist of charging robot paths that collectively charge each of the UAVs. The problem is solved by first formulating the rendezvous planning problem to recharge each UAV once using both an integer linear program and a transformation to the Travelling Salesman Problem. The methods are then leveraged to plan recurring rendezvous' over longer horizons using fixed horizon and receding horizon strategies. Simulation results using realistic vehicle and battery models demonstrate the feasibility and robustness of the proposed approach.
Neil Mathew, Stephen L. Smith 0001, Steven Lake Waslander
IEEE Trans. Robotics3
2014 Outlier rejection for visual odometry using parity space methods
abstract
Typically, random sample consensus (RANSAC) approaches are used to perform outlier rejection for visual odometry, however the use of RANSAC can be computationally expensive. The parity space approach (PSA) provides methodology to perform computationally efficient consistency checks for observations, without having to explicitly compute the system state. This work presents two outlier based rejection techniques, Group Parity Outlier Rejection and Parity Initialized RANSAC, which use the parity space approach to perform rapid outlier rejection. Experiments demonstrate the proposed approaches are able to compute solutions with increased accuracy and improved run-time when compared to RANSAC.
Arun Das 0005, Steven Lake Waslander
ICRA2
2014 Multi channel generalized-ICP
abstract
Current state of the art scan registration algorithms which use only position information often fall victim to correspondence ambiguity and degeneracy in the optimization solutions. Other methods which use additional channels, such as color or intensity, often use only a small fraction of the available information and ignore the underlying structural information of the added channels. The proposed method incorporates the additional channels directly into the scan registration formulation to provide information within the plane of the surface. This is achieved by calculating the uncertainty both along and perpendicular to the local surface at each point and calculating nearest neighbour correspondences in the higher dimensional space. The proposed method reduces instances of degenerate transformation estimates and improves both registration accuracy and convergence rate. The method is tested on the Ford Vision and Lidar dataset using both color and intensity channels as well as on Microsoft Kinect data obtained from the University of Waterloo campus.
James Servos, Steven Lake Waslander
ICRA2
2014 Model-aided state estimation for quadrotor micro air vehicles amidst wind disturbances
abstract
This paper extends the recently developed Model-Aided Visual-Inertial Fusion (MA-VIF) technique for quadrotor Micro Air Vehicles (MAV) to deal with wind disturbances. The wind effects are explicitly modelled in the quadrotor dynamic equations excluding the unobservable wind velocity component. This is achieved by a nonlinear observability of the dynamic system with wind effects. We show that using the developed model, the vehicle pose and two components of the wind velocity vector can be simultaneously estimated with a monocular camera and an inertial measurement unit. We also show that the MA-VIF is reasonably tolerant to wind disturbances, even without explicit modelling of wind effects and explain the reasons for this behaviour. Experimental results using a Vicon motion capture system are presented to demonstrate the effectiveness of the proposed method and validate our claims.
Dinuka M. W. Abeywardena, Gamini Dissanayake, Steven Lake Waslander, Sarath Kodagoda
IROS4
2014 MPC based collaborative adaptive cruise control with rear end collision avoidance
abstract
This paper presents a model predictive control (MPC) based approach to improve a recently developed class of collaborative adaptive cruise control (CACC) schemes. The PID structure used previously is replaced with MPC, which is able to accommodate actuator limits and parameter estimation. In addition to the regular CACC functionalities, rear end collision control is also incorporated. This approach is able to avoid rear end collisions with the following car, as long as it can still maintain the safe distance with the preceding vehicle. Simulation results are presented which demonstrate the validity of the approach.
Feyyaz Emre Sancar, Baris Fidan, Jan Paul Huissoon, Steven Lake Waslander
Intelligent Vehicles Symposium4
2014 Pneumatic trail based slip angle observer with Dugoff tire model
abstract
Autonomous driving requires reliable and accurate vehicle control at the limits of tire performance, which is only possible if accurate slip angle estimates are available. Recent methods have demonstrated the value of pneumatic trail for estimating slip angle in the non-linear region using the Fiala tire model. We present an improved slip angle estimation method based on the pneumatic trail method, which incorporates both lateral and longitudinal acceleration effects through the use of the Dugoff tire model. The proposed method offers significant improvements over existing methods, where longitudinal effects of the road-tire were assumed negligible. The results are demonstrated using CarSim, which relies on empirical data models for tire modelling and therefore presents a useful evaluation of the method.
Sirui Song, Michael Chi Kam Chun, Jan Paul Huissoon, Steven Lake Waslander
Intelligent Vehicles Symposium4
2014 Modular deep Recurrent Neural Network: Application to quadrotors
abstract
A modular deep Recurrent Neural Network (RNN) is introduced to facilitate the process of deploying various architectures of RNNs, and to automatically compute derivatives for gradient-based learning methods. The modularity leads to a set of new architectures, one of which includes feedforward inter-layer connections. By adding feedforward inter-layer connections in a multi-layer RNN, it is observed that the capability of the RNN to learn and model high-order dynamics and nonlinearities is significantly improved. The problem of vanishing/exploding gradient in space for a multilayer RNN is also alleviated using feedforward connections. These results are demonstrated using a quadrotor case study, for which a model of the altitude dynamics is learned with our particular network structure, while existing methods are unable to generalize as quickly or at all.
Nima Mohajerin, Steven Lake Waslander
SMC2
2014 Optimal Path Planning in Cooperative Heterogeneous Multi-robot Delivery Systems
Neil Mathew, Stephen L. Smith 0001, Steven Lake Waslander
WAFR3
2014 Scale recovery in multicamera cluster SLAM with non-overlapping fields of view
Michael J. Tribou, Steven Lake Waslander, David Wang 0001
Comput. Vis. Image Underst.2
2013 3D scan registration using the Normal Distributions Transform with ground segmentation and point cloud clustering
abstract
The Normal Distributions Transform (NDT) scan registration algorithm models the environment as a set of Gaussian distributions and generates the Gaussians by discretizing the environment into voxels. With the standard approach, the NDT algorithm has a tendency to have poor convergence performance for even modest initial transformation error. In this work, a segmented greedy cluster NDT (SGC-NDT) variant is proposed, which uses natural features in the environment to generate Gaussian clusters for the NDT algorithm. By segmenting the ground plane and clustering the remaining features, the SGC-NDT approach results in a smooth and continuous cost function which guarantees that the optimization will converge. Experiments show that the SGC-NDT algorithm results in scan registrations with higher accuracy and better convergence properties when compared against other state-of-the- art methods for both urban and forested environments.
Arun Das 0005, James Servos, Steven Lake Waslander
ICRA3
2013 A graph-based approach to multi-robot rendezvous for recharging in persistent tasks
abstract
This paper addresses the problem of maintaining persistence in coordinated tasks performed by a team of autonomous robots. We introduce a dedicated team of charging robots to service a team of primary working robots. Given that the trajectories of the working robots are known within a planning interval, the objective is to plan routes for the charging robots such that they rendezvous with and recharge all working robots to guarantee their continuous operation. To this end, the working robot trajectories are discretized to form a finite set of recharging points at which rendezvous can occur. The problem is formulated as a directed acyclic graph with vertex partitions containing sets of charging points for each working robot. Solutions consist of paths through the graph for each of the charging robots. The problem is shown to be NP-hard and a mixed integer linear program formulation is presented and solved for small problem instances. Finally, it is shown that while the optimal solution is not computationally feasible for large problem sizes, it is possible to graphically transform the single charging robot problem to a Traveling Salesman Problem, for which existing heuristic and approximation algorithms can be applied. Simulation results are presented for both single and multiple charging robot scenarios.
Neil Mathew, Stephen L. Smith 0001, Steven Lake Waslander
ICRA3
2013 Underwater stereo SLAM with refraction correction
abstract
This work presents a method for underwater stereo localization and mapping for detailed inspection tasks. The method generates dense, geometrically accurate reconstructions of underwater environments by compensating for image distortions due to refraction. A refractive model of the camera and enclosure is calculated offline using calibration images and produces non-linear epipolar curves for use in stereo matching. An efficient block matching algorithm traverses the precalculated epipolar curves to find pixel correspondences and depths are calculated using pixel ray tracing. Finally the depth maps are used to perform dense simultaneous localization and mapping to generate a 3D model of the environment. The localization and mapping algorithm incorporates refraction corrected ray tracing to improve map quality. The method is shown to improve overall depth map quality over existing methods and to generate high quality 3-D reconstructions.
James Servos, Michael Smart, Steven Lake Waslander
IROS3
2012 A nonlinear path following controller for an underactuated unmanned surface vessel
abstract
This work presents a novel path following controller for underactuated unmanned surface vessels (USVs) that is both provably stable and intuitive to tune. The approach consists of a navigation component that computes a desired heading angle to ensure the USV will arrive at the path, and a nonlinear controller to guarantee exponential tracking of surge velocity and heading. Additionally, ultimate boundedness of the unactuated sway velocity is proven. Simulation results are presented to show numerically that the controller works as expected in the ideal case. Outdoor experimental results are presented, using a GPS and compass as sensors, showing the practical feasibility of the approach in the presence of sensor noise, disturbances, and unmodeled dynamics.
John Michael Daly, Michael J. Tribou, Steven Lake Waslander
IROS3
2012 Scan registration with multi-scale k-means normal distributions transform
abstract
The normal distributions transform (NDT) scan registration algorithm has been shown to produce good results, however, has a tendency to converge to a local minimum if the initial parameter error is large. In order to improve the convergence basin for NDT, a multi-scale k-means NDT (MSKM-NDT) variant is proposed. This approach divides the point cloud using k-means clustering and performs the optimization step at multiple scales of cluster sizes. The k-means clustering approach guarantees that the optimization will converge, as it resolves the issue of discontinuities in the cost function found in the standard NDT algorithm. The optimization step of the NDT algorithm is performed over a decreasing scale, which greatly improves the basin of convergence. Experiments show that this approach can be used to register partially overlapping scans with large initial transformation error.
Arun Das 0005, Steven Lake Waslander
IROS2
2011 Coordinated landing of a quadrotor on a skid-steered ground vehicle in the presence of time delays
abstract
This work presents a control technique to autonomously coordinate a landing between a quadrotor UAV and a skid-steered UGV. Local controllers to feedback linearize the models are presented, and a joint decentralized controller is developed to coordinate a rendezvous for the two vehicles. The effects of time delays on closed loop stability are examined using a Retarded Functional Differential Equation (RFDE) formulation of the problem, and delay margins are determined for particular closed loop setups. Simulation results are presented, which demonstrate the feasibility of this approach for autonomous outdoor coordinated landing.
John Michael Daly, Steven Lake Waslander
IROS3
2009 Aerodynamics and control of autonomous quadrotor helicopters in aggressive maneuvering
abstract
Quadrotor helicopters have become increasingly important in recent years as platforms for both research and commercial unmanned aerial vehicle applications. This paper extends previous work on several important aerodynamic effects impacting quadrotor flight in regimes beyond nominal hover conditions. The implications of these effects on quadrotor performance are investigated and control techniques are presented that compensate for them accordingly. The analysis and control systems are validated on the Stanford Testbed of Autonomous Rotorcraft for Multi-Agent Control quadrotor helicopter testbed by performing the quadrotor equivalent of the stall turn aerobatic maneuver. Flight results demonstrate the accuracy of the aerodynamic models and improved control performance with the proposed control schemes.
Haomiao Huang, Gabriel M. Hoffmann, Steven Lake Waslander, Claire J. Tomlin
ICRA3
2009 Stanford Testbed of Autonomous Rotorcraft for Multi-Agent Control
abstract
The Stanford Testbed of Autonomous Rotorcraft for Multi-Agent Control, a fleet of quadrotor helicopters, has been developed as a testbed for novel algorithms that enable autonomous operation of aerial vehicles. The testbed has been used to validate multiple algorithms such as reactive collision avoidance, collision avoidance through Nash Bargaining, path planning, cooperative search and aggressive maneuvering. This article briefly describes the algorithms presented and provides references for a more in-depth formulation, and the accompanying movie shows the demonstration of the algorithms on the testbed.
Gabriel M. Hoffmann, Steven Lake Waslander, Michael P. Vitus, Haomiao Huang, Jeremy H. Gillula, Vijay Pradeep, Claire J. Tomlin
IROS2
2008 Lump-Sum Markets for Air Traffic Flow Control With Competitive Airlines
abstract
Air traffic flow control during adverse weather conditions is managed by the Federal Aviation Administration in today's air traffic system, although it is the individual airlines that are in the best position to assess the costs of disruptions to scheduled operations. To improve the efficiency of resource allocation, a market mechanism is proposed that enables airlines to participate directly in the flow control decision-making process. Since airlines can be expected to behave strategically, a lump-sum market mechanism is used for which existence of a Nash equilibrium and a bound on the worst case efficiency loss have been shown for agents that anticipate the effects of their own bids on resource prices. The convergence properties of this mechanism are studied for a two-player game with linear utilities, which reveals that restricting the airline bid update step-size can result in a wider range of stable bidding processes. The mechanism is then applied to an air traffic flow control scenario for multiple airports in the northeastern United States, which demonstrates the feasibility of performing market-based resource allocation within the time horizon for reliable weather predictions.
Steven Lake Waslander, Kaushik Roy 0007, Ramesh Johari, Claire J. Tomlin
Proc. IEEE1
2005 Multi-agent quadrotor testbed control design: integral sliding mode vs. reinforcement learning
abstract
The Stanford Testbed of Autonomous Rotorcraft for Multi-Agent Control (STARMAC) is a multi-vehicle testbed currently comprised of two quadrotors, also called X4-flyers, with capacity for eight. This paper presents a comparison of control design techniques, specifically for outdoor altitude control, in and above ground effect, that accommodate the unique dynamics of the aircraft. Due to the complex airflow induced by the four interacting rotors, classical linear techniques failed to provide sufficient stability. Integral sliding mode and reinforcement learning control are presented as two design techniques for accommodating the nonlinear disturbances. The methods both result in greatly improved performance over classical control techniques.
Steven Lake Waslander, Gabriel M. Hoffmann, Jung Soon Jang, Claire J. Tomlin
IROS1