Stewart Worrall 0002

dblp:33/555-2 · DBLP profile ↗
← Back
45ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0001-7940-4742ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 3 first-author · 10 since 2021Systems, architecture and hardware · 10 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 What demands attention in urban street scenes? From scene understanding towards road safety: A survey of vision-driven datasets and studies
abstract
Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To facilitate the use of these improvements for road safety, this survey systematically categorizes the critical elements that demand attention in traffic scenarios and comprehensively analyzes available vision-driven tasks and datasets. Compared to existing surveys that focus on isolated domains, our taxonomy categorizes attention-worthy traffic entities into two main groups, namely anomalies (abnormal entities) and pertinent entities (normal but critical elements), integrating eleven categories and twenty-three subclasses. It establishes connections between inherently related fields and provides a unified analytical framework. Based on the proposed taxonomy, our survey highlights the analysis of 40 vision-driven tasks and the comprehensive examinations and visualizations of 78 available datasets, including their basic characteristics, sensor settings, label design, visualization practices, annotation schemas, and the resulting implications. The cross-domain investigation reveals substantial variations in benchmark quality across tasks, with recurring limitations including uneven task coverage, imbalanced distributions, inconsistent or insufficient annotations, and limited multimodal and cross-task support. Our article further outlines promising solutions from the perspectives of task formulation, benchmark evaluations, dataset adoption and future dataset construction. The integrated taxonomy, comprehensive analysis, and recapitulatory tables provide researchers with a holistic overview of this rapidly evolving field, guiding strategic resource selection, and highlighting critical yet underexplored areas.
Yaoqi Huang, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002
Eng. Appl. Artif. Intell.4
2025 Animal Interaction with Autonomous Mobility Systems: Designing for Multi-Species Coexistence
abstract
Autonomous mobility systems increasingly operate in environments shared with animals, from urban pets to wildlife.However, their design has largely focused on human interaction, with limited understanding of how non-human species perceive, respond to, or are affected by these systems.Motivated by research in Animal-Computer Interaction (ACI) and more-than-human design, this study investigates animal interactions with autonomous mobility through a multi-method approach combining a scoping review (45 articles), online ethnography (39 YouTube videos and 11 Reddit discussions), and expert interviews (8 participants).Our analysis surfaces five key areas of concern: Physical Impact (e.g., collisions, failures to detect), Behavioural Effects (e.g., avoidance, stress), Accessibility Concerns (particularly for service animals), Ethics and Regulations, and Urban Disturbance.We conclude with design and policy directions aimed at supporting multispecies coexistence in the age of autonomous systems.This work underscores the importance of incorporating non-human perspectives to ensure safer, more inclusive futures for all species.
Tram Thi Minh Tran, Xinyan Yu 0005, Marius Hoggenmüller, Callum Parker, Paul Schmitt, Julie Stephany Berrio, Stewart Worrall 0002, Martin Tomitsch
AutomotiveUI7
2025 Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
abstract
To operate safely, autonomous vehicles (AVs) need to detect and handle unexpected objects or anomalies on the road. While significant research exists for anomaly detection and segmentation in 2D, research progress in 3D is underexplored. Existing datasets lack high-quality multimodal data that are typically found in AVs. This paper presents a novel dataset for anomaly segmentation in driving scenarios. To the best of our knowledge, it is the first publicly available dataset focused on road anomaly segmentation with dense 3D semantic labeling, incorporating both LiDAR and camera data, as well as sequential information to enable anomaly detection across various ranges. This capability is critical for the safe navigation of autonomous vehicles. We adapted and evaluated several baseline models for 3D segmentation, highlighting the challenges of 3D anomaly detection in driving environments. Our dataset and evaluation code will be openly available, facilitating the testing and performance comparison of different approaches.
Alexey Nekrasov 0001, Malcolm Burdorf, Stewart Worrall 0002, Bastian Leibe, Julie Stephany Berrio
CVPR3
2025 Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X Collaboration
abstract
Vehicle-to-everything (V2X) collaborative perception has emerged as a promising solution to address the limitations of single-vehicle perception systems. However, existing V2X datasets are limited in scope, diversity, and quality. To address these gaps, we present Mixed Signals, a comprehensive V2X dataset featuring 45.1k point clouds and 240.6k bounding boxes collected from three connected autonomous vehicles (CAVs) equipped with two different configurations of LiDAR sensors, plus a roadside unit with dual LiDARs. Our dataset provides point clouds and bounding box annotations across 10 classes, ensuring reliable data for perception training. We provide detailed statistical analysis on the quality of our dataset and extensively benchmark existing V2X methods on it. The Mixed Signals dataset is ready-to-use, with precise alignment and consistent annotations across time and viewpoints. Dataset website is available at https://mixedsignalsdataset.cs.cornell.edu/.
Katie Luo, Minh-Quan Dao, Mark E. Campbell, Wei-Lun Chao, Kilian Q. Weinberger, Ezio Malis, Vincent Frémont, Bharath Hariharan, Mao Shan, Stewart Worrall 0002, Julie Stephany Berrio
ICCV11
2024 InverseMatrixVT3D: An Efficient Projection Matrix-Based Approach for 3D Occupancy Prediction
abstract
This paper introduces InverseMatrixVT3D, an efficient method for transforming multi-view image features into 3D feature volumes for 3D semantic occupancy prediction. Existing methods for constructing 3D volumes often rely on depth estimation, device-specific operators, or transformer queries, which hinders the widespread adoption of 3D occupancy models. In contrast, our approach leverages two projection matrices to store the static mapping relationships and matrix multiplications to efficiently generate global Bird’s Eye View (BEV) features and local 3D feature volumes. Specifically, we achieve this by performing matrix multiplications between multi-view image feature maps and two sparse projection matrices. We introduce a sparse matrix handling technique for the projection matrices to optimize GPU memory usage. Moreover, a global-local attention fusion module is proposed to integrate the global BEV features with the local 3D feature volumes to obtain the final 3D volume. We also employ a multi-scale supervision mechanism to enhance performance further. Extensive experiments performed on the nuScenes and SemanticKITTI datasets reveal that our approach not only stands out for its simplicity and effectiveness but also achieves the top performance in detecting vulnerable road users (VRU), crucial for autonomous driving and road safety. The code has been made available at: https://github.com/DanielMing123/InverseMatrixVT3D
Zhenxing Ming, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002
IROS4
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions (i) which information should be exchanged over the V2X network, and (ii) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code is released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IV6
2024 Label-Efficient 3D Object Detection For Road-Side Units
abstract
Occlusion presents a significant challenge for safety-critical applications such as autonomous driving. Collaborative perception has recently attracted a large research interest thanks to the ability to enhance the perception of autonomous vehicles via deep information fusion with intelligent roadside units (RSU), thus minimizing the impact of occlusion. While significant advancement has been made, the data-hungry nature of these methods creates a major hurdle for their realworld deployment, particularly due to the need for annotated RSU data. Manually annotating the vast amount of RSU data required for training is prohibitively expensive, given the sheer number of intersections and the effort involved in annotating point clouds. We address this challenge by devising a label-efficient object detection method for RSU based on unsupervised object discovery. Our paper introduces two new modules: one for object discovery based on a spatial temporal aggregation of point clouds, and another for refinement. Furthermore, we demonstrate that fine-tuning on a small portion of annotated data allows our object discovery models to narrow the performance gap with, or even surpass, fully supervised models. Extensive experiments are carried out in simulated and real-world datasets to evaluate our method†.
Minh-Quan Dao, Holger Caesar, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Vincent Frémont, Ezio Malis
IV5
2024 Safety Driver Attention on Autonomous Vehicle Operation Based on Head Pose and Vehicle Perception
abstract
Despite the continual advances in Advanced Driver Assistance Systems (ADAS) and the development of high-level autonomous vehicles (AV), there is a consensus that for the short to medium term, there is a requirement for a human supervisor to handle the edge cases that inevitably arise. Given this requirement, the state of the autonomous vehicle operator (referred to as the safety driver) must be monitored to ensure their contribution to the vehicle's safe operation. This paper introduces a dual-source approach integrating data from an infrared camera facing the safety driver and vehicle perception systems to produce a metric for safety driver alertness to promote and ensure safe operator behaviour. The infrared camera detects the safety driver’s head, enabling the calculation of head orientation, which is relevant as the head typically moves according to the individual's focus of attention. By incorporating environmental data from the perception system, it becomes possible to determine whether the safety driver observes objects in the surroundings. Experiments were conducted using data collected in Sydney, Australia, simulating AV operations in an urban environment. Our results demonstrate that the proposed system effectively determines a metric for the attention levels of the safety driver, enabling interventions such as warnings or reducing autonomous functionality as appropriate. The results indicate reduced awareness on subsequent laps during the study, demonstrating the "automation complacency" phenomenon. This comprehensive solution shows promise in contributing to ADAS and AVs’ overall safety and efficiency in a real-world setting.
Santiago Gerling Konrad, Julie Stephany Berrio, Mao Shan, Favio R. Masson, Eduardo M. Nebot, Stewart Worrall 0002
IV6
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions 1) which information should be exchanged over the V2X network and 2) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code will be released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IEEE Trans. Intell. Transp. Syst.6
2023 Viewer-Centred Surface Completion for Unsupervised Domain Adaptation in 3D Object Detection
abstract
Every autonomous driving dataset has a different configuration of sensors, originating from distinct geographic regions and covering various scenarios. As a result, 3D detectors tend to overfit the datasets they are trained on. This causes a drastic decrease in accuracy when the detectors are trained on one dataset and tested on another. We observe that lidar scan pattern differences form a large component of this reduction in performance. We address this in our approach, SEE-VCN, by designing a novel viewer-centred surface completion network (VCN) to complete the surfaces of objects of interest within an unsupervised domain adaptation framework, SEE [1]. With SEE-VCN, we obtain a unified representation of objects across datasets, allowing the network to focus on learning geometry, rather than overfitting on scan patterns. By adopting a domain-invariant representation, SEE-VCN can be classed as a multi-target domain adaptation approach where no annotations or re-training is required to obtain 3D detections for new scan patterns. Through extensive experiments, we show that our approach outperforms previous domain adaptation methods in multiple domain adaptation settings. Our code and data are available at https://github.com/darrenjkt/SEE-VCN.
Darren Tsai, Julie Stephany Berrio, Mao Shan, Eduardo M. Nebot, Stewart Worrall 0002
ICRA5
2023 My Eyes Speak: Improving Perceived Sociability of Autonomous Vehicles in Shared Spaces Through Emotional Robotic Eyes
abstract
The ability of autonomous vehicles (AVs) to interact socially with pedestrians poses a significant impact on their integration with urban traffic. This is particularly important for vehicle-pedestrian shared spaces due to increased social requirements in comparison to vehicular roads. Current pedestrian experience in shared spaces suffers from negative attitudes towards AVs and the consequently low acceptability of AVs in these spaces. HRI work shows that the acceptability of robots in public spaces can be positively impacted by their perceived sociability (i.e., possessing social skills), which can be enhanced by their ability to express emotions. Inspired by this approach, we follow a systematic process to design emotional expressions for AVs using the headlight ("eye'') area and investigate their impact on perceived sociability of AVs in shared spaces, by conducting expert focus groups (N=12) and an online video-based user study (N=106). Our findings confirm that the perceived sociability of AVs can be enhanced by emotional expressions indicated through emotional eyes. We further discuss implications of our findings for improving pedestrian experience and attitude in shared spaces and highlight opportunities to use AVs' emotional expressions as a new external communication strategy for future research.
Yiyuan Wang 0001, Senuri Wijenayake, Marius Hoggenmüller, Luke Hespanhol, Stewart Worrall 0002, Martin Tomitsch
Proc. ACM Hum. Comput. Interact.5
2022 Towards Collision-Free Probabilistic Pedestrian Motion Prediction for Autonomous Vehicles
abstract
Autonomous vehicle navigation in shared pedestrian environments requires the ability to predict future crowd motion as well as understand human behaviour. However, most existing methods predict pedestrian future motion without considering potential collisions within the crowd. Furthermore, most current predictive models are tested on datasets that assume full observability of the crowd by relying on a top-down view, which does not reflect the real-world use case of autonomous vehicles due to the inherent limitations of on-board sensors such as visual occlusion. Inspired by prior works, we propose a pedestrian motion prediction model trained via contrastive learning, improving prediction accuracy as well as forecasting collision-free trajectories. Additionally, we propose a method for implementing a predictor using a multi-pedestrian probabilistic tracker, which fuses multiple on-board sensors to track pedestrians in 3D space. Through comprehensive experiments on both aerial view and driving datasets collected in a real-world urban environment, we show that our proposed method improves on state of art methods with better prediction accuracy and more socially acceptable prediction trajectories.
Kunming Li, Mao Shan, Stuart Eiffert, Stewart Worrall 0002, Eduardo M. Nebot
IV4
2022 Camera-LIDAR Integration: Probabilistic Sensor Fusion for Semantic Mapping
abstract
An automated vehicle operating in an urban environment must be able to perceive and recognise objects and obstacles in a three-dimensional world for navigation and path planning. In order to plan and execute accurate and sophisticated driving maneuvers, a high-level contextual understanding of the surroundings is essential. Due to the recent progress in image processing, it is now possible to obtain high definition semantic information in 2D from monocular cameras, though cameras cannot reliably provide the high accuracy 3D information provided by lasers. The fusion of these two sensor modalities can overcome the shortcomings of each individual sensor, though there are a number of important challenges that need to be addressed in a probabilistic manner. In this paper we address the common, yet challenging, LIDAR/camera/semantic fusion problems which are seldom approached in a wholly probabilistic manner. Our approach is capable of using a multi-sensor platform to build a three-dimensional semantic voxelized map that considers the uncertainty of all of the processes involved. We present a probabilistic pipeline that incorporates uncertainty from the sensor readings (cameras, LIDAR, IMU and wheel encoders), compensation for the motion of the vehicle, and heuristic label probabilities for the semantic images depicted inFig. 1. We also present a novel and efficient viewpoint validation algorithm to check for occlusions within the camera frame. A probabilistic projection is performed from the camera images to the LIDAR point cloud. Each labelled LIDAR scan then feeds into an octree map-building algorithm that updates the class probabilities of the map voxels every time a new observation is available. We validate our approach using a set of qualitative and quantitative experiments using the USyd Campus Dataset.1These tests demonstrate the usefulness of a probabilistic sensor fusion approach by evaluating the performance of the perception system in a typical autonomous vehicle application.
Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.3
2022 Long-Term Map Maintenance Pipeline for Autonomous Vehicles
abstract
For autonomous vehicles to operate persistently in a typical urban environment, it is essential to have high accuracy position information. This requires a mapping and localisation system that can adapt to changes over time. A localisation approach based on a single-survey map will not be suitable for long-term operation as it does not incorporate variations in the environment. In this paper, we present new algorithms to maintain a featured-based map. A map maintenance pipeline is proposed that can continuously update a map with the most relevant features taking advantage of the changes in the surroundings. Our pipeline detects and removes transient features based on their geometrical relationships with the vehicle’s pose. Newly identified features became part of a new feature map and are assessed by the pipeline as candidates for the localisation map. By purging out-of-date features and adding newly detected features, we continually update the prior map to more accurately represent the most recent environment. We have validated our approach using the USyd Campus Dataset, which includes more than 18 months of data. The results presented demonstrate that our maintenance pipeline produces a resilient map which can provide sustained localisation performance over time.
Julie Stephany Berrio, Stewart Worrall 0002, Mao Shan, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.2
2021 Context-Based Interface Prototyping: Understanding the Effect of Prototype Representation on User Feedback
abstract
The rise of autonomous systems in cities, such as automated vehicles (AVs), requires new approaches for prototyping and evaluating how people interact with those systems through context-based user interfaces, such as external human-machine interfaces (eHMIs). In this paper, we present a comparative study of three prototype representations (real-world VR, computer-generated VR, real-world video) of an eHMI in a mixed-methods study with 42 participants. Quantitative results show that while the real-world VR representation results in higher sense of presence, no significant differences in user experience and trust towards the AV itself were found. However, interview data shows that participants focused on different experiential and perceptual aspects in each of the prototype representations. These differences are linked to spatial awareness and perceived realism of the AV behaviour and its context, affecting in turn how participants assess trust and the eHMI. The paper offers guidelines for prototyping and evaluating context-based interfaces through simulations.
Marius Hoggenmüller, Martin Tomitsch, Luke Hespanhol, Tram Thi Minh Tran, Stewart Worrall 0002, Eduardo M. Nebot
CHI5
2021 Attentional-GCNN: Adaptive Pedestrian Trajectory Prediction towards Generic Autonomous Vehicle Use Cases
abstract
Autonomous vehicle navigation in shared pedestrian environments requires the ability to predict future crowd motion both accurately and with minimal delay. Understanding the uncertainty of the prediction is also crucial. Most existing approaches however can only estimate uncertainty through repeated sampling of generative models. Additionally, most current predictive models are trained on datasets that assume complete observability of the crowd using an aerial view. These are generally not representative of real-world usage from a vehicle perspective, and can lead to the underestimation of uncertainty bounds when the on-board sensors are occluded. Inspired by prior work in motion prediction using spatio-temporal graphs, we propose a novel Graph Convolutional Neural Network (GCNN)-based approach, Attentional-GCNN, which aggregates information of implicit interaction between pedestrians in a crowd by assigning attention weight in edges of the graph. Our model can either output a probabilistic distribution or faster deterministic prediction, demonstrating applicability to autonomous vehicle use cases where either speed or accuracy with uncertainty bounds are required. To further improve the training of predictive models, we propose an automatically labelled pedestrian dataset collected from an intelligent vehicle platform representative of real-world use. Through experiments on a number of datasets, we show our proposed method achieves an improvement over the state of the art by 10% on Average Displacement Error (ADE) and 12% on Final Displacement Error (FDE) with fast inference speeds.
Kunming Li, Stuart Eiffert, Mao Shan, Francisco Gomez-Donoso, Stewart Worrall 0002, Eduardo M. Nebot
ICRA5
2020 Using a 3D CNN for Rejecting False Positives on Pedestrian Detection
abstract
Self-driving cars are becoming slowly but surely the future of transport. Nonetheless, in order to achieve fully automatic operation, several challenges are still needed to be tackled. One of the main goals that is currently being pursued is a very accurate scene understanding and object detection. In this regard, the most accurate object detectors are image-based. However, these methods yield critical flaws that make them prone to error in some specific scenarios. For instance, actual objects would be detected in depictions of such objects. The urban environment is strewn with these cases. Namely, in billboards and advertisements. However, most of the self-driving cars feature a lidar that provides 3D perception. This sensor could help to disambiguate the cases mentioned before.In this paper, we combine the accuracy of 2D deep learning object detectors with a 3D Convolutional Neural Network (3D CNN) for rejecting false positives on pedestrian detection. First, the object detector provides all the detected pedestrians in the scene, and then the 3D CNN is in charge of rejecting or verify the detections. Our proposal is tested on two well-known publicly available datasets and provides up to 84% accuracy.
Francisco Gomez-Donoso, Edmanuel Cruz, Miguel Cazorla, Stewart Worrall 0002, Eduardo M. Nebot
IJCNN4
2020 Automated Evaluation of Semantic Segmentation Robustness for Autonomous Driving
abstract
One of the fundamental challenges in the design of perception systems for autonomous vehicles is validating the performance of each algorithm under a comprehensive variety of operating conditions. In the case of vision-based semantic segmentation, there are known issues when encountering new scenarios that are sufficiently different to the training data. In addition, even small variations in environmental conditions, such as illumination and precipitation, can affect the classification performance of the segmentation model. Given the reliance on visual information, these effects often translate into poor semantic pixel classification which can potentially lead to catastrophic consequences when driving autonomously. This paper presents a novel method for analyzing the robustness of semantic segmentation models and provides a number of metrics to evaluate the classification performance over a variety of environmental conditions. The process incorporates an additional sensor (lidar) to automate the process and improve the system integrity, eliminating the need for labor-intensive hand labeling of validation data. The experimental results are presented based on multiple datasets collected at different times of the year with different environmental conditions. We extract the ”Road” class using the lidar to demonstrate the concepts, but this could be extended to other classes with different feature detection algorithms. These results show that the semantic segmentation performance varies depending on the weather, camera parameters, and existence of shadows. The results also demonstrate how the metrics can be used to compare and validate the performance after making improvements to a model, and compare the performance of different networks.
Wei Zhou 0025, Julie Stephany Berrio, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.3
2020 Naturalistic Driver Intention and Path Prediction Using Recurrent Neural Networks
abstract
Understanding the intentions of drivers at intersections is a critical component for autonomous vehicles. Urban intersections that do not have traffic signals are a common epicenter of highly variable vehicle movement and interactions. We present a method for predicting driver intent at urban intersections through multi-modal trajectory prediction with uncertainty. Our method is based on recurrent neural networks combined with a mixture density network output layer. To consolidate the multi-modal nature of the output probability distribution, we introduce a clustering algorithm that extracts the set of possible paths that exist in the prediction output and ranks them according to probability. To verify the method's performance and generalizability, we present a real-world dataset that consists of over 23 000 vehicles traversing five different intersections, collected using a vehicle-mounted lidar-based tracking system. An array of metrics is used to demonstrate the performance of the model against several baselines.
Alex Zyner, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.2
2019 Uncertainty Estimation for Projecting Lidar Points onto Camera Images for Moving Platforms
abstract
Combining multiple sensors for advanced perception is a crucial requirement for autonomous vehicle navigation. Heterogeneous sensors are used to obtain rich information about the surrounding environment. The combination of the camera and lidar sensors enables precise range information that can be projected onto the visual image data. This gives a high level understanding of the scene which can be used to enable context based algorithms such as collision avoidance and navigation. The main challenge when combining these sensors is aligning the data into a common domain. This can be difficult due to the errors in the intrinsic calibration of the camera, extrinsic calibration between the camera and the lidar and errors resulting from the motion of the platform. In this paper, we examine the algorithms required to provide motion correction for scanning lidar sensors. The error resulting from the projection of the lidar measurements into a consistent odometry frame is not possible to remove entirely, and as such it is essential to incorporate the uncertainty of this projection when combining the two different sensor frames. This work proposes a novel framework for the prediction of the uncertainty of lidar measurements (in 3D) projected in to the image frame (in 2D) for moving platforms. The proposed approach fuses the uncertainty of the motion correction with uncertainty resulting from errors in the extrinsic and intrinsic calibration. By incorporating the main components of the projection error, the uncertainty of the estimation process is better represented. Experimental results for our motion correction algorithm and the proposed extended uncertainty model are demonstrated using real-world data collected on an electric vehicle equipped with wide-angle cameras covering a 180-degree field of view and a 16-beam scanning lidar.
Charika De Alvis, Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
ICRA3
2019 Identifying robust landmarks in feature-based maps
abstract
To operate in an urban environment, an automated vehicle must be capable of accurately estimating its position within a global map reference frame. This is necessary for optimal path planning and safe navigation. To accomplish this over an extended period of time, the global map requires long term maintenance. This includes the addition of newly observable features and the removal of transient features belonging to dynamic objects. The latter is especially important for the long-term use of the map as matching against a map with features that no longer exist can result in incorrect data associations, and consequently erroneous localisation. This paper addresses the problem of removing features from the map that correspond to objects that are no longer observable/present in the environment. This is achieved by assigning a single score which depends on the geometric distribution and characteristics when the features are re-detected (or not) on different occasions. Our approach not only eliminates ephemeral features, but can also be used as a reduction algorithm for highly dense maps. We tested our approach using half a year of weekly drives over the same 500 metre section of road in an urban environment. The results presented demonstrate the validity of the long term approach to map maintenance.
Julie Stephany Berrio, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IV3
2019 Updating the visibility of a feature-based map for long-term maintenance
abstract
Mobile vehicles operating in urban navigation applications can achieve high integrity localisation with high accuracy by using maps of the surroundings. To accomplish this, the map should always have an accurate representation of the environment. Thus, it is necessary to detect and remove the map components that no longer exist in the current environment. This maintains the map compactness and dependability while simplifying the data association problem. This paper addresses the problem of deletion of transient map components by taking advantage of the geometric connection between the map and agent poses in order to establish and update the visibility of each feature. Once the map is created an initial visibility vector is associated with every map element and updated over time. The visibility of a map element which no longer exists is reduced and ultimately removed from the map. We demonstrate our approach in a 2D feature-based map composed of poles and corners extracted from information provided by a Iidar sensor. The experimental results show the map update using a seven-month data set collected in the University of Sydney campus.
Julie Stephany Berrio, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IV3
2019 Extended Vehicle Tracking with Probabilistic Spatial Relation Projection and Consideration of Shape Feature Uncertainties
abstract
This work focuses on a novel probabilistic approach for extended vehicle tracking, where multiple spatially distributed measurements can originate from the target, and kinematic state and geometry variables are estimated jointly. Prominent shape features extracted from raw measurement points contain spatial uncertainties due to noise in sensor measurements, the feature extraction process, approximation error of shape hypothesis, partial vision occlusion, to name a few. This work proposes a novel tracking paradigm that respects the variant spatial measurement model subject to changes in target pose and sensor viewpoint. This is achieved through probabilistic projection of the spatial measurement points to the predicted measurement sources on the visible side(s) of the target shape. The spatial uncertainties in the shape features are probabilistically modelled and incorporated in the unscented Kalman filter based estimation. The proposed approach is validated with field experiment results using cameras and a laser range scanner.
Mao Shan, Charika De Alvis, Stewart Worrall 0002, Eduardo M. Nebot
IV3
2019 Metrics for the Evaluation of localisation Robustness
abstract
Robustness and safety are crucial properties for the real-world application of autonomous vehicles. One of the most critical components of any autonomous system is localisation. During the last 20 years there has been significant progress in this area with the introduction of very efficient algorithms for mapping, localisation and SLAM. Many of these algorithms present impressive demonstrations for a particular domain, but fail to operate reliably with changes to the operating environment. The aspect of robustness has not received enough attention and localisation systems for self-driving vehicle applications are seldom evaluated for their robustness. In this paper we propose novel metrics to effectively quantify localisation robustness with or without an accurate ground truth. The experimental results present a comprehensive analysis of the application of these metrics against a number of well known localisation strategies.
Siqi Yi, Stewart Worrall 0002, Eduardo M. Nebot
IV2
2018 Automated Process for Incorporating Drivable Path into Real-Time Semantic Segmentation
abstract
Vision systems are widely used in autonomous vehicle systems due to the rich information that camera sensors provide of the surrounding environment. This paper presents an automatic algorithm to obtain the drivable path of a vehicle operating in urban roads with or without clear lane markings. The developed system projects trajectories obtained during human operation of the vehicle and utilizes these to generate automatic labels for training a semantic based path prediction model. The system segments an urban scenario into 13 categories including vehicles, pedestrian, undrivable road, other categories relevant to urban roads, and a new class for a path proposal. The drivable path information is essential particularly in unstructured scenarios, and is critical for an intelligent vehicle system to make sound driving decisions. The path proposal category is a car-width drivable lane estimated to be safe to drive for the vehicle under consideration. The data collection, model training and inference process requires only images from a monocular camera and odometry from a low-cost IMU combined with a wheel encoder. The algorithm has been successfully demonstrated on the Sydney University campus, which is a challenging environment without clear road markings. The algorithm was demonstrated to run in real-time, proving its applicability for intelligent vehicles.
Wei Zhou 0025, Stewart Worrall 0002, Alex Zyner, Eduardo M. Nebot
ICRA2
2018 Octree map based on sparse point cloud and heuristic probability distribution for labeled images
abstract
To navigate through urban roads, an automated vehicle must be able to perceive and recognize objects in a three-dimensional environment. A high level contextual understanding of the surroundings is necessary to execute accurate driving maneuvers. This paper presents a novel approach to build three dimensional semantic octree maps from lidar scans and the output of a convolutional neural network (CNN) to obtain the labels of the environment. We present a heuristic method to associate uncertainties to the labels from the images based on a combination of the labels themselves, score maps retrieved by the CNN and the raw images. These uncertainties and the camera-lidar calibration parameters for multiple cameras are considered in the projection of the labels and their uncertainties into the point cloud. Every labeled lidar scan works as an input to an octree map building algorithm that calculates and updates the label probabilities of the voxels in the map. This paper also presents a qualitative and quantitative evaluation of accuracy, analyzing projection in single lidar scans and complete maps built with our probabilistic octree framework.
Julie Stephany Berrio, Wei Zhou 0025, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IROS4
2018 Pedestrian Dynamic and Kinematic Information Obtained from Vision Sensors
abstract
The estimation and prediction of pedestrian motion is of fundamental importance in ITS applications. Most existing solutions have utilized a particular type of sensor for perception such as cameras (stereo, monocular, infrared) or other modalities such as a laser range finder or radar. The advent of wearable devices with inertial sensors have led to the development of systems capable of the robust inference of pedestrian intention. Unfortunately, these devices do not have communications capabilities to broadcast this information to all vehicles in proximity, and also this strategy requires functioning devices on all pedestrians to work. This paper presents a robust perception method that is able to extract dynamic pedestrian information with accuracy comparable to that of typical gyroscopes and accelerometers installed in wearable devices. Experimental results are presented to demonstrate the potential for obtaining very comprehensive dynamic information from limbs representing the skeleton of a pedestrian. This work also demonstrates the accuracy of vision based systems by comparing these results to the rotation and acceleration measured directly on the pedestrian using a wearable device. The contributions of this paper demonstrate that it is possible to significantly improve both the detection and estimation of pedestrian intention by incorporating dynamic information obtained from vision sensors.
Santiago Gerling Konrad, Mao Shan, Favio R. Masson, Stewart Worrall 0002, Eduardo M. Nebot
Intelligent Vehicles Symposium4
2017 Long short term memory for driver intent prediction
abstract
Advanced Driver Assistance Systems have been shown to greatly improve road safety. However, existing systems are typically reactive with an inability to understand complex traffic scenarios. We present a method to predict driver intention as the vehicle enters an intersection using a Long Short Term Memory (LSTM) based Recurrent Neural Network (RNN). The model is learnt using the position, heading and velocity fused from GPS, IMU and odometry data collected by the ego-vehicle. In this paper we focus on determining the earliest possible moment in which we can classify the driver's intention at an intersection. We consider the outcome of this work an essential component for all levels of road vehicle automation.
Alex Zyner, Stewart Worrall 0002, James R. Ward, Eduardo M. Nebot
Intelligent Vehicles Symposium2
2016 A Flexible System Architecture for Acquisition and Storage of Naturalistic Driving Data
abstract
Innovation in intelligent transportation systems relies on analysis of high-quality data. In this paper, we describe the design principles behind our data management infrastructure. The principles we adopt place an emphasis on flexibility and maintainability. This is achieved by breaking up code into a modular design that can be run on many independent processes. Message passing over a publish-subscribe network enables interprocess communication and promotes data-driven execution. By following these principles, rapid prototyping and experimentation with new sensing modalities and algorithms are possible. The communication library underpinning our proposed architecture is compared against several popular communication libraries. Features designed into the system make it decentralized, robust to failure, and amenable to scaling across multiple machines with minimal configuration. Code written using the proposed architecture is compact, transparent, and easy to maintain. Experimentation shows that our proposed architecture offers a high performance when compared against alternative communication libraries.
Asher Bender, James R. Ward, Stewart Worrall 0002, Marcelo L. Moreyra, Santiago Gerling Konrad, Favio R. Masson, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.3
2015 An Unsupervised Approach for Inferring Driver Behavior From Naturalistic Driving Data
abstract
Intelligent transportation systems are able to collect large volumes of high-resolution data. The amount of data collected by these systems can quickly overwhelm the ability of human analysts to draw meaningful conclusions from the data, particularly in large-scale multivehicle field trials. As advanced driver assistance systems develop, they will also be required to form a rich and high-level understanding of the world from the data they receive, including the behavior of the driver. These applications motivate the need for unsupervised tools capable of forming a high-level summary of low-level driving data. This paper presents an unsupervised method for converting naturalistic driving data into high-level behaviors. The proposed method works in two steps. In the first step, inertial data are automatically decomposed into linear segments. In the second step, the segments are assigned to high-level driving behaviors. The proposed method is computationally efficient and completely unsupervised and requires minimal preprocessing. Although the method is unsupervised, the clusters produced exhibit high-level patterns that can easily be associated with driving behaviors such as braking, turning, accelerating, and coasting. The effectiveness of the proposed algorithms is demonstrated in an offline application where the objective is to summarize inertial data into driving behaviors. The method is also demonstrated in an online application where the aim is to infer the current driving behavior using only inertial data. Both experiments were conducted using driving data collected in natural driving conditions.
Asher Bender, Gabriel Agamennoni, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.4
2015 Delayed-State Nonparametric Filtering in Cooperative Tracking
abstract
This paper presents a novel nonparametric approach toward delayed-state filtering for cooperative tracking. Standard parametric cooperative localization/tracking approaches are generally aimed at problems that can be easily parameterized and/or are limited to incorporate only real-time measurements. This paper provides a nonparametric yet computationally tractable alternative that is suitable for tracking cases where real-time observations are not always possible, e.g., in a sparse mesh network. The proposed delayed-state cooperative particle filter features forward filtering and backward smoothing to incorporate measurements that are received with time delays. A record of historical marginal states is kept for each mobile node within a sliding time window, instead of the high-dimensional joint state. Essentially, it replaces the importance sampling in traditional particle filters by a Gibbs sampler, which is a Markov chain Monte Carlo method, to fuse all available egocentric and internode relative observations into the global position estimate, thus alleviating the high-dimensionality problems in cooperative tracking. The performance of the proposed approach is evaluated in a multiagent simulation, and experimental results from a large-scale multivehicle industrial operation clearly demonstrate that the proposed approach effectively facilitates the tracking of mobile nodes without position awareness, through the use of relative range, negative detection, and time-delayed measurements.
Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Robotics2
2014 Bayesian model-based sequence segmentation for inferring primitives in driving-behavioral data
Gabriel Agamennoni, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
FUSION3
2014 Nonparametric cooperative tracking in mobile Ad-Hoc networks
abstract
This paper presents a new nonparametric approach for cooperative localisation and tracking by fusing relative range measurements between mobile network equipped nodes. Standard approaches based on parametric methods are known to be limited to problems that contain Gaussian properties. This paper overcomes this limitation by proposing a novel particle filter based cooperative tracking approach that is suitable for mobile ad-hoc networks (MANETs). The filter maintains the marginal state of every node instead of a joint state of the group, and updates the estimates using a Gibbs sampler, which is known as a Markov chain Monte Carlo (MCMC) method. The performance of the proposed algorithm is demonstrated by examining 16 mobile nodes moving randomly while sharing information with neighbouring nodes in a MANET. The results show that the proposed approach facilitates the tracking of mobile nodes that do not have egocentric position information available. It also outperforms the EKF in systems with non-Gaussian properties.
Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
ICRA2
2014 Vehicle collision probability calculation for general traffic scenarios under uncertainty
abstract
Vehicle-to-vehicle (V2V) communication systems allow vehicles to share state information with one another to improve safety and efficiency of transportation networks. One of the key applications of such a system is in the prediction and avoidance of collisions between vehicles. If a method to do this is to succeed it must be robust to measurement uncertainty. The method should also be general enough that it does not rely on constraints on vehicle motion for the accuracy of its predictions. It should work for all interactions between vehicles and not just a select subset. This paper presents a method for collision probability calculation that addresses these problems.
James R. Ward, Gabriel Agamennoni, Stewart Worrall 0002, Eduardo M. Nebot
Intelligent Vehicles Symposium3
2014 Using Delayed Observations for Long-Term Vehicle Tracking in Large Environments
abstract
The tracking of vehicles over large areas with limited position observations is of significant importance in many industrial applications. This paper presents algorithms for long-term vehicle motion estimation based on a vehicle motion model that incorporates the properties of the working environment and information collected by other mobile agents and fixed infrastructure collection points. The prediction algorithm provides long-term estimates of vehicle positions using speed and timing profiles built for a particular environment and considering the probability of a vehicle stopping. A limited number of data collection points distributed around the field are used to update the estimates, with negative information (no communication) also used to improve the prediction. This paper introduces the concept of observation harvesting, a process in which peer-to-peer communication between vehicles allows egocentric position updates to be relayed among vehicles and finally conveyed to the collection point for an improved position estimate. Positive and negative communication information is incorporated into the fusion stage, and a particle filter is used to incorporate the delayed observations harvested from vehicles in the field to improve the position estimates. The contributions of this work enable the optimization of fleet scheduling using discrete observations. Experimental results from a typical large-scale mining operation are presented to validate the algorithms.
Mao Shan, Stewart Worrall 0002, Favio R. Masson, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.2
2013 Anomaly detection in driving behaviour by road profiling
abstract
This paper presents a statistical method for detecting anomalous driving behavior by analyzing cross-sectional profiles of the road. A profile captures the way vehicles normally traverse the road; a statistical hypothesis test determines whether the observed behavior is anomalous. Experimental results on genuine data collected by a fleet of vehicles demonstrates the potential of this new method.
Gabriel Agamennoni, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
Intelligent Vehicles Symposium3
2013 Robust non-linear smoothing for vehicle state estimation
abstract
This paper presents a robust, non-linear smoothing algorithm and develops the theory behind it. This algorithm is extremely robust to outliers and missing data and handles state-dependent noise. Implementing it is straightforward as it consists mainly of two sub-routines: (a) the Rauch-Tung-Striebel recursions, or Kalman smoother; and (b) a backtracking line search strategy. The computational load grows linearly with the number of data because the algorithm preserves the underlying structure of the problem. Global convergence to a local optimum is guaranteed, under mild assumptions.
Gabriel Agamennoni, Stewart Worrall 0002, James R. Ward, Eduardo M. Nebot
Intelligent Vehicles Symposium2
2013 Towards mapping of dynamic environments with FMCW radar
abstract
Frequency-modulated continuous waveform (FMCW) microwave and millimetre-wave radar is an attractive sensor for intelligent transport systems due to its reliable all-weather performance. This paper discusses issues involved in the design of FMCW radar mapping systems for use in collision avoidance in large vehicles operating in dynamic environments. The performance characteristics of radar are examined before an analysis is made of traditional grid-based and feature-based mapping approaches, both conceptually and in terms of implementation. The probability hypothesis density (PHD) filter is discussed as a potentially superior approach for radar mapping in dynamic environments.
Bryan Clarke, Stewart Worrall 0002, Graham M. Brooker, Eduardo M. Nebot
Intelligent Vehicles Symposium2
2013 Vehicle operation safety monitoring using context based metrics: A case study
abstract
This paper presents results of the deployment of vehicle safety systems developed by the Intelligent Vehicles and Safety Systems Group at the Australian Centre for Field Robotics. The technology was deployed within an active mine site in Australia undergoing standard open-cut mining operations. Data were collected from the vehicles and analysed in order to assess the overall safety and performance of the vehicle operations within the mine. Metrics calculated include distributions of vehicle paths along stretches of road for driving line analysis, and statistics around vehicle events such as overspeed and proximity to other vehicles.
James R. Ward, Stewart Worrall 0002, Gabriel Agamennoni, Eduardo M. Nebot
Intelligent Vehicles Symposium2
2013 Fault detection for vehicular ad-hoc wireless networks
abstract
An increasing number of intelligent transportation applications require robust and reliable wireless communication. To achieve the required quality of service it is necessary to implement redundancy in the critical path which includes the radio software and hardware. In a real-world application there are many things that can cause the communication between two vehicles to degrade or stop completely. This paper describes a novel technique for detecting degradation or failure of communication links by comparing the performance of the radios to a probabilistic model built using data collected in the field. The results show that this techinique can successfully detect when there is partial or complete failure to communicate due to damage to the external components such as antennas, connectors and cables.
Stewart Worrall 0002, Gabriel Agamennoni, James R. Ward, Eduardo M. Nebot
Intelligent Vehicles Symposium1
2013 Probabilistic Long-Term Vehicle Motion Prediction and Tracking in Large Environments
abstract
Vehicle position tracking and prediction over large areas is of significant importance in many industrial applications, such as mining operations. In a small area, this can easily be achieved by providing vehicles with a constant communication link to a control center and having the vehicles broadcast their position. The problem dramatically changes when vehicles operate within a large environment of potentially hundreds of square kilometers and in difficult terrain. This paper presents algorithms for long-term vehicle motion prediction and tracking based on a multiple-model approach. It incorporates a probabilistic vehicle model that includes the structure of the environment. The prediction algorithm evaluates the vehicle position using acceleration, speed, and timing profiles built for the particular environment and considers the probability that the vehicle will stop. A limited number of data collection points distributed around the field are used to update the vehicle position estimate when in communication range, and prediction is used at points in between. A particle filter is used to estimate the vehicle position using both positive and negative information (whether communication is possible) in the fusion stage. The algorithms presented are validated with experimental results using data collected from a large-scale mining operation.
Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.2
2012 Sensor modelling for radar-based occupancy mapping
abstract
This paper addresses the issue of creating a sensor model for a new short-range 24GHz close proximity detection (CPD) radar. The CPD radar is designed to provide improved situational awareness to the driver of a large vehicle. It is able to detect light vehicles and other targets at ranges from 2.2m to 45m within an arc of 160° azimuth, but radar measurements contain data from noise and clutter which must be filtered out. Dynamic thresholds such as constant false-alarm rate (CFAR) processors [19] do not work well with short measurement vectors or targets that occupy multiple measurement bins. In this paper, a new method is used where measurements of environmental noise, clutter and targets are used to calculate the false alarm and target detection probabilities for each bin and develop a fixed detection threshold for each bin. This filter is used to construct a sensor model which maps measurement power to probability of bin occupancy, which is then used to generate an occupancy grid map of the environment from CPD radar measurements.
Bryan Clarke, Stewart Worrall 0002, Graham M. Brooker, Eduardo M. Nebot
IROS2
2012 Improving situational awareness with radar information
abstract
This paper addresses the issue of improving situational awareness for large vehicles operating in all weather conditions. A new short-range 24GHz close proximity detection (CPD) microwave radar technology for providing situational awareness to the driver of large vehicles is presented. A series of experiments are performed in which the radar is used to detect a variety of static targets. The CPD radar is able to detect light vehicles at ranges from 2.2m to 17.6m within a horizontal arc of 160° with superior angular resolution to many existing low-cost radars. From this data, key characteristics of the radar's performance and the radar cross-section of light vehicles at 24GHz are calculated. The data gathered allows the design of future tests and the development of a model of the radar that allows a better understanding of its performance characteristics and can be used for occupancy grid mapping.
Bryan Clarke, Stewart Worrall 0002, Graham M. Brooker, Javier Martinez, Eduardo M. Nebot
Intelligent Vehicles Symposium2
2008 A probabilistic method for detecting impending vehicle interactions
abstract
In mining operations it is advantageous to be able to predict the future movements of nearby vehicles. For autonomous mining, this can be used for localised, short term path planning and risk assessment. For semi-autonomous or non-autonomous mining, this can be used for collision avoidance, situational awareness and risk assessment of maneuvers between a human operated vehicle, and another vehicle (operated by a human or otherwise). This paper introduces a probabilistic approach to predicting vehicle movements, in particular, the time until two vehicle paths intersect. Results are shown using real data collected from the operation of two separate fleets of vehicles.
Stewart Worrall 0002, Eduardo M. Nebot
ICRA1
2007 Using Non-Parametric Filters and Sparse Observations to Localise a Fleet of Mining Vehicles
abstract
Mining operations generally involve a large number of expensive vehicles, and for the efficient management of these vehicles it is very beneficial to know their location at all times. The current procedure for vehicle localisation in mines is to provide the mine with complete wireless network coverage to facilitate the broadcasting of vehicle positions. This paper examines an alternative method of localisation that does not require the expense of a radio network with full mine coverage. Two different non-parametric filter approaches are presented to estimate the location of the vehicles. A comparison of the two filters is also presented with experimental results using data collected in two operational mines.
Stewart Worrall 0002, Eduardo M. Nebot
ICRA1