Mao Shan

dblp:130/3040 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-5032-6581ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Blinking Beyond EAR: A Stable Eyelid Angle Metric for Driver Drowsiness Detection and Data Augmentation
Mathis Wolter, Julie Stephany Berrio, Mao Shan
IV3
2026 What demands attention in urban street scenes? From scene understanding towards road safety: A survey of vision-driven datasets and studies
abstract
Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To facilitate the use of these improvements for road safety, this survey systematically categorizes the critical elements that demand attention in traffic scenarios and comprehensively analyzes available vision-driven tasks and datasets. Compared to existing surveys that focus on isolated domains, our taxonomy categorizes attention-worthy traffic entities into two main groups, namely anomalies (abnormal entities) and pertinent entities (normal but critical elements), integrating eleven categories and twenty-three subclasses. It establishes connections between inherently related fields and provides a unified analytical framework. Based on the proposed taxonomy, our survey highlights the analysis of 40 vision-driven tasks and the comprehensive examinations and visualizations of 78 available datasets, including their basic characteristics, sensor settings, label design, visualization practices, annotation schemas, and the resulting implications. The cross-domain investigation reveals substantial variations in benchmark quality across tasks, with recurring limitations including uneven task coverage, imbalanced distributions, inconsistent or insufficient annotations, and limited multimodal and cross-task support. Our article further outlines promising solutions from the perspectives of task formulation, benchmark evaluations, dataset adoption and future dataset construction. The integrated taxonomy, comprehensive analysis, and recapitulatory tables provide researchers with a holistic overview of this rapidly evolving field, guiding strategic resource selection, and highlighting critical yet underexplored areas.
Yaoqi Huang, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002
Eng. Appl. Artif. Intell.3
2025 Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X Collaboration
abstract
Vehicle-to-everything (V2X) collaborative perception has emerged as a promising solution to address the limitations of single-vehicle perception systems. However, existing V2X datasets are limited in scope, diversity, and quality. To address these gaps, we present Mixed Signals, a comprehensive V2X dataset featuring 45.1k point clouds and 240.6k bounding boxes collected from three connected autonomous vehicles (CAVs) equipped with two different configurations of LiDAR sensors, plus a roadside unit with dual LiDARs. Our dataset provides point clouds and bounding box annotations across 10 classes, ensuring reliable data for perception training. We provide detailed statistical analysis on the quality of our dataset and extensively benchmark existing V2X methods on it. The Mixed Signals dataset is ready-to-use, with precise alignment and consistent annotations across time and viewpoints. Dataset website is available at https://mixedsignalsdataset.cs.cornell.edu/.
Katie Luo, Minh-Quan Dao, Mark E. Campbell, Wei-Lun Chao, Kilian Q. Weinberger, Ezio Malis, Vincent Frémont, Bharath Hariharan, Mao Shan, Stewart Worrall 0002, Julie Stephany Berrio
ICCV10
2024 InverseMatrixVT3D: An Efficient Projection Matrix-Based Approach for 3D Occupancy Prediction
abstract
This paper introduces InverseMatrixVT3D, an efficient method for transforming multi-view image features into 3D feature volumes for 3D semantic occupancy prediction. Existing methods for constructing 3D volumes often rely on depth estimation, device-specific operators, or transformer queries, which hinders the widespread adoption of 3D occupancy models. In contrast, our approach leverages two projection matrices to store the static mapping relationships and matrix multiplications to efficiently generate global Bird’s Eye View (BEV) features and local 3D feature volumes. Specifically, we achieve this by performing matrix multiplications between multi-view image feature maps and two sparse projection matrices. We introduce a sparse matrix handling technique for the projection matrices to optimize GPU memory usage. Moreover, a global-local attention fusion module is proposed to integrate the global BEV features with the local 3D feature volumes to obtain the final 3D volume. We also employ a multi-scale supervision mechanism to enhance performance further. Extensive experiments performed on the nuScenes and SemanticKITTI datasets reveal that our approach not only stands out for its simplicity and effectiveness but also achieves the top performance in detecting vulnerable road users (VRU), crucial for autonomous driving and road safety. The code has been made available at: https://github.com/DanielMing123/InverseMatrixVT3D
Zhenxing Ming, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002
IROS3
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions (i) which information should be exchanged over the V2X network, and (ii) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code is released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IV4
2024 Label-Efficient 3D Object Detection For Road-Side Units
abstract
Occlusion presents a significant challenge for safety-critical applications such as autonomous driving. Collaborative perception has recently attracted a large research interest thanks to the ability to enhance the perception of autonomous vehicles via deep information fusion with intelligent roadside units (RSU), thus minimizing the impact of occlusion. While significant advancement has been made, the data-hungry nature of these methods creates a major hurdle for their realworld deployment, particularly due to the need for annotated RSU data. Manually annotating the vast amount of RSU data required for training is prohibitively expensive, given the sheer number of intersections and the effort involved in annotating point clouds. We address this challenge by devising a label-efficient object detection method for RSU based on unsupervised object discovery. Our paper introduces two new modules: one for object discovery based on a spatial temporal aggregation of point clouds, and another for refinement. Furthermore, we demonstrate that fine-tuning on a small portion of annotated data allows our object discovery models to narrow the performance gap with, or even surpass, fully supervised models. Extensive experiments are carried out in simulated and real-world datasets to evaluate our method†.
Minh-Quan Dao, Holger Caesar, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Vincent Frémont, Ezio Malis
IV4
2024 Safety Driver Attention on Autonomous Vehicle Operation Based on Head Pose and Vehicle Perception
abstract
Despite the continual advances in Advanced Driver Assistance Systems (ADAS) and the development of high-level autonomous vehicles (AV), there is a consensus that for the short to medium term, there is a requirement for a human supervisor to handle the edge cases that inevitably arise. Given this requirement, the state of the autonomous vehicle operator (referred to as the safety driver) must be monitored to ensure their contribution to the vehicle's safe operation. This paper introduces a dual-source approach integrating data from an infrared camera facing the safety driver and vehicle perception systems to produce a metric for safety driver alertness to promote and ensure safe operator behaviour. The infrared camera detects the safety driver’s head, enabling the calculation of head orientation, which is relevant as the head typically moves according to the individual's focus of attention. By incorporating environmental data from the perception system, it becomes possible to determine whether the safety driver observes objects in the surroundings. Experiments were conducted using data collected in Sydney, Australia, simulating AV operations in an urban environment. Our results demonstrate that the proposed system effectively determines a metric for the attention levels of the safety driver, enabling interventions such as warnings or reducing autonomous functionality as appropriate. The results indicate reduced awareness on subsequent laps during the study, demonstrating the "automation complacency" phenomenon. This comprehensive solution shows promise in contributing to ADAS and AVs’ overall safety and efficiency in a real-world setting.
Santiago Gerling Konrad, Julie Stephany Berrio, Mao Shan, Favio R. Masson, Eduardo M. Nebot, Stewart Worrall 0002
IV3
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions 1) which information should be exchanged over the V2X network and 2) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code will be released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IEEE Trans. Intell. Transp. Syst.4
2023 Viewer-Centred Surface Completion for Unsupervised Domain Adaptation in 3D Object Detection
abstract
Every autonomous driving dataset has a different configuration of sensors, originating from distinct geographic regions and covering various scenarios. As a result, 3D detectors tend to overfit the datasets they are trained on. This causes a drastic decrease in accuracy when the detectors are trained on one dataset and tested on another. We observe that lidar scan pattern differences form a large component of this reduction in performance. We address this in our approach, SEE-VCN, by designing a novel viewer-centred surface completion network (VCN) to complete the surfaces of objects of interest within an unsupervised domain adaptation framework, SEE [1]. With SEE-VCN, we obtain a unified representation of objects across datasets, allowing the network to focus on learning geometry, rather than overfitting on scan patterns. By adopting a domain-invariant representation, SEE-VCN can be classed as a multi-target domain adaptation approach where no annotations or re-training is required to obtain 3D detections for new scan patterns. Through extensive experiments, we show that our approach outperforms previous domain adaptation methods in multiple domain adaptation settings. Our code and data are available at https://github.com/darrenjkt/SEE-VCN.
Darren Tsai, Julie Stephany Berrio, Mao Shan, Eduardo M. Nebot, Stewart Worrall 0002
ICRA3
2022 Towards Collision-Free Probabilistic Pedestrian Motion Prediction for Autonomous Vehicles
abstract
Autonomous vehicle navigation in shared pedestrian environments requires the ability to predict future crowd motion as well as understand human behaviour. However, most existing methods predict pedestrian future motion without considering potential collisions within the crowd. Furthermore, most current predictive models are tested on datasets that assume full observability of the crowd by relying on a top-down view, which does not reflect the real-world use case of autonomous vehicles due to the inherent limitations of on-board sensors such as visual occlusion. Inspired by prior works, we propose a pedestrian motion prediction model trained via contrastive learning, improving prediction accuracy as well as forecasting collision-free trajectories. Additionally, we propose a method for implementing a predictor using a multi-pedestrian probabilistic tracker, which fuses multiple on-board sensors to track pedestrians in 3D space. Through comprehensive experiments on both aerial view and driving datasets collected in a real-world urban environment, we show that our proposed method improves on state of art methods with better prediction accuracy and more socially acceptable prediction trajectories.
Kunming Li, Mao Shan, Stuart Eiffert, Stewart Worrall 0002, Eduardo M. Nebot
IV2
2022 Multimobile Robot Cooperative Localization Using Ultrawideband Sensor and GPU Acceleration
abstract
To tackle the poor localization accuracy of multimobile robots caused by non-line-of-sight (NLOS) errors in a complex indoor environment and to meet the real-time requirement, this article proposes a multimobile robot cooperative localization system using ultrawideband (UWB) sensor and GPU hardware acceleration. First, a UWB multinode ranging network is established to obtain the relative distance information between robots and anchors. Then, the line-of-sight (LOS) and NLOS errors in distance information are effectively mitigated by using the proposed UWB ranging error mitigation algorithm based on the Bayesian filter. A cooperative particle filter (PF) localization algorithm based on the Gibbs sampling is designed to estimate the position information of each robot at any time. Finally, in order to improve the real-time performance of the collaborative localization system, a parallel Gibbs collaborative localization algorithm that can be accelerated by GPU is proposed considering the characteristics of GPU hardware and CUDA programming model. The experimental results of three TurtleBot2 mobile robots in real scene show that the proposed multimobile robot cooperative localization system using UWB technology can estimate the position information of each robot robustly and accurately, and the localization accuracy is superior to that of the popular extended Kalman filter (EKF) and PF algorithms. It is shown through further evaluations that the proposed parallel algorithm achieves about 3.2 times acceleration effect in the scenarios of three mobile robots. The speed gain is found more significant with more robots, which substantially improves the real-time performance of the cooperative localization system. In the test with seven mobile robots, the speedup is as high as 11.9, that is, the execution time of the algorithm is only 8.39% of that of the original algorithm. Note to Practitioners—The purpose of this article is to improve the accuracy and real-time indoor multimobile robot cooperative localization, but the method proposed in this article is also applicable to outdoor multimobile robot cooperative localization. The existing methods for indoor cooperative localization of mobile robots usually use Bluetooth, infrared, RFID, and other technologies to establish a wireless sensor network (WSN) and then combine Karman filter or particle filter (PF) to achieve cooperative localization, which is difficult to achieve low-cost and high-precision real-time localization. In this article, a new method of cooperative localization is proposed, which uses ultrawideband (UWB) ranging network with high penetration and high precision to obtain accurate distance information, then weakens NLOS error by the Bayesian filtering to further improve the accuracy of distance information, and, finally, uses a novel cooperative localization approach to realize fast and high-precision indoor multimobile robot localization. In addition, by redesigning the collaborative localization algorithm in parallel, the real-time performance of the algorithm is improved while ensuring high accuracy. The collaborative localization experiment of three mobile robots in the real scene shows that the proposed algorithm can effectively improve the localization accuracy, and the real-time performance of collaborative localization is significantly improved. However, the algorithm still has some limitations when mitigating UWB ranging errors in highly obstructed environments, and the algorithm parallelization framework can be further improved for higher real-time performance. In the future, we will further improve the robustness of cooperative localization and apply this algorithm to cooperative control of multiple mobile robots, such as formation control and cooperative search.
Jing Xin, Guo Xie, Mao Shan, Peng Li 0007, Kaiyuan Gao
IEEE Trans Autom. Sci. Eng.4
2022 Camera-LIDAR Integration: Probabilistic Sensor Fusion for Semantic Mapping
abstract
An automated vehicle operating in an urban environment must be able to perceive and recognise objects and obstacles in a three-dimensional world for navigation and path planning. In order to plan and execute accurate and sophisticated driving maneuvers, a high-level contextual understanding of the surroundings is essential. Due to the recent progress in image processing, it is now possible to obtain high definition semantic information in 2D from monocular cameras, though cameras cannot reliably provide the high accuracy 3D information provided by lasers. The fusion of these two sensor modalities can overcome the shortcomings of each individual sensor, though there are a number of important challenges that need to be addressed in a probabilistic manner. In this paper we address the common, yet challenging, LIDAR/camera/semantic fusion problems which are seldom approached in a wholly probabilistic manner. Our approach is capable of using a multi-sensor platform to build a three-dimensional semantic voxelized map that considers the uncertainty of all of the processes involved. We present a probabilistic pipeline that incorporates uncertainty from the sensor readings (cameras, LIDAR, IMU and wheel encoders), compensation for the motion of the vehicle, and heuristic label probabilities for the semantic images depicted inFig. 1. We also present a novel and efficient viewpoint validation algorithm to check for occlusions within the camera frame. A probabilistic projection is performed from the camera images to the LIDAR point cloud. Each labelled LIDAR scan then feeds into an octree map-building algorithm that updates the class probabilities of the map voxels every time a new observation is available. We validate our approach using a set of qualitative and quantitative experiments using the USyd Campus Dataset.1These tests demonstrate the usefulness of a probabilistic sensor fusion approach by evaluating the performance of the perception system in a typical autonomous vehicle application.
Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.2
2022 Long-Term Map Maintenance Pipeline for Autonomous Vehicles
abstract
For autonomous vehicles to operate persistently in a typical urban environment, it is essential to have high accuracy position information. This requires a mapping and localisation system that can adapt to changes over time. A localisation approach based on a single-survey map will not be suitable for long-term operation as it does not incorporate variations in the environment. In this paper, we present new algorithms to maintain a featured-based map. A map maintenance pipeline is proposed that can continuously update a map with the most relevant features taking advantage of the changes in the surroundings. Our pipeline detects and removes transient features based on their geometrical relationships with the vehicle’s pose. Newly identified features became part of a new feature map and are assessed by the pipeline as candidates for the localisation map. By purging out-of-date features and adding newly detected features, we continually update the prior map to more accurately represent the most recent environment. We have validated our approach using the USyd Campus Dataset, which includes more than 18 months of data. The results presented demonstrate that our maintenance pipeline produces a resilient map which can provide sustained localisation performance over time.
Julie Stephany Berrio, Stewart Worrall 0002, Mao Shan, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.3
2021 Attentional-GCNN: Adaptive Pedestrian Trajectory Prediction towards Generic Autonomous Vehicle Use Cases
abstract
Autonomous vehicle navigation in shared pedestrian environments requires the ability to predict future crowd motion both accurately and with minimal delay. Understanding the uncertainty of the prediction is also crucial. Most existing approaches however can only estimate uncertainty through repeated sampling of generative models. Additionally, most current predictive models are trained on datasets that assume complete observability of the crowd using an aerial view. These are generally not representative of real-world usage from a vehicle perspective, and can lead to the underestimation of uncertainty bounds when the on-board sensors are occluded. Inspired by prior work in motion prediction using spatio-temporal graphs, we propose a novel Graph Convolutional Neural Network (GCNN)-based approach, Attentional-GCNN, which aggregates information of implicit interaction between pedestrians in a crowd by assigning attention weight in edges of the graph. Our model can either output a probabilistic distribution or faster deterministic prediction, demonstrating applicability to autonomous vehicle use cases where either speed or accuracy with uncertainty bounds are required. To further improve the training of predictive models, we propose an automatically labelled pedestrian dataset collected from an intelligent vehicle platform representative of real-world use. Through experiments on a number of datasets, we show our proposed method achieves an improvement over the state of the art by 10% on Average Displacement Error (ADE) and 12% on Final Displacement Error (FDE) with fast inference speeds.
Kunming Li, Stuart Eiffert, Mao Shan, Francisco Gomez-Donoso, Stewart Worrall 0002, Eduardo M. Nebot
ICRA3
2019 Uncertainty Estimation for Projecting Lidar Points onto Camera Images for Moving Platforms
abstract
Combining multiple sensors for advanced perception is a crucial requirement for autonomous vehicle navigation. Heterogeneous sensors are used to obtain rich information about the surrounding environment. The combination of the camera and lidar sensors enables precise range information that can be projected onto the visual image data. This gives a high level understanding of the scene which can be used to enable context based algorithms such as collision avoidance and navigation. The main challenge when combining these sensors is aligning the data into a common domain. This can be difficult due to the errors in the intrinsic calibration of the camera, extrinsic calibration between the camera and the lidar and errors resulting from the motion of the platform. In this paper, we examine the algorithms required to provide motion correction for scanning lidar sensors. The error resulting from the projection of the lidar measurements into a consistent odometry frame is not possible to remove entirely, and as such it is essential to incorporate the uncertainty of this projection when combining the two different sensor frames. This work proposes a novel framework for the prediction of the uncertainty of lidar measurements (in 3D) projected in to the image frame (in 2D) for moving platforms. The proposed approach fuses the uncertainty of the motion correction with uncertainty resulting from errors in the extrinsic and intrinsic calibration. By incorporating the main components of the projection error, the uncertainty of the estimation process is better represented. Experimental results for our motion correction algorithm and the proposed extended uncertainty model are demonstrated using real-world data collected on an electric vehicle equipped with wide-angle cameras covering a 180-degree field of view and a 16-beam scanning lidar.
Charika De Alvis, Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
ICRA2
2019 Extended Vehicle Tracking with Probabilistic Spatial Relation Projection and Consideration of Shape Feature Uncertainties
abstract
This work focuses on a novel probabilistic approach for extended vehicle tracking, where multiple spatially distributed measurements can originate from the target, and kinematic state and geometry variables are estimated jointly. Prominent shape features extracted from raw measurement points contain spatial uncertainties due to noise in sensor measurements, the feature extraction process, approximation error of shape hypothesis, partial vision occlusion, to name a few. This work proposes a novel tracking paradigm that respects the variant spatial measurement model subject to changes in target pose and sensor viewpoint. This is achieved through probabilistic projection of the spatial measurement points to the predicted measurement sources on the visible side(s) of the target shape. The spatial uncertainties in the shape features are probabilistically modelled and incorporated in the unscented Kalman filter based estimation. The proposed approach is validated with field experiment results using cameras and a laser range scanner.
Mao Shan, Charika De Alvis, Stewart Worrall 0002, Eduardo M. Nebot
IV1
2018 Pedestrian Dynamic and Kinematic Information Obtained from Vision Sensors
abstract
The estimation and prediction of pedestrian motion is of fundamental importance in ITS applications. Most existing solutions have utilized a particular type of sensor for perception such as cameras (stereo, monocular, infrared) or other modalities such as a laser range finder or radar. The advent of wearable devices with inertial sensors have led to the development of systems capable of the robust inference of pedestrian intention. Unfortunately, these devices do not have communications capabilities to broadcast this information to all vehicles in proximity, and also this strategy requires functioning devices on all pedestrians to work. This paper presents a robust perception method that is able to extract dynamic pedestrian information with accuracy comparable to that of typical gyroscopes and accelerometers installed in wearable devices. Experimental results are presented to demonstrate the potential for obtaining very comprehensive dynamic information from limbs representing the skeleton of a pedestrian. This work also demonstrates the accuracy of vision based systems by comparing these results to the rotation and acceleration measured directly on the pedestrian using a wearable device. The contributions of this paper demonstrate that it is possible to significantly improve both the detection and estimation of pedestrian intention by incorporating dynamic information obtained from vision sensors.
Santiago Gerling Konrad, Mao Shan, Favio R. Masson, Stewart Worrall 0002, Eduardo M. Nebot
Intelligent Vehicles Symposium2
2016 Visual tracking via random partition image hashing
abstract
In this paper, we propose a discriminative and robust appearance model based on features extracted from a random partition image hashing algorithm to account for severe occlusion and disappearance. We divide the original image into multiple sub-blocks with random positions and scales. Hash functions are used to map blocks into compact binary codes, with which more effective target matching can be achieved. The tracking task is then formulated by producing a confidence map for the target and background, and obtaining the best samples using maximum a posteriori estimate. Experimental results demonstrate that our tracker can achieve more accurate tracking results in situations of occlusion, out-of-view, and violent motion blur when compared with most of state-of-the-art competing algorithms. Besides, the proposed tracking algorithm is able to run in real time.
Mingyang Guan, Changyun Wen, Kwang-Yong Lim, Mao Shan, Paul Tan, Cheng-Leong Ng, Ying Zou 0002
ICARCV4
2016 Probabilistic trajectory estimation based leader following for multi-robot systems
abstract
The paper is concerned with the multi-robot leader-following problem in the presence of frequent dropouts in vision detection. In many scenarios, for instance a structured environment, it is inevitable to experience outage of vision detection due to reasons such as the target moving out of view, vision occlusion, motion blurring, etc. The paper proposes a Bayesian trajectory estimation based leader-following approach that can offer accurate path following given intermittent vision observations. The follower robot estimates the trajectory of the leader robot based on the noise-corrupted odometry information of both robots, and inter-robot relative observations based on detection of fiducial markers using an RGBD camera. A linear trajectory-following control method is employed to track a historical pose of the leader robot on the estimated trajectory. Results are obtained based on evaluating the proposed leader-following approach in tests with a zig-zag shaped trajectory and with a trajectory that contains sharp turns.
Mao Shan, Ying Zou 0002, Mingyang Guan, Changyun Wen, Kwang-Yong Lim, Cheng-Leong Ng, Paul Tan
ICARCV1
2016 Image-based visual tracking adaptive control for mobile robots
abstract
In this paper, we deal with the problem of image-based visual tracking for mobile robots. It is noted that the presence of actuator dynamics and the unknown target motion increases the complexity of the system model and makes the design of the controller more difficult. To solve this problem, we propose an adaptive control approach via utilizing backstepping technique and extended state observer (ESO). For the controller design, the adaptive technique is employed to estimate the bound of the target motion and the adaptive law is derived from the Lyapunov stability theory. We also adopt two ESOs to estimate and compensate for the disturbances affecting the mobile robot dynamics and the wheel actuator dynamics. It is shown that the proposed adaptive controller guarantees the boundness of all the signals in the closed-loop system and enables the tracking errors to exponentially converge to a compact set which is adjustable.
Ying Zou 0002, Changyun Wen, Mao Shan, Mingyang Guan
ICARCV3
2015 Delayed-State Nonparametric Filtering in Cooperative Tracking
abstract
This paper presents a novel nonparametric approach toward delayed-state filtering for cooperative tracking. Standard parametric cooperative localization/tracking approaches are generally aimed at problems that can be easily parameterized and/or are limited to incorporate only real-time measurements. This paper provides a nonparametric yet computationally tractable alternative that is suitable for tracking cases where real-time observations are not always possible, e.g., in a sparse mesh network. The proposed delayed-state cooperative particle filter features forward filtering and backward smoothing to incorporate measurements that are received with time delays. A record of historical marginal states is kept for each mobile node within a sliding time window, instead of the high-dimensional joint state. Essentially, it replaces the importance sampling in traditional particle filters by a Gibbs sampler, which is a Markov chain Monte Carlo method, to fuse all available egocentric and internode relative observations into the global position estimate, thus alleviating the high-dimensionality problems in cooperative tracking. The performance of the proposed approach is evaluated in a multiagent simulation, and experimental results from a large-scale multivehicle industrial operation clearly demonstrate that the proposed approach effectively facilitates the tracking of mobile nodes without position awareness, through the use of relative range, negative detection, and time-delayed measurements.
Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Robotics1
2014 Nonparametric cooperative tracking in mobile Ad-Hoc networks
abstract
This paper presents a new nonparametric approach for cooperative localisation and tracking by fusing relative range measurements between mobile network equipped nodes. Standard approaches based on parametric methods are known to be limited to problems that contain Gaussian properties. This paper overcomes this limitation by proposing a novel particle filter based cooperative tracking approach that is suitable for mobile ad-hoc networks (MANETs). The filter maintains the marginal state of every node instead of a joint state of the group, and updates the estimates using a Gibbs sampler, which is known as a Markov chain Monte Carlo (MCMC) method. The performance of the proposed algorithm is demonstrated by examining 16 mobile nodes moving randomly while sharing information with neighbouring nodes in a MANET. The results show that the proposed approach facilitates the tracking of mobile nodes that do not have egocentric position information available. It also outperforms the EKF in systems with non-Gaussian properties.
Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
ICRA1
2014 Using Delayed Observations for Long-Term Vehicle Tracking in Large Environments
abstract
The tracking of vehicles over large areas with limited position observations is of significant importance in many industrial applications. This paper presents algorithms for long-term vehicle motion estimation based on a vehicle motion model that incorporates the properties of the working environment and information collected by other mobile agents and fixed infrastructure collection points. The prediction algorithm provides long-term estimates of vehicle positions using speed and timing profiles built for a particular environment and considering the probability of a vehicle stopping. A limited number of data collection points distributed around the field are used to update the estimates, with negative information (no communication) also used to improve the prediction. This paper introduces the concept of observation harvesting, a process in which peer-to-peer communication between vehicles allows egocentric position updates to be relayed among vehicles and finally conveyed to the collection point for an improved position estimate. Positive and negative communication information is incorporated into the fusion stage, and a particle filter is used to incorporate the delayed observations harvested from vehicles in the field to improve the position estimates. The contributions of this work enable the optimization of fleet scheduling using discrete observations. Experimental results from a typical large-scale mining operation are presented to validate the algorithms.
Mao Shan, Stewart Worrall 0002, Favio R. Masson, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.1
2013 Probabilistic Long-Term Vehicle Motion Prediction and Tracking in Large Environments
abstract
Vehicle position tracking and prediction over large areas is of significant importance in many industrial applications, such as mining operations. In a small area, this can easily be achieved by providing vehicles with a constant communication link to a control center and having the vehicles broadcast their position. The problem dramatically changes when vehicles operate within a large environment of potentially hundreds of square kilometers and in difficult terrain. This paper presents algorithms for long-term vehicle motion prediction and tracking based on a multiple-model approach. It incorporates a probabilistic vehicle model that includes the structure of the environment. The prediction algorithm evaluates the vehicle position using acceleration, speed, and timing profiles built for the particular environment and considers the probability that the vehicle will stop. A limited number of data collection points distributed around the field are used to update the vehicle position estimate when in communication range, and prediction is used at points in between. A particle filter is used to estimate the vehicle position using both positive and negative information (whether communication is possible) in the fusion stage. The algorithms presented are validated with experimental results using data collected from a large-scale mining operation.
Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.1