EDBT 2026 Demo / reviewers in the wild / expert
Huijing Zhao
dblp:95/967
· DBLP profile ↗
91ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0001-9245-3039ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 6 first-author · 6 since 2021Systems, architecture and hardware · 30 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11Human-computer interaction and ubiquitous computing · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TerraX: Visual Terrain Classification Enhanced by Vision-Language ModelsabstractVisual Terrain Classification (VTC) plays a vital role in enabling unmanned ground vehicles to understand complex environments. Existing research relies on image-label pairs annotated by static label sets, where semantic ambiguity and high annotation costs constrain fine-grained terrain characterization. These limitations hinder the model’s adaptation to real-world terrain diversity and restrict its applicability. To address these issues, we propose TerraX, a vision-language learning framework that integrates multi-modal image-label-text data, unifying structured annotations with fine-grained natural language descriptions. The framework introduces a composite dataset TerraData, an evaluation benchmark suite TerraBench, and a CLIP-based visual terrain classification model TerraCLIP. TerraData aggregates multi-source terrain images from public and self-collected datasets, annotated through a VLM-based vision-language data annotation pipeline. TerraBench defines three evaluation benchmarks to systematically assess model robustness and adaptability in real-world terrain classification scenarios. Built on the CLIP model, TerraCLIP utilizes multi-granularity contrastive loss and LoRA fine-tuning to enhance understanding for terrain categories and attributes, and incorporates confidence-weighted inference for accurate predictions. Extensive experiments across benchmarks and real-world platforms demonstrate that our approach significantly enhances VTC performance, highlighting its potential for deployment in complex environments. Hongze Li, Xuchuan Huang, Xinhai Chang, Huijing Zhao |
IROS | 5 |
| 2025 | How to Enhance the Interpretability of Learning-Based Motion Planning for Intelligent Vehicles - A SurveyabstractWith the advancement of deep learning, the learning-based motion planning (MP) approach exhibits immense potential in intelligent vehicles (IVs). Because the principle and framework of the learning-based MP method differ from the traditional MP methods, exploring effective strategies to enhance interpretability plays an important role. This survey fills the gaps in the IV field’s learning-based motion planning and interpretability enhancement. Our study aims to explore two fundamental inquiries. Firstly, how can we design learning-based MP to achieve high performance? Secondly, how can we enhance the interpretability of learning-based MP? To this end, this paper provides an extensive overview of more than 200 papers employed in learning-based MP techniques within the last 10 years. By summarizing these techniques, a taxonomy for integrating learning-based MP techniques into an IV architecture is presented as three modes: learning-based key-module generator, learning-based trajectory generator, and learning-based policy generator. Interpretability enhancement has different considerations for different modes. Additionally, we compile a summary of resources utilized in learning-based MP. Finally, we discuss critical challenges and make suggestions. Tao Wu 0001, Huijing Zhao, Xin Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Uncertainty-aware Deep Imitation Learning and Deployment for Autonomous Navigation through Crowded IntersectionsabstractNavigation through crowded intersections is a challenge for autonomous vehicles, where uncertainty arises from interaction with other road users, encountering new scenes and weathers, etc. Recent end-to-end autonomous control deep models learned from human drivers have shown promising driving performance, whereas they are not as transparent and safe as traditional rule-based systems. When facing situations that they are unfamiliar with or uncertain about, the deep models’ predictions could be unsafe and untrustworthy. Without the ability to identify these situations and issue warnings beforehand, cascading errors of deep models may result in catastrophes. Therefore, this work combines the strengths of both data-driven and traditional rule-based approaches to achieve better driving quality and safety. We propose a heterogeneity uncertainty quantification method based on imitation learning, where both data and model uncertainties of the lateral and longitudinal control tasks are quantified. We also propose a policy deployment strategy where a safety indicator is developed upon estimated uncertainty to bridge the data-driven performance layer and the rule-based fallback layer. We learned from human driving demonstrations and conducted extensive closed-loop tests. Results demonstrate the effectiveness and importance of the proposed uncertainty quantification method and policy deployment strategy. Huijing Zhao |
IROS | 3 |
| 2023 | Re-Thinking Classification Confidence with Model Quality QuantificationabstractDeep neural networks using for real-world classification task require high reliability and robustness. However, the Softmax output by the last layer of network is often over-confident. We propose a novel confidence estimation method by considering model quality for deep classification models. Two metrics, MQ-Repres and MQ-Discri are developed accordingly to evaluate the model quality, and also provide a new confidence estimation called MQ-Conf for online inference. We demonstrate the capability of the proposed method by the$3D$semantic segmentation tasks using three different deep networks. Through confusion analysis and feature visualization we show the rationality and reliability of the model quality quantification method.11This work is supported by the National Natural Science Foundation of China under Grant U22A2061 and High-performance Computing Platform of Peking University. Yancheng Pan, Huijing Zhao |
IROS | 2 |
| 2023 | Joint Imitation Learning of Behavior Decision and Control for Autonomous Intersection NavigationabstractModern autonomous driving systems face substantial challenges when navigating dense intersections due to the high uncertainty introduced by other road users. Due to the complexity of the task, the autonomous vehicle needs to generate policies at multiple levels of abstraction. However, previous deep imitation learning methods focused on learning control policies while using simple rule-based behavior models. To bridge this gap and achieve human-like driving, we develop a hierarchy of high-level behavior decision and low-level control, where both policies are jointly learned from human demonstrations based on imitation learning. Over 60 hours of driving data from 10 drivers at six intersections was collected. The proposed method is extensively evaluated in challenging intersection scenarios. Empirical results demonstrate the method's superior performance over baselines in terms of task completion and control quality. We demonstrate the importance of learning human-like behavior decisions as well as joint learning of behavior and control policies. The capability of imitating different driving styles is also illustrated. Huijing Zhao |
IROS | 2 |
| 2023 | An Active and Contrastive Learning Framework for Fine-Grained Off-Road Semantic SegmentationabstractOff-road semantic segmentation with fine-grained labels is necessary for autonomous vehicles to understand driving scenes, as the coarse-grained road detection cannot satisfy off-road vehicles with various mechanical properties. Pixel-wise annotation of fine-grained labels in off-road scenes is very hard because a large part of the pixels could suffer from severe semantic ambiguity. Furthermore, semantic properties of off-road scenes can be very changeable due to various precipitations, temperature, defoliation, etc. To address these challenges, this research proposes an active and contrastive learning-based method. A few image patches are annotated mainly to distinguish semantic differences rather than semantic categories, which can greatly reduce the burden of manual annotation. A feature representation is learnt using the contrastive pairs of image patches, and semantic categories are adaptively modeled from the data. To actively adapt to new scenes, a risk evaluation method is developed to discover and select hard frames with high-risk predictions for supplementary labeling, to update the model efficiently. Extensive experiments and analyses are conducted on self-developed and public datasets. Experimental results demonstrate that fine-grained semantic segmentation can be learned with only dozens of weakly labeled frames, and the model can efficiently adapt across scenes by weak supervision, while achieving competitive performance with the typical fully supervised ones. Biao Gao, Xijun Zhao, Huijing Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Understanding the Challenges When 3D Semantic Segmentation Faces Class Imbalanced and OOD Dataabstract3D semantic segmentation (3DSS) is an essential process in the creation of a safe autonomous driving system. However, deep learning models for 3D semantic segmentation often suffer from the class imbalance problem and out-of-distribution (OOD) data. In this study, we explore how the class imbalance problem affects 3DSS performance and whether the model can detect the category prediction correctness, or whether data is ID or OOD. For these purposes, we conduct two experiments using four representative 3DSS models and five trust scoring methods, and conduct both a confusion and feature analysis of each class. Furthermore, a data augmentation method for the 3D LiDAR dataset is proposed to create a new dataset based on SemanticKITTI and SemanticPOSS, called AugKITTI. We propose the wPre metric and TSD for a more in-depth analysis of the results, and follow are proposals with an insightful discussion. Based on the experimental results, we find that: 1) classes are not only imbalanced in their data size but also in the basic properties of each semantic category; 2) intraclass diversity and interclass ambiguity make class learning difficult and greatly limit the models’ performance, creating the challenges of semantic and data gaps; 3) trust scores are unreliable for classes whose features are confused with other classes. For 3DSS models, those misclassified ID classes and OODs may also be given high trust scores, making the 3DSS predictions unreliable, and leading to the challenges in judging 3DSS result trustworthiness. All of these outcomes point to several research directions for improving the performance and reliability of the 3DSS models used for real-world applications. Yancheng Pan, Fan Xie 0007, Huijing Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Contrastive-Learning-Based Method for Alert-Scene CategorizationabstractWhether it’s a driver warning or an autonomous driving system, ADAS needs to decide when to alert the driver of danger or take over control. This research formulates the problem as an alert-scene categorization one and proposes a method using contrastive learning. Given a front-view video of a driving scene, a set of anchor points is marked by a human driver, where an anchor point indicates that the semantic attribute of the current scene is different from that of the previous one. The anchor frames are then used to generate contrastive image pairs to train a feature encoder and obtain a scene similarity measure, so as to expand the distance of the scenes of different categories in the feature space. Each scene category is explicitly modeled to capture the meta pattern on the distribution of scene similarity values, which is then used to infer scene categories. Experiments are conducted using front-view videos that were collected during driving at a cluttered dynamic campus. The scenes are categorized into no alert, longitudinal alert, and lateral alert. The results are studied at both feature encoding, category modeling, and reasoning aspects. By comparing precision with two full supervised end-to-end baseline models, the proposed method demonstrates competitive or superior performance. However, it remains still questions: how to generate ground truth data and how to evaluate performance in ambiguous situations, which leads to future works. Shaochi Hu, Hanwei Fan, Biao Gao, Huijing Zhao |
IV | 4 |
| 2022 | Driver Identification Through Heterogeneity Modeling in Car-Following SequencesabstractIntra-driver and inter-driver heterogeneity has been confirmed to exist in human driving behaviors by many studies. This research proposes a driver identification method by modeling such heterogeneities in car following sequences. It is assumed that all drivers share a pool of driver states; under each state, a car-following data sequence obeys a specific probability distribution in feature space; each driver has his/her own probability distribution over the states, called driver profile, which characterize the intra-driver heterogeneity, while the difference between the driver profile of different drivers depicts the inter-driver heterogeneity. Thus, the driver profile can be used to distinguish a driver from others. Based on the assumption, a method of driver identification is proposed to take both intra- and inter-driver heterogeneity into consideration, and a method is developed to jointly learn parameters in behavioral feature extractor, driver states, and driver profiles. Experiments demonstrate the performance of the proposed method in driver identification on naturalistic car-following data: accuracy of 82.3% is achieved in an 8-driver experiment using 10 car-following sequences of duration 15 seconds for online inference. The potential of fast registration of new drivers is demonstrated and discussed. Zhezhang Ding, Donghao Xu, Chenfeng Tu, Huijing Zhao, Mathieu Moze, François Aioun, Franck Guillemard |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Are We Hungry for 3D LiDAR Data for Semantic Segmentation? A Survey of Datasets and Methodsabstract3D semantic segmentation is a fundamental task for robotic and autonomous driving applications. Recent works have been focused on using deep learning techniques, whereas developing fine-annotated 3D LiDAR datasets is extremely labor intensive and requires professional skills. The performance limitation caused by insufficient datasets is called data hunger problem. This research provides a comprehensive survey on the question: are we hungry for 3D LiDAR data for semantic segmentation? The studies are conducted at three levels. First, a broad review to the main 3D LiDAR datasets is conducted, followed by a statistical analysis on three representative datasets to gain an in-depth view on the datasets’ size, diversity and quality, which are the critical factors in learning deep models. Second, an organized survey of 3D semantic segmentation methods is given with a focus on the mainstream of the latest research trend using deep learning techniques, followed by a systematic survey to the existing efforts to solve the data hunger problem. Finally, an insightful discussion of the remaining problems on both methodological and datasets’ viewpoints, and the open questions on dataset bias, domain and semantic gap are given, leading to potential topics in future works. To the best of our knowledge, this is the first work to study the data hunger problem for 3D semantic segmentation using deep learning techniques, which are addressed in both methodological and dataset review, and we share findings and discussions through a comprehensive dataset analysis. Biao Gao, Yancheng Pan, Chengkun Li, Sibo Geng, Huijing Zhao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Scene-Aware Error Modeling of LiDAR/Visual Odometry for Fusion-Based Vehicle LocalizationabstractLocalization is an essential technique in mobile robotics. In a complex environment, it is necessary to fuse different localization modules to obtain more robust results, in which the error model plays a paramount role. However, exteroceptive sensor-based odometries(ESOs), such as LiDAR/visual odometry, often deliver results with scene-related error, which is difficult to model accurately. To address this problem, this research designs a scene-aware error model for ESO, based on which a multimodal localization fusion framework is developed. In addition, an end-to-end learning method is proposed to train this error model using sparse global poses such as results from global positioning system(GPS) and inertial measurement unit(IMU). The proposed method is realized for error modeling of LiDAR/visual odometry, and the results are fused with dead reckoning to examine the performance of vehicle localization. Experiments are conducted using both simulation and real-world data of experienced and unexperienced environments, and the experimental results demonstrate that with the learned scene-aware error models, vehicle localization accuracy can be largely improved and shows adaptiveness in unexperienced scenes. Xiaoliang Ju, Donghao Xu, Huijing Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Survey of Deep RL and IL for Autonomous Driving Policy LearningabstractAutonomous driving (AD) agents generate driving policies based on online perception results, which are obtained at multiple levels of abstraction, e.g., behavior planning, motion planning and control. Driving policies are crucial to the realization of safe, efficient and harmonious driving behaviors, where AD agents still face substantial challenges in complex scenarios. Due to their successful application in fields such as robotics and video games, the use of deep reinforcement learning (DRL) and deep imitation learning (DIL) techniques to derive AD policies have witnessed vast research efforts in recent years. This paper is a comprehensive survey of this body of work, which is conducted at three levels: First, a taxonomy of the literature studies is constructed from the system perspective, among which five modes of integration of DRL/DIL models into an AD architecture are identified. Second, the formulations of DRL/DIL models for conducting specified AD tasks are comprehensively reviewed, where various designs on the model state and action spaces and the reinforcement learning rewards are covered. Finally, an in-depth review is conducted on how the critical issues of AD applications regarding driving safety, interaction with other traffic participants and uncertainty of the environment are addressed by the DRL/DIL models. To the best of our knowledge, this is the first survey to focus on AD policy learning using DRL/DIL, which is addressed simultaneously from the system, task-driven and problem-driven perspectives. We share and discuss findings, which may lead to the investigation of various topics in the future. Huijing Zhao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Fine-Grained Off-Road Semantic Segmentation and Mapping via Contrastive LearningabstractRoad detection or traversability analysis has been a key technique for a mobile robot to traverse complex off-road scenes. The problem has been mainly formulated in early works as a binary classification one, e.g. associating pixels with road or non-road labels. Whereas understanding scenes with fine-grained labels are needed for off-road robots, as scenes are very diverse, and the various mechanical performance of off-road robots may lead to different definitions of safe regions to traverse. How to define and annotate fine-grained labels to achieve meaningful scene understanding for a robot to traverse off-road is still an open question. This research proposes a contrastive learning based method. With a set of human-annotated anchor patches, a feature representation is learned to discriminate regions with different traversability, a method of fine-grained semantic segmentation and mapping is subsequently developed for off-road scene understanding. Experiments are conducted on a dataset of three driving segments that represent very diverse off-road scenes. An anchor accuracy of 89.8% is achieved by evaluating the matching with human-annotated image patches in cross-scene validation. Examined by associated 3D LiDAR data, the fine-grained segments of visual images are demonstrated to have different levels of toughness and terrain elevation, which represents their semantical meaningfulness. The resultant maps contain both fine-grained labels and confidence values, providing rich information to support a robot traversing complex off-road scenes. Biao Gao, Shaochi Hu, Xijun Zhao, Huijing Zhao |
IROS | 4 |
| 2021 | Learning From Naturalistic Driving Data for Human-Like Autonomous Highway DrivingabstractDriving in a human-like manner is important for an autonomous vehicle to be a smart and predictable traffic participant. To achieve this goal, parameters of the motion planning module should be carefully tuned, which needs great effort and expert knowledge. In this study, a method of learning cost parameters of a motion planner from naturalistic driving data is proposed. The learning is achieved by encouraging the selected trajectory to approximate the human driving trajectory under the same traffic situation. The employed motion planner follows a widely accepted methodology that first samples candidate trajectories in the trajectory space, then select the one with minimal cost as the planned trajectory. Moreover, in addition to traditional factors such as comfort, efficiency and safety, the cost function is proposed to incorporate incentive of behavior decision like a human driver, so that both lane change decision and motion planning are coupled into one framework. Two types of lane incentive cost — heuristic and learning based — are proposed and implemented. To verify the validity of the proposed method, a data set is developed by using the naturalistic trajectory data of human drivers collected on the motorways in Beijing, containing samples of lane changes to the left and right lanes, and car followings. Experiments are conducted with respect to both lane change decision and motion planning, and promising results are achieved. Donghao Xu, Zhezhang Ding, Huijing Zhao, Mathieu Moze, François Aioun, Franck Guillemard |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Cross Scene Prediction via Modeling Dynamic Correlation using Latent Space Shared Auto-EncodersabstractThis work addresses on the following problem: given a set of unsynchronized history observations of two scenes that are correlative on their dynamic changes, the purpose is to learn a cross-scene predictor, so that with the observation of one scene, a robot can onlinely predict the dynamic state of the other. A method is proposed to solve the problem via modeling dynamic correlation using latent space shared auto-encoders. Assuming that the inherent correlation of scene dynamics can be represented by shared latent space, where a common latent state is reached if the observations of both scenes are at an approximate time. A learning model is developed by connecting two auto-encoders through the latent space, and a prediction model is built by concatenating the encoder of the input scene with the decoder of the target one. Simulation datasets are generated imitating the dynamic flows at two adjacent gates of a campus, where the dynamic changes are triggered by a common working and teaching schedule. Similar scenarios can also be found at successive intersections on a single road, gates of a subway station, etc. Accuracy of cross-scene prediction is examined at various conditions of scene correlation and pairwise observations. Potentials of the proposed method are demonstrated by comparing with conventional end-to-end methods and linear predictions. Shaochi Hu, Donghao Xu, Huijing Zhao |
IROS | 3 |
| 2020 | SemanticPOSS: A Point Cloud Dataset with Large Quantity of Dynamic Instancesabstract3D semantic segmentation is one of the key tasks for autonomous driving system. Recently, deep learning models for 3D semantic segmentation task have been widely researched, but they usually require large amounts of training data. However, the present datasets for 3D semantic segmentation are lack of point-wise annotation, diversiform scenes and dynamic objects. In this paper1, we propose the SemanticPOSS dataset, which contains 2988 various and complicated LiDAR scans with large quantity of dynamic instances. The data is collected in Peking University and uses the same data format as SemanticKITTI. In addition, we evaluate several typical 3D semantic segmentation models on our SemanticPOSS dataset. Experimental results show that SemanticPOSS can help to improve the prediction accuracy of dynamic objects as people, car in some degree. SemanticPOSS will be published at www.poss.pku.edu.cn. Yancheng Pan, Biao Gao, Jilin Mei, Sibo Geng, Chengkun Li, Huijing Zhao |
IV | 6 |
| 2020 | Off-road Autonomous Vehicles Traversability Analysis and Trajectory Planning Based on Deep Inverse Reinforcement LearningabstractTerrain traversability analysis is a fundamental issue to achieve the autonomy of a robot at off-road environments. Geometry-based and appearance-based methods have been studied in decades, while behavior-based methods exploiting learning from demonstration (LfD) are new trends. Behavior-based methods learn cost functions that guide trajectory planning in compliance with experts' demonstrations, which can be more scalable to various scenes and driving behaviors. This research proposes a method of off-road traversability analysis and trajectory planning using Deep Maximum Entropy Inverse Reinforcement Learning. To incorporate the vehicle's kinematics while solving the problem of exponential increase of state-space complexity, two convolutional neural networks, i.e., RL ConvNet and Svf ConvNet, are developed to encode kinematics into convolution kernels and achieve efficient forward reinforcement learning. We conduct experiments in off-road environments. Scene maps are generated using 3D LiDAR data, and expert demonstrations are either the vehicle's real driving trajectories at the scene or synthesized ones to represent specific behaviors such as crossing negative obstacles. Different cost functions of traversability analysis are learned and tested at various scenes of capability in guiding the trajectory planning of different behaviors. We also demonstrate the peformance and computation efficiency of the proposed method. Ruoyu Sun 0003, Donghao Xu, Huijing Zhao |
IV | 5 |
| 2020 | Semantic Segmentation of 3D LiDAR Data in Dynamic Scene Using Semi-Supervised LearningabstractThis work studies the semantic segmentation of 3D LiDAR data in dynamic scenes for autonomous driving applications. A system of semantic segmentation using 3D LiDAR data, including range image segmentation, sample generation, inter-frame data association, track-level annotation, and semi-supervised learning, is developed. To reduce the considerable requirement of fine annotations, a CNN-based classifier is trained by considering both supervised samples with manually labeled object classes and pairwise constraints, where a data sample is composed of a segment as the foreground and neighborhood points as the background. A special loss function is designed to account for both annotations and constraints, where the constraint data are encouraged to be assigned to the same semantic class. A dataset containing 1838 frames of LiDAR data, 39 934 pairwise constraints and 57 927 human annotations is developed. The performance of the method is examined extensively. The qualitative and quantitative experiments show that the combination of a few annotations and large amount of constraint data significantly enhances the effectiveness and scene adaptability, resulting in greater than 10% improvement. Jilin Mei, Biao Gao, Donghao Xu, Xijun Zhao, Huijing Zhao |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2019 | Camera and LiDAR Fusion for On-road Vehicle Tracking with Reinforcement LearningabstractWe formulate camera and LiDAR fusion tracking as a sequential decision-making process. With our deep reinforcement learning framework, we try to optimize the tracking trajectory to be as accurate, smooth, and long as possible. In contrast to traditional fusion algorithms involving complex feature and strategy design and hyperparameters tuned for different scenarios, our fusion agent can learn the confidence of each input by tracking the results from raw observation in a data-driven fashion. Given the input states of different sensors, our approach chooses one input with a higher expected cumulative reward as the observation of a Kalman filter to iteratively predict the target position. The expected cumulative reward is estimated with a convolutional neural network, trained with a modified DQN algorithm, which takes inputs from both LiDAR and a camera. Through case studies and quantitative result evaluation on our dataset from the 4th Ring Road in Beijing, our algorithm is validated to achieve more accurate and robust tracking performance. Yongkun Fang, Huijing Zhao, Hongbin Zha, Xijun Zhao |
IV | 2 |
| 2019 | Off-Road Drivable Area Extraction Using 3D LiDAR DataabstractWe propose a method for off-road drivable area extraction using 3D LiDAR data with the goal of autonomous driving application. A specific deep learning framework is designed to deal with the ambiguous area, which is one of the main challenges in the off-road environment. To reduce the considerable demand for human-annotated data for network training, we utilize the information from vast quantities of vehicle paths and auto-generated obstacle labels. Using these auto-generated annotations, the proposed network can be trained using weakly supervised or semi-supervised methods, which can achieve better performance with fewer human annotations. The experiments on our dataset illustrate the reasonability of our framework and the validity of our weakly and semi-supervised methods. Biao Gao, Yancheng Pan, Xijun Zhao, Huijing Zhao |
IV | 6 |
| 2019 | Learning Scene Adaptive Covariance Error Model of LiDAR Scan Matching for Fusion Based LocalizationabstractLocalization is an essential technique for many robotic tasks such as mapping and navigation. Scan matching has been fused with other sensors to solve the problem at GPS restricted areas, where an accurate error model describing matching precision at various scenes is indispensable. We proposed an end-to-end method to learn a scene adaptive error model of LiDAR scan matching. A CNN (Convolutional Neural Network) is learnt to map from a LiDAR scan to an information matrix of the matching result, and a localization framework is proposed to fuse the results of LiDAR scan matching based on its error model. Experiments are conducted using both simulated and real world data, where the former is to validate the proposed method of its adaptability at various simple but typical scenes, while the later is to examine the method's practicability at real world environments. We demonstrate the performance of learning covariance error model, and examine the localization accuracy by comparing with other traditional methods. Efficiency of the proposed method is demonstrated. Xiaoliang Ju, Donghao Xu, Xijun Zhao, Huijing Zhao |
IV | 5 |
| 2019 | Supervised Learning for Semantic Segmentation of 3D LiDAR DataabstractThis work studies a supervised learning method using 3D LiDAR data for autonomous driving applications. A system of semantic segmentation, including range image segmentation, sample generation, track-level annotation and supervised learning, is developed. The formation and content of a data sample is studied intensively to address the specialty of 3D LiDAR data, which can be represented at a Cartesian or a 2D polar coordinate system, and composed of a segment as the foreground and/or the neighborhood points as the background. A CNN-based classifier is trained to map a given sample to an object label. Qualitative and quantitative experiments show that the background information and multiple feature map fusion significantly improve the performance of the classifier. Jilin Mei, Jiayu Chen 0006, Xijun Zhao, Huijing Zhao |
IV | 5 |
| 2019 | On-Road Vehicle Tracking Using Part-Based Particle FilterabstractIn this paper, we propose a part-based particle filter for on-road vehicle tracking. The proposed model combines a part-based strategy with a particle filter. By introducing a hidden state representing the center position of the vehicle, particles corresponding to vehicle parts sharing the same motion can be collectively updated in an efficient manner. By using a pre-trained appearance and geometric model, the tracker can distinguish parts with rich information from invalid parts to make more precise predictions. Meanwhile, some prior knowledge about the motion patterns of vehicles in a well-structured on-road environment is learned and can be used to infer measurement and motion models to improve tracking performance and efficiency. Experiments were conducted using the real data collected in Beijing to examine the performance of the method in different situations in terms of both its advantages and challenges. The collected Beijing highway dataset for on-road vehicle tracking will be made publicly available. We compare our method with the state-of-the-art approaches. The results demonstrate that the proposed algorithm is able to handle occlusion and the aspect ratio changes in the on-road vehicle tracking problem. Yongkun Fang, Chao Wang 0060, Xijun Zhao, Huijing Zhao, Hongbin Zha |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2019 | Aware of Scene Vehicles - Probabilistic Modeling of Car-Following Behaviors in Real-World TrafficabstractHeterogeneity exists in car following behaviors due to the driver's habit, fatigue, distraction, or surrounding traffic. This research proposes a method of modeling and reasoning heterogeneous car-following behaviors based on a stochastic system, where scene vehicles are involved explicitly in addition to the traditional leader-follower pair in characterizing driving situations, a hidden variable (driver state) is introduced to conjugate driving situations to heterogeneous models in predicting a driver's acceleration control, and the dynamic procedure is described using a dynamic Bayesian network. Experiments are conducted using a large set of naturalistic driving data that were collected by driving an instrumented vehicle on the multi-lane motorways in Beijing, where four distinctive driver states are learnt from data, characterizing the car-following procedure with normal, slow responsive, strong and prompt responsive, and unresponsive behavioral styles. By using the proposed scene-aware multi-state model for acceleration prediction, the error is reduced to 0.19 m/s2in average compared with 0.29 m/s2of a single-state model. Influence of scene vehicles on a driver state and subsequently on velocity control is verified based on the data. To the best of our knowledge, this is the first work that explicitly incorporates scene vehicles as influential factors in a probabilistic approach for modeling and reasoning heterogeneous car-following behaviors, and the performance is demonstrated on a large set of naturalistic driving data. Donghao Xu, Huijing Zhao, Franck Guillemard, Stéphane Géronimi, François Aioun |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | A Human-like Trajectory Planning Method by Learning from Naturalistic Driving DataabstractTrajectory planning has generally been framed as finding the lowest cost one from a set of trajectory candidates, where the cost function has been hand-crafted with carefully tuned parameters by experts. Such methods have technological feasibility of achieving vehicle autonomy, while the resultant behaviors could be much different with those of human drivers. This research proposes a humanlike trajectory planning method by learning from naturalistic driving data. A cost function is formulated by incorporating not only the components on comfort, efficiency and safety, but also lane incentive by referring to a human driver's lane change decisions. Coefficients of the cost components are learnt by correlating the probability of a trajectory being selected with its distance (i.e. similarity) to the human driven one at the same driving situation. A data set is developed by using the naturalistic data of human drivers on the motorways in Beijing, containing samples of lane changes to the left and right lanes, and car followings. Experiments are conducted on three aspects: 1) lane change trajectory planning to a given target lane; 2) lane change trajectory planning with simultaneous decision of a target lane; and 3) trajectory planning with simultaneous decision of maneuver. Promising results are presented. Donghao Xu, Huijing Zhao, Mathieu Moze, François Aioun, Franck Guillemard |
Intelligent Vehicles Symposium | 3 |
| 2018 | Naturalistic Lane Change Analysis for Human-Like Trajectory GenerationabstractHuman-like driving is of great significance for safety and comfort of autonomous vehicles, but existing trajectory planning methods for on-road vehicles rarely take the similarity with human behavior into consideration. From a representative trajectory-generation-based planning algorithm, this paper analyzes the systematic deviation of the generated trajectories from human trajectories, and proposes a new scheme of trajectory generation by compensating the deviation using a deviation profile learned from data. Experimental results show that the proposed trajectory generator is able to fit the human trajectories considerably better than the original one with only one additional degree of freedom. When used for online trajectory planning, with the same level of computational complexity, the proposed generator is able to generate trajectories that are more human-like than original generator does, which provides basis for autonomous vehicle to perform human-like trajectory planning. Donghao Xu, Zhezhang Ding, Huijing Zhao, Mathieu Moze, François Aioun, Franck Guillemard |
Intelligent Vehicles Symposium | 3 |
| 2018 | Scene-Adaptive Off-Road Detection Using a Monocular CameraabstractThis paper studies vision-based road detection for a robot's path following in off-road environments. We define the problem as detecting the region in front of the robot that is mechanically traversable (i.e., mechanical traversability), that is apt to be chosen by a human to drive through (i.e., human selection), and that extends for a distance to show the road's direction, shape, or even network of the intersection ahead (i.e., far-field capability). An algorithm framework is designed that contains two parts: inference and learning. In inference, the problem is formulated as a consecutive road type classification and road region segmentation to address the diversity of terrain surfaces. In model learning, the robot is first driven by a human being, with image samples on the track of the robot being collected that meet the prerequisites of both mechanical traversability and human selection. Evaluation measures are defined to examine the three requirements of mechanical traversability, human selection, and far-field capability. The performances of the above aspects are demonstrated on a data set using LiDAR, track and manual references, which will be released together with this publication. Jilin Mei, Huijing Zhao, Hongbin Zha |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Ego-centric traffic behavior understanding through multi-level vehicle trajectory analysisabstractThis study proposes a multi-level trajectory analysis method for modeling traffic behavior from an ego-centric view, where on-road vehicle trajectories are collected based on the authors' previous studies of an on-board system consisting of multiple 2D lidar sensors. From an input set of trajectories, a set of hot regions (topics) that trajectory points most frequently present are first discovered using a sticky HDP-HMM; then, the major paths of the trajectories' transitions across different hot regions are extracted by recursively mining frequent subsequences of topics; and finally, paths are modeled using a hierarchical hidden Markov model (HHMM), where the intra-path dynamics is represented using an HMM, in which each state corresponds to a hot region, while the inter-path transition is assumed to be Markovian. The model could be used for behavior prediction, i.e. whenever a vehicle is detected in a scene, predicting which route it will probably follow and how its trajectory will probably develop over time, which is essential to interpreting the potential risks for longer time horizons. Experiments are conducted using a large set of vehicle trajectories collected from motorways in Beijing, and promising results are presented. Donghao Xu, Huijing Zhao, Jinshi Cui, Hongbin Zha, Franck Guillemard, Stéphane Géronimi, François Aioun |
ICRA | 3 |
| 2017 | On-road vehicle tracking using part-based particle filterabstractIn this paper, we propose a part-based particle filter for on-road vehicle tracking. The proposed model takes part-based strategies into account in a particle filter. By introducing a hidden state vehicle center position, vehicle parts particles can be updated efficiently as a whole sharing same motion. With a pre-trained appearance and geometric model, tracker can distinguish parts with rich information from invalid parts to make a more precise prediction. Meanwhile some priori knowledge about the moving pattern of vehicles in well-structured on-road environment is learned, and can be used in the inference of measurement model and motion model to improve tracking performance and efficiency. Experiments were conducted with real data collected in Beijing to examine the performance in different situations on both the advantages and challenges. The Beijing highway dataset for on-road vehicle tracking will be opened to the society. We compare our method with the state-of-the-art approaches. Result demonstrate that the proposed algorithm are able to handle occlusion and aspect ratio change in on-road vehicle tracking problem. Yongkun Fang, Chao Wang 0060, Huijing Zhao, Hongbin Zha |
IROS | 3 |
| 2017 | Scene-aware driver state understanding in car-following behaviorsabstractThis research represents the heterogeneity in car following by a hidden variable driver state, which could change due to the driver's habit, fatigue, distraction, influence of surrounding traffic etc, resulting in the heterogeneous behaviors of such as fast or slow, strong or weak response to the same level of stimuli. A probabilistic method of driver state understanding is proposed by modeling and reasoning the heterogeneity in car-following behaviors, and the influence of surrounding traffic is addressed explicitly in addition to the leader-follower pair aiming at applications in crowded real-world traffic. Experiments are conducted by using the on-road trajectory data that were collected from motorways in Beijing, where four distinctive driver states and corresponding car-following models are learnt. With online understanding of driver state, the particular car-following model is used to predict the drivers velocity control, where results of improved accuracy are demonstrated. Donghao Xu, Huijing Zhao, Franck Guillemard, Stéphane Géronimi, François Aioun |
Intelligent Vehicles Symposium | 2 |
| 2017 | On-Road Vehicle Trajectory Collection and Scene-Based Lane Change Analysis: Part IIabstractThis two-part paper aims to study lane change behaviors at the tactical level from an on-road perspective. Compared with longitudinal driving tasks, a lane change is more complicated because this task has more interactions with surrounding vehicles; thus, there are more potential risks during this procedure. Based on the results from Part I on an on-road vehicle trajectory collection, this part investigates lane change extraction and scene-based behavior analysis, and it has a particular focus on understanding the interactions between an ego and surrounding vehicles during the procedure. We claim that this paper provides the following novel contributions: 1) an automatic method is proposed for extracting lane change segments from a continuous driving sequence by modeling and recognizing patterns in a steering angle; 2) a lane change database at the trajectory level is generated, which reflects the interactions between an ego and the surrounding vehicles during the procedures; and 3) we present findings from analyzing lane change procedures using real-world data on the axes of both the ego's trajectory and interactions with the scene vehicles. To the authors' knowledge, this is the first lane change behavior study from an on-road perspective that addresses the vehicle interactions in real-world traffic at the trajectory level. Qiqi Zeng, Yuping Lin, Donghao Xu, Huijing Zhao, Franck Guillemard, Stéphane Géronimi, François Aioun |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | On-Road Vehicle Trajectory Collection and Scene-Based Lane Change Analysis: Part IabstractThis two-part paper aims to study lane change behaviors at the tactical level from an on-road perspective, with a special focus on analyzing the interactions between an ego and surrounding vehicles during the procedure. Part I addresses vehicle trajectory collection, whereas Part II addresses lane change extraction and scene-based behavioral analysis. Different from the general technique of moving object detection and tracking, trajectory collection for tactical driving behavior study is required to have the properties of consistency, completeness, continuity, and accuracy. This paper proposes a system of on-road vehicle trajectory collection, where an instrumented vehicle is developed with multiple horizontal 2-D lidars that have 360° coverage. The software is developed by fitting the laser points of all lidars on a vehicle model using a coupled estimation of features and reliability along frames to achieve accurate state estimations of occluded data and robust data association in multiviewpoint sensing. The performance is investigated extensively, and a large trajectory set is developed through on-road driving at the Fourth Ring Road in Beijing for a total distance of 64 km, with more than 5700 environmental trajectories with a total length of over 19 h. The performance is demonstrated to be of high quality in terms of the required properties. To the authors' knowledge, this is the first system that is able to automatically collect all-around vehicle trajectories during on-road driving and to demonstrate good performance in providing a high-quality database for driving behavior studies from an on-road perspective that addresses vehicle interactions in real-world traffic at the trajectory level. Huijing Zhao, Chao Wang 0060, Yuping Lin, Franck Guillemard, Stéphane Géronimi, François Aioun |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Multimodal information fusion for urban scene understanding
Philippe Xu, Franck Davoine, Jean-Baptiste Bordes, Huijing Zhao, Thierry Denoeux |
Mach. Vis. Appl. | 4 |
| 2016 | Object Discovery: Soft Attributed Graph MiningabstractWe categorize this research in terms of its contribution to both graph theory and computer vision. From the theoretical perspective, this study can be considered as the first attempt to formulate the idea of mining maximal frequent subgraphs in the challenging domain of messy visual data, and as a conceptual extension to the unsupervised learning of graph matching. We define a soft attributed pattern (SAP) to represent the common subgraph pattern among a set of attributed relational graphs (ARGs), considering both their structure and attributes. Regarding the differences between ARGs with fuzzy attributes and conventional labeled graphs, we propose a new mining strategy that directly extracts the SAP with the maximal graph size without applying node enumeration. Given an initial graph template and a number of ARGs, we develop an unsupervised method to modify the graph template into the maximal-size SAP. From a practical perspective, this research develops a general platform for learning the category model (i.e., the SAP) from cluttered visual data (i.e., the ARGs) without labeling "what is where," thereby opening the possibility for a series of applications in the era of big visual data. Experiments demonstrate the superior performance of the proposed method on RGB/RGB-D images and videos. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Probabilistic Inference for Occluded and Multiview On-road Vehicle DetectionabstractVisual-based approaches have been extensively studied for on-road vehicle detection; however, it faces great challenges as the visual appearance of a vehicle may greatly change across different viewpoints and as a partial observation sometimes happens due to occlusions from infrastructure or scene dynamics and/or a limited camera vision field. This paper presents a visual-based on-road vehicle detection algorithm for a multilane traffic scene. A probabilistic inference framework based on part models is proposed to overcome the challenges from a multiview and partial observation. Geometric models are learned for each dominant viewpoint to describe the configuration of vehicle parts and their spatial relations in probabilistic representations. Viewpoint maps are generated based on the knowledge of the road structure and driving patterns, which provide a prediction of the viewpoints of a vehicle whenever it happens at a certain location. Extensive experiments are conducted using an onboard camera on multilane motor ways in Beijing. A large-scale data set that contains more than 30 000 labeled ground truths for both fully and partially observed vehicles in different viewpoints across various traffic density scenes is developed. The data set will be opened to the society together with this publication. Chao Wang 0060, Yongkun Fang, Huijing Zhao, Chunzhao Guo, Seiichi Mita, Hongbin Zha |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | Visual-based on-road vehicle detection: A transnational experiment and comparisonabstractAs a key technique in ADAS (Advanced Driving Assistant System) or autonomous driving systems, visual-based on-road vehicle detection has been studied widely, while it faces still great challenges, among which are the complexity, diversity and unpredictable changes of the real-world environments. In the authors' previous work, an algorithm was developed in a probabilistic inference framework with its focus on solving the multi-view and occlusion problems at multi-lane motor way scenes. In this research, we seek to answer the questions: how efficient is the system during a long-term operation across a large area of changed conditions? To this end, a large scale experiment is conducted, where three testing data sets are developed containing the samples of more than 30,000 on Beijing's ring roads, 800 on Nagoya's fast road, and 3,000 on Nagoya's downtown streets, and the performance of visual-based vehicle detection concerning the multi-view and occlusion problems across extensive regions and at transnational environments are studied. We present our preliminary findings in this paper, which leads to a more extensive study in future work. Chao Wang 0060, Huijing Zhao, Chunzhao Guo, Seiichi Mita, Hongbin Zha |
Intelligent Vehicles Symposium | 2 |
| 2015 | From RGB-D Images to RGB Images: Single Labeling for Mining Visual ModelsabstractMining object-level knowledge, that is, building a comprehensive category model base, from a large set of cluttered scenes presents a considerable challenge to the field of artificial intelligence. How to initiate model learning with the least human supervision (i.e., manual labeling) and how to encode the structural knowledge are two elements of this challenge, as they largely determine the scalability and applicability of any solution. In this article, we propose a model-learning method that starts from a single-labeled object for each category, and mines further model knowledge from a number of informally captured, cluttered scenes. However, in these scenes, target objects are relatively small and have large variations in texture, scale, and rotation. Thus, to reduce the model bias normally associated with less supervised learning methods, we use the robust 3D shape in RGB-D images to guide our model learning, then apply the properly trained category models to both object detection and recognition in more conventional RGB images. In addition to model training for their own categories, the knowledge extracted from the RGB-D images can also be transferred to guide model learning for a new category, in which only RGB images without depth information in the new category are provided for training. Preliminary testing shows that the proposed method performs as well as fully supervised learning methods. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2014 | When 3D Reconstruction Meets Ubiquitous RGB-D Imagesabstract3D reconstruction from a single image is a classical problem in computer vision. However, it still poses great challenges for the reconstruction of daily-use objects with irregular shapes. In this paper, we propose to learn 3D reconstruction knowledge from informally captured RGB-D images, which will probably be ubiquitously used in daily life. The learning of 3D reconstruction is defined as a category modeling problem, in which a model for each category is trained to encode category-specific knowledge for 3D reconstruction. The category model estimates the pixel-level 3D structure of an object from its 2D appearance, by taking into account considerable variations in rotation, 3D structure, and texture. Learning 3D reconstruction from ubiquitous RGB-D images creates a new set of challenges. Experimental results have demonstrated the effectiveness of the proposed approach. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
CVPR | 4 |
| 2014 | Attributed Graph Mining and Matching: An Attempt to Define and Extract Soft Attributed PatternsabstractGraph matching and graph mining are two typical areas in artificial intelligence. In this paper, we define the soft attributed pattern (SAP) to describe the common subgraph pattern among a set of attributed relational graphs (ARGs), considering both the graphical structure and graph attributes. We propose a direct solution to extract the SAP with the maximal graph size without node enumeration. Given an initial graph template and a number of ARGs, we modify the graph template into the maximal SAP among the ARGs in an unsupervised fashion. The maximal SAP extraction is equivalent to learning a graphical model (i.e. an object model) from large ARGs (i.e. cluttered RGB/RGB-D images) for graph matching, which extends the concept of "unsupervised learning for graph matching." Furthermore, this study can be also regarded as the first known approach to formulating "maximal graph mining" in the graph domain of ARGs. Our method exhibits superior performance on RGB and RGB-D images. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
CVPR | 4 |
| 2014 | Calibration method for multiple 2D LIDARs systemabstractMany robotic and mobile mapping systems have been developed using multiple 2D LIDARs (briefly multi-LIDAR system) to sense environment. In such systems, extrinsic calibration of all LIDARs is essential for making collaborative use of the data from different sensors. This research aims at developing a calibration method for multi-LIDAR systems at the general scene, such as an outdoor place or an underground parking-lot, without modification to environment by putting calibration targets. In this paper, the calibration method is proposed by aligning the 3D data of different LIDARs. They are concerned at two-levels: 1) reference calibration, i.e. finding the transformation from a reference LIDAR to the platform frame; 2) multi-LIDAR calibration, i.e. finding the LIDARs' relative geometries by referring to the reference one. The method is examined in calibrating the multiple 2D LIDARs on an intelligent vehicle platform POSS-V, where the data collected through a driving in an underground parking-lot are registered to find sensors' geometry. Calibration accuracy is examined by comparing with a CAD model of the scene, which was measured by using a total station. Mengwen He, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ICRA | 2 |
| 2014 | Start from minimum labeling: Learning of 3D object models and point labeling from a large and complex environmentabstractA large category model base can provide object-level knowledge for various perception tasks of the intelligent vehicle system. The automatic and efficient construction of such a model base is highly desirable but challenging. This paper presents a novel semi-supervised approach to discover possible prototype models of 3D object structures from the point cloud of a large and complex environment, given a limited number of seeds in an object category. Our method incrementally trains the models while simultaneously collecting object samples. Considering the bias problem of model learning caused by bias accumulation in a sample collection, we propose to gradually differentiate the standard category model into several sub-category models to represent different intra-category structural styles. Thus, new sub-categories are discovered and modeled, old models are improved, and redundant models for similar structures are deleted iteratively during the learning process. This multiple-model strategy provides several interactive options for the category boundary to deal with the bias problem. Experimental results demonstrate the effectiveness and high efficiency of our approach to model mining from “big point cloud data”. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ICRA | 4 |
| 2014 | On-road vehicle detection through part model learning and probabilistic inferenceabstractVisual based approach has been studied extensively for on-road vehicle detection, while it faces great challenges, as visual appearance of a vehicle may change greatly across different viewpoints, and partial observation happens sometime due to occlusions from infrastructure or scene dynamics, and/or limited camera vision field. Inspired by the works on part-based detection, this research proposes a probabilistic framework for on-road vehicle detection, where focus is cast on vehicle pose inference on the set of part instances by addressing the issues of partial observation and varying viewpoints. To this end, geometric models describing the configuration of vehicle parts as well as their spatial relations in probabilistic representations are learned for each dominant viewpoint, and viewpoint maps are generated on each typical road structure, which provide probabilistic prediction to the viewpoints of a vehicle at each location at ego frame. Experiments have been conducted using a data set that was developed in the authors' previous work on the ring roads in Beijing. Viewpoint-discriminative part appearance models (VDPAM) and viewpoint-discriminative part-based geometric models (VDPGM) are learned on the image samples of the data set, and the road structure-based probabilistic viewpoint maps (RSPVM) are generated by taking the statistics of the Lidar-based vehicle detection results. On-road vehicle detection is examined using an on-road video stream that has been labelled with ground truth. Experimental results are presented and efficiency on detecting the partially observed vehicles on varying viewpoints is demonstrated. Chao Wang 0060, Huijing Zhao, Chunzhao Guo, Seiichi Mita, Hongbin Zha |
IROS | 2 |
| 2014 | Monocular visual localization using road structural featuresabstractPrecise localization is an essential issue for autonomous driving applications, where GPS-based systems are challenged to meet requirements such as lane-level accuracy. This paper introduces a new visual-based localization approach in dynamic traffic environments, focusing on and exploiting properties of structured roads like straight roads or intersections. Such environments show several line segments on lane markings, curbs, poles, building edges, etc., which demonstrate the road's longitude, latitude and vertical directions. Based on this observation, we define a Road Structural Feature (RSF) as sets of segments along three perpendicular axes together with feature points. At each video frame, the proper road structure (or multiple road structures in case of an intersection) is predicted based on the geometric information given by a 2D map. The RSF is then detected from line segments and points extracted from the image, and used to estimate the pose of the vehicle. Experiments are conducted using video streams collected on major roads in downtown Beijing, which are structured and with intense dynamic traffic. GPS/IMU data have been collected and synchronized with the video streams as a reference in validation. The results show good performance compared with that of a more traditional visual odometry method. Future work will be addressed on using visual approach to improve GPS localization accuracy. Huijing Zhao, Franck Davoine, Jinshi Cui, Hongbin Zha |
Intelligent Vehicles Symposium | 2 |
| 2014 | Efficient Closed-Loop Multiple-View RegistrationabstractRegistering multiple views is an essential and challenging problem for many intelligent transportation applications that employ a mobile sensing platform or consist of multiple stationary sensors. In this paper a novel algorithm is presented for multiple-view registration under a loop closure constraint. Different from most existing methods, which use general optimization techniques, our method studies the mechanism of adjusting the poses of views in a loop and provides a highly efficient and accurate solution. We prove that translation vectors can be decoupled if the same point set is used in each view to associate the previous and subsequent views, leading to our solution for such decouplable cases. If this condition does not hold, an exact solution of translation vectors is provided when rotation parameters are given, which results in our iterative solution for general cases by updating rotation and translation alternately. In our method, the effect of the accumulated pose error in a loop can be distributed to all views efficiently through loop factors, and only a few iterations are needed. Most important of all, in each iteration our method has linear computational complexity with respect to the number of views, which is much superior to that of state-of-the-art methods. A series of experiments was conducted, involving simulation of thousands of views and real vehicle-borne sensing data that include 65 371 point pairs in 352 views. Experimental results show that our proposed method is not only stable and highly efficient but also provides competitive accuracy relative to existing methods. Xiaowei Shao, Huijing Zhao, Xuelong Li 0001, Ryosuke Shibasaki |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2013 | Category Modeling from Just a Single Labeling: Use Depth Information to Guide the Learning of 2D ModelsabstractAn object model base that covers a large number of object categories is of great value for many computer vision tasks. As artifacts are usually designed to have various textures, their structure is the primary distinguishing feature between different categories. Thus, how to encode this structural information and how to start the model learning with a minimum of human labeling become two key challenges for the construction of the model base. We design a graphical model that uses object edges to represent object structures, and this paper aims to incrementally learn this category model from one labeled object and a number of casually captured scenes. However, the incremental model learning may be biased due to the limited human labeling. Therefore, we propose a new strategy that uses the depth information in RGBD images to guide the model learning for object detection in ordinary RGB images. In experiments, the proposed method achieves superior performance as good as the supervised methods that require the labeling of all target objects. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao |
CVPR | 5 |
| 2013 | Learning Graph Matching: Oriented to Category Modeling from Cluttered ScenesabstractAlthough graph matching is a fundamental problem in pattern recognition, and has drawn broad interest from many fields, the problem of learning graph matching has not received much attention. In this paper, we redefine the learning of graph matching as a model learning problem. In addition to conventional training of matching parameters, our approach modifies the graph structure and attributes to generate a graphical model. In this way, the model learning is oriented toward both matching and recognition performance, and can proceed in an unsupervised fashion. Experiments demonstrate that our approach outperforms conventional methods for learning graph matching. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ICCV | 4 |
| 2013 | Unsupervised 3D category discovery and point labeling from a large urban environmentabstractThe building of an object-level knowledge base is the foundation of a new methodology for many perception tasks in artificial intelligence, and is an area that has received increasing attention in recent years. In this paper, we propose, for the first time, to mine category shape patterns directly from a large urban environment, thus constructing a category structure base. Conventionally, category patterns are learned from a large collection of object samples, but automatic object collection requires prior knowledge of category structures. To solve this chicken-and-egg problem, we learn shape patterns from raw segmentations, and then refine these segmentations based on the pattern knowledge. In the process, we solve two challenging problems of knowledge mining. First, as some categories have large intra-category structure variations, we design an entropy-based method to determine the structure variation for each category, in order to establish the correct range of sample collection. Second, because incorrect segmentation is unavoidable without prior knowledge, we propose a novel unsupervised method that uses a pattern competition strategy to identify and subtract shape patterns formed by incorrectly segmented objects. This ensures that shape patterns are meaningful at the object level. Experimental results demonstrated the effectiveness of the proposed method for category structure mining in a large urban environment. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ICRA | 4 |
| 2013 | Pairwise LIDAR calibration using multi-type 3D geometric features in natural sceneabstractIt has become a well-known technology that 3D measurement of a large environment could be achieved by using a number of 2D LIDARs on a mobile platform. In such a system, calibration is essential for making collaborative use of different LIDAR data, while existing methods usually require modifications to the environments, such as putting calibration targets, or rely on special facilities, which is labor intensive and put many restrictions to potential applications. This research aims at developing a calibration method for multiple 2D LIDAR sensing systems, which could be conducted in a general outdoor environment using the features of a nature scene. Special focus is cast on solving the noisy sensing in a complex environment and the occlusions caused by largely different sensor viewpoints. A multi-type geometric feature based calibration algorithm is proposed, which extracts the features such as points, lines, planes and quadrics from the 3D points of each LIDAR sensing. Transformation parameters from each sensor to the frame on a moving platform is estimated by matching the multi-type features. Experiments are conducted using the data sets of an intelligent vehicle platform (POSS-V) through a driving in the campus of Peking University. Results of calibrating two LIDAR sensors with largely different viewpoints are presented, and the accuracy and robustness concerning noisy feature extractions are examined intensively. Mengwen He, Huijing Zhao, Franck Davoine, Jinshi Cui, Hongbin Zha |
IROS | 2 |
| 2013 | Lane change trajectory prediction by using recorded human driving dataabstractBeing able to predict the trajectory of a human driver's potential lane change behavior in urban high way scenario is crucial for lane change risk assessment task. A good prediction of the driver's lane change trajectory makes it possible to evaluate the risk and warn the driver beforehand. Rather than generating such a trajectory only using a mathematical model, this paper develops a lane change trajectory prediction approach based on real human driving data stored in a database. In real-time, the system generates parametric trajectories by interpolating k human lane change trajectory instances from the pre-collected database that are similar to the current driving situation. In order to build this real lane change database, a human lane change data collection vehicle platform is developed. Extensive experiments have been carried out in urban highway environments to build a significant database with more than 200 lane changes. Real results show that this approach produces lane change trajectories that are quite similar to real ones which makes it suitable to predict humanlike lane change maneuvers. Huijing Zhao, Philippe Bonnifait, Hongbin Zha |
Intelligent Vehicles Symposium | 2 |
| 2013 | Laser-based tracking of multiple interacting pedestrians via on-line learning
Xuan Song 0001, Jinshi Cui, Huijing Zhao, Hongbin Zha, Ryosuke Shibasaki |
Neurocomputing | 3 |
| 2013 | Unsupervised skeleton extraction and motion capture from 3D deformable matching
Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao |
Neurocomputing | 5 |
| 2013 | A fully online and unsupervised system for large and high-density area surveillance: Tracking, semantic scene learning and abnormality detectionabstractFor reasons of public security, an intelligent surveillance system that can cover a large, crowded public area has become an urgent need. In this article, we propose a novel laser-based system that can simultaneously perform tracking, semantic scene learning, and abnormality detection in a fully online and unsupervised way. Furthermore, these three tasks cooperate with each other in one framework to improve their respective performances. The proposed system has the following key advantages over previous ones: (1) It can cover quite a large area (more than 60×35m), and simultaneously perform robust tracking, semantic scene learning, and abnormality detection in a high-density situation. (2) The overall system can vary with time, incrementally learn the structure of the scene, and perform fully online abnormal activity detection and tracking. This feature makes our system suitable for real-time applications. (3) The surveillance tasks are carried out in a fully unsupervised manner, so that there is no need for manual labeling and the construction of huge training datasets. We successfully apply the proposed system to the JR subway station in Tokyo, and demonstrate that it can cover an area of 60×35m, robustly track more than 150 targets at the same time, and simultaneously perform online semantic scene learning and abnormality detection with no human intervention. Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2013 | An online system for multiple interacting targets tracking: Fusion of laser and vision, tracking and learningabstractMultitarget tracking becomes significantly more challenging when the targets are in close proximity or frequently interact with each other. This article presents a promising online system to deal with these problems. The novelty of this system is that laser and vision are integrated with tracking and online learning to complement each other in one framework: when the targets do not interact with each other, the laser-based independent trackers are employed and the visual information is extracted simultaneously to train some classifiers online for “possible interacting targets”. When the targets are in close proximity, the classifiers learned online are used alongside visual information to assist in tracking. Therefore, this mode of cooperation not only deals with various tough problems encountered in tracking, but also ensures that the entire process can be completely online and automatic. Experimental results demonstrate that laser and vision fully display their respective advantages in our system, and it is easy for us to obtain a good trade-off between tracking accuracy and the time-cost factor. Xuan Song 0001, Huijing Zhao, Jinshi Cui, Xiaowei Shao, Ryosuke Shibasaki, Hongbin Zha |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | Tracking Generic Human Motion via Fusion of Low- and High-Dimensional ApproachesabstractTracking generic human motion is highly challenging due to its high-dimensional state space and the various motion types involved. In order to deal with these challenges, a fusion formulation which integrates low- and high-dimensional tracking approaches into one framework is proposed. The low-dimensional approach successfully overcomes the high-dimensional problem of tracking the motions with available training data by learning motion models, but it only works with specific motion types. On the other hand, although the high-dimensional approach may recover the motions without learned models by sampling directly in the pose space, it lacks robustness and efficiency. Within the framework, the two parallel approaches, low- and high-dimensional, are fused via a probabilistic approach at each time step. This probabilistic fusion approach ensures that the overall performance of the system is improved by concentrating on the respective advantages of the two approaches and resolving their weak points. The experimental results, after qualitative and quantitative comparisons, demonstrate the effectiveness of the proposed approach in tracking generic human motion. Jinshi Cui, Ye Liu 0002, Yuandong Xu, Huijing Zhao, Hongbin Zha |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2012 | Fusion of low-and high-dimensional approaches by trackers sampling for generic human motion tracking
Ye Liu 0002, Jinshi Cui, Huijing Zhao, Hongbin Zha |
ICPR | 3 |
| 2012 | Laser-based intelligent surveillance and abnormality detection in extremely crowded scenariosabstractAbnormal activity detection plays a crucial role in surveillance applications, and a surveillance system that can perform robustly in the extremely crowded area has become an urgent need for public security. In this paper, we propose a novel laser-based system which can simultaneously perform the tracking, semantic scene learning and abnormality detection in the large and crowded environment. In our system, a novel abnormality detection model is proposed, and it considers and combines various factors that will influence human activity. Moreover, this model intensively investigate the relationship between pedestrians' social behaviors and their walking scenarios. We successfully applied the proposed system to the JR subway station of Tokyo, which can cover a 60×35m area, robustly track more than 180 targets at the same time and simultaneously perform the online semantic scene learning and abnormality detection with no human intervention. Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Hongbin Zha |
ICRA | 5 |
| 2012 | A real-time motion planner with trajectory optimization for autonomous vehiclesabstractIn this paper, an efficient real-time autonomous driving motion planner with trajectory optimization is proposed. The planner first discretizes the plan space and searches for the best trajectory based on a set of cost functions. Then an iterative optimization is applied to both the path and speed of the resultant trajectory. The post-optimization is of low computational complexity and is able to converge to a higher-quality solution within a few iterations. Compared with the planner without optimization, this framework can reduce the planning time by 52% and improve the trajectory quality. The proposed motion planner is implemented and tested both in simulation and on a real autonomous vehicle in three different scenarios. Experiments show that the planner outputs high-quality trajectories and performs intelligent driving behaviors. Wenda Xu, Junqing Wei, John M. Dolan, Huijing Zhao, Hongbin Zha |
ICRA | 4 |
| 2012 | Computing object-based saliency in urban scenes using laser sensingabstractIt becomes a well-known technology that a low-level map of complex environment containing 3D laser points can be generated using a robot with laser scanners. Given a cloud of 3D laser points of an urban scene, this paper proposes a method for locating the objects of interest, e.g. traffic signs or road lamps, by computing object-based saliency. Our major contributions are: 1) a method for extracting simple geometric features from laser data is developed, where both range images and 3D laser points are analyzed; 2) an object is modeled as a graph used to describe the composition of geometric features; 3) a graph matching based method is developed to locate the objects of interest on laser data. Experimental results on real laser data depicting urban scenes are presented; efficiency as well as limitations of the method are discussed. Yipu Zhao, Mengwen He, Huijing Zhao, Franck Davoine, Hongbin Zha |
ICRA | 3 |
| 2012 | A system of automated training sample generation for visual-based car detectionabstractThis paper presents a system to automatically generate car sample dataset for visual-based car detector training. The dataset contains multi-view car samples labeled with the car's pose, so that a view-discriminative training and car detection is also available. There are mainly two parts in the system: laser-based car detection and tracking generates motion trajectories of on-road cars, and then visual samples are extracted by fusing the detection and tracking results with visual-based detection. A multi-modal sensor system is developed for the omni-directional data collection on a test-bed vehicle. By processing the data of experiment conducted on the freeway of Beijing, a large number of multi-view car samples with pose information were generated. The samples' quality is evaluated by applying it in a visual car detector's training and testing procedure. Chao Wang 0060, Huijing Zhao, Franck Davoine, Hongbin Zha |
IROS | 2 |
| 2012 | Learning lane change trajectories from on-road driving dataabstractLane change is one of the most principle driving behaviors on structure roads. It frequently happens in daily driving. A key issue in lane change technique is trajectory planning, where a set of trajectories describing possible vehicle motions are generated by applying a parametric function, and by uniformly sampling the end states in configuration space; the trajectories are then examined to find an optimal one for execution. However, such a trajectory set has poor efficiency due to the large sample number. Many trajectories in this set seldom happen in real human driving behaviors. In this research, lane change trajectories are collected from real driving data of different drivers. Their statistics are analyzed, through which, a simplified trajectory set is generated. Experiment results show that the trajectory set has much less number of samples but can still guarantee to cover usual lane change behaviors of human being. Huijing Zhao, Franck Davoine, Hongbin Zha |
Intelligent Vehicles Symposium | 2 |
| 2012 | Omni-directional detection and tracking of on-road vehicles using multiple horizontal laser scannersabstractThis research aims at generating an omnidirectional perception at the host vehicle's surroundings, extracting accurate and continuous motion trajectories of the nearby vehicles using low cost laser scanners. A system of detecting and tracking on-road vehicles using multiple laser scanners is developed, where focuses are cast on solving data association of simultaneous measurements from multiple sensors at different viewpoints, and state estimation in case of partial observations in dense dynamic situations. Experimental results in freeways in Beijing are presented, system efficiency is demonstrated, where motion trajectories describing driving behaviors such as overtaking, lane changing and other interactions between driving objects are captured. In addition, the accuracy in vehicle detection and tracking is examined using a reference vehicle with a ground truth GPS. Huijing Zhao, Chao Wang 0060, Franck Davoine, Jinshi Cui, Hongbin Zha |
Intelligent Vehicles Symposium | 1 |
| 2012 | Detection and Tracking of Moving Objects at Intersections Using a Network of Laser ScannersabstractIn our previous work, we reported a system that monitors an intersection using a network of horizontal laser scanners. This paper focuses on an algorithm for moving-object detection and tracking, given a sequence of distributed laser scan data of an intersection. The goal is to detect each moving object that enters the intersection; estimate state parameters such as size; and track its location, speed, and direction while it passes through the intersection. This work is unique, to the best of the authors' knowledge, in that the data is novel, which provides new possibilities but with great challenges; the algorithm is the first proposal that uses such data in detecting and tracking all moving objects that pass through a large crowded intersection with focus on achieving robustness to partial observations, some of which result from occlusions, and on performing correct data associations in crowded situations. Promising results are demonstrated using experimental data from real intersections, whereby, for 1063 objects moving through an intersection over 20 min, 988 are perfectly tracked from entrance to exit with an excellent tracking ratio of 92.9%. System advantages, limitations, and future work are discussed. Huijing Zhao, Jie Sha, Yipu Zhao, Junqiang Xi, Jinshi Cui, Hongbin Zha, Ryosuke Shibasaki |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | TripVista: Triple Perspective Visual Trajectory Analytics and its application on microscopic traffic data at a road intersectionabstractIn this paper, we present an interactive visual analytics system, Triple Perspective Visual Trajectory Analytics (TripVista), for exploring and analyzing complex traffic trajectory data. The users are equipped with a carefully designed interface to inspect data interactively from three perspectives (spatial, temporal and multi-dimensional views). While most previous works, in both visualization and transportation research, focused on the macro aspects of traffic flows, we develop visualization methods to investigate and analyze microscopic traffic patterns and abnormal behaviors. In the spatial view of our system, traffic trajectories with various presentation styles are directly interactive with user brushing, together with convenient pattern exploration and selection through ring-style sliders. Improved ThemeRiver, embedded with glyphs indicating directional information, and multiple scatterplots with time as horizontal axes illustrate temporal information of the traffic flows. Our system also harnesses the power of parallel coordinates to visualize the multi-dimensional aspects of the traffic trajectory data. The above three view components are linked closely and interactively to provide access to multiple perspectives for users. Experiments show that our system is capable of effectively finding both regular and abnormal traffic flow patterns. Hanqi Guo 0001, Zuchao Wang, Huijing Zhao, Xiaoru Yuan |
PacificVis | 4 |
| 2011 | Tracking Generic Human Motion via Fusion of Low- and High-Dimensional Approaches
Yuandong Xu, Jinshi Cui, Huijing Zhao, Hongbin Zha |
BMVC | 3 |
| 2011 | A novel laser-based system: Fully online detection of abnormal activity via an unsupervised methodabstractAbnormal activity detection plays a crucial role in surveillance applications, and such system has become an urgent need for public security. In this paper, we propose a novel laser-based system, which can perform the online detection of abnormal activity with an unsupervised way. The proposed system has the following key features that make it advantageous over previous ones: (1) It can cover quite a large and crowded area, such as subway station, public square, intersection and etc. (2) The overall system can vary with time period, incrementally learn the behavior pattern of pedestrians and perform the fully online detection of abnormal activity. This feature makes our system be quite suitable for the real-time applications. (3) The abnormal activity detection is carried out with a fully unsupervised way, there is no need for manual labelling and constructing the huge training datasets. We successfully applied the proposed system into the JR subway station of Tokyo, which can cover a 60×35m area, track more 150 targets at the same time and simultaneously perform the robust detection of abnormal activity with no human intervention. Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ICRA | 4 |
| 2011 | A vehicle model for micro-traffic simulation in dynamic urban scenariosabstractIn order to improve energy efficiency of transport systems, eco-driving strategies are studied world-widely. However, most literatures on eco-driving based on traditional traffic flow models, are greatly simplified, and can not evaluate the effects on detailed driving behaviors. By referring to robot motion planning approaches, in this research a microscopic vehicle model is developed and it can represent different driving behaviors, such as aggressive or conservative driving; a collision detection algorithm is proposed that takes O(1) time to check for a trajectory's collision, enabling realtime planning; and a traffic simulation system is developed by incorporating traffic rules, so that the driving behaviors such as observing or not observing traffic rules can also be represented. Experiments are conducted on the simulation platform, and the performance of different driving behaviors on travel time, mileage, comfort and eco is studied. Wenda Xu, Wen ZhaYao, Huijing Zhao, Hongbin Zha |
ICRA | 3 |
| 2011 | 3D crowd surveillance and analysis using laser range scannersabstractIn this study, we present a novel system for crowd surveillance and quantified analysis based on laser range scanners. By mounting a laser scanner at a swinging platform, the spatial information of passengers inside the area of interest can be reconstructed in a form of 3D points. Multiple laser scanners are integrated together by semi-auto calibration procedures. Background map is generated through histogram analysis of scan maps, and is further applied for 3D moving object detection. An improved version of mean-shift clustering algorithm is proposed to extract individual passengers efficiently. In addition, quantified crowdness analysis is conducted from different aspects to indicate the situation inside the surveillance area according to the extraction results of passengers. The proposed system was tested in a central subway station in Tokyo and experimental results demonstrate the effectiveness of our proposed system. Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki, Kiyoshi Sakamoto |
IROS | 2 |
| 2010 | An online approach: Learning-Semantic-Scene-by-Tracking and Tracking-by-Learning-Semantic-SceneabstractLearning the knowledge of scene structure and tracking a large number of targets are both active topics of computer vision in recent years, which plays a crucial role in surveillance, activity analysis, object classification and etc. In this paper, we propose a novel system which simultaneously performs the Learning-Semantic-Scene and Tracking, and makes them supplement each other in one framework. The trajectories obtained by the tracking are utilized to continually learn and update the scene knowledge via an online un-supervised learning. On the other hand, the learned knowledge of scene in turn is utilized to supervise and improve the tracking results. Therefore, this “adaptive learning-tracking loop” can not only perform the robust tracking in high density crowd scene, dynamically update the knowledge of scene structure and output semantic words, but also ensures that the entire process is completely automatic and online. We successfully applied the proposed system into the JR subway station of Tokyo, which can dynamically obtain the semantic scene structure and robustly track more than 150 targets at the same time. Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Jinshi Cui, Ryosuke Shibasaki, Hongbin Zha |
CVPR | 3 |
| 2010 | Fusion of laser and vision for multiple targets tracking via on-line learningabstractMulti-target tracking becomes significantly more challenging when the targets are in close proximity or frequently interact with each other. This paper presents a promising tracking system to deal with these problems. The novelty of this system is that laser and vision, tracking and learning are integrated and can complement each other in one framework: when the targets do not interact with each other, the laser-based independent trackers are employed and the visual information is extracted simultaneously to train some classifiers for the “possible interacting targets”. When the targets are in close proximity, the learned classifiers and visual information are used to assist in tracking. Therefore, this mode of co-operation between them not only deals with various tough problems encountered in the tracking, but also ensures that the entire process can be completely on-line and automatic. Experimental results demonstrated that laser and vision fully display their respective advantages in our system, and it is easy for us to obtain a perfect trade-off between tracking accuracy and time-cost. Xuan Song 0001, Huijing Zhao, Jinshi Cui, Xiaowei Shao, Ryosuke Shibasaki, Hongbin Zha |
ICRA | 2 |
| 2010 | Scene understanding in a large dynamic environment through a laser-based sensingabstractIt became a well known technology that a map of complex environment containing low-level geometric primitives (such as laser points) can be generated using a robot with laser scanners. This research is motivated by the need of obtaining semantic knowledge of a large urban outdoor environment after the robot explores and generates a low-level sensing data set. An algorithm is developed with the data represented in a range image, while each pixel can be converted into a 3D coordinate. Using an existing segmentation method that models only geometric homogeneities, the data of a single object of complex geometry, such as people, cars, trees etc., is partitioned into different segments. Such a segmentation result will greatly restrict the capability of object recognition. This research proposes a framework of simultaneous segmentation and classification of range image, where the classification of each segment is conducted based on its geometric properties, and homogeneity of each segment is evaluated conditioned on each object class. Experiments are presented using the data of a large dynamic urban outdoor environment, and performance of the algorithm is evaluated. Huijing Zhao, Yipu Zhao, Hongbin Zha |
ICRA | 1 |
| 2010 | Segmentation and classification of range image from an intelligent vehicle in urban environmentabstractAs the rapid development of sensing and mapping techniques, it becomes a well-known technology that a map of complex environment can be generated using a robot carrying sensors. However, most of the existing researches represent environments directly using the integration of point clouds or other low-level geometric primitives. It remains an open problem to automatically convert these low-level map representations to semantic descriptions in order to effectively support high-level decision of a robot. Based on another representation of 3D point clouds, i.e. range image, this paper proposes a framework of segmentation and classification of range image, the objective of which is to annotate class labels to the data clusters that are obtained through a graph-based segmentation. Experimental results are presented and evaluated demonstrating that the proposed algorithm has efficiency in understanding the semantic knowledge of a large dynamic urban outdoor environment. Huijing Zhao, Yipu Zhao, Hongbin Zha |
IROS | 2 |
| 2010 | Multiple people extraction using 3D range sensorabstractWe propose a novel system for extracting multiple people in crowded scenes by employing LED 3D range sensor. This new kind of device can capture depth image at a relatively high frame speed, which makes the extraction of dynamic objects possible. However, it suffers from the strong noise caused by the environment illumination. Here we propose a novel method for multiple people extraction by integrating an improved version of mean-shift clustering algorithm and total variation based denoising technique. When handling the depth image, it is considered both as a point cloud as well as a 2D image. In this way, different properties of the depth image are sufficiently exploited. The proposed method can fully automatically extract multiple people in a cluttered environment, even from a mobile platform. The experiment conducted at the platform of a railway station demonstrates the effectiveness of our proposed algorithm. Xiaowei Shao, Kyoichiro Katabira, Ryosuke Shibasaki, Huijing Zhao |
SMC | 4 |
| 2009 | Moving object classification using horizontal laser scan dataabstractMotivated by two potential applications, i.e. enhancing driving safety and traffic data collection, a system has been developed using a single-layer horizontal laser scanner as the major sensor for both localization and perception of the surroundings in a large dynamic urban environment. This research focuses on a classification method, that given a stream of laser measurements, classify the moving object into either a person, a group of people, a bicycle or a car. In this research, a number of features are defined after examining the property of data appearance. A classification method is proposed after examining the likelihood measures between each pair of feature and class. Experimental results are presented, demonstrating that the algorithm has efficiency with respect to both driving safety and traffic data collection in highly dynamic environment. Huijing Zhao, Quanshi Zhang, Masaki Chiba, Ryosuke Shibasaki, Jinshi Cui, Hongbin Zha |
ICRA | 1 |
| 2009 | Combining Laser-Scanning Data and Images for Target Tracking and Scene Modeling
Hongbin Zha, Huijing Zhao, Jinshi Cui, Xuan Song 0001, Xianghua Ying |
ISRR | 2 |
| 2009 | A Laser-Scanner-Based Approach Toward Driving Safety and Traffic Data CollectionabstractThis work is motivated by the following two potential applications: 1) enhancing driving safety and 2) collecting traffic data in a large dynamic urban environment. A laser-scanner-based approach is proposed. The problem is formulated as a simultaneous localization and mapping (SLAM) with object tracking and classification, where the focus is on managing a mixture of data from both dynamic and static objects in a highly dynamic environment. A trajectory-oriented closure is also proposed using the sporadically available global positioning system (GPS) measurements in urban areas to assist for global accuracy, particularly when the vehicle makes a noncyclical measurement in a large outdoor environment. Experiments are conducted using the data that were collected along a course near 4.5 km in a highly dynamic environment. Possibilities of the approaches toward the two potential applications are demonstrated, and avenues for future works are discussed. Huijing Zhao, Masaki Chiba, Ryosuke Shibasaki, Xiaowei Shao, Jinshi Cui, Hongbin Zha |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2008 | Probabilistic Detection-based Particle Filter for Multi-target TrackingabstractIn this paper, we present a Probabilistic Detection-based Particle Filter (PD-PF) for tracking a variable number of interacting targets. When the objects do not interact with each other, our method performs like the deterministic detection-base methods. When the objects are in close proximity, the interactions and occlusions are modelled by a mixed proposal constructed by probabilistic detections and information from dynamic models. Specially, prior of detection-reliability minimizes the influence of non-detection or false alarm in the tracking. Moreover, we run independent PD-PF for each target, such that particles are sampled in a small state space, thus our method not only obtains a better approximation of posterior than joint particle filter or independent particle filter when interactions occur, but also has an acceptable computational complexity. Different evaluations demonstrate the validity and efficiency of the proposed method. 1 Xuan Song 0001, Jinshi Cui, Hongbin Zha, Huijing Zhao |
BMVC | 4 |
| 2008 | Vision-Based Multiple Interacting Targets Tracking via On-Line Supervised Learning
Xuan Song 0001, Jinshi Cui, Hongbin Zha, Huijing Zhao |
ECCV (3) | 4 |
| 2008 | Tracking interacting targets with laser scanner via on-line supervised learningabstractSuccessful multi-target tracking requires locating the targets and labeling their identities. For the laser based tracking system, the latter becomes significantly more challenging when the targets frequently interact with each other. This paper presents a novel on-line supervised learning based method for tracking interacting targets with laser scanner. When the targets do not interact with each other, we collect samples and train a classifier for each target. When the targets are in close proximity, we use these classifiers to assist in tracking. Different evaluations demonstrate that this method has a better tracking performance than previous methods when interactions occur, and can maintain correct tracking under various complex tracking situations. Xuan Song 0001, Jinshi Cui, Xulei Wang, Huijing Zhao, Hongbin Zha |
ICRA | 4 |
| 2008 | SLAM in a dynamic large outdoor environment using a laser scannerabstractIn this research, we propose a method of SLAM in a dynamic large outdoor environment using a laser scanner. Focus are cast on solving two major problems: 1) achieving global accuracy especially in non-cyclical environment, 2) tackling a mixture of data from both dynamic and static objects. Algorithms are developed, where GPS data and control inputs are used to diagnose pose error and guide to achieve a global accuracy; Classification of laser points and objects are conducted not in an independent module but across the processing in a framework of SLAM with moving object detection and tracking. Experiments are conducted using the data from two test-bed vehicles, and performance of the algorithms are demonstrated. Huijing Zhao, Masaki Chiba, Ryosuke Shibasaki, Xiaowei Shao, Jinshi Cui, Hongbin Zha |
ICRA | 1 |
| 2008 | Tracking a variable number of pedestrians in crowded scenes by using laser range scannersabstractWe propose a novel system for tracking a variable number of pedestrians in crowded scenes by exploiting laser range scanners. Based on the specific pattern generated by walking feet in the spatio-temporal domain, a walking model is constructed and applied to track pedestrians. To track interactive targets, an algorithm based on Interactive Multiple Particle Filters (IMPF) is proposed, whose computation load increases linearly with the number of targets. To handle a variable number of pedestrians, spatio-temporal correlation analysis in combination with a mean shift based clustering technique is proposed. Compared with camera-based surveillance methods, our system provides a novel technique for automatically tracking a large number of pedestrians in a relatively large area. The experiments, in which over 2600 pedestrians were tracked in 10 minutes at a 60 m × 20 m subway station, show the effectiveness of our proposed algorithm. Xiaowei Shao, Kyoichiro Katabira, Ryosuke Shibasaki, Huijing Zhao, Yuri Nakagawa |
SMC | 4 |
| 2008 | Multi-modal tracking of people using laser scanners and video camera
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
Image Vis. Comput. | 3 |
| 2007 | Monitoring a populated environment using single-row laser range scanners from a mobile platformabstractIn this research, we proposed a system of detecting and monitoring pedestrians' motion trajectories at a populated and wide environment, such as exhibition hall, supermarket etc., using the horizontally profiling single-row laser range scanners on a mobile platform. A simplified walking model is defined to track the rhythmic swing feet at the ground level. Pedestrians are recognized by detecting the braided styles, which is a typical appearance that could discriminate the data of moving feet with other mobile and motionless objects. Two experiments are conducted. One is at the laboratory environment, the purpose of which is to examine the algorithm in details. Another is at an exhibition hall, a populated and wide environment, the purpose is to examine whether the system could be applied for practical needs. It is a big challenge, while the system did well. Pedestrians in the exhibition hall at the moment of measurement are detected. Their motion trajectories are extracted, and associated to the background map, which is made of the motionless objects, and covers the whole exhibition hall. Huijing Zhao, Xiaowei Shao, Kyoichiro Katabira, Ryosuke Shibasaki |
ICRA | 1 |
| 2007 | Detection and tracking of multiple pedestrians by using laser range scannersabstractWe propose a novel system for tracking multiple pedestrians in a crowded scene by exploiting single-row laser range scanners that measure distances of surrounding objects. A walking model is built to describe the periodicity of the movement of the feet in the spatial-temporal domain, and a mean-shift clustering technique in combination with spatial- temporal correlation analysis is applied to detect pedestrians. Based on the walking model, particle filter is employed to track multiple pedestrians. Compared with camera-based methods, our system provides a novel technique to track multiple pedestrians in a relatively large area. The experiments, in which over 300 pedestrians were tracked in 5 minutes, show the validity of the proposed system. Xiaowei Shao, Huijing Zhao, Katsuyuki Nakamura, Kyoichiro Katabira, Ryosuke Shibasaki, Yuri Nakagawa |
IROS | 2 |
| 2007 | Laser-based detection and tracking of multiple people in crowds
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
Comput. Vis. Image Underst. | 3 |
| 2006 | Fusion of Detection and Matching Based Approaches for Laser Based Multiple People TrackingabstractMost of visual tracking algorithms have been achieved by matching-based searching strategies or detection-based data association algorithms. In this paper, our objective is to analysis laser scan image sequences to track multiple people in a crowded environment. Due to the poor features provided by laser scan images, neither of the above two approaches can achieves good tracking. To address the problem, we propose a novel multiple-target tracking algorithm fusing both detection and matching based strategies. First, target to detected measurement data association is incorporated to the joint state proposal, to form a mixture proposal that combines information from the dynamic model and the detected measurements. And then, we utilize a MCMC sampling step to obtain a more efficient multi-target filter. Our approach has been applied to the real laser scan image data. Evaluations show that the proposed method is a robust and effective multi-target tracking algorithm. Jinshi Cui, Huijing Zhao, Ryosuke Shibasaki |
CVPR (1) | 2 |
| 2006 | Laser-based Interacting People Tracking Using Multi-level ObservationsabstractLaser based people tracking systems have been developed for mobile robotics and intelligent surveillance areas. Existing systems rely on simple laser point clustering methods to extract object locations. However, when dealing with multiple interacting people, laser points of different persons are often interlaced and undistinguishable due to measurement noise and they can not provide reliable features. It causes current systems quite fragile and unreliable. In this paper, we try to explore potentials from multi-level observations including weakly detected features, stably extracted features and foreground points. For inference, detection incorporated joint particle filter is used. And stably extracted features are utilized to properly estimate parameters of dynamic model for each target. In real experiments, we obtain raw data from multiple registered laser scanners, which measure two legs for each people. Evaluations with real data show that the proposed method is more robust and effective than existing approaches Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
IROS | 3 |
| 2006 | Analyzing Pedestrians' Walking Patterns Using Single-Row Laser Range ScannersabstractWe propose a novel system for analyzing pedestrians' walking patterns by exploiting single-row laser range scanners that measure distances of surrounding objects by reflecting eye-safe laser beams. A walking model is built in the spatial-temporal domain to describe the periodicity of the movement of the feet, and a mean-shift technique is applied to recover model parameters. Compared with camera-based methods, our system provides a novel technique to analyze the behavior of pedestrians. The experiments show the validity of the algorithm. Xiaowei Shao, Huijing Zhao, Katsuyuki Nakamura, Ryosuke Shibasaki, Rong Zhang 0004, Zhengkai Liu |
SMC | 2 |
| 2005 | Tracking multiple people using laser and visionabstractWe present a novel system that aims at reliably detecting and tracking multiple people in an open area. Multiple single-row laser scanners and one video camera are utilized. Feet trajectory tracking based on registration of distance information from multiple laser scanners and visual body region tracking based on color histogram are combined in a Bayesian formulation. Results from tests in a real environment are reported to demonstrate that the system can detect and track multiple people simultaneously with reliable and real-time performance. Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
IROS | 3 |
| 2005 | A novel system for tracking pedestrians using multiple single-row laser-range scannersabstractWe propose a novel system for tracking pedestrians in a wide and open area, such as a shopping mall and exhibition hall, using a number of single-row laser-range scanners (LD-A), which have a profiling rate of 10 Hz and a scanning angle of 270/spl deg/. LD-As are set directly on the floor doing horizontal scanning at an elevation of about 20 cm above the ground, so that horizontal cross sections of the surroundings, containing moving feet of pedestrians as well as still objects, are obtained in a rectangular coordinate system of real dimension. The data of moving feet are extracted through background subtraction by the client computers that control each LD-A, and sent to a server computer, where they are spatially and temporally integrated into a global coordinate system. A simplified pedestrian's walking model based on the typical appearance of moving feet is defined and a tracking method utilizing Kalman filter is developed to track pedestrian's trajectories. The system is evaluated through both real experiment and computer simulation. A real experiment is conducted in an exhibition hall, where three LD-As are used covering an area of about 60/spl times/60 m/sup 2/. Changes in visitors' flow during the whole exhibition day are analyzed, where in the peak hour, about 100 trajectories are extracted simultaneously. On the other hand, a computer simulation is conducted to quantitatively examine system performance with respect to different crowd density. Huijing Zhao, Ryosuke Shibasaki |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2003 | Reconstructing a textured CAD model of an urban environment using vehicle-borne laser range scanners and line cameras
Huijing Zhao, Ryosuke Shibasaki |
Mach. Vis. Appl. | 1 |
| 2003 | A vehicle-borne urban 3-D acquisition system using single-row laser range scannersabstractIn this research, a novel vehicle-borne system of measuring three-dimensional (3-D) urban data using single-row laser range scanners is proposed. Two single-row laser range scanners are mounted on the roof of a vehicle, doing horizontal and vertical profiling respectively. As the vehicle moves ahead, a horizontal and a vertical range profile of the surroundings are captured at each odometer trigger. The freedom of vehicle motion is reduced from six to three by assuming that the ground surface is flat and smooth so resulting in the vehicle moving on almost the same horizontal plane. Horizontal range profiles, which have an overwhelming overlay between successive ones, are registered to trace vehicle location and attitude. Vertical range profiles are aligned to the coordinate system of the horizontal one according to the physical geometry between the pair of laser range scanners, and subsequently to a global coordinate system to make up 3-D data. An experiment is conducted where 3-D data of a real urban scene is obtained by registering and integrating 2412 horizontal and vertical range profiles. Two ground truths are used in examination. They are the outputs of a GPS/INS/Odometer based positioning system and a 1:500 digital map of the testing site. Accuracy and efficiency of the method in measuring 3-D urban scene is demonstrated. Huijing Zhao, Ryosuke Shibasaki |
IEEE Trans. Syst. Man Cybern. Part B | 1 |