Haoran Li 0010

dblp:50/10038-10 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-2559-9585ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TeViR: Text-to-Video Reward With Diffusion Models for Efficient Reinforcement Learning
abstract
Developing scalable and generalizable reward engineering for reinforcement learning (RL) is crucial for creating general-purpose agents, especially in the challenging domain of robotic manipulation. While recent advances in reward engineering with vision–language models (VLMs) have shown promise, their sparse reward nature significantly limits sample efficiency. This article introduces text-to-video reward (TeViR), a novel method that leverages a pretrained text-to-video diffusion model to generate dense rewards by comparing the predicted image sequence with current observations. Experimental results across 13 simulation and real-world robotic tasks demonstrate that TeViR outperforms traditional methods leveraging sparse rewards and other state-of-the-art (SOTA) methods, achieving better sample efficiency and performance without ground truth environmental rewards. TeViR’s ability to efficiently guide agents in complex environments highlights its potential to advance RL applications in robotic manipulation.
Yuhui Chen, Haoran Li 0010, Zhennan Jiang, Haowei Wen, Dongbin Zhao
IEEE Trans. Syst. Man Cybern. Syst.2
2025 Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward Representation
abstract
Online unsupervised reinforcement learning (URL) can discover diverse skills via reward-free pre-training and exhibits impressive downstream task adaptation abilities through further fine-tuning. However, online URL methods face challenges in achieving zero-shot generalization, i.e., directly applying pre-trained policies to downstream tasks without additional planning or learning. In this paper, we propose a novel Dual-Value Forward-Backward representation (DVFB) framework with a contrastive entropy intrinsic reward to achieve both zero-shot generalization and fine-tuning adaptation in online URL. On the one hand, we demonstrate that poor exploration in forward-backward representations can lead to limited data diversity in online URL, impairing successor measures, and ultimately constraining generalization ability. To address this issue, the DVFB framework learns successor measures through a skill value function while promoting data diversity through an exploration value function, thus enabling zero-shot generalization. On the other hand, and somewhat surprisingly, by employing a straightforward dual-value fine-tuning scheme combined with a reward mapping technique, the pre-trained policy further enhances its performance through fine-tuning on downstream tasks, building on its zero-shot performance. Through extensive multi-task generalization experiments, DVFB demonstrates both superior zero-shot generalization (outperforming on all 12 tasks) and fine-tuning adaptation (leading on 10 out of 12 tasks) abilities, surpassing state-of-the-art URL methods.
Jingbo Sun 0001, Songjun Tu, Haoran Li 0010, Xin Liu 0039, Yaran Chen, Dongbin Zhao
ICLR4
2025 LeAffordNav: Enhancing Open-vocabulary Mobile Manipulation with LLM-guided Exploration and Affordance-aware Navigation
abstract
Open-vocabulary mobile manipulation is a fundamental task for robotic assistants. However, Inefficient exploration and hand-off errors between different skills pose significant challenges to completing mobile manipulation tasks. In this paper, we propose a novel method named LeAffordNav, which is composed of LLM-guided exploration and Affordance-aware Navigation to address these challenges. LLM-guided exploration introduces LLMs to combine commonsense inference and frontier-based exploration, and achieves the balance between exploration and finding the target object. Considering the manipulability of the robot arm and the accessibility of the robot, we propose Affordance-aware Navigation which predicts the affordance of the mobile manipulation to reduce the hand-off errors between navigation and manipulation. Experiments on the HomeRobot benchmark show that LeAffordNav achieves new state-of-the-art performance, with a 20% higher success rate than the previous best. The code is available at https: //github.com/Cyuanwen/LeAffordNav.
Yuanwen Chen, Haoran Li 0010, Yaran Chen, Dongbin Zhao
ICME2
2025 GAPartManip: A Large-Scale Part-Centric Dataset for Material-Agnostic Articulated Object Manipulation
abstract
Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception and pose detection. However, in real-world environments, these methods often face challenges due to imperfect depth perception, such as with transparent lids and reflective handles. Moreover, they generally lack the diversity in partbased interactions required for flexible and adaptable manipulation. To address these challenges, we introduced a largescale part-centric dataset for articulated object manipulation that features both photo-realistic material randomizations and detailed annotations of part-oriented, scene-level actionable interaction poses. We evaluated the effectiveness of our dataset by integrating it with several state-of-the-art methods for depth estimation and interaction pose prediction. Additionally, we proposed a novel modular framework that delivers superior and robust performance for generalizable articulated object manipulation. Our extensive experiments demonstrate that our dataset significantly improves the performance of depth perception and actionable interaction pose prediction in both simulation and real-world scenarios. More information and demos can be found at: https://pku-epic.github.io/GAPartManip/.
Wenbo Cui, Chengyang Zhao, Songlin Wei, Jiazhao Zhang, Yaran Chen, Haoran Li 0010, He Wang 0010
ICRA7
2025 Sample-Efficient Unsupervised Policy Cloning from Ensemble Self-Supervised Labeled Videos
abstract
Current advanced policy learning methodologies have demonstrated the ability to develop expert-level strategies when provided enough information. However, their requirements, including task-specific rewards, action-labeled expert trajectories, and huge environmental interactions, can be expensive or even unavailable in many scenarios. In contrast, humans can efficiently acquire skills within a few trials and errors by imitating easily accessible internet videos, in the absence of any other supervision. In this paper, we try to let machines replicate this efficient watching-and-learning process through Unsupervised Policy from Ensemble Self-supervised labeled Videos (UPESV), a novel framework to efficiently learn policies from action-free videos without rewards and any other expert supervision. UPESV trains a video labeling model to infer the expert actions in expert videos through several organically combined self-supervised tasks. Each task performs its duties, and they together enable the model to make full use of both action-free videos and reward-free interactions for robust dynamics understanding and advanced action prediction. Simultaneously, UPESV clones a policy from the labeled expert videos, in turn collecting environmental interactions for self-supervised tasks. After a sample-efficient, unsupervised, and iterative training process, UPESV obtains an advanced policy based on a robust video labeling model. Extensive experiments in sixteen challenging procedurally generated environments demonstrate that the proposed UPESV achieves state-of-the-art interaction-limited policy learning performance (outperforming five current advanced baselines on 12/16 tasks) without exposure to any other supervision except for videos.
Xin Liu 0039, Yaran Chen, Haoran Li 0010
ICRA3
2025 Consistency Policy with Categorical Critic for Autonomous Driving
Haoran Li 0010, Dongbin Zhao
AAMAS3
2025 Advancing Object-Goal Navigation through LLM-enhanced Object Affinities Transfer
abstract
Object-goal navigation requires mobile robots to efficiently locate targets with visual and spatial information, yet existing methods struggle with generalization in unseen environments. Heuristic approaches with naive metrics fail in complex layouts, while graph-based and learning-based methods suffer from environmental biases and limited generalization. Although Large Language Models (LLMs) as planners or agents offer a rich knowledge base, they are cost-inefficient and lack targeted historical experience. To address these challenges, we propose the LLM-enhanced Object Affinities Transfer (LOAT) framework, integrating LLM-derived semantics with learning-based approaches to leverage experiential object affinities for better generalization in unseen settings. LOAT employs a dual-module strategy: one module accesses LLMs’ vast knowledge, and the other applies learned object semantic relationships, dynamically fusing these sources based on context. Evaluations in AI2-THOR and Habitat simulators show significant improvements in navigation success and efficiency, and real-world deployment demonstrates the zero-shot ability of LOAT to enhance object-goal navigation systems.
Mengying Lin, Shugao Liu, Dingxi Zhang, Yaran Chen, Zhaoran Wang 0001, Haoran Li 0010, Dongbin Zhao
IROS6
2025 Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
abstract
Humans can efficiently extract knowledge and learn skills from the videos within only a few trials and errors. However, it poses a big challenge to replicate this learning process for autonomous agents, due to the complexity of visual input, the absence of action or reward signals, and the limitations of interaction steps. In this paper, we propose a novel, unsupervised, and sample-efficient framework to achieve imitation learning from videos (ILV), named Behavior Cloning from Videos via Latent Representations (BCV-LR). BCV-LR extracts action-related latent features from high-dimensional video inputs through self-supervised tasks, and then leverages a dynamics-based unsupervised objective to predict latent actions between consecutive frames. The pre-trained latent actions are fine-tuned and efficiently aligned to the real action space online (with collected interactions) for policy behavior cloning. The cloned policy in turn enriches the agent experience for further latent action finetuning, resulting in an iterative policy improvement that is highly sample-efficient. We conduct extensive experiments on a set of challenging visual tasks, including both discrete control and continuous control. BCV-LR enables effective (even expert-level on some tasks) policy performance with only a few interactions, surpassing state-of-the-art ILV baselines and reinforcement learning methods (provided with environmental rewards) in terms of sample efficiency across 24/28 tasks. To the best of our knowledge, this work for the first time demonstrates that videos can support extremely sample-efficient visual policy learning, without the need to access any other expert supervision.
Xin Liu 0039, Haoran Li 0010, Dongbin Zhao
NeurIPS2
2025 FusionNav: Enhancing Zero-Shot Object-Goal Navigation via 3D Semantic Fusion and Farsight Value Reasoning
abstract
Zero-Shot Object-Goal Navigation (ZSON) tasks require agents to efficiently find target objects in unfamiliar environments, demanding strong semantic understanding and generalization. We propose FusionNav, a novel zero-shot navigation framework that combines 3D point cloud semantics with farsight value estimation. By integrating both local semantic information and semantic cues from regions outside the agent’s current field of view, FusionNav enables reasoning about unexplored areas and improves navigation efficiency. Experiments on the HM3D and MP3D benchmarks show that FusionNav outperforms strong baselines in both success rate and path efficiency. Moreover, FusionNav can be directly deployed on real-world robots without additional training or fine-tuning, and operates efficiently with moderate computational resources. These results demonstrate the effectiveness and practicality of FusionNav for real-world zero-shot object-goal navigation.
Shugao Liu, Haoran Li 0010, Dongbin Zhao
SMC3
2025 Balancing State Exploration and Skill Diversity in Unsupervised Skill Discovery
abstract
Unsupervised skill discovery seeks to acquire different useful skills without extrinsic reward via unsupervised reinforcement learning (RL), with the discovered skills efficiently adapting to multiple downstream tasks in various ways. However, recent advanced skill discovery methods struggle to well balance state exploration and skill diversity, particularly when the potential skills are rich and hard to discern. In this article, we propose contrastive dynamic skill discovery (ComSD) which generates diverse and exploratory unsupervised skills through a novel intrinsic incentive, named contrastive dynamic reward. It contains a particle-based exploration reward to make agents access far-reaching states for exploratory skill acquisition, and a novel contrastive diversity reward to promote the discriminability between different skills. Moreover, a novel dynamic weighting mechanism between the above two rewards is proposed to balance state exploration and skill diversity, which further enhances the quality of the discovered skills. Extensive experiments and analysis demonstrate that ComSD can generate diverse behaviors at different exploratory levels for multijoint robots, enabling state-of-the-art adaptation performance on challenging downstream tasks. It can also discover distinguishable and far-reaching exploration skills in the challenging tree-like 2-D maze.
Xin Liu 0039, Yaran Chen, Guixing Chen, Haoran Li 0010, Dongbin Zhao
IEEE Trans. Cybern.4
2025 Cross-Domain Random Pretraining With Prototypes for Reinforcement Learning
abstract
Unsupervised cross-domain reinforcement learning (RL) pretraining shows great potential for challenging continuous visual control but poses a big challenge. In this article, we propose cross-domain random pretraining with prototypes (CRPTpro), a novel, efficient, and effective self-supervised cross-domain RL pretraining framework. CRPTpro decouples data sampling from encoder pretraining, proposing decoupled random collection to easily and quickly generate a qualified cross-domain pretraining dataset. Moreover, a novel prototypical self-supervised algorithm is proposed to pretrain an effective visual encoder that is generic across different domains. Without finetuning, the cross-domain encoder can be implemented for challenging downstream tasks defined in different domains, either seen or unseen. Compared with recent advanced methods, CRPTpro achieves better performance on downstream policy learning without extra training on exploration agents for data collection, greatly reducing the burden of pretraining. We conduct extensive experiments across multiple challenging continuous visual-control domains, including balance control, robot locomotion, and manipulation. CRPTpro significantly outperforms the next best Proto-RL(C) on 11/12 cross-domain downstream tasks with only 54.5% wall-clock pretraining time, exhibiting state-of-the-art pretraining performance with greatly improved pretraining efficiency.
Xin Liu 0039, Yaran Chen, Haoran Li 0010, Boyu Li 0003, Dongbin Zhao
IEEE Trans. Syst. Man Cybern. Syst.3
2024 ATV3D: 3D Object Detection from Attention-based Three-view Representation
abstract
In the fields of autonomous driving and robot perception, the majority of methods are designed for onboard camera object detection, while there are fewer methods specifically tailored to environmental cameras. However, environmental cameras have the capability to capture a significant amount of road geometry and vehicle position information, which can enhance the safety of autonomous driving. Nevertheless, there is a difference in perspective between environmental cameras and onboard cameras, resulting in poorer performance of many methods designed for onboard camera 3D object detection when apply to environmental camera. In this paper, we propose a 3D Object Detection Algorithm from Attention-based Three-view Representation (ATV3D). The algorithm projects the 2D image features onto three orthogonal views (left view, front view, bird’s eye view) to achieve a representation of the 3D information. Compared to voxel-based 3D detection methods, our proposed approach retains the ability to capture 3D features while reducing computational complexity. During the process of three-view representation, we design a feature projection module based on attention. Unlike inverse perspective mapping that requires precise camera parameters, the attention can implicitly learn the mapping relationship from 2D images to the three-view planes. This enables the extraction and transformation of image features without the calibrated camera parameters, effectively addressing challenges associated with obtaining camera parameters for environmental cameras and their susceptibility to natural factors. The experimental results on the DAIR-V2X dataset demonstrate that our method achieves a 3D detection mean average precision (mAP) of 73.6%, surpassing the performance of previous calibration-free environmental camera methods. Furthermore, our method achieves the highest detection accuracy on the indoor multi-view robot dataset Neurons Perception, providing evidence of its outstanding detection performance.
Yaran Chen, Haoran Li 0010, Yunzhen Zhao, Pengfei Hu 0004
IJCNN3
2024 High-quality Synthetic Data is Efficient for Model-based Offline Reinforcement Learning
abstract
Recent work has found that two types of dataset characteristics including the dataset’s coverage and data quality are critical for offline reinforcement learning (RL). To improve the policy, model-based offline RL tries to generate reliable synthetic data to expand the dataset’s coverage based on trained forward and backward dynamics models. However, the characteristic of synthetic data’s quality is ignoring, which raises a question of whether augmenting high-quality synthetic data is efficient for offline RL agents. Motivated by this, we propose a novel forward High-quality Imagination and backward Reliable Check (HIRC), which is an effective data augmentation method to generate high-quality and reliable synthetic data. Specifically, we construct a value-guided forward model to generate high-quality imaginary trajectories, and employ a backward model for reliable checking to obtain synthetic data that better match with pre-collected offline transitions. In other words, the proposed HIRC method can generate high-quality synthetic data on the premise of reliability, which can be combined with model-free offline RL methods. Experimental results on the D4RL benchmark demonstrate that high-quality synthetic data generated by HIRC boosts the performance of a base agent TD3_BC. Especially, HIRC with such a base agent achieves better scores against recent popular model-free and model-based offline RL methods.
Kaixuan Xu, Weixin Zhao, Haoran Li 0010, Dongbin Zhao
IJCNN5
2024 Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
abstract
With high-dimensional state spaces, visual reinforcement learning (RL) faces significant challenges in exploitation and exploration, resulting in low sample efficiency and training stability. As a time-efficient diffusion model, although consistency models have been validated in online state-based RL, it is still an open question whether it can be extended to visual RL. In this paper, we investigate the impact of non-stationary distribution and the actor-critic framework on consistency policy in online RL, and find that consistency policy was unstable during the training, especially in visual RL with the high-dimensional state space. To this end, we suggest sample-based entropy regularization to stabilize the policy training, and propose a consistency policy with prioritized proximal experience regularization (CP3ER) to improve sample efficiency. CP3ER achieves new state-of-the-art (SOTA) performance in 21 tasks across DeepMind control suite and Meta-world. To our knowledge, CP3ER is the first method to apply diffusion/consistency models to visual RL and demonstrates the potential of consistency models in visual RL.
Haoran Li 0010, Zhennan Jiang, Yuhui Chen, Dongbin Zhao
NeurIPS1
2023 NeuronsMAE: A Novel Multi-Agent Reinforcement Learning Environment for Cooperative and Competitive Multi-Robot Tasks
abstract
Multi-agent reinforcement learning (MARL) has achieved remarkable success in various challenging problems. Meanwhile, more and more benchmarks have emerged and provided some standards to evaluate the algorithms in different fields. On the one hand, the virtual MARL environments lack knowledge of real-world tasks and actuator abilities. On the other hand, the current task-specified multi-robot platform has poor support for the universality of multi-agent reinforcement learning algorithms and lacks support for transferring from simulation to the real environment. Bridging the gap between the virtual MARL environments and the real multi-robot platform becomes the key to promoting the practicability of MARL algorithms. This paper proposes a novel MARL environment for real multi-robot tasks named NeuronsMAE (Neurons Multi-Agent Environment). This environment supports cooperative and competitive multi-robot tasks and is configured with rich parameter interfaces to study the multi-agent policy transfer from simulation to reality. With this platform, we evaluate various popular MARL algorithms and build a new MARL benchmark for multi-robot tasks. We hope that this platform will facilitate the research and application of MARL algorithms for real robot tasks. Information about the benchmark and the open-source code are released at https://github.com/DRL-CASIA/NeuronsMAE.
Guangzheng Hu, Haoran Li 0010, Yuanheng Zhu, Dongbin Zhao
IJCNN2
2022 Neurons Perception Dataset for RoboMaster AI Challenge
abstract
From virtual game to physical robot, games have witnessed the development of artificial intelligence (AI) technology, especially the data-driven technology represented by deep learning. Compared with virtual games, a physical robot game such as RoboMaster AI challenge needs to build a complete closed-loop architecture composed of perception, planning, control, and decision-making to support autonomous confrontation. Perception, as the eye of the robot, its performance in the complex environment depends on a massive dataset. Although there are many open perception datasets, these datasets are difficult to meet the needs of RoboMaster AI challenge due to the high dynamics of the task, the distinctiveness of the objects, and limited computing resources. In this paper, we release a dataset named Neurons11Neurons is a team dedicated to promoting the development of robot with deep neural network. We will release the code and dataset at https://github.com/DRL-CASIA/NeuronsDataset. perception dataset for RoboMaster AI challenge, which covers 3 tasks including monocular depth estimation, lightweight object detection, and multi-view 3D object detection, and makes up the data blank in this field. In addition, we also evaluate State-Of-The-Art (SOTA) methods on each task, hoping to provide an impartial benchmark for the development of perception algorithm.
Haoran Li 0010, Zicheng Duan, Yaran Chen, Dongbin Zhao
IJCNN1
2022 BiFNet: Bidirectional Fusion Network for Road Segmentation
abstract
Multisensor fusion-based road segmentation plays an important role in the intelligent driving system since it provides a drivable area. The existing mainstream fusion method is mainly to feature fusion in the image space domain which causes the perspective compression of the road and damages the performance of the distant road. Considering the bird's eye views (BEVs) of the LiDAR remains the space structure in the horizontal plane, this article proposes a bidirectional fusion network (BiFNet) to fuse the image and BEV of the point cloud. The network consists of two modules: 1) the dense space transformation (DST) module, which solves the mutual conversion between the camera image space and BEV space and 2) the context-based feature fusion module, which fuses the different sensors information based on the scenes from corresponding features. This method has achieved competitive results on the KITTI dataset.
Haoran Li 0010, Yaran Chen, Dongbin Zhao
IEEE Trans. Cybern.1
2022 Boost 3-D Object Detection via Point Clouds Segmentation and Fused 3-D GIoU-L₁ Loss
abstract
The 3-D object detection is crucial for many real-world applications, attracting many researchers’ attention. Beyond 2-D object detection, 3-D object detection usually needs to extract appearance, depth, position, and orientation information from light detection and ranging (LiDAR) and camera sensors. However, due to more degrees of freedom and vertices, existing detection methods that directly transform from 2-D to 3-D still face several challenges, such as exploding increase of anchors’ number and inefficient or hard-to-optimize objective. To this end, we present a fast segmentation method for 3-D point clouds to reduce anchors, which can largely decrease the computing cost. Moreover, taking advantage of 3-D generalized Intersection of Union (GIoU) and$L_{1}$losses, we propose a fused loss to facilitate the optimization of 3-D object detection. A series of experiments show that the proposed method has alleviated the abovementioned issues effectively.
Yaran Chen, Haoran Li 0010, Ruiyuan Gao 0001, Dongbin Zhao
IEEE Trans. Neural Networks Learn. Syst.2
2021 IA-CNN: A generalised interpretable convolutional neural network with attention mechanism
abstract
In recent years, convolutional neural network (CNN) has been widely used in security, autonomous driving, and healthcare. Even though CNN has achieved a great performance, the results produced by CNN are difficult to explain and sometimes irresponsible. The black-box nature of CNN makes it lack trust. In this paper, we propose an attention based CNN structure, named IA -CNN, which highly improves the interpretability of the CNN models. Each feature map of the last conv-layer only has one response (one key point) of the target object, which is directly connected to the output. We also combine the attention mechanism to weakly supervise the last conv-layer. In this way, our model can clearly show that which features the model extracted are the keys to the output prediction. Meanwhile, our IA-CNN structure can be used in various classical models with higher performance in the fine-grained classification and comparative performance in the ordinary classification task. Note that our IA-CNN structure is an end-to-end model, the last conv-layer of which can extract key points from images automatically and is connected to the output prediction linearly.
Zhisong Zhang, Yaran Chen, Haoran Li 0010
IJCNN3
2020 RailNet: An Information Aggregation Network for Rail Track Segmentation
abstract
As the basis of scenes understanding for the track inspection task, track segmentation is challenging due to the various illumination conditions, track crossing, and plant coverage. Since the rail has a strong shape prior, strict rail spacing and special distribution in the image, making full use of the spatial information of the rail features becomes an important factor to improve the accuracy of rail segmentation. In this paper, an information aggregation module is proposed to enhance the spatial relationship between pixels of the rail features. In other words, this module expands the receptive field. Furthermore, we build an information aggregation network based on this module, which is called as RailNet. Finally, the RailNet is evaluated in an open train track dataset. Experimental results show that RailNet can achieve the best performance so far in the dataset of trains.
Haoran Li 0010, Dongbin Zhao, Yaran Chen
IJCNN1
2020 Deep Reinforcement Learning-Based Automatic Exploration for Navigation in Unknown Environment
abstract
This paper investigates the automatic exploration problem under the unknown environment, which is the key point of applying the robotic system to some social tasks. The solution to this problem via stacking decision rules is impossible to cover various environments and sensor properties. Learning-based control methods are adaptive for these scenarios. However, these methods are damaged by low learning efficiency and awkward transferability from simulation to reality. In this paper, we construct a general exploration framework via decomposing the exploration process into the decision, planning, and mapping modules, which increases the modularity of the robotic system. Based on this framework, we propose a deep reinforcement learning-based decision algorithm that uses a deep neural network to learning exploration strategy from the partial map. The results show that this proposed algorithm has better learning efficiency and adaptability for unknown environments. In addition, we conduct the experiments on the physical robot, and the results suggest that the learned policy can be well transferred from simulation to the real robot.
Haoran Li 0010, Dongbin Zhao
IEEE Trans. Neural Networks Learn. Syst.1
2019 Deep Kalman Filter with Optical Flow for Multiple Object Tracking
abstract
Deep matching and Kalman filter-based multiple object tracking (DK-tracking) have been demonstrated to be promising. However, most of existing DK-tracking trackers assume that objects are slow-varying movement with a constant velocity. The assumption is hard to be satisfied in the real world, especially in the image space due to the sight distance. In this paper, we propose a novel multiple object tracking method combining deep feature matching, Kalman filter and flow information, which is called DK-flow-tracking, to improve tracking performance. In DK-flow-tracking, optical flow in consecutive frames is used to provide accurate object motion information for guiding Kalman filter to track objects. Experiments are performed on public datasets: MOT2016, MOT2017, and the proposed method achieves better performances compared to the DK-tracking with the assumption of a constant velocity movement.
Yaran Chen, Dongbin Zhao, Haoran Li 0010
SMC3
2018 A temporal-based deep learning method for multiple objects detection in autonomous driving
abstract
This paper proposes a novel vision-based object detection method in autonomous driving, which introduces the temporal information into the deep learning-based detection method for moving object detection. Vision-based object detection is a critical technology for autonomous driving. The objects in the real world such as driving cars, don't have great changes in their positions and velocities. So the position change of objects between two consecutive frames is not large. This is usually ignored by traditional works, which usually use object detection methods on still-images to detect moving objects. Considering the relationship among consecutive frames (temporal information), we present a robust and real-time tracking method following image detection to refine the object detection results. Based on the three key attributes (distances, sizes and positions), the tracking method aims to build the association between the detected objects on the current frame and those in previous frames. The proposed object detection with temporal information dramatically improves the performance of existing object detection algorithms based on stillimage. With the proposed method, we won the champion in the preceding vehicle detection task in 2017 intelligent vehicle future challenge(2017 IVFC)1.
Yaran Chen, Dongbin Zhao, Haoran Li 0010, Dong Li 0016, Ping Guo 0002
IJCNN3