EDBT 2026 Demo / reviewers in the wild / expert
Zhiheng Li 0001
dblp:89/6935-1
· DBLP profile ↗
35ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0002-1523-1114ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 10 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Cross-Modal Consistency for Sensor Fault Detection in Autonomous Driving
Shiteng Cao, Zhiheng Li 0001 |
IV | 3 |
| 2026 | Energy-Efficient Speed Planning for Automated Truck Platoon With Uncertainty-Aware Trajectory PredictionabstractFor fully automated heavy-duty truck platoons, the uncertain behavior of surrounding traffic participants makes platooning control and stability extremely complex. This may lead to a deterioration in fuel efficiency and could even result in collisions. To mitigate the impact of traffic participants’ uncertainty and enhance fuel economy, this paper proposes an energy-efficient speed planning method for automated truck platoons considering uncertainty-aware trajectory prediction. In trajectory prediction uncertainty modeling, perceptual uncertainty is incorporated into the loss function of the prediction module to reduce overconfidence in long-horizon predictions. To generate accurate long-horizon platoon speed trajectories that better align with vehicle dynamic characteristics, vehicle response delays are identified and and integrated into the planning process. With varying levels of prediction uncertainty and response delays, the platoon speed planning method adapts more safely and proactively to multi-vehicle interaction scenarios. Experimental results from two representative traffic flow conditions demonstrate that the proposed method effectively enhances both safety and energy efficiency in complex traffic conditions, particularly in cut-in events. Renzong Lian, Binghong Jiang, Zhiheng Li 0001, Junqing Wei, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | ABCSR: An Adaptive Blur-Aware Computational Framework for Efficient Text Image Super-ResolutionabstractScene text recognition is a fundamental task with significant implications for applications such as autonomous driving and robotics. However, unpredictable outdoor imaging conditions (e.g., motion blur, defocus) pose substantial challenges for accurate text recognition. To address this, traditional methods typically employ a text super-resolution network to restore image quality before recognition. However, this approach treats all images indiscriminately, resulting in high computational costs and degrading the quality of originally clear images. In this work, we propose the Adaptive Bluraware Computational Super-Resolution (ABCSR) framework, which integrates a blur-aware module to dynamically select computation routes based on the image’s actual blur state. Simultaneously, We also introduce a parallel computing strategy that implicitly eliminates the inference time added by the blur-aware module, significantly reducing computational overhead and latency. To our knowledge, this is the first framework to adaptively select super-resolution computational routes based on the degree of image blur. Experiments show that ABCSR framework reduces computational overhead by at least 76.92% and achieves at least 2.04× speed-up (up to 8.83×), without sacrificing reconstruction quality or recognition performance. Yawen Qiu, Qinyu Wang 0002, Zhiheng Li 0001 |
ECAI | 3 |
| 2025 | EP-ViT: Leveraging Encoding Priors to Eliminate Spatio-Temporal Redundancy in Vision TransformerabstractVision Transformer (ViT) has shown strong performance in video analysis tasks. However, conventional frame-by-frame processing results in substantial redundant computation due to repeated processing of spatiotemporally similar content across frames. Existing methods often utilize extra modules to detect spatiotemporal redundant regions for computational reduction. However, the detection process incurs additional computation and latency. To address this, we propose the Encoding Prior Vision Transformer (EP-ViT), which leverages prior information in video encoding to identify redundant regions without incurring additional computational cost. Furthermore, we introduce a redundancy-aware attention mechanism that reuses tokens with unchanged features across frames. Experimental results demonstrate that EP-ViT reduces the Transformer’s computational cost by 58.91% without compromising accuracy, surpassing state-of-the-art methods. Yawen Qiu, Qinyu Wang 0002, Zhiheng Li 0001 |
ECAI | 3 |
| 2025 | A delay-robust method for enhanced real-time reinforcement learning
Bo Xia, Bo Yuan 0003, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001 |
Neural Networks | 4 |
| 2025 | Agent-Based Space Teleoperation: Mitigating Time Delays With Deep Reinforcement LearningabstractSpace teleoperation significantly extends human reach in space missions. However, traditional approaches are constrained by factors, such as the reliance on accurate dynamic models and the risk of operator fatigue during prolonged tasks. Additionally, while data-driven intelligent approaches reduce the need for prior knowledge, they have yet to adequately address the time delay issues inherent in these systems. To overcome these challenges, we introduce the belief state actor-critic (BSAC) method, the first deep reinforcement learning approach tailored for space teleoperation capture tasks within a bilateral control framework. We first establish a generalized agent-based architecture for space teleoperation, shifting decision-making from human operators to autonomous agents. Following a comprehensive analysis of the time delay challenges, we propose the BSAC algorithm, which integrates state augmentation and belief state techniques to mitigate the effects of delays in teleoperated Markov decision processes. Extensive experiments are conducted on the MuJoCo simulation platform, modeling a real hardware system across various scenarios. The learned policies are then successfully transferred and validated in a real-world setup, demonstrating the effectiveness and robustness of BSAC. In summary, our results support the feasibility of agent-based frameworks capable of overcoming time delay challenges in space teleoperation. Bo Xia, Xianru Tian, Bo Yuan 0003, Chunju Yang, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | StreamMOTP: Streaming and Unified Framework for Joint 3D Multi-Object Tracking and Trajectory Prediction
Jiaheng Zhuang, Guoan Wang, Siyu Zhang 0002, Xiyang Wang 0002, Hangning Zhou, Ziyao Xu 0003, Chi Zhang 0067, Zhiheng Li 0001 |
ACCV (2) | 8 |
| 2024 | Dynamic Modeling for Reinforcement Learning with Random Delay
Yalou Yu, Bo Xia, Minzhi Xie, Zhiheng Li 0001, Xuwqian Wang |
ICANN (4) | 4 |
| 2024 | D3D: Conditional Diffusion Model for Decision-Making Under Random Frame DroppingabstractThe occurrence of frame drops due to issues such as corrupted communications or malfunctioning sensors presents a significant challenge to an agent’s decision-making, especially in remote control scenarios. Classical reinforcement learning (RL) usually assumes a continuous data stream without frame drops and relies heavily on online interactions, which is time-consuming, resource-intensive, and often impractical in certain scenarios. Consequently, the performance of RL may deteriorate significantly in face of non-negligible frame drops. To tackle this challenge caused by frame dropping, We propose Conditional Diffusion Model for Decision-Making under Random Frame Dropping (D3D), an offline algorithm that can effectively enhance performance robustness in frame dropping scenarios. D3D addresses this issue through a two-phase approach: 1) During the policy generation phase, D3D adopts a return-conditional diffusion model for decision making rather than the temporal difference learning, whose policy is derived using offline datasets of return-labeled trajectories without information loss. 2) When frame dropping occurs during evaluation, D3D seamlessly substitutes the missing state with its corresponding prediction in the horizon made by the diffusion model. Extensive experiments are conducted on MuJoCo and Adroit tasks to validate D3D’s robustness and efficiency. The results demonstrate that D3D consistently outperforms state-of-the-art RL algorithms, especially excelling on tasks featuring severe drop rates. Bo Xia, Yifu Luo, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001 |
RO-MAN | 5 |
| 2024 | Mixed Attention Network for Cross-domain Sequential RecommendationabstractIn modern recommender systems, sequential recommendation leverages chronological user behaviors to make effective next-item suggestions, which suffers from data sparsity issues, especially for new users. One promising line of work is the cross-domain recommendation, which trains models with data across multiple domains to improve the performance in data-scarce domains. Recent proposed cross-domain sequential recommendation models such as PiNet and DASL have a common drawback relying heavily on overlapped users in different domains, which limits their usage in practical recommender systems. In this paper, we propose a M ixed A ttention N etwork (MAN) with local and global attention modules to extract the domain-specific and cross-domain information. Firstly, we propose a local/global encoding layer to capture the domain-specific/cross-domain sequential pattern. Then we propose a mixed attention layer with item similarity attention, sequence-fusion attention, and group-prototype attention to capture the local/global item similarity, fuse the local/global item sequence, and extract the user groups across different domains, respectively. Finally, we propose a local/global prediction layer to further evolve and combine the domain-specific and cross-domain interests. Experimental results on two real-world datasets (each with two domains) demonstrate the superiority of our proposed model. Further study also illustrates that our proposed method and components are model-agnostic and effective, respectively. The code and data are available at https://github.com/Guanyu-Lin/MAN. Guanyu Lin, Chen Gao 0001, Yu Zheng 0010, Jianxin Chang, Yanan Niu, Yang Song 0008, Kun Gai, Zhiheng Li 0001, Depeng Jin, Yong Li 0008, Meng Wang 0001 |
WSDM | 8 |
| 2024 | Inverse Learning with Extremely Sparse Feedback for RecommendationabstractModern personalized recommendation services often rely on user feedback, either explicit or implicit, to improve the quality of services. Explicit feedback refers to behaviors like ratings, while implicit feedback refers to behaviors like user clicks. However, in the scenario of full-screen video viewing experiences like Tiktok and Reels, the click action is absent, resulting in unclear feedback from users, hence introducing noises in modeling training. Existing approaches on de-noising recommendation mainly focus on positive instances while ignoring the noise in a large amount of sampled negative feedback. In this paper, we propose a meta-learning method to annotate the unlabeled data from loss and gradient perspectives, which considers the noises in both positive and negative instances. Specifically, we first propose anInverse Dual Loss (IDL) to boost the true label learning and prevent the false label learning. Then we further propose anInverse Gradient (IG) method to explore the correct updating gradient and adjust the updating based on meta-learning. Finally, we conduct extensive experiments on both benchmark and industrial datasets where our proposed method can significantly improve AUC by 9.25% against state-of-the-art methods. Further analysis verifies the proposed inverse learning framework is model-agnostic and can improve a variety of recommendation backbones. The source code, along with the best hyper-parameter settings, is available at this link: https://github.com/Guanyu-Lin/InverseLearning. Guanyu Lin, Chen Gao 0001, Yu Zheng 0010, Yinfeng Li, Jianxin Chang, Yanan Niu, Yang Song 0008, Kun Gai, Zhiheng Li 0001, Depeng Jin, Yong Li 0008 |
WSDM | 9 |
| 2024 | Solving time-delay issues in reinforcement learning via transformers
Bo Xia, Zaihui Yang, Minzhi Xie, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001, Bin Liang 0001 |
Appl. Intell. | 6 |
| 2024 | Predictive Vehicle Stability Assessment Using Lyapunov Exponent Under Extreme ConditionsabstractUnder extreme conditions, vehicles may encounter critical instability and cause traffic accidents due to the tire force saturation. In such cases, accurately predicting the vehicle instability is conducive to vehicle safety because drivers or vehicle controllers can be alerted and take early interventions to ensure driving safety. However, the existing stability assessment methods tend to be conservative, hard to quantify, and often ignore the coupled longitudinal and lateral dynamics, as well as the nonlinear characteristics of tires. Simultaneously, under extreme operating conditions, the assessment of vehicle potential risk imposes higher demands on the prediction accuracy of vehicle motion states. To address these 2 issues, this paper proposes a predictive vehicle stability assessment method using 3-dimensional Lyapunov exponents (3D-LEs) for a nonlinear vehicle system. Firstly, a nonlinear 8-degree-of-freedom vehicle dynamics model is constructed for an electric vehicle, aiming to capture the coupling dynamic characteristics and the tire force saturation under extreme conditions. To minimize the simulation-reality disparities, the vehicle parameters are automatically calibrated through Bayesian optimization using field test data. Secondly, to predict the potential risk of vehicle instability precisely, a physics-informed neural network based state prediction module is established for the vehicle stability assessment system. The ordinary differential equations of the vehicle system are integrated into neural networks to obtain physically consistent predictions of vehicle dynamic motion. Finally, the 3D-LEs, encompassing lateral motion, yaw motion, and roll motion, are employed to concurrently evaluate vehicle stability. Experimental results demonstrate that the predictive vehicle stability assessment method accurately evaluates the stability of predicted state sequences, enabling safer and more stable control under extreme conditions. Renzong Lian, Zhiheng Li 0001, Wenchang Li, Jingwei Ge, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Rethinking Feature Context in Learning Image-Guided Depth Completion
Yikang Ding, Pengzhi Li, Dihe Huang, Zhiheng Li 0001 |
ICANN (3) | 4 |
| 2023 | Denoising Point Clouds with Intensity and Spatial Features in Rainy WeatherabstractLiDAR is important for 3D vision in autonomous vehicles, but rain causes inaccurate LiDAR point clouds due to reflection and scattering. Rain noise removal without loss of environmental features becomes an inevitable challenge. This paper presents a novel point cloud denoising method with intensity and spatial features to solve the problem. It utilizes a weighted edge-preserving filter to recover distorted contours and intensities of point clouds due to the reflection of the surface attached by raindrops. A low-intensity filtering method is also proposed to remove low-intensity noise due to the reflection of rainfall. In addition, a semi-synthetic rainy point cloud dataset with point-wise annotations is created, which benefits the research on improving LiDAR perception in adverse weather. Our method outperforms existing methods in terms of precision when it achieves a high recall of 99.28%. Using denoised data by our method can improve target detection accuracy by 5.37%. It is also faster than the state-of-the-art methods and shows the potential for use in snowy weather, making it suitable for all-weather LiDAR applications. Haozheng Han, Xin Jin 0002, Zhiheng Li 0001 |
ICIP | 3 |
| 2023 | Gravity-Shift-VIO: Adaptive Acceleration Shift and Multi-Modal Fusion with Transformer in Visual-Inertial OdometryabstractVisual-inertial odometry (VIO) estimates the 6-degree-of-freedom (6-DoF) ego-motion of an agent based on sequential data from cameras and inertial measurement units (IMUs). The acceleration measured through IMUs is affected by gravity, which is typically addressed by initialization methods in traditional VIO approaches. However, this problem has not received much attention in recent end-to-end deep learning methods. For raw accelerations, gravity causes overlapping be-tween different motion patterns, degenerating the representation embedding, which limits the performance of pose estimation. In this paper, we propose Gravity-Shift-VIO, an attention-based approach that addresses this issue by adaptively shifting the acceleration vector before the representation embedding. Further, a cross-frame multimodal transformer is introduced to fuse multimodal information. Experimentation on the KITTI dataset shows that Gravity-Shift-VIO exhibits strong performance and shows promising results in terms of ego-motion estimation. Further ablation study indicates that the Gravity-Shift- Viois highly effective in reducing the overlap of acceleration representation caused by gravity. And the cross-frame transformer effectively improves the multi-sensor fusion and time-series feature extraction. Zhiheng Li 0001, Xin Jin 0002 |
IJCNN | 3 |
| 2023 | Towards Practical Consistent Video Depth EstimationabstractMonocular depth estimation algorithms aim to explore the possible links between 2D and 3D data, but challenges remain for existing methods to predict consistent depth from a casual video. Relying on camera poses and the optical flow in the time-consuming test-time training phases makes these methods fail in many scenarios and cannot be used for practical applications. In this work, we present a data-driven post-processing method to overcome these challenges and achieve online processing. Based on a deep recurrent network, our method takes the adjacent original and optimized depth map as inputs to learn temporal consistency from the dataset and achieves higher depth accuracy. Our approach can be applied to multiple single-frame depth estimation models and used for various real-world scenes in real-time. In addition, to tackle the lack of a temporally consistent video depth training dataset of dynamic scenes, we propose an approach to generate the training video sequences dataset from a single image based on inferring motion field. To the best of our knowledge, this is the first data-driven plug-and-play method to improve the temporal consistency of depth estimation for casual videos. Extensive experiments on three datasets and three depth estimation models show that our method outperforms the state-of-the-art methods. Pengzhi Li, Yikang Ding, Linge Li, Jingwei Guan, Zhiheng Li 0001 |
ICMR | 5 |
| 2023 | Overcoming Delayed Feedback via Overlook Decision MakingabstractReinforcement learning is one of the most general paradigms to solve sequential decision making issues on the assumption that the action selection and environmental feedback are instantaneous, however, unfortunately this assumption is rarely true with regard to such ubiquitous delays in real-world system which could degrade the performance of reinforcement learning algorithms. The most common solution to solve a fixed delay problem is to design a forward dynamic model which is used to predict the newest state by recursively iterating over long steps so that a predicted state can be got and it would be taken as the agent's observation to make the newest decision. However, there exists cumulative errors during the iterative process which make long-term prediction inaccurate and further affect agent's decision. Motivated by the goal to reduce cumulative errors, we propose a new algorithm named Multi-step Prediction model with Delayed Observation(MPDO), aiming at accurately predicting future state at longer horizons for better decision making. Our approach includes two parts: a multi-step prediction model and a strategy training based on proximal policy optimization algorithms(PPO). Our model only needs a small amount of data to conduct dynamic modeling quickly, and the accuracy of prediction and iteration speed are higher than traditional methods. Experiments on Gym and MuJoCo show that MPDO achieves higher performance in such different tasks with different delays compared with other state-of-the-art methods, which verify our method's effectiveness. Yalou Yu, Bo Xia, Minzhi Xie, Xueqian Wang 0001, Zhiheng Li 0001, Yongzhe Chang |
SMC | 5 |
| 2023 | Dual-interest Factorization-heads Attention for Sequential RecommendationabstractAccurate user interest modeling is vital for recommendation scenarios. One of the effective solutions is the sequential recommendation that relies on click behaviors, but this is not elegant in the video feed recommendation where users are passive in receiving the streaming contents and return skip or no-skip behaviors. Here skip and no-skip behaviors can be treated as negative and positive feedback, respectively. With the mixture of positive and negative feedback, it is challenging to capture the transition pattern of behavioral sequence. To do so, FeedRec has exploited a shared vanilla Transformer, which may be inelegant because head interaction of multi-heads attention does not consider different types of feedback. In this paper, we propose Dual-interest Factorization-heads Attention for Sequential Recommendation (short for DFAR) consisting of feedback-aware encoding layer, dual-interest disentangling layer and prediction layer. In the feedback-aware encoding layer, we first suppose each head of multi-heads attention can capture specific feedback relations. Then we further propose factorization-heads attention which can mask specific head interaction and inject feedback information so as to factorize the relation between different types of feedback. Additionally, we propose a dual-interest disentangling layer to decouple positive and negative interests before performing disentanglement on their representations. Finally, we evolve the positive and negative interests by corresponding towers whose outputs are contrastive by BPR loss. Experiments on two real-world datasets show the superiority of our proposed method against state-of-the-art baselines. Further ablation study and visualization also sustain its effectiveness. We release the source code here: https://github.com/tsinghua-fib-lab/WWW2023-DFAR. Guanyu Lin, Chen Gao 0001, Yu Zheng 0010, Jianxin Chang, Yanan Niu, Yang Song 0008, Zhiheng Li 0001, Depeng Jin, Yong Li 0008 |
WWW | 7 |
| 2023 | Mastering Arterial Traffic Signal Control With Multi-Agent Attention-Based Soft Actor-Critic ModelabstractRecent studies have made dozens of attempts to apply multi-agent deep reinforcement learning (MARL) for large-scale traffic signal control. However, most related studies have ignored how to master arterial traffic signal control. We cannot easily extract useful information and search solution space because the arterial traffic control problem has large state-action spaces. Here we tackle these issues by proposing a multi-agent attention-base soft actor-critic (MASAC) model to master arterial traffic control. Specifically, we implement the attention mechanism in the actor and critic network to enhance traffic information extraction ability. More importantly, we are the first to apply the soft actor-critic (SAC) algorithm to train the arterial traffic control model to search more solution spaces. Testing results indicate that the MASAC method significantly outperforms existing MARL algorithms and the multiband-based method. These findings can help researchers to design better model structures for other MARL problems. Feng Mao, Zhiheng Li 0001, Yilun Lin 0002, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Adaptive Range Guided Multi-view Depth Estimation with Normal Ranking Loss
Yikang Ding, Dihe Huang, Kai Zhang 0012, Zhiheng Li 0001, Wensen Feng |
ACCV (1) | 5 |
| 2022 | Enhancing Multi-View Stereo with Contrastive Matching and Weighted Focal LossabstractLearning-based multi-view stereo (MVS) methods have made impressive progress and surpassed traditional methods in recent years. However, their accuracy and completeness are still struggling. In this paper, we propose a new method to enhance the performance of existing networks inspired by contrastive learning and feature matching. First, we propose a Contrast Matching Loss (CML), which treats the correct matching points in depth-dimension as positive sample and other points as negative samples, and computes the contrastive loss based on the similarity of features. We further propose a Weighted Focal Loss (WFL) for better classification capability, which weakens the contribution of low-confidence pixels in unimportant areas to the loss according to predicted confidence. Extensive experiments performed on DTU, Tanks and Temples and BlendedMVS datasets show our method achieves state-of-the-art performance and significant improvement over baseline network. Yikang Ding, Dihe Huang, Zhiheng Li 0001, Kai Zhang 0012 |
ICIP | 4 |
| 2022 | Dual Contrastive Network for Sequential RecommendationabstractWidely applied in today's recommender systems, sequential recommendation predicts the next interacted item for a given user via his/her historical item sequence. However, sequential recommendation suffers data sparsity issue like most recommenders. To extract auxiliary signals from the data, some recent works exploit self-supervised learning to generate augmented data via dropout strategy, which, however, leads to sparser sequential data and obscure signals. In this paper, we propose D ual C ontrastive N etwork (DCN) to boost sequential recommendation, from a new perspective of integrating auxiliary user-sequence for items. Specifically, we propose two kinds of contrastive learning. The first one is the dual representation contrastive learning that minimizes the distances between embeddings and sequence-representations of users/items. The second one is the dual interest contrastive learning which aims to self-supervise the static interest with the dynamic interest of next item prediction via auxiliary training. We also incorporate the auxiliary task of predicting next user for a given item's historical user sequence, which can capture the trends of items preferred by certain types of users. Experiments on benchmark datasets verify the effectiveness of our proposed method. Further ablation study also illustrates the boosting effect of the proposed components upon different sequential models. Guanyu Lin, Chen Gao 0001, Yinfeng Li, Yu Zheng 0010, Zhiheng Li 0001, Depeng Jin, Yong Li 0008 |
SIGIR | 5 |
| 2022 | An RSU Deployment Strategy Based on Traffic Demand in Vehicular Ad Hoc Networks (VANETs)abstractThe rapid development of connected automatic vehicle (CAV) technology makes vehicularad hocnetworks (VANETs) an urgently needed research field. It includes vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) message flows. A roadside unit (RSU) is an important infrastructure for V2I communication and provides roadside information services for CAVs. However, an unoptimal RSU deployment may result in RSUs failing to improve the efficiency of VANETs and compromising the capability of service to most vehicles. Motivated by this observation, this study focuses on balancing the two objectives of efficiency and coverage and establishing an RSU deployment strategy based on traffic demand. In detail, this model optimizes both the average data delivery delay in VANETs and the number of vehicles covered by RSUs. The effectiveness of the method is verified by simulation in a 4 km${\times }4$km virtual road network. We also found that: 1) if 25% of the road segments in the road network are covered by RSUs, most vehicles can be served, and the delay of VANETs can be reduced; 2) compared with the road network with low traffic demand, more RSUs need to be deployed in the road network with high traffic demand to achieve the same effect; and 3) early RSU investment is more cost effective. Our method can provide a reference for the areas where RSU investments should be made and the priority of the areas. Haiyang Yu 0002, Runkun Liu, Zhiheng Li 0001, Yilong Ren, Han Jiang 0003 |
IEEE Internet Things J. | 3 |
| 2022 | Three Principles to Determine the Right-of-Way for AVs: Safe Interaction With HumansabstractAutonomous vehicles (AVs) are widely believed to be good for improving transportation safety and efficiency. However, recent fatal accidents by some of their prototypes remind us that there are no operationalizable and quantitative safe driving strategies available for an AV in a wide range of situations to avoid collisions. In contrast with many recent studies that focused on ethical considerations when AVs are facing unavoidable harms, we study how to proactively prevent collisions by setting up a set of decision rules for AVs to determine the right-of-way efficiently. Notably, we summarize three essential principles for AVs designing to increase driving safety, and establish a rule-based nine-step communication-decision model to implement them. Our method is constructed by analyzing how human drivers solve potential conflicts. The decision rules are designed to be ambiguity-free and readily computable with the least communication so that human drivers and AVs could easily understand each other in terms of their behaviors and intentions of. We have demonstrated the effectiveness of our method by comparing it with some alternative approaches. Li Li 0013, Can Zhao 0004, Xiao Wang 0002, Zhiheng Li 0001, Long Chen 0005, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | A Flexible and Explainable Vehicle Motion Prediction and Inference Framework Combining Semi-Supervised AOG and ST-LSTMabstractAccurate trajectory prediction of surrounding vehicles is important for automated vehicles. To solve several existing problems of maneuver-based trajectory prediction, we propose four targeted solutions and establish a trajectory prediction model that integrates semi-supervised And-or Graph (AOG) and Spatio-temporal LSTM (ST-LSTM). To reduce the dependence on the well-labeled dataset, we introduce the concept of sub-maneuvers to improve the classifications of vehicle movements based on the given rough maneuver labels. AOG is used as the backbone of the probabilistic motion inference considering sub-maneuvers. We only define the basic units and inference logics of AOG and design a semi-supervised approach to directly learn the sub-maneuvers and the inference model structure from the training data, without manually specifying the structure (layers and nodes) of the inference model. This approach helps to avoid excessive artificial design or biases. The learned hierarchical motion inference model improves the interpretability of the overall trajectory prediction process. To utilize vehicle interaction information and further yield more accurate prediction, we adopt two different methods to consider vehicle interaction in the two sub-models (maneuver recognition and trajectory prediction). The experiment on NGSIM I-80 dataset shows that the maneuver-based model proposed in this paper (AOG-ST and refined AOG-ST-TB) performs more accurate trajectory prediction results. Although the AOG-ST seems clumsy and slow, we show that it is a flexible and quick model for trajectory prediction for various driving scenarios through the discussion and experiment. Shengzhe Dai, Zhiheng Li 0001, Li Li 0013, Nanning Zheng 0001, Shuofeng Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Harmonious Lane Changing via Deep Reinforcement LearningabstractIn this paper, we study how to learn a harmonious deep reinforcement learning (DRL) based lane-changing strategy for autonomous vehicles without Vehicle-to-Everything (V2X) communication support. The basic framework of this paper can be viewed as a multi-agent reinforcement learning in which different agents will exchange their strategies after each round of learning to reach a zero-sum game state. Unlike cooperation driving, harmonious driving only relies on individual vehicles’ limited sensing results to balance overall and individual efficiency. Specifically, we propose a well-designed reward that combines individual efficiency with overall efficiency for harmony, instead of only emphasizing individual interests like competitive strategy. Testing results show that competitive strategy often leads to selfish lane change behaviors, anarchy of crowd, and thus the degeneration of traffic efficiency. In contrast, the proposed harmonious strategy can promote traffic efficiency in both free flow and traffic jam than the competitive strategy. This interesting finding indicates that we should take care of the reward setting for reinforcement learning-based AI robots (e.g., automated vehicles) design, when the utilities of these robots are not strictly in alignment. Jianming Hu, Zhiheng Li 0001, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Systematic Solution of Human Driving Behavior Modeling and Simulation for Automated Vehicle StudiesabstractThough automated vehicles (AVs) are believed to play a crucial role in future transport, human driving vehicles will share the road with automated vehicles for a relatively long period. So, we need to enable automated vehicles to run along with human drivers especially when they may have conflicts in the right of way. One key problem is how to appropriately model human driving behaviors and quickly simulate their actions when training/testing automated vehicles. Many existing models were originally built for traffic flow studies and may not be suitable for automated vehicles studies. In this paper, we propose a set of new principles of human driving behaviors modeling and simulations. Then, we propose a Data-Driven Simulator (D2Sim) model for human behavior learning, description, and vehicle interaction simulation. In contrast to conventional microscopic traffic flow models, the D2Sim is a trajectory generation model that accepts rich driving environment information (e.g., lane geometry, crosswalks, traffic signals, surrounding vehicles, etc.). Different from many empirical trajectory records replay models, we can arbitrarily set the long-term intentions of the simulated vehicles and intentionally design the corner cases that had not been observed in practice. In addition, the D2Sim adopts adversarial learning to comprehend complex yet stochastic human driving behaviors from empirical data. Testing results show that the proposed model can quickly generate high-resolution trajectory data for training and testing. Wenqin Zhong, Shen Li 0001, Zhiheng Li 0001, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Investigating the dynamic memory effect of human drivers via ON-LSTM
Shengzhe Dai, Zhiheng Li 0001, Li Li 0013, Dongpu Cao, Xingyuan Dai, Yilun Lin 0002 |
Sci. China Inf. Sci. | 2 |
| 2017 | Parking Like a Human: A Direct Trajectory Planning SolutionabstractParking control problems remain to be fully solved for autonomous vehicles. Existing approaches usually first design a reference parking trajectory that does not exactly match vehicle dynamic constraints and then apply certain online negative feedback control to make the vehicle roughly track this reference trajectory. In this paper, we propose a novel trajectory planning method that directly links the actual parking trajectories and the steering actions to find the best parking trajectory. Tests show that this new approach has high reliability and less computation cost. Moreover, we also discuss how to counter with trajectory planning errors that are caused by model uncertainty in this paper. We show that an appropriate combination of feedforward trajectory planning and online feedback control can solve such problems. Wei Liu 0085, Zhiheng Li 0001, Li Li 0013, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Trend Modeling for Traffic Time Series Analysis: An Integrated StudyabstractThis paper discusses the trend modeling for traffic time series. First, we recount two types of definitions for a long-term trend that appeared in previous studies and illustrate their intrinsic differences. We show that, by assuming an implicit temporal connection among the time series observed at different days/locations, the PCA trend brings several advantages to traffic time series analysis. We also describe and define the so-called short-term trend that cannot be characterized by existing definitions. Second, we sequentially review the role that trend modeling plays in four major problems in traffic time series analysis: abnormal data detection, data compression, missing data imputation, and traffic prediction. The relations between these problems are revealed, and the benefit of detrending is explained. For the first three problems, we summarize our findings in the last ten years and try to provide an integrated framework for future study. For traffic prediction problem, we present a new explanation on why prediction accuracy can be improved at data points representing the short-term trends if the traffic information from multiple sensors can be appropriately used. This finding indicates that the trend modeling is not only a technique to specify the temporal pattern but is also related to the spatial relation of traffic time series. Li Li 0013, Xiaonan Su, Yi Zhang 0029, Yuetong Lin, Zhiheng Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2013 | Freeway Travel-Time Estimation Based on Temporal-Spatial Queueing ModelabstractTravel time serves as a fundamental measurement for transportation systems and becomes increasingly important to both drivers and traffic operators. Existing speed interpolation algorithms use the average speed time series collected from upstream and downstream detectors to estimate the travel time of a road link. Such approaches often result in inaccurate estimations or even systematic bias, particularly when the real travel times quickly vary. To get rid of this problem, Coifman proposed a creative interpolation algorithm based on kinetic-wave models. This algorithm reconstructs vehicle trajectories according to the velocities and the headways of vehicles. However, it sometimes gives significant biased estimation, particularly when jams emerge from somewhere between the upstream and downstream detectors. To make an amendment, we design a new algorithm based on the temporal-spatial queueing model to describe the fast travel-time variations using only the speed and headway time series that is measured at upstream and downstream detectors. Numerical studies show that this new interpolation algorithm could better utilize the dynamic traffic flow information that is embedded in the speed/headway time series in some special cases. Li Li 0013, Xiqun Chen, Zhiheng Li 0001, Lei Zhang 0118 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2012 | Phase Diagram Analysis Based on a Temporal-Spatial Queueing ModelabstractIn this paper, we propose a simple temporal-spatial queueing model to quantitatively address some typical congestion patterns that were observed around on/off-ramps. In particular, we examine three prime factors that play important roles in ramping traffic scenarios: the time τinfor a vehicle to join a jam queue, the time τoutfor this vehicle to depart from this jam queue, and the time intervalTfor the ramping vehicle to merge into the mainline. Based on Newell's simplified car-following model, we show how τinchanges with the main road flow rateqmain. Meanwhile,Tis the reciprocal of the ramping road flow rateqramp. Thus, we analytically derive the macroscopic phase diagram plotted on theqmain-versus-qrampplane and τin-versus-Tplane based on the proposed model. Further study shows that the new queueing model not only reserves the merits of Newell's model on the microscopic level but helps quantify the contributions of these parameters in characterizing macroscopic congestion patterns as well. Previous approaches distinguished phases merely through simulations, but our model could derive analytical boundaries for the phases. The phase transition conditions obtained by this model agree well with simulations and empirical observations. These findings help reveal the origins of some well-known phenomena during traffic congestion. Xiqun Chen, Li Li 0013, Zhiheng Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2009 | Quantization Errors of Uniformly Quantized fGn and fBm SignalsabstractIn this letter, we show that under the assumption of high resolution, the quantization errors of fGn and fBm signals with uniform quantizer can be treated as uncorrelated white noises. Zhiheng Li 0001, Yudong Chen 0001, Li Li 0013, Yi Zhang 0029 |
IEEE Signal Process. Lett. | 1 |
| 2004 | Spatial-temporal traffic data analysis based on global data management using MASabstractThe spatial-temporal traffic data analysis based on global data management is a newly developed and crucial approach to help traffic managers having the global view of urban traffic status in the level of road network, which is very clearly useful in traffic control and route guidance. The multiagent systems are used in traffic data management with full consideration of the characteristics of traffic data and the cooperation and workflow among them. In software implementation of data management, the agent-based common object request broker architecture is adopted taking the distributed urban traffic data in the large area under network environments into account. Based on the global traffic data, the approach of visualized spatial-temporal analysis is then induced. The similarity of traffic data is analyzed first for each link and its profile is achieved to undertake the primary processing of urban traffic data. Furthermore, analysis results are shown on the basis of the geographic information systems for transportation. The two types of visualization, pseudocolor and contour maps, are adopted in the demonstration to display the traffic status graphically and its changing frames. Among the applications in some big cities in China, the case of urban traffic analysis for Beijing is studied to demonstrate the implementation of the approach. Yi Zhang 0029, Zhiheng Li 0001, Dongcheng Hu |
IEEE Trans. Intell. Transp. Syst. | 3 |