Zhenning Li 0001

dblp:16/8388-1 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
31since 2021 · last 2026
0000-0002-0877-6829ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 11 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Predict and Resist: Long-Term Accident Anticipation Under Sensor Noise
abstract
Accident anticipation is essential for proactive and safe autonomous driving, where even a brief advance warning can enable critical evasive actions. However, two key challenges hinder real-world deployment: (1) noisy or degraded sensory inputs from weather, motion blur, or hardware limitations, and (2) the need to issue timely yet reliable predictions that balance early alerts with false-alarm suppression. We propose a unified framework that integrates diffusion-based denoising with a time-aware actor-critic model to address these challenges. The diffusion module reconstructs noise-resilient image and object features through iterative refinement, preserving critical motion and interaction cues under sensor degradation. In parallel, the actor-critic architecture leverages long-horizon temporal reasoning and time-weighted rewards to determine the optimal moment to raise an alert, aligning early detection with reliability. Experiments on three benchmark datasets (DAD, CCD, A3D) demonstrate state-of-the-art accuracy and significant gains in mean time-to-accident, while maintaining robust performance under Gaussian and impulse noise. Qualitative analyses further show that our model produces earlier, more stable, and human-aligned predictions in both routine and highly complex traffic scenarios, highlighting its potential for real-world, safety-critical deployment.
Xingcheng Liu, Bin Rao 0003, Yanchen Guan, Chengyue Wang 0001, Haicheng Liao, Jiaxun Zhang, Chengyu Lin 0003, Meixin Zhu, Zhenning Li 0001
AAAI9
2026 Differentiable Semantic Meta-Learning Framework for Long-Tail Motion Forecasting in Autonomous Driving
abstract
Long-tail motion forecasting is a core challenge for autonomous driving, where rare yet safety-critical events-such as abrupt maneuvers and dense multi-agent interactions-dominate real-world risk. Existing approaches struggle in these scenarios because they rely on either non-interpretable clustering or model-dependent error heuristics, providing neither a differentiable notion of “tailness” nor a mechanism for rapid adaptation. We propose SAML, a Semantic-Aware Meta-Learning framework that introduces the first differentiable definition of tailness for motion forecasting. SAML quantifies motion rarity via semantically meaningful intrinsic (kinematic, geometric, temporal) and interactive (local and global risk) properties, which are fused by a Bayesian Tail Perceiver into a continuous, uncertainty-aware Tail Index. This Tail Index drives a meta-memory adaptation module that couples a dynamic prototype memory with an MAML-based cognitive set mechanism, enabling fast adaptation to rare or evolving patterns. Experiments on nuScenes, NGSIM, and HighD show that SAML achieves state-of-the-art overall accuracy and substantial gains on top 1-5% worst-case events, while maintaining high efficiency. Our findings highlight semantic meta-learning as a pathway toward robust and safety-critical motion forecasting.
Bin Rao 0003, Chengyue Wang 0001, Haicheng Liao, Qianfang Wang, Yanchen Guan, Jiaxun Zhang, Xingcheng Liu, Meixin Zhu, Kanye Ye Wang, Zhenning Li 0001
AAAI10
2026 Fusion Filtering for RIS-Assisted Vehicle Localization With Unknown Inputs: A Privacy-Preserving Output-Mask Strategy
abstract
This article investigates the privacy-preserving fusion filtering problem for vehicle localization subject to unknown inputs. High-accuracy and privacy-preserving localization is essential for the safe and reliable operation of intelligent transportation systems. In practice, vehicle localization is often challenged by measurement bias caused by nonline-of-sight (NLOS) propagation, unknown inputs arising from uncertainties or acceleration/deceleration maneuvers, and risks of signal and location privacy leakage. To address these issues, a privacy-preserving fusion filtering framework is proposed by integrating the reconfigurable intelligent surface (RIS) technique, an unknown-input estimation method, and an output-mask mechanism. A unified measurement model is first developed to represent both line-of-sight (LOS) and NLOS scenarios, and RISs are employed to construct virtual LOS paths to mitigate NLOS effects. An output-mask-based privacy-preserving strategy is then designed to prevent eavesdroppers from inferring vehicle locations or signal characteristics while maintaining the required filtering performance. Based on this model, an unknown-input estimator and a privacy-preserving filter are constructed, and the impacts of NLOS propagation, unknown inputs, and masking on filtering performance are analyzed. The associated gain matrix parameters are obtained by solving the corresponding optimization problems. Furthermore, an RIS-assisted privacy-preserving fusion filtering algorithm is developed to exploit multisource measurements and enhance localization robustness and accuracy. Simulation results demonstrate the effectiveness of the proposed method.
Kaiqun Zhu, Zidong Wang 0001, Xinhu Zheng, Zhiyong Cui, Zhenning Li 0001, Keqiang Li 0002
IEEE Trans. Ind. Informatics5
2026 Enhancing Trust Management System for Connected Autonomous Vehicles Using Machine Learning Methods: A Survey
abstract
Connected Autonomous Vehicles (CAVs) operate in dynamic, open, and multi-domain networks, rendering them vulnerable to various threats. Trust Management Systems (TMS) systematically organize the essential steps in the trust mechanism, identifying malicious nodes against both internal and external threats, while ensuring reliable decision-making for more cooperative tasks. Recent advances in machine learning (ML) offer significant potential to enhance TMS, particularly for the stringent requirements of CAVs, such as CAV nodes moving at varying speeds and exhibiting opportunistic and intermittent network behavior. Those features distinguish ML-based TMS from social networks, static IoT, and Social IoT. This survey proposes a novel three-layer ML-based TMS framework for CAVs in the vehicle-road-cloud integration system, comprising a trust data layer, a trust calculation layer, and a trust incentive layer. A six-dimensional taxonomy of objectives is proposed. Furthermore, the principles of ML methods for each module in each layer are analyzed. Then, recent studies are categorized based on traffic scenarios that are against the proposed objectives. The survey concludes by proposing future research directions that address open issues and align with current trends. An accompanying repository of state-of-the-art works and open-source projects is available at:https://github.com/octoberzzzzz/ML-based-TMS-CAV-Survey
Qian Xu 0020, Lei Zhang 0110, Zhenning Li 0001
IEEE Trans. Intell. Transp. Syst.4
2025 NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous Driving
abstract
Accurate trajectory prediction is essential for the safety and efficiency of autonomous driving. Traditional models often struggle with real-time processing, capturing non-linearity and uncertainty in traffic environments, efficiency in dense traffic, and modeling temporal dynamics of interactions. We introduce NEST (Neuromodulated Small-world Hypergraph Trajectory Prediction), a novel framework that integrates Small-world Networks and hypergraphs for superior interaction modeling and prediction accuracy. This integration enables the capture of both local and extended vehicle interactions, while the Neuromodulator component adapts dynamically to changing traffic conditions. We validate the NEST model on several real-world datasets, including nuScenes, MoCAD, and HighD. The results consistently demonstrate that NEST outperforms existing methods in various traffic scenarios, showcasing its exceptional generalization capability, efficiency, and temporal foresight. Our comprehensive evaluation illustrates that NEST significantly improves the reliability and operational efficiency of autonomous driving systems, making it a robust solution for trajectory prediction in complex traffic environments.
Chengyue Wang 0001, Haicheng Liao, Bonan Wang, Yanchen Guan, Bin Rao 0003, Ziyuan Pu, Zhiyong Cui, Cheng-Zhong Xu 0001, Zhenning Li 0001
AAAI9
2025 AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction
abstract
Accurately predicting the future trajectories of traffic agents is essential in autonomous driving. However, due to the inherent imbalance in trajectory distributions, tail data in natural datasets often represents more complex and hazardous scenarios. Existing studies typically rely solely on a base model's prediction error, without considering the diversity and uncertainty of long-tail trajectory patterns. We propose an adaptive momentum and decoupled contrastive learning framework (AMD), which integrates unsupervised and supervised contrastive learning strategies. By leveraging an improved momentum contrast learning (MoCo-DT) and decoupled contrastive learning (DCL) module, our framework enhances the model's ability to recognize rare and complex trajectories. Additionally, we design four types of trajectory random augmentation methods and introduce an online iterative clustering strategy, allowing the model to dynamically update pseudo-labels and better adapt to the distributional shifts in long-tail data. We propose three different criteria to define long-tail trajectories and conduct extensive comparative experiments on the nuScenes and ETH$/$UCY datasets. The results show that AMD not only achieves optimal performance in long-tail trajectory prediction but also demonstrates outstanding overall prediction accuracy.
Bin Rao 0003, Haicheng Liao, Yanchen Guan, Chengyue Wang 0001, Bonan Wang, Jiaxun Zhang, Zhenning Li 0001
ICCV7
2025 SAH-Drive: A Scenario-Aware Hybrid Planner for Closed-Loop Vehicle Trajectory Generation
abstract
Reliable planning is crucial for achieving autonomous driving. Rule-based planners are efficient but lack generalization, while learning-based planners excel in generalization yet have limitations in real-time performance and interpretability. In long-tail scenarios, these challenges make planning particularly difficult. To leverage the strengths of both rule-based and learning-based planners, we proposed the Scenario-Aware Hybrid Planner (SAH-Drive) for closed-loop vehicle trajectory planning. Inspired by human driving behavior, SAH-Drive combines a lightweight rule-based planner and a comprehensive learning-based planner, utilizing a dual-timescale decision neuron to determine the final trajectory. To enhance the computational efficiency and robustness of the hybrid planner, we also employed a diffusion proposal number regulator and a trajectory fusion module. The experimental results show that the proposed method significantly improves the generalization capability of the planning system, achieving state-of-the-art performance in interPlan, while maintaining computational efficiency without incurring substantial additional runtime.
Yuqi Fan 0003, Zhiyong Cui, Zhenning Li 0001, Yilong Ren, Haiyang Yu 0002
ICML3
2025 DRIVE: Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous Driving
abstract
Recent advancements in autonomous driving have seen a paradigm shift towards end-to-end learning paradigms, which map sensory inputs directly to driving actions, thereby enhancing the robustness and adaptability of autonomous vehicles. However, these models often sacrifice interpretability, posing significant challenges to trust, safety, and regulatory compliance. To address these issues, we introduce DRIVE – Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous Driving, a comprehensive framework designed to improve the dependability and stability of explanations in end-to-end unsupervised autonomous driving models. Our work specifically targets the inherent instability problems observed in the Driving through the Concept Gridlock (DCG) model, which undermine the trustworthiness of its explanations and decisionmaking processes. We define four key attributes of DRIVE: consistent interpretability, stable interpretability, consistent output, and stable output. These attributes collectively ensure that explanations remain reliable and robust across different scenarios and perturbations. Through extensive empirical evaluations, we demonstrate the effectiveness of our framework in enhancing the stability and dependability of explanations, thereby addressing the limitations of current models. Our contributions include an in-depth analysis of the dependability issues within the DCG model, a rigorous definition of DRIVE with its fundamental properties, a framework to implement DRIVE, and novel metrics for evaluating the dependability of concept-based explainable autonomous driving models. These advancements lay the groundwork for the development of more reliable and trusted autonomous driving systems, paving the way for their broader acceptance and deployment in real-world applications. “We can only see a short distance ahead, but we can see plenty there that needs to be done.” – Alan Turing
Songning Lai, Tianlang Xue, Hongru Xiao, Lijie Hu, Jiemin Wu, Ninghui Feng, Runwei Guan, Haicheng Liao, Zhenning Li 0001, Yutao Yue
ICRA9
2025 Beyond Patterns: Harnessing Causal Logic for Autonomous Driving Trajectory Prediction
abstract
Accurate trajectory prediction has long been a major challenge for autonomous driving (AD). Traditional data-driven models predominantly rely on statistical correlations, often overlooking the causal relationships that govern traffic behavior. In this paper, we introduce a novel trajectory prediction framework that leverages causal inference to enhance predictive robustness, generalization, and accuracy. By decomposing the environment into spatial and temporal components, our approach identifies and mitigates spurious correlations, uncovering genuine causal relationships. We also employ a progressive fusion strategy to integrate multimodal information, simulating human-like reasoning processes and enabling real-time inference. Evaluations on five real-world datasets—ApolloScape, nuScenes, NGSIM, HighD, and MoCAD—demonstrate our model's superiority over existing state-of-the-art (SOTA) methods, with improvements in key metrics such as RMSE and FDE. Our findings highlight the potential of causal reasoning to transform trajectory prediction, paving the way for robust AD systems.
Bonan Wang, Haicheng Liao, Chengyue Wang 0001, Bin Rao 0003, Yanchen Guan, Guyang Yu, Jiaxun Zhang, Songning Lai, Cheng-Zhong Xu 0001, Zhenning Li 0001
IJCAI10
2025 Eyes on the Road, Mind Beyond Vision: Context-Aware Multi-modal Enhanced Risk Anticipation
abstract
Accurate accident anticipation remains challenging when driver cognition and dynamic road conditions are underrepresented in predictive models. In this paper, we propose CAMERA (Context-Aware Multi-modal Enhanced Risk Anticipation), a multi-modal framework integrating dashcam video, textual annotations, and driver attention maps for robust accident anticipation. Unlike existing methods that rely on static or environment-centric thresholds, CAMERA employs an adaptive mechanism guided by scene complexity and gaze entropy, reducing false alarms while maintaining high recall in dynamic, multi-agent traffic scenarios. A hierarchical fusion pipeline with Bi-GRU (Bidirectional GRU) captures spatio-temporal dependencies, while a Geo-Context Vision-Language module translates 3D spatial relationships into interpretable, human-centric alerts. Evaluations on the DADA-2000 and benchmarks show that CAMERA achieves state-of-the-art performance, improving accuracy and lead time. These results demonstrate the effectiveness of modeling driver attention, contextual description, and adaptive risk thresholds to enable more reliable accident anticipation.
Jiaxun Zhang, Haicheng Liao, Yumu Xie, Chengyue Wang 0001, Yanchen Guan, Bin Rao 0003, Zhenning Li 0001
ACM Multimedia7
2025 Alternating interaction fusion of Image-Point cloud for Multi-Modal 3D object detection
Guofa Li, Haifeng Lu, Jie Li 0042, Zhenning Li 0001, Qingkun Li, Xiangyun Ren
Adv. Eng. Informatics4
2025 WAKE: Towards Robust and Physically Feasible Trajectory Prediction for Autonomous Vehicles With WAvelet and KinEmatics Synergy
abstract
Addressing the pervasive challenge of imperfect data in autonomous vehicle (AV) systems, this study pioneers an integrated trajectory prediction model, WAKE, that fuses physics-informed methodologies with sophisticated machine learning techniques. Our model operates in two principal stages: the initial stage utilizes a Wavelet Reconstruction Network to accurately reconstruct missing observations, thereby preparing a robust dataset for further processing. This is followed by the Kinematic Bicycle Model which ensures that reconstructed trajectory predictions adhere strictly to physical laws governing vehicular motion. The integration of these physics-based insights with a subsequent machine learning stage, featuring a Quantum Mechanics-Inspired Interaction-aware Module, allows for sophisticated modeling of complex vehicle interactions. This fusion approach not only enhances the prediction accuracy but also enriches the model's ability to handle real-world variability and unpredictability. Extensive tests using specific versions of MoCAD, NGSIM, HighD, INTERACTION, and nuScenes datasets featuring missing observational data, have demonstrated the superior performance of our model in terms of both accuracy and physical feasibility, particularly in scenarios with significant data loss-up to 75% missing observations. Our findings underscore the potency of combining physics-informed models with advanced machine learning frameworks to advance autonomous driving technologies, aligning with the interdisciplinary nature of information fusion.
Chengyue Wang 0001, Haicheng Liao, Zhenning Li 0001, Cheng-Zhong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Secure Observer-Based Collision-Free Control for Autonomous Vehicles Under Non-Gaussian Noises
abstract
This article is concerned with the secure collision-free tracking control problem for autonomous vehicles with uncertainties, where system signals are transmitted through constrained communication networks. In such open and uncertain environments, the control performance of vehicles is seriously affected by privacy leakage, non-Gaussian noise, and obstacles. The aim of this research is to propose a tracking control scheme that ensures security, mean-square boundedness, and collision-free performance concurrently. Initially, to safeguard the privacy of the transmitted data and to enable secure tracking control, a dynamic encoding-based ElGamal encryption mechanism is introduced, which is further embedded in the design of the observer-based tracking controller. Subsequently, a collision-free chance-constrained index is proposed for achieving real-time obstacle avoidance by comprehensively considering the influence of stochastic noises. A thorough analysis is conducted to examine the impact of non-Gaussian noise and unmeasurable states on the performance of collision-free tracking control. Sufficient conditions are derived to guarantee the desired performance, and the corresponding control inputs are obtained by solving certain optimization problems subject to chance constraints. Finally, an illustrative example is provided to validate the effectiveness of the proposed secure collision-free tracking controller.
Kaiqun Zhu, Zidong Wang 0001, Zhenning Li 0001, Cheng-Zhong Xu 0001
IEEE Trans. Ind. Informatics3
2025 Digital Twin-Based Driver Risk-Aware Predictive Mobility Analytics for Real-Time Situational Awareness Through Cooperative Sensing
abstract
Traffic safety risk significantly impacts road users in urban mobility systems, making it important in transportation management decision-making. Current mobility management strategies predominantly focus on macro-level monitoring through traffic sensing infrastructure, and struggle to capture network-wide, real-time safety risks due to the limited spread of vehicle-based sensors. To address this, we propose a Digital Twin-based Driver Risk-aware Predictive Mobility Analytics (DT-DIMA) system. The DT-DIMA system integrates real-time traffic information from pan-tilt-cameras (PTCs), synchronizes this data into a digital twin to accurately replicate the physical world, and predicts network-wide mobility and safety risks in real time. The system’s innovation lies in its integration of spatial-temporal modeling, simulation, and online control modules. Tested and evaluated under normal traffic conditions and incidental situations (e.g., unexpected accidents, pre-planned work zones) in a simulated testbed in Brooklyn, New York, DT-DIMA demonstrated mean absolute percentage errors (MAPEs) ranging from 9.40% to 13.12% in estimating network-level traffic volume and MAPEs from 2.12% to 12.97% in network-level safety risk prediction. In addition, the highly accurate safety risk prediction enables PTCs to preemptively monitor road segments with high driving risks before incidents take place. Such proactive PTC surveillance creates around a 5-minute lead time in capturing traffic incidents. The DT-DIMA system enables transportation managers to understand mobility not only in terms of traffic patterns but also driver-experienced safety risks, allowing for proactive resource allocation in response to various traffic situations.
Tao Li 0046, Zilin Bian, Haozhe Lei, Fan Zuo, Ya-Ting Yang, Quanyan Zhu, Zhenning Li 0001, Zhibin Chen 0001, Kaan Özbay
IEEE Trans. Intell. Transp. Syst.7
2025 Cross-Driver Domain Generalization for Improved Drowsiness Recognition Based on EEG Signals
abstract
Designing brain-computer interface systems for electroencephalogram (EEG)-based driver drowsiness recognition remains a significant challenge due to the significant variation in EEG signals across subjects and recording sessions. To address this problem, this paper develops a novel two-stage information transfer strategy framework for domain generalization. The framework has two domain mappers to reduce the distribution differences of EEG features from different individuals, a mapper mix block for generating hybrid mapping features, and a domain adversarial neural network (DANN) for drowsiness recognition based on hybrid EEG features. In the process of DANN to capture common features, we additionally employ two models based on self-attention mechanism to capture domain-invariant attention relationships between electrode channels and between frequency bands. Experimental results show that the proposed framework achieves an average accuracy of 81.34% in the leave-one-out cross validation for driver drowsiness recognition, which is higher than the state-of-the-art model with the number of 79.37%. In addition, we explore the impact of EEG features from different frequency bands and brain regions on this cross-subject task. The results show that EEG features from delta, theta and alpha bands can achieve much better performance than the other two bands, and features from the frontal lobe region perform better than the other regions. These findings reveal domain-invariant features and their relationships with brain regions and frequency bands, enhancing our understanding of the underlying messages of EEG signals.
Guofa Li, Delin Ouyang, Qingkun Li, Zhenning Li 0001, Shengbo Eben Li, Cristina Olaverri-Monreal
IEEE Trans. Intell. Transp. Syst.5
2025 Chain-of-Thought Guided Multimodal Large Language Models for Scene-Aware Accident Anticipation in Autonomous Driving
abstract
Accurately anticipating traffic accidents is a fundamental task for the safe and effective deployment of autonomous vehicles (AVs). However, existing models primarily rely on dashcam footage and often fail to generalize across varied driving scenarios due to their dependence on visual data and the rarity of high-risk events in datasets. These limitations undermine their robustness and reduce practical applicability in dynamic, unpredictable environments. To address these challenges, this study proposes a novel approach, termed MLTA, which integrates multimodal learning with the hypergraph attention network to hierarchically extract and capture cross-modal interaction. It leverages LLava-next, a multimodal large language model (MLLM) guided by the Chain-of-Thought (CoT) prompting paradigm, to produce context-aware interpretations of traffic scenes. This is further enhanced by a human-inspired attention mechanism that mimics the decision-making priorities of experienced human drivers. This combination enables more accurate identification of critical elements in a scene, improving both prediction precision and timeliness. Extensive experiments on four real-world datasets—DAD, A3D, CCD, and DADA-2000—show that our approach consistently outperforms state-of-the-art (SOTA) methods, demonstrating strong adaptability and robustness in complex driving environments.
Haicheng Liao, Bin Rao 0003, Chengyue Wang 0001, Shengbo Eben Li, Cheng-Zhong Xu 0001, Zhenning Li 0001
IEEE Trans. Intell. Transp. Syst.8
2025 Minds on the Move: Decoding Trajectory Prediction in Autonomous Driving With Cognitive Insights
abstract
In mixed autonomous driving environments, accurately predicting the future trajectories of surrounding vehicles is crucial for the safe operation of autonomous vehicles (AVs). In driving scenarios, a vehicle’s trajectory is determined by the decision-making process of human drivers. However, existing models primarily focus on the inherent statistical patterns in the data, often neglecting the critical aspect of understanding the decision-making processes of human drivers. This oversight results in models that fail to capture the true intentions of human drivers, leading to suboptimal performance in long-term trajectory prediction. To address this limitation, we introduce a Cognitive-Informed Transformer (CITF) that incorporates a cognitive concept, Perceived Safety, to interpret drivers’ decision-making mechanisms. Perceived Safety encapsulates the varying risk tolerances across drivers with different driving behaviors. Specifically, we develop a Perceived Safety-aware Module that includes a Quantitative Safety Assessment for measuring the subject risk levels within scenarios, and Driver Behavior Profiling for characterizing driver behaviors. Furthermore, we present a novel module, Leanformer, designed to capture social interactions among vehicles. CITF demonstrates significant performance improvements on three well-established datasets. In terms of long-term prediction, it surpasses existing benchmarks by 12.0% on the NGSIM, 28.2% on the HighD, and 20.8% on the MoCAD dataset. Additionally, its robustness in scenarios with limited or missing data is evident, surpassing most state-of-the-art (SOTA) baselines, and paving the way for real-world applications.
Haicheng Liao, Chengyue Wang 0001, Kaiqun Zhu, Yilong Ren, Bolin Gao, Shengbo Eben Li, Cheng-Zhong Xu 0001, Zhenning Li 0001
IEEE Trans. Intell. Transp. Syst.8
2025 Human-Machine Shared Control Approach for the Takeover of Cooperative Adaptive Cruise Control
abstract
Cooperative Adaptive Cruise Control (CACC) often requires human takeover for tasks such as exiting a freeway. Direct human takeover can pose significant risks, especially given the close-following strategy employed by CACC, which might cause drivers to feel unsafe and execute hard braking, potentially leading to collisions. This research aims to develop a CACC takeover controller that ensures a smooth transition from automated to human control. The proposed CACC takeover maneuver employs an indirect human-machine shared control approach, modeled as a Stackelberg competition where the machine acts as the leader and the human as the follower. The machine guides the human to respond in a manner that aligns with the machine’s expectations, aiding in maintaining following stability. Additionally, the human reaction function is integrated into the machine’s predictive control system, moving beyond a simple “prediction-planning” pipeline to enhance planning optimality. The controller has been verified to 1) enable a smooth takeover maneuver of CACC; 2) ensure string stability in the condition that the platoon has less than 6 CAVs and human control authority is less than 40%; 3) enhance both perceived and actual safety through machine interventions; and 4) reduce the impact on upstream traffic by up to 60%.
Haoran Wang 0002, Zhexi Lian, Zhenning Li 0001, Arno Eichberger, Jia Hu 0003, Yongyu Chen, Yongji Gao
IEEE Trans. Intell. Transp. Syst.3
2025 Emotions in Fandom Crowdfunding: Investigating How Online Interactions Affect Collaborative Monetary Activities
abstract
Fandom crowdfunding, where fans collectively raise funds for idols, fosters dynamic interactions within fandom communities, evoking a range of emotions. Despite the prevalence of such activities, the specific emotions involved and their effects on participant behavior remain underexplored. Addressing this, our mixed-methods study—encompassing observations, interviews, and analysis of crowdfunding data—investigated emotions during fandom crowdfunding and their influence on behavior across crowdfunding stages: planning, support, encouragement, realization, and auditing. We identified 10 key emotions related to idols and the community, finding these emotions crucial in shaping participant actions. Our findings highlight the dual impact of fandom crowdfunding on the community’s internal dynamics and its relationships with idols and broader society. We propose design recommendations for enhancing fandom crowdfunding and suggest how general crowdfunding can benefit from insights gained from the fandom context, offering a novel understanding of emotions in collaborative monetary activities.
Molly Zhuangtong Huang, Zhicong Lu, Caishi Huang, Zhenning Li 0001, Hantao Zhao, Xiaobo Zhou 0002, Dazhao Cheng, Kanye Ye Wang
ACM Trans. Comput. Hum. Interact.4
2025 SA-TP$^{2}$: A Safety-Aware Trajectory Prediction and Planning Model for Autonomous Driving
abstract
Trajectory prediction and planning remain key challenges for autonomous vehicles (AVs), particularly in complex and dynamic environments. Existing methods, typically based on static safety metrics like Time-to-Collision (TTC), fail to account for the evolving nature of risk in real-world traffic. This paper proposes a novel Safety-Aware Trajectory Prediction and Planning (SA-TP$^{2}$) model, which introduces an adaptive driver risk field to simulate human- like risk perception and decision-making. By dynamically modeling risk as a continuous variable, SA-TP$^{2}$adjusts vehicle trajectories in real-time, accounting for interactions with other agents, road conditions, and environmental uncertainties. The model integrates imitation learning (IL), rule-based strategies, and physics-informed neural networks (PINNs) to ensure safe, efficient, and human-compatible behavior. A Linformer-based architecture and Temporal Hypergraph Convolution Network (THGCN) are introduced to optimize computational efficiency, enabling real-time operation in resource-constrained environments. Experimental results on benchmark datasets including NGSIM, HighD, MoCAD, and NuScenes demonstrate that SA-TP$^{2}$achieves SOTA performance in trajectory prediction. Additionally, extensive closed-loop testing on the NuPlan and CommonRoad platforms further confirms that SA-TP$^{2}$outperforms existing baselines, paving the way for safer navigation of autonomous driving systems.
Haicheng Liao, Zhenning Li 0001, Kaiqun Zhu, Keqiang Li 0002, Cheng-Zhong Xu 0001
IEEE Trans. Robotics2
2024 BAT: Behavior-Aware Human-Like Trajectory Prediction for Autonomous Driving
abstract
The ability to accurately predict the trajectory of surrounding vehicles is a critical hurdle to overcome on the journey to fully autonomous vehicles. To address this challenge, we pioneer a novel behavior-aware trajectory prediction model (BAT) that incorporates insights and findings from traffic psychology, human behavior, and decision-making. Our model consists of behavior-aware, interaction-aware, priority-aware, and position-aware modules that perceive and understand the underlying interactions and account for uncertainty and variability in prediction, enabling higher-level learning and flexibility without rigid categorization of driving behavior. Importantly, this approach eliminates the need for manual labeling in the training process and addresses the challenges of non-continuous behavior labeling and the selection of appropriate time windows. We evaluate BAT's performance across the Next Generation Simulation (NGSIM), Highway Drone (HighD), Roundabout Drone (RounD), and Macao Connected Autonomous Driving (MoCAD) datasets, showcasing its superiority over prevailing state-of-the-art (SOTA) benchmarks in terms of prediction accuracy and efficiency. Remarkably, even when trained on reduced portions of the training data (25%), our model outperforms most of the baselines, demonstrating its robustness and efficiency in predicting vehicle trajectories, and the potential to reduce the amount of data required to train autonomous vehicles, especially in corner cases. In conclusion, the behavior-aware model represents a significant advancement in the development of autonomous vehicles capable of predicting trajectories with the same level of proficiency as human drivers. The project page is available on our GitHub.
Haicheng Liao, Zhenning Li 0001, Huanming Shen, Wenxuan Zeng, Dongping Liao, Guofa Li, Cheng-Zhong Xu 0001
AAAI2
2024 Less is More: Efficient Brain-Inspired Learning for Autonomous Driving Trajectory Prediction
abstract
Accurately and safely predicting the trajectories of surrounding vehicles is essential for fully realizing autonomous driving (AD). This paper presents the Human-Like Trajectory Prediction model (HLTP++), which emulates human cognitive processes to improve trajectory prediction in AD. HLTP++ incorporates a novel teacher-student knowledge distillation framework. The “teacher” model, equipped with an adaptive visual sector, mimics the dynamic allocation of attention human drivers exhibit based on factors like spatial orientation, proximity, and driving speed. On the other hand, the “student” model focuses on real-time interaction and human decision-making, drawing parallels to the human memory storage mechanism. Furthermore, we improve the model’s efficiency by introducing a new Fourier Adaptive Spike Neural Network (FA-SNN), allowing for faster and more precise predictions with fewer parameters. Evaluated using the NGSIM, HighD, and MoCAD benchmarks, HLTP++ demonstrates superior performance compared to existing models, which reduces the predicted trajectory error with over 11% on the NGSIM dataset and 25% on the HighD datasets. Moreover, HLTP++ demonstrates strong adaptability in challenging environments with incomplete input data. This marks a significant stride in the journey towards fully AD systems.
Haicheng Liao, Yongkang Li 0003, Zhenning Li 0001, Chengyue Wang 0001, Guofa Li, Chunlin Tian, Zilin Bian, Kaiqun Zhu, Zhiyong Cui, Jia Hu 0003
ECAI3
2024 Human Observation-Inspired Trajectory Prediction for Autonomous Driving in Mixed-Autonomy Traffic Environments
abstract
In the burgeoning field of autonomous vehicles (AVs), trajectory prediction remains a formidable challenge, especially in mixed autonomy environments. Traditional approaches often rely on computational methods such as time-series analysis. Our research diverges significantly by adopting an interdisciplinary approach that integrates principles of human cognition and observational behavior into trajectory prediction models for AVs. We introduce a novel “adaptive visual sector” mechanism that mimics the dynamic allocation of attention human drivers exhibit based on factors like spatial orientation, proximity, and driving speed. Additionally, we develop a “dynamic traffic graph” using Convolutional Neural Networks (CNN) and Graph Attention Networks (GAT) to capture spatio-temporal dependencies among agents. Benchmark tests on the NGSIM, HighD, and MoCAD datasets reveal that our model (GAVA) outperforms state-of-the-art baselines by at least 15.2%, 19.4%, and 12.0%, respectively. Our findings underscore the potential of leveraging human cognition principles to enhance the proficiency and adaptability of trajectory prediction algorithms in AVs.
Haicheng Liao, Shangqian Liu, Yongkang Li 0003, Zhenning Li 0001, Chengyue Wang 0001, Yunjian Li, Shengbo Eben Li, Cheng-Zhong Xu 0001
ICRA4
2024 CDSTraj: Characterized Diffusion and Spatial-Temporal Interaction Network for Trajectory Prediction in Autonomous Driving
Haicheng Liao, Xuelin Li, Yongkang Li 0003, Hanlin Kong, Chengyue Wang 0001, Bonan Wang, Yanchen Guan, Kahou Tam, Zhenning Li 0001
IJCAI9
2024 MFTraj: Map-Free, Behavior-Driven Trajectory Prediction for Autonomous Driving
Haicheng Liao, Zhenning Li 0001, Chengyue Wang 0001, Huanming Shen, Dongping Liao, Bonan Wang, Guofa Li, Cheng-Zhong Xu 0001
IJCAI2
2024 A Cognitive-Driven Trajectory Prediction Model for Autonomous Driving in Mixed Autonomy Environments
Haicheng Liao, Zhenning Li 0001, Chengyue Wang 0001, Bonan Wang, Hanlin Kong, Yanchen Guan, Guofa Li, Zhiyong Cui
IJCAI2
2024 Physics-Informed Trajectory Prediction for Autonomous Driving under Missing Observation
Haicheng Liao, Chengyue Wang 0001, Zhenning Li 0001, Yongkang Li 0003, Bonan Wang, Guofa Li, Cheng-Zhong Xu 0001
IJCAI3
2024 When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language Models
abstract
As autonomous driving systems increasingly become part of daily transportation, the ability to accurately anticipate and mitigate potential traffic accidents is paramount. Traditional accident anticipation models primarily utilizing dashcam videos are adept at predicting when an accident may occur but fall short in localizing the incident and identifying involved entities. Addressing this gap, this study introduces a novel framework that integrates Large Language Models (LLMs) to enhance predictive capabilities across multiple dimensions-what, when, and where accidents might occur. We develop an innovative chain-based attention mechanism that dynamically adjusts to prioritize high-risk elements within complex driving scenes. This mechanism is complemented by a three-stage model that processes outputs from smaller models into detailed multimodal inputs for LLMs, thus enabling a more nuanced understanding of traffic dynamics. Empirical validation on the DAD, CCD, and A3D datasets demonstrates superior performance in Average Precision (AP) and Mean Time-To-Accident (mTTA), establishing new benchmarks for accident prediction technology. Our approach not only advances the technological framework for autonomous driving safety but also enhances human-AI interaction, making predictive insights generated by autonomous systems more intuitive and actionable.
Haicheng Liao, Yongkang Li 0003, Chengyue Wang 0001, Yanchen Guan, Kahou Tam, Chunlin Tian, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001
ACM Multimedia9
2024 CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus Attentions
abstract
Accurately and promptly predicting accidents among surrounding traffic agents from camera footage is crucial for the safety of autonomous vehicles (AVs). This task presents substantial challenges stemming from the unpredictable nature of traffic accidents, their long-tail distribution, the intricacies of traffic scene dynamics, and the inherently constrained field of vision of onboard cameras. To address these challenges, this study introduces a novel accident anticipation framework for AVs, termed CRASH. It seamlessly integrates five components: object detector, feature extractor, object-aware module, context-aware module, and multi-layer fusion. Specifically, we develop the object-aware module to prioritize high-risk objects in complex and ambiguous environments by calculating the spatial-temporal relationships between traffic agents. In parallel, the context-aware is also devised to extend global visual information from the temporal to the frequency domain using the Fast Fourier Transform (FFT) and capture fine-grained visual features of potential objects and broader context cues within traffic scenes. To capture a wider range of visual cues, we further propose a multi-layer fusion that dynamically computes the temporal dependencies between different scenes and iteratively updates the correlations between different visual features for accurate and timely accident prediction. Evaluated on real-world datasets-Dashcam Accident Dataset (DAD), Car Crash Dataset (CCD), and AnAn Accident Detection (A3D) datasets-our model surpasses existing top baselines in critical evaluation metrics like Average Precision (AP) and mean Time-To-Accident (mTTA). Importantly, its robustness and adaptability are particularly evident in challenging driving scenarios with missing or limited training data, demonstrating significant potential for application in real-world autonomous driving systems.
Haicheng Liao, Huanming Shen, Chengyue Wang 0001, Chunlin Tian, Kahou Tam, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001
ACM Multimedia9
2024 L-TLA: A Lightweight Driver Distraction Detection Method Based on Three-Level Attention Mechanisms
abstract
Driver distraction is a significant factor leading to traffic accidents. Detecting driver distraction is crucial for the development of advanced driver assistance systems (ADAS). With the development of deep learning techniques, advanced computer vision technologies have been continuously applied for driver distraction detection. To date, most distraction detection approaches cannot well be adapted to the distraction behaviors that are not included in the training dataset. To address this problem, we propose a lightweight driver distraction detection method using semisupervised contrastive learning. Unlike other studies that rely on large-scale models, a lightweight vision transformer with convolutional neural network (CNN) obtained by knowledge distillation is adapted to extract features, and the design of the dual-stream backbone network increases the generalization ability without increasing the computational burden. Furthermore, the combination of three-level attention mechanisms (i.e., channel-level, spatial-level, and batch-level) enhances the representative power of the model. Both depth and RGB datasets are used to train and test our proposed method. The experimental results show that our method shows superior performance in comparison with other state-of-the-art methods. Its lightweight architecture is suitable for practical applications. This study contributes to the development of ADAS and provides a new perspective on driver distraction detection.
Zizheng Guo 0004, Lin Zhang 0035, Zhenning Li 0001, Guofa Li
IEEE Trans. Reliab.4
2022 An Intelligent Train Operation Method Based on Event-Driven Deep Reinforcement Learning
abstract
Train operation control in urban railways is challenging due to its high dynamics, complex environment, and level of comfort and safety. To address these challenges, in this article, the authors propose a new deep reinforcement-based train operation (DRTO) method which includes: 1) A deterministic deep reinforcement learning algorithm, 2) a dynamic incentive system, which is used to ensure safe operation in a multitrain environment, and 3) an event-driven method, which is used to improve the DRTO performance based on an event-driven strategy. To evaluate the performance, we thoroughly compare the proposed method with other operation control solutions on both synthetic and real datasets. Our results demonstrate that DRTO is effective in: 1) Decreasing the energy consumption of train operation, 2) increasing passenger comfort, and 3) achieving a good tradeoff between efficiency and safety. In addition, the effectiveness of the event-driven strategy and the dynamic incentive system is demonstrated in the experiments.
Leong Hou U, Mingliang Zhou 0001, Zhenning Li 0001
IEEE Trans. Ind. Informatics4
2019 Taxi-Based Mobility Demand Formulation and Prediction Using Conditional Generative Adversarial Network-Driven Learning Approaches
abstract
In this paper, a deep learning (DL) framework was proposed to predict the taxi-passenger demand while the spatial, the temporal, and external dependencies were considered simultaneously. The proposed DL framework combined a modified density-based spatial clustering algorithm with noise (DBSCAN) and a conditional generative adversarial network (CGAN) model. More specifically, the modified DBSCAN model was applied to produce a number of sub-networks considering the spatial correlation of taxi pick-up events in the road network. And the CGAN model, fed with the historical taxi passenger demand and other conditional information, was capable to predict the taxi-passenger demands. The proposed CGAN model was made up with two long short-term memory (LSTM) neural networks, which are termed as the generative network G and the discriminative network D, respectively. Adversarial training process was conducted to the two LSTMs. In the numerical experiment, different model layouts were compared. It was found that different network layouts provided reasonable accuracy. With limited training data, more LSTM layers in the generator network resulted in not only higher accuracy, but also more difficulties in training. Comparisons were also conducted between the proposed prediction model and four typical approaches, including the moving average method, the autoregressive integrated moving method, the neural network model, and the LSTM neural network model. The comparison results showed that the proposed model outperformed all the other methods. And the repeated experiment indicated that the proposed CGAN model provided significant better predictions than the LSTM model did. Future research was recommended to include more datasets for testing the model and more information for improving predictive performance.
Hao Yu 0031, Zhenning Li 0001, Guohui Zhang 0001, Pan Liu 0013, Jin-Fu Yang, Yin Yang 0002
IEEE Trans. Intell. Transp. Syst.3