Bin Rao 0003

dblp:28/9838-3 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0008-4907-914XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Autonomous driving · 52% Transfer learning and domain adaptation · 14% Generative modeling · 7%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
trajectory prediction
3.642026
Differentiable Semantic Meta-Learning Framework for Long-Tail Motion Forecasting in Autonomous Driving · AAAI 2026
Beyond Patterns: Harnessing Causal Logic for Autonomous Driving Trajectory Prediction · IJCAI 2025
AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction · ICCV 2025
Robotics › Autonomous driving › risk assessment
accident anticipation
1.922026
Predict and Resist: Long-Term Accident Anticipation Under Sensor Noise · AAAI 2026
Eyes on the Road, Mind Beyond Vision: Context-Aware Multi-modal Enhanced Risk Anticipation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model › denoising
diffusion-based denoising
1.012026
Predict and Resist: Long-Term Accident Anticipation Under Sensor Noise · AAAI 2026
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation
1.012026
Differentiable Semantic Meta-Learning Framework for Long-Tail Motion Forecasting in Autonomous Driving · AAAI 2026
Machine learning › Transfer learning and domain adaptation
meta-learning
1.012026
Differentiable Semantic Meta-Learning Framework for Long-Tail Motion Forecasting in Autonomous Driving · AAAI 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction · ICCV 2025
Machine learning › Graph learning › hypergraph learning
hypergraph neural network
0.912025
NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous Driving · AAAI 2025
Robotics › Autonomous driving
interaction modeling
0.912025
NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous Driving · AAAI 2025
Robotics › Autonomous driving › trajectory prediction
long-tail trajectory prediction
0.912025
AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction · ICCV 2025
Computer vision › Vision and language
multimodal fusion
0.912025
Eyes on the Road, Mind Beyond Vision: Context-Aware Multi-modal Enhanced Risk Anticipation · ACM Multimedia 2025
Machine learning › Reinforcement learning
actor-critic methods
0.312026
Predict and Resist: Long-Term Accident Anticipation Under Sensor Noise · AAAI 2026
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.312025
Beyond Patterns: Harnessing Causal Logic for Autonomous Driving Trajectory Prediction · IJCAI 2025
Machine learning › Reinforcement learning
neuromodulation
0.312025
NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous Driving · AAAI 2025
Data mining
clustering
0.312025
AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction · ICCV 2025

Methods — techniques the papers use, named apart from their topics

time-weighted reward · 1.0prototype memory · 1.0meta-learning · 1.0diffusion model · 1.0bayesian inference · 1.0actor-critic · 1.0MAML · 1.0trajectory augmentation · 0.9small-world network · 0.9neuromodulation · 0.9momentum contrast · 0.9hypergraph neural network · 0.9contrastive learning · 0.9
YearPublicationVenuePosition
2026 Predict and Resist: Long-Term Accident Anticipation Under Sensor Noise
abstract
Accident anticipation is essential for proactive and safe autonomous driving, where even a brief advance warning can enable critical evasive actions. However, two key challenges hinder real-world deployment: (1) noisy or degraded sensory inputs from weather, motion blur, or hardware limitations, and (2) the need to issue timely yet reliable predictions that balance early alerts with false-alarm suppression. We propose a unified framework that integrates diffusion-based denoising with a time-aware actor-critic model to address these challenges. The diffusion module reconstructs noise-resilient image and object features through iterative refinement, preserving critical motion and interaction cues under sensor degradation. In parallel, the actor-critic architecture leverages long-horizon temporal reasoning and time-weighted rewards to determine the optimal moment to raise an alert, aligning early detection with reliability. Experiments on three benchmark datasets (DAD, CCD, A3D) demonstrate state-of-the-art accuracy and significant gains in mean time-to-accident, while maintaining robust performance under Gaussian and impulse noise. Qualitative analyses further show that our model produces earlier, more stable, and human-aligned predictions in both routine and highly complex traffic scenarios, highlighting its potential for real-world, safety-critical deployment.
Xingcheng Liu, Bin Rao 0003, Yanchen Guan, Chengyue Wang 0001, Haicheng Liao, Jiaxun Zhang, Chengyu Lin 0003, Meixin Zhu, Zhenning Li 0001
AAAI2
2026 Differentiable Semantic Meta-Learning Framework for Long-Tail Motion Forecasting in Autonomous Driving
abstract
Long-tail motion forecasting is a core challenge for autonomous driving, where rare yet safety-critical events-such as abrupt maneuvers and dense multi-agent interactions-dominate real-world risk. Existing approaches struggle in these scenarios because they rely on either non-interpretable clustering or model-dependent error heuristics, providing neither a differentiable notion of “tailness” nor a mechanism for rapid adaptation. We propose SAML, a Semantic-Aware Meta-Learning framework that introduces the first differentiable definition of tailness for motion forecasting. SAML quantifies motion rarity via semantically meaningful intrinsic (kinematic, geometric, temporal) and interactive (local and global risk) properties, which are fused by a Bayesian Tail Perceiver into a continuous, uncertainty-aware Tail Index. This Tail Index drives a meta-memory adaptation module that couples a dynamic prototype memory with an MAML-based cognitive set mechanism, enabling fast adaptation to rare or evolving patterns. Experiments on nuScenes, NGSIM, and HighD show that SAML achieves state-of-the-art overall accuracy and substantial gains on top 1-5% worst-case events, while maintaining high efficiency. Our findings highlight semantic meta-learning as a pathway toward robust and safety-critical motion forecasting.
Bin Rao 0003, Chengyue Wang 0001, Haicheng Liao, Qianfang Wang, Yanchen Guan, Jiaxun Zhang, Xingcheng Liu, Meixin Zhu, Kanye Ye Wang, Zhenning Li 0001
AAAI1
2025 NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous Driving
abstract
Accurate trajectory prediction is essential for the safety and efficiency of autonomous driving. Traditional models often struggle with real-time processing, capturing non-linearity and uncertainty in traffic environments, efficiency in dense traffic, and modeling temporal dynamics of interactions. We introduce NEST (Neuromodulated Small-world Hypergraph Trajectory Prediction), a novel framework that integrates Small-world Networks and hypergraphs for superior interaction modeling and prediction accuracy. This integration enables the capture of both local and extended vehicle interactions, while the Neuromodulator component adapts dynamically to changing traffic conditions. We validate the NEST model on several real-world datasets, including nuScenes, MoCAD, and HighD. The results consistently demonstrate that NEST outperforms existing methods in various traffic scenarios, showcasing its exceptional generalization capability, efficiency, and temporal foresight. Our comprehensive evaluation illustrates that NEST significantly improves the reliability and operational efficiency of autonomous driving systems, making it a robust solution for trajectory prediction in complex traffic environments.
Chengyue Wang 0001, Haicheng Liao, Bonan Wang, Yanchen Guan, Bin Rao 0003, Ziyuan Pu, Zhiyong Cui, Cheng-Zhong Xu 0001, Zhenning Li 0001
AAAI5
2025 AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction
abstract
Accurately predicting the future trajectories of traffic agents is essential in autonomous driving. However, due to the inherent imbalance in trajectory distributions, tail data in natural datasets often represents more complex and hazardous scenarios. Existing studies typically rely solely on a base model's prediction error, without considering the diversity and uncertainty of long-tail trajectory patterns. We propose an adaptive momentum and decoupled contrastive learning framework (AMD), which integrates unsupervised and supervised contrastive learning strategies. By leveraging an improved momentum contrast learning (MoCo-DT) and decoupled contrastive learning (DCL) module, our framework enhances the model's ability to recognize rare and complex trajectories. Additionally, we design four types of trajectory random augmentation methods and introduce an online iterative clustering strategy, allowing the model to dynamically update pseudo-labels and better adapt to the distributional shifts in long-tail data. We propose three different criteria to define long-tail trajectories and conduct extensive comparative experiments on the nuScenes and ETH$/$UCY datasets. The results show that AMD not only achieves optimal performance in long-tail trajectory prediction but also demonstrates outstanding overall prediction accuracy.
Bin Rao 0003, Haicheng Liao, Yanchen Guan, Chengyue Wang 0001, Bonan Wang, Jiaxun Zhang, Zhenning Li 0001
ICCV1
2025 Beyond Patterns: Harnessing Causal Logic for Autonomous Driving Trajectory Prediction
abstract
Accurate trajectory prediction has long been a major challenge for autonomous driving (AD). Traditional data-driven models predominantly rely on statistical correlations, often overlooking the causal relationships that govern traffic behavior. In this paper, we introduce a novel trajectory prediction framework that leverages causal inference to enhance predictive robustness, generalization, and accuracy. By decomposing the environment into spatial and temporal components, our approach identifies and mitigates spurious correlations, uncovering genuine causal relationships. We also employ a progressive fusion strategy to integrate multimodal information, simulating human-like reasoning processes and enabling real-time inference. Evaluations on five real-world datasets—ApolloScape, nuScenes, NGSIM, HighD, and MoCAD—demonstrate our model's superiority over existing state-of-the-art (SOTA) methods, with improvements in key metrics such as RMSE and FDE. Our findings highlight the potential of causal reasoning to transform trajectory prediction, paving the way for robust AD systems.
Bonan Wang, Haicheng Liao, Chengyue Wang 0001, Bin Rao 0003, Yanchen Guan, Guyang Yu, Jiaxun Zhang, Songning Lai, Cheng-Zhong Xu 0001, Zhenning Li 0001
IJCAI4
2025 Eyes on the Road, Mind Beyond Vision: Context-Aware Multi-modal Enhanced Risk Anticipation
abstract
Accurate accident anticipation remains challenging when driver cognition and dynamic road conditions are underrepresented in predictive models. In this paper, we propose CAMERA (Context-Aware Multi-modal Enhanced Risk Anticipation), a multi-modal framework integrating dashcam video, textual annotations, and driver attention maps for robust accident anticipation. Unlike existing methods that rely on static or environment-centric thresholds, CAMERA employs an adaptive mechanism guided by scene complexity and gaze entropy, reducing false alarms while maintaining high recall in dynamic, multi-agent traffic scenarios. A hierarchical fusion pipeline with Bi-GRU (Bidirectional GRU) captures spatio-temporal dependencies, while a Geo-Context Vision-Language module translates 3D spatial relationships into interpretable, human-centric alerts. Evaluations on the DADA-2000 and benchmarks show that CAMERA achieves state-of-the-art performance, improving accuracy and lead time. These results demonstrate the effectiveness of modeling driver attention, contextual description, and adaptive risk thresholds to enable more reliable accident anticipation.
Jiaxun Zhang, Haicheng Liao, Yumu Xie, Chengyue Wang 0001, Yanchen Guan, Bin Rao 0003, Zhenning Li 0001
ACM Multimedia6
2025 Chain-of-Thought Guided Multimodal Large Language Models for Scene-Aware Accident Anticipation in Autonomous Driving
abstract
Accurately anticipating traffic accidents is a fundamental task for the safe and effective deployment of autonomous vehicles (AVs). However, existing models primarily rely on dashcam footage and often fail to generalize across varied driving scenarios due to their dependence on visual data and the rarity of high-risk events in datasets. These limitations undermine their robustness and reduce practical applicability in dynamic, unpredictable environments. To address these challenges, this study proposes a novel approach, termed MLTA, which integrates multimodal learning with the hypergraph attention network to hierarchically extract and capture cross-modal interaction. It leverages LLava-next, a multimodal large language model (MLLM) guided by the Chain-of-Thought (CoT) prompting paradigm, to produce context-aware interpretations of traffic scenes. This is further enhanced by a human-inspired attention mechanism that mimics the decision-making priorities of experienced human drivers. This combination enables more accurate identification of critical elements in a scene, improving both prediction precision and timeliness. Extensive experiments on four real-world datasets—DAD, A3D, CCD, and DADA-2000—show that our approach consistently outperforms state-of-the-art (SOTA) methods, demonstrating strong adaptability and robustness in complex driving environments.
Haicheng Liao, Bin Rao 0003, Chengyue Wang 0001, Shengbo Eben Li, Cheng-Zhong Xu 0001, Zhenning Li 0001
IEEE Trans. Intell. Transp. Syst.2