EDBT 2026 Demo / reviewers in the wild / expert
Seth Z. Zhao
dblp:313/9781
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0004-4727-492XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Autonomous driving · 39% Question answering and dialogue systems · 17% Robot navigation and mapping · 16% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Autonomous driving
collaborative perception |
1.7 | 2 | 2025 | TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction · ICCV 2025 V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction · ICCV 2025 |
Natural language and speech › Question answering and dialogue systems
retrieval-augmented reasoning |
1.0 | 1 | 2026 | Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning · AAAI 2026 |
Computer vision › Vision and language › visual question answering
driving question answering |
0.9 | 1 | 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving · ICML 2025 |
Robotics › Autonomous driving
driving scene understanding |
0.9 | 1 | 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving · ICML 2025 |
Robotics › Robot navigation and mapping
embodied AI simulation |
0.9 | 1 | 2025 | Towards Autonomous Micromobility through Scalable Urban Simulation · CVPR 2025 |
Robotics › Autonomous driving › interaction modeling
interaction reasoning |
0.9 | 1 | 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving · ICML 2025 |
Natural language and speech › Question answering and dialogue systems
multimodal question answering |
0.9 | 1 | 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving · ICML 2025 |
Machine learning › Learning paradigms
multi-task learning |
0.9 | 1 | 2025 | TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction · ICCV 2025 |
Robotics › Autonomous driving
trajectory prediction |
0.9 | 1 | 2025 | V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction · ICCV 2025 |
Robotics › Robot navigation and mapping › mobile robot navigation › outdoor navigation
urban navigation |
0.9 | 1 | 2025 | Towards Autonomous Micromobility through Scalable Urban Simulation · CVPR 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2026 | Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked reconstruction |
0.3 | 1 | 2025 | TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 1.0large language model reasoning · 1.0transformer · 0.9masked reconstruction learning · 0.9masked reconstruction · 0.9hierarchical urban generation · 0.9gradient conflict suppression · 0.9fine-tuning · 0.9early/late/intermediate fusion · 0.9asynchronous scene sampling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented ReasoningabstractUnderstanding and adhering to traffic regulations is essential for autonomous vehicles to ensure safety and trustworthiness. However, traffic regulations are complex, context-dependent, and differ between regions, posing a major challenge to conventional rule-based decision-making approaches. We present an interpretable, regulation-aware decision-making framework, DriveReg, which enables autonomous vehicles to understand and adhere to region-specific traffic laws and safety guidelines. The framework integrates a Retrieval Augmented Generation (RAG)-based Traffic Regulation Retrieval Agent, which retrieves relevant rules from regulatory documents based on the current situation, and a Large Language Model (LLM)-powered Reasoning Agent that evaluates actions for legal compliance and safety. Our design emphasizes interpretability to enhance transparency and trustworthiness. To support systematic evaluation, we introduce DriveReg Scenarios Dataset, a comprehensive dataset of driving scenarios across Boston, Singapore, and Los Angeles, with both hypothesized text-based cases and real-world driving data, specifically constructed and annotated to evaluate models’ capacity for regulation understanding and reasoning. We validate our framework on the DriveReg Scenarios Dataset and real-world deployment, demonstrating strong performance and robustness across diverse environments. Tianhui Cai, Zewei Zhou, Haoxuan Ma, Seth Z. Zhao, Zhiwen Wu, Xu Han 0014, Zhiyu Huang, Jiaqi Ma 0003 |
AAAI | 5 |
| 2025 | Towards Autonomous Micromobility through Scalable Urban SimulationabstractMicromobility, which utilizes lightweight mobile machines moving in urban public spaces - such as delivery robots and electric wheelchairs - emerges as a promising alternative to vehicular mobility. Current micromobility depends mostly on human manual operation (in-person or remote control), which raises safety and efficiency concerns when navigating busy urban environments full of unpredictable obstacles and pedestrians. Assisting humans with AI agents in maneuvering micromobility devices presents a viable solution for enhancing safety and efficiency. In this work, we present a scalable urban simulation solution to advance autonomous micromobility. First, we build URBAN-SIM – a high-performance robot learning platform for large-scale training of embodied agents in interactive urban scenes. URBAN-SIM contains three critical modules: Hierarchical Urban Generation pipeline, Interactive Dynamics Generation strategy, and Asynchronous Scene Sampling scheme, to improve the diversity, realism, and efficiency of robot learning in simulation. Then, we propose URBAN-BENCH – a suite of essential tasks and benchmarks to gauge various capabilities of the AI agents in achieving autonomous micromobility. URBAN-BENCH includes eight tasks based on three core skills of the agents: Urban Locomotion, Urban Navigation, and Urban Traverse. We evaluate four robots with heterogeneous embodiments, such as the wheeled and legged robots, across these tasks. Experiments on diverse terrains and urban structures reveal each robot’s strengths and limitations. Project page: https://metadriverse.github.io/urban-sim/. Wayne Wu, Honglin He, Chaoyuan Zhang, Jack He, Seth Z. Zhao, Quanyi Li, Bolei Zhou |
CVPR | 5 |
| 2025 | V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionabstractVehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores temporal cues and temporal tasks (e.g., temporal perception and prediction). In this paper, we focus on the spatio-temporal fusion in V2X scenarios and design one-step and multi-step communication strategies (when to transmit) as well as examine their integration with three fusion strategies - early, late, and intermediate (what to transmit), providing comprehensive benchmarks with 11 fusion models (how to fuse). Furthermore, we propose V2XPnP, a novel intermediate fusion framework within one-step communication for end-to-end perception and prediction. Our framework employs a unified Transformer-based architecture to effectively model complex spatio-temporal relationships across multiple agents, frames, and high-definition maps. Moreover, we introduce the V2XPnP Sequential Dataset that supports all V2X collaboration modes and addresses the limitations of existing real-world datasets, which are restricted to single-frame or single-mode cooperation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods in both perception and prediction tasks. Zewei Zhou, Hao Xiang 0001, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Tianhui Cai, Johnson Liu, Maheswari Bajji, Xin Xia 0007, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
ICCV | 4 |
| 2025 | TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionabstractEnd-to-end training of multi-agent systems offers significant advantages in improving multi-task performance. However, training such models remains challenging and requires extensive manual design and monitoring. In this work, we introduce TurboTrain, a novel and efficient training framework for multi-agent perception and prediction. TurboTrain comprises two key components: a multi-agent spatiotemporal pretraining scheme based on masked reconstruction learning and a balanced multi-task learning strategy based on gradient conflict suppression. By streamlining the training process, our framework eliminates the need for manually designing and tuning complex multi-stage training pipelines, substantially reducing training time and improving performance. We evaluate TurboTrain on a real-world cooperative driving dataset, V2XPnP-Seq, and demonstrate that it further improves the performance of state-of-the-art multi-agent perception and prediction models. Our results highlight that pretraining effectively captures spatiotemporal multi-agent features and significantly benefits downstream tasks. Moreover, the proposed balanced multi-task learning strategy enhances detection and prediction. Zewei Zhou, Seth Z. Zhao, Tianhui Cai, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
ICCV | 2 |
| 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in DrivingabstractLanguage models uncover unprecedented abilities in analyzing driving scenarios, owing to their limitless knowledge accumulated from text-based pre-training. Naturally, they should particularly excel in analyzing rule-based interactions, such as those triggered by traffic laws, which are well documented in texts. However, such interaction analysis remains underexplored due to the lack of dedicated language datasets that address it. Therefore, we propose Waymo Open Motion Dataset-Reasoning (WOMD-Reasoning), a comprehensive large-scale Q&As dataset built on WOMD focusing on describing and reasoning traffic rule-induced interactions in driving scenarios. WOMD-Reasoning also presents by far the largest multi-modal Q&A dataset, with 3 million Q&As on real-world driving scenarios, covering a wide range of driving topics from map descriptions and motion status descriptions to narratives and analyses of agents' interactions, behaviors, and intentions. To showcase the applications of WOMD-Reasoning, we design Motion-LLaVA, a motion-language model fine-tuned on WOMD-Reasoning. Quantitative and qualitative evaluations are performed on WOMD-Reasoning dataset as well as the outputs of Motion-LLaVA, supporting the data quality and wide applications of WOMD-Reasoning, in interaction predictions, traffic rule compliance plannings, etc. The dataset and its vision modal extension are available on https://waymo.com/open/download/. The codes & prompts to build it are available on https://github.com/yhli123/WOMD-Reasoning. Cunxin Fan, Chongjian Ge, Seth Z. Zhao, Chenran Li, Chenfeng Xu, Huaxiu Yao, Masayoshi Tomizuka, Bolei Zhou, Chen Tang 0001, Mingyu Ding |
ICML | 4 |
| 2025 | CooPre: Cooperative Pretraining for V2X Cooperative PerceptionabstractExisting Vehicle-to-Everything (V2X) cooperative perception methods rely on accurate multi-agent 3D annotations. Nevertheless, it is time-consuming and expensive to collect and annotate real-world data, especially for V2X systems. In this paper, we present a self-supervised learning framwork for V2X cooperative perception, which utilizes the vast amount of unlabeled 3D V2X data to enhance the perception performance. Specifically, multi-agent sensing information is aggregated to form a holistic view and a novel proxy task is formulated to reconstruct the LiDAR point clouds across multiple connected agents to better reason multi-agent spatial correlations. Besides, we develop a V2X bird-eye-view (BEV) guided masking strategy which effectively allows the model to pay attention to 3D features across heterogeneous V2X agents (i.e., vehicles and infrastructure) in the BEV space. Noticeably, such a masking strategy effectively pretrains the 3D encoder with a multi-agent LiDAR point cloud reconstruction objective and is compatible with mainstream cooperative perception backbones. Our approach, validated through extensive experiments on representative datasets (i.e., V2X-Real, V2V4Real, and OPV2V) and multiple state-of-the-art cooperative perception methods (i.e., AttFuse, F-Cooper, and V2X-ViT), leads to a performance boost across all V2X settings. Notably, CooPre achieves a 4% mAP improvement on V2X-Real dataset and surpasses baseline performance using only 50% of the training data, highlighting its data efficiency. Additionally, we demonstrate the framework’s powerful performance in cross-domain transferability and robustness under challenging scenarios. The code will be made publicly available at https://github.com/ucla-mobility/CooPre. Seth Z. Zhao, Hao Xiang 0001, Chenfeng Xu, Xin Xia 0007, Bolei Zhou, Jiaqi Ma 0003 |
IROS | 1 |
| 2025 | AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningabstractRecent advancements in Vision-Language-Action (VLA) models have shown promise for end-to-end autonomous driving by leveraging world knowledge and reasoning capabilities. However, current VLA models often struggle with physically infeasible action outputs, complex model structures, or unnecessarily long reasoning. In this paper, we propose AutoVLA, a novel VLA model that unifies reasoning and action generation within a single autoregressive generation model for end-to-end autonomous driving. AutoVLA performs semantic reasoning and trajectory planning directly from raw visual inputs and language instructions. We tokenize continuous trajectories into discrete, feasible actions, enabling direct integration into the language model. For training, we employ supervised fine-tuning to equip the model with dual thinking modes: fast thinking (trajectory-only) and slow thinking (enhanced with chain-of-thought reasoning). To further enhance planning performance and efficiency, we introduce a reinforcement fine-tuning method based on Group Relative Policy Optimization (GRPO), reducing unnecessary reasoning in straightforward scenarios. Extensive experiments across real-world and simulated datasets and benchmarks, including nuPlan, nuScenes, Waymo, and CARLA, demonstrate the competitive performance of AutoVLA in both open-loop and closed-loop settings. Qualitative results showcase the adaptive reasoning and accurate planning capabilities of AutoVLA in diverse scenarios. Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
NeurIPS | 3 |
| 2024 | Pre-training on Synthetic Driving Data for Trajectory PredictionabstractAccumulating substantial volumes of real-world driving data proves pivotal in the realm of trajectory forecasting for autonomous driving. Given the heavy reliance of current trajectory forecasting models on data-driven methodologies, we aim to tackle the challenge of learning general trajectory forecasting representations under limited data availability. We propose a pipeline-level solution to mitigate the issue of data scarcity in trajectory forecasting. The solution is composed of two parts: firstly, we adopt HD map augmentation and trajectory synthesis for generating driving data, and then we learn representations by pre-training on them. Specifically, we apply vector transformations to reshape the maps, and then employ a rule-based model to generate trajectories on both original and augmented scenes; thus enlarging the driving data without collecting additional real ones. To foster the learning of general representations within this augmented dataset, we comprehensively explore the different pre-training strategies, including extending the concept of a Masked AutoEncoder (MAE) for trajectory forecasting. Without bells and whistles, our proposed pipeline-level solution is general, simple, yet effective: we conduct extensive experiments to demonstrate the effectiveness of our data expansion and pre-training strategies, which outperform the baseline prediction model by large margins, e.g. 5.04%, 3.84% and 8.30% in terms of MR6, minADE6and minFDE6. The pre-training dataset and the codes for pre-training and fine-tuning are released at https://github.com/yhli123/Pretraining_on_Synthetic_Driving_Data_for_Trajectory_Prediction. Seth Z. Zhao, Chenfeng Xu, Chen Tang 0001, Chenran Li, Mingyu Ding, Masayoshi Tomizuka |
IROS | 2 |
| 2022 | Multimodal Semantic Mismatch Detection in Social Media PostsabstractShort videos have become the most popular form of social media in recent years. In this work, we focus on the threat scenario where video, audio, and their text description are semantically mismatched to mislead the audience. We develop self-supervised methods to detect semantic mismatch across multiple modalities, namely video, audio and text. We use state-of-the-art language, video and audio models to extract dense features from each modality, and explore transformer architecture together with contrastive learning methods on a dataset of one million Twitter posts from 2021 to 2022. Our best-performing method benefits from the robustness of Noise-Contrastive loss and the context provided by fusing modalities together using a cross-transformer. It outperforms state-of-the-art by over 9% in accuracy. We further characterize the performance of our system on topic-specific datasets containing COVID-19 and Russia-Ukraine related tweets, and shows that it outperforms state-of-the-art by over 17% in accuracy. Seth Z. Zhao, Avideh Zakhor, John F. Canny |
MMSP | 2 |