Jiazhi Yang

dblp:305/2099 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Autonomous driving · 41% Generative modeling · 33% Video understanding and tracking · 7%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.932025
ReSim: Reliable World Simulation for Autonomous Driving · NeurIPS 2025
Generalized Predictive Model for Autonomous Driving · CVPR 2024
Decoupled Diffusion Sparks Adaptive Scene Generation · ICCV 2025
Robotics › Autonomous driving › driving model learning
driving world model
1.622025
ReSim: Reliable World Simulation for Autonomous Driving · NeurIPS 2025
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability · NeurIPS 2024
Robotics › Autonomous driving › scenario generation
controllable traffic scene generation
0.912025
Decoupled Diffusion Sparks Adaptive Scene Generation · ICCV 2025
Machine learning › Generative modeling
scene generation
0.912025
Decoupled Diffusion Sparks Adaptive Scene Generation · ICCV 2025
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception
0.812024
Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.812024
Generalized Predictive Model for Autonomous Driving · CVPR 2024
Robotics › Robot navigation and mapping
sensor fusion
0.812024
Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Video understanding and tracking
video prediction
0.812024
Generalized Predictive Model for Autonomous Driving · CVPR 2024
Computer vision › 3D vision
view transformation
0.812024
Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Robotics › Autonomous driving
planning for self-driving vehicles
0.712023
Planning-oriented Autonomous Driving · CVPR 2023
Machine learning › Reinforcement learning
policy evaluation
0.312025
ReSim: Reliable World Simulation for Autonomous Driving · NeurIPS 2025
Robotics › Motion planning and robot control
motion planning
0.212024
Generalized Predictive Model for Autonomous Driving · CVPR 2024
Machine learning › Generative modeling
video generation
0.212024
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

video generation · 0.9reward estimation · 0.9partial noise-masking · 0.9noise-aware schedule · 0.9diffusion transformer · 0.9decoupled diffusion · 0.9temporal reasoning blocks · 0.8reward modeling · 0.8latent replacement · 0.8action-conditioned prediction · 0.8
YearPublicationVenuePosition
2026 Integrated component and damage recognition for post-earthquake reinforced concrete structures via image segmentation
abstract
Accurate and efficient damage recognition of reinforced concrete structures is essential for structural health monitoring and recovery planning. However, most existing deep learning-based approaches treat damage recognition as an isolated visual task, without explicitly considering component-level context or the engineering relevance of different failure modes. This study proposes a multi-task framework, guided by the geometric and engineering characteristics of different damage types, that integrates component recognition with damage recognition to achieve fine-grained post-earthquake component assessment. Component localization is achieved through instance segmentation, region-level damage is obtained via semantic segmentation, and cracks are identified via orientation-aware bounding box detection. To address data scarcity, a data augmentation strategy termed foreground-background regrouping is introduced, which leverages structural domain knowledge to generate physically realistic damage samples. Experimental results demonstrate the effectiveness of the framework, yielding accurate component localization and damage segmentation. The framework shows strong potential for automated, rapid post-earthquake component assessment and decision-making support.
Zhilin Bai, Dujian Zou, Tiejun Liu, Zichao Que, Jiazhi Yang, Ao Zhou 0002
Eng. Appl. Artif. Intell.6
2026 A Multiagent Reinforcement Learning-Based Offloading and Resource Allocation for Vehicle Edge Computing
abstract
In the Internet of Vehicles (IoV), vehicles have the capability to offload their computational tasks to the Mobile Edge Computing (MEC) servers in order to reduce service delay. However, the majority of existent task offloading and computational resource allocation schemes are static and lack consideration of the heterogeneous nature of tasks. Furthermore, in scenarios involving both collaborative and competitive resource utilization, there remains considerable room for performance enhancement for delay-sensitive tasks. To address these challenges, this paper proposes a novel Global Heterogeneous Multi-Agent Reinforcement Learning (GHMARL) that is an enhancement to the general MARL. In GHMARL, each vehicle and MEC server is represented by an agent, and intelligent collaboration and dynamic resource allocation are employed to balance resource usage and delay performance. In particular, GHMARL introduces a global Critic network and a local Critic network, working in a collaborative manner. The former is responsible for guiding the overall system performance to ensure the service delay performance, and the latter is responsible for MEC servers’ performance to ensure the resource usage efficiency. Simulation results demonstrate that, in comparison with alternative schemes, GHMARL significantly enhances the overall system performance, particularly with regard to resource usage efficiency. Furthermore, GHMARL offers distinct advantages in balancing task delay and resource consumption under various system dynamics, making it a robust solution to address issues of resource wastage and delay violations for IoV systems.
Jiazhi Yang, Nan Ma 0014, Pei Xiao 0001
IEEE Internet Things J.3
2025 Decoupled Diffusion Sparks Adaptive Scene Generation
abstract
Controllable scene generation could reduce the cost of diverse data collection substantially for autonomous driving. Prior works formulate the traffic layout generation as predictive progress, either by denoising entire sequences at once or by iteratively predicting the next frame. However, full sequence denoising hinders online reaction, while the latter's short-sighted next-frame prediction lacks precise goal-state guidance. Further, the learned model struggles to generate complex or challenging scenarios due to a large number of safe and ordinal driving behaviors from open datasets. To overcome these, we introduce Nexus, a decoupled scene generation framework that improves reactivity and goal conditioning by simulating both ordinal and challenging scenarios from fine-grained tokens with independent noise states. At the core of the decoupled pipeline is the integration of a partial noise-masking training strategy and a noise-aware schedule that ensures timely environmental updates throughout the denoising process. To complement challenging scenario generation, we collect a dataset consisting of complex corner cases. It covers 540 hours of simulated data, including high-risk interactions such as cut-in, sudden braking, and collision. Nexus achieves superior generation realism while preserving reactivity and goal orientation, with a 40% reduction in displacement error. We further demonstrate that Nexus improves closed-loop planning by 20% through data augmentation and showcase its capability in safety-critical data generation.
Yunsong Zhou, Naisheng Ye, William Ljungbergh, Tianyu Li 0004, Jiazhi Yang, Zetong Yang, Hongzi Zhu, Christoffer Petersson, Hongyang Li 0001
ICCV5
2025 ReSim: Reliable World Simulation for Autonomous Driving
abstract
How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusively on real-world driving data composed mainly of safe expert trajectories, struggle to follow hazardous or non-expert behaviors, which are rare in such data. This limitation restricts their applicability to tasks such as policy evaluation. In this work, we address this challenge by enriching real-world human demonstrations with diverse non-expert data collected from a driving simulator (e.g., CARLA), and building a controllable world model trained on this heterogeneous corpus. Starting with a video generator featuring diffusion transformer architecture, we devise several strategies to effectively integrate conditioning signals and improve prediction controllability and fidelity. The resulting model, ReSim, enables Reliable Simulation of diverse open-world driving scenarios under various actions, including hazardous non-expert ones. To close the gap between high-fidelity simulation and applications that require reward signals to judge different actions, we introduce a Video2Reward module that estimates reward from ReSim’s simulated future. Our ReSim paradigm achieves up to 44% higher visual fidelity, improves controllability for both expert and non-expert actions by over 50%, and boosts planning and policy selection performance on NAVSIM by 2% and 25%, respectively.
Jiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 0005, Yuqian Shao, Xiaosong Jia, Hongyang Li 0001, Andreas Geiger 0001, Xiangyu Yue 0001, Li Chen 0008
NeurIPS1
2024 Generalized Predictive Model for Autonomous Driving
abstract
In this paper, we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model, we ac-quire massive data from the web and pair it with diverse and high-quality text descriptions. The resultant dataset accumulates over 2000 hours of driving videos, spanning areas all over the world with diverse weather conditions and traffic scenarios. Inheriting the merits from recent latent diffusion models, our model, dubbed GenAD, handles the challenging dynamics in driving scenes with novel tem-poral reasoning blocks. We showcase that it can general-ize to various unseen driving datasets in a zero-shot man-ner, surpassing general or driving-specific video prediction counterparts. Furthermore, GenAD can be adapted into an action-conditioned prediction model or a motion planner, holding great potential for real-world driving applications.
Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen 0008, Tianyu Li 0004, Bo Dai 0002, Kashyap Chitta, Penghao Wu, Ping Luo 0002, Jun Zhang 0106, Andreas Geiger 0001, Yu Qiao 0001, Hongyang Li 0001
CVPR1
2024 Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
abstract
World models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction fidelity of critical details, and action controllability for flexible application. In this paper, we present Vista, a generalizable driving world model with high fidelity and versatile controllability. Based on a systematic diagnosis of existing methods, we introduce several key ingredients to address these limitations. To accurately predict real-world dynamics at high resolution, we propose two novel losses to promote the learning of moving instances and structural information. We also devise an effective latent replacement approach to inject historical frames as priors for coherent long-horizon rollouts. For action controllability, we incorporate a versatile set of controls from high-level intentions (command, goal point) to low-level maneuvers (trajectory, angle, and speed) through an efficient learning strategy. After large-scale training, the capabilities of Vista can seamlessly generalize to different scenarios. Extensive experiments on multiple datasets show that Vista outperforms the most advanced general-purpose video generator in over 70% of comparisons and surpasses the best-performing driving world model by 55% in FID and 27% in FVD. Moreover, for the first time, we utilize the capacity of Vista itself to establish a generalizable reward for real-world action evaluation without accessing the ground truth actions.
Shenyuan Gao, Jiazhi Yang, Li Chen 0008, Kashyap Chitta, Yihang Qiu, Andreas Geiger 0001, Jun Zhang 0106, Hongyang Li 0001
NeurIPS2
2024 Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe
abstract
Learning powerful representations in bird's-eye-view (BEV) for perception tasks is trending and drawing extensive attention both from industry and academia. Conventional approaches for most autonomous driving algorithms perform detection, segmentation, tracking, etc., in a front or perspective view. As sensor configurations get more complex, integrating multi-source information from different sensors and representing features in a unified view come of vital importance. BEV perception inherits several advantages, as representing surrounding scenes in BEV is intuitive and fusion-friendly; and representing objects in BEV is most desirable for subsequent modules as in planning and/or control. The core problems for BEV perception lie in (a) how to reconstruct the lost 3D information via view transformation from perspective view to BEV; (b) how to acquire ground truth annotations in BEV grid; (c) how to formulate the pipeline to incorporate features from different sources and views; and (d) how to adapt and generalize algorithms as sensor configurations vary across different scenarios. In this survey, we review the most recent works on BEV perception and provide an in-depth analysis of different solutions. Moreover, several systematic designs of BEV approach from the industry are depicted as well. Furthermore, we introduce a full suite of practical guidebook to improve the performance of BEV perception tasks, including camera, LiDAR and fusion inputs. At last, we point out the future research directions in this area. We hope this report will shed some light on the community and encourage more research effort on BEV perception.
Hongyang Li 0001, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Jiazhi Yang, Hanming Deng, Hao Tian 0006, Enze Xie, Jiangwei Xie, Li Chen 0008, Tianyu Li 0004, Yang Li 0189, Yulu Gao, Xiaosong Jia, Si Liu 0001, Jianping Shi, Dahua Lin, Yu Qiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.9
2023 Planning-oriented Autonomous Driving
abstract
Modern autonomous driving system is characterized as modular tasks in sequential order, i.e., perception, prediction, and planning. In order to perform a wide diversity of tasks and achieve advanced-level intelligence, contemporary approaches either deploy standalone models for individual tasks, or design a multi-task paradigm with separate heads. However, they might suffer from accumulative errors or deficient task coordination. Instead, we argue that a favorable framework should be devised and optimized in pursuit of the ultimate goal, i.e., planning of the self-driving car. Oriented at this, we revisit the key components within perception and prediction, and prioritize the tasks such that all these tasks contribute to planning. We introduce Unified Autonomous Driving (UniAD), a comprehensive framework up-to-date that incorporates full-stack driving tasks in one network. It is exquisitely devised to leverage advantages of each module, and provide complementary feature abstractions for agent interaction from a global perspective. Tasks are communicated with unified query interfaces to facilitate each other toward planning. We instantiate UniAD on the challenging nuScenes benchmark. With extensive ablations, the effectiveness of using such a philosophy is proven by substantially outperforming previous state-of-the-arts in all aspects. Code and models are public.
Yihan Hu 0001, Jiazhi Yang, Li Chen 0008, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Wenhai Wang, Lewei Lu, Xiaosong Jia, Jifeng Dai, Yu Qiao 0001, Hongyang Li 0001
CVPR2
2023 Off-Axis Four-Reflection Optical Structure for Lightweight Single-Band Bathymetric LiDAR
abstract
A traditional bathymetric LiDAR (light detection and ranging) has disadvantages such as large volume, heavy weight, necessity for airport and runway, and high cost for operation. For these reasons, this paper presents an off-axis four-reflection optical structure for single-band (532 nm) bathymetric LiDAR carried on UAV (Unmanned Aerial Vehicle). This optical system fully considers characteristics of the laser echo energy under different water conditions, which relate with the optical system parameters, such as peak power of laser emission, field of view (FOV), receiver aperture area, etc. The proposed optical system designs the objective lens, which are composed of one APD detector and two PMT detectors, the primary mirror, the second mirror, the plane mirror and the third mirror, two split field mirrors that separate the echo signals from shallow water, medium water and deep water, respectively. This proposed optical system was verified in laboratory tank, swimming pool, Lijiang River, lake, and the Qiaogang Sea Bay. It is found that the maximum water depth measured can reach 25.0 m with an error less than 0.1 m averagely. The dimension and weight of this LiDAR reach 90mm×160mm×90mm, and 10.25 kg, respectively, which is lightest and smallest bathymetric LiDAR worldwide.
Guoqing Zhou 0001, Jiasheng Xu, Haocheng Hu, Zhexian Liu, Haotian Zhang 0014, Xiang Zhou 0002, Jiazhi Yang, Xueqin Nong, Naihui Song, Guoshuai Jia, Hanjiang Xiong, Yiqiang Zhao
IEEE Trans. Geosci. Remote. Sens.8