Chaojun Ni

dblp:368/3306 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0008-7092-8310ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
abstract
Closed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these methods struggle with rendering novel trajectory, such as lane changes. Recent works have demonstrated that integrating world model knowledge alleviates these issues. Despite their efficiency, these approaches still encounter difficulties in the accurate representation of more complex maneuvers, with multi-lane shifts being a notable example. Therefore, we introduce ReconDreamer, which enhances driving scene reconstruction through incremental integration of world model knowledge. Specifically, DriveRestorer is proposed to mitigate artifacts via online restoration. This is complemented by a progressive data update strategy designed to ensure high-quality rendering for more complex maneuvers. To the best of our knowledge, ReconDreamer is the first method to effectively render in large maneuvers. Experimental results demonstrate that ReconDreamer outperforms Street Gaussians in the NTA-IoU, NTL-IoU, and FID, with relative improvements by 24.87%, 6.72%, and 29.97%. Furthermore, ReconDreamer surpasses DriveDreamer4D with PVG during large maneuver rendering, as verified by a relative improvement of 195.87% in the NTA-IoU metric and a user study.
Chaojun Ni, Guosheng Zhao, Wenkang Qin, Guan Huang 0003, Yuyin Chen, Xueyang Zhang, Yifei Zhan, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Xingang Wang 0003, Wenjun Mei
CVPR1
2025 HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
abstract
Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from existing videos, thereby lacking flexibility. To address this, we propose HumanDreamer, a decoupled human video generation framework that first generates diverse poses from text prompts and then leverages these poses to generate human-motion videos. Specifically, we propose MotionVid, the largest dataset for human-motion pose generation. Based on the dataset, we present MotionDiT, which is trained to generate structured human-motion poses from text prompts. Besides, a novel LAMA loss is introduced, which together contribute to a significant improvement in FID by 62.4%, along with respective enhancements in R-precision for top1, top2, and top3 by 41.8%, 26.3%, and 18.3%, thereby advancing both the Text-to-Pose control accuracy and FID metrics. Our experiments across various Pose-to-Video baselines demonstrate that the poses generated by our method can produce diverse and high-quality human-motion videos. Furthermore, our model can facilitate other downstream tasks, such as pose sequence prediction and 2D-3D motion lifting.
Chaojun Ni, Guosheng Zhao, Zhiqin Yang, Muyang Zhang, Xinze Chen, Guan Huang 0003, Lihong Liu, Xingang Wang 0003
CVPR3
2025 DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation
abstract
Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which are largely confined to forward-driving scenarios. Consequently, these methods face limitations when rendering complex maneuvers (e.g., lane change, acceleration, deceleration). Recent advancements in autonomous-driving world models have demonstrated the potential to generate diverse driving videos. However, these approaches remain constrained to 2D video generation, inherently lacking the spatiotemporal coherence required to capture intricacies of dynamic driving environments. In this paper, we introduce DriveDreamer4D, which enhances 4D driving scene representation leveraging world model priors. Specifically, we utilize the world model as a data machine to synthesize novel trajectory videos, where structured conditions are explicitly leveraged to control the spatial-temporal consistency of traffic elements. Besides, the cousin data training strategy is proposed to facilitate merging real and synthetic data for optimizing 4DGS. To our knowledge, DriveDreamer4D is the first to utilize video generation models for improving 4D reconstruction in driving scenarios. Experimental results reveal that DriveDreamer4D significantly enhances generation quality under novel trajectory views, achieving a relative improvement in FID by 32.1%, 46.4%, and 16.3% compared to PVG, S3Gaussian, and Deformable-GS. Moreover, DriveDreamer4D markedly enhances the spatiotemporal coherence of driving agents, which is verified by a comprehensive user study and the relative increases of 22.6%, 43.5%, and 15.6% in the NTA-IoU metric.
Guosheng Zhao, Chaojun Ni, Xueyang Zhang, Guan Huang 0003, Xinze Chen, Youyi Zhang, Wenjun Mei, Xingang Wang 0003
CVPR2
2025 WonderTurbo: Generating Interactive 3D World in 0.72 Seconds
abstract
Interactive 3D generation is gaining momentum and capturing extensive attention for its potential to create immersive virtual experiences. However, a critical challenge in current 3D generation technologies lies in achieving real-time interactivity. To address this issue, we introduce WonderTurbo, the first real-time interactive 3D scene generation framework capable of generating novel perspectives of 3D scenes within 0.72 seconds. Specifically, WonderTurbo accelerates both geometric and appearance modeling in 3D scene generation. In terms of geometry, we propose StepSplat, an innovative method that constructs efficient 3D geometric representations through dynamic updates, each taking only 0.26 seconds. Additionally, we design QuickDepth, a lightweight depth completion module that provides consistent depth input for StepSplat, further enhancing geometric accuracy. For appearance modeling, we develop FastPaint, a 2-steps diffusion model tailored for instant inpainting, which focuses on maintaining spatial appearance consistency. Experimental results demonstrate that WonderTurbo achieves a remarkable 15X speedup compared to baseline methods, while preserving excellent spatial consistency and delivering high-quality output.
Chaojun Ni, Haoyun Li, Guosheng Zhao, Wenkang Qin, Guan Huang 0003, Wenjun Mei
ICCV1
2025 ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
Guosheng Zhao, Chaojun Ni, Wenkang Qin, Guan Huang 0003, Xingang Wang 0003
ICCV3
2023 Feature Adaptive YOLO for Remote Sensing Detection in Adverse Weather Conditions
abstract
Target detection in remote sensing has been one of the most challenging tasks in the past few decades. However, the detection performance in adverse weather conditions still needs to be satisfactory, mainly caused by the low-quality image features and the fuzzy boundary information. This work proposes a novel framework called Feature Adaptive YOLO (FA-YOLO). Specifically, we present a Hierarchical Feature Enhancement Module (HFEM), which adaptively performs feature-level enhancement to tackle the adverse impacts of different weather conditions. Then, we propose an Adaptive receptive Field enhancement Module (AFM) that dynamically adjusts the receptive field of the features and thus can enrich the context information for feature augmentation. In addition, we introduce Deformable Gated Head (DG-Head) which reduces the clutter caused by adverse weather. Experimental results on RTTS and two synthetic datasets demonstrate that our proposed FA-YOLO significantly outperforms other state-of-the-art target detection models.
Chaojun Ni, Wenhui Jiang 0001, Qishou Zhu, Yuming Fang 0001
VCIP1