Yingshuang Zou

dblp:342/7291 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0002-2881-0518ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with Transformers
abstract
Compared with previous 3D reconstruction methods like Nerf, recent Generalizable 3D Gaussian Splatting (G-3DGS) methods demonstrate impressive efficiency even in the sparse-view setting. However, the promising reconstruction performance of existing G-3DGS methods relies heavily on accurate multi-view feature matching, which is quite challenging. Especially for the scenes that have many non-overlapping areas between various views and contain numerous similar regions, the matching performance of existing methods is poor and the reconstruction precision is limited. To address this problem, we develop a strategy that utilizes a predicted depth confidence map to guide accurate local feature matching. In addition, we propose to utilize the knowledge of existing monocular depth estimation models as prior to boost the depth estimation precision in non-overlapping areas between views. Combining the proposed strategies, we present a novel G-3DGS method named TranSplat, which obtains the best performance on both the RealEstate10K and ACID benchmarks while maintaining competitive speed and presenting strong cross-dataset generalization ability.
Chuanrui Zhang, Yingshuang Zou, Zhuoling Li, Minmin Yi, Haoqian Wang
AAAI2
2025 UniScene: Unified Occupancy-centric Driving Scene Generation
abstract
Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/.
Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014
CVPR4
2025 DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation
abstract
Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. A key challenge lies in finding an efficient and generalizable geometric representation that seamlessly connects temporal and spatial synthesis. To address this, we propose DiST-4D, the first disentangled spatiotemporal diffusion framework for 4D driving scene generation, which leverages metric depth as the core geometric representation. DiST-4D decomposes the problem into two diffusion processes: DiST-T, which predicts future metric depth and multi-view RGB sequences directly from past observations, and DiST-S, which enables spatial NVS by training only on existing viewpoints while enforcing cycle consistency. This cycle consistency mechanism introduces a forward-backward rendering constraint, reducing the generalization gap between observed and unseen viewpoints. Metric depth is essential for both accurate reliable forecasting and accurate spatial NVS, as it provides a view-consistent geometric representation that generalizes well to unseen perspectives. Experiments demonstrate that DiST-4D achieves state-of-the-art performance in both temporal prediction and NVS tasks, while also delivering competitive performance in planning-related evaluations.
Jiazhe Guo, Yikang Ding, Xiwu Chen, Bohan Li 0015, Yingshuang Zou, Xiaoyang Lyu, Feiyang Tan, Xiaojuan Qi 0001, Hao Zhao 0002
ICCV6
2025 Leveraging 2D Annotations for Cost-Effective Dynamic Urban Scene Reconstruction
abstract
Accurate reconstruction of dynamic objects in urban scenes remains a challenging task. Existing methods typically rely on complex 3D annotations to identify dynamic objects, which are expensive and difficult to obtain. In this paper, we propose a novel approach to address this challenge by leveraging accessible and effective 2D annotations. We utilize a 2D foundation model to identify dynamic objects within the image and lift this knowledge to 3D space, obtaining the 3D point cloud of dynamic objects. To capture the rigid motion patterns of dynamic objects, we introduce an object-level motion model together with the local aggregation strategy. Additionally, we propose a structural consistency constraint to maintain the consistency of dynamic objects during motion, significantly improving the accuracy of scene reconstruction. Extensive experiments demonstrate that our method achieves comparable reconstruction performance to methods relying on 3D annotations, providing a cost-effective and accurate solution for dynamic urban scene reconstruction.
Chuming Wang, Yingshuang Zou, Haoqian Wang
ICME2
2025 SLGaussian: Fast Language Gaussian Splatting in Sparse Views
abstract
3D semantic field learning is crucial for applications like autonomous navigation, AR/VR, and robotics, where accurate comprehension of 3D scenes from limited viewpoints is essential. Existing methods struggle under sparse view conditions, relying on inefficient per-scene multi-view optimizations, which are impractical for many real-world tasks. To address this, we propose SLGaussian, a feed-forward method for constructing 3D semantic fields from sparse viewpoints, allowing direct inference of 3DGS-based scenes. By ensuring consistent SAM segmentations through video tracking and using low-dimensional indexing for high-dimensional CLIP features, SLGaussian efficiently embeds language information in 3D space, offering a robust solution for accurate 3D scene understanding under sparse view conditions. In experiments on two-view sparse 3D object querying and segmentation in the LERF and 3D-OVS datasets, SLGaussian outperforms existing methods in chosen IoU, Localization Accuracy, and mIoU. Moreover, our model achieves scene inference in under 30 seconds and open-vocabulary querying in just 0.011 seconds per query.
Kangjie Chen, BingQuan Dai, Minghan Qin, Dongbin Zhang, Peihao Li 0003, Yingshuang Zou, Haoqian Wang
ACM Multimedia6
2024 M2Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation
Yingshuang Zou, Yikang Ding, Xi Qiu, Haoqian Wang
ECCV (46)1