Shangjin Zhai

dblp:318/2381 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-7741-5595ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Robot navigation and mapping · 58% 3D vision · 25% Generative modeling · 15%
Computer graphics and multimedia
3 papers
Geometric modeling and processing · 46% Visual content generation and editing · 23% Rendering · 23%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › depth estimation
depth completion
1.622025
Depth Completion With Multiple Balanced Bases and Confidence for Dense Monocular SLAM · IEEE Trans. Vis. Comput. Graph. 2025
Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras · ICRA 2024
Robotics › Robot navigation and mapping
SLAM
1.422025
Depth Completion With Multiple Balanced Bases and Confidence for Dense Monocular SLAM · IEEE Trans. Vis. Comput. Graph. 2025
VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM · ICRA 2022
Computer vision › 3D vision
depth estimation
1.022025
Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras · ICRA 2024
Depth Completion With Multiple Balanced Bases and Confidence for Dense Monocular SLAM · IEEE Trans. Vis. Comput. Graph. 2025
Robotics › Robot navigation and mapping › SLAM › dense SLAM
dense monocular SLAM
0.912025
Depth Completion With Multiple Balanced Bases and Confidence for Dense Monocular SLAM · IEEE Trans. Vis. Comput. Graph. 2025
Machine learning › Generative modeling
diffusion model
0.912025
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation · CVPR 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation · CVPR 2025
Geometric modeling and processing
3d reconstruction
0.912025
PGSR: Planar-Based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction · IEEE Trans. Vis. Comput. Graph. 2025
Rendering
gaussian splatting
0.912025
PGSR: Planar-Based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction · IEEE Trans. Vis. Comput. Graph. 2025
Visual content generation and editing › scene authoring
scene generation
0.912025
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation · CVPR 2025
Geometric modeling and processing
surface reconstruction
0.912025
PGSR: Planar-Based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction · IEEE Trans. Vis. Comput. Graph. 2025
Robotics › Robot navigation and mapping › visual odometry
visual-inertial odometry
0.822024
Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency Situations · IEEE Trans. Vis. Comput. Graph. 2022
Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras · ICRA 2024
Robotics › Robot navigation and mapping › SLAM
dense SLAM
0.812024
Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras · ICRA 2024
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.812024
Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras · ICRA 2024
Robotics › Robot navigation and mapping
localization
0.612022
Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency Situations · IEEE Trans. Vis. Comput. Graph. 2022
Robotics › Robot navigation and mapping › localization
map-based localization
0.612022
Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency Situations · IEEE Trans. Vis. Comput. Graph. 2022
Robotics › Robot navigation and mapping › SLAM
planar SLAM
0.612022
VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM · ICRA 2022
Robotics › Robot navigation and mapping › SLAM › multi-sensor SLAM
visual-inertial SLAM
0.612022
VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM · ICRA 2022
Computer vision › 3D vision
novel view synthesis
0.312025
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation · CVPR 2025
Computer vision › Video understanding and tracking
spatiotemporal consistency
0.312025
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation · CVPR 2025
Virtual and augmented reality
augmented reality
0.212022
Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency Situations · IEEE Trans. Vis. Comput. Graph. 2022
Virtual and augmented reality › augmented reality
mobile augmented reality
0.212022
Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency Situations · IEEE Trans. Vis. Comput. Graph. 2022

Methods — techniques the papers use, named apart from their topics

autoregressive generation · 1.73d warping · 1.7bundle adjustment · 1.4feature tracking · 1.1unbiased depth rendering · 0.9planar-based gaussian splatting · 0.9geometric regularization · 0.9depth weight factors · 0.9confidence map prediction · 0.9sliding window optimization · 0.8multi-basis depth representation · 0.8fisheye camera model · 0.8structure from motion · 0.6ray casting · 0.6homography constraint · 0.6
YearPublicationVenuePosition
2025 StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
abstract
Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging. To address this, we propose StarGen, a novel framework that employs a pre-trained video diffusion model in an autoregressive manner for long-range scene generation. The generation of each video clip is conditioned on the 3D warping of spatially adjacent images and the tem-porally overlapping image from previously generated clips, improving spatiotemporal consistency in long-range scene generation with precise pose control. The spatiotemporal condition is compatible with various input conditions, facilitating diverse tasks, including sparse view interpolation, perpetual view generation, and layout-conditioned city generation. Quantitative and qualitative evaluations demonstrate StarGen’s superior scalability, fidelity, and pose accuracy compared to state-of-the-art methods. Project page: https://zju3dv.github.io/StarGen.
Shangjin Zhai, Zhichao Ye, Weijian Xie, Danpeng Chen, Nan Wang 0020, Haomin Liu, Guofeng Zhang 0001
CVPR1
2025 PGSR: Planar-Based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction
abstract
Recently, 3D Gaussian Splatting (3DGS) has attracted widespread attention due to its high-quality rendering, and ultra-fast training and rendering speed. However, due to the unstructured and irregular nature of Gaussian point clouds, it is difficult to guarantee geometric reconstruction accuracy and multi-view consistency simply by relying on image reconstruction loss. Although many studies on surface reconstruction based on 3DGS have emerged recently, the quality of their meshes is generally unsatisfactory. To address this problem, we propose a fast planar-based Gaussian splatting reconstruction representation (PGSR) to achieve high-fidelity surface reconstruction while ensuring high-quality rendering. Specifically, we first introduce an unbiased depth rendering method, which directly renders the distance from the camera origin to the Gaussian plane and the corresponding normal map based on the Gaussian distribution of the point cloud, and divides the two to obtain the unbiased depth. We then introduce single-view geometric, multi-view photometric, and geometric regularization to preserve global geometric accuracy. We also propose a camera exposure compensation model to cope with scenes with large illumination variations. Experiments on indoor and outdoor scenes show that the proposed method achieves fast training and rendering while maintaining high-fidelity rendering and geometric reconstruction, outperforming 3DGS-based and NeRF-based methods.
Danpeng Chen, Weicai Ye, Yifan Wang 0025, Weijian Xie, Shangjin Zhai, Nan Wang 0020, Haomin Liu, Hujun Bao, Guofeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.6
2025 Depth Completion With Multiple Balanced Bases and Confidence for Dense Monocular SLAM
abstract
Dense SLAM based on monocular cameras does indeed have immense application value in the field of AR/VR, especially when it is performed on a mobile device. In this article, we propose a novel method that integrates a light-weight depth completion network into a sparse SLAM system using a multi-basis depth representation, so that dense mapping can be performed online even on a mobile phone. Specifically, we present a specifically optimized multi-basis depth completion network, called BBC-Net, tailored to the characteristics of traditional sparse SLAM systems. BBC-Net can predict multiple balanced bases and a confidence map from a monocular image with sparse points generated by off-the-shelf keypoint-based SLAM systems. The final depth is a linear combination of predicted depth bases that can be easily optimized by tuning the corresponding weights. To seamlessly incorporate the weights into traditional SLAM optimization and ensure efficiency and robustness, we design a set of depth weight factors, which makes our network a versatile plug-in module, facilitating easy integration into various existing sparse SLAM systems and significantly enhancing global depth consistency through bundle adjustment. To verify the portability of our method, we integrate BBC-Net into two representative SLAM systems. The experimental results on various datasets show that the proposed method achieves better performance in monocular dense mapping than the state-of-the-art methods. We provide an online demo running on a mobile phone, which verifies the efficiency and mapping quality of the proposed method in real-world scenarios.
Weijian Xie, Guanyi Chu, Quanhao Qian, Yihao Yu, Danpeng Chen, Shangjin Zhai, Nan Wang 0020, Hujun Bao, Guofeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.7
2024 Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras
abstract
We propose a real-time visual-inertial dense SLAM system that utilizes the online data streams from back-to-back dual fisheye cameras setup, providing 360◦coverage of the environment. Firstly, we employ a sliding-window-based front-end to estimate real-time poses from the binocular fisheye images and IMU data. Then, we implement a lightweight panoramic depth completion network based on multi-basis depth representation. The network takes panoramic images (obtained by stitching dual-fisheye images with extrinsics and intrinsic parameters) and sparse depths (generated by the front-end local tracking) as input and predicts multiple depth bases along with corresponding confidence as output. The final dense depth is the linear combination of the multiple depth bases. Thanks to the multi-basis depth representation, we can continuously optimize the 360° depth with the traditional optimizer to achieve higher global consistency in depth. We conducted experiments on both simulated and real-world datasets to evaluate our method. The results demonstrate that the proposed method outperforms SoTA methods in terms of depth prediction and 3D reconstruction. In addition, we develop a demo that can run on a mobile to demonstrate the real-time capabilities of our method.
Weijian Xie, Guanyi Chu, Quanhao Qian, Yihao Yu, Shangjin Zhai, Danpeng Chen, Nan Wang 0020, Hujun Bao, Guofeng Zhang 0001
ICRA5
2022 VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM
abstract
In this paper, we propose a tightly-coupled SLAM system fused with RGB, Depth, IMU and structured plane information. Traditional sparse points based SLAM systems always maintain a mass of map points to model the environment. Huge number of map points bring us a high computational complexity, making it difficult to be deployed on mobile devices. On the other hand, planes are common structures in man-made environment especially in indoor environments. We usually can use a small number of planes to represent a large scene. So the main purpose of this article is to decrease the high complexity of sparse points based SLAM. We build a lightweight back-end map which consists of a few planes and map points to achieve efficient bundle adjustment (BA) with an equal or better accuracy. We use homography constraints to eliminate the parameters of numerous plane points in the optimization and reduce the complexity of BA. We separate the parameters and measurements in homography and point-to-plane constraints and compress the measurements part to further effectively im-prove the speed of BA. We also integrate the plane information into the whole system to realize robust planar feature extraction, data association, and global consistent planar reconstruction. Finally, we perform an ablation study and compare our method with similar methods in simulation and real environment data. Our system achieves obvious advantages in accuracy and efficiency. Even if the plane parameters are involved in the optimization, we effectively simplify the back-end map by using planar structures. The global bundle adjustment is nearly 2 times faster than the sparse points based SLAM algorithm.
Danpeng Chen, Weijian Xie, Shangjin Zhai, Nan Wang 0020, Hujun Bao, Guofeng Zhang 0001
ICRA4
2022 Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency Situations
abstract
In this paper, we present a novel monocular visual-inertial odometry system with pre-built maps deployed on the remote server, which can robustly run in real-time on a mobile device even in high latency situations. By tightly coupling VIO with geometric priors from pre-built maps, our system can tolerate the high latency and low frequency of global localization service, which is especially suitable for practical applications when the localization service is deployed on the remote server. Firstly, sparse point clouds are obtained from the dense mesh by the ray casting method according to the localization results. The dense mesh can be reconstructed from the point clouds generated by Structure-from-Motion. We directly use the sparse point clouds in feature tracking and state update to suppress drift. In the process of feature tracking, the high local accuracy of VIO is fully utilized to effectively remove outliers and make our system robust. The experiments on EurocMav datasets and simulation datasets show that compared with state-of-the-art methods, our method can achieve better results in terms of both precision and robustness. The effectiveness of the proposed method is further demonstrated through a real-time AR demo on a mobile phone with the aid of visual localization on the remote server.
Hujun Bao, Weijian Xie, Quanhao Qian, Danpeng Chen, Shangjin Zhai, Nan Wang 0020, Guofeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.5