Kaizhi Yang

dblp:256/7058 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hie4DGS: Hierarchical 4D Gaussian Splatting From Monocular Dynamic Video
abstract
Monocular dynamic video reconstruction is a typical ill-posed problem due to the limited observations and complex 3D motions. Despite the recent advances in dynamic 3D Gaussian splatting techniques, most of them still struggle with the monocular setting, since they heavily rely on geometric cues from multiple cameras or ignore the structural coherence among the optimized 3D Gaussains. To address this, we propose Hie4DGS, a novel hierarchical structure representation to model the complex dynamic motions from monocular dynamic videos. Specifically, we decompose the motions of a dynamic scene into groups of multiple structure granularities and progressively compose them to derive the motion of each 3D Gaussian. Building on this representation, we leverage hierarchical semantic segmentation to group Gaussians and initialize their motion using depth and tracking priors within each group. Additionally, we introduce a structure rendering loss that enforces consistency between the learned motion structure and semantic priors, further reducing motion ambiguity. Compared to the state-of-the-art dynamic Gaussian methods, we achieve significant improvement in rendering quality on monocular video datasets featuring complex real-world motions.
Kaizhi Yang, Xiaoxiao Long, Xuejin Chen
IEEE Trans. Vis. Comput. Graph.2
2025 IMLS-Splatting: Efficient Mesh Reconstruction from Multi-view Images via Point Representation
abstract
Multi-view mesh reconstruction has long been a challenging problem in graphics and computer vision. In contrast to recent volumetric rendering methods that generate meshes through post-processing, we propose an end-to-end mesh optimization approach called IMLS-Splatting. Our method leverages the sparsity and flexibility of point clouds to efficiently represent the underlying surface. To achieve this, we introduce a splatting-based differentiable Implicit Moving-Least Squares (IMLS) algorithm that enables the fast conversion of point clouds into SDFs and texture fields, optimizing both mesh reconstruction and rasterization. Additionally, the IMLS representation ensures that the reconstructed SDF and mesh maintain continuity and smoothness without the need for extra regularization. With this efficient pipeline, our method enables the reconstruction of highly detailed meshes in approximately 11 minutes, supporting high-quality rendering and achieving state-of-the-art reconstruction performance. Our code is available at https://github.com/SilenKZYoung/IMLS-Splatting.
Kaizhi Yang, Liu Dai, Isabella Liu, Xiaoshuai Zhang, Xiaoyan Sun 0001, Xuejin Chen, Zexiang Xu, Hao Su 0001
ACM Trans. Graph.1
2024 MovingParts: Motion-based 3D Part Discovery in Dynamic Radiance Field
abstract
We present MovingParts, a NeRF-based method for dynamic scene reconstruction and part discovery. We consider motion as an important cue for identifying parts, that all particles on the same part share the common motion pattern. From the perspective of fluid simulation, existing deformation-based methods for dynamic NeRF can be seen as parameterizing the scene motion under the Eulerian view, i.e., focusing on specific locations in space through which the fluid flows as time passes. However, it is intractable to extract the motion of constituting objects or parts using the Eulerian view representation. In this work, we introduce the dual Lagrangian view and enforce representations under the Eulerian/Lagrangian views to be cycle-consistent. Under the Lagrangian view, we parameterize the scene motion by tracking the trajectory of particles on objects. The Lagrangian view makes it convenient to discover parts by factorizing the scene motion as a composition of part-level rigid motions. Experimentally, our method can achieve fast and high-quality dynamic scene reconstruction from even a single moving camera, and the induced part-based representation allows direct applications of part tracking, animation, 3D scene editing, etc.
Kaizhi Yang, Xiaoshuai Zhang, Zhiao Huang, Xuejin Chen, Zexiang Xu, Hao Su 0001
ICLR1
2024 GaussianPro: 3D Gaussian Splatting with Progressive Propagation
abstract
3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain texture-less surfaces, SfM techniques fail to produce enough points in these surfaces and cannot provide good initialization for 3DGS. As a result, 3DGS suffers from difficult optimization and low-quality renderings. In this paper, inspired by classic multi-view stereo (MVS) techniques, we propose GaussianPro, a novel method that applies a progressive propagation strategy to guide the densification of the 3D Gaussians. Compared to the simple split and clone strategies used in 3DGS, our method leverages the priors of the existing reconstructed geometries of the scene and utilizes patch matching to produce new Gaussians with accurate positions and orientations. Experiments on both large-scale and small-scale scenes validate the effectiveness of our method. Our method significantly surpasses 3DGS on the Waymo dataset, exhibiting an improvement of 1.15dB in terms of PSNR. Codes and data are available at https://github.com/kcheng1021/GaussianPro.
Xiaoxiao Long, Kaizhi Yang, Yao Yao 0008, Wei Yin 0006, Yuexin Ma, Wenping Wang 0001, Xuejin Chen
ICML3
2023 DPF-Net: Combining Explicit Shape Priors in Deformable Primitive Field for Unsupervised Structural Reconstruction of 3D Objects
abstract
Unsupervised methods for reconstructing structures face significant challenges in capturing the geometric details with consistent structures among diverse shapes of the same category. To address this issue, we present a novel unsupervised structural reconstruction method, named DPF-Net, based on a new Deformable Primitive Field (DPF) representation, which allows for high-quality shape reconstruction using parameterized geometric primitives. We design a two-stage shape reconstruction pipeline which consists of a primitive generation module and a primitive deformation module to approximate the target shape of each part progressively. The primitive generation module estimates the explicit orientation, position, and size parameters of parameterized geometric primitives, while the primitive deformation module predicts a dense deformation field based on a parameterized primitive field to recover shape details. The strong shape prior encoded in parameterized geometric primitives enables our DPF-Net to extract high-level structures and recover fine-grained shape details consistently. The experimental results on three categories of objects in diverse shapes demonstrate the effectiveness and generalization ability of our DPF-Net on structural reconstruction and shape segmentation.
Qingyao Shuai, Chi Zhang 0044, Kaizhi Yang, Xuejin Chen
ICCV3
2022 Fast Generation of Deceptive Jamming Signal Against Spaceborne SAR Based on Spatial Frequency Domain Interpolation
abstract
This article proposes a novel method for the generation of the deceptive jamming signal against spaceborne SAR, which can calculate the jammer’s frequency response (JFR) efficiently. The first advantage of this algorithm is that the real-time computational complexity is independent of the number of fake scattering units, so it is particularly suitable for generating scenes containing a large number of scatterers. Besides, the imaging quality of fake targets is also improved significantly, especially in the case of squint geometries. First, using the Taylor series expansion and approximations, we derive the recurrence relationship of the slant range between the adjacent pulse repetition intervals. Based on this relationship, the JFR at each azimuth time is mapped to a special spatial spectrum, which is the 2-D Fourier transform along the range and azimuth dimensions of the JFR matrix at the time of the first pulse arrival. Therefore, after the initial JFR is calculated, the subsequent JFRs can be quickly obtained by sinc function interpolation. Then, the validity and effective region of this algorithm are estimated by the theoretical analysis of the slant range error. The simulation results and a computational complexity analysis verify that the algorithm demonstrates superior performance in terms of imaging quality and efficiency.
Kaizhi Yang, Fangfang Ma, Da Ran, Guojing Li
IEEE Trans. Geosci. Remote. Sens.1
2021 Learning Scale-Adaptive Representations for Point-Level LiDAR Semantic Segmentation
abstract
A large number of objects with various scales and categories in autonomous driving scenes pose a great challenge to LiDAR semantic segmentation. Voxel-based 3D convolutional networks have been widely employed by existing state-of-the-art methods to extract features with different spatial scales. However, the voxel network architecture limits its effectiveness in combining multi-scale features for point-level discrimination. In this paper, we propose point-wise prediction by taking the geometric structure of the original point cloud into account. We propose a Scale-Adaptive Fusion (SAF) module that progressively and selectively fuses multi-scale features to deal with scale variations across objects adaptively. Moreover, we propose a novel Local Point Refinement (LPR) module to address the quantization loss problem of voxel-based methods. Our approach achieves state-of-the-art performance on three public datasets, i.e., Semantic-KITTI, Semantic-POSS, and nuScenes dataset, while greatly improving the computational and memory efficiency.
Tongfeng Zhang, Kaizhi Yang, Xuejin Chen
3DV2
2021 Deep 3D Modeling of Human Bodies from Freehand Sketching
Kaizhi Yang, Jintao Lu, Siyu Hu, Xuejin Chen
MMM (2)1
2021 Unsupervised learning for cuboid shape abstraction via joint segmentation from point clouds
abstract
Representing complex 3D objects as simple geometric primitives, known as shape abstraction, is important for geometric modeling, structural analysis, and shape synthesis. In this paper, we propose an unsupervised shape abstraction method to map a point cloud into a compact cuboid representation. We jointly predict cuboid allocation as part segmentation and cuboid shapes and enforce the consistency between the segmentation and shape abstraction for self-learning. For the cuboid abstraction task, we transform the input point cloud into a set of parametric cuboids using a variational auto-encoder network. The segmentation network allocates each point into a cuboid considering the point-cuboid affinity. Without manual annotations of parts in point clouds, we design four novel losses to jointly supervise the two branches in terms of geometric similarity and cuboid compactness. We evaluate our method on multiple shape collections and demonstrate its superiority over existing shape abstraction methods. Moreover, based on our network architecture and learned representations, our approach supports various applications including structured shape generation, shape interpolation, and structural shape clustering.
Kaizhi Yang, Xuejin Chen
ACM Trans. Graph.1