EDBT 2026 Demo / reviewers in the wild / expert
Kaiqiang Xiong
dblp:343/8355
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-9744-1173ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video RepresentationabstractVolumetric video enables immersive free viewpoint rendering of dynamic real world scenes, yet existing methods struggle with long sequences and complex motions, often leading to temporal instability and visual artifacts. To address these challenges, we propose ATGS, a Gaussian splatting based framework for volumetric video reconstruction. Our key insight is that explicitly tracking long term complex motion with individual Gaussian primitives is inherently unstable. Instead, we organize Gaussians around time conditioned anchors that localize their spatial and temporal support, thereby reducing long range motion complexity. We further introduce a temporal windowing strategy to activate only anchors relevant to the queried time, which improves scalability and temporal coherence. In addition, to ensure spatial and temporal stability, we design a compact set of multi level anchor features that encode global features, local spatial features, and local temporal features, jointly constraining Gaussian generation. Extensive experiments demonstrate that ATGS consistently outperforms prior methods on long sequence volumetric videos with complex motions. Project page: https://github.com/WuJH2001/ATGS. Kaiqiang Xiong, Xiang Li 0225, Xiaoyun Zheng, Chao Wang 0037, Ronggang Wang |
ACM Trans. Graph. | 5 |
| 2025 | LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature DecouplingabstractDue to the complex and highly dynamic motions in the real world, synthesizing dynamic videos from multi-view inputs for arbitrary viewpoints is challenging. Previous works based on neural radiance field or 3D Gaussian splatting are limited to modeling fine-scale motion, greatly restricting their application. In this paper, we introduce LocalDyGS, which consists of two parts to adapt our method to both large-scale and fine-scale motion scenes: 1) We decompose a complex dynamic scene into streamlined local spaces defined by seeds, enabling global modeling by capturing motion within each local space. 2) We decouple static and dynamic features for local space motion modeling. A static feature shared across time steps captures static information, while a dynamic residual field provides time-specific features. These are combined and decoded to generate Temporal Gaussians, modeling motion within each local space. As a result, we propose a novel dynamic scene reconstruction framework to model highly dynamic real-world scenes more realistically. Our method not only demonstrates competitive performance on various fine-scale datasets compared to state-of-the-art (SOTA) methods, but also represents the first attempt to model larger and more complex highly dynamic scenes. Project page: https://wujh2001.github.io/LocalDyGS/. Jianbo Jiao, Luyang Tang, Kaiqiang Xiong, Jinbo Yan, Runling Liu, Ronggang Wang |
ICCV | 6 |
| 2025 | Swift4D: Adaptive divide-and-conquer Gaussian Splatting for compact and efficient reconstruction of dynamic sceneabstractNovel view synthesis has long been a practical but challenging task, although the introduction of numerous methods to solve this problem, even combining advanced representations like 3D Gaussian Splatting, they still struggle to recover high-quality results and often consume too much storage memory and training time.
In this paper we propose Swift4D, a divide-and-conquer 3D Gaussian Splatting method that can handle static and dynamic primitives separately, achieving a good trade-off between rendering quality and efficiency, motivated by the fact that most of the scene is the static primitive and does not require additional dynamic properties. Concretely, we focus on modeling dynamic transformations only for the dynamic primitives which benefits both efficiency and quality. We first employ a learnable decomposition strategy to separate the primitives, which relies on an additional parameter to classify primitives as static or dynamic. For the dynamic primitives, we employ a compact multi-resolution 4D Hash mapper to transform these primitives from canonical space into deformation space at each timestamp, and then mix the static and dynamic primitives to produce the final output. This divide-and-conquer method facilitates efficient training and reduces storage redundancy. Our method not only achieves state-of-the-art rendering quality while being 20× faster in training than previous SOTA methods with a minimum storage requirement of only 30MB on real-world datasets. Luyang Tang, Jinbo Yan, Kaiqiang Xiong, Ronggang Wang |
ICLR | 7 |
| 2025 | SAP: Exact Sorting in Splatting via Screen-Aligned PrimitivesabstractRecently, 3D Gaussian Splatting (3DGS) has achieved state-of-the-art rendering results. However, its efficiency relies on simplifications that disregard the thickness of Gaussian primitives and their overlapping interactions. These simplifications can lead to popping artifacts due to inaccurate sorting, thereby affecting the rendering quality. In this paper, we propose Screen-Aligned Primitives (SAP), an anisotropic kernel that generates primitives parallel to the image plane for each view. Our rasterization pipeline enables full per-pixel ordering in real time. Since the primitives are parallel for a given viewpoint, a single global sorting operation suffices for correct per-pixel depth ordering. We formulate 3D reconstruction as a combination of a 3D-consistent decoder and 2D view-specific primitives, and further propose a highly efficient decoder to ensure 3D consistency. Moreover, within our framework, the primitive function values remain consistent between view space and screen space, allowing arbitrary radial basis functions (RBFs) to represent the scene without introducing projection errors. Experiments on diverse datasets demonstrate that our method achieves state-of-the-art rendering quality while maintaining real-time performance. Zhanke Wang, Kaiqiang Xiong, Ronggang Wang |
NeurIPS | 3 |
| 2025 | SFR-GS: Spatial-Frequency Domain Regularization for 3D Gaussian Splatting
Guanhua Wu, Kaiqiang Xiong, Zhanke Wang, Ronggang Wang |
PRCV (10) | 3 |
| 2025 | MVD-HuGaS: Human Gaussians from a Single Image via 3D Human Multi-View Diffusion Prior
Kaiqiang Xiong, Jianbo Jiao, Huachen Gao, Ronggang Wang |
PRCV (10) | 1 |
| 2024 | Surface-Centric Modeling for High-Fidelity Generalizable Neural Surface Reconstruction
Shihe Shen, Kaiqiang Xiong, Huachen Gao, Jianbo Jiao, Ronggang Wang |
ECCV (32) | 3 |
| 2024 | Disentangled Generation and Aggregation for Robust Radiance Fields
Shihe Shen, Huachen Gao, Wangze Xu, Luyang Tang, Kaiqiang Xiong, Jianbo Jiao, Ronggang Wang |
ECCV (49) | 6 |
| 2024 | FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth ConsistencyabstractLearning neural radiance fields (NeRF) without camera poses has been widely studied. However, recent methods lack explicit and effective supervision for pose estimation, resulting in ambiguous optimization of camera pose and NeRF geometry during joint training, particularly in scenarios involving large camera movements. In this paper, we propose FDCNeRF that leverages the direction information contained in the RGB-based optical flow and depth-based virtual flow as a direct guidance for camera pose optimization to reduce pose-geometry ambiguity. Additionally, we introduce Adaptive Pose-Aware Sampling (APAS) to replace the previous random ray sampling strategy, which reduces the difficulty of pose learning in early stages and preserves the diversity of rays in later stages. Experiments on the challenging Tanks and Temples dataset demonstrate that our method achieves state-of-the-art results in both novel view synthesis quality and pose estimation accuracy. Huachen Gao, Shihe Shen, Kaiqiang Xiong, Zhirui Gao, Yugui Xie, Ronggang Wang |
ICASSP | 4 |
| 2024 | High Fidelity Aggregated Planar Prior Assisted PatchMatch Multi-View StereoabstractThe quality of 3D models reconstructed by PatchMatch Multi-View Stereo remains a challenging problem due to unreliable photometric consistency in object boundaries and textureless areas. Since textureless areas usually exhibit strong planarity, previous methods used planar prior to improve the reconstruction performance. However, their planar prior ignores the depth discontinuity at the object boundary, making the boundary inaccurate (not sharp). In addition, due to the unreliable planar models in large-scale low-textured objects, the reconstruction results are incomplete. To address the above issues, we introduce the segmentation generated from Segment Anything Model into PatchMatch. Using segmentation to determine whether the depth is continuous based on the characteristics of segmentation and depth sharing boundaries. Then we construct Boundary Plane that fits the object boundary and Object Plane to increase consistency of planes in large-scale textureless objects. Finally, we use a probability graph model to calculate Aggregated Prior guided by Multiple Planes and embed it into the matching cost. The experimental results indicate that our method achieves SOTA in boundary sharpness on ETH3D and improves the completeness of weakly textured objects. Rongjie Wang 0004, Rui Peng 0011, Zhe Zhang 0049, Kaiqiang Xiong, Ronggang Wang |
ACM Multimedia | 5 |
| 2023 | CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive LearningabstractUnsupervised Multi-View Stereo (MVS) methods have achieved promising progress recently. However, previous methods primarily depend on the photometric consistency assumption, which may suffer from two limitations: indistinguishable regions and view-dependent effects, e.g., low-textured areas and reflections. To address these issues, in this paper, we propose a new dual-level contrastive learning approach, named CL-MVSNet. Specifically, our model integrates two contrastive branches into an unsupervised MVS framework to construct additional supervisory signals. On the one hand, we present an image-level contrastive branch to guide the model to acquire more context awareness, thus leading to more complete depth estimation in indistinguishable regions. On the other hand, we exploit a scene-level contrastive branch to boost the representation ability, improving robustness to view-dependent effects. Moreover, to recover more accurate 3D geometry, we introduce an ℒ0.5 photometric consistency loss, which encourages the model to focus more on accurate points while mitigating the gradient penalty of undesirable ones. Extensive experiments on DTU and Tanks&Temples benchmarks demonstrate that our approach achieves state-of-the-art performance among all end-to-end unsupervised MVS frameworks and outperforms its supervised counterpart by a considerable margin without fine-tuning. Kaiqiang Xiong, Tianxing Feng, Jianbo Jiao, Feng Gao 0014, Ronggang Wang |
ICCV | 1 |
| 2023 | Context-Guided Multi-view Stereo with Depth Back-Projection
Tianxing Feng, Kaiqiang Xiong, Ronggang Wang |
MMM (2) | 3 |