VLDB 2026 Research / reviewers in the wild / expert
Guangkai Xu
dblp:313/2105
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-9669-0381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
3D vision · 59% Generative modeling · 26% Segmentation and scene understanding · 15% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.6 | 2 | 2025 | DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation · AAAI 2025 Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation · NeurIPS 2024 |
Computer vision › 3D vision
3d reconstruction |
1.5 | 2 | 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction · ICCV 2025 FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models · ICCV 2023 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
1.5 | 2 | 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025 FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models · ICCV 2023 |
Computer vision › 3D vision
depth estimation |
1.1 | 3 | 2025 | FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models · ICCV 2023 DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation · AAAI 2025 Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints · ICRA 2024 |
Computer vision › 3D vision
camera calibration |
0.9 | 1 | 2025 | DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation · AAAI 2025 |
Machine learning › Generative modeling › diffusion model › diffusion-based perception
diffusion-based dense prediction |
0.9 | 1 | 2025 | DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation · AAAI 2025 |
Machine learning › Generative modeling › diffusion model › diffusion model training
diffusion model fine-tuning |
0.9 | 1 | 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025 |
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction |
0.9 | 1 | 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.9 | 1 | 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025 |
Computer vision › 3D vision › camera calibration
monocular camera calibration |
0.9 | 1 | 2025 | DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation · AAAI 2025 |
Computer vision › Segmentation and scene understanding › semantic segmentation
few-shot segmentation |
0.8 | 1 | 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation · NeurIPS 2024 |
Computer vision › 3D vision › 3d scene reconstruction
indoor scene reconstruction |
0.8 | 1 | 2024 | Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints · ICRA 2024 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.8 | 1 | 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation · NeurIPS 2024 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction › neural surface reconstruction
neural implicit surface reconstruction |
0.8 | 1 | 2024 | Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints · ICRA 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation · NeurIPS 2024 |
Computer vision › 3D vision › depth estimation
monocular depth prior |
0.2 | 1 | 2024 | Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints · ICRA 2024 |
Methods — techniques the papers use, named apart from their topics
temporal motion modeling · 0.9one-step fine-tuning · 0.9diffusion prior · 0.9diffusion model · 0.9RANSAC · 0.9two-stage training · 0.8self-supervised consistency constraints · 0.8neural implicit surface · 0.8in-context segmentation · 0.8KV fusion · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map GenerationabstractMonocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data. Recently, diffusion models trained on expansive datasets have been confirmed to maintain the capability to generate diverse, high-quality images. This success suggests a strong potential of the models to effectively understand varied visual information. In this work, we leverage the comprehensive visual knowledge embedded in pre-trained diffusion models to enable more robust and accurate monocular camera intrinsic estimation. Specifically, we reformulate the problem of estimating the four degrees of freedom (4-DoF) of camera intrinsic parameters as a dense incident map generation task. The map details the angle of incidence for each pixel in the RGB image, and its format aligns well with the paradigm of diffusion models. The camera intrinsic then can be derived from the incident map with a simple non-learning RANSAC algorithm during inference. Moreover, to further enhance the performance, we jointly estimate a depth map to provide extra geometric information for the incident map estimation. Extensive experiments on multiple testing datasets demonstrates that our model achieves state-of-the-art performance, gaining up to a 40% reduction in prediction errors. Besides, the experiments also show that the precise camera intrinsic and depth maps estimated by our pipeline can greatly benefit practical applications such as 3D reconstruction from a single in-the-wild image. Xiankang He, Guangkai Xu, Dongyan Guo |
AAAI | 2 |
| 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction
Songyan Zhang, Yongtao Ge, Jinyuan Tian, Guangkai Xu, Hao Chen 0041, Chunhua Shen |
ICCV | 4 |
| 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?abstractExtensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for a few dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) As a special case of the diffusion scheduler by setting its hyper-parameters, the multi-step generation can be simplified to a one-step fine-tuning paradigm without any loss of performance, while significantly speeding up inference. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific supervision can be beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailored for dense visual perception tasks exploiting diffusion priors. Different from the previous multi-step methods, our paradigm offers a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for task-specific supervision, which can be critical for improving the fine-grained details of predictions. Comprehensive experiments on a diverse set of dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method. Code: https://github.com/aim-uofa/GenPercept Guangkai Xu, Yongtao Ge, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen 0041, Chunhua Shen |
ICLR | 1 |
| 2024 | Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraintsabstract3D scene reconstruction from 2D images has been a long-standing task. Instead of estimating per-frame depth maps and fusing them in 3D, recent researches leverage the neural implicit surface as a global representation for 3D reconstruction. Equipped with data-driven pre-trained geometric cues, these methods have demonstrated promising performance. However, the inevitable inaccurate estimation of priors can lead to suboptimal reconstruction quality, particularly in some geometrically complex regions. In this paper, we propose a two-stage training process to further improve the reconstruction quality. It decouples the view-dependent and view-independent colors, and leverages two novel consistency constraints to enhance detail reconstruction performance without requiring extra priors. Additionally, we introduce an essential mask scheme to adaptively influence the selection of supervision constraints, thereby improving performance in a self-supervised paradigm. Experiments on synthetic and real-world datasets show the capability of reducing the side effects of inaccurately estimated priors and achieving high-quality scene reconstruction with rich geometric details. Liqin Lu, Jintao Rong 0001, Guangkai Xu, Linlin Ou |
ICRA | 4 |
| 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic SegmentationabstractThe Diffusion Model has not only garnered noteworthy achievements in the realm of image generation
but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data.
Drawing from the extensive potential unveiled by the Diffusion Model in both semantic correspondence and open vocabulary segmentation, our work initiates an investigation into employing the Latent Diffusion Model for Few-shot Semantic Segmentation.
Recently, inspired by the in-context learning ability of large language models, Few-shot Semantic Segmentation has evolved into In-context Segmentation tasks, morphing into a crucial element in assessing generalist segmentation models.
In this context, we concentrate
on Few-shot Semantic Segmentation,
establishing a solid foundation for the future development of a Diffusion-based generalist model for segmentation. Our initial focus lies in understanding how to facilitate interaction between the query image and the support image, resulting in the proposal of a KV fusion method within the self-attention framework.
Subsequently, we delve deeper into optimizing the infusion of information from the support mask and simultaneously re-evaluating how to provide reasonable supervision from the query mask.
Based on our analysis, we establish a simple and effective framework named DiffewS, maximally retaining the original Latent Diffusion Model's generative framework and effectively utilizing the pre-training prior. Experimental results demonstrate that our method significantly outperforms the previous SOTA models in multiple settings. Muzhi Zhu, Yang Liu 0357, Zekai Luo, Chenchen Jing, Hao Chen 0041, Guangkai Xu, Chunhua Shen |
NeurIPS | 6 |
| 2023 | FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Modelsabstract3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but may face catastrophic failures due to the reliance on accurate pixel correspondence across views, while the latter mitigates these issues by learning 2D or 3D representation directly. However, without a largescale video or 3D training data, it can hardly be generalized to diverse real-world scenarios due to the presence of tens of millions or even billions of optimization parameters in the deep network.Recently, robust monocular depth estimation models trained with large-scale datasets have been proven to possess weak 3D geometry prior, but they are insufficient for reconstruction due to the unknown camera parameters, the affine-invariant property, and inter-frame inconsistency. To address these issues, we propose a novel test-time optimization approach that can transfer the robustness of affine- invariant depth models such as LeReS to challenging diverse scenes while ensuring inter-frame consistency, with only dozens of parameters to optimize per video frame. Specifically, our approach involves freezing the pre-trained affine-invariant depth model’s depth predictions, rectifying them by optimizing the unknown scale-shift values with a geometric consistency alignment module, and employing the resulting scale-consistent depth maps to robustly obtain camera poses and achieve dense scene reconstruction, even in low-texture regions. Experiments show that our method achieves state-of-the-art cross-dataset reconstruction on five zero-shot testing datasets. Code is available at: https://aim-uofa.github.io/FrozenRecon/ Guangkai Xu, Wei Yin 0006, Hao Chen 0041, Chunhua Shen, Feng Zhao 0004 |
ICCV | 1 |