VLDB 2026 Research / reviewers in the wild / expert
Kangkan Wang
dblp:148/2119
· DBLP profile ↗
23ranked-venue papers
14as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionabstractLong-tail recognition remains challenging for pre-trained foundation models like CLIP, which often suffer from performance degradation under imbalanced data. This stems not only from the overfitting/underfitting issues during fine-tuning but, more fundamentally, from the inherent bias inherited from the long-tail distribution of their massive pre-training datasets. To address this, we propose HGLTR (Hierarchy-Guided Long-Tail Recognition), a novel framework that calibrates pre-trained models by injecting objective class hierarchy knowledge. We argue that the semantic proximity defined by a hierarchy provides a robust, data-independent prior to counteract model bias. Our method is specifically designed for vision-language models' dual-modality architecture. At the feature level, we align image embeddings with a hierarchy-guided text similarity structure. At the classifier level, we employ a distillation loss to regularize predictions using soft labels derived from the hierarchy. This dual-level injection effectively transfers knowledge from head to tail classes. Experiments on ImageNet-LT, Places-LT, and iNaturalist 2018 demonstrate that HGLTR achieves state-of-the-art performance, particularly in tail-classes accuracy, highlighting the importance of leveraging structural priors to calibrate foundation models for real-world data. Jinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan, Zijian Tao, Songcan Chen, Kangkan Wang |
AAAI | 7 |
| 2026 | Dynamic view synthesis with topologically-varying neural radiance fields from sparse input views
Kangkan Wang, Kejie Wei, Shaoyuan Li |
Neurocomputing | 1 |
| 2026 | GSH3D: Efficient 3D Gaussian human generation from 2D image collections
Kangkan Wang, Miao Zhao, Shao-Yuan Li |
Knowl. Based Syst. | 1 |
| 2026 | One-shot novel view and pose human image synthesis via 3D prior guided diffusion model
Shenjian Gong, Kangkan Wang, Shanshan Zhang 0001, Jian Yang 0003 |
Pattern Recognit. | 2 |
| 2025 | InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention Priorabstract3D Gaussian Splatting based 3D editing has demonstrated impressive performance in recent years. However, the multi-view editing often exhibits significant local inconsistency, especially in areas of non-rigid deformation, which lead to local artifacts, texture blurring, or semantic variations in edited 3D scenes. We also found that the existing editing methods, which rely entirely on text prompts make the editing process a "one-shot deal", making it difficult for users to control the editing degree flexibly. In response to these challenges, we present InterGSEdit, a novel framework for high-quality 3DGS editing via interactively selecting key views with users' preferences. We propose a CLIP-based Semantic Consistency Selection (CSCS) strategy to adaptively screen a group of semantically consistent reference views for each user-selected key view. Then, the cross-attention maps derived from the reference views are used in a weighted Gaussian Splatting unprojection to construct the 3D Geometry-Consistent Attention Prior ($GAP^{3D}$). We project $GAP^{3D}$ to obtain 3D-constrained attention, which are fused with 2D cross-attention via Attention Fusion Network (AFN). AFN employs an adaptive attention strategy that prioritizes 3D-constrained attention for geometric consistency during early inference, and gradually prioritizes 2D cross-attention maps in diffusion for fine-grained features during the later inference. Extensive experiments demonstrate that InterGSEdit achieves state-of-the-art performance, delivering consistent, high-fidelity 3DGS editing with improved user experience. Minghao Wen, Shengjie Wu, Kangkan Wang, Dong Liang 0008 |
ICCV | 3 |
| 2025 | Robust contrastive knowledge distillation for long-tailed noisy class labels
Shao-Yuan Li, Jinpeng Zheng, Mingguang Zhang, Shaofang Li, Kangkan Wang |
Knowl. Based Syst. | 6 |
| 2025 | Prototypes as Anchors: Tackling Unseen Noise for online continual learning
Shaoyuan Li, Sheng-Jun Huang, Songcan Chen, Kangkan Wang |
Neural Networks | 5 |
| 2025 | CloCap-GS: Clothed Human Performance Capture With 3D Gaussian SplattingabstractCapturing the human body and clothing from videos has obtained significant progress in recent years, but several challenges remain to be addressed. Previous methods reconstruct the 3D bodies and garments from videos with self-rotating human motions or capture the body and clothing separately based on neural implicit fields. However, the reconstruction methods for self-rotating motions may cause instable tracking on dynamic videos with arbitrary human motions, while implicit fields based methods are limited to inefficient rendering and low quality synthesis. To solve these problems, we propose a new method, called CloCap-GS, for clothed human performance capture with 3D Gaussian Splatting. Specifically, we align 3D Gaussians with the deforming geometries of body and clothing, and leverage photometric constraints formed by matching Gaussians renderings with input video frames to recover temporal deformations of the dense template geometry. The geometry deformations and Gaussians properties of both the body and clothing are optimized jointly, achieving both dense geometry tracking and novel-view synthesis. In addition, we introduce a physics-aware material-varying cloth model to preserve physically-plausible cloth dynamics and body-clothing interactions that is pre-trained in a self-supervised manner without preparing training data. Compared with the existing methods, our method improves the accuracy of dense geometry tracking and quality of novel-view synthesis for a variety of daily garment types (e.g., loose clothes). Extensive experiments in both quantitative and qualitative evaluations demonstrate the effectiveness of CloCap-GS on real sparse-view or monocular videos. Kangkan Wang, Jian Yang 0003, Guofeng Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | A Pose-Aware Auto-Augmentation Framework for 3D Human Pose and Shape Estimation from Partial Point Clouds
Kangkan Wang, Shihao Yin, Chenghao Fang 0002 |
PRCV (6) | 1 |
| 2023 | Clothed Human Performance Capture with a Double-layer Neural Radiance FieldsabstractThis paper addresses the challenge of capturing performance for the clothed humans from sparse-view or monocular videos. Previous methods capture the performance of full humans with a personalized template or recover the garments from a single frame with static human poses. However, it is inconvenient to extract cloth semantics and capture clothing motion with one-piece template, while single frame-based methods may suffer from instable tracking across videos. To address these problems, we propose a novel method for human performance capture by tracking clothing and human body motion separately with a double-layer neural radiance fields (NeRFs). Specifically, we propose a double-layer NeRFsfor the body and garments, and track the densely deforming template of the clothing and body by jointly optimizing the deformation fields and the canonical double-layer NeRFs. In the optimization, we introduce a physics-aware cloth simulation network which can help generate physically plausible cloth dynamics and body-cloth interactions. Compared with existing methods, our method is fully differentiable and can capture both the body and clothing motion robustly from dynamic videos. Also, our method represents the clothing with an independent NeRFs, allowing us to model implicit fields of general clothes feasibly. The experimental evaluations validate its effectiveness on real multi-view or monocular videos. Kangkan Wang, Guofeng Zhang 0001, Suxu Cong, Jian Yang 0003 |
CVPR | 1 |
| 2023 | NerfCap: Human Performance Capture With Dynamic Neural Radiance FieldsabstractThis paper addresses the challenge of human performance capture from sparse multi-view or monocular videos. Given a template mesh of the performer, previous methods capture the human motion by non-rigidly registering the template mesh to images with 2D silhouettes or dense photometric alignment. However, the detailed surface deformation cannot be recovered from the silhouettes, while the photometric alignment suffers from instability caused by appearance variation in the videos. To solve these problems, we propose NerfCap, a novel performance capture method based on the dynamic neural radiance field (NeRF) representation of the performer. Specifically, a canonical NeRF is initialized from the template geometry and registered to the video frames by optimizing the deformation field and the appearance model of the canonical NeRF. To capture both large body motion and detailed surface deformation, NerfCap combines linear blend skinning with embedded graph deformation. In contrast to the mesh-based methods that suffer from fixed topology and texture, NerfCap is able to flexibly capture complex geometry and appearance variation across the videos, and synthesize more photo-realistic images. In addition, NerfCap can be pre-trained end to end in a self-supervised manner by matching the synthesized videos with the input videos. Experimental results on various datasets show that NerfCap outperforms prior works in terms of both surface reconstruction accuracy and novel-view synthesis quality. Kangkan Wang, Sida Peng, Xiaowei Zhou 0001, Jian Yang 0003, Guofeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | 3D human pose and shape estimation with dense correspondence from a single depth image
Kangkan Wang, Guofeng Zhang 0001, Jian Yang 0003 |
Vis. Comput. | 1 |
| 2021 | Adaptive Multi-Domain Learning for Outdoor 3d Human Pose and Shape EstimationabstractIt is an extremely challenging task to estimate 3D human pose and shape in outdoor scenes for which we can hardly obtain precise ground truth data for training. Previous methods usually use multiple datasets collected at different scenes to train their models, including those collected in laboratories with precise ground truth and those collected at outdoor scenes with estimated or even no ground truth. Since data from different scenes are included in training, it is necessary to handle the domain difference problem, which unfortunately has never been considered by previous works. In this paper, we first point out this problem and then address it via a novel cascade multi-domain learning module (CMDL), where multiple adapters are employed to extract more discriminative features for different domains. We show that our method with CMDL outperforms previous methods in outdoor scenes. In principle, the proposed CMDL module can be easily applied on top of any arbitrary 3D human pose and shape approach. Zhaoyang Gui, Shanshan Zhang 0001, Kangkan Wang, Jian Yang 0003, Pong C. Yuen |
ICASSP | 3 |
| 2021 | Unsupervised Detailed Human Shape Estimation from Multi-view Color Images
Huayu Zheng, Kangkan Wang, Jian Yang 0003 |
ICIG (2) | 2 |
| 2021 | Parametric Model Estimation for 3D Clothed Humans from Point CloudsabstractThis paper presents a novel framework to estimate parametric model- s for 3D clothed humans from partial point clouds. It is a challenging problem due to factors such as arbitrary human shape and pose, large variations in clothing details, and significant missing data. Existing methods mainly focus on estimating the parametric model of undressed bodies or reconstructing the non-parametric 3D shapes from point clouds. In this paper, we propose a hierarchical regression framework to learn the parametric model of detailed human shapes from partial point clouds of a single depth frame. Benefiting from the favorable ability of deep neural networks to model nonlinearity, the proposed framework cascades several successive regression networks to estimate the parameters of detailed 3D human body models in a coarse-to-fine manner. Specifically, the first global regression network extracts global deep features of point clouds to obtain an initial estimation of the undressed human model. Based on the initial estimation, the local regression network then refines the undressed human model by using the local features of neighborhood points of human joints. Finally, the clothing details are inferred as an additive displacement on the refined undressed model using the vertex-level regression network. The experimental results demonstrate that the proposed hierarchical regression approach can accurately predict detailed human shapes from partial point clouds and outperform prior works in the recovery accuracy of 3D human models. Kangkan Wang, Huayu Zheng, Guofeng Zhang 0001, Jian Yang 0003 |
ISMAR | 1 |
| 2021 | High accuracy and geometry-consistent confidence prediction network for multi-view stereo
Zhaoxin Li, Xiaoge Zhang 0003, Kangkan Wang, Hao Jiang 0013 |
Comput. Graph. | 3 |
| 2021 | Learning Dense Correspondences for Non-Rigid Point Clouds With Two-Stage RegressionabstractWe propose a novel deep learning method to predict dense correspondences for partial point clouds of non-rigidly deformable targets. Dense correspondences are learned in the form of vertex displacements of a template mesh towards the point clouds. A two-stage regression framework is proposed to estimate accurate displacement vectors, including the global and local regression networks. Specifically, the global regression network estimates global displacements from the global features of the template mesh and point clouds through a graph CNN based hierarchical encoder-decoder network. Based on the initial displacements, a mesh can be generated that fits to the point clouds roughly. In the local regression network, a local feature embedding layer fuses local features of point clouds with graph features on the generated mesh through an attention mechanism. Consequently, the embedded local features are employed to refine the correspondences in local regions of the targets by predicting the increments of vertex displacements. Our method is further generalized to correspondence estimation on unseen real data with a robust fine-tuning method. The experimental results on diverse datasets of various deformable subjects (e.g., human bodies, animals, and hands) demonstrate that the proposed approach can accurately and robustly estimate dense correspondences from non-rigid point clouds. Kangkan Wang, Guofeng Zhang 0001, Huayu Zheng, Jian Yang 0003 |
IEEE Trans. Image Process. | 1 |
| 2021 | Dynamic human body reconstruction and motion tracking with low-cost depth cameras
Kangkan Wang, Guofeng Zhang 0001, Jian Yang 0003, Hujun Bao |
Vis. Comput. | 1 |
| 2020 | Sequential 3D Human Pose and Shape Estimation From Point CloudsabstractThis work addresses the problem of 3D human pose and shape estimation from a sequence of point clouds. Existing sequential 3D human shape estimation methods mainly focus on the template model fitting from a sequence of depth images or the parametric model regression from a sequence of RGB images. In this paper, we propose a novel sequential 3D human pose and shape estimation framework from a sequence of point clouds. Specifically, the proposed framework can regress 3D coordinates of mesh vertices at different resolutions from the latent features of point clouds. Based on the estimated 3D coordinates and features at the low resolution, we develop a spatial-temporal mesh attention convolution (MAC) to predict the 3D coordinates of mesh vertices at the high resolution. By assigning specific attentional weights to different neighboring points in the spatial and temporal domains, our spatial-temporal MAC can capture structured spatial and temporal features of point clouds. We further generalize our framework to the real data of human bodies with a weakly supervised fine-tuning method. The experimental results on SURREAL, Human3.6M, DFAUST and the real detailed data demonstrate that the proposed approach can accurately recover the 3D body model sequence from a sequence of point clouds. Kangkan Wang, Jin Xie 0001, Guofeng Zhang 0001, Jian Yang 0003 |
CVPR | 1 |
| 2020 | 3D Human Body Shape and Pose Estimation from Depth Image
Kangkan Wang, Jian Yang 0003 |
PRCV (1) | 2 |
| 2017 | Templateless Non-Rigid Reconstruction and Motion Tracking With a Single RGB-D CameraabstractWe present a novel templateless approach for nonrigid reconstruction and motion tracking using a single RGB-D camera. Without any template prior, our system achieves accurate reconstruction and tracking for considerably deformable objects. To robustly register the input sequence of partial depth scans with dynamic motion, we propose an efficient local-to-global hierarchical optimization framework inspired by the idea of traditional structure-from-motion. Our proposed framework mainly consists of two stages, local nonrigid bundle adjustment and global optimization. To eliminate error accumulation during the nonrigid registration of loop motion sequences, we split the full sequence into several segments and apply local nonrigid bundle adjustment to align each segment locally. Global optimization is then adopted to combine all segments and handle the drift problem through loop-closure constraint. By fitting to the input partial data, a deforming 3D model sequence of dynamic objects is finally generated. Experiments on both synthetic and real test data sets and comparisons with state of the art demonstrate that our approach can handle considerable motions robustly and efficiently, and reconstruct high-quality 3D model sequences without drift. Kangkan Wang, Guofeng Zhang 0001, Shihong Xia |
IEEE Trans. Image Process. | 1 |
| 2014 | A Two-Stage Framework for 3D FaceReconstruction from RGBD ImagesabstractThis paper proposes a new approach for 3D face reconstruction with RGBD images from an inexpensive commodity sensor. The challenges we face are: 1) substantial random noise and corruption are present in low-resolution depth maps; and 2) there is high degree of variability in pose and face expression. We develop a novel two-stage algorithm that effectively maps low-quality depth maps to realistic face models. Each stage is targeted toward a certain type of noise. The first stage extracts sparse errors from depth patches through the data-driven local sparse coding, while the second stage smooths noise on the boundaries between patches and reconstructs the global shape by combining local shapes using our template-based surface refinement. Our approach does not require any markers or user interaction. We perform quantitative and qualitative evaluations on both synthetic and real test sets. Experimental results show that the proposed approach is able to produce high-resolution 3D face models with high accuracy, even if inputs are of low quality, and have large variations in viewpoint and face expression. Kangkan Wang, Xianwang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Robust 3D Reconstruction With an RGB-D CameraabstractWe present a novel 3D reconstruction approach using a low-cost RGB-D camera such as Microsoft Kinect. Compared with previous methods, our scanning system can work well in challenging cases where there are large repeated textures and significant depth missing problems. For robust registration, we propose to utilize both visual and geometry features and combine SFM technique to enhance the robustness of feature matching and camera pose estimation. In addition, a novel prior-based multicandidates RANSAC is introduced to efficiently estimate the model parameters and significantly speed up the camera pose estimation under multiple correspondence candidates. Even when serious depth missing occurs, our method still can successfully register all frames together. Loop closure also can be robustly detected and handled to eliminate the drift problem. The missing geometry can be completed by combining multiview stereo and mesh deformation techniques. A variety of challenging examples demonstrate the effectiveness of the proposed approach. Kangkan Wang, Guofeng Zhang 0001, Hujun Bao |
IEEE Trans. Image Process. | 1 |