Chen Zhao 0025

dblp:81/3-25 · DBLP profile ↗
← Back
17ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0001-9843-6758ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 8 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Monte Carlo Diffusion for Generalizable Learning-Based RANSAC
abstract
Random Sample Consensus (RANSAC) is a fundamental approach for robustly estimating parametric models from noisy data. Existing learning-based RANSAC methods utilize deep learning to enhance the robustness of RANSAC against outliers. However, these approaches are trained and tested on the data generated by the same algorithms, leading to limited generalization to out-of-distribution data during inference. Therefore, in this paper, we introduce a novel diffusion-based paradigm that progressively injects noise into ground-truth data, simulating the noisy conditions for training learning-based RANSAC. To enhance data diversity, we incorporate Monte Carlo sampling into the diffusion paradigm, approximating diverse data distributions by introducing different types of randomness at multiple stages. We evaluate our approach in the context of feature matching through comprehensive experiments on the ScanNet and MegaDepth datasets. The experimental results demonstrate that our Monte Carlo diffusion mechanism significantly improves the generalization ability of learning-based RANSAC. We also develop extensive ablation studies that highlight the effectiveness of key components in our framework.
Chen Zhao 0025, Wei Ke 0003, Tong Zhang 0023
AAAI2
2026 Pose without Guesses: Generalizable Object Pose Estimation From a Single Reference
Chen Zhao 0025, Tong Zhang 0023, Zheng Dang, Mathieu Salzmann
Int. J. Comput. Vis.1
2025 BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation
Yuanhong Yu 0003, Chen Zhao 0025, Junhao Yu, Jiaqi Yang 0002, Ruizhen Hu, Yujun Shen, Xiaowei Zhou 0001, Sida Peng
ICCV3
2025 Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesis
abstract
3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its generalization to novel viewpoints. In this paper, we address this overfitting issue by introducing Self-Ensembling Gaussian Splatting (SE-GS). We achieve self-ensembling by incorporating an uncertainty-aware perturbation strategy during training. A $\mathbfΔ$-model and a $\mathbfΣ$-model are jointly trained on the available images. The $\mathbfΔ$-model is dynamically perturbed based on rendering uncertainty across training steps, generating diverse perturbed models with negligible computational overhead. Discrepancies between the $\mathbfΣ$-model and these perturbed models are minimized throughout training, forming a robust ensemble of 3DGS models. This ensemble, represented by the $\mathbfΣ$-model, is then used to generate novel-view images during inference. Experimental results on the LLFF, Mip-NeRF360, DTU, and MVImgNet datasets demonstrate that our approach enhances NVS quality under few-shot training conditions, outperforming existing state-of-the-art methods. The code is released at: https://sailor-z.github.io/projects/SEGS.html.
Chen Zhao 0025, Xuan Wang 0009, Tong Zhang 0023, Saqib Javed, Mathieu Salzmann
ICCV1
2024 Unsupervised 3D Keypoint Discovery with Multi-View Geometry
abstract
Analyzing and training 3D body posture models depend heavily on the availability of joint labels that are commonly acquired through laborious manual annotation of body joints or via marker-based joint localization using carefully curated markers and capturing systems. However, such annotations are not always available, especially for people performing unusual activities. In this paper, we propose an algorithm that learns to discover 3D keypoints on human bodies from multiple-view images without any supervision or labels other than the constraints multiple-view geometry provides. To ensure that the discovered 3D keypoints are meaningful, they are re-projected to each view to estimate the person’s mask that the model itself has initially estimated without supervision. Our approach discovers more interpretable and accurate 3D keypoints compared to other state-of-the-art unsupervised approaches on Human3.6M and MPI-INF-3DHP benchmark datasets.
Sina Honari, Chen Zhao 0025, Mathieu Salzmann, Pascal Fua
3DV2
2024 LocPoseNet: Robust Location Prior for Unseen Object Pose Estimation
abstract
Object location prior is critical for the standard 6D object pose estimation setting. The prior can be used to initialize the 3D object translation and facilitate 3D object rotation estimation. Unfortunately, the object detectors that are used for this purpose do not generalize to unseen objects. Therefore, existing 6D pose estimation methods for unseen objects either assume the ground-truth object location to be known or yield inaccurate results when it is unavailable. In this paper, we address this problem by developing a method, LocPoseNet, able to robustly learn location prior for unseen objects. Our method builds upon a template matching strategy, where we propose to distribute the reference kernels and convolve them with a query to efficiently compute multi-scale correlations. We then introduce a novel translation estimator, which decouples scale-aware and scale-robust features to predict different object location parameters. Our method outperforms existing works by a large margin on LINEMOD and GenMOP. We further construct a challenging synthetic dataset, which allows us to highlight the better robustness of our method to various noise sources. Our project website is at: https://sailorz.github.io/projects/3DV2024_LocPoseNet.html.
Chen Zhao 0025, Yinlin Hu, Mathieu Salzmann
3DV1
2024 HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields
abstract
Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus, existing methods often rely on intermediate 3D shape representations to increase performance. These representations are typically explicit, such as 3D point clouds or meshes, and thus provide information in the direct surroundings of the intermediate hand pose estimate. To address this, we in-troduce HOISDF, a Signed Distance Field (SDF) guided hand-object pose estimation network, which jointly exploits hand and object SDFs to provide a global, implicit repre-sentation over the complete reconstruction volume. Specif-ically, the role of the SDFs is threefold: equip the visual encoder with implicit shape information, help to encode hand-object interactions, and guide the hand and object pose regression via SDF-based sampling and by augmenting the feature representations. We show that HOISDF achieves state-of-the-art results on hand-object pose esti-mation benchmarks (DexYCB and H03Dv2). Code is avail-able at https://github.com/amathislabIHOISDF.
Haozhe Qi, Chen Zhao 0025, Mathieu Salzmann, Alexander Mathis
CVPR2
2024 DVMNet: Computing Relative Pose for Unseen Objects Beyond Hypotheses
abstract
Determining the relative pose of an object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically approximate the continuous pose representation with a large number of discrete pose hypotheses, which incurs a computationally expensive process of scoring each hypothesis at test time. By contrast, we present a Deep Voxel Matching Network (DVMNet) that eliminates the need for pose hypotheses and computes the relative object pose in a single pass. To this end, we map the two input RGB images, reference and query, to their respective voxelized 3D representations. We then pass the resulting voxels through a pose estimation module, where the voxels are aligned and the pose is computed in an end-to-end fashion by solving a least-squares problem. To enhance robustness, we introduce a weighted closest voxel algorithm capable of mitigating the impact of noisy voxels. We conduct extensive experiments on the CO3D, LINEMOD, and Objaverse datasets, demonstrating that our method delivers more accurate relative pose estimates for novel objects at a lower computational cost compared to state-of-the-art methods. Our code is released at: https://github.com/sailor-z/DVMNet/.
Chen Zhao 0025, Tong Zhang 0023, Zheng Dang, Mathieu Salzmann
CVPR1
2024 3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation
abstract
Prior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where only a single reference view of the object is available. Our goal then is to estimate the relative object pose between this reference view and a query image that depicts the object in a different pose. In this scenario, robust generalization is imperative due to the presence of unseen objects during testing and the large-scale object pose variation between the reference and the query. To this end, we present a new hypothesis-and-verification framework, in which we generate and evaluate multiple pose hypotheses, ultimately selecting the most reliable one as the relative object pose. To measure reliability, we introduce a 3D-aware verification that explicitly applies 3D transformations to the 3D object representations learned from the two input images. Our comprehensive experiments on the Objaverse, LINEMOD, and CO3D datasets evidence the superior accuracy of our approach in relative pose estimation and its robustness in large-scale pose variations, when dealing with unseen objects.
Chen Zhao 0025, Tong Zhang 0023, Mathieu Salzmann
ICLR1
2022 Unsupervised Learning of 3D Semantic Keypoints with Mutual Reconstruction
Haocheng Yuan, Chen Zhao 0025, Shichao Fan, Jiaxi Jiang, Jiaqi Yang 0002
ECCV (2)2
2022 Fusing Local Similarities for Retrieval-Based 3D Orientation Estimation of Unseen Objects
Chen Zhao 0025, Yinlin Hu, Mathieu Salzmann
ECCV (1)1
2022 Rotation invariant point cloud analysis: Where local geometry meets global topology
Chen Zhao 0025, Jiaqi Yang 0002, Angfan Zhu, Zhiguo Cao 0001, Xin Li 0005
Pattern Recognit.1
2021 Progressive Correspondence Pruning by Consensus Learning
abstract
Correspondence pruning aims to correctly remove false matches (outliers) from an initial set of putative correspondences. The pruning process is challenging since putative matches are typically extremely unbalanced, largely dominated by outliers, and the random distribution of such outliers further complicates the learning process for learning-based methods. To address this issue, we propose to progressively prune the correspondences via a local-to-global consensus learning procedure. We introduce a "pruning" block that lets us identify reliable candidates among the initial matches according to consensus scores estimated using local-to-global dynamic graphs. We then achieve progressive pruning by stacking multiple pruning blocks sequentially. Our method outperforms state-of-the-arts on robust line fitting, camera pose estimation and retrieval-based image localization benchmarks by significant margins and shows promising generalization ability to different datasets and detector/descriptor combinations.
Chen Zhao 0025, Yixiao Ge, Feng Zhu 0006, Rui Zhao 0001, Hongsheng Li 0001, Mathieu Salzmann
ICCV1
2020 Sparse-to-Dense Depth Completion Revisited: Sampling Strategy and Graph Construction
Haipeng Xiong, Ke Xian, Chen Zhao 0025, Zhiguo Cao 0001, Xin Li 0005
ECCV (21)4
2020 Image Feature Correspondence Selection: A Comparative Study and a New Contribution
abstract
Image feature correspondence selection is pivotal to many computer vision tasks from object recognition to 3D reconstruction. Although many correspondence selection algorithms have been developed in the past decade, there still lacks an in-depth evaluation and comparison in the open literature, which makes it difficult to choose the appropriate algorithm for a specific application. This paper attempts to fill this gap by evaluating eight competing correspondence selection algorithms including both classical methods and current state-of-the-art ones. In addition to preselected correspondences, we have compared different combinations of detector and descriptor on four standard datasets. The diversity of those datasets cover a wide range of uncertainty factors including zoom, rotation, blur, viewpoint change, JPEG compression, light change, different rendering styles and multiple structures. We have measured the quality of competing correspondence selection algorithms in terms of four performance metrics -i.e., precision, recall, F-measure and efficiency. Moreover, we propose to combine the strengths of eight competing methods by combining their correspondence selection results. Extensive experimental results are reported to demonstrate the superiority of several fusion strategies to individual methods, which suggests the possibility of adaptively combining those methods for even better performance.
Chen Zhao 0025, Zhiguo Cao 0001, Jiaqi Yang 0002, Ke Xian, Xin Li 0005
IEEE Trans. Image Process.1
2019 NM-Net: Mining Reliable Neighbors for Robust Feature Correspondences
abstract
Feature correspondence selection is pivotal to many feature-matching based tasks in computer vision. Searching spatially k-nearest neighbors is a common strategy for extracting local information in many previous works. However, there is no guarantee that the spatially k-nearest neighbors of correspondences are consistent because the spatial distribution of false correspondences is often irregular. To address this issue, we present a compatibility-specific mining method to search for consistent neighbors. Moreover, in order to extract and aggregate more reliable features from neighbors, we propose a hierarchical network named NM-Net with a series of graph convolutions that is insensitive to the order of correspondences. Our experimental results have shown the proposed method achieves the state-of-the-art performance on four datasets with various inlier ratios and varying numbers of feature consistencies.
Chen Zhao 0025, Zhiguo Cao 0001, Xin Li 0005, Jiaqi Yang 0002
CVPR1
2018 Scalable Multi-Consistency Feature Matching with Non-Cooperative Games
abstract
Correspondence selection aiming at seeking correct relationships between two images is a fundamental and critical task in computer vision. This paper attempts to select consistent correspondences in the context of dynamic scenarios where multiple matching consistencies are normally incorporated. To this end, we present a grid-based game-theoretic matching (Grid-GTM) method which is divided into three processes, i.e., grid matching, local games and enrichment. Specifically, grid matching translates the multi-consistency problem into several independent single-consistency problems to decrease difficulties of selection and boost the efficiency. Local games extended under the guidance of a novel payoff function guarantee that mismatches are effectively removed. Enrichment is added to recover correct matches neglected by local games. Crucially, our approach achieves the state-of-the-art performance compared with seven algorithms in comprehensive evaluations. In addition, we construct a dataset that involves multiple consistencies under three different scenes in this paper.
Chen Zhao 0025, Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001
ICIP1