Chi Xu 0002

dblp:34/117-2 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0002-5301-9376ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose Estimation
abstract
Estimating the 3D poses of hands and objects from a single RGB image is a fundamental yet challenging problem, with broad applications in augmented reality and human-computer interaction. Existing methods largely rely on visual cues alone, often producing results that violate physical constraints such as interpenetration or non-contact. Recent efforts to incorporate physics reasoning typically depend on post-optimization or non-differentiable physics engines, which compromise visual consistency and end-to-end trainability. To overcome these limitations, we propose a novel framework that jointly integrates visual and physical cues for hand-object pose estimation. This integration is achieved through two key ideas: 1) joint visual-physical cue learning: The model is trained to extract 2D visual cues and 3D physical cues, thereby enabling more comprehensive representation learning for hand-object interactions; 2) candidate pose aggregation: A novel refinement process that aggregates multiple diffusion-generated candidate poses by leveraging both visual and physical predictions, yielding a final estimate that is visually consistent and physically plausible. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches in both pose accuracy and physical plausibility.
Jun Zhou 0026, Chi Xu 0002, Kaifeng Tang, Yuting Ge, Tingrui Guo, Li Cheng 0001
AAAI2
2025 A Coarse-to-Fine Multi-Hypothesis Method for Ambiguous Hand Pose Estimation
abstract
In hand pose estimation, challenges such as occlusion often result in partial observation of a human hand, making it difficult to uniquely determine the hand pose, thus leading to ambiguity in certain hand regions. Heatmap-based methods may struggle with locating ambiguous joints and end up violating physiological constraints in their predictions. Parametric model based single-solution methods often fail to adequately address this ambiguity issue due to the inherent one-to-many mappings between input and output, resulting in unstable regression. While some existing multi-hypothesis methods have improved diversity by directly modeling the distribution of ambiguous hypotheses, their localization accuracy still falls short compared to the recent single-solution methods. To achieve quality results in both diversity and accuracy, we propose a novel multi-hypothesis approach for hand pose estimation, by progressively integrating heatmap information into the distribution of ambiguous poses using a RANSAC-like strategy. It starts with a conditional-flow model to provide an initial estimate of a coarse distribution over ambiguous joint poses. This is followed by randomly sampling multiple hypotheses, projecting each of them onto 2D heatmap plane, and employing consensus checks to identify unambiguous joints that adhere to skeletal constraints. Joint features are then resampled, with mismatches due to incorrect estimations being eliminated. Finally, we refine the distribution of ambiguous poses using graph neural networks and attention mechanisms. Extensive empirical experiments are carried out, where our approach are carefully examined both qualitatively and quantitatively. It is shown to not only produce more diverse & feasible pose hypotheses than existing multi-hypothesis methods, but also achieves accurate localization results comparable to the state-of-the-art single-solution methods.
Yuting Ge, Chi Xu 0002, Li Cheng 0001
IEEE Trans. Image Process.2
2025 A Hough Voting-Based 2-Point RANSAC Solution to the Perspective-n-Point Problem
abstract
Perspective- $n$ -point is a fundamental problem in multi-view geometry, yet two critical challenges persist: 1) The issues of high outlier rate and near degenerate cases exert a substantial impact on the robustness of existing P $n$ P methods. In the worst-case where both issues are in presence, existing methods tend to either produce erroneous results or become computationally prohibitive. 2) Conventionally, the hypothetical pose with the maximum inlier-set is assumed to be correct. However, it remains unclear whether this assumption holds when the outlier rate approaches ultra-high levels, and along this line what is the maximum amount of outliers that can be robustly handled. To address these challenges, this paper proposes a novel Hough voting based 2-point RANSAC solution. To our knowledge, it is the first P $n$ P solution capable of accurately and efficiently handling high outlier rates in near-degenerate cases. Extensive empirical evaluations have been conducted using the proposed approach, with a particular focus on a systematic examination under ultra-high outlier rates. The results show that, on random synthetic data, our approach works robustly even when dealing with up to 99% outliers. Meanwhile on real-world datasets, the maximum inlier-set assumption oftentimes fails when the outlier rate exceeds 97%, as the incorrect hypothetical poses may yield more inliers than the ground-truths. Our dataset and source code are to be made available at https://github.com/xuchi7/RPnP_plusplus.
Chi Xu 0002, Tingrui Guo, Li Cheng 0001
IEEE Trans. Image Process.1
2025 Hand Gesture Recognition From an Open-Set Perspective
abstract
Existing hand gesture recognition methods predominantly rely on a close-set assumption, which in essence limits the viewpoints, gesture categories, and hand shapes at test time to closely resemble those seen during training. This requirement is however rarely met in practice, as images are often captured from unconstrained viewpoints, with novel gestures and unseen hand shapes that can differ significantly from the training data. This motivates us to investigate an open-set hand gesture recognition problem, where hand gestures are still recognizable from unconstrained viewpoints, and novel gesture classes and hand shapes can be incrementally learned with just a few examples. To address this, we propose a viewpoint influence elimination network that extracts view-independent features, significantly improving performance in scenarios with unconstrained viewpoints. Moreover, a joint-weighted classification scheme is introduced to augment the cosine similarity metric for evaluating few-shot incremental learning of novel gestures and shapes. Finally, as existing hand gesture recognition datasets primarily adhere to the close-set assumption, a new hand gesture recognition dataset, OHG, is introduced in this paper, that includes a wide range of viewpoints, diverse gesture classes, and distinct hand shapes. Experimental hand gesture recognition results demonstrate the superior performance of our approach in both unconstrained viewpoint and few-shot incremental learning scenarios.
Jun Zhou 0026, Chi Xu 0002, Li Cheng 0001
IEEE Trans. Multim.2
2024 Bi-directional attention based RGB-D fusion for category-level object pose and shape estimation
Kaifeng Tang, Chi Xu 0002
Multim. Tools Appl.2
2024 Realistic Depth Image Synthesis for 3D Hand Pose Estimation
abstract
The training of depth image-based hand pose estimation model typically relies on real-life datasets which are expected to be 1) largescale and cover a diverse range of hand poses and hand shapes, and 2) always come with high-precision annotations. However, existing datasets in reality are rather limited in the above regards due to multitude practical constraints, with time and cost being the major concerns. This observation motivates us to propose an alternative approach, where hand pose model is primarily trained with synthesized hand depth images that closely mimicking the characteristic noise patterns of a specific depth camera make under consideration. It is achieved by firstly mapping a Gaussian distributed variable to certain specific non-i.i.d. (independent and identically distributed) depth noise pattern, and then transforming a “vanilla” noise-free synthetic depth image to a realistic-looking image. Extensive empirical experiments demonstrate that our approach is capable of generating camera-specific realistic-looking hand depth images with precise annotations; comparing to entirely relying on annotated real images, a hand pose model with better performance is obtained by using only a small fraction (10%) of annotated real images as well as our synthesized images.
Jun Zhou 0026, Chi Xu 0002, Yuting Ge, Li Cheng 0001
IEEE Trans. Multim.2
2020 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001
ECCV (14)5
2018 Correction to: Lie-X: Depth Image Based Articulated Object Pose Estimation, Tracking, and Action Recognition on Lie Groups
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Yu Zhang 0004, Zoe Bichler, Suresh Jesuthasan, Adam Claridge-Chang, Ajay Sriram Mathuru, Wenlong Tang, Peixin Zhu, Li Cheng 0001
Int. J. Comput. Vis.1
2017 Lie-X: Depth Image Based Articulated Object Pose Estimation, Tracking, and Action Recognition on Lie Groups
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Yu Zhang 0004, Li Cheng 0001
Int. J. Comput. Vis.1
2017 Pose Estimation from Line Correspondences: A Complete Analysis and a Series of Solutions
abstract
In this paper we deal with the camera pose estimation problem from a set of 2D/3D line correspondences, which is also known as PnL (Perspective-n-Line) problem. We carry out our study by comparing PnL with the well-studied PnP (Perspective-n-Point) problem, and our contributions are three-fold: (1) We provide a complete 3D configuration analysis for P3L, which includes the well-known P3P problem as well as several existing analyses as special cases. (2) By exploring the similarity between PnL and PnP, we propose a new subset-based PnL approach as well as a series of linear-formulation-based PnL approaches inspired by their PnP counterparts. (3) The proposed linear-formulation-based methods can be easily extended to deal with the line and point features simultaneously.
Chi Xu 0002, Lilian Zhang, Li Cheng 0001, Reinhard Koch
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Hand action detection from ego-centric depth sequences with error-correcting Hough transform
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Li Cheng 0001
Pattern Recognit.1
2016 Estimate Hand Poses Efficiently from Single Depth Images
abstract
This paper aims to tackle the practically very challenging problem of efficient and accurate hand pose estimation from single depth images. A dedicated two-step regression forest pipeline is proposed: given an input hand depth image, step one involves mainly estimation of 3D location and in-plane rotation of the hand using a pixel-wise regression forest. This is utilized in step two which delivers final hand estimation by a similar regression forest model based on the entire hand image patch. Moreover, our estimation is guided by internally executing a 3D hand kinematic chain model. For an unseen test image, the kinematic model parameters are estimated by a proposed dynamically weighted scheme. As a combined effect of these proposed building blocks, our approach is able to deliver more precise estimation of hand poses. In practice, our approach works at 15.6 frame-per-second (FPS) on an average laptop when implemented in CPU, which is further sped-up to 67.2 FPS when running on GPU. In addition, we introduce and make publicly available a data-glove annotated depth image dataset covering various hand shapes and gestures, which enables us conducting quantitative analyses on real-world hand images. The effectiveness of our approach is verified empirically on both synthetic and the annotated real-world datasets for hand pose estimation, as well as related applications including part-based labeling and gesture classification. In addition to empirical studies, the consistency property of our approach is also theoretically analyzed.
Chi Xu 0002, Ashwin Nanjappa, Xiaowei Zhang 0002, Li Cheng 0001
Int. J. Comput. Vis.1
2013 Efficient Hand Pose Estimation from a Single Depth Image
abstract
We tackle the practical problem of hand pose estimation from a single noisy depth image. A dedicated three-step pipeline is proposed: Initial estimation step provides an initial estimation of the hand in-plane orientation and 3D location, Candidate generation step produces a set of 3D pose candidate from the Hough voting space with the help of the rotational invariant depth features, Verification step delivers the final 3D hand pose as the solution to an optimization problem. We analyze the depth noises, and suggest tips to minimize their negative impacts on the overall performance. Our approach is able to work with Kinect-type noisy depth images, and reliably produces pose estimations of general motions efficiently (12 frames per second). Extensive experiments are conducted to qualitatively and quantitatively evaluate the performance with respect to the state-of-the-art methods that have access to additional RGB images. Our approach is shown to deliver on par or even better results.
Chi Xu 0002, Li Cheng 0001
ICCV1
2012 Robust and Efficient Pose Estimation from Line Correspondences
Lilian Zhang, Chi Xu 0002, Kok-Meng Lee, Reinhard Koch
ACCV (3)2
2012 A Robust O(n) Solution to the Perspective-n-Point Problem
abstract
We propose a noniterative solution for the Perspective-n-Point ({\rm P}n{\rm P}) problem, which can robustly retrieve the optimum by solving a seventh order polynomial. The central idea consists of three steps: 1) to divide the reference points into 3-point subsets in order to achieve a series of fourth order polynomials, 2) to compute the sum of the square of the polynomials so as to form a cost function, and 3) to find the roots of the derivative of the cost function in order to determine the optimum. The advantages of the proposed method are as follows: First, it can stably deal with the planar case, ordinary 3D case, and quasi-singular case, and it is as accurate as the state-of-the-art iterative algorithms with much less computational time. Second, it is the first noniterative {\rm P}n{\rm P} solution that can achieve more accurate results than the iterative algorithms when no redundant reference points can be used (n\le 5). Third, large-size point sets can be handled efficiently because its computational complexity is O(n).
Shiqi Li 0002, Chi Xu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2