Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jinghong Zheng 0002

dblp:54/5344-2 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0000-7996-8927ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 80% Generative modeling · 10% Face, body and person analysis · 6%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth estimation
2.332025
CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching · CVPR 2025
Self-Distilled Depth Refinement with Noisy Poisson Fusion · NeurIPS 2024
Diffusion-Augmented Depth Prediction with Sparse Annotations · ACM Multimedia 2023
Computer vision › 3D vision
3d human pose estimation
1.922026
3D Hand Pose Estimation via Articulated Anchor-to-Joint 3D Local Regressors · IEEE Trans. Pattern Anal. Mach. Intell. 2026
PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space · NeurIPS 2025
Computer vision › 3D vision › pose estimation
3d hand pose estimation
1.722026
3D Hand Pose Estimation via Articulated Anchor-to-Joint 3D Local Regressors · IEEE Trans. Pattern Anal. Mach. Intell. 2026
A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image · CVPR 2023
Computer vision › 3D vision › depth estimation
monocular depth estimation
1.522025
CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching · CVPR 2025
Diffusion-Augmented Depth Prediction with Sparse Annotations · ACM Multimedia 2023
Computer vision › 3D vision › pose estimation › 3d hand pose estimation
monocular 3d hand pose estimation
1.012026
3D Hand Pose Estimation via Articulated Anchor-to-Joint 3D Local Regressors · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision › 3d human pose estimation
2d-to-3d pose lifting
0.912025
PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space · NeurIPS 2025
Machine learning › Generative modeling
flow matching
0.912025
CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching · CVPR 2025
Computer vision › 3D vision › depth estimation
video depth estimation
0.912025
CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching · CVPR 2025
Computer vision › 3D vision › depth estimation
depth map refinement
0.812024
Self-Distilled Depth Refinement with Noisy Poisson Fusion · NeurIPS 2024
Image and video processing
image restoration
0.812024
Self-Distilled Depth Refinement with Noisy Poisson Fusion · NeurIPS 2024
Computer vision › 3D vision › depth estimation › deep depth estimation
diffusion-based depth estimation
0.712023
Diffusion-Augmented Depth Prediction with Sparse Annotations · ACM Multimedia 2023
Machine learning › Generative modeling
diffusion model
0.712023
Diffusion-Augmented Depth Prediction with Sparse Annotations · ACM Multimedia 2023
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.712023
A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image · CVPR 2023
Computer vision › 3D vision › pose estimation › 3d hand pose estimation
interacting hand pose estimation
0.712023
A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image · CVPR 2023
Computer vision › 3D vision › depth estimation › deep depth estimation
depth foundation model
0.312025
CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching · CVPR 2025
Computer vision › Face, body and person analysis
human pose estimation
0.312025
PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space · NeurIPS 2025
Computer vision › 3D vision › 3d human pose estimation
single-image 3d pose estimation
0.312025
PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation
0.212024
Self-Distilled Depth Refinement with Noisy Poisson Fusion · NeurIPS 2024
Robotics › Autonomous driving
perception
0.212023
Diffusion-Augmented Depth Prediction with Sparse Annotations · ACM Multimedia 2023
Machine learning › Deep learning architectures and training
transformer
0.212023
A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image · CVPR 2023

Methods — techniques the papers use, named apart from their topics

transformer · 1.7articulated anchor points · 1.0anchor-to-joint local regressors · 1.0non-uniform sampling · 0.9latent temporal stabilizer · 0.9inversion by direct iteration · 0.9graph neural network · 0.9flow matching · 0.9depth-aware feature lifting · 0.9anchor-feature interaction decoder · 0.9self-distillation · 0.8poisson fusion · 0.8edge-guided loss · 0.8
YearPublicationVenuePosition
2026 3D Hand Pose Estimation via Articulated Anchor-to-Joint 3D Local Regressors
abstract
In this paper, we propose to address monocular 3D hand pose estimation from a single RGB or depth image via articulated anchor-to-joint 3D local regressors, in form of A2J-Transformer+. The key idea is to make the local regressors (i.e., anchor points) in 3D space be aware of hand's local fine details and global articulated context jointly, to facilitate predicting their 3D offsets toward hand joints with linear weighted aggregation for joint localization. Our intuition is that, local fine details help to estimate accurate offset but may suffer from the issues including serious occlusion, confusing similar patterns, and overfitting risk. On the other hand, hand's global articulated context can essentially provide additional descriptive clues and constraints to alleviate these issues. To set anchor points adaptively in 3D space, A2J-Transformer+ runs in a 2-stage manner. At the first stage, since the input modality property anchor points distribute more densely on X-Y plane, it leads to lower prediction accuracy along Z direction compared with those in the X and Y directions. To alleviate this, at the second stage anchor points are set near the joints yielded by the first stage evenly along X, Y, and Z directions. This treatment brings two main advantages: (1) balancing the prediction accuracy along X, Y, and Z directions, and (2) ensuring the anchor-joint offsets are of small values relatively easy to estimate. Wide-range experiments on three RGB hand datasets (InterHand2.6 M, HO-3D V2 and RHP) and three depth hand datasets (NYU, ICVL and HANDS 2017) verify A2J-Transformer+'s superiority and generalization ability for different modalities (i.e., RGB and depth) and hand cases (i.e., single hand, interacting hands, and hand-object interaction), even outperforming model-based manners. The test on ITOP dataset reveals that, A2J-Transformer+ can also be applied to 3D human pose estimation task.
Changlong Jiang, Yang Xiao 0007, Jinghong Zheng 0002, Haohong Kuang, Cunlin Wu, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching
abstract
Depth estimation is a fundamental task in 3D vision. An ideal depth estimation model is expected to embrace meticulous detail, temporal consistency, and high efficiency. Although existing foundation models can perform well in certain specific aspects, most of them fall short of fulfilling all the above requirements simultaneously. In this paper, we present CH3Depth, an efficient and flexible model for depth estimation with flow matching to address this challenge. Specifically, 1) we reframe the optimization objective of flow matching as the Inversion by Direct Iteration (InDI) to improve accuracy. 2) To enhance efficiency, we propose non-uniform sampling to achieve better prediction with fewer sampling steps. 3) We design the Latent Temporal Stabilizer (LTS) to enhance temporal consistency by aggregating latent codes of adjacent frames, enabling our method to be lightweight and compatible for video depth estimation. CH3Depth achieves state-of-the-art performance in zero-shot evaluations across multiple image and video datasets, excelling in prediction accuracy, efficiency, and temporal consistency, highlighting its potential as the next foundation model for depth estimation.
Jiaqi Li 0007, Yiran Wang 0005, Jinghong Zheng 0002, Junrui Zhang 0003, Liao Shen, Tianqi Liu 0003, Zhiguo Cao 0001
CVPR3
2025 PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space
abstract
3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental limitations: inevitable error propagation from input predicted 2D pose to 3D predictions and inherent difficulties in handling self-occlusion cases. In this paper, we propose PandaPose, a 3D human pose lifting approach via propagating 2D pose prior to 3D anchor space as the unified intermediate representation. Specifically, our 3D anchor space comprises: (1) Joint-wise 3D anchors in the canonical coordinate system, providing accurate and robust priors to mitigate 2D pose estimation inaccuracies. (2) Depth-aware joint-wise feature lifting that hierarchically integrates depth information to resolve self-occlusion ambiguities. (3) The anchor-feature interaction decoder that incorporates 3D anchors with lifted features to generate unified anchor queries encapsulating joint-wise 3D anchor set, visual cues and geometric depth information. The anchor queries are further employed to facilitate anchor-to-joint ensemble prediction. Experiments on three well-established benchmarks (i.e., Human3.6M, MPI-INF-3DHP and 3DPW) demonstrate the superiority of our proposition. The substantial reduction in error by 14.7% compared to SOTA methods on the challenging conditions of Human3.6M and qualitative comparisons further showcase the effectiveness and robustness of our approach.
Jinghong Zheng 0002, Changlong Jiang, Yang Xiao 0007, Jiaqi Li 0007, Haohong Kuang, Ran Wang 0005, Zhiguo Cao 0001, Joey Tianyi Zhou
NeurIPS1
2025 UniPose: Unified Cross-Modality Pose Prior Propagation Towards RGB-D Data for Weakly Supervised 3D Human Pose Estimation
Jinghong Zheng 0002, Changlong Jiang, Jiaqi Li 0007, Haohong Kuang, Tingbing Yan
PRCV (7)1
2024 Self-Distilled Depth Refinement with Noisy Poisson Fusion
abstract
Depth refinement aims to infer high-resolution depth with fine-grained edges and details, refining low-resolution results of depth estimation models. The prevailing methods adopt tile-based manners by merging numerous patches, which lacks efficiency and produces inconsistency. Besides, prior arts suffer from fuzzy depth boundaries and limited generalizability. Analyzing the fundamental reasons for these limitations, we model depth refinement as a noisy Poisson fusion problem with local inconsistency and edge deformation noises. We propose the Self-distilled Depth Refinement (SDDR) framework to enforce robustness against the noises, which mainly consists of depth edge representation and edge-based guidance. With noisy depth predictions as input, SDDR generates low-noise depth edge representations as pseudo-labels by coarse-to-fine self-distillation. Edge-based guidance with edge-guided gradient loss and edge-based fusion loss serves as the optimization objective equivalent to Poisson fusion. When depth maps are better refined, the labels also become more noise-free. Our model can acquire strong robustness to the noises, achieving significant improvements in accuracy, edge quality, efficiency, and generalizability on five different benchmarks. Moreover, directly training another model with edge labels produced by SDDR brings improvements, suggesting that our method could help with training robust refinement models in future works.
Jiaqi Li 0007, Yiran Wang 0005, Jinghong Zheng 0002, Zihao Huang 0001, Ke Xian, Zhiguo Cao 0001, Jianming Zhang 0001
NeurIPS3
2023 A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image
abstract
3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-the state-of-the-art depth-based 3D single hand pose estimation method-to RGB domain under interacting hand condition. Our key idea is to equip A2J with strong local-global aware ability to well capture interacting hands' local fine details and global articulated clues among joints jointly. To this end, A2J is evolved under Transformer's non-local encoding-decoding framework to build A2J- Transformer. It holds 3 main advantages over A2J. First, self-attention across local anchor points is built to make them global spatial context aware to better capture joints' articulation clues for resisting occlusion. Secondly, each anchor point is regarded as learnable query with adaptive feature learning for facilitating pattern fitting capacity, instead of having the same local representation with the others. Last but not least, anchor point locates in 3D space instead of 2D as in A2J, to leverage 3D pose prediction. Experiments on challenging InterHand 2.6M demonstrate that, A2J-Transformer can achieve state-of-the-art model-free performance (3.38mm MPJPE advancement in 2-hand case) and can also be applied to depth domain with strong generalization. The code is avaliable at https://github.com/ChanglongJiangGit/A2J-Transformer.
Changlong Jiang, Yang Xiao 0007, Cunlin Wu, Jinghong Zheng 0002, Zhiguo Cao 0001, Joey Tianyi Zhou
CVPR5
2023 Diffusion-Augmented Depth Prediction with Sparse Annotations
abstract
Depth estimation aims to predict dense depth maps. In autonomous driving scenes, sparsity of annotations makes the task challenging. Supervised models produce concave objects due to insufficient structural information. They overfit to valid pixels and fail to restore spatial structures. Self-supervised methods are proposed for the problem. Their robustness is limited by pose estimation, leading to erroneous results in natural scenes. In this paper, we propose a supervised framework termed Diffusion-Augmented Depth Prediction (DADP). We leverage the structural characteristics of diffusion model to enforce depth structures of depth models in a plug-and-play manner. An object-guided integrality loss is also proposed to further enhance regional structure integrality by fetching objective information. We evaluate DADP on three driving benchmarks and achieve significant improvements in depth structures and robustness. Our work provides a new perspective on depth estimation with sparse annotations in autonomous driving scenes.
Jiaqi Li 0007, Yiran Wang 0005, Zihao Huang 0001, Jinghong Zheng 0002, Ke Xian, Zhiguo Cao 0001, Jianming Zhang 0001
ACM Multimedia4
2022 Reward Shaping-based Double Deep Q-networks for Unmanned Surface Vessel Navigation and Obstacle Avoidance
abstract
In this paper, a method for navigation and obstacle avoidance of unmanned surface vessel (USV) based on reinforcement learning and reward shaping is proposed. This approach uses double deep Q networks (DDQN) to make decisions based on the continuous states observed from sensors in USV. In addition, a new reward function is designed based on prior knowledge to accelerate the convergence of the algorithm and improve the performance. For training the neural networks, a simulation platform is developed, in which a 3 degree of freedom mathematical model describes USV dynamic system and two-dimension actions are required to control USV. Simulation results on the platform demonstrate the DDQN hoists USV’s capabilities of navigation and obstacle avoidance, and reward shaping technique improves the speed of convergence.
Zihan Gan, Jinghong Zheng 0002, Renzhi Lu
IECON2