Zhenhua Tang 0001

dblp:65/349-1 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-9189-4983ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 SNS-Grasp: Semantic-guided Noise Scaling for Grasp Generation
abstract
While diffusion models show promise for intent-based grasp generation, their isotropic noise schedules struggle with joint-specific sensitivity and task-aware variability. This limitation leads to grasps with suboptimal semantic alignment or physical feasibility. To address this challenge, we propose Semantic-guided Noise Scaling for grasp generation (SNS-Grasp), a novel framework that integrates two key innovations. First, the Semantic-guided Noise Scaling Diffusion (SNS-Diff) module generates intent-aware grasps by replacing isotropic noise with anisotropic modulation, dynamically adapting to task semantics and joint-specific sensitivity. Specifically, SNS-Diff leverages a pretrained Intent Recognizer to extract task-aware confidence scores and joint-specific gradient sensitivities from the interaction context. These signals adjust the noise scaling during denoising, downweighting perturbations for semantically critical joints to ensure semantic alignment. Second, the Fine-grained Grasp Refinement (FGR) module establishes dynamic joint-vertex coupling through fine-grained hand-object spatial relationships, enabling iterative optimization of physically executable grasps. Extensive experiments on OakInk and GRAB demonstrate SNS-Grasp's superior performance in semantic accuracy and physical feasibility, with robust generalization to unseen objects.
Zhenhua Tang 0001, Yudian Zheng, Yuzhang Zhong, Haolun Li 0001, Yanbin Hao, Chi-Man Pun
AAAI1
2026 Human-Structure-Aware Token Position Embedding for Tokenized Pose Estimation
abstract
Tokenized pose estimation (TPE) has demonstrated remarkable performance in lightweight human pose estimation (HPE) models. However, existing TPE methods typically initialize keypoint tokens randomly, without explicitly incorporating human structure priors. These priors play a vital role in HPE by effectively mitigating common challenges such as occlusion and ambiguity. To this end, we propose a Structure-Aware Keypoint Position Embedding (SAKPE). This embedding explicitly encodes inherent structural properties of the human body, such as symmetry and order, into the positional coordinates of keypoint tokens. It also employs learnable scale and offset factors to adapt to diverse human poses, thereby fully exploiting the geometric constraints among keypoints. Furthermore, to better leverage the positional relationships among patch tokens, we introduce a Layer-adaptive Hybrid Patch Position Embedding (LHPPE). It dynamically fuses absolute and relative position embeddings of patch tokens based on attention distributions across Transformer layers, enabling the model to learn both absolute and relative positional information adaptively. Taking the two together, we propose a novel position embedding method for pose estimation, named Human-structure-aware Token Position Embedding (HTPE). It significantly improves the performance of various TPE models. Extensive experiments on COCO, CrowdPose, and OCHuman show that HTPE achieves state-of-the-art (SOTA) performance among lightweight methods, with a negligible increase in parameters and FLOPs. Notably, it demonstrates consistent improvements under occlusion,, achieving up to 3.3 AP gains. The source code can be found in https://github.com/guzejungithub/HTPE.
Zejun Gu, Zhong-Qiu Zhao, Henghui Ding, Hao Shen 0006, Zhenhua Tang 0001, Zhao Zhang 0001, De-Shuang Huang
IEEE Trans. Image Process.5
2025 RAGG: Retrieval-Augmented Grasp Generation Model
abstract
Intent-based grasp generation inherently involves challenges such as manipulation ambiguity and modality gaps. To address these, we propose a novel Retrieval-Augmented Grasp Generation model (RAGG). Our key insight is that when humans manipulate new objects, they initially mimic the interaction patterns observed in similar objects, then progressively adjust hand-object contact. Consequently, we develop RAGG as a two-stage approach, encompassing retrieval-guided generation and structurally stable grasp refinement. In the first stage, we propose a Retrieval-Augmented Diffusion Model (ReDim), which identifies the most relevant interaction instance from a knowledge base to explicitly guide grasp generation, thereby mitigating ambiguity and bridging modality gaps to ensure semantically correct manipulation. In the second stage, we introduce a Progressive Refinement Network (PRN) with Kolmogorov-Arnold Network (KAN) layers to refine the generated coarse grasp, employing a Structural Similarity Index loss to constrain the spatial relationship between the hand and the object, thus ensuring the stability of the grasp. Extensive experiments on the OakInk and GRAB benchmarks demonstrate that RAGG achieves superior results compared to state-of-the-art approach, indicating not only better physical feasibility and controllability but also strong generalization and interpretability for unseen objects.
Zhenhua Tang 0001, Bin Zhu 0006, Yanbin Hao, Chong-Wah Ngo, Richang Hong
AAAI1
2025 MotionRefineNet: Fine-Grained Pose Sequence Smoothing and Refinement
abstract
Capturing human motion with existing monocular estimators often results in large errors when dealing with rare poses, occlusions, truncations, and frame blurring, leading to jitter and long-term drift. Although previous methods have introduced post-processing networks for pose refinement, they struggle to balance global smoothing and fine-grained correction. In this work, we propose MotionRefineNet, which leverages the synergy and complementarity between long- and short-term features in the temporal domain and high- and low-frequency features in the frequency domain to address these challenges. The temporal branch is designed as a hierarchical motion structure to learn multi-time scale features, where long-term features learn motion smoothness, and short-term features capture local rapid changes. The frequency branch employs different frequency band learning strategies based on the degrees of freedom (DoF) of body parts. For body parts with low DoF, the focus is on low-frequency features that represent overall motion trends and regular actions. For body parts with high DoF, we design a filter to adaptively extract useful information from all frequency bands, including subtle motion changes in the high-frequency bands. Extensive experiments on multiple datasets and estimators demonstrate that MotionRefineNet outperforms existing methods in refining 2D, 3D, and SMPL poses, achieving superior pose smoothing and deviation correction. Our code is available at: https://github.com/Wheels319/MotionRefineNet.
Haolun Li 0001, Weihuang Liu, Jiateng Liu, Zhenhua Tang 0001, Chi-Man Pun, Qiguang Miao, Feng Xu 0005, Hao Gao 0005
ACM Multimedia4
2024 JPA: A Joint-Part Attention for Mitigating Overfocusing on 3D Human Pose Estimation
Dengqing Yang, Zhenhua Tang 0001, Jinmeng Wu, Shuo Wang 0008, Lechao Cheng, Yanbin Hao
PRCV (6)2
2024 Space-View Decoupled 3D Gaussians for Novel-View Synthesis of Mirror Reflections
Zhenwu Wang, Zhuopeng Li, Zhenhua Tang 0001, Yanbin Hao, Huasen He
PRICAI (4)3
2024 FTCM: Frequency-Temporal Collaborative Module for Efficient 3D Human Pose Estimation in Video
abstract
Capturing cross-pose correlation from a sequence of frame-level 2D poses is essential for 3D human pose estimation (3D-HPE) in the video. Recent studies have shown the promising potential of modeling the pose relation with feature-mixing operations on the temporal domain. However, they seldom consider the interaction across poses in the frequency domain. This paper studies a Frequency-Temporal Collaborative Module (FTCM) to explore the feasibility of encoding the cross-pose correlations in both frequency and temporal domains. FTCM aims to jointly capture the global and local cross-pose correlations with a more lightweight network model. Specifically, FTCM splits the pose features into two groups along the channel dimension and separately models the frequency and temporal interactions across poses with different feature-mixing operations in parallel. To achieve this goal, we purposely design two pose-mixing units, i.e., the frequency pose-mixing (FPM) and the temporal pose-mixing (TPM). Particularly, FPM is designed to reap the global correlations among different pose frequencies with the representation obtained by converting the original pose signals with Fast Fourier transform (FFT). Unlike the pose-mixing used by previous methods like Transformers that influences an individual pose with all other poses, TPM locally calibrates the pose with dynamics aggregated within several adjacent poses in the temporal domain, explicitly weighting neighboring poses more with respect to the far-away ones so as to enforce a strict locality constraint. Besides, the group strategy significantly reduces the model complexity. To verify the effectiveness of FTCM, we conduct extensive experiments on two benchmarks (i.e., Human3.6M and MPI-INF-3DHP). Experimental results not only exhibit favorable accuracy/complexity trade-offs of our FTCM but also show superior or comparable performance to state-of-the-art methods on both datasets. The code and model are publicly available at:https://github.com/zhenhuat/FTCM.
Zhenhua Tang 0001, Yanbin Hao, Jia Li 0013, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.1
2023 3D Human Pose Estimation with Spatio-Temporal Criss-Cross Attention
abstract
Recent transformer-based solutions have shown great success in 3D human pose estimation. Nevertheless, to calculate the joint-to-joint affinity matrix, the computational cost has a quadratic growth with the increasing number of joints. Such drawback becomes even worse especially for pose estimation in a video sequence, which necessitates spatio-temporal correlation spanning over the entire video. In this paper, we facilitate the issue by decomposing correlation learning into space and time, and present a novel Spatio-Temporal Criss-cross attention (STC) block. Technically, STC first slices its input feature into two partitions evenly along the channel dimension, followed by performing spatial and temporal attention respectively on each partition. STC then models the interactions between joints in an identical frame and joints in an identical trajectory simultaneously by concatenating the outputs from attention layers. On this basis, we devise STCFormer by stacking multiple STC blocks and further integrate a new Structure-enhanced Positional Embedding (SPE) into STCFormer to take the structure of human body into consideration. The embedding function consists of two components: spatio-temporal convolution around neighboring joints to capture local structure, and part-aware embedding to indicate which part each joint belongs to. Extensive experiments are conducted on Human3.6M and MPI-INF-3DHP benchmarks, and superior results are reported when comparing to the state-of-the-art approaches. More remarkably, STCFormer achieves to-date the best published performance: 40.5mm P1 error on the challenging Human3.6M dataset.
Zhenhua Tang 0001, Zhaofan Qiu, Yanbin Hao, Richang Hong, Ting Yao 0003
CVPR1
2023 RepEPnP: Weakly Supervised 3D Human Pose Estimation with EPnP Algorithm
abstract
This paper describes an end-to-end weakly super-vised framework for estimating 3D human pose from a single image. The model is trained by projecting 3D pose to 2D pose for matching ground-truth 2D pose for supervision. To obtain accu-rate projection from 3D pose to 2D pose, a mathematical camera model based on intrinsic and extrinsic camera parameters is used. Specifically, we use EPnP algorithm to estimate extrinsic transformation matrix to transform the estimated 3D pose to be reprojected back to 2D pose. The advantage of this projection is that it requires no training and it is robust to the diversity of training datasets. We further constrain the pose generation using an adversarial generative network, where a transformer is used as the 3D pose generator. Transformer can use self-attention mechanism to establish dependencies between each joint and predict pose based on important joints. Based on our reprojection method, our method achieves competitive results on Human3.6M and MPI-INF-3DHP among weakly supervised methods. The experiments also demonstrate our model's generalization ability for wild images.
Huaijing Lai, Zhenhua Tang 0001, Xiaoyan Zhang 0002
IJCNN2
2023 MLP-JCG: Multi-Layer Perceptron With Joint-Coordinate Gating for Efficient 3D Human Pose Estimation
abstract
Various structural relations/dependencies exist among human body joints, which makes it possible to estimate 3D poses from 2D sources. The current research on 3D human pose estimation (3D-HPE for short) mainly focuses on structural information from a specific perspective. However, this information cannot facilitate 2D-to-3D pose lifting. This paper presents a novel and efficient multi-layer perceptron with a joint-coordinate gating (MLP-JCG) model, exploring and utilizing both the local and global structural information to perform 3D pose estimations. Specifically, MLP-JCG contains two independent MLP blocks, i.e., joint-mixing MLP and coordinate-mixing MLP, which solely act on the joint and coordinate axes in modelling their local structural information. For the global structural information, we first explore two kinds of global statistics from the pose matrix embeddings, which are referred to as the dynamics aggregated along the joint/coordinate axis. Then, we propose two kinds of gating units to elementwisely contextualize the features learned from MLP blocks. All the model components are designed based on MLP, making the MLP-JCG easy to implement and train. We conduct experiments on three 3D-HPE benchmarks, and the results demonstrate the superior effectiveness and efficiency of the proposed approach.
Zhenhua Tang 0001, Jia Li 0013, Yanbin Hao, Richang Hong
IEEE Trans. Multim.1
2019 An Articulated Structure-aware Network for 3D Human Pose Estimation
abstract
In this paper, we propose a new end-to-end articulated structure-aware network to regress 3D joint coordinates from the given 2D joint detections. The proposed method is capable of dealing with hard joints well that usually fail existing methods. Specifically, our framework cascades a refinement network with a basic network for two types of joints, and employs a attention module to simulate a camera projection model. In addition, we propose to use a random enhancement module to intensify the constraints between joints. Experimental results on the Human3.6M and HumanEva databases demonstrate the effectiveness and flexibility of the proposed network, and errors of hard joints and bone lengths are significantly reduced, compared with state-of-the-art approaches.
Zhenhua Tang 0001, Xiaoyan Zhang 0002, Junhui Hou
ACML1
2019 3D human pose estimation via human structure-aware fully connected network
Xiaoyan Zhang 0002, Zhenhua Tang 0001, Junhui Hou, Yanbin Hao
Pattern Recognit. Lett.2
2019 Pose-Based Composition Improvement for Portrait Photographs
abstract
This paper studies the composition in portrait paintings and develops an algorithm to improve the composition of portrait photographs based on example portrait paintings. A study of portrait paintings shows that the placement of the face and the figure is pose-related. Based on this observation, this paper develops an algorithm to improve the composition of a portrait photograph by learning the placement of the face and the figure from an example portrait painting. This example portrait painting is selected based on the similarity of its figure pose to that of the input photograph. This similarity measure is modeled as a graph matching problem. Finally, space cropping is performed using an optimization function to assign a similar location for each body part of the figure in the photograph with that of the figure in the example portrait painting. The experimental results demonstrate the effectiveness of the proposed method. A user study shows that the proposed pose-based composition improvement is preferred more than rule-based methods and learning-based methods.
Xiaoyan Zhang 0002, Zhuopeng Li, Martin Constable, Kap Luk Chan, Zhenhua Tang 0001, Gaoyang Tang
IEEE Trans. Circuits Syst. Video Technol.5