Zhongqun Zhang

dblp:237/8882 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-0884-8711ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Force-Aware 3D Contact Modeling for Stable Grasp Generation
abstract
Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, specifically the contact force, leading to reduced stability of the grasp. In this paper, we focus on stable grasp generation using explicit contact force predictions. First, we define a force-aware contact representation by transforming the normal force value into discrete levels and encoding it using a one-hot vector. Next, we introduce force-aware stability constraints. We define the stability problem as an acceleration minimization task and explicitly relate stability with contact geometry by formulating the underlying physical constraints. Finally, we present a pose optimizer that systematically integrates our contact representation and stability constraints to enable stable grasp generation. We show that these constraints can help identify key contact points for stability which provide effective initialization and guidance for optimization towards a stable grasp. Experiments are carried out on two public benchmarks, showing that our method brings about 20% improvement in stability metrics and adapts well to novel objects.
Zhuo Chen 0028, Zhongqun Zhang, Yihua Cheng, Ales Leonardis, Hyung Jin Chang
AAAI2
2026 RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
abstract
Gaze redirection methods aim to generate realistic human face images with controllable eye movement. However, recent methods often struggle with 3D consistency, efficiency, or quality, limiting their practical applications. In this work, we propose RTGaze, a real-time and high-quality gaze redirection method. Our approach learns a gaze-controllable facial representation from face images and gaze prompts, then decodes this representation via neural rendering for gaze redirection. Additionally, we distill face geometric priors from a pretrained 3D portrait generator to enhance generation quality. We evaluate RTGaze both qualitatively and quantitatively, demonstrating state-of-the-art performance in efficiency, redirection accuracy, and image quality across multiple datasets. Our system achieves real-time, 3D-aware gaze redirection with a feedforward network (~0.06 sec/image), making it 800× faster than the previous state-of-the-art 3D-aware methods.
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
AAAI2
2025 Single-view Image to Novel-view Generation for Hand-Object Interactions
abstract
Hand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view hand-object interaction images from a single image. To this end, we first train a 2D diffusion prior. Given the camera pose in novel views, our approach transfers the camera information into explicit hand representations, including hand depth and skeleton images. We propose a global hand embedding to control the diffusion model based on these hand representations. We then learn a 3D Gaussian splatting for novel-view rendering using the diffusion prior. However, occluded objects present a persistent challenge. To address this issue, we further introduce local hand embedding, where a contact field is defined in the 3D Gaussian Splatting. We leverage contact information to guide the rendering in the contact field. Extensive experiments on the HO3D and DexYCB datasets demonstrate that our method significantly outperforms state-of-the-art novel-view synthesis for hand-object interactions.
Zhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, Jifei Song
AAAI1
2025 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimation
abstract
3D and 2D gaze estimation share the fundamental objective of capturing eye movements but are traditionally treated as two distinct research domains. In this paper, we introduce a novel cross-task few-shot 2D gaze estimation approach, aiming to adapt a pre-trained 3D gaze estimation network for 2D gaze prediction on unseen devices using only a few training images. This task is highly challenging due to the domain gap between 3D and 2D gaze, unknown screen poses, and limited training data. To address these challenges, we propose a novel framework that bridges the gap between 3D and 2D gaze. Our framework contains a physics-based differentiable projection module with learnable parameters to model screen poses and project 3D gaze into 2D gaze. The framework is fully differentiable and can integrate into existing 3D gaze networks without modifying their original architecture. Additionally, we introduce a dynamic pseudo-labelling strategy for flipped images, which is particularly challenging for 2D labels due to unknown screen poses. To overcome this, we reverse the projection process by converting 2D labels to 3D space, where flipping is performed. Notably, this 3D space is not aligned with the camera coordinate system, so we learn a dynamic transformation matrix to compensate for this misalignment. We evaluate our method on MPIIGaze, EVE, and GazeCapture datasets, collected respectively on laptops, desktop computers, and mobile devices. The superior performance highlights the effectiveness of our approach, and demonstrates its strong potential for real-world applications.
Yihua Cheng, Hengfei Wang, Zhongqun Zhang, Boeun Kim, Feng Lu 0005, Hyung Jin Chang
CVPR3
2025 Multi-Hypothesis 3D Hand Mesh Recovering from a Single Blurry Image
abstract
Recovery of 3D hand mesh from blurry hand images is challenging due to the ambiguity. Most existing works attempt to solve this issue by exploiting physical and temporal constraints. However, those works ignore the fact that multiple feasible solutions exist. In this paper, we propose a two-stage Multi-Hypothesis Hand Mesh Recovery network, consisting of a generation and selection model. In the first stage, the generation model explicitly extracts the temporal information with an unfolder. Then, a multi-hypothesis Transformer generates multiple diverse hypotheses with a lightweight hypothesis embedding set. In the second stage, the selection model selects a subset of good-quality hypotheses. We additionally combine the classifying and ranking loss to better align with the target of the selection model. Extensive experiments show that the proposed method produces much more accurate results on blurry images. Source code is available at https://github.com/RandSF/Multi_Hypothesis_BlurHandNet.
Rongyu Chen, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
ICME3
2024 NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object Interaction
abstract
Modeling hand-object interactions is a fundamentally challenging task in 3D computer vision. Despite remarkable progress that has been achieved in this field, existing methods still fail to synthesize the hand-object interaction photo-realistically, suffering from degraded rendering quality caused by the heavy mutual occlusions between the hand and the object, and inaccurate hand-object pose estimation. To tackle these challenges, we present a novel free-viewpoint rendering framework, Neural Contact Radiance Field (NCRF), to reconstruct hand-object interactions from a sparse set of videos. In particular, the proposed NCRF framework consists of two key components: (a) A contact optimization field that predicts an accurate contact field from 3D query points for achieving desirable contact between the hand and the object. (b) A hand-object neural radiance field to learn an implicit hand-object representation in a static canonical space, in concert with the specifically designed hand-object motion field to produce observation-to-canonical correspondences. We jointly learn these key components where they mutually help and regularize each other with visual and geometric constraints, producing a high-quality hand-object reconstruction that achieves photorealistic novel view synthesis. Extensive experiments on HO3D and DexYCB datasets show that our approach outperforms the current state-of-the-art in terms of both rendering quality and pose estimation accuracy.
Zhongqun Zhang, Jifei Song, Eduardo Pérez-Pellitero, Yiren Zhou, Hyung Jin Chang, Ales Leonardis
3DV1
2024 NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
Zhongqun Zhang, Hengfei Wang, Ziwei Yu, Yihua Cheng, Angela Yao, Hyung Jin Chang
ECCV (28)1
2024 TextGaze: Gaze-Controllable Face Generation with Natural Language
abstract
Generating face image with specific gaze information has attracted considerable attention in recent years. Existing approaches typically input gaze values directly for face generation, which is unnatural and requires annotated gaze datasets for training, thereby limiting its application. In this paper, we present a novel gaze-controllable face generation task that overcomes these limitations. Our approach inputs textual descriptions that describe human gaze and head behavior and generates corresponding face images. Our work first introduces a text-of-gaze dataset containing over 90k text descriptions spanning a dense distribution of gaze and head poses. We further propose a gaze-controllable text-to-face method. Our method contains a sketch-conditioned face diffusion module and a model-based sketch diffusion module. We define a face sketch based on facial landmarks and eye segmentation map. It provides a structured and detailed foundation for generating facial images. The face diffusion module generates face images from the face sketch, and the sketch diffusion module employs a 3D face model to generate face sketch from text description. Experiments on the FFHQ dataset show the effectiveness of our method. Our dataset is available at https://github.com/hengfei-wang/TextGaze.
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
ACM Multimedia2
2024 Multi-Modal Gaze Following in Conversational Scenarios
abstract
Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides crucial cues for determining human behavior. This suggests that we can further improve gaze following considering audio cues. In this paper, we explore gaze following tasks in conversational scenarios. We propose a novel multimodal gaze following framework based on our observation "audiences tend to focus on the speaker". We first leverage the correlation between audio and lips, and classify speakers and listeners in a scene. We then use the identity information to enhance scene images and propose a gaze candidate estimation network. The network estimates gaze candidates from enhanced scene images and we use MLP to match subjects with candidates as classification tasks. Existing gaze following datasets focus on visual images while ignore audios. To evaluate our method, we collect a conversational dataset, VideoGazeSpeech (VGS), which is the first gaze following dataset including images and audio. Our method significantly outperforms existing methods in VGS datasets. The visualization result also prove the advantage of audio cues in gaze following tasks. Our work will inspire more researches in multi-modal gaze following estimation.
Zhongqun Zhang, Nora Horanyi, Jaewon Moon, Yihua Cheng, Hyung Jin Chang
WACV2
2023 High-Fidelity Eye Animatable Neural Radiance Fields for Human Face
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
BMVC2
2022 S2Contact: Graph-Based Network for 3D Hand-Object Contact Estimation with Semi-supervised Learning
Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng 0001, Hyung Jin Chang
ECCV (1)2
2022 Towards Generic 3D Tracking in RGBD Videos: Benchmark and Baseline
Zhongqun Zhang, Zhe Li 0008, Hyung Jin Chang, Ales Leonardis, Feng Zheng 0001
ECCV (22)2