Jinlu Zhang 0001

dblp:130/5411-1 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-8142-5470ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action Recognition
abstract
Human action recognition in low-light environments is crucial for various real-world applications. However, the existing methods overlook the full utilization of brightness information throughout the training phase, leading to suboptimal performance. To address this issue, we propose OwlSight, a biomimetic framework with whole-stage illumination enhancement to interact with action classification for accurate dark video human action recognition. Specifically, OwlSight incorporates a Time-Consistency Module (TCM) to capture shallow spatiotemporal features meanwhile maintaining temporal coherence, which are then processed by a Luminance Adaptation Module (LAM) to dynamically adjust the brightness based on the input luminance distribution. Furthermore, a Reflect Augmentation Module (RAM) is presented to maximize illumination utilization and simultaneously enhance action recognition via two interactive paths. Additionally, we build a large-scale datasetDark-101, which comprises 21,030 dark videos across 101 action categories, significantly surpassing the existing datasets (e.g., ARID1.5 and Dark-48) in scale and diversity. Our method establishes new state-of-the-art (SOTA) performance across all benchmarks, achieving Top-1 accuracies of 99.27% on ARID (a 2.0% improvement), 94.85% on ARID1.5 (a 5.36% improvement), 48.24% on Dark-48 (a 1.56% improvement), and 53.85% on Dark-101 (a 1.72% improvement), demonstrating its superior effectiveness in challenging dark video environments.
Shihao Cheng, Jinlu Zhang 0001, Aoran Xiao, Zhigang Tu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
abstract
Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main challenges of this task: precise human-object relation reasoning, affordance parsing for any object, and detailed human interaction pose synthesis aligning description and object geometry. In this work, we propose a novel zero-shot 3D HOI generation framework without training on specific datasets, leveraging the knowledge from large-scale pre-trained models. Specifically, the human-object relations are inferred from large language models (LLMs) to initialize object properties and guide the optimization process. Then we utilize a pre-trained 2D image diffusion model to parse unseen objects and extract contact points, avoiding the limitations imposed by existing 3D asset knowledge. The initial human pose is generated by sampling multiple hypotheses through multi-view SDS based on the input text and object geometry. Finally, we introduce a detailed optimization to generate fine-grained, precise, and natural interaction, enforcing realistic 3D contact between the 3D object and the involved body parts, including hands in grasping. This is achieved by distilling human-level feedback from LLMs to capture detailed human-object relations from the text instruction. Extensive experiments validate the effectiveness of our approach compared to prior works, particularly in terms of the fine-grained nature of interactions and the ability to handle open-set 3D objects. Project page: jinluzhang.site/projects/interactanything.
Jinlu Zhang 0001, Yixin Chen 0003, Yizhou Wang 0001, Siyuan Huang 0001
CVPR1
2025 Expressive Keypoints for Skeleton-Based Action Recognition via Progressive Skeleton Evolution
abstract
In the realm of skeleton-based human action recognition, the traditional methods which rely on coarse body keypoints fall short of capturing subtle human actions. In this work, we propose Expressive Keypoints that incorporates hand and foot details to form a fine-grained skeletal representation, to improve the discriminative ability for existing models in discerning intricate human actions. However, the increased computational cost from processing nearly three times more joints becomes a new challenge. To address this, we present the Progressive Skeleton Evolution strategy, which significantly improves efficiency while preserving the benefits of fine-grained keypoints. The core idea involves utilizing learnable mapping matrices, semantically initialized to progressively downsample keypoints and prioritize prominent joints by allocating importance weights. Additionally, a plug-and-play Instance Pooling module is exploited to extend our approach to multi-person scenarios without surging computation cost. Extensive experimental results over seven datasets demonstrate the superiority of our method compared to the state-of-the-arts for skeleton-based human action recognition. Code has been made available at https://github.com/YijieYang23/PSE-GCN.
Jinlu Zhang 0001, Bo Du 0001, Zhigang Tu 0001
IEEE Trans. Image Process.2
2024 Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene Affordance
abstract
Despite significant advancements in text-to-motion syn-thesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models capable of jointly modeling natural language, 3D scenes, and human motion, and (ii) the generative models' in-tensive data requirements contrasted with the scarcity of comprehensive, high-quality, language-scene-motion datasets. To tackle these issues, we introduce a novel two-stage frame-work that employs scene affordance as an intermediate representation, effectively linking 3D scene grounding and conditional motion generation. Our framework comprises an Affordance Diffusion Model (ADM) for predicting ex-plicit affordance map and an Affordance-to-Motion Diffusion Model (AMDM) for generating plausible human motions. By leveraging scene affordance maps, our method overcomes the difficulty in generating human motion under multimodal condition signals, especially when training with limited data lacking extensive language-scene-motion pairs. Our exten-sive experiments demonstrate that our approach consistently outperforms all baselines on established benchmarks, in-cluding HumanML3D and HUMANISE. Additionally, we validate our model's exceptional generalization capabilities on a specially curated evaluation set featuring previously unseen descriptions and scenes.
Yixin Chen 0003, Baoxiong Jia, Puhao Li, Jinlu Zhang 0001, Jingze Zhang, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Siyuan Huang 0001
CVPR5
2024 Human Motion Generation: A Survey
abstract
Human motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying the foundation for increasing interest in human motion generation. Most research within this field focuses on generating human motions based on conditional signals, such as text, audio, and scene contexts. While significant advancements have been made in recent years, the task continues to pose challenges due to the intricate nature of human motion and its implicit relationship with conditional signals. In this survey, we present a comprehensive literature review of human motion generation, which, to the best of our knowledge, is the first of its kind in this field. We begin by introducing the background of human motion and generative models, followed by an examination of representative methods for three mainstream sub-tasks: text-conditioned, audio-conditioned, and scene-conditioned human motion generation. Additionally, we provide an overview of common datasets and evaluation metrics. Lastly, we discuss open problems and outline potential future research directions. We hope that this survey could provide the community with a comprehensive glimpse of this rapidly evolving field and inspire novel ideas that address the outstanding challenges.
Wentao Zhu 0004, Xiaoxuan Ma 0001, Dongwoo Ro, Hai Ci, Jinlu Zhang 0001, Jiaxin Shi, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 PHRIT: Parametric Hand Representation with Implicit Template
abstract
We propose PHRIT, a novel approach for parametric hand mesh modeling with an implicit template that combines the advantages of both parametric meshes and implicit representations. Our method represents deformable hand shapes using signed distance fields (SDFs) with part-based shape priors, utilizing a deformation field to execute the deformation. The model offers efficient high-fidelity hand reconstruction by deforming the canonical template at infinite resolution. Additionally, it is fully differentiable and can be easily used in hand modeling since it can be driven by the skeleton and shape latent codes. We evaluate PHRIT on multiple downstream tasks, including skeleton-driven hand reconstruction, shapes from point clouds, and singleview 3D reconstruction, demonstrating that our approach achieves realistic and immersive hand modeling with state- of-the-art performance.
Zhisheng Huang, Yujin Chen, Jinlu Zhang 0001, Zhigang Tu 0001
ICCV4
2022 MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video
abstract
Recent transformer-based solutions have been introduced to estimate 3D human pose from 2D keypoint sequence by considering body joints among all frames globally to learn spatio-temporal correlation. We observe that the motions of different joints differ significantly. However, the previous methods cannot efficiently model the solid inter-frame correspondence of each joint, leading to insufficient learning of spatial-temporal correlation. We propose MixSTE (Mixed Spatio-Temporal Encoder), which has a temporal transformer block to separately model the temporal motion of each joint and a spatial transformer block to learn inter-joint spatial correlation. These two blocks are utilized alternately to obtain better spatio-temporal feature encoding. In addition, the network output is extended from the central frame to entire frames of the input video, thereby improving the coherence between the input and output sequences. Extensive experiments are conducted on three benchmarks (i.e. Human3.6M, MPI-INF-3DHP, and HumanEva). The results show that our model outperforms the state-of-the-art approach by 10.9% P-MPJPE and 7.6% MPJPE. The code is available at https://github.com/JinluZhang1126/MixSTE.
Jinlu Zhang 0001, Zhigang Tu 0001, Jianyu Yang 0002, Yujin Chen, Junsong Yuan 0001
CVPR1
2022 Uncertainty-Aware 3D Human Pose Estimation from Monocular Video
abstract
Estimating the 3D human pose from the monocular video is challenging mainly due to the depth ambiguity and inaccurate 2D detected keypoints. To quantify the depth uncertainty of 3D human pose via the neural network, we imbue the uncertainty modeling to depth prediction by using evidential deep learning (EDL). Meanwhile, to calibrate the distribution uncertainty of the 2D detection, we explore a probabilistic representation to model the realistic distribution. Specifically, we exploit the EDL to measure the depth prediction uncertainty of the network, and decompose the x-y coordinates into individual distributions to model the deviation uncertainty of the inaccurate 2D keypoints. Then we optimize the depth uncertainty parameters and calibrate the 2D deviations to obtain accurate 3D human poses. Besides, to provide effective latent features for uncertainty learning, we design an encoder which combines graph convolutional network (GCN) and transformer to learn discriminative spatio-temporal representations. Extensive experiments are conducted on three benchmarks (Human3.6M, MPI-INF-3DHP, and HumanEva-I) and the comprehensive results show that our model surpasses the state-of-the-arts by a large margin.
Jinlu Zhang 0001, Yujin Chen, Zhigang Tu 0001
ACM Multimedia1