He Zhang 0015

dblp:24/2058-15 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-7280-6746ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Human Pose Estimation with General Contact
abstract
Existing human pose estimation methods seldom consider the impact or constraint of different types of contact. In this paper, we elaborate on the impact of both body-scene contact and self-contact on pose estimation and refer to them as general contact. First, we extend existing datasets by calculating additional contact labels for general contact inference. Moreover, based on the extended dataset, we present the first network to predict dense general contact from a single RGB image. Finally, we develop a novel optimization method that successfully utilizes the inferred general contact information for accurate 3D pose estimation. Our results show that knowledge of contact can provide strong constraints and resolve pose ambiguity, thus significantly improving human pose estimation accuracy, especially for challenging poses that cannot be well handled by existing methods. Experimental results and comparisons further demonstrate the effectiveness of the proposed method. Our results are even more reasonable than certain pseudo-ground truth determined from multi-view images.
He Zhang 0015, Jianhui Zhao 0002, Fan Li 0023, Yitian Wu, Shuangpeng Sun, Yaohua Wu, Tao Yu 0007
Comput. Vis. Media1
2024 MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors
abstract
Foot contact is an important cue for human motion capture, understanding, and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However, these approaches either suffer from low accuracy or are only designed for small-range and slow motion. There is still a lack of a vision-pressure multimodal dataset with large-range and fast human motion, as well as accurate and dense foot-contact annotation. To fill this gap, we propose a Multimodal MoCap Dataset with Vision and Pressure sensors, named MMVP. MMVP provides accurate and dense plantar pressure signals synchronized with RGBD observations, which is especially useful for both plausible shape estimation, robust pose fitting without foot drifting, and accurate global translation tracking. To validate the dataset, we propose an RGBD-P SMPL fitting method and also a monocular-video-based baseline framework, VP-MoCap, for human motion capture. Experiments demonstrate that our RGBD-P SMPL Fitting results significantly outperform pure visual motion capture. Moreover, VP-MoCap outperforms SOTA methods in foot-contact and global translation estimation accuracy. We believe the configuration of the dataset and the baseline frameworks will stimulate the research in this direction and also provide a good reference for MoCap applications in various domains. Project page: https://metaverse-ai-lab-thu.github.io/MMVP-Dataset/
He Zhang 0015, Shenghao Ren, Haolei Yuan, Jianhui Zhao 0002, Fan Li 0023, Shuangpeng Sun, Zhenghao Liang, Tao Yu 0007, Qiu Shen, Xun Cao
CVPR1
2024 Bisection Window and Homogeneity Principle- Based Local Contrast Measure for Infrared Sea-Sky Line Detection
abstract
Fast and accurate infrared (IR) sea–sky line (SSL) detection could greatly benefit the efficiency of target detection under maritime. However, the traditional SSL detection algorithms are slightly inferior in detection accuracy, while the convolutional neural network (CNN)-based algorithms have high standard hardware and dataset requirements, which are difficult to satisfy in some practical scenes. In this article, a novel concise and intuitive SSL detection algorithm named bisection window and homogeneity principle-based local contrast measure (BHLCM) is proposed. First, a bisection local contrast window (BLCW) is proposed based on the local contrast measure (LCM) and be used for the search of the patches that contain SSL segments along a set of variable preset vertical paths. Then, based on the analysis of the physical characteristics of areas near SSL, the homogeneity principle is proposed to remove the patches containing false SSL segments. Finally, the midpoints of the remaining SSL segments are used to fit the final SSL through random sample consensus (RANSAC). Experimental results based on four IR image sequences with more than 3000 IR images illustrate that compared with the state-of-the-art algorithms, BHLCM not only achieves precision comparable to the most accurate one but also is significantly ahead in speed among most of the algorithms compared. In addition, speed tests about BHLCM based on different hardware platforms show that real-time detection could still be achieved even without high-performance graphics cards. The code and dataset are available at BHLCMhttps://github.com/FJsRepo/BHLCM.
Jianhui Zhao 0002, Fan Li 0023, Yongfei Wang, He Zhang 0015
IEEE Geosci. Remote. Sens. Lett.5
2024 HVTR++: Image and Pose Driven Human Avatars Using Hybrid Volumetric-Textural Rendering
abstract
Recent neural rendering methods have made great progress in generating photorealistic human avatars. However, these methods are generally conditioned only on low-dimensional driving signals (e.g., body poses), which are insufficient to encode the complete appearance of a clothed human. Hence they fail to generate faithful details. To address this problem, we exploit driving view images (e.g., in telepresence systems) as additional inputs. We propose a novel neural rendering pipeline, Hybrid Volumetric-Textural Rendering (HVTR++), which synthesizes 3D human avatars from arbitrary driving poses and views while staying faithful to appearance details efficiently and at high quality. First, we learn to encode the driving signals of pose and view image on a dense UV manifold of the human body surface and extract UV-aligned features, preserving the structure of a skeleton-based parametric model. To handle complicated motions (e.g., self-occlusions), we then leverage the UV-aligned features to construct a 3D volumetric representation based on a dynamic neural radiance field. While this allows us to represent 3D geometry with changing topology, volumetric rendering is computationally heavy. Hence we employ only a rough volumetric representation using a pose- and image-conditioned downsampled neural radiance field (PID-NeRF), which we can render efficiently at low resolutions. In addition, we learn 2D textural features that are fused with rendered volumetric features in image space. The key advantage of our approach is that we can then convert the fused features into a high-resolution, high-quality avatar by a fast GAN-based textural renderer. We demonstrate that hybrid rendering enables HVTR++ to handle complicated motions, render high-quality avatars under user-controlled poses/shapes, and most importantly, be efficient at inference time. Our experimental results also demonstrate state-of-the-art quantitative results.
Tao Hu 0006, Linjie Luo, Tao Yu 0007, Zerong Zheng, He Zhang 0015, Yebin Liu, Matthias Zwicker
IEEE Trans. Vis. Comput. Graph.6
2023 Infrared Small Dim Target Detection Under Maritime Near Sea-Sky Line Based on Regional-Division Local Contrast Measure
abstract
Infrared (IR) small dim target detection near the sea-sky line (SSL) is crucial for enhancing the early warning capability of maritime vehicles. However, the interferences caused by the strong contrast have not been properly addressed. Consequently, a specially designed algorithm Regional-Division Local Contrast Measure (RDLCM) that focuses on the detection of infrared small dim targets appearing near the SSL is proposed. First, an SSL detection module based on a lightweight Convolutional Neural Network (CNN) is devised to achieve fast pixel-level SSL detection. Then, a set of regional-division windows (RDW) are designed according to the strong grayscale contrast distribution around the SSL, through the division of the effective regions, the RDWs could realize the potential extraction and refinement of the IR small dim targets that appear near the SSL. Experiments on three IR image sequences demonstrate that the proposed algorithm achieves the best detection accuracy among the classical and the state-of-the-art algorithms in comparison, and runs at 44 frames per second (FPS) which could meet real-time requirements. The code and dataset are available at RDLCM.
Fan Li 0023, Jianhui Zhao 0002, Jie Tong, He Zhang 0015
IEEE Geosci. Remote. Sens. Lett.5
2023 Controllable Free Viewpoint Video Reconstruction Based on Neural Radiance Fields and Motion Graphs
abstract
In this paper, we propose a controllable high-quality free viewpoint video generation method based on the motion graph and neural radiance fields (NeRF). Different from existing pose-driven NeRF or time/structure conditioned NeRF works, we propose to first construct a directed motion graph of the captured sequence. Such a sequence-motion-parameterization strategy not only enables flexible pose control for free viewpoint video rendering but also avoids redundant calculation of similar poses and thus improves the overall reconstruction efficiency. Moreover, to support body shape control without losing the realistic free viewpoint rendering performance, we improve the vanilla NeRF by combining explicit surface deformation and implicit neural scene representations. Specifically, we train a local surface-guided NeRF for each valid frame on the motion graph, and the volumetric rendering was only performed in the local space around the real surface, thus enabling plausible shape control ability. As far as we know, our method is the first method that supports both realistic free viewpoint video reconstruction and motion graph-based user-guided motion traversal. The results and comparisons further demonstrate the effectiveness of the proposed method.
He Zhang 0015, Fan Li 0023, Jianhui Zhao 0002, Dongming Shen, Yebin Liu, Tao Yu 0007
IEEE Trans. Vis. Comput. Graph.1
2022 HVTR: Hybrid Volumetric-Textural Rendering for Human Avatars
abstract
We propose a novel neural rendering pipeline, Hybrid Volumetric-Textural Rendering (HVTR), which synthesizes virtual human avatars from arbitrary poses efficiently and at high quality. First, we learn to encode articulated human motions on a dense UV manifold of the human body surface. To handle complicated motions (e.g., self-occlusions), we then leverage the encoded information on the UV manifold to construct a 3D volumetric representation based on a dynamic pose-conditioned neural radiance field. While this allows us to represent 3D geometry with changing topology, volumetric rendering is computationally heavy. Hence we employ only a rough volumetric representation using a pose-conditioned downsampled neural radiance field (PD-NeRF), which we can render efficiently at low resolutions. In addition, we learn 2D textural features that are fused with rendered volumetric features in image space. The key advantage of our approach is that we can then convert the fused features into a high-resolution, high-quality avatar by a fast GAN-based textural renderer. We demonstrate that hybrid rendering enables HVTR to handle complicated motions, render high-quality avatars under user-controlled poses/shapes and even loose clothing, and most importantly, be efficient at inference time. Our experimental results also demonstrate state-of-the-art quantitative results. More results are available at our project page: https://www.cs.umd.edu/~taohu/hvtr/
Tao Hu 0006, Tao Yu 0007, Zerong Zheng, He Zhang 0015, Yebin Liu, Matthias Zwicker
3DV4
2022 DoubleField: Bridging the Neural Surface and Radiance Fields for High-fidelity Human Reconstruction and Rendering
abstract
We introduce DoubleField, a novel framework combining the merits of both surface field and radiance field for high-fidelity human reconstruction and rendering. Within DoubleField, the surface field and radiance field are associated together by a shared feature embedding and a surface-guided sampling strategy. Moreover, a view-to-view transformer is introduced to fuse multi-view features and learn view-dependent features directly from high-resolution inputs. With the modeling power of DoubleField and the view-to-view transformer, our method significantly improves the reconstruction quality of both geometry and appearance, while supporting direct inference, scene-specific high-resolution finetuning, and fast rendering. The efficacy of DoubleField is validated by the quantitative evaluations on several datasets and the qualitative results in a real-world sparse multi-view system, showing its superior capability for high-quality human model reconstruction and photo-realistic free-viewpoint human rendering. Data and source code will be made public for the research purpose.
Ruizhi Shao, Hongwen Zhang 0001, He Zhang 0015, Mingjia Chen, Yan-Pei Cao 0001, Tao Yu 0007, Yebin Liu
CVPR3