Hang Ye 0002

dblp:40/11094-2 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-6646-1059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 40% Trustworthy machine learning · 18% Reinforcement learning · 18%
Computer graphics and multimedia
2 papers
Geometric modeling and processing · 50% Computational photography and imaging · 25% Visual content generation and editing · 25%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d human reconstruction
0.912025
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data · NeurIPS 2025
Computer vision › 3D vision › human body modeling › 3d human modeling
human avatar
0.912025
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling · CVPR 2025
Geometric modeling and processing › 3d reconstruction
3d human reconstruction
0.912025
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data · NeurIPS 2025
Geometric modeling and processing › shape modeling › human body modeling
clothed human modeling
0.912025
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling · CVPR 2025
Computational photography and imaging › image-based modeling › 3d reconstruction from images
single-view 3d reconstruction
0.912025
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data · NeurIPS 2025
Visual content generation and editing › image generation › text-to-image generation
text-to-image diffusion
0.912025
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data · NeurIPS 2025
Machine learning › Reinforcement learning
imitation learning
0.712023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.712023
Denoising Masked Autoencoders Help Robust Classification · ICLR 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness › robust learning
robust classification
0.712023
Denoising Masked Autoencoders Help Robust Classification · ICLR 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Denoising Masked Autoencoders Help Robust Classification · ICLR 2023
Computer vision › 3D vision
3d human pose estimation
0.612022
Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection · ECCV (6) 2022
Computer vision › Face, body and person analysis
human pose estimation
0.612022
Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection · ECCV (6) 2022
Computer vision › 3D vision › camera calibration › camera model
orthographic projection
0.612022
Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection · ECCV (6) 2022
Computer vision › Video understanding and tracking
human motion prediction
0.212023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023
Robotics › Autonomous driving › trajectory prediction
multi-person motion prediction
0.212023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

texture refinement · 1.7multi-view reconstruction · 1.7linear blend skinning · 1.7free-form generation · 1.7diffusion model · 1.7masked image modeling · 0.7generative adversarial imitation learning · 0.7denoising · 0.7cognitive hierarchy framework · 0.7behavioral cloning · 0.7
YearPublicationVenuePosition
2025 FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
abstract
Achieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skinning (LBS) of minimally-clothed human models like SMPL to model deformation. However, they struggle to handle loose clothing, such as long dresses, where the canonicalization process becomes ill-defined when the clothing is far from the body, leading to disjointed and fragmented results. To overcome this limitation, we propose FreeCloth, a novel hybrid framework to model challenging clothed humans. Our core idea is to use dedicated strategies to model different regions, depending on whether they are close to or distant from the body. Specifically, we segment the human body into three categories: unclothed, deformed, and generated. We simply replicate unclothed regions that require no deformation. For deformed regions close to the body, we leverage LBS to handle the deformation. As for the generated regions, which correspond to loose clothing areas, we introduce a novel free-form, part-aware generator to model them, as they are less affected by movements. This free-form generation paradigm brings enhanced flexibility and expressiveness to our hybrid framework, enabling it to capture the intricate geometric details of challenging loose clothing, such as skirts and dresses. Experimental results on the benchmark dataset featuring loose clothing demonstrate that FreeCloth achieves state-of-the-art performance with superior visual fidelity and realism, particularly in the most challenging cases.
Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Wentao Zhu 0004, Yizhou Wang 0001
CVPR1
2025 GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
abstract
Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities in human postures and inconsistency in human textures. In addition, the scarcity of high-quality human data intensifies the challenge. To address these problems, we propose a Generalizable image-to-3D huMAN reconstruction framework, dubbed GeneMAN, building upon a comprehensive multi-source collection of high-quality human data, including 3D scans, multi-view videos, single photos, and our generated synthetic human data. GeneMAN encompasses three key modules. 1) Without relying on parametric human models (e.g., SMPL), GeneMAN first trains a human-specific text-to-image diffusion model and a view-conditioned diffusion model, serving as GeneMAN 2D human prior and 3D human prior for reconstruction, respectively. 2) With the help of the pretrained human prior models, the Geometry Initialization-&-Sculpting pipeline is leveraged to recover high-quality 3D human geometry given a single image. 3) To achieve high-fidelity 3D human textures, GeneMAN employs the Multi-Space Texture Refinement pipeline, consecutively refining textures in the latent and the pixel spaces. Extensive experimental results demonstrate that GeneMAN could generate high-quality 3D human models from a single image input, outperforming prior state-of-the-art methods. Notably, GeneMAN could reveal much better generalizability in dealing with in-the-wild images, often yielding high-quality 3D human models in natural poses with common items, regardless of the body proportions in the input images.
Wentao Wang 0009, Hang Ye 0002, Fangzhou Hong, Xue Yang 0005, Jianfu Zhang 0003, Yizhou Wang 0001, Ziwei Liu 0002, Liang Pan
NeurIPS2
2023 Denoising Masked Autoencoders Help Robust Classification
Quanlin Wu, Hang Ye 0002, Yuntian Gu, Huishuai Zhang, Liwei Wang 0001, Di He 0001
ICLR2
2023 Social Motion Prediction with Cognitive Hierarchies
abstract
Humans exhibit a remarkable capacity for anticipating the actions of others and planning their own actions accordingly. In this study, we strive to replicate this ability by addressing the social motion prediction problem. We introduce a new benchmark, a novel formulation, and a cognition-inspired framework. We present Wusi, a 3D multi-person motion dataset under the context of team sports, which features intense and strategic human interactions and diverse pose distributions. By reformulating the problem from a multi-agent reinforcement learning perspective, we incorporate behavioral cloning and generative adversarial imitation learning to boost learning efficiency and generalization. Furthermore, we take into account the cognitive aspects of the human social action planning process and develop a cognitive hierarchy framework to predict strategic human social interactions. We conduct comprehensive experiments to validate the effectiveness of our proposed dataset and approach.
Wentao Zhu 0004, Jason Qin, Yuke Lou, Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Yizhou Wang 0001
NeurIPS4
2022 Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection
Hang Ye 0002, Wentao Zhu 0004, Chunyu Wang 0001, Rujie Wu, Yizhou Wang 0001
ECCV (6)1