Hsuan-I Ho

dblp:207/9893 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0001-8683-7538ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Gaussian Wardrobe: Compositional 3D Gaussian Avatars for Free-Form Virtual Try-On
abstract
We introduce Gaussian Wardrobe, a novel framework to digitalize compositional 3D neural avatars from multi-view videos. Existing methods for 3D neural avatars typically treat the human body and clothing as an inseparable entity. However, this paradigm fails to capture the dynamics of complex free-form garments and limits the reuse of clothing across different individuals. To overcome these problems, we develop a novel, compositional 3D Gaussian representation to build avatars from multiple layers of free-form garments. The core of our method is decomposing neural avatars into bodies and layers of shape-agnostic neural garments. To achieve this, our framework learns to disentangle each garment layer from multi-view videos and canonicalizes it into a shape-independent space. In experiments, our method models photorealistic avatars with high-fidelity dynamics, achieving new state-of-the-art performance on novel pose synthesis benchmarks. In addition, we demonstrate that the learned compositional garments contribute to a versatile digital wardrobe, enabling a practical virtual try-on application where clothing can be freely transferred to new subjects.
Jie Song 0006, Hsuan-I Ho, Manuel Kaufmann, Tianjian Jiang
3DV3
2025 PHD: Personalized 3D Human Body Fitting with Point Diffusion
abstract
We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-efficient, requiring only synthetic data for training, and serves as a versatile plug-and-play module that can be seamlessly integrated with existing 3D pose estimators to enhance their performance. Project page: https://phd-pose.github.io/
Hsuan-I Ho, Po-Chen Wu, Ivan Shugurov, Chengcheng Tang, Abhay Mittal, Sizhe An, Manuel Kaufmann, Linguang Zhang
ICCV1
2025 PriorAvatar: Efficient and Robust Avatar Creation from Monocular Video Using Learned Priors
abstract
High-fidelity avatar reconstruction from monocular videos faces significant challenges due to imperfect foreground segmentation and inaccurate body poses. Existing methods typically depend on additive components, such as explicit background modeling, which introduce additional overhead and reduce the flexibility of avatar reconstruction. We argue that these challenges need to be addressed fundamentally. To this end, we propose leveraging a learned 3D human prior to guide the reconstruction of 3D avatars, dubbed PriorAvatar, without increasing model complexity. At the core of our method is a learned 3D prior, which consists of a multi-person feature codebook that stores the 3D shapes and appearances derived from human scans. These latent features are complemented by a shared U-Net decoder that converts them into a set of renderable 3D Gaussians. During reconstruction, the learned 3D prior allows for fitting to unseen subjects in the monocular videos by fine-tuning with 2D photometric losses using 3D Gaussians. This approach ensures that the reconstruction process effectively utilizes the learned latent spaces while minimizing discrepancies with the 2D observations. In our experiments, we demonstrate the efficiency and robustness of our novel reconstruction scheme, as evidenced by its state-of-the-art quantitative and qualitative performance without relying on complex regularizers or additional model enhancements. The results of ablation studies further verify the effectiveness of incorporating a learned human prior for monocular avatar reconstruction.
Tianjian Jiang, Hsuan-I Ho, Manuel Kaufmann, Jie Song 0006
SIGGRAPH Asia2
2024 4D-DRESS: A 4D Dataset of Real-World Human Clothing with Semantic Annotations
abstract
The studies of human clothing for digital avatars have predominantly relied on synthetic datasets. While easy to collect, synthetic data often fall short in realism and fail to capture authentic clothing dynamics. Addressing this gap, we introduce 4D-DRESS, the first real-world 4D dataset advancing human clothing research with its high-quality 4D textured scans and garment meshes. 4D-DRESS captures 64 outfits in 520 human motion sequences, amounting to 78k textured scans. Creating a real-world clothing dataset is challenging, particularly in annotating and segmenting the extensive and complex 4D human scans. To address this, we develop a semi-automatic 4D human parsing pipeline. We efficiently combine a human-in-the-loop process with automation to accurately label 4D scans in diverse garments and body movements. Leveraging precise annotations and high-quality garment meshes, we establish several benchmarks for clothing simulation and reconstruction. 4D-DRESS offers realistic and challenging data that complements synthetic sources, paving the way for advancements in research of lifelike human clothing.
Wenbo Wang 0007, Hsuan-I Ho, Boxiang Rong, Artur Grigorev 0002, Jie Song 0006, Juan Jose Zarate, Otmar Hilliges
CVPR2
2024 SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion
abstract
A long-standing goal of 3D human reconstruction is to cre-ate lifelike and fully detailed 3D humans from single-view images. The main challenge lies in inferring unknown body shapes, appearances, and clothing details in areas not visi-ble in the images. To address this, we propose SiTH, a novel pipeline that uniquely integrates an image-conditioned dif-fusion model into a 3D mesh reconstruction workflow. At the core of our method lies the decomposition of the chal-lenging single-view reconstruction problem into generative hallucination and reconstruction subproblems. For the for-mer, we employ a powerful generative diffusion model to hallucinate unseen back-view appearance based on the in-put images. For the latter, we leverage skinned body meshes as guidance to recover full-body texture meshes from the in-put and back-view images. SiTH requires as few as 500 3D human scans for training while maintaining its generality and robustness to diverse images. Extensive evaluations on two 3D human benchmarks, including our newly created one, highlighted our method's superior accuracy and per-ceptual quality in 3D textured human reconstruction.
Hsuan-I Ho, Jie Song 0006, Otmar Hilliges
CVPR1
2024 HSR: Holistic 3D Human-Scene Reconstruction from Monocular Videos
Lixin Xue, Chengwei Zheng, Fangjinhua Wang, Tianjian Jiang, Hsuan-I Ho, Manuel Kaufmann, Jie Song 0006, Otmar Hilliges
ECCV (72)6
2023 Learning Locally Editable Virtual Humans
abstract
In this paper, we propose a novel hybrid representation and end-to-end trainable network architecture to model fully editable and customizable neural avatars. At the core of our work lies a representation that combines the modeling power of neural fields with the ease of use and inherent 3D consistency of skinned meshes. To this end, we construct a trainable feature codebook to store local geometry and texture features on the vertices of a deformable body model, thus exploiting its consistent topology under articulation. This representation is then employed in a generative auto-decoder architecture that admits fitting to unseen scans and sampling of realistic avatars with varied appearances and geometries. Furthermore, our representation allows local editing by swapping local features between 3D assets. To verify our method for avatar creation and editing, we contribute a new highquality dataset, dubbed CustomHumans, for training and evaluation. Our experiments quantitatively and qualitatively show that our method generates diverse detailed avatars and achieves better model fitting performance compared to state-of-the-art methods. Our code and dataset are available at https://ait.ethz.ch/customhumans.
Hsuan-I Ho, Lixin Xue, Jie Song 0006, Otmar Hilliges
CVPR1
2021 Render In-between: Motion Guided Video Synthesis for Action Interpolation
Hsuan-I Ho, Xu Chen 0025, Jie Song 0006, Otmar Hilliges
BMVC1
2020 READ: Reciprocal Attention Discriminator for Image-to-Video Re-identification
Minho Shim, Hsuan-I Ho, Jinhyung Kim, Dongyoon Wee
ECCV (14)2
2020 Learning from Dances: Pose-Invariant Re-Identification for Multi-Person Tracking
abstract
Most existing multi-person tracking approaches rely on appearance based re-identification (re-ID) to resolve fragmented tracklets. However, simply using appearance information could be insufficient for videos containing severe pose changes, such as sports or dance videos. With the goal of learning pose-invariant representations, we propose an end-to-end deep learning framework Sparse-Temporal ReID Network. Our proposed network not only realizes human pose disentanglement in an image recovery manner, but also makes efficient linkages between the identical subjects via a unique Sparse temporal identity sampling technique across time steps. Experimental results demonstrate the effectiveness of our proposed method on both multi-view re-ID benchmarks and our newly collected dance video dataset DanceReID1.
Hsuan-I Ho, Minho Shim, Dongyoon Wee
ICASSP1
2018 Summarizing First-Person Videos from Third Persons' Points of Views
Hsuan-I Ho, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang
ECCV (15)1