Yihua Huang 0002

dblp:50/4147-2 · also Yi-Hua Huang 0002 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-2208-8280ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HoloMobile: Photorealistic Avatar on Mobile via Compact Dynamic Gaussians from a Monocular Video
abstract
Human avatars, defined as animatable 3D digital people, have diverse applications, including teleconferencing, motion guidance, remote care, and gaming. With the widespread adoption of smart mobile devices, the demand for real-time viewing of dynamic digital humans on mobile platforms has grown significantly. However, existing avatar reconstruction methods, such as MagicStream and ExAvatar, struggle to produce high-fidelity avatars from monocular videos captured by mobile devices. Furthermore, most related research focuses on individual algorithmic components and does not provide an end-to-end pipeline that jointly addresses avatar reconstruction, 4D data transmission, and mobile rendering. In this work, we introduce HoloMobile, the first end-to-end mobile avatar system with two key capabilities: (1) High-fidelity avatar reconstruction from monocular videos captured on mobile phones, using a hybrid explicit-implicit model. (2) Efficient polynomial compression, transmission, and web-based rendering of dynamic 3D avatars on mobile devices. With avatar reconstruction performed on a PC or cloud with a single consumer-grade GPU, HoloMobile provides an end-to-end solution for mobile users: from data capture to streaming and real-time viewing. Evaluation on 20 self-captured avatars reveals that HoloMobile achieves an average PSNR of 30.36 dB, a compression rate of 99.57%, and over 200 FPS for mobile rendering, demonstrating superiority over existing baselines. The demo video of HoloMobile is available at https://youtu.be/kfOFV7AiP5A.
Yuting He 0006, Yihua Huang 0002, Xiaojuan Qi 0001, Zhenyu Yan 0002, Guoliang Xing
MobiSys2
2026 AniGen: Unified S3 Fields for Animatable 3D Asset Generation
abstract
Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodied agents, and animation production. While recent 3D generative models can synthesize visually plausible shapes from images, the results are typically static. Obtaining usable rigs via post-hoc auto-rigging is brittle and often produces skeletons that are topologically inconsistent with the generated geometry. We present AniGen , a unified framework that directly generates animate-ready 3D assets conditioned on a single image. Our key insight is to represent shape, skeleton, and skinning as mutually consistent S 3 Fields (Shape, Skeleton, Skin) defined over a shared spatial domain. To enable the robust learning of these fields, we introduce two technical innovations: (i) a confidence-decaying skeleton field that explicitly handles the geometric ambiguity of bone prediction at Voronoi boundaries, and (ii) a dual skin feature field that decouples skinning weights from specific joint counts, allowing a fixed-architecture network to predict rigs of arbitrary complexity. Built upon a two-stage flow-matching pipeline, AniGen first synthesizes a sparse structural scaffold and then generates dense geometry and articulation in a structured latent space. Extensive experiments demonstrate that AniGen substantially outperforms state-of-the-art sequential baselines in rig validity and animation quality, generalizing effectively to in-the-wild images across diverse categories including animals, humanoids, and machinery. Homepage : https://yihua7.github.io/AniGen_web/
Yihua Huang 0002, Zixin Zou, Yuting He 0006, Chirui Chang, Cheng-Feng Pu, Ziyi Yang 0008, Yan-Pei Cao 0001, Xiaojuan Qi 0001
ACM Trans. Graph.1
2026 MesoSplats: Texture Synthesis With Gaussian Splatting
abstract
Texture is fundamental to high-fidelity rendering of 3D digital assets, directly influencing scene detail and visual realism. Existing methods typically adopt 2D texture mapping, where texture images are either manually created or synthesized from exemplars. While advances in texture synthesis have improved 2D texture quality, 2D representations remain inadequate for modeling volumetric meso-structure textures with complex geometry. Methods targeting meso-structure textures often struggle to capture high-frequency details and lack real-time rendering capabilities, limiting their practical use. We propose MesoSplats, a neural implicit method for extracting and synthesizing meso-structure textures using 3D Gaussian splatting. Given the multi-view images containing the meso-structure geometric details, our approach supports texture extraction, synthesis, and real-time rendering. We introduce a mesh-Gaussian hybrid representation that decouples geometry into a coarse base mesh and embedded 3D Gaussians, guided by initial point cloud constraints to enhance reconstruction fidelity. Local implicit texture features are sampled from the base mesh surface and further refined through a proposed Consistency Tuning strategy, which enforces alignment between the reconstruction and sampling spaces. To boost texture synthesis quality, we incorporate a tileability-aware patch-matching algorithm alongside a smoothness regularization on the latent feature space to ensure spatial coherence. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method.
Jing-Wen Yang 0002, Jie Yang 0038, Yihua Huang 0002, Yongliang Yang 0002, Yan-Pei Cao 0001, Lin Gao 0004
IEEE Trans. Vis. Comput. Graph.3
2025 Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training
abstract
In-home resistance training (RT) is a convenient and effective way to maintain health and well-being. However, incorrect exercise execution can result in unintended muscle engagement and an increased risk of injury. Without access to professional coaching, an accurate muscle-aware motion feedback system becomes essential for safe and effective training. However, existing visual language models (VLMs) struggle to provide accurate and effective muscle-aware movement guidance due to their limited understanding of RT motion and the absence of related expert knowledge. In this work, we introduce Myo-Trainer, the first vision-based muscle-aware motion feedback system that uses explicit muscle-aware motion analysis and domain-specific expert knowledge to provide corrective guidance on muscle engagement and movement execution. Also, we propose a novel DAGCN-Former network that integrates both spatial and temporal modeling capabilities to capture the complex dynamics of human RT motion. Experiments involving 26 subjects and 1000+ minutes of RT demonstrate that Myo-Trainer improves the accuracy of motion analysis by 17.22%, achieves a 2.5x reduced inference latency and a BertScore of 85.88% of generated feedback compared to those provided by experienced certified trainers, outperforming existing solutions. Additionally, Myo-Trainer received higher satisfaction ratings from participants compared to other AI trainers and video tutorials, highlighting its potential for real-world applications.
Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Bufang Yang, Siyang Jiang, Yihua Huang 0002, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001
MobiCom6
2024 SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes
abstract
Novel view synthesis for dynamic scenes is still a challenging problem in computer vision and graphics. Recently, Gaussian splatting has emerged as a robust technique to represent static scenes and enable high-quality and real-time novel view synthesis. Building upon this technique, we propose a new representation that explicitly decomposes the motion and appearance of dynamic scenes into sparse control points and dense Gaussians, respectively. Our key idea is to use sparse control points, significantly fewer in number than the Gaussians, to learn compact 6 DoF transformation bases, which can be locally interpolated through learned interpolation weights to yield the motion field of 3D Gaussians. We employ a deformation MLP to predict time-varying 6 DoF transformations for each control point, which reduces learning complexities, enhances learning abilities, and facilitates obtaining temporal and spatial coherent motion patterns. Then, we jointly learn the 3D Gaussians, the canonical space locations of control points, and the deformation MLP to reconstruct the appearance, geometry, and dynamics of 3D scenes. During learning, the location and number of control points are adaptively adjusted to accommodate varying motion complexities in different regions, and an ARAP loss following the principle of as rigid as possible is developed to enforce spatial continuity and local rigidity of learned motions. Finally, thanks to the explicit sparse motion representation and its decomposition from appearance, our method can enable user-controlled motion editing while retaining high-fidelity appearances. Extensive experiments demonstrate that our approach outperforms existing approaches on novel view synthesis with a high rendering speed and enables novel appearance-preserved motion editing applications.
Yihua Huang 0002, Yang-Tian Sun, Ziyi Yang 0008, Xiaoyang Lyu, Yan-Pei Cao 0001, Xiaojuan Qi 0001
CVPR1
2024 Splatter a Video: Video Gaussian Representation for Versatile Processing
abstract
Video representation is a long-standing problem that is crucial for various downstream tasks, such as tracking, depth prediction, segmentation, view synthesis, and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D representations that are ill-suited for manipulation tasks. To address these challenges, we introduce a novel explicit 3D representation—video Gaussian representation—that embeds a video into 3D Gaussians. Our proposed representation models video appearance in a 3D canonical space using explicit Gaussians as proxies and associates each Gaussian with 3D motions for video motion. This approach offers a more intrinsic and explicit representation than layered atlas or volumetric pixel matrices. To obtain such a representation, we distill 2D priors, such as optical flow and depth, from foundation models to regularize learning in this ill-posed setting. Extensive applications demonstrate the versatility of our new video representation. It has been proven effective in numerous video processing tasks, including tracking, consistent video depth and feature refinement, motion and appearance editing, and stereoscopic video generation.
Yang-Tian Sun, Yihua Huang 0002, Xiaoyang Lyu, Yan-Pei Cao 0001, Xiaojuan Qi 0001
NeurIPS2
2024 Spec-Gaussian: Anisotropic View-Dependent Appearance for 3D Gaussian Splatting
abstract
The recent advancements in 3D Gaussian splatting (3D-GS) have not only facilitated real-time rendering through modern GPU rasterization pipelines but have also attained state-of-the-art rendering quality. Nevertheless, despite its exceptional rendering quality and performance on standard datasets, 3D-GS frequently encounters difficulties in accurately modeling specular and anisotropic components. This issue stems from the limited ability of spherical harmonics (SH) to represent high-frequency information. To overcome this challenge, we introduce Spec-Gaussian, an approach that utilizes an anisotropic spherical Gaussian (ASG) appearance field instead of SH for modeling the view-dependent appearance of each 3D Gaussian. Additionally, we have developed a coarse-to-fine training strategy to improve learning efficiency and eliminate floaters caused by overfitting in real-world scenes. Our experimental results demonstrate that our method surpasses existing approaches in terms of rendering quality. Thanks to ASG, we have significantly improved the ability of 3D-GS to model scenes with specular and anisotropic components without increasing the number of 3D Gaussians. This improvement extends the applicability of 3D GS to handle intricate scenarios with specular and anisotropic surfaces.
Ziyi Yang 0008, Yang-Tian Sun, Yihua Huang 0002, Xiaoyang Lyu, Shaohui Jiao, Xiaojuan Qi 0001, Xiaogang Jin 0001
NeurIPS4
2024 NeRF-Texture: Synthesizing Neural Radiance Field Textures
abstract
Texture synthesis is a fundamental problem in computer graphics that would benefit various applications. Existing methods are effective in handling 2D image textures. In contrast, many real-world textures contain meso-structure in the 3D geometry space, such as grass, leaves, and fabrics, which cannot be effectively modeled using only 2D image textures. We propose a novel texture synthesis method with Neural Radiance Fields (NeRF) to capture and synthesize textures from given multi-view images. In the proposed NeRF texture representation, a scene with fine geometric details is disentangled into the meso-structure textures and the underlying base shape. This allows textures with meso-structure to be effectively learned as latent features situated on the base shape, which are fed into a NeRF decoder trained simultaneously to represent the rich view-dependent appearance. Using this implicit representation, we can synthesize NeRF-based textures through patch matching of latent features. However, inconsistencies between the metrics of the reconstructed content space and the latent feature space may compromise the synthesis quality. To enhance matching performance, we further regularize the distribution of latent features by incorporating a clustering constraint. In addition to generating NeRF textures over a planar domain, our method can also synthesize NeRF textures over curved surfaces, which are practically useful. Experimental results and evaluations demonstrate the effectiveness of our approach.
Yihua Huang 0002, Yan-Pei Cao 0001, Yukun Lai, Ying Shan, Lin Gao 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Learning Critically: Selective Self-Distillation in Federated Learning on Non-IID Data
abstract
Federated learning (FL) enables multiple clients to collaboratively train a global model while keeping local data decentralized. Data heterogeneity (non-IID) across clients has imposed significant challenges to FL, which makes local models re-optimize towards their own local optima and forget the global knowledge, resulting in performance degradation and convergence slowdown. Many existing works have attempted to address the non-IID issue by adding an extra global-model-based regularizing item to the local training but without an adaption scheme, which is not efficient enough to achieve high performance with deep learning models. In this paper, we propose a Selective Self-Distillation method for Federated learning (FedSSD), which imposes adaptive constraints on the local updates by self-distilling the global model’s knowledge and selectively weighting it by evaluating the credibility at both the class and sample level. The convergence guarantee of FedSSD is theoretically analyzed and extensive experiments are conducted on three public benchmark datasets, which demonstrates that FedSSD achieves better generalization and robustness in fewer communication rounds, compared with other state-of-the-art FL methods.
Yuting He 0008, Yiqiang Chen 0001, Xiaodong Yang 0005, Hanchao Yu, Yihua Huang 0002, Yang Gu 0001
IEEE Trans. Big Data5
2024 3DGSR: Implicit Surface Reconstruction with 3D Gaussian Splatting
abstract
In this paper, we present an implicit surface reconstruction method with 3D Gaussian Splatting (3DGS), namely 3DGSR, that allows for accurate 3D reconstruction with intricate details while inheriting the high efficiency and rendering quality of 3DGS. The key insight is to incorporate an implicit signed distance field (SDF) within 3D Gaussians for surface modeling, and to enable the alignment and joint optimization of both SDF and 3D Gaussians. To achieve this, we design coupling strategies that align and associate the SDF with 3D Gaussians, allowing for unified optimization and enforcing surface constraints on the 3D Gaussians. With alignment, optimizing the 3D Gaussians provides supervisory signals for SDF learning, enabling the reconstruction of intricate details. However, this only offers sparse supervisory signals to the SDF at locations occupied by Gaussians, which is insufficient for learning a continuous SDF. Then, to address this limitation, we incorporate volumetric rendering and align the rendered geometric attributes (depth, normal) with that derived from 3DGS. In sum, these two designs allow SDF and 3DGS to be aligned, jointly optimized, and mutually boosted. Our extensive experimental results demonstrate that our 3DGSR enables high-quality 3D surface reconstruction while preserving the efficiency and rendering quality of 3DGS. Besides, our method competes favorably with leading surface reconstruction techniques while offering a more efficient learning process and much better rendering qualities.
Xiaoyang Lyu, Yang-Tian Sun, Yihua Huang 0002, Xiuzhe Wu, Ziyi Yang 0008, Jiangmiao Pang, Xiaojuan Qi 0001
ACM Trans. Graph.3
2023 Neural Radiance Fields From Sparse RGB-D Images for High-Quality View Synthesis
abstract
The recently proposed neural radiance fields (NeRF) use a continuous function formulated as a multi-layer perceptron (MLP) to model the appearance and geometry of a 3D scene. This enables realistic synthesis of novel views, even for scenes with view dependent appearance. Many follow-up works have since extended NeRFs in different ways. However, a fundamental restriction of the method remains that it requires a large number of images captured from densely placed viewpoints for high-quality synthesis and the quality of the results quickly degrades when the number of captured views is insufficient. To address this problem, we propose a novel NeRF-based framework capable of high-quality view synthesis using only a sparse set of RGB-D images, which can be easily captured using cameras and LiDAR sensors on current consumer devices. First, a geometric proxy of the scene is reconstructed from the captured RGB-D images. Renderings of the reconstructed scene along with precise camera parameters can then be used to pre-train a network. Finally, the network is fine-tuned with a small number of real captured images. We further introduce a patch discriminator to supervise the network under novel views during fine-tuning, as well as a 3D color prior to improve synthesis quality. We demonstrate that our method can generate arbitrary novel views of a 3D scene from as few as 6 RGB-D images. Extensive experiments show the improvements of our method compared with the existing NeRF-based methods, including approaches that also aim to reduce the number of input images.
Yu-Jie Yuan, Yukun Lai, Yihua Huang 0002, Leif Kobbelt, Lin Gao 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Multiscale Mesh Deformation Component Analysis With Attention-Based Autoencoders
abstract
Deformation component analysis is a fundamental problem in geometry processing and shape understanding. Existing approaches mainly extract deformation components in local regions at a similar scale while deformations of real-world objects are usually distributed in a multi-scale manner. In this article, we propose a novel method to exact multiscale deformation components automatically with a stacked attention-based autoencoder. The attention mechanism is designed to learn to softly weight multi-scale deformation components in active deformation regions, and the stacked attention-based autoencoder is learned to represent the deformation components at different scales. Quantitative and qualitative evaluations show that our method outperforms state-of-the-art methods. Furthermore, with the multiscale deformation components extracted by our method, the user can edit shapes in a coarse-to-fine fashion which facilitates effective modeling of new shapes.
Jie Yang 0038, Lin Gao 0004, Qingyang Tan, Yihua Huang 0002, Shihong Xia, Yukun Lai
IEEE Trans. Vis. Comput. Graph.4
2022 StylizedNeRF: Consistent 3D Scene Stylization as Stylized NeRF via 2D-3D Mutual Learning
abstract
3D scene stylization aims at generating stylized images of the scene from arbitrary novel views following a given set of style examples, while ensuring consistency when rendered from different views. Directly applying methods for image or video stylization to 3D scenes cannot achieve such consistency. Thanks to recently proposed neural radiance fields (NeRF), we are able to represent a 3D scene in a consistent way. Consistent 3D scene stylization can be effectively achieved by stylizing the corresponding NeRF. However, there is a significant domain gap between style examples which are 2D images and NeRF which is an implicit volumetric representation. To address this problem, we propose a novel mutual learning framework for 3D scene stylization that combines a 2D image stylization network and NeRF to fuse the stylization ability of 2D stylization network with the 3D consistency of NeRF. We first pre-train a standard NeRF of the 3D scene to be stylized and replace its color prediction module with a style network to obtain a stylized NeRF. It is followed by distilling the prior knowledge of spatial consistency from NeRF to the 2D stylization network through an introduced consistency loss. We also introduce a mimic loss to supervise the mutual learning of the NeRF style module and fine-tune the 2D stylization decoder. In order to further make our model handle ambiguities of 2D stylization results, we introduce learnable latent codes that obey the probability distributions conditioned on the style. They are attached to training samples as conditional inputs to better learn the style module in our novel stylized NeRF. Experimental results demonstrate that our method is superior to existing approaches in both visual quality and long-range consistency.
Yihua Huang 0002, Yu-Jie Yuan, Yukun Lai, Lin Gao 0004
CVPR1