Yuting He 0006

dblp:167/1989-6 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0003-1954-952XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HoloMobile: Photorealistic Avatar on Mobile via Compact Dynamic Gaussians from a Monocular Video
abstract
Human avatars, defined as animatable 3D digital people, have diverse applications, including teleconferencing, motion guidance, remote care, and gaming. With the widespread adoption of smart mobile devices, the demand for real-time viewing of dynamic digital humans on mobile platforms has grown significantly. However, existing avatar reconstruction methods, such as MagicStream and ExAvatar, struggle to produce high-fidelity avatars from monocular videos captured by mobile devices. Furthermore, most related research focuses on individual algorithmic components and does not provide an end-to-end pipeline that jointly addresses avatar reconstruction, 4D data transmission, and mobile rendering. In this work, we introduce HoloMobile, the first end-to-end mobile avatar system with two key capabilities: (1) High-fidelity avatar reconstruction from monocular videos captured on mobile phones, using a hybrid explicit-implicit model. (2) Efficient polynomial compression, transmission, and web-based rendering of dynamic 3D avatars on mobile devices. With avatar reconstruction performed on a PC or cloud with a single consumer-grade GPU, HoloMobile provides an end-to-end solution for mobile users: from data capture to streaming and real-time viewing. Evaluation on 20 self-captured avatars reveals that HoloMobile achieves an average PSNR of 30.36 dB, a compression rate of 99.57%, and over 200 FPS for mobile rendering, demonstrating superiority over existing baselines. The demo video of HoloMobile is available at https://youtu.be/kfOFV7AiP5A.
Yuting He 0006, Yihua Huang 0002, Xiaojuan Qi 0001, Zhenyu Yan 0002, Guoliang Xing
MobiSys1
2026 A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
abstract
Multimodal human action recognition (HAR) utilizes complementary data for activity classification. Built on traditional HAR tasks, recent advances in Large Language Models (LLMs) enable detailed descriptions and causal reasoning of human actions, advancing new tasks of human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially multimodal Large Vision-Language Models (LVLMs), struggle with modalities other than RGB images, like depth, IMU, ormmWave, due to a lack of large-scale datasets in these task domains. Existing HAR datasets provide only coarse-grained annotations, in-sufficient for depicting the detailed action dynamics required in HAU and HARn tasks. Simply combining annotations and generating captions with LLMs often lacks necessary logical and spatiotemporal consistency. In this paper, we introduce CUHK-X, a large-scale multi-modal dataset and benchmarks for HAR, HAU, and HARn. It includes 64,267 samples of 40 actions performed by 30 participants across two indoor environments, covering diverse daily scenarios. To address the challenge of spatiotemporal inconsistencies in captions, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences. CUHK-X also includes three benchmarks with six tasks to evaluate state-of-the-art models. Experimental results show average accuracies of 76.52% for HAR, 40.76% for HAU, and 70.25% for HARn. This large-scale multimodal dataset aims to empower the research community to apply, develop, and adapt data-intensive learning techniques for a wide range of human activity-related tasks.
Siyang Jiang, Mu Yuan, Bufang Yang, Lilin Xu, Yang Li 0147, Yuting He 0006, Liran Dong, Wenrui Lu, Zhenyu Yan 0002, Xiaofan Jiang 0001, Wei Gao 0006, Hongkai Chen 0001, Guoliang Xing
MobiSys8
2026 AniGen: Unified S3 Fields for Animatable 3D Asset Generation
abstract
Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodied agents, and animation production. While recent 3D generative models can synthesize visually plausible shapes from images, the results are typically static. Obtaining usable rigs via post-hoc auto-rigging is brittle and often produces skeletons that are topologically inconsistent with the generated geometry. We present AniGen , a unified framework that directly generates animate-ready 3D assets conditioned on a single image. Our key insight is to represent shape, skeleton, and skinning as mutually consistent S 3 Fields (Shape, Skeleton, Skin) defined over a shared spatial domain. To enable the robust learning of these fields, we introduce two technical innovations: (i) a confidence-decaying skeleton field that explicitly handles the geometric ambiguity of bone prediction at Voronoi boundaries, and (ii) a dual skin feature field that decouples skinning weights from specific joint counts, allowing a fixed-architecture network to predict rigs of arbitrary complexity. Built upon a two-stage flow-matching pipeline, AniGen first synthesizes a sparse structural scaffold and then generates dense geometry and articulation in a structured latent space. Extensive experiments demonstrate that AniGen substantially outperforms state-of-the-art sequential baselines in rig validity and animation quality, generalizing effectively to in-the-wild images across diverse categories including animals, humanoids, and machinery. Homepage : https://yihua7.github.io/AniGen_web/
Yihua Huang 0002, Zixin Zou, Yuting He 0006, Chirui Chang, Cheng-Feng Pu, Ziyi Yang 0008, Yan-Pei Cao 0001, Xiaojuan Qi 0001
ACM Trans. Graph.3
2025 Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training
abstract
In-home resistance training (RT) is a convenient and effective way to maintain health and well-being. However, incorrect exercise execution can result in unintended muscle engagement and an increased risk of injury. Without access to professional coaching, an accurate muscle-aware motion feedback system becomes essential for safe and effective training. However, existing visual language models (VLMs) struggle to provide accurate and effective muscle-aware movement guidance due to their limited understanding of RT motion and the absence of related expert knowledge. In this work, we introduce Myo-Trainer, the first vision-based muscle-aware motion feedback system that uses explicit muscle-aware motion analysis and domain-specific expert knowledge to provide corrective guidance on muscle engagement and movement execution. Also, we propose a novel DAGCN-Former network that integrates both spatial and temporal modeling capabilities to capture the complex dynamics of human RT motion. Experiments involving 26 subjects and 1000+ minutes of RT demonstrate that Myo-Trainer improves the accuracy of motion analysis by 17.22%, achieves a 2.5x reduced inference latency and a BertScore of 85.88% of generated feedback compared to those provided by experienced certified trainers, outperforming existing solutions. Additionally, Myo-Trainer received higher satisfaction ratings from participants compared to other AI trainers and video tutorials, highlighting its potential for real-world applications.
Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Bufang Yang, Siyang Jiang, Yihua Huang 0002, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001
MobiCom1
2024 Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance Training
abstract
Resistance training is widely incorporated in exercise programs, including in-home fitness and rehabilitation. However, improper motion patterns and muscle stimulation can undermine the safety of the subjects, making precise monitoring essential. Existing solutions primarily focus on correcting motion patterns with difficulties assessing muscle contraction levels. In this work, we introduce MyoTrainer, which provides muscle-aware motion descriptions and personalized feedback in natural language. Taking a person's exercise video as input, MyoTrainer first utilizes pose estimation models to capture motion sequences in real-time. A GCN-Former model has been developed for fine-grained motion analysis, which includes action recognition, incorrect movement pattern detection, and muscle contraction intensity estimation. Additionally, MyoTrainer integrates fitness and physiotherapeutic domain knowledge to deliver personalized, professional feedback. Extensive evaluations show that our system outperforms existing solutions in all recognition tasks and a survey indicates 88.9% of users find the generated feedback to be beneficial.
Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Di Duan, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001
SenSys1