Yichen Peng

dblp:239/9307 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-8544-3905ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Sensing Your Vocals: Exploring the Activity of Vocal Cord Muscles for Pitch Assessment Using Electromyography and Ultrasonography
abstract
Vocal training is difficult because the muscles that control pitch, resonance, and phonation are internal and invisible to learners. This paper investigates how Electromyography (EMG) and ultrasonic imaging (UI) can make these muscles observable for training purposes. We report three studies. First, we analyze the EMG and UI data from 16 singers (beginners, experienced & professionals), revealing differences among three vocal groups of the muscle control proficiency. Second, we use the collected data to create a system that visualizes an expert’s muscle activity as reference. This system is tested in a user study with 12 novices, showing that EMG highlighted muscle activation nuances, while UI provided insights into vocal cord length and dynamics. Third, to compare our approach to traditional methods (audio analysis and coach instructions), we conducted a focus group study with 15 experienced singers. Our results suggest that EMG is promising for improving vocal skill development and enhancing feedback systems. We conclude the paper with a detailed comparison of the analyzed modalities (EMG, UI and traditional methods), resulting in recommendations to improve vocal muscle training systems.
Kanyu Chen, Rebecca Panskus, Erwin Wu, Yichen Peng, Daichi Saito, Emiko Kamiyama, Ruiteng Li, Chen-Chieh Liao, Karola Marky, Kato Akira, Hideki Koike, Kai Kunze
CHI4
2026 SoleCoach: Sole Pressure and IMU-based MLLMs for Skill Coaching
abstract
In sports training, individualized skill assessment and feedback are essential for athletes to master complex movements and enhance performance. Existing approaches for generating coaching comments primarily rely on externally captured pose information, which limits their applicability in outdoor sports such as skiing that involve large-scale movement. To address this challenge, we propose a method for presenting athletes’ postures and generating coaching feedback solely based on foot pressure and IMU data collected from insole sensors. In our approach, a large language model directly interprets foot pressure signals to provide actionable coaching, thereby supporting independent practice. Through model evaluation and user studies, we demonstrate that the proposed method generates expert-level feedback and outperforms pose-based approaches. Furthermore, the user study shows that the feedback helps athletes identify body parts requiring correction and enhances their motivation for training.
Toshihiro Hirano, Hitoshi Yoshihara, Yichen Peng, Chen-Chieh Liao, Erwin Wu, Hideki Koike
CHI3
2026 MuscleGolfAR: Embodied versus Detached Visualization for Motions and Inferred Muscle Activations in Augmented Reality
abstract
Precise kinematic and dynamic representations are equally critical for fine-grained motor skills such as golf, yet the latter remains underexplored. This work introduces MuscleGolfAR, an augmented reality (AR) training system integrating swing motions and muscle activities. To support cost-effective inference of muscle activations, we construct a multimodal dataset encompassing posture, electromyography (EMG), and plantar pressure. The system displays inferred EMG through embodied (first-person perspective) and detached (third-person perspective) visualizations. User studies are subsequently conducted to evaluate the impact of these strategies on training effectiveness. Results reveal that embodied visualizations enhance ownership over augmented feedback, while detached visualizations facilitate a holistic comprehension of the whole body. Furthermore, individuals' preferences for these strategies correlate with their practice habits and skill proficiencies, offering new insights for the design of AR-based motor skill training systems.
Ruofan Liu 0001, Chen-Chieh Liao, Takuya Takahashi, Yichen Peng, Erwin Wu, Hideki Koike
IEEE Trans. Vis. Comput. Graph.4
2025 PiaMuscle: Improving Piano Skill Acquisition by Cost-effectively Estimating and Visualizing Activities of Miniature Hand Muscles
Ruofan Liu 0001, Yichen Peng, Takanori Oku, Chen-Chieh Liao, Erwin Wu, Shinichi Furuya, Hideki Koike
CHI2
2025 LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model
abstract
Image rendering from line drawings is vital in design and image generation technologies reduce costs, yet professional line drawings demand preserving complex details. Text prompts struggle with accuracy, and image translation struggles with consistency and fine-grained control. We present LineArt, a framework that transfers complex appearance onto detailed design drawings, facilitating design and artistic creation. It generates high-fidelity appearance while preserving structural accuracy by simulating hierarchical visual cognition and integrating human artistic experience to guide the diffusion process. LineArt overcomes the limitations of current methods in terms of difficulty in fine-grained control and style degradation in design drawings. It requires no precise 3D modeling, physical property specifications, or network training, making it more convenient for design tasks. LineArt consists of two stages: a multi-frequency lines fusion module to supplement the input design drawing with detailed structural information and a two-part painting process for Base Layer Shaping and Surface Layer Coloring. We also present a new design drawing dataset, ProLines, for evaluation. The experiments show that LineArt performs better in accuracy, realism, and material precision compared to SOTAs. Project page: https://meaoxixi.github.io/LineArt/.
Hongzhen Li, Yichen Peng, Haoran Xie 0002, Xi Yang 0017
CVPR4
2025 From Pose to Muscle: Multimodal Learning for Piano Hand Muscle Electromyography
abstract
Muscle coordination is fundamental when humans interact with the world. Reliable estimation of hand muscle engagement can serve as a source of internal feedback, supporting the development of embodied intelligence and the acquisition of dexterous skills. However, contemporary electromyography (EMG) sensing techniques either require prohibitively expensive devices or are constrained to gross motor movements, which inherently involve large muscles. On the other hand, EMGs exhibit dependency on individual anatomical variability and task-specific contexts, resulting in limited generalization. In this work, we preliminarily investigate the latent pose-EMG correspondence using a general EMG gesture dataset. We further introduce a multimodal dataset, PianoKPM Dataset, and a hand muscle estimation framework, PianoKPM Net, to facilitate high-fidelity EMG inference. Subsequently, our approach is compared against reproducible competitive baselines. The generalization and adaptation across unseen users and tasks are evaluated by quantifying the training set scale and the included data amount.
Ruofan Liu 0001, Yichen Peng, Takanori Oku, Chen-Chieh Liao, Erwin Wu, Shinichi Furuya, Hideki Koike
NeurIPS2
2024 EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
abstract
We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME), a new mesh-level holistic co-speech dataset. BEAT2 combines a MoShed SMPL-X body with FLAME head parameters and further refines the modeling of head, neck, and finger movements, offering a community-standardized, high-quality 3D motion captured dataset. EMAGE leverages masked body gesture priors during training to boost inference performance. It involves a Masked Audio Gesture Transformer, facilitating joint training on audio-to-gesture generation and masked gesture reconstruction to effectively encode audio and body gesture hints. Encoded body hints from masked gestures are then separately employed to generate facial and body movements. Moreover, EMAGE adaptively merges speech features from the audio's rhythm and content and utilizes four compositional VQ-VAEs to enhance the results' fidelity and diversity. Experiments demonstrate that EMAGE generates holistic gestures with state-of-the-art performance and is flexible in accepting predefined spatial-temporal gesture inputs, generating complete, audio-synchronized results. Our code and dataset are available.1
Giorgio Becherini, Yichen Peng, Mingyang Su, Xuefei Zhe, Naoya Iwamoto, Michael J. Black
CVPR4
2022 BEAT: A Large-Scale Semantic and Emotional Multi-modal Dataset for Conversational Gestures Synthesis
Naoya Iwamoto, Yichen Peng, Zhengqing Li, Elif Bozkurt
ECCV (7)4
2022 DualFace: Two-stage drawing guidance for freehand portrait sketching
abstract
Special skills are required in portrait painting, such as imagining geometric structures and facial detail for final portrait designs. This makes it a difficult task for users, especially novices without prior artistic training, to draw freehand portraits with high-quality details. In this paper, we propose dualFace, a portrait drawing interface to assist users with different levels of drawing skills to complete recognizable and authentic face sketches. Inspired by traditional artist workflows for portrait drawing, dualFace gives two-stages of drawing assistance to provide global and local visual guidance. The former helps users draw contour lines for portraits (i.e., geometric structure), and the latter helps users draw details of facial parts, which conform to the user-drawn contour lines. In the global guidance stage, the user draws several contour lines, and dualFace then searches for several relevant images from an internal database and displays the suggested face contour lines on the background of the canvas. In the local guidance stage, we synthesize detailed portrait images with a deep generative model from user-drawn contour lines, and then use the synthesized results as detailed drawing guidance. We conducted a user study to verify the effectiveness of dualFace, which confirms that dualFace significantly helps users to produce a detailed portrait sketch.
Yichen Peng, Tomohiro Hibino, Chunqi Zhao, Haoran Xie 0002, Tsukasa Fukusato, Kazunori Miyata
Comput. Vis. Media2
2021 Two-Stage Motion Editing Interface for Character Animation
abstract
In this paper, we propose a user interface that enables users to intuitively retrieve relevant motions from a database and edit them by drawing motion trajectories on the screen. This system consists of two-stage operations to provide global-level and local-level motion editing: a global stage that enables users to design the body movement in virtual space roughly, and a local stage that enables users to design detailed movements such as limbs movement. We verified the proposed system with character animation editing with both global and local stages.
Yichen Peng, Chunqi Zhao, Tsukasa Fukusato, Haoran Xie 0002, Kazunori Miyata
SCA1