VLDB 2026 Research / reviewers in the wild / expert
Erwin Wu
dblp:230/7572
· DBLP profile ↗
21ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-6723-2864ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 13 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sensing Your Vocals: Exploring the Activity of Vocal Cord Muscles for Pitch Assessment Using Electromyography and UltrasonographyabstractVocal training is difficult because the muscles that control pitch, resonance, and phonation are internal and invisible to learners. This paper investigates how Electromyography (EMG) and ultrasonic imaging (UI) can make these muscles observable for training purposes. We report three studies. First, we analyze the EMG and UI data from 16 singers (beginners, experienced & professionals), revealing differences among three vocal groups of the muscle control proficiency. Second, we use the collected data to create a system that visualizes an expert’s muscle activity as reference. This system is tested in a user study with 12 novices, showing that EMG highlighted muscle activation nuances, while UI provided insights into vocal cord length and dynamics. Third, to compare our approach to traditional methods (audio analysis and coach instructions), we conducted a focus group study with 15 experienced singers. Our results suggest that EMG is promising for improving vocal skill development and enhancing feedback systems. We conclude the paper with a detailed comparison of the analyzed modalities (EMG, UI and traditional methods), resulting in recommendations to improve vocal muscle training systems. Kanyu Chen, Rebecca Panskus, Erwin Wu, Yichen Peng, Daichi Saito, Emiko Kamiyama, Ruiteng Li, Chen-Chieh Liao, Karola Marky, Kato Akira, Hideki Koike, Kai Kunze |
CHI | 3 |
| 2026 | SoleCoach: Sole Pressure and IMU-based MLLMs for Skill CoachingabstractIn sports training, individualized skill assessment and feedback are essential for athletes to master complex movements and enhance performance. Existing approaches for generating coaching comments primarily rely on externally captured pose information, which limits their applicability in outdoor sports such as skiing that involve large-scale movement. To address this challenge, we propose a method for presenting athletes’ postures and generating coaching feedback solely based on foot pressure and IMU data collected from insole sensors. In our approach, a large language model directly interprets foot pressure signals to provide actionable coaching, thereby supporting independent practice. Through model evaluation and user studies, we demonstrate that the proposed method generates expert-level feedback and outperforms pose-based approaches. Furthermore, the user study shows that the feedback helps athletes identify body parts requiring correction and enhances their motivation for training. Toshihiro Hirano, Hitoshi Yoshihara, Yichen Peng, Chen-Chieh Liao, Erwin Wu, Hideki Koike |
CHI | 5 |
| 2026 | MuscleGolfAR: Embodied versus Detached Visualization for Motions and Inferred Muscle Activations in Augmented RealityabstractPrecise kinematic and dynamic representations are equally critical for fine-grained motor skills such as golf, yet the latter remains underexplored. This work introduces MuscleGolfAR, an augmented reality (AR) training system integrating swing motions and muscle activities. To support cost-effective inference of muscle activations, we construct a multimodal dataset encompassing posture, electromyography (EMG), and plantar pressure. The system displays inferred EMG through embodied (first-person perspective) and detached (third-person perspective) visualizations. User studies are subsequently conducted to evaluate the impact of these strategies on training effectiveness. Results reveal that embodied visualizations enhance ownership over augmented feedback, while detached visualizations facilitate a holistic comprehension of the whole body. Furthermore, individuals' preferences for these strategies correlate with their practice habits and skill proficiencies, offering new insights for the design of AR-based motor skill training systems. Ruofan Liu 0001, Chen-Chieh Liao, Takuya Takahashi, Yichen Peng, Erwin Wu, Hideki Koike |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | PiaMuscle: Improving Piano Skill Acquisition by Cost-effectively Estimating and Visualizing Activities of Miniature Hand Muscles
Ruofan Liu 0001, Yichen Peng, Takanori Oku, Chen-Chieh Liao, Erwin Wu, Shinichi Furuya, Hideki Koike |
CHI | 5 |
| 2025 | From Pose to Muscle: Multimodal Learning for Piano Hand Muscle ElectromyographyabstractMuscle coordination is fundamental when humans interact with the world. Reliable estimation of hand muscle engagement can serve as a source of internal feedback, supporting the development of embodied intelligence and the acquisition of dexterous skills. However, contemporary electromyography (EMG) sensing techniques either require prohibitively expensive devices or are constrained to gross motor movements, which inherently involve large muscles. On the other hand, EMGs exhibit dependency on individual anatomical variability and task-specific contexts, resulting in limited generalization. In this work, we preliminarily investigate the latent pose-EMG correspondence using a general EMG gesture dataset. We further introduce a multimodal dataset, PianoKPM Dataset, and a hand muscle estimation framework, PianoKPM Net, to facilitate high-fidelity EMG inference. Subsequently, our approach is compared against reproducible competitive baselines. The generalization and adaptation across unseen users and tasks are evaluated by quantifying the training set scale and the included data amount. Ruofan Liu 0001, Yichen Peng, Takanori Oku, Chen-Chieh Liao, Erwin Wu, Shinichi Furuya, Hideki Koike |
NeurIPS | 5 |
| 2025 | ColorizeDiffusion: Improving Reference-Based Sketch Colorization with Latent Diffusion ModelabstractDiffusion models have achieved great success in dual-conditioned image generation. However, they still face significant challenges in image-guided sketch colorization, where reference and sketch images usually exhibit different semantics and spatial structures. This mismatch, termed “distribution shift” in this peper, results in various artifacts and degrades the colorization quality. To address this issue, we conducted thorough investigations into the image-prompted latent diffusion model and developed a two-stage training framework to mitigate the effects of distribution shift based on our analysis. Comprehensive quantitative comparisons, qualitative evaluations, and user studies were performed to demonstrate the superiority of our proposed methods. Additionally, ablation studies were conducted to assess the impact of the distribution shift and the selection of reference embeddings. Codes are made publicly available at https://github.com/tellurion-kanata/colorizeDiffusion. Dingkun Yan, Erwin Wu, Yuma Nishioka, Issei Fujishiro, Suguru Saito |
WACV | 3 |
| 2024 | SolePoser: Full Body Pose Estimation using a Single Pair of Insole SensorabstractWe propose SolePoser, a real-time 3D pose estimation system that leverages only a single pair of insole sensors. Unlike conventional methods relying on fixed cameras or bulky wearable sensors, our approach offers minimal and natural setup requirements. The proposed system utilizes pressure and IMU sensors embedded in insoles to capture the body weight’s pressure distribution at the feet and its 6 DoF acceleration. This information is used to estimate the 3D full-body joint position by a two-stream transformer network. A novel double-cycle consistency loss and a cross-attention module are further introduced to learn the relationship between 3D foot positions and their pressure distributions. We also introduced two different datasets of sports and daily exercises, offering 908k frames across eight different activities. Our experiments show that our method’s performance is on par with top-performing approaches, which utilize more IMUs and even outperform third-person-view camera-based methods in certain scenarios. Erwin Wu, Rawal Khirodkar, Hideki Koike, Kris Makoto Kitani |
UIST | 1 |
| 2024 | ARpenSki: Augmenting Ski Training with Direct and Indirect Postural VisualizationabstractAlpine skiing is a popular winter sport, and several systems have been proposed to enhance training and improve efficiency. However, many existing systems rely on simulation-based environments, which suffer from drawbacks such as a gap between real skiing and the lack of body ownership. To address these limitations, we present ARpenSki, a novel augmented reality (AR) ski training system that employs a see-through head mounted display (HMD) to deliver augmented visual training cues that may be applied on real slopes. The proposed AR system provides a transparent view of the lower half of the field of vision, where we implemented three different AR-based direct and indirect postural visualization methods. We conducted an user study to investigate the influence of different visual cues in the AR environment. Our results indicate that a simple AR visualization of the user’s spine (Figure 1.2) yields the most favorable training performance, surpassing conventional visualizations by 7% improvement in the user’s posture. Building upon these promising findings, we further tested our system on real slopes and showed the potential of a real AR skiing application. Erwin Wu, Chen-Chieh Liao, Hideki Koike |
VR | 2 |
| 2023 | OmniSense: Exploring Novel Input Sensing and Interaction Techniques on Mobile Device with an Omni-Directional CameraabstractAn omni-directional (360°) camera captures the entire viewing sphere surrounding its optical center. Such cameras are growing in use to create highly immersive content and viewing experiences. When such a camera is held by a user, the view includes the user’s hand grip, finger, body pose, face, and the surrounding environment, providing a complete understanding of the visual world and context around it. This capability opens up numerous possibilities for rich mobile input sensing. In OmniSense, we explore the broad input design space for mobile devices with a built-in omni-directional camera and broadly categorize them into three sensing pillars: i) near device ii) around device and iii) surrounding device. In addition we explore potential use cases and applications that leverage these sensing capabilities to solve user needs. Following this, we develop a working system to put these concepts into action, by leveraging these sensing capabilities to enable potential use cases and applications. We studied the system in a technical evaluation and a preliminary user study to gain initial feedback and insights. Collectively these techniques illustrate how a single, omni-purpose sensor on a mobile device affords many compelling ways to enable expressive input, while also affording a broad range of novel applications that improve user experience during mobile interaction. Hui-Shyong Yeo, Erwin Wu, Daehwa Kim, Hyungil Kim, Seoyoung Oh, Luna Takagi, Woontack Woo, Hideki Koike, Aaron J. Quigley |
CHI | 2 |
| 2023 | PianoSyncAR: Enhancing Piano Learning through Visualizing Synchronized Hand Pose Discrepancies in Augmented RealityabstractMotor skill acquisition involves learning from spatiotemporal discrepancies between target and self-generated motions. However, in dexterous skills with numerous degrees of freedom, understanding and correcting these motor errors are challenging. This issue becomes crucial for experienced individuals who seek for mastering and sophisticating their skills, where even subtle errors need to be minimized. To enable efficient optimization of body posture in piano learning, we present PianoSyncAR, an augmented reality system that superimposes the time-varying complex hand postures of a teacher over the hand of a learner. Through a user study with 12 pianists, we demonstrate several advantages of the proposed system over conventional tablet-screen, which implicate the potential of AR training as a complementary tool for video-based skill learning in piano playing. Ruofan Liu 0001, Erwin Wu, Chen-Chieh Liao, Hayato Nishioka, Shinichi Furuya, Hideki Koike |
ISMAR | 2 |
| 2023 | Marker-removal Networks to Collect Precise 3D Hand Data for RGB-based Estimation and its Application in PianoabstractHand pose analysis is a key step to understanding dexterous hand performances of many high-level skills, such as playing the piano. Currently, most accurate hand tracking systems are using fabric-/marker-based sensing that potentially disturbs users’ performance. On the other hand, markerless computer vision-based methods rely on a precise bare-hand dataset for training, which is difficult to obtain. In this paper, we collect a large-scale high precision 3D hand pose dataset with a small workload using a marker-removal network (MR-Net). The proposed MR-Net translates the marked-hand images to realistic bare-hand images, and the corresponding 3D postures are captured by a motion capture thus few manual annotations are required. A baseline estimation network PiaNet is introduced and we report the accuracy of various metrics together with a blind qualitative test to show the practical effect. Erwin Wu, Hayato Nishioka, Shinichi Furuya, Hideki Koike |
WACV | 1 |
| 2022 | A Distance Learning System With Shareable Physical Information For Ski TrainingabstractDistance learning for skill learning is still inadequate because the perceptual information provided to the user is limited. This study proposed a framework for a distance learning system with real-time feedback to share physical information between the student and the teacher. As an initial trial to use this framework, we developed a prototype of a distance learning system for skiing including visual feedback, and verified its operation. The result indicated the system could be applied enough to a distance learning system. Shigeharu Ono, Hideaki Kanai, Erwin Wu, Hideki Koike |
VRST | 3 |
| 2021 | SPinPong - Virtual Reality Table Tennis Skill Acquisition using Visual, Haptic and Temporal CuesabstractLearning an advanced skill in sports requires a huge amount of practice and players also have to overcome both physical difficulties and the dullness of repetitive training. Returning a fast spin shot in table tennis could be taken as an example, as athletes need to judge the spin type and decide the racket pose within a second, which is difficult for beginners. Therefore, in this paper, we show how to design an intuitive training system to acquire this specific skill using different cues in Virtual Reality (VR). Using VR, we can easily provide visual information, attach haptic devices, and distort the speed of time, however, it is difficult to decide which types of information could benefit the training. In an initial study, by comparing real world training with VR training, we showed the effect of VR training and obtained some insights about augmentation for training spin shots. The training system was then improved by adding three new conditions using different visualizations and temporal distortions, as well as a haptic racket for creating realistic feedback. Finally, we performed a detailed experiment, which suggest a significant improvement of skill for each condition compared to the baseline, while a qualitative evaluation indicates that both users' motivation and their understanding of spin are increased by using our system. Erwin Wu, Mitski Piekenbrock, Takuto Nakamura, Hideki Koike |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Back-Hand-Pose: 3D Hand Pose Estimation for a Wrist-worn Camera via Dorsum Deformation NetworkabstractThe automatic recognition of how people use their hands and fingers in natural settings -- without instrumenting the fingers -- can be useful for many mobile computing applications. To achieve such an interface, we propose a vision-based 3D hand pose estimation framework using a wrist-worn camera. The main challenge is the oblique angle of the wrist-worn camera, which makes the fingers scarcely visible. To address this, a special network that observes deformations on the back of the hand is required. We introduce DorsalNet, a two-stream convolutional neural network to regress finger joint angles from spatio-temporal features of the dorsal hand region (the movement of bones, muscle, and tendons). This work is the first vision-based real-time 3D hand pose estimator using visual features from the dorsal hand region. Our system achieves a mean joint-angle error of 8.81 degree for user-specific models and 9.77 degree for a general model. Further evaluation shows that our system outperforms previous work with an average of 20% higher accuracy in recognizing dynamic gestures, and achieves a 75% accuracy of detecting 11 different grasp types. We also demonstrate 3 applications which employ our system as a control device, an input device, and a grasped object recognizer. Erwin Wu, Ye Yuan 0007, Hui-Shyong Yeo, Aaron J. Quigley, Hideki Koike, Kris Makoto Kitani |
UIST | 1 |
| 2019 | OmniGlobe: An Interactive I/O System For Symmetric 360-Degree Video CommunicationabstractVideo communication systems have been suffered from the narrow field of view. To solve this limitation, one study proposed symmetric 360° video communication system by combining an omnidirectional camera and a hemispherical display. However, the system still had several issues, e.g., the invisibility of hemisphere which was at the opposite side from a user caused the inconvenience of observing the remote environment. To solve these issues, we introduce OmniGlobe, a novel symmetric full 360° video communication system which incorporates an omnidirectional camera, a full spherical display, and several visual or interactive techniques. Based on an experiment, we could indicate that our system is effective in reducing the inconvenience of observing the remote environment and increased the remote space awareness and user's gaze awareness to support remote collaboration. We also discuss the takeaways, limitations and application areas in our system which help improve the system. Zhengqing Li, Shio Miyafuji, Erwin Wu, Hideaki Kuzuoka, Naomi Yamashita, Hideki Koike |
Conference on Designing Interactive Systems | 3 |
| 2019 | Opisthenar: Hand Poses and Finger Tapping Recognition by Observing Back of Hand Using Embedded Wrist CameraabstractWe introduce a vision-based technique to recognize static hand poses and dynamic finger tapping gestures. Our approach employs a camera on the wrist, with a view of the opisthenar (back of the hand) area. We envisage such cameras being included in a wrist-worn device such as a smartwatch, fitness tracker or wristband. Indeed, selected off-the-shelf smartwatches now incorporate a built-in camera on the side for photography purposes. However, in this configuration, the fingers are occluded from the view of the camera. The oblique angle and placement of the camera make typical vision-based techniques difficult to adopt. Our alternative approach observes small movements and changes in the shape, tendons, skin and bones on the opisthenar area. We train deep neural networks to recognize both hand poses and dynamic finger tapping gestures. While this is a challenging configuration for sensing, we tested the recognition with a real-time user test and achieved a high recognition rate of 89.4% (static poses) and 67.5% (dynamic gestures). Our results further demonstrate that our approach can generalize across sessions and to new users. Namely, users can remove and replace the wrist-worn device while new users can employ a previously trained system, to a certain degree. We conclude by demonstrating three applications and suggest future avenues of work based on sensing the back of the hand. Hui-Shyong Yeo, Erwin Wu, Aaron J. Quigley, Hideki Koike |
UIST | 2 |
| 2019 | VR Ski Coach: Indoor Ski Training System Visualizing Difference from Leading SkierabstractThe training of skiing is difficult because of environmental requirements and teaching methods. Therefore, we propose a virtual reality ski training system using an indoor ski simulator. The system is based on a simple indoor ski simulator with two trackers to capture the motion of skis. Users can control the skis in the virtual ski slope we provided and train their skills with a replay of a professional skier. The training system consists of three modules: a coach replay system for reviewing pro-skiers's motion; a time control system that can be used to watch the detailed motion of both the coach and the user; and a visualization of the angle of the skis to compare the difference of motions between the users and the coach. Takayuki Nozawa, Erwin Wu, Hideki Koike |
VR | 2 |
| 2019 | Real-time Human Motion Forecasting using a RGB CameraabstractThis paper propose a real-time human motion forecasting system which visualize the future pose in virtual reality using a RGB camera. Our system consists of three parts: 2D pose estimation from RGB frames using a residual neural network, 2D pose forecasting using a recurrent neural network, and 3D recovery from the predicted 2D pose using a residual linear network. To improve the prediction learning quantity of temporal feature, we propose a special method using lattice optical flow for the joints movement estimation. After fitting the skeleton, a predicted 3d model of target human will be built 0.5s in advance in a 30-fps video. Erwin Wu, Hideki Koike |
VR | 1 |
| 2019 | FuturePose - Mixed Reality Martial Arts Training Using Real-Time 3D Human Pose Forecasting With a RGB CameraabstractIn this paper, we propose a novel mixed reality martial arts training system using deep learning based real-time human pose forecasting. Our training system is based on 3D pose estimation using a residual neural network with input from a RGB camera, which captures the motion of a trainer. The student wearing a head mounted display can see the virtual model of the trainer and his forecasted future pose. The pose forecasting is based on recurrent networks, to improve the learning quantity of the motion's temporal feature, we use a special lattice optical flow method for the joints movement estimation. We visualize the real-time human motion by a generated human model while the forecasted pose is shown by a red skeleton model. In our experiments, we evaluated the performance of our system when predicting 15 frames ahead in a 30-fps video (0.5s forecasting), the accuracies were acceptable since they are equal to or even outperforms some methods using depth IR cameras or fabric technologies, user studies showed that our system is helpful for beginners to understand martial arts and the usability is comfortable since the motions were captured by RGB camera. Erwin Wu, Hideki Koike |
WACV | 1 |
| 2018 | OmniEyeball: An Interactive I/O Device For 360-Degree Video CommunicationabstractWe propose OmniEyeball (OEB), which is a novel interactive 360° image I/O system combining a spherical display system with an omnidirectional camera. We also present our experimental design of a user interface on the OEB, including a vision-based touch detection technique as well as several visual and interactive features. Our proposed techniques may contribute to solving the weak awareness of the opposite side of the spherical display as well as the workload caused by walking around in the 360° symmetric video communication. Zhengqing Li, Shio Miyafuji, Erwin Wu, Toshiki Sato, Hideaki Kuzuoka, Hideki Koike |
ISS | 3 |
| 2018 | Real-time human motion forecasting using a RGB cameraabstractWe propose a real-time human motion forecasting system which visualize the future pose in virtual reality using a RGB camera. Our system consists of three parts: 2D pose estimation from RGB frames using a residual neural network, 2D pose forecasting using a recurrent neural network, and 3D recovery from the predicted 2D pose using a residual linear network. To improve the prediction learning quantity of temporal feature, we propose a special method using lattice optical flow for the joints movement estimation. After fitting the skeleton, a predicted 3d model of target human will be built 0.5s in advance in a 30-fps video. Erwin Wu, Hideki Koike |
VRST | 1 |