Hayato Nishioka

dblp:227/5539 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-8323-8927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Interpretable Visualization of Expertise-Dependent Motor Skills Toward Supporting Piano Practice
abstract
The quality of piano performance depends on nuanced timing, articulation, and dynamic control, but practice feedback is often summary-based and hard to act on. We introduce Profy, a weakly supervised system that learns from take-level labels derived from aggregated listener ratings (expert-labeled vs. amateur-labeled) to produce time-aligned highlights for review during piano practice. We collected synchronized 1 kHz key-motion and audio from 73 pianists and used 1,083 valid takes for modeling and evaluation. The model outputs clip-level predictions together with evidence scores on a shared resampled model time base for visualization. On 20 amateur clips from short technique studies annotated by 21 expert pianists, the displayed highlight score aligns with passages that expert pianists marked for review despite training without localized labels (Pearson r=0.61, ROC-AUC 0.75). Rather than summarizing a take with a single global score, Profy helps learners decide where to inspect next by supporting scrubbing, looping, and focused replay of time-localized passages associated with expert-amateur differences.
Kazuki Kawamura 0001, Fujiki Nakamura, Hayato Nishioka, Momoko Shioki, Shinichi Furuya, Jun Rekimoto
DIS3
2026 Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key Motion
abstract
Detailed measurements of piano key motion capture touch, timing, and dynamic control, providing crucial performance insights. Such expressive gestures are overlooked in MIDI, which only records pitch onset, duration, and velocity. Here, we introduce a novel transcription technique that directly maps audio from expressive piano performance to continuous piano key motion. User studies reveal a preference to the transcribed key motion trajectories over MIDI in representing sound, and over 80% accuracy in matching transcribed trajectories to audio from contrasting piano expressions. Follow-up interviews further indicate that the visualised trajectories can reveal subtle performance nuances and provide actionable guidance for both teaching and practice. An interface example for pedagogy and performance analysis utilising our technique is also illustrated. By providing a physically grounded performance representation that musicians can interpret and act upon, this work establishes a foundation for future interactive tools in music pedagogy, performance feedback, and embodied musical learning.
Jingjing Tang 0002, Shinichi Furuya, Hayato Nishioka, Momoko Shioki, Geraint A. Wiggins, György Fazekas, Vincent K. M. Cheung
CHI3
2025 LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
abstract
Research in music understanding has extensively explored composition-level attributes such as key, genre, and instrumentation through advanced representations, leading to cross-modal applications using large language models. However, aspects of musical performance such as stylistic expression and technique remain underexplored, along with the potential of using large language models to enhance educational outcomes with customized feedback. To bridge this gap, we introduce LLaQo, a Large Language Query-based music coach that leverages audio language modeling to provide detailed and formative assessments of music performances. We also introduce instruction-tuned query-response datasets that cover a variety of performance dimensions from pitch accuracy to articulation, as well as contextual performance understanding (such as difficulty and performance techniques). Utilizing AudioMAE encoder and Vicuna-7b LLM backend, our model achieved state-of-the-art (SOTA) results in predicting teachers’ performance ratings, as well as in identifying piece difficulty and playing techniques. Textual responses from LLaQo was moreover rated significantly higher compared to other baseline models in a user study using audio-text matching. Our proposed model can thus provide informative answers to open-ended questions related to musical performance from audio data.
Huan Zhang 0001, Vincent K. M. Cheung, Hayato Nishioka, Simon Dixon, Shinichi Furuya
ICASSP3
2023 PianoSyncAR: Enhancing Piano Learning through Visualizing Synchronized Hand Pose Discrepancies in Augmented Reality
abstract
Motor skill acquisition involves learning from spatiotemporal discrepancies between target and self-generated motions. However, in dexterous skills with numerous degrees of freedom, understanding and correcting these motor errors are challenging. This issue becomes crucial for experienced individuals who seek for mastering and sophisticating their skills, where even subtle errors need to be minimized. To enable efficient optimization of body posture in piano learning, we present PianoSyncAR, an augmented reality system that superimposes the time-varying complex hand postures of a teacher over the hand of a learner. Through a user study with 12 pianists, we demonstrate several advantages of the proposed system over conventional tablet-screen, which implicate the potential of AR training as a complementary tool for video-based skill learning in piano playing.
Ruofan Liu 0001, Erwin Wu, Chen-Chieh Liao, Hayato Nishioka, Shinichi Furuya, Hideki Koike
ISMAR4
2023 Marker-removal Networks to Collect Precise 3D Hand Data for RGB-based Estimation and its Application in Piano
abstract
Hand pose analysis is a key step to understanding dexterous hand performances of many high-level skills, such as playing the piano. Currently, most accurate hand tracking systems are using fabric-/marker-based sensing that potentially disturbs users’ performance. On the other hand, markerless computer vision-based methods rely on a precise bare-hand dataset for training, which is difficult to obtain. In this paper, we collect a large-scale high precision 3D hand pose dataset with a small workload using a marker-removal network (MR-Net). The proposed MR-Net translates the marked-hand images to realistic bare-hand images, and the corresponding 3D postures are captured by a motion capture thus few manual annotations are required. A baseline estimation network PiaNet is introduced and we report the accuracy of various metrics together with a blind qualitative test to show the practical effect.
Erwin Wu, Hayato Nishioka, Shinichi Furuya, Hideki Koike
WACV2
2018 Augmented jump: a backpack multirotor system for jumping ability augmentation
abstract
This paper introduces Augmented Jump, a backpack multirotor system for jumping ability augmentation. Augmented Jump hovers and supports users' weight by a constant upward power of thrust. Users can jump higher and stay in the air for a longer time than usual with Augmented Jump. We designed and developed our first proof-of-concept prototype that can be controlled as an octocopter and support user's weight by 50kg at maximum. In our experiments, it is found that the system enabled the user to perform jumping in simulated 75% reduced gravity. From user study, the results showed that our system was effective for extending the height and the duration of jumping.
Takumi Takahashi, Keisuke Shiro, Akira Matsuda, Ryo Komiyama, Hayato Nishioka, Kazunori Hori, Yoshio Ishiguro, Takashi Miyaki, Jun Rekimoto
UbiComp5