Masatoshi Hamanaka

dblp:65/2942 · DBLP profile ↗
← Back
16ranked-venue papers
11as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 AI-Based Composition Tools for Composing A School Song with Student Participation
Masatoshi Hamanaka
MMM (4)1
2025 BandNaviHD: Band-Member Backtrack Interface Based on Member History Information
abstract
A band-member-backtracking application called BandNaviHD has been developed that enables a user to discover new songs and bands by tracing musicians who have played in different bands. Previous similarity-based song recommendation systems only retrieve similar songs, but many retrieved results are not songs by musicians the user likes. BandNaviHD enables a user to discover other bands in which a musician has played.
Masatoshi Hamanaka
CBMI1
2025 Demonstration of Visualizer for Beats and Scratches of Breaking DJ Performances
Masatoshi Hamanaka
ICEC1
2025 Automatic Fingering Saxophone Quartet System
Gou Koutaki, Masatoshi Hamanaka
ICEC2
2025 Implementation of Visualizer for Beats and Scratches
Masatoshi Hamanaka
ACM Multimedia1
2025 RoboSax Melody Slot Machine
abstract
RoboSax Melody Slot Machine is a mobile-to-acoustic performance system that connects an interactive tablet interface to live saxophone performance. Participants select melodic fragments on an iPad-based Melody Slot Machine, where related musical phrases are generated through GTTM-based melody morphing rather than arbitrary recombination. The selected melody is transmitted as MIDI to RoboSax, a robotic saxophone mechanism that actuates the instrument’s fingerings, while a human performer provides breath, tonguing, phrasing, dynamics, and tone color. This division of roles allows participants without instrumental training to influence the musical structure of a live performance while preserving the embodied expressiveness of acoustic wind performance. For SIGGRAPH Appy Hour, the work presents a participatory experience in which mobile interaction becomes immediately audible as live acoustic sound. By combining structurally coherent melody variation, audience-controlled interaction, and human–robot performance, RoboSax Melody Slot Machine explores how mobile music apps can extend beyond the screen and become part of a shared performative space.
Masatoshi Hamanaka, Gou Koutaki
ACM Multimedia1
2025 Real-Time Visualizer for Turntablist Performance
Masatoshi Hamanaka
MMM (5)1
2025 SyncViolinist: Music-Oriented Violin Motion Generation Based on Bowing and Fingering
abstract
Automatically generating realistic musical performance motion can greatly enhance digital media production, often involving collaboration between professionals and musicians. However, capturing the intricate body, hand, and finger movements required for accurate musical performances is challenging. Existing methods often fall short due to the complex mapping between audio and motion, typically requiring additional inputs like scores or MIDI data. In this work, we present SyncViolinist, a multi-stage end-to-end framework that generates synchronized violin performance motion solely from audio input. Our method overcomes the challenge of capturing both global and finegrained performance features through two key modules: a bowing/fingering module and a motion generation module. The bowing/fingering module extracts detailed playing information from the audio, which the motion generation module uses to create precise, coordinated body motions reflecting the temporal granularity and nature of the violin performance. We demonstrate the effectiveness of SyncViolinist with significantly improved qualitative and quantitative results from unseen violin performance audio, outperforming state-of-the-art methods. Extensive subjective evaluations involving professional violinists further validate our approach. The code and dataset are available at https://github.com/Kakanat/SyncViolinist.
Hiroki Nishizawa, Keitaro Tanaka, Asuka Hirata, Shugo Yamaguchi, Masatoshi Hamanaka, Shigeo Morishima
WACV6
2024 Music Scope Pad: Video Selecting Application by Natural Movement in VR Space
abstract
This paper describes Music Scope Pad, a novel video selecting application that enables us to select videos without having to click a mouse or touch a screen. Existing video players enable us to see and hear only one video at a time, and thus we have to sample videos individually to select the video we want to watch from numerous new music videos, which involves a large number of mouse and screen-touch operations. The main advantage of Music Scope Pad is that it detects natural movements, such as head or hand movements, when users are listening to sounds and enables users to focus on a particular sound source that they want to hear. By moving their head left or right, users can hear the source from a frontal position as the tablet detects changes in the direction they are facing. By putting their hand behind their ear, users can focus on a particular sound source.
Masatoshi Hamanaka
CBMI1
2024 Implementation of Melody Slot Machines
Masatoshi Hamanaka
MMM (4)1
2022 Sound Scope Pad: Controlling a VR Concert with Natural Movement
abstract
We developed Sound Scope Pad, an application that provides an active music listening experience that combines AI, virtual reality, and spatial acoustics. Users can emphasize the sounds of certain performers by turning their head to the left or right or bringing their hands closer to their ears to find and focus on the performer they want to listen to. In the Sound Scope Headphones that we previously built, the user’s head direction was detected by an accelerometer mounted on the arch of the headphones. In Sound Scope Pad, the head direction is detected by combining the angle information detected by the acceleration gyro sensor of a tablet and the angle information of the head recognized from the front camera image of the tablet.
Masatoshi Hamanaka
ICMI1
2021 Audio-Oriented Video Interpolation Using Key Pose
abstract
This paper describes a deep learning-based method for long-term video interpolation that generates intermediate frames between two music performance videos of a person playing a specific instrument. Recent advances in deep learning techniques have successfully generated realistic images with high-fidelity and high-resolution in short-term video interpolation. However, there is still room for improvement in long-term video interpolation due to lack of resolution and temporal consistency of the generated video. Particularly in music performance videos, the music and human performance motion need to be synchronized. We solved these problems by using human poses and music features essential for music performance in long-term video interpolation. By closely matching human poses with music and videos, it is possible to generate intermediate frames that synchronize with the music. Specifically, we obtain the human poses of the last frame of the first video and the first frame of the second video in the performance videos to be interpolated as key poses. Then, our encoder–decoder network estimates the human poses in the intermediate frames from the obtained key poses, with the music features as the condition. In order to construct an end-to-end network, we utilize a differentiable network that transforms the estimated human poses in vector form into the human pose in image form, such as human stick figures. Finally, a video-to-video synthesis network uses the stick figures to generate intermediate frames between two music performance videos. We found that the generated performance videos were of higher quality than the baseline method through quantitative experiments.
Takayuki Nakatsuka, Yukitaka Tsuchiya, Masatoshi Hamanaka, Shigeo Morishima
Int. J. Pattern Recognit. Artif. Intell.3
2019 Melody Slot Machine: A Controllable Holographic Virtual Performer
abstract
This paper describes the "Melody Slot Machine," an interactive music system that enables control over virtual performers. Conventional virtual players focus on what kind of output performance is given to the input performance, and the performance output is difficult to control. The Melody Slot Machine enables the user to select the melody to be played next by the virtual player by rotating a dial. Furthermore, the performer is projected on a holographic display, and the user can feel as if a real virtual player is there. To achieve this, the system needs to change the melody and the performance video.
Masatoshi Hamanaka
ACM Multimedia1
2019 Proposal of an Annotation Method for Integrating Musical Technique Knowledge Using a GTTM Time-Span Tree
Nami Iino, Mayumi Shimada, Takuichi Nishimura, Hideaki Takeda 0001, Masatoshi Hamanaka
MMM (1)5
2016 Tree-structured probabilistic model of monophonic written music based on the generative theory of tonal music
abstract
This paper presents a probabilistic formulation of music language modelling based on the generative theory of tonal music (GTTM) named probabilistic GTTM (PGTTM). GTTM is a well-known music theory that describes the tree structure of written music in analogy with the phrase structure grammar of natural language. To develop a computational music language model incorporating GTTM and a machine-learning framework for data-driven music grammar induction, we construct a generative model of monophonic music based on probabilistic context-free grammar, in which the time-span tree proposed in GTTM corresponds to the parse tree. Applying the techniques of natural language processing, we also derive supervised and unsupervised learning algorithms based on the maximal-likelihood estimation, and a Bayesian inference algorithm based on the Gibbs sampling. Despite the conceptual simplicity of the model, we found that the model automatically acquires music grammar from data and reproduces time-span trees of written music as accurately as an analyser that required elaborate manual parameter tuning.
Eita Nakamura, Masatoshi Hamanaka, Keiji Hirata 0001, Kazuyoshi Yoshii
ICASSP2
2003 A Learning-Based Jam Session System that Imitates a Player's Personality Model
Masatoshi Hamanaka, Masataka Goto, Hideki Asoh, Nobuyuki Otsu
IJCAI1