Masahiro Hamasaki

dblp:75/6053 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-3085-7446ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-authorHuman-computer interaction and ubiquitous computing · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author
YearPublicationVenuePosition
2026 RAG-Enhanced Prompt Compression with Need-Oriented Knowledge for Dialog Based Embodied Navigation
Hiroaki Shimoma, Sudesna Chakraborty, Takeshi Morita 0001, Aoi Ohta, Masaki Asada, Shusaku Egami, Takanori Ugai, Masahiro Hamasaki
ICAART (4)8
2026 HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
abstract
Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured knowledge. Integrating these two in a complementary manner enables the development of reliable and verifiable AI systems. In particular, knowledge graph question answering (KGQA) has attracted attention as a means to reduce LLM hallucinations and to leverage knowledge beyond the training data. However, existing KGQA benchmark datasets are biased toward encyclopedic knowledge, limited to a single modality, and lack fine-grained spatiotemporal data, which limits their applicability to real-world scenarios targeted by Embodied AI. We introduce HOME-KGQA, a novel KGQA benchmark dataset built on a multimodal KG of daily household activities. HOME-KGQA consists of complex, multi-hop natural language questions paired with graph database query languages. Compared to existing benchmarks, it includes more challenging questions that involve multi-level spatiotemporal reasoning, multimodal grounding, and aggregate functions. Experimental results show that the LLM-based KGQA methods fail to achieve performance comparable to that on existing datasets when evaluated on HOME-KGQA. This highlights significant challenges that should be addressed for the real-world deployment of KGQA systems. Our dataset is available at https://github.com/aistairc/home-kgqa
Shusaku Egami, Aoi Ohta, Tomoki Tsujimura, Masaki Asada, Tatsuya Ishigaki, Ken Fukuda, Masahiro Hamasaki, Hiroya Takamura
LREC7
2026 A Case Study of a Transparent and Controllable Music Recommender System with Multi-relational Layers
abstract
Abstract In recommending songs to users, various types of relationships can be considered, such as songs liked by users with similar preferences or songs that are acoustically similar to those the target user already likes. Providing explanations for recommendations based on such relationships improves transparency and trust, but users currently have no control over which relationships are emphasized. To solve this problem, we extend an existing recommendation method based on a graph convolutional network (GCN) by representing each relationship as a separate graph layer with adjustable weights. By applying this method, we implemented a song recommender system with three types of relationships (user preference similarity, acoustic similarity, and creator commonality) on a music web service called “Kiite.” On the service, four types of recommendation results are displayed, depending on which relationships are emphasized and to what degree. The recommender system offers both transparency and controllability in that users can freely switch between the four recommendation result types. An analysis of over two years of usage logs demonstrates the effectiveness of combining transparency and controllability in music recommendation.
Kosetsu Tsukuda, Keisuke Ishida, Takumi Takahashi, Masahiro Hamasaki, Masataka Goto
MMM (1)4
2025 Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
abstract
Preferential Bayesian optimization (PBO) is a variant of Bayesian optimization that observes relative preferences (e.g., pairwise comparisons) instead of direct objective values, making it especially suitable for human-in-the-loop scenarios. However, real-world optimization tasks often involve inequality constraints, which existing PBO methods have not yet addressed. To fill this gap, we propose constrained preferential Bayesian optimization (CPBO), an extension of PBO that incorporates inequality constraints for the first time. Specifically, we present a novel acquisition function for this purpose. Our technical evaluation shows that our CPBO method successfully identifies optimal solutions by focusing on exploring feasible regions. As a practical application, we also present a designer-in-the-loop system for banner ad design using CPBO, where the objective is the designer's subjective preference, and the constraint ensures a target predicted click-through rate. We conducted a user study with professional ad designers, demonstrating the potential benefits of our approach in guiding creative design under real-world constraints.
Koki Iwai, Yusuke Kumagae, Yuki Koyama 0001, Masahiro Hamasaki, Masataka Goto
IJCAI4
2025 Kiite World: Socializing Map-Based Music Exploration Through Playlist Sharing and Synchronized Listening
abstract
Abstract Numerous systems have been proposed for placing songs on a map to enable music exploration, but existing systems assume that users explore alone and thus lack social interactions, which has been identified as a significant issue for these systems. In this paper, we describe “Kiite World,” a web service that enables social-aware music exploration. Kiite World has over 440,000 songs placed on a map and lets users perform the following social interactions while moving their avatars: (1) Users can publish “My Kiite World,” where songs from their created playlists are displayed on the map, and they can visit each other’s “My Kiite Worlds” to explore songs on the map. (2) The activities of all users exploring songs on Kiite World are visualized in real time, enabling users to synchronize with interested users and explore songs while listening to music together. (3) Any user can easily host music events where she listens to her favorite songs together with other users while they synchronize with her. Analysis of user behavior logs over seven months revealed several reusable insights on the usefulness of incorporating social aspects into map-based music exploration (e.g., users often like songs that are farther from their original interests as a result of exploring songs in other users’ “My Kiite Worlds.”).
Kosetsu Tsukuda, Takumi Takahashi, Keisuke Ishida, Masahiro Hamasaki, Masataka Goto
MMM (2)4
2025 SingDistVis: interactive Overview+Detail visualization for F0 trajectories of numerous singers singing the same song
abstract
Abstract This paper describes SingDistVis, an information visualization technique for fundamental frequency (F0) trajectories of large-scale singing data where numerous singers sing the same song. SingDistVis allows to explore F0 trajectories interactively by combining two views: OverallView and DetailedView. OverallView visualizes a distribution of the F0 trajectories of the song in a time-frequency heatmap. When a user specifies an interesting part, DetailedView zooms in on the specified part and visualizes singing assessment (rating) results. Here, it displays high-rated singings in red and low-rated singings in blue. When the user clicks on a particular singing, the audio source is played and its F0 trajectory through the song is displayed in OverallView. We selected heatmap-based visualization for OverallView to provide an overview of a large-scale F0 dataset, and polyline-based visualization for DetailedView to provide a more precise representation of a small number of particular F0 trajectories. This paper introduces a subjective experiment using 1,000 singing voices to determine suitable visualization parameters. Then, this paper presents user evaluations where we asked participants to compare visualization results of four types of Overview+Detail designs and concluded that the presented design archived better evaluations than other designs in all the seven questions. Finally, this paper describes a user experiment in which eight participants compare SingDistVis with a baseline implementation in exploring interested singing voices and concludes that the proposed SingDistVis archived better evaluations in nine of the questions.
Takayuki Itoh, Tomoyasu Nakano, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto
Multim. Tools Appl.4
2025 Exploring the effectiveness of user-driven intent-based recommendation models implemented in a real-world music web service
abstract
Abstract This paper explores the effectiveness of a flexible song recommendation function implemented in a music web service. The function allows users to create recommendation models, which we refer to as intent-based recommendation models (IBRMs), according to their intents. For example, a user can develop IBRMs for “cool songs,” “songs for concentrating on work,” and so on, and receive recommendations from each of the IBRMs according to her intents. The key novelty of this work lies in the architecture that enables users to explicitly construct and maintain multiple personalized recommendation models in parallel, each specialized for a particular intent. This user-driven approach contrasts with conventional systems that rely on a single, system-controlled recommendation model per user. To develop an IBRM, the user first initializes it by choosing seed songs and then repeatedly updates it by giving feedback based on whether recommended songs are relevant to the user’s intent. In the case study using the real-world web service “Kiite,” we analyze 1,116 IBRMs created by 417 users and show key characteristics of those IBRMs (e.g., it is meaningful to enable users to create their own IBRMs, because the created IBRMs generate largely different recommendation results from one another). These findings demonstrate the effectiveness and practical value of enabling users to control intent-specific recommendation behavior through the proposed IBRM framework.
Kosetsu Tsukuda, Keisuke Ishida, Kento Watanabe, Masahiro Hamasaki, Masataka Goto
Multim. Tools Appl.4
2023 Content-Based Music-Image Retrieval Using Self- and Cross-Modal Feature Embedding Memory
abstract
This paper describes a method based on deep metric learning for content-based cross-modal retrieval of a piece of music and its representative image (i.e., a music audio signal and its cover art image). We train music and image encoders so that the embeddings of a positive music-image pair lie close to each other, while those of a random pair lie far from each other, in a shared embedding space. Furthermore, we propose a mechanism called self- and cross-modal feature embedding memory, which stores both the music and image embeddings of any previous iterations in memory and enables the encoders to mine informative pairs for training. To perform such training, we constructed a dataset containing 78,325 music-image pairs. We demonstrate the effectiveness of the proposed mechanism on this dataset: specifically, our mechanism outperforms baseline methods by ×1.93 ∼ 3.38 for the mean reciprocal rank, ×2.19 ∼ 3.56 for recall@50, and 528 ∼ 891 ranks for the median rank.
Takayuki Nakatsuka, Masahiro Hamasaki, Masataka Goto
WACV2
2020 Interactive deep singing-voice separation based on human-in-the-loop adaptation
abstract
This paper presents a deep-learning-based interactive system separating the singing voice from input polyphonic music signals. Although deep neural networks have been successful for singing voice separation, no approach using them allows any user interaction for improving the separation quality. We present a framework that allows a user to interactively fine-tune the deep neural model at run time to adapt it to the target song. This is enabled by designing unified networks consisting of two U-Net architectures based on frequency spectrogram representations: one for estimating the spectrogram mask that can be used to extract the singing-voice spectrogram from the input polyphonic spectrogram; the other for estimating the fundamental frequency (F0) of the singing voice. Although it is not easy for the user to edit the mask, he or she can iteratively correct errors in part of the visualized F0 trajectory through simple interaction. Our unified networks leverage the user-corrected F0 to improve the rest of the F0 trajectory through the model adaptation, which results in better separation quality. We validated this approach in a simulation experiment showing that the F0 correction can improve the quality of singing-voice separation. We also conducted a pilot user study with an expert musician, who used our system to produce a high-quality singing-voice separation result.
Tomoyasu Nakano, Yuki Koyama 0001, Masahiro Hamasaki, Masataka Goto
IUI3
2020 Audio-visual object removal in 360-degree videos
abstract
Abstract We present a novel concept audio–visual object removal in 360-degree videos, in which a target object in a 360-degree video is removed in both the visual and auditory domains synchronously. Previous methods have solely focused on the visual aspect of object removal using video inpainting techniques, resulting in videos with unreasonable remaining sounds corresponding to the removed objects. We propose a solution which incorporates direction acquired during the video inpainting process into the audio removal process. More specifically, our method identifies the sound corresponding to the visually tracked target object and then synthesizes a three-dimensional sound field by subtracting the identified sound from the input 360-degree video. We conducted a user study showing that our multi-modal object removal supporting both visual and auditory domains could significantly improve the virtual reality experience, and our method could generate sufficiently synchronous, natural and satisfactory 360-degree videos.
Ryo Shimamura, Yuki Koyama 0001, Takayuki Nakatsuka, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto, Shigeo Morishima
Vis. Comput.6
2018 Collaboration in N-th Order Derivative Creation
Shiori Hironaka, Kosetsu Tsukuda, Masahiro Hamasaki, Masataka Goto
ICWSM3
2017 Classifying derivative works with search, text, audio and video features
abstract
Users of video-sharing sites often search for derivative works of music, such as live versions, covers, and remixes. Audio and video content are both important for retrieval: “karaoke” specifies audio content (instrumental version) and video content (animated lyrics). Although YouTube's text search is fairly reliable, many search results do not match the exact query. We introduce an algorithm to classify YouTube videos by category of derivative work. Based on a standard pipeline for video-based genre classification, it combines search, text, and video features with a novel set of audio features derived from audio fingerprints. A baseline approach is outperformed by the search and text features alone, and combining these with video and audio features performs best of all, reducing the audio content error rate from 25% to 15%.
Jordan B. L. Smith, Masahiro Hamasaki, Masataka Goto
ICME2
2017 QueryShare: Working Together to Facilitate Exploratory Multimedia Searches without Skill in Creating
abstract
This paper describes a music exploratory search interface called QueryShare, which provides query searching and recommendation functions for query sharing among users. Most people are not expert users who know how to use various music metadata that include automatically estimated musical features to represent their own information needs as a query. Therefore, it is difficult for them to enter a complex query for music content retrieval. The original feature of our proposed interface is to make users share every query as a public web page. This feature enables users to use search queries, find recommended queries, and revise existing queries. Beginners can use an applicable query, which is more complicated than they might create on their own. Experts can readily reuse a query (web page) of their own making. The interface assists users in finding results for interesting queries and in performing music exploratory search without skills to create complex queries. We developed a prototype system as a web application for music videos on the most popular Japanese video sharing service. Users can search for over 360,000 music videos using our system. Results of a preliminary user study demonstrated that users found the query creation interesting and that they were interested in seeing and using queries created by other users, although some users hesitated to share their queries.
Masahiro Hamasaki, Masataka Goto
OpenSym1
2016 Why Did You Cover That Song?: Modeling N-th Order Derivative Creation with Content Popularity
abstract
Many amateur creators now create derivative works and put them on the web. Although there are several factors that inspire the creation of derivative works, such factors cannot usually be observed on the web. In this paper, we propose a model for inferring latent factors from sequences of derivative work posting events. We assume a sequence to be a stochastic process incorporating the following three factors: (1) the original work's attractiveness, (2) the original work's popularity, and (3) the derivative work's popularity. To characterize content popularity, we use content ranking data and incorporate rank-biased popularity based on the creators' browsing behavior. Our main contributions are three-fold: (1) to the best of our knowledge, this is the first study modeling derivative creation activity, (2) by using a real-world dataset of music-related derivative work creation to evaluate our model, we showed the effectiveness of adopting all three factors to model derivative creation activity and onsidering creators' browsing behavior, and (3) we carried out qualitative experiments and showed that our model is useful in analyzing derivative creation activity in terms of category characteristics, temporal development of factors that trigger derivative work posting events, etc.
Kosetsu Tsukuda, Masahiro Hamasaki, Masataka Goto
CIKM2
2016 PlaylistPlayer: An Interface Using Multiple Criteria to Change the Playback Order of a Music Playlist
abstract
We propose a novel interface that allows the user to interactively change the playback order of multiple songs by choosing one or more criteria. The criteria include not only the song's title and artist name but also its content automatically estimated by music/singing signal processing and artist-level social analysis. The artist-level social information is discovered from Wikipedia and DBpedia. With regard to manipulating playback order, existing interfaces typically allow the user to change it manually or automatically by choosing one of a few types of criteria. The proposed interface, on the other hand, deals with nine properties and multiple integrations of them (e.g., vocal gender and beats per minute). To realize the ordering by multiple criteria, a distance matrix is computed from the criteria vectors and is then used to estimate paths for ascending, descending, and random orders by applying principle component analysis or to estimate a path for a smooth order by solving the travelling salesman problem.
Tomoyasu Nakano, Jun Kato 0001, Masahiro Hamasaki, Masataka Goto
IUI3
2013 Songrium: a music browsing assistance service based on visualization of massive open collaboration within music content creation community
abstract
This paper describes a music browsing assistance service, Songrium (http://songrium.jp), that helps a user enjoy songs while seeing visualization of open collaboration. Songrium focuses on open collaboration for music content creation on the most popular Japanese video-sharing service. Since this open collaboration generates more than half a million video clips with a rich variety of music content, we call it massive open collaboration. To develop a shared understanding of this collaboration we have analyzed, we developed Songrium that visualizes relations among both original songs and derivative works generated from the collaboration. Songrium also features a social annotation framework to verbalize and share various relations among songs, and a flexible ranking mechanism to find interesting songs. After we launched Songrium in August 2012, more than 7,000 users have used our service in which over 98,000 songs and 520,000 derivative works have automatically been registered. We hope Songrium will not only encourage creators to create more derivative works, but also attract consumers to participate in the collaboration as creators.
Masahiro Hamasaki, Masataka Goto
OpenSym1
2011 Social Infobox: collaborative knowledge construction by social property tagging
abstract
We propose a novel style of social tagging to construct knowledge collaboratively called Social Property Tagging and introduce the prototype system Social Infobox. Structured data is useful for computer system, however defining structure of knowledge for representing data semantics is usually a costly and time consuming task. In general, data structures are constructed by experts of knowledge engineering. Our method aims to construct not only structured data but also structure of data collaboratively by simple user input.
Masahiro Hamasaki, Masataka Goto, Hideaki Takeda 0001
CSCW1
2010 Zuzie: Collaborative Storytelling Based on Multiple Compositions
Yoshiyuki Nakamura, Maiko Kobayakawa, Chisato Takami, Yuta Tsuruga, Hidekazu Kubota, Masahiro Hamasaki, Takuichi Nishimura, Takeshi Sunaga
ICIDS6
2009 Familial collaborations in a museum
abstract
Studies of interactive systems in museums have raised important design considerations, but so far have failed to address sufficiently the particularities of family interaction and co-operation. This paper introduces qualitative video-based observations of Japanese families using an interactive portable guide system in a museum. Results show how unexpected usage can occur through particularities of interaction between family members. The paper highlights the necessity to more fully consider familial relationships in HCI.
Tom Hope, Yoshiyuki Nakamura, Atsushi Nobayashi, Shota Fukuoka, Masahiro Hamasaki, Takuichi Nishimura
CHI6
2009 TrBagg: A Simple Transfer Learning Method and its Application to Personalization in Collaborative Tagging
abstract
The aim of transfer learning is to improve prediction accuracy on a target task by exploiting the training examples for tasks that are related to the target one. Transfer learning has received more attention in recent years, because this technique is considered to be helpful in reducing the cost of labeling. In this paper, we propose a very simple approach to transfer learning: TrBagg, which is the extension of bagging. TrBagg is composed of two stages: Many weak classifiers are first generated as in standard bagging, and these classifiers are then filtered based on their usefulness for the target task. This simplicity makes it easy to work reasonably well without severe tuning of learning parameters. Further, our algorithm equips an algorithmic scheme to avoid negative transfer. We applied TrBagg to personalized tag prediction tasks for social bookmarks. Our approach has several convenient characteristics for this task such as adaptation to multiple tasks with low computational cost.
Toshihiro Kamishima, Masahiro Hamasaki, Shotaro Akaho
ICDM2
2009 Network Analysis of an Emergent Massively Collaborative Creation Community: How Can People Create Videos Collaboratively without Collaboration?
Masahiro Hamasaki, Hideaki Takeda 0001, Tom Hope, Takuichi Nishimura
ICWSM1
2007 POLYPHONET: An advanced social network extraction system from the Web
Yutaka Matsuo, Junichiro Mori, Masahiro Hamasaki, Takuichi Nishimura, Hideaki Takeda 0001, Kôiti Hasida, Mitsuru Ishizuka
J. Web Semant.3
2006 Spinning Multiple Social Networks for Semantic Web
Yutaka Matsuo, Masahiro Hamasaki, Yoshiyuki Nakamura, Takuichi Nishimura, Kôiti Hasida, Hideaki Takeda 0001, Junichiro Mori, Danushka Bollegala, Mitsuru Ishizuka
AAAI2
2006 Doing Community: Co-construction of Meaning and Use with Interactive Information Kiosks
Tom Hope, Masahiro Hamasaki, Yutaka Matsuo, Yoshiyuki Nakamura, Noriyuki Fujimura, Takuichi Nishimura
UbiComp2
2006 Context-Aware Weblog to Enhance Communication among Participants in a Conference
Kosuke Numa, Hideaki Takeda 0001, Takuichi Nishimura, Yutaka Matsuo, Masahiro Hamasaki, Noriyuki Fujimura, Keisuke Ishida, Tom Hope, Yoshiyuki Nakamura, Satoshi Fujiyoshi, Kazuya Sakamoto, Hiroshi Nagata, Osamu Nakagawa, Eiji Shinbori
WEBIST (1)5
2006 POLYPHONET: an advanced social network extraction system from the web
abstract
Social networks play important roles in the Semantic Web: knowledge management, information retrieval, ubiquitous computing, and so on. We propose a social network extraction system called POLYPHONET, which employs several advanced techniques to extract relations of persons, detect groups of persons, and obtain keywords for a person. Search engines, especially Google, are used to measure co-occurrence of information and obtain Web documents.Several studies have used search engines to extract social networks from the Web, but our research advances the following points: First, we reduce the related methods into simple pseudocodes using Google so that we can build up integrated systems. Second, we develop several new algorithms for social networking mining such as those to classify relations into categories, to make extraction scalable, and to obtain and utilize person-to-word relations. Third, every module is implemented in POLYPHONET, which has been used at four academic conferences, each with more than 500 participants. We overview that system. Finally, a novel architecture called Super Social Network Mining is proposed; it utilizes simple modules using Google and is characterized by scalability and Relate-Identify processes: Identification of each entity and extraction of relations are repeated to obtain a more precise social network.
Yutaka Matsuo, Junichiro Mori, Masahiro Hamasaki, Keisuke Ishida, Takuichi Nishimura, Hideaki Takeda 0001, Kôiti Hasida, Mitsuru Ishizuka
WWW3
2004 Discovering Relationships Among Catalogs
Ryutaro Ichise, Masahiro Hamasaki, Hideaki Takeda 0001
Discovery Science2
2004 A Hybrid Algorithm for Alignment of Concept Hierarchies
Ryutaro Ichise, Masahiro Hamasaki, Hideaki Takeda 0001
EKAW2
2004 A Multi-strategy Approach for Catalog Integration
Ryutaro Ichise, Masahiro Hamasaki, Hideaki Takeda 0001
PRICAI2
2004 Metadata-Driven Personal Knowledge Publishing
Ikki Ohmukai, Hideaki Takeda 0001, Masahiro Hamasaki, Kosuke Numa, Shin Adachi
ISWC3
2003 Neighborhood Matchmaker Method: A Decentralized Optimization Algorithm for Personal Human Network
Masahiro Hamasaki, Hideaki Takeda 0001
KES1