Masataka Goto

dblp:85/1753 · DBLP profile ↗
← Back
151ranked-venue papers
20as first author
19since 2021 · last 2026
0000-0003-1167-0977ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 97 · 16 first-author · 9 since 2021Artificial intelligence and machine learning · 47 · 8 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 25 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 A Case Study of a Transparent and Controllable Music Recommender System with Multi-relational Layers
abstract
Abstract In recommending songs to users, various types of relationships can be considered, such as songs liked by users with similar preferences or songs that are acoustically similar to those the target user already likes. Providing explanations for recommendations based on such relationships improves transparency and trust, but users currently have no control over which relationships are emphasized. To solve this problem, we extend an existing recommendation method based on a graph convolutional network (GCN) by representing each relationship as a separate graph layer with adjustable weights. By applying this method, we implemented a song recommender system with three types of relationships (user preference similarity, acoustic similarity, and creator commonality) on a music web service called “Kiite.” On the service, four types of recommendation results are displayed, depending on which relationships are emphasized and to what degree. The recommender system offers both transparency and controllability in that users can freely switch between the four recommendation result types. An analysis of over two years of usage logs demonstrates the effectiveness of combining transparency and controllability in music recommendation.
Kosetsu Tsukuda, Keisuke Ishida, Takumi Takahashi, Masahiro Hamasaki, Masataka Goto
MMM (1)5
2025 Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
abstract
Preferential Bayesian optimization (PBO) is a variant of Bayesian optimization that observes relative preferences (e.g., pairwise comparisons) instead of direct objective values, making it especially suitable for human-in-the-loop scenarios. However, real-world optimization tasks often involve inequality constraints, which existing PBO methods have not yet addressed. To fill this gap, we propose constrained preferential Bayesian optimization (CPBO), an extension of PBO that incorporates inequality constraints for the first time. Specifically, we present a novel acquisition function for this purpose. Our technical evaluation shows that our CPBO method successfully identifies optimal solutions by focusing on exploring feasible regions. As a practical application, we also present a designer-in-the-loop system for banner ad design using CPBO, where the objective is the designer's subjective preference, and the constraint ensures a target predicted click-through rate. We conducted a user study with professional ad designers, demonstrating the potential benefits of our approach in guiding creative design under real-world constraints.
Koki Iwai, Yusuke Kumagae, Yuki Koyama 0001, Masahiro Hamasaki, Masataka Goto
IJCAI5
2025 Kiite World: Socializing Map-Based Music Exploration Through Playlist Sharing and Synchronized Listening
abstract
Abstract Numerous systems have been proposed for placing songs on a map to enable music exploration, but existing systems assume that users explore alone and thus lack social interactions, which has been identified as a significant issue for these systems. In this paper, we describe “Kiite World,” a web service that enables social-aware music exploration. Kiite World has over 440,000 songs placed on a map and lets users perform the following social interactions while moving their avatars: (1) Users can publish “My Kiite World,” where songs from their created playlists are displayed on the map, and they can visit each other’s “My Kiite Worlds” to explore songs on the map. (2) The activities of all users exploring songs on Kiite World are visualized in real time, enabling users to synchronize with interested users and explore songs while listening to music together. (3) Any user can easily host music events where she listens to her favorite songs together with other users while they synchronize with her. Analysis of user behavior logs over seven months revealed several reusable insights on the usefulness of incorporating social aspects into map-based music exploration (e.g., users often like songs that are farther from their original interests as a result of exploring songs in other users’ “My Kiite Worlds.”).
Kosetsu Tsukuda, Takumi Takahashi, Keisuke Ishida, Masahiro Hamasaki, Masataka Goto
MMM (2)5
2025 A Data-Driven Method for Analyzing and Quantifying Lyrics-Dance Motion Relationships
abstract
Kento Watanabe, Masataka Goto. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kento Watanabe, Masataka Goto
NAACL (Long Papers)2
2025 SingDistVis: interactive Overview+Detail visualization for F0 trajectories of numerous singers singing the same song
abstract
Abstract This paper describes SingDistVis, an information visualization technique for fundamental frequency (F0) trajectories of large-scale singing data where numerous singers sing the same song. SingDistVis allows to explore F0 trajectories interactively by combining two views: OverallView and DetailedView. OverallView visualizes a distribution of the F0 trajectories of the song in a time-frequency heatmap. When a user specifies an interesting part, DetailedView zooms in on the specified part and visualizes singing assessment (rating) results. Here, it displays high-rated singings in red and low-rated singings in blue. When the user clicks on a particular singing, the audio source is played and its F0 trajectory through the song is displayed in OverallView. We selected heatmap-based visualization for OverallView to provide an overview of a large-scale F0 dataset, and polyline-based visualization for DetailedView to provide a more precise representation of a small number of particular F0 trajectories. This paper introduces a subjective experiment using 1,000 singing voices to determine suitable visualization parameters. Then, this paper presents user evaluations where we asked participants to compare visualization results of four types of Overview+Detail designs and concluded that the presented design archived better evaluations than other designs in all the seven questions. Finally, this paper describes a user experiment in which eight participants compare SingDistVis with a baseline implementation in exploring interested singing voices and concludes that the proposed SingDistVis archived better evaluations in nine of the questions.
Takayuki Itoh, Tomoyasu Nakano, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto
Multim. Tools Appl.5
2025 Exploring the effectiveness of user-driven intent-based recommendation models implemented in a real-world music web service
abstract
Abstract This paper explores the effectiveness of a flexible song recommendation function implemented in a music web service. The function allows users to create recommendation models, which we refer to as intent-based recommendation models (IBRMs), according to their intents. For example, a user can develop IBRMs for “cool songs,” “songs for concentrating on work,” and so on, and receive recommendations from each of the IBRMs according to her intents. The key novelty of this work lies in the architecture that enables users to explicitly construct and maintain multiple personalized recommendation models in parallel, each specialized for a particular intent. This user-driven approach contrasts with conventional systems that rely on a single, system-controlled recommendation model per user. To develop an IBRM, the user first initializes it by choosing seed songs and then repeatedly updates it by giving feedback based on whether recommended songs are relevant to the user’s intent. In the case study using the real-world web service “Kiite,” we analyze 1,116 IBRMs created by 417 users and show key characteristics of those IBRMs (e.g., it is meaningful to enable users to create their own IBRMs, because the created IBRMs generate largely different recommendation results from one another). These findings demonstrate the effectiveness and practical value of enabling users to control intent-specific recommendation behavior through the proposed IBRM framework.
Kosetsu Tsukuda, Keisuke Ishida, Kento Watanabe, Masahiro Hamasaki, Masataka Goto
Multim. Tools Appl.5
2023 Lyric App Framework: A Web-based Framework for Developing Interactive Lyric-driven Musical Applications
abstract
Lyric videos have become a popular medium to convey lyrical content to listeners, but they present the same content whenever they are played and cannot adapt to listeners’ preferences. Lyric apps, as we name them, are a new form of lyric-driven visual art that can render different lyrical content depending on user interaction and address the limitations of static media. To open up this novel design space for programmers and musicians, we present Lyric App Framework, a web-based framework for building interactive graphical applications that play musical pieces and show lyrics synchronized with playback. We designed the framework to provide a streamlined development experience for building production-ready lyric apps with creative coding libraries of choice. We held programming contests twice and collected 52 examples of lyric apps, enabling us to reveal eight representative categories, confirm the framework’s effectiveness, and report lessons learned.
Jun Kato 0001, Masataka Goto
CHI2
2023 CatAlyst: Domain-Extensible Intervention for Preventing Task Procrastination Using Large Generative Models
abstract
CatAlyst uses generative models to help workers’ progress by influencing their task engagement instead of directly contributing to their task outputs. It prompts distracted workers to resume their tasks by generating a continuation of their work and presenting it as an intervention that is more context-aware than conventional (predetermined) feedback. The prompt can function by drawing their interest and lowering the hurdle for resumption even when the generated continuation is insufficient to substitute their work, while recent human-AI collaboration research aiming at work substitution depends on a stable high accuracy. This frees CatAlyst from domain-specific model-tuning and makes it applicable to various tasks. Our studies involving writing and slide-editing tasks demonstrated CatAlyst’s effectiveness in helping workers swiftly resume tasks with a lowered cognitive load. The results suggest a new form of human-AI collaboration where large generative models publicly available but imperfect for each individual domain can contribute to workers’ digital well-being.
Riku Arakawa, Hiromu Yakura, Masataka Goto
CHI3
2023 U-Beat: A Multi-Scale Beat Tracking Model Based on Wave-U-Net
abstract
In this paper, we propose a multi-scale model for beat tracking based on the Wave-U-Net model. The proposed model learns multi-scale features by repeatedly resampling feature maps via a series of downsampling blocks and upsampling blocks. With the U-shape structure, we observe that global features are summarized at the bottom blocks. Then, these global features guide feature upsampling for predicting beats with a steady tempo. The local features learned in the downsampling blocks are combined with the upsampled features for predicting beats precisely. Besides the features learned from the waveform, we also combine spectral features at a middle level in the model. Experimental results show that beat tracking performance is improved by combining spectral features.
Tian Cheng 0001, Masataka Goto
ICASSP2
2023 Content-Based Music-Image Retrieval Using Self- and Cross-Modal Feature Embedding Memory
abstract
This paper describes a method based on deep metric learning for content-based cross-modal retrieval of a piece of music and its representative image (i.e., a music audio signal and its cover art image). We train music and image encoders so that the embeddings of a positive music-image pair lie close to each other, while those of a random pair lie far from each other, in a shared embedding space. Furthermore, we propose a mechanism called self- and cross-modal feature embedding memory, which stores both the music and image embeddings of any previous iterations in memory and enables the encoders to mine informative pairs for training. To perform such training, we constructed a dataset containing 78,325 music-image pairs. We demonstrate the effectiveness of the proposed mechanism on this dataset: specifically, our mechanism outperforms baseline methods by ×1.93 ∼ 3.38 for the mean reciprocal rank, ×2.19 ∼ 3.56 for recall@50, and 528 ∼ 891 ranks for the median rank.
Takayuki Nakatsuka, Masahiro Hamasaki, Masataka Goto
WACV3
2022 BeParrot: Efficient Interface for Transcribing Unclear Speech via Respeaking
abstract
Transcribing speech from audio files to text is an important task not only for exploring the audio content in text form but also for utilizing the transcribed data as a source to train speech models, such as automated speech recognition (ASR) models. A post-correction approach has been frequently employed to reduce the time cost of transcription where users edit errors in the recognition results of ASR models. However, this approach assumes clear speech and is not designed for unclear speech (such as speech with high levels of noise or reverberation), which severely degrades the accuracy of ASR and requires many manual corrections. To construct an alternative approach to transcribe unclear speech, we introduce the idea of respeaking, which has primarily been used to create captions for television programs in real time. In respeaking, a proficient human respeaker repeats the heard speech as shadowing, and their utterances are recognized by an ASR model. While this approach can be effective for transcribing unclear speech, one problem is that respeaking is a highly cognitively demanding task and extensive training is often required to become a respeaker. We address this point with BeParrot, the first interface designed for respeaking that allows novice users to benefit from respeaking without extensive training through two key features: parameter adjustment and pronunciation feedback. Our user study involving 60 crowd workers demonstrated that they could transcribe different types of unclear speech 32.2 % faster with BeParrot than with a conventional approach without losing the accuracy of transcriptions. In addition, comments from the workers supported the design of the adjustment and feedback features, exhibiting a willingness to continue using BeParrot for transcription tasks. Our work demonstrates how we can leverage recent advances in machine learning techniques to overcome the area that is still challenging for computers themselves with the help of a human-in-the-loop approach.
Riku Arakawa, Hiromu Yakura, Masataka Goto
IUI3
2022 BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions
abstract
Many design tasks involve parameter adjustment, and designers often struggle to find desirable parameter value combinations by manipulating sliders back and forth. For such a multi-dimensional search problem, Bayesian optimization (BO) is a promising technique because of its intelligent sampling strategy; in each iteration, BO samples the most effective points considering both exploration (i.e., prioritizing unexplored regions) and exploitation (i.e., prioritizing promising regions), enabling efficient searches. However, existing BO-based design frameworks take the initiative in the design process and thus are not flexible enough for designers to freely explore the design space using their domain knowledge. In this paper, we propose a novel design framework, BO as Assistant, which enables designers to take the initiative in the design process while also benefiting from BO’s sampling strategy. The designer can manipulate sliders as usual; the system monitors the slider manipulation to automatically estimate the design goal on the fly and then asynchronously provides unexplored-yet-promising suggestions using BO’s sampling strategy. The designer can choose to use the suggestions at any time. This framework uses a novel technique to automatically extract the necessary information to run BO by observing slider manipulation without requesting additional inputs. Our framework is domain-agnostic, demonstrated by applying it to photo color enhancement, 3D shape design for personal fabrication, and procedural material design in computer graphics.
Yuki Koyama 0001, Masataka Goto
UIST2
2022 Deep Learning Approaches in Topics of Singing Information Processing
abstract
Singing, the vocal production of musical tones, is one of the most important elements of music. Addressing the needs of real-world applications, the study of technologies related to singing voices has become an increasingly active area of research. In this paper, we provide a comprehensive overview of the recent developments in the field of singing information processing, specifically in the topics of singing skill evaluation, singing voice synthesis, singing voice separation, and lyrics synchronization and transcription. We will especially focus on deep learning approaches including modern representation learning techniques for singing voices. We will also provide an overview of contributions in public datasets for singing voice research.
Chitralekha Gupta, Haizhou Li 0001, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Singer Diarization for Polyphonic Music With Unison Singing
abstract
This paper introduces a new framework for singer diarization, which is a technique to reveal who sings when in songs with multiple singers. Although various techniques have been developed to analyze and extract features of singing voices in musical audio signals, most of them assume that a song is sung by a single singer, and singer diarization for multiple singers has not been well studied in the field of singing information processing. To deal with multiple speakers in speech analysis, speaker diarization has been explored to handle overlapped speech voices, but cannot handle singing voices well because of acoustic differences between singing and speech voices. This paper therefore proposes a new diarization framework specialized in singing voices. To achieve high accuracy in overlap detection, this paper proposes a novel acoustic feature named Cosacorr score, which is helpful in estimating whether a song is sung by more than one singer. After extracting singing voices from polyphonic music by using a singing voice separation technique, the framework adopts an existing ArcFace technique to extract discriminative singer representations from short segments of the separated singing voices. The framework is evaluated by using a new private dataset of unison singing voices, which is constructed using commercially available compact discs (CDs). The experimental results show that the proposed framework outperformed the baseline method for speaker diarization in terms of diarization error rate (DER).
Hitoshi Suda, Daisuke Saito, Satoru Fukayama, Tomoyasu Nakano, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Self-Supervised Contrastive Learning for Singing Voices
abstract
This study introduces self-supervised contrastive learning to acquire feature representations of singing voices. To acquire robust representations in an unsupervised manner, regular self-supervised contrastive learning trains neural networks to make the feature representation of a sample close to those of its computationally transformed versions. Similarly, we employ two transformations—pitch shifting and time stretching—considering the nature of singing voices. Nevertheless, we use them reversely: we train networks to push away representations of the transformed versions. The networks then attempt to discriminate changes in vocal timbres introduced by pitch shifting without time stretching and those in singing expressions introduced by time stretching without pitch shifting. Consequently, the acquired representations become attentive to vocal timbre and singing expression. This was confirmed through a singer identification task, where we trained a classifier to learn the relationship between the feature representations to the corresponding singer labels of 500 singers. As a result, the employed transformations helped the classifier improve the classification accuracy by 9.12% (top-1 accuracy: 63.08%) compared with the case where the feature representations fed to the classifier were acquired without the transformations (top-1 accuracy: 53.96%). Furthermore, the proposed approach can be extended to acquire feature representations attentive to either vocal timbre or singing expression but not to the other by changing how the transformations are incorporated. We particularly explored the characteristics of such vocal timbre- or singing expression-oriented feature representations against song genre, singer gender, and vocal technique, and confirmed that they successfully capture different aspects of singing voices.
Hiromu Yakura, Kento Watanabe, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 An automated system recommending background music to listen to while working
abstract
Abstract Many people listen to music while working nowadays. However, conventional recommendation systems that are designed for playing songs matching user preferences cannot be applied for such a situation. This is because previous research showed that listeners’ concentration can be negatively affected not only by music that listeners strongly dislike but also by music that the listeners strongly like. Therefore, when we consider a recommendation system to be used while working, it is desirable to avoid both songs the user likes very much and songs the user dislikes very much. Given this background, we propose FocusMusicRecommender, a system designed specifically for recommending music to listen to while working. It summarizes songs automatically and plays them successively in order to enable users to give not only “dislike (very much)” feedback via a “skip” button but also “like (very much)” feedback via a “keep listening” button. The feedback is then combined with the users’ concentration level that is estimated from their behavioral history during the playback of the corresponding song, which allows the system to obtain preference information that distinguishes between “like” and “like very much” without burdening the user who is working. Based on the preference information, the system estimates the preference levels of unplayed songs and prioritizes the songs for subsequent playback by also considering the user’s current concentration level. Our experiments showed the validity and effectiveness of the proposed method, including the accuracy of the concentration level estimation. Moreover, our user study verified the suitability of the recommendation results from both the observed behavior and obtained comments of the participants.
Hiromu Yakura, Tomoyasu Nakano, Masataka Goto
User Model. User Adapt. Interact.3
2021 Tool- and Domain-Agnostic Parameterization of Style Transfer Effects Leveraging Pretrained Perceptual Metrics
abstract
Current deep learning techniques for style transfer would not be optimal for design support since their "one-shot" transfer does not fit exploratory design processes. To overcome this gap, we propose parametric transcription, which transcribes an end-to-end style transfer effect into parameter values of specific transformations available in an existing content editing tool. With this approach, users can imitate the style of a reference sample in the tool that they are familiar with and thus can easily continue further exploration by manipulating the parameters. To enable this, we introduce a framework that utilizes an existing pretrained model for style transfer to calculate a perceptual style distance to the reference sample and uses black-box optimization to find the parameters that minimize this distance. Our experiments with various third-party tools, such as Instagram and Blender, show that our framework can effectively leverage deep learning techniques for computational design support.
Hiromu Yakura, Yuki Koyama 0001, Masataka Goto
IJCAI3
2021 Interactive Exploration-Exploitation Balancing for Generative Melody Composition
abstract
Recent content creation systems allow users to generate various high-quality content (e.g., images, 3D models, and melodies) by just specifying a parameter set (e.g., a latent vector of a deep generative model). The task here is to search for an appropriate parameter set that produces the desired content. To facilitate this task execution, researchers have investigated user-in-the-loop optimization, where the system samples candidate solutions, asks the user to provide preferential feedback on them, and iterates this procedure until finding the desired solution. In this work, we investigate a novel approach to enhance this interactive process: allowing users to control the sampling behavior. More specifically, we allow users to adjust the balance between exploration (i.e., favoring diverse samples) and exploitation (i.e., favoring focused samples) in each iteration. To evaluate how this approach affects the user experience and optimization behavior, we implement it into a melody composition system that combines a deep generative model with Bayesian optimization. Our experiments suggest that this approach could improve the user’s engagement and optimization performance.
Yuki Koyama 0001, Masataka Goto, Takeo Igarashi
IUI3
2021 Atypical Lyrics Completion Considering Musical Audio Signals
Kento Watanabe, Masataka Goto
MMM (1)2
2020 Lyric Video Analysis Using Text Detection and Tracking
Shota Sakaguchi, Jun Kato 0001, Masataka Goto, Seiichi Uchida
DAS3
2020 Enhancing Participation Experience in VR Live Concerts by Improving Motions of Virtual Audience Avatars
abstract
While participating in live concerts is a promising application of virtual reality (VR), it falls short of our participation experience in the real world. In particular, to increase the engagement of participants, previous studies emphasized the importance of social experience among audience members, such as the sense of co-presence elicited by sharing physical reactions or body movements synchronized with music. In this respect, a common strategy in existing platforms is to present avatars of remote human participants in a VR venue and make every avatar imitate movements of the corresponding participant. However, this strategy implicitly assumes that a not small number of users connect simultaneously to watch the same content and thus is not applicable when only a few users gather or a user is watching alone. Therefore, with the aim of providing better experience to a user who participates in live concerts as one of the audience, we examine computational approaches to enhancing the sense of co-presence through virtual audience avatars. We propose four methods of presenting avatar movements: copying the user’s own movements, copying other users’ movements, repeating beat-synchronous movements, and synthesizing machine-learning-based movements. We compare their effectiveness in a user experiment and discuss application scenarios and design implications that open up new ways of active media consumption in VR environments.
Hiromu Yakura, Masataka Goto
ISMAR2
2020 Interactive deep singing-voice separation based on human-in-the-loop adaptation
abstract
This paper presents a deep-learning-based interactive system separating the singing voice from input polyphonic music signals. Although deep neural networks have been successful for singing voice separation, no approach using them allows any user interaction for improving the separation quality. We present a framework that allows a user to interactively fine-tune the deep neural model at run time to adapt it to the target song. This is enabled by designing unified networks consisting of two U-Net architectures based on frequency spectrogram representations: one for estimating the spectrogram mask that can be used to extract the singing-voice spectrogram from the input polyphonic spectrogram; the other for estimating the fundamental frequency (F0) of the singing voice. Although it is not easy for the user to edit the mask, he or she can iteratively correct errors in part of the visualized F0 trajectory through simple interaction. Our unified networks leverage the user-corrected F0 to improve the rest of the F0 trajectory through the model adaptation, which results in better separation quality. We validated this approach in a simulation experiment showing that the F0 correction can improve the quality of singing-voice separation. We also conducted a pilot user study with an expert musician, who used our system to produce a high-quality singing-voice separation result.
Tomoyasu Nakano, Yuki Koyama 0001, Masahiro Hamasaki, Masataka Goto
IUI4
2020 Drum Synthesis and Rhythmic Transformation with Adversarial Autoencoders
abstract
Creative rhythmic transformations of musical audio refer to automated methods for manipulation of temporally-relevant sounds in time. This paper presents a method for joint synthesis and rhythm transformation of drum sounds through the use of adversarial autoencoders (AAE). Users may navigate both the timbre and rhythm of drum patterns in audio recordings through expressive control over a low-dimensional latent space. The model is based on an AAE with Gaussian mixture latent distributions that introduce rhythmic pattern conditioning to represent a wide variety of drum performances. The AAE is trained on a dataset of bar-length segments of percussion recordings, along with their clustered rhythmic pattern labels. The decoder is conditioned during adversarial training for mixing of data-driven rhythmic and timbral properties. The system is trained with over 500000 bars from 5418 tracks in popular datasets covering various musical genres. In an evaluation using real percussion recordings, the reconstruction accuracy and latent space interpolation between drum performances are investigated for audio generation conditioned by target rhythmic patterns.
Maciej Tomczak, Masataka Goto, Jason Hockman
ACM Multimedia2
2020 Explainable Recommendation for Repeat Consumption
abstract
Displaying appropriate explanations for recommended items is of vital importance for improving the persuasiveness and user satisfaction of recommender systems. Although a user often consumes the same item repeatedly in some domains such as music and restaurants, existing studies have focused on generating explanations for recommending novel items. In this paper, we describe the concept of explainable recommendation for repeatedly consumed items. Because of the high proportion of repeat consumption in music listening, we suggest nine kinds of explanations for song recommendations according to three factors: personal, social, and item factors. From the results of an online survey involving 622 participants, we evaluate the usefulness of these explanations.
Kosetsu Tsukuda, Masataka Goto
RecSys2
2020 Query/Task Satisfaction and Grid-based Evaluation Metrics Under Different Image Search Intents
abstract
People use web image search with various search intents: from serious demands for work to just passing time by browsing images of a favorite actor. Such a diversity of intents can influence user satisfaction and evaluation metrics, both of which are important factors for providing a better image search environment. In this paper, we investigate this influence by using a publicly available one-month field study dataset. With respect to satisfaction, we take into consideration both query-level and task-level satisfaction provided by search users. Regarding the evaluation metrics, we use grid-based evaluation metrics that incorporate user behavior specific to image search. The results of our analysis indicate that both query/task satisfaction and grid-based evaluation metrics are influenced by the image search intent. Based on the results, we show possibilities to support users' search processes according to their search intents. We also discuss that there is still room for improvement in evaluation metrics through the development of intent-aware evaluation metrics in image search.
Kosetsu Tsukuda, Masataka Goto
SIGIR2
2020 Bayesian Singing Transcription Based on a Hierarchical Generative Model of Keys, Musical Notes, and F0 Trajectories
abstract
This article describes automatic singing transcription (AST) that estimates a human-readable musical score of a sung melody represented with quantized pitches and durations from a given music audio signal. To achieve the goal, we propose a statistical method for estimating the musical score by quantizing a trajectory of vocal fundamental frequencies (F0s) in the time and frequency directions. Since vocal F0 trajectories considerably deviate from the pitches and onset times of musical notes specified in musical scores, the local keys and rhythms of musical notes should be taken into account. In this article we propose a Bayesian hierarchical hidden semi-Markov model (HHSMM) that integrates a musical score model describing the local keys and rhythms of musical notes with an F0 trajectory model describing the temporal and frequency deviations of an F0 trajectory. Given an F0 trajectory, a sequence of musical notes, that of local keys, and the temporal and frequency deviations can be estimated jointly by using a Markov chain Monte Carlo (MCMC) method. We investigated the effect of each component of the proposed model and showed that the musical score model improves the performance of AST.
Ryo Nishikimi, Eita Nakamura, Masataka Goto, Katsutoshi Itoyama, Kazuyoshi Yoshii
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Sequential gallery for interactive visual design optimization
abstract
Visual design tasks often involve tuning many design parameters. For example, color grading of a photograph involves many parameters, some of which non-expert users might be unfamiliar with. We propose a novel user-in-the-loop optimization method that allows users to efficiently find an appropriate parameter set by exploring such a high-dimensional design space through much easier two-dimensional search subtasks. This method, called sequential plane search , is based on Bayesian optimization to keep necessary queries to users as few as possible. To help users respond to plane-search queries, we also propose using a gallery-based interface that provides options in the two-dimensional subspace arranged in an adaptive grid view. We call this interactive framework Sequential Gallery since users sequentially select the best option from the options provided by the interface. Our experiment with synthetic functions shows that our sequential plane search can find satisfactory solutions in fewer iterations than baselines. We also conducted a preliminary user study, results of which suggest that novices can effectively complete search tasks with Sequential Gallery in a photo-enhancement scenario.
Yuki Koyama 0001, Issei Sato, Masataka Goto
ACM Trans. Graph.3
2020 Audio-visual object removal in 360-degree videos
abstract
Abstract We present a novel concept audio–visual object removal in 360-degree videos, in which a target object in a 360-degree video is removed in both the visual and auditory domains synchronously. Previous methods have solely focused on the visual aspect of object removal using video inpainting techniques, resulting in videos with unreasonable remaining sounds corresponding to the removed objects. We propose a solution which incorporates direction acquired during the video inpainting process into the audio removal process. More specifically, our method identifies the sound corresponding to the visually tracked target object and then synthesizes a three-dimensional sound field by subtracting the identified sound from the input 360-degree video. We conducted a user study showing that our multi-modal object removal supporting both visual and auditory domains could significantly improve the virtual reality experience, and our method could generate sufficiently synchronous, natural and satisfactory 360-degree videos.
Ryo Shimamura, Yuki Koyama 0001, Takayuki Nakatsuka, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto, Shigeo Morishima
Vis. Comput.7
2019 Zero-mean Convolutional Network with Data Augmentation for Sound Level Invariant Singing Voice Separation
abstract
We address an issue of separating singing voices from polyphonic music signals regardless of sound level variance of the mixture input. Using a standard separation quality assessment tool BSS Eval 4.0, we found that the separation quality of a singing voice separation (SVS) system based on a dilatable Convolutional Neural Network (CNN) decreases under different sound levels. Even if this SVS system is comparable to state-of-the-art SVS systems, it is vulnerable to the issue of sound level variance. We therefore investigate four methods of making the CNN-based SVS system invariant to different sound levels - two types of data augmentation, frame normalization, and zero-mean convolution. By testing all 15 combinations of the four methods, we found that all combinations can improve the sound level invariance and analyzed the best combinations. To the best of our knowledge, this is the first SVS work systematically investigating sound level variance.
Kin Wah Edward Lin, Masataka Goto
ICASSP2
2019 Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention Mechanism
abstract
This paper describes neural singing transcription that estimates a sequence of musical notes directly from the audio signal of singing voice in an end-to-end manner without time-aligned training data. A conventional approach to singing transcription is to perform vocal F0 estimation followed by musical note estimation. The performance of this approach, however, is severely limited because the F0 estimation errors propagate to the note estimation step and rich acoustic information cannot be used. In addition, it is difficult and time-consuming to split continuous signals of singing voices into segments corresponding to musical notes for making precise time-aligned transcriptions. To solve these problems, we use an encoder-decoder model with an attention mechanism that can automatically learn an input-output alignment and mapping, even from non-aligned training data. The main challenge of our study is to estimate temporal categories (note values) in addition to instantaneous categories (pitches). We thus propose a novel loss function for the attention weights of time-aligned notes for semi-supervised alignment training. By gradually reducing the weight of the loss function, a better input-output alignment can be learned much more quickly. We showed that our method performed well for isolated singing voice in popular music.
Ryo Nishikimi, Eita Nakamura, Satoru Fukayama, Masataka Goto, Kazuyoshi Yoshii
ICASSP4
2019 Transdrums: A Drum Pattern Transfer System Preserving Global Pattern Structure
abstract
This paper presents TransDrums, which is a system that transfers drum patterns from a drum-pattern-source song (D-song) to a base song (B-song) and synthesizes the audio with the substituted drum pattern. Typical drum parts consist of multiple drum patterns that are concatenated to form a structure by, for example, inserting fill-in patterns at structural boundaries. The previous system that replaced the drum parts was not able to form such a structure. Therefore, we propose TransDrums, which extracts and transfers multiple drum patterns to form the structure. It takes two songs as the input and extracts multiple typical drum patterns from each song. It then makes pairs of those patterns between B-song and D-song and replaces them using the counterpart drum patterns to synthesize audio with the altered drum pattern. To achieve the key idea of properly replacing a drum phrase in B-song with that in D-song, it is necessary to model the structure of the drum parts by analyzing the transition probabilities between the typical drum patterns. The appropriate pairs are determined so that the sum of the Jensen-Shannon divergence between the transition probabilities is minimized. Our experimental results show that TransDrums can generate audio to change by altering the drum patterns with the structure.
Shun Sawada, Satoru Fukayama, Masataka Goto, Keiji Hirata 0001
ICASSP3
2019 Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model
abstract
This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical language model that represents the characteristics of each part in order to solve the ambiguity of part assignment and estimate a musically-natural score. We propose a factorial hidden semi-Markov model that consists of three language models corresponding to the three guitar parts (three latent chains) and an acoustic model of a mixture spectrogram (emission model). The language model for rhythm guitar represents a homophonic sequence of musical notes (chord sequence) and those for lead and bass guitars represent a monophonic sequence of musical notes in a higher and lower frequency range respectively. The acoustic model represents a spectrogram as a sum of low-rank spectrograms of the three guitar parts approximated by NMF. Given a spectrogram, we estimate the note sequences using Gibbs sampling. We show that our model outperforms a state-of-the-art multi-pitch detection method in the accuracy and naturalness of the transcribed scores.
Kentaro Shibata, Ryo Nishikimi, Satoru Fukayama, Masataka Goto, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii
ICASSP4
2019 Audio-Based Automatic Generation of a Piano Reduction Score by Considering the Musical Structure
Hirofumi Takamori, Takayuki Nakatsuka, Satoru Fukayama, Masataka Goto, Shigeo Morishima
MMM (2)4
2019 Query-by-Dancing: A Dance Music Retrieval System Based on Body-Motion Similarity
Shuhei Tsuchida, Satoru Fukayama, Masataka Goto
MMM (1)3
2019 DualDiv: diversifying items and explanation styles in explainable hybrid recommendation
abstract
In recommender systems, item diversification and explainable recommendations improve users' satisfaction. Unlike traditional explainable recommendations that display a single explanation for each item, explainable hybrid recommendations display multiple explanations for each item and are, therefore, more beneficial for users. When multiple explanations are displayed, one problem is that similar sets of explanation styles (ESs) such as user-based, item-based, and popularity-based may be displayed for similar items. Although item diversification has been studied well, the question of how to diversify the ESs remains underexplored. In this paper, we propose a method for diversifying ESs and a framework, called DualDiv, that recommends items by diversifying both the items and the ESs. Our experimental results show that DualDiv can increase the diversity of the items and the ESs without largely reducing the recommendation accuracy.
Kosetsu Tsukuda, Masataka Goto
RecSys2
2019 ABCPRec: Adaptively Bridging Consumer and Producer Roles for User-Generated Content Recommendation
abstract
In Web services dealing with user-generated content (UGC), a user can have two roles: a role of a consumer and that of a producer. Since most item recommendation models have only considered the role of a user as a consumer, how to leverage the two roles to improve UGC recommendation accuracy has been underexplored. In this paper, based on the state-of-the-art UGC recommendation method called CPRec (consumer and producer based recommendation), we propose ABCPRec (adaptively bridging CPRec). Unlike CPRec, which assumes that the two roles of a user are always related to each other, ABCPRec adaptively bridges the two roles according to the similarity between her nature as a consumer and that as a producer. This enables the model to learn each user's characteristics as both a consumer and a producer and to recommend items to each user more accurately. By using two real-world datasets, we showed that our proposed method significantly outperformed comparative methods in terms of AUC.
Kosetsu Tsukuda, Satoru Fukayama, Masataka Goto
SIGIR3
2019 Precomputed optimal one-hop motion transition for responsive character animation
abstract
Characters in interactive 3D applications are often animated by creating transitions from one motion clip to another in response to user input. It is not trivial, however, to achieve quick, natural-looking transitions between two arbitrary motion clips, especially when the two motions are dissimilar. To tackle this problem, we present a simple framework called optimal one-hop motion transition , which creates quick, natural-looking transitions on the fly without requiring careful manual specifications. The key ideas are (1) to insert a short intermediate motion clip, called a hop , between the source and destination motion clips, and (2) to select such a hop motion clip and its temporal alignment in an optimal way by solving a search problem. In the search problem, our framework tries to balance the naturalness of the resulting transitions and the responsiveness to user input. This search can be precomputed and the results can be stored in a lookup table, making the runtime cost to play an optimal transition negligible. We demonstrate that our framework is easily integrated into a widely used game engine, and that it greatly improves the quality of transitions in practical scenarios.
Yuki Koyama 0001, Masataka Goto
Vis. Comput.2
2018 OptiMo: Optimization-Guided Motion Editing for Keyframe Character Animation
abstract
The mission of animators is to create nuanced, high-quality character motions. To achieve this, the careful editing of animation curves---curves that determine how a series of keyframed poses are interpolated over time---is an important task. Manual editing affords full and precise control, but requires tedious and nonintuitive trials and errors. Numerical optimization can automate such exploration; however, automatic solutions cannot always be perfect, and it is difficult for animators to control optimization owing to its black-box behavior. In this paper, we present a new framework called optimization-guided motion editing, which is aimed at maintaining a sense of full control while utilizing the power of optimization. We have designed interactions and developed a set of mathematical formulations to enable them. We discuss the framework's potential by demonstrating several usage scenarios with our proof-of-concept system, named OptiMo.
Yuki Koyama 0001, Masataka Goto
CHI2
2018 Music Structure Boundary Detection and Labelling by a Deconvolution of Path-Enhanced Self-Similarity Matrix
abstract
We propose a music structure analysis method that converts a path-enhanced self-similarity matrix (SSM) into a block-enhanced SSM using non-negative matrix factor 2-D deconvolution (NMF2D). With a non-negative constraint, the deconvolution intuitively corresponds to the repeated stripes in the path-enhanced SSM. Then the block-enhanced SSM is constructed without any clustering technique. We fuse block-enhanced SSMs obtained using different parameters, resulting in better and more robust results. Discussion shows that the proposed method can be a potential tool for analysing music structure at different scales.
Tian Cheng 0001, Jordan B. L. Smith, Masataka Goto
ICASSP3
2018 Retrieval of Song Lyrics from Sung Queries
abstract
Retrieving the lyrics of a sung recording from a database of text documents is a research topic that has not received much attention so far. Such a retrieval system has many practical applications, e.g. for karaoke applications or for indexing large song databases by their lyric content. We present a new method for lyrics retrieval. An acoustic model trained on singing is used to obtain phoneme probabilities from sung queries, which are then mapped to phoneme sequences. These are compared against lines of textual lyrics in a large corpus in order to retrieve the best-matching song. The approach is tested on three sung datasets. Lyrics are retrieved from a set of 300 possible songs (12,000 lines of lyrics). The results are highly encouraging and could be used further to perform automatic lyrics alignment and keyword spotting for large databases of songs, or for retrieving lyrics from the internet.
Anna M. Kruspe, Masataka Goto
ICASSP2
2018 Instlistener: An Expressive Parameter Estimation System Imitating Human Performances of Monophonic Musical Instruments
abstract
We present InstListener, a system that takes an expressive monophonic solo instrument performance by a human performer as the input and imitates its audio recordings by using an existing MIDI (Musical Instrument Digital Interface) synthesizer. It automatically analyzes the input and estimates, for each musical note, expressive performance parameters such as the timing, duration, discrete semitone-level pitch, amplitude, continuous pitch contour, and continuous amplitude contour. The system uses an iterative process to estimate and update those parameters by analyzing both the input and output of the system so that the output from the MIDI synthesizer can be similar enough to the input. Our evaluation results showed that the iterative parameter estimation improved the accuracy of imitating of the input performance and thus increased the naturalness and expressiveness of the output performance.
Zhengshan Shi, Tomoyasu Nakano, Masataka Goto
ICASSP3
2018 Nonnegative Tensor Factorization for Source Separation of Loops in Audio
abstract
The prevalence of exact repetition in loop-based music makes it an opportune target for source separation. Nonnegative factorization approaches have been used to model the repetition of looped content, and kernel additive modeling has leveraged periodicity within a piece to separate looped background elements. We propose a novel method of leveraging periodicity in a factorization model: we treat the two-dimensional spectrogram as a three-dimensional tensor, and use nonnegative tensor factorization to estimate the component spectral templates, rhythms and loop recurrences in a single step. Testing our method on synthesized loop-based examples, we find that our algorithm mostly exceeds the performance of competing methods, with a reduction in execution cost. We discuss limitations of the algorithm as we demonstrate its potential to analyze larger and more complex songs.
Jordan B. L. Smith, Masataka Goto
ICASSP2
2018 Collaboration in N-th Order Derivative Creation
Shiori Hironaka, Kosetsu Tsukuda, Masahiro Hamasaki, Masataka Goto
ICWSM4
2018 Intelligent Music Interfaces
abstract
Automatic music-understanding technologies (automatic analysis of music signals) make possible the creation of intelligent music interfaces that enrich music experiences and open up new ways of listening to music. In the past, it was common to listen to music in a somewhat passive manner; in the future, people will be able to enjoy music in a more active manner by using music technologies. Listening to music through active interactions is called active music listening. In this keynote speech I first introduce active music listening interfaces demonstrating how end users can benefit from music-understanding technologies based on signal processing and/or machine learning. By analyzing the music structure (chorus sections), for example, the SmartMusicKIOSK interface enables people to access their favorite part of a song directly (skipping other parts) while viewing a visual representation of the song's structure. I then introduce our recent challenge of deploying such research-level music interfaces as web services open to the public. Those services augment people's understanding of music, enable music-synchronized control of computer-graphics animation and robots, and provide various bird's-eye views on a large music collection. In the future, further advances in music-understanding technologies and music interfaces based on them will make interaction between people and music even more active and enriching.
Masataka Goto
IUI1
2018 FocusMusicRecommender: A System for Recommending Music to Listen to While Working
abstract
This paper proposes FocusMusicRecommender, an automated system recommending background music to listen to while working. Recommendation systems matching user preferences have been widely researched even though research has shown that music that listeners strongly like is not suitable background music because it interferes with their concentration. FocusMusicRecommender plays songs that users may "neither like nor dislike" instead of "like very much." It is designed to by default summarize a song automatically so that users can give "like very much" feedback by pressing a "keep listening" button or "dislike very much" feedback by pressing a "skip" button. It uses this feedback, along with users» concentration levels estimated from their behavior history, to distinguish between the preference levels "like" and "like very much." It then estimates the preference levels of unplayed songs and selects the most suitable song by considering the user»s current concentration level. The effectiveness of the proposed feedback method and suitability of the recommendation results were verified experimentally and in user studies. Furthermore, it is confirmed that the proposed method can estimate the user»s concentration level more accurately than the previous methods.
Hiromu Yakura, Tomoyasu Nakano, Masataka Goto
IUI3
2018 Songle Sync: A Large-Scale Web-based Platform for Controlling Various Devices in Synchronization with Music
abstract
This paper presents Songle Sync, a web-based platform on which hundreds of Internet-connected devices - including smartphones, computers, and other physical computing devices - can be controlled to synchronize with music playback. It uses music-understanding technologies to dynamically synthesize music-driven multimedia performances from a musical piece of choice. To simultaneously control hundreds of devices, a conventional architecture keeps always-on connections between them. However, it does not scale and suffers from latency and jitter issues when there are various devices with potentially unstable networks. We address this with a novel autonomous control architecture in which each device is notified of forthcoming musical events (e.g., beats and chorus sections) to automatically drive various changes in multimedia performances. Moreover, we provide a development kit of an event-driven multimedia framework for JavaScript, example programs, and an interactive tutorial. To evaluate the platform, we compared latencies, jitters, and amounts of network traffic between ours and the conventional architecture. To examine use cases in the wild, we deployed the platform to drive over a hundred of a variety of devices. We also developed a web browser-based application for a multimedia performance with music playback. It provided audiences of hundreds with a bring-your-own-device experience of synchronized animations on smartphones. In addition, the development kit was used in a two-day hackathon. We report lessons learned from these studies and discuss the future of the Internet of Musical Things.
Jun Kato 0001, Masa Ogata, Takahiro Inoue, Masataka Goto
ACM Multimedia4
2018 A Melody-Conditioned Lyrics Language Model
abstract
Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano
NAACL-HLT4
2018 DeployGround: A Framework for Streamlined Programming from API playgrounds to Application Deployment
abstract
Interactive web pages for learning programming languages and application programming interfaces (APIs), called “playgrounds,” allow programmers to run and edit example codes in place. Despite the benefits of this live programming experience, programmers need to leave the playground at some point and restart the development from scratch in their own programming environments. This paper proposes “DeployGround,” a framework for creating web-based tutorials that streamlines learning APIs on playgrounds and developing and deploying applications. As a case study, we created a web-based tutorial for browser-based and Node.js-based JavaScript APIs. A preliminary user study found appreciation of the streamlined and social workflow of the DeplovGround framework.
Jun Kato 0001, Masataka Goto
VL/HCC2
2018 Decomposing Images into Layers with Advanced Color Blending
abstract
Abstract Digital paintings are often created by compositing semi‐transparent layers using various advanced color‐blend modes, such as “color‐burn,” “multiply,” and “screen,” which can produce interesting non‐linear color effects. We propose a method of decomposing an input image into layers with such advanced color blending. Unlike previous layer‐decomposition methods, which typically support only linear color‐blend modes, ours can handle any user‐specified color‐blend modes. To enable this, we generalize a previous color‐unblending formulation, in which only a specific layering model was considered. We also introduce several techniques for adapting our generalized formulation to practical use, such as the post‐processing for refining smoothness. Our method lets users explore possible decompositions to find the one that matches for their purposes by manipulating the target color‐blend mode and desired color distribution for each layer, as well as the number of layers. Thus, the output of our method is a layered, easily editable image composition organized in a way that digital artists are familiar with. Our method is useful for remixing existing illustrations, flexibly editing single‐layer paintings, and bringing physically painted media (e.g., oil paintings) into a digital workflow.
Yuki Koyama 0001, Masataka Goto
Comput. Graph. Forum2
2017 Automatic System for Editing Dance Videos Recorded Using Multiple Cameras
Shuhei Tsuchida, Satoru Fukayama, Masataka Goto
ACE3
2017 f3.js: A Parametric Design Tool for Physical Computing Devices for Both Interaction Designers and End-users
abstract
Although the exploration of design alternatives is crucial for interaction designers and customization is required for end-users, the current development tools for physical computing devices have focused on single versions of an artifact. We propose the parametric design of devices including their enclosure layouts and programs to address this issue. A Web-based design tool called f3.js is presented as an example implementation, which allows devices assembled from laser-cut panels with sensors and actuator modules to be parametrically created and customized. It enables interaction designers to write code with dedicated APIs, declare parameters, and interactively tune them to produce the enclosure layouts and programs. It also provides a separate user interface for end-users that allows parameter tuning and dynamically generates instructions for device assembly. The parametric design approach and the tool were evaluated through two user studies with interaction designers, university students, and end-users.
Jun Kato 0001, Masataka Goto
Conference on Designing Interactive Systems2
2017 Strummer: An interactive guitar chord practice system
abstract
Musical instrument playing is a skill many people desire to acquire, and learners now have a wide variety of learning materials. However, their volume is enormous, and novice learners may easily get lost in which songs they should practice first. We develop Strummer: an interactive multimedial system for guitar practice. Strummer provide data-driven and personalized practice for learners in order to identify important and easy-to-learn chords and songs. This practice design is intended to encourage smooth skill transfers to songs that learners even have not seen. Our user study confirms the benefits and possible improvements of the Strummer system. In particular, participants expressed their positive impressions on lessons provided by the system.
Shunya Ariga, Masataka Goto, Koji Yatani
ICME2
2017 Classifying derivative works with search, text, audio and video features
abstract
Users of video-sharing sites often search for derivative works of music, such as live versions, covers, and remixes. Audio and video content are both important for retrieval: “karaoke” specifies audio content (instrumental version) and video content (animated lyrics). Although YouTube's text search is fairly reliable, many search results do not match the exact query. We introduce an algorithm to classify YouTube videos by category of derivative work. Based on a standard pipeline for video-based genre classification, it combines search, text, and video features with a novel set of audio features derived from audio fingerprints. A baseline approach is outperformed by the search and text features alone, and combining these with video and audio features performs best of all, reducing the audio content error rate from 25% to 15%.
Jordan B. L. Smith, Masahiro Hamasaki, Masataka Goto
ICME3
2017 LyriSys: An Interactive Support System for Writing Lyrics Based on Topic Transition
abstract
This paper presents LyriSys, a novel lyric-writing support system. Previous systems for lyric writing can fully automatically only generate a single line of lyrics that satisfies given constraints on accent and syllable patterns or an entire lyric. In contrast to such systems, LyriSys allows users to create and revise their work incrementally in a trial-and-error manner. Through fine-grained interactions with the system, the user can create the specifications of the musical structure and the story of the lyrics in terms of the verse-bridge-chorus structure, the number of lines, words and syllables, and most importantly, the transition over semantic topics such as "scene", "dark" and "sweet love". This paper provides an overview of the design of the system and its user interface and describes how the writing process is guided by a state-of-the-art probabilistic generative topic model that is trained without supervision. The system works for both Japanese and English.
Kento Watanabe, Yuichiroh Matsubayashi, Kentaro Inui, Tomoyasu Nakano, Satoru Fukayama, Masataka Goto
IUI6
2017 Taste or Addiction?: Using Play Logs to Infer Song Selection Motivation
Kosetsu Tsukuda, Masataka Goto
PAKDD (2)2
2017 QueryShare: Working Together to Facilitate Exploratory Multimedia Searches without Skill in Creating
abstract
This paper describes a music exploratory search interface called QueryShare, which provides query searching and recommendation functions for query sharing among users. Most people are not expert users who know how to use various music metadata that include automatically estimated musical features to represent their own information needs as a query. Therefore, it is difficult for them to enter a complex query for music content retrieval. The original feature of our proposed interface is to make users share every query as a public web page. This feature enables users to use search queries, find recommended queries, and revise existing queries. Beginners can use an applicable query, which is more complicated than they might create on their own. Experts can readily reuse a query (web page) of their own making. The interface assists users in finding results for interesting queries and in performing music exploratory search without skills to create complex queries. We developed a prototype system as a web application for music videos on the most popular Japanese video sharing service. Users can search for over 360,000 music videos using our system. Results of a preliminary user study demonstrated that users found the query creation interesting and that they were interested in seeing and using queries created by other users, although some users hesitated to share their queries.
Masahiro Hamasaki, Masataka Goto
OpenSym2
2016 Why Did You Cover That Song?: Modeling N-th Order Derivative Creation with Content Popularity
abstract
Many amateur creators now create derivative works and put them on the web. Although there are several factors that inspire the creation of derivative works, such factors cannot usually be observed on the web. In this paper, we propose a model for inferring latent factors from sequences of derivative work posting events. We assume a sequence to be a stochastic process incorporating the following three factors: (1) the original work's attractiveness, (2) the original work's popularity, and (3) the derivative work's popularity. To characterize content popularity, we use content ranking data and incorporate rank-biased popularity based on the creators' browsing behavior. Our main contributions are three-fold: (1) to the best of our knowledge, this is the first study modeling derivative creation activity, (2) by using a real-world dataset of music-related derivative work creation to evaluate our model, we showed the effectiveness of adopting all three factors to model derivative creation activity and onsidering creators' browsing behavior, and (3) we carried out qualitative experiments and showed that our model is useful in analyzing derivative creation activity in terms of category characteristics, temporal development of factors that trigger derivative work posting events, etc.
Kosetsu Tsukuda, Masahiro Hamasaki, Masataka Goto
CIKM3
2016 Modeling Discourse Segments in Lyrics Using Repeated Patterns
abstract
This study proposes a computational model of the discourse segments in lyrics to understand and to model the structure of lyrics. To test our hypothesis that discourse segmentations in lyrics strongly correlate with repeated patterns, we conduct the first large-scale corpus study on discourse segments in lyrics. Next, we propose the task to automatically identify segment boundaries in lyrics and train a logistic regression model for the task with the repeated pattern and textual features. The results of our empirical experiments illustrate the significance of capturing repeated patterns in predicting the boundaries of discourse segments in lyrics.
Kento Watanabe, Yuichiroh Matsubayashi, Naho Orita, Naoaki Okazaki, Kentaro Inui, Satoru Fukayama, Tomoyasu Nakano, Jordan B. L. Smith, Masataka Goto
COLING9
2016 Music emotion recognition with adaptive aggregation of Gaussian process regressors
abstract
This paper describes a novel method for estimating the emotions elicited by a piece of music from its acoustic signals. Previous research in this field has centered on finding effective acoustic features and regression methods to relate features to emotions. The state-of-the-art method is based on a multi-stage regression, which aggregates the results from different regressors trained with training data. However, after training, the aggregation happens in a fixed way and cannot be adapted to acoustic signals with different musical properties. We propose a method that adapts the aggregation by taking into account new acoustic signal inputs. Since we cannot know the emotions elicited by new inputs beforehand, we need a way of adapting the aggregation weights. We do so by exploiting the deviation observed in the training data using Gaussian process regressions. We confirmed with an experiment comparing different aggregation approaches that our adaptive aggregation is effective in improving recognition accuracy.
Satoru Fukayama, Masataka Goto
ICASSP2
2016 An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer
abstract
This paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice timbre expression words, such as "Age" and "Gender", and they usually need to be manually assigned to individual singers' singing voices through listening. To make it possible to automatically estimate them from given singer's singing voices, an acoustic feature to well capture only each singer's voice timbre is extracted with a Gaussian mixture model trained using parallel data between singing voices sung by many pre-stored target singers and same voices sung by a reference singer. Then, the voice timbre evaluation values are estimated from the extracted feature using regression models. The experimental results showed that the proposed method is capable of accurately estimating those values for some expression words, such as "Age" and "Gender", and nonlinear regression is effective for the expression words, "Powerfulness" and "Uniqueness."
Soichi Yamane, Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001
ICASSP5
2016 Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separation
abstract
This paper presents a robust variant of nonnegative matrix factorization (NMF) based on complex Student's t distributions (t-NMF) for source separation of single-channel audio signals. The Itakura-Saito divergence NMF (Gaussian NMF) is justified for this purpose under an assumption that the complex spectra of source signals and those of the mixture signal are complex Gaussian distributed (the additiv-ity of power spectra holds). In fact, however, the source spectra are often heavy-tailed distributed. When the source spectra are complex Cauchy distributed, for example, the mixture spectra are also complex Cauchy distributed (the additivity of amplitude spectra holds). Using the complex t distribution that includes the complex Gaussian and Cauchy distributions as its special cases, we propose t-NMF as a unified extension of Gaussian NMF and Cauchy NMF. Furthermore, we propose the corresponding variant of positive semidefinite tensor factorization based on multivariate complex t distributions (t-PSDTF). The experimental results showed that while t-NMF and t-PSDTF were comparative to Gaussian counterparts in terms of peak performance, they worked much better on average because they are insensitive to initialization and tend to avoid local optima.
Kazuyoshi Yoshii, Katsutoshi Itoyama, Masataka Goto
ICASSP3
2016 A soundtrack generation system to synchronize the climax of a video clip with music
abstract
In this paper, we present a soundtrack generation system that can automatically add a soundtrack with the length and climax points aligned to those of a video clip. Adding a soundtrack to a video clip is an important process in video editing. Editors tend to add chorus sections to the climax points of the video clip by replacing and concatenating musical segments. However, this process is time-consuming. Our system automatically detects climaxes of both the video clips and music based on feature extraction and analysis. This enables the system to add a soundtrack in which the climax is synchronized to the climax of the video clip. We evaluated the generated soundtracks through a subjective evaluation.
Haruki Sato, Tatsunori Hirai, Tomoyasu Nakano, Masataka Goto, Shigeo Morishima
ICME4
2016 PlaylistPlayer: An Interface Using Multiple Criteria to Change the Playback Order of a Music Playlist
abstract
We propose a novel interface that allows the user to interactively change the playback order of multiple songs by choosing one or more criteria. The criteria include not only the song's title and artist name but also its content automatically estimated by music/singing signal processing and artist-level social analysis. The artist-level social information is discovered from Wikipedia and DBpedia. With regard to manipulating playback order, existing interfaces typically allow the user to change it manually or automatically by choosing one of a few types of criteria. The proposed interface, on the other hand, deals with nine properties and multiple integrations of them (e.g., vocal gender and beats per minute). To realize the ordering by multiple criteria, a distance matrix is computed from the criteria vectors and is then used to estimate paths for ascending, descending, and random orders by applying principle component analysis or to estimate a path for a smooth order by solving the travelling salesman problem.
Tomoyasu Nakano, Jun Kato 0001, Masahiro Hamasaki, Masataka Goto
IUI4
2015 TextAlive: Integrated Design Environment for Kinetic Typography
abstract
This paper presents TextAlive, a graphical tool that allows interactive editing of kinetic typography videos in which lyrics or transcripts are animated in synchrony with the corresponding music or speech. While existing systems have allowed the designer and casual user to create animations, most of them do not take into account synchronization with audio signals. They allow predefined motions to be applied to objects and parameters to be tweaked, but it is usually impossible to extend the predefined set of motion algorithms within these systems. We therefore propose an integrated design environment featuring (1) GUIs that designers can use to create and edit animations synchronized with audio signals, (2) integrated tools that programmers can use to implement animation algorithms, and (3) a framework for bridging the interfaces for designers and programmers. A preliminary user study with designers, programmers, and casual users demonstrated its capability in authoring various kinetic typography videos.
Jun Kato 0001, Tomoyasu Nakano, Masataka Goto
CHI3
2015 A feedback framework for improved chord recognition based on NMF-based approximate note transcription
abstract
This paper presents a feedback framework that can improve chord recognition for music audio signals by performing approximate note transcription with Bayesian non-negative matrix factorization (NMF) using prior knowledge on chords. Although the names and note compositions of chords are intrinsically linked with each other (e.g., C major chords are highly likely to include C, E, and G notes, and those notes are highly likely to be in C major chords), chord recognition and note transcription (multipitch analysis) have been studied independently. To solve this chicken-and-egg problem, our framework iterates chord recognition and approximate note transcription using each other's results. More specifically, we first perform approximate note transcription based on Bayesian NMF that forces basis spectra to respectively correspond to different semitone-level pitches covering the whole range. We then execute chord recognition based on Bayesian hidden Markov models (HMMs) that use chroma features obtained from the activation patterns of those pitches. To improve note transcription, we again perform Bayesian NMF that encourages certain kinds of pitches in each chord region to be activated. Experimental results showed that our feedback framework gradually improved the accuracy of chord recognition.
Satoshi Maruo, Kazuyoshi Yoshii, Katsutoshi Itoyama, Matthias Mauch, Masataka Goto
ICASSP5
2015 Songle Widget: Making Animation and Physical Devices Synchronized with Music Videos on the Web
abstract
This paper describes a web-based multimedia development framework, Songle Widget, that makes it possible to control computer-graphic animation and physical devices such as lighting devices and robots in synchronization with music publicly available on the web. To avoid the difficulty of time-consuming manual annotation, Songle Widget makes it easy to develop web-based applications with rigid music synchronization by leveraging music-understanding technologies. Four types of musical elements (music structure, hierarchical beat structure, melody line, and chords) have been automatically annotated for more than 920,000 songs on music-or video-sharing services and can readily be used by music-synchronized applications. Since errors are inevitable when elements are annotated automatically, Songle Widget takes advantage of a user-friendly crowdsourcing interface that enables users to correct them. This is effective when applications require error-free annotation. We made Songle Widget open to the public, and its capabilities and usefulness have been demonstrated in seven music-synchronized applications.
Masataka Goto, Kazuyoshi Yoshii, Tomoyasu Nakano
ISM1
2015 Musical Similarity and Commonness Estimation Based on Probabilistic Generative Models
abstract
This paper proposes a novel concept we call musical commonness, which is the similarity of a song to a set of songs, in other words, its typicality. This commonness can be used to retrieve representative songs from a song set (e.g., songs released in the 80s or 90s). Previous research on musical similarity has compared two songs but has not evaluated the similarity of a song to a set of songs. The methods presented here for estimating the similarity and commonness of polyphonic musical audio signals are based on a unified framework of probabilistic generative modeling of four musical elements (vocal timbre, musical timbre, rhythm, and chord progression). To estimate the commonness, we use a generative model trained from a song set instead of estimating musical similarities of all possible song-pairs by using a model trained from each song. In experimental evaluation, we used 3278 popular music songs. Estimated song-pair similarities are comparable to ratings by a musician at the 0.1% significance level for vocal and musical timbre, at the 1% level for rhythm, and the 5% level for chord progression. Results of commonness evaluation show that the higher the musical commonness is, the more similar a song is to songs of a song set.
Tomoyasu Nakano, Kazuyoshi Yoshii, Masataka Goto
ISM3
2015 ExploratoryVideoSearch: A Music Video Search System Based on Coordinate Terms and Diversification
abstract
Many people search and watch music videos on video sharing websites. Although a vast variety of videos are uploaded, current search systems on video sharing websites allow users to search a limited range of music videos for an input query. Especially when a user does not have enough knowledge of a query, the problem gets worse because the user cannot customize the range by changing the query or adding some keywords to the original query. In this paper, we propose a music video search system, called ExploratoryVideoSearch, that is coordinate term aware and diversity aware. Our system focuses on artist name queries to search videos on YouTube and has two novel functions: (1) given an artist name query, the system shows a search result for the artist as well as those for its coordinate terms, and (2) the system diversifies search results for the query and its coordinate terms, and allows users to interactively change the diversity level. Coordinate terms are obtained by utilizing the Million Song Dataset and Wikipedia, while search results are diversified based on tags attached to YouTube music videos. ExploratoryVideoSearch enables users to search a wide variety of music videos without requiring deep knowledge about a query.
Kosetsu Tsukuda, Masataka Goto
ISM2
2015 AutoGuitarTab: Computer-Aided Composition of Rhythm and Lead Guitar Parts in the Tablature Space
abstract
We present AutoGuitarTab, a system for generating realistic guitar tablature given an input symbolic chord and key sequence. Our system consists of two modules: AutoRhythmGuitar and AutoLeadGuitar. The first of these generates rhythm guitar tablatures which outline the input chord sequence in a particular style (using Markov chains to ensure playability) and performs a structural analysis to produce a structurally consistent composition. AutoLeadGuitar generates lead guitar parts in distinct musical phrases, guiding the pitch classes towards chord tones and steering the evolution of the rhythmic and melodic intensity according to user preference. Experimentally, we uncover musician-specific trends in guitar playing style, and demonstrate our system's ability to produce playable, realistic and style-specific tablature using a combination of algorithmic, user-surveyed and expert evaluation techniques.
Matt McVicar, Satoru Fukayama, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Automated choreography synthesis using a Gaussian process leveraging consumer-generated dance motions
abstract
We propose a novel method of automatically generating dance choreography using machine learning. In a typical approach to automatic choreography, a dance is constructed by concatenating segments of existing dances which are maximally correlated to the target audio features with connectivity constraints. However, researchers using this approach are unable to produce dances with much variety, since the set of examples used in these experiments (usually motion-capture of existing choreographies) is limited and costly to produce. To solve this issue, we propose a probabilistic model which maps beat structures to dance movements using a Gaussian process, trained with a large amount of consumer-generated dance motion obtained from the web. The main contribution of our work is the combination of two approaches: the previously mentioned correlation based approach which seeks for relationships between music and dance, and a machine learning approach which is based on human motion modeling. Inspection of the generated dances proves that our method can generate choreographies with different characters by switching the training dataset, and highlights opportunities in training with further dance motions on the web to generate more expressive dance choreography.
Satoru Fukayama, Masataka Goto
Advances in Computer Entertainment2
2014 Sharedo: to-do list interface for human-agent task sharing
abstract
In this paper, we propose a to-do list interface for sharing tasks between human and multiple agents including robots and software personal assistants. While much work on software architectures aims to achieve efficient (semi-)autonomous task coordination among human and agents, little work on user interfaces can be found for user-oriented flexible task coordination. Instead, most of the existing human-agent interfaces are designed to command a single agent to handle specific kinds of tasks. Meanwhile, our interface is designed to be a platform to share any kinds of tasks between users and multiple agents. When agents can handle the task, they ask for details and permission to execute it. Otherwise, they try supporting users or just keep silent. New tasks can be registered not only by humans but also by agents when errors occur that can only be fixed by human users. We present the interaction design and implementation of the interface, Sharedo, with three example agents, followed by brief user feedback collected from a preliminary user study.
Jun Kato 0001, Daisuke Sakamoto, Takeo Igarashi, Masataka Goto
HAI4
2014 Regression approaches to perceptual age control in singing voice conversion
abstract
The perceptual age of a singing voice is the age of the singer as perceived by the listener, and is one of the notable characteristics that determines perceptions of a song. In this paper, we describe a novel voice timbre control technique based on the perceptual age for singing voice conversion (SVC). Singers can sing expressively by controlling prosody and voice timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome the limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice timbre of an arbitrary source singer into that of an arbitrary target singer. However, it is still difficult to intuitively control singing voice characteristics by manipulating parameters corresponding to specific physical traits, such as gender and age. In this paper, we develop a technique for controlling the voice timbre based on perceptual age that maintains the singer's individuality. The experimental results show that the proposed voice timbre control method makes it possible to change the singer's perceptual age while not having an adverse effect on the perceived individuality.
Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001
ICASSP4
2014 Leveraging repetition for improved automatic lyric transcription in popular music
abstract
Transcribing lyrics from musical audio is a challenging research problem which has not benefited from many advances made in the related field of automatic speech recognition, owing to the prevalent musical accompaniment and differences between the spoken and sung voice. However, one aspect of this problem which has yet to be exploited by researchers is that significant portions of the lyrics will be repeated throughout the song. In this paper we investigate how this information can be leveraged to form a consensus transcription with improved consistency and accuracy. Our results show that improvements can be gained using a variety of techniques, and that relative gains are largest under the most challenging and realistic experimental conditions.
Matt McVicar, Daniel P. W. Ellis, Masataka Goto
ICASSP3
2014 Timbre replacement of harmonic and drum components for music audio signals
abstract
This paper presents a system that allows users to customize an audio signal of polyphonic music (input), without using musical scores, by replacing the frequency characteristics of harmonic sounds and the timbres of drum sounds with those of another audio signal of polyphonic music (reference). To develop the system, we first use a method that can separate the amplitude spectra of the input and reference signals into harmonic and percussive spectra. We characterize frequency characteristics of the harmonic spectra by two envelopes tracing spectral dips and peaks roughly, and the input harmonic spectra are modified such that their envelopes become similar to those of the reference harmonic spectra. The input and reference percussive spectrograms are further decomposed into those of individual drum instruments, and we replace the timbres of those drum instruments in the input piece with those in the reference piece. Through the subjective experiment, we show that our system can replace drum timbres and frequency characteristics adequately.
Tomohiko Nakamura, Hirokazu Kameoka, Kazuyoshi Yoshii, Masataka Goto
ICASSP4
2014 Vocal timbre analysis using latent Dirichlet allocation and cross-gender vocal timbre similarity
abstract
This paper presents a vocal timbre analysis method based on topic modeling using latent Dirichlet allocation (LDA). Although many works have focused on analyzing characteristics of singing voices, none have dealt with “latent” characteristics (topics) of vocal timbre, which are shared by multiple singing voices. In the work described in this paper, we first automatically extracted vocal timbre features from polyphonic musical audio signals including vocal sounds. The extracted features were used as observed data, and mixing weights of multiple topics were estimated by LDA. Finally, the semantics of each topic were visualized by using a word-cloud-based approach. Experimental results for a singer identification task using 36 songs sung by 12 singers showed that our method achieved a mean reciprocal rank of 0.86. We also proposed a method for estimating cross-gender vocal timbre similarity by generating pitch-shifted (frequency-warped) signals of every singing voice. Experimental results for a cross-gender singer retrieval task showed that our method discovered interesting similar pitch-shifted singers.
Tomoyasu Nakano, Kazuyoshi Yoshii, Masataka Goto
ICASSP3
2014 Cultivating vocal activity detection for music audio signals in a circulation-type crowdsourcing ecosystem
abstract
This paper presents a crowdsourcing-based self-improvement framework of vocal activity detection (VAD) for music audio signals. A standard approach to VAD is to train a vocal-and-non-vocal classifier by using labeled audio signals (training set) and then use that classifier to label unseen signals. Using this technique, we have developed an online music-listening service called Songle that can help users better understand music by visualizing automatically estimated vocal regions and pitches of arbitrary songs existing on the Web. The accuracy of VAD is limited, however, because in general the acoustic characteristics of the training set are different from those of real songs on the Web. To overcome this limitation, we adapt a classifier by leveraging vocal regions and pitches corrected by volunteer users. UnlikeWikipedia-type crowdsourcing, our Songle-based framework can amplify user contributions: error corrections made for a limited number of songs improve VAD for all songs. This gives better music listening experiences to all users as non-monetary rewards.
Kazuyoshi Yoshii, Hiromasa Fujihara, Tomoyasu Nakano, Masataka Goto
ICASSP4
2014 A Multi-Touch DJ Interface with Remote Audience Feedback
abstract
Current DJ interfaces lack direct support for typical digital communication common in social media. We present a novel DJ interface for live internet broadcast performances with remote audience feedback integration. Our multi-touch interface is designed for a table top display, featuring a time-line based visualization. Two studies are presented involving seven DJs, culminating in four live broadcasts gathering and analyzing data to better understand both the DJ and audience perspective. This study is one of the first to look closer at DJs and remote audiences. We present useful insight for future interaction design between DJs and remote audiences, and interface integrated audience feedback.
Lasse Farnung Laursen, Masataka Goto, Takeo Igarashi
ACM Multimedia2
2014 Modeling Structural Topic Transitions for Automatic Lyrics Generation
Kento Watanabe, Yuichiroh Matsubayashi, Kentaro Inui, Masataka Goto
PACLIC4
2014 AutoMashUpper: automatic creation of multi-song music mashups
abstract
In this paper we present a system, AutoMashUpper, for making multi-song music mashups. Central to our system is a measure of “mashability” calculated between phrase sections of an input song and songs in a music collection. We define mashability in terms of harmonic and rhythmic similarity and a measure of spectral balance. The principal novelty in our approach centres on the determination of how elements of songs can be made fit together using key transposition and tempo modification, rather than based on their unaltered properties. In this way, the properties of two songs used to model their mashability can be altered with respect to transformations performed to maximize their perceptual compatibility. AutoMashUpper has a user interface to allow users to control the parameterization of the mashability estimation. It allows users to define ranges for key shifts and tempo as well as adding, changing or removing elements from the created mashups. We evaluate AutoMashUpper by its ability to reliably segment music signals into phrase sections, and also via a listening test to examine the relationship between estimated mashability and user enjoyment.
Matthew E. P. Davies, Philippe Hamel, Kazuyoshi Yoshii, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Infinite kernel linear prediction for joint estimation of spectral envelope and fundamental frequency
abstract
This paper presents a new probabilistic formulation of linear prediction (LP) for jointly estimating the spectral envelope and fundamental frequency (F0) of a speech signal. A main problem of classical LP is that the peaks of the estimated envelope are highly biased toward the harmonic partials of a speech spectrum. To solve this problem, we propose a nonparametric Bayesian model called infinite kernel linear prediction (IKLP) based on a Gaussian process with multiple kernel learning. Our model can represent the periodicity of a speech signal by using a weighted sum of infinitely many periodic kernels that correspond to different F0s. We put a gamma process prior on the positive weights of those kernels and perform sparse learning to determine a predominant kernel indicating the F0 at the same time of spectral envelope estimation. The experimental results showed that our model can estimate spectral envelopes and F0s of speech and singing signals while identifying pitched segments.
Kazuyoshi Yoshii, Masataka Goto
ICASSP2
2013 Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture Signals
abstract
This paper presents a new class of tensor factorization called positive semidefinite tensor factorization (PSDTF) that decomposes a set of positive semidefinite (PSD) matrices into the convex combinations of fewer PSD basis matrices. PSDTF can be viewed as a natural extension of nonnegative matrix factorization. One of the main problems of PSDTF is that an appropriate number of bases should be given in advance. To solve this problem, we propose a nonparametric Bayesian model based on a gamma process that can instantiate only a limited number of necessary bases from the infinitely many bases assumed to exist. We derive a variational Bayesian algorithm for closed-form posterior inference and a multiplicative update rule for maximum-likelihood estimation. We evaluated PSDTF on both synthetic data and real music recordings to show its superiority.
Kazuyoshi Yoshii, Ryota Tomioka, Daichi Mochihashi, Masataka Goto
ICML (3)4
2013 Evaluation of a singing voice conversion method based on many-to-many eigenvoice conversion
abstract
In this paper, we evaluate our proposed singing voice conver-sion method from various perspectives. To enable singers to freely control their voice timbre of singing voice, we have pro-posed a singing voice conversion method based on many-to-many eigenvoice conversion (EVC) that enables to convert the voice timbre of an arbitrary source singer into that of another arbitrary target singer using a probabilistic model. Further-more, to easily develop training data consisting of multiple par-allel data sets between a single reference singer and many other singers, a technique for efficiently and effectively generating the parallel data sets from nonparallel singing voice data sets of many singers using a singing-to-singing synthesis system have been proposed. However, we have never conducted sufficient investigations into the effectiveness of these proposed methods. In this paper, we conduct both objective and subjective eval-uations to carefully investigate the effectiveness of proposed methods. Moreover, the differences between singing voice con-version and speaking voice conversion are also analyzed. Ex-perimental results show that our proposed method succeeds in enabling people to control their own voice timbre by using only an extremely small amount of the target singing voice. Index Terms: singing voice, voice conversion, eigenvoice con-version, singing-to-singing synthesis, performance evaluation
Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001
INTERSPEECH4
2013 An investigation of acoustic features for singing voice conversion based on perceptual age
abstract
In this paper, we investigate the acoustic features that can be modified to control the perceptual age of a singing voice. Singers can sing expressively by controlling prosody and vocal timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome this limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, it is still difficult to intu-itively control singing voice characteristics by manipulating pa-rameters corresponding to specific physical traits, such as gen-der and age. In this paper, we focus on controlling the perceived age of the singer and, as a first step, perform an investigation of the factors that play a part in the listener’s perception of the singer’s age. The experimental results demonstrate that 1) the perceptual age of singing voices corresponds relatively well to the actual age of the singer, 2) speech analysis/synthesis pro-cessing and statistical voice conversion processing don’t cause adverse effects on the perceptual age of singing voices, and 3) prosodic features have a larger effect on the perceptual age than spectral features.
Kazuhiro Kobayashi, Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001
INTERSPEECH5
2013 Multimedia information retrieval: music and audio
abstract
No abstract available.
Markus Schedl, Emilia Gómez, Masataka Goto
ACM Multimedia3
2013 Songrium: a music browsing assistance service based on visualization of massive open collaboration within music content creation community
abstract
This paper describes a music browsing assistance service, Songrium (http://songrium.jp), that helps a user enjoy songs while seeing visualization of open collaboration. Songrium focuses on open collaboration for music content creation on the most popular Japanese video-sharing service. Since this open collaboration generates more than half a million video clips with a rich variety of music content, we call it massive open collaboration. To develop a shared understanding of this collaboration we have analyzed, we developed Songrium that visualizes relations among both original songs and derivative works generated from the collaboration. Songrium also features a social annotation framework to verbalize and share various relations among songs, and a flexible ranking mechanism to find interesting songs. After we launched Songrium in August 2012, more than 7,000 users have used our service in which over 98,000 songs and 520,000 derivative works have automatically been registered. We hope Songrium will not only encourage creators to create more derivative works, but also attract consumers to participate in the collaboration as creators.
Masahiro Hamasaki, Masataka Goto
OpenSym2
2012 VocaListener and VocaWatcher: Imitating a human singer by using signal processing
abstract
In this paper, we describe three singing information processing systems, VocaListener, VocaListener2, and VocaWatcher, that imitate singing expressions of the voice and face of a human singer. VocaListener can synthesize natural singing voices by analyzing and imitating the pitch and dynamics of the human singing. VocaListener2 imitates temporal timbre changes in addition to the pitch and dynamics. In synchronization with the synthesized singing voices, VocaWatcher can generate realistic facial motions of a humanoid robot, the HRP-4C, by analyzing and imitating facial motions of a human singing that are recorded by a single video camera. These systems that focus on “imitation” are not only promising for representing human-like naturalness, but also useful for providing intuitive control means.
Masataka Goto, Tomoyasu Nakano, Shuuji Kajita, Yosuke Matsusaka, Shinichiro Nakaoka, Kazuhito Yokoi
ICASSP1
2012 Unsupervised music understanding based on nonparametric Bayesian models
abstract
This paper presents a new research framework for unsupervised music understanding. Our goal is to recognize musical notes from polyphonic audio signals and simultaneously induce grammatical patterns from the recognized notes by integrating probabilistic acoustic and language models. Given music audio signals, both models could be jointly trained in a self-organizing manner without manually specifying the numbers of musical notes and grammatical patterns. In this paper, we introduce our nonparametric Bayesian acoustic and language models for multipitch analysis and chord progression analysis and discuss issues for integrating these models. We then provide a novel overview of various acoustic and language models whose underlying concepts are useful for implementing the framework.
Kazuyoshi Yoshii, Masataka Goto
ICASSP2
2012 PodCastle: Collaborative Training of Language Models on the Basis of Wisdom of Crowds
abstract
This paper presents a language-model training method for improving automatic transcription of online spoken contents. Unlike previously studied LVCSR tasks such as broadcast news and lectures, large-sized task-specific corpora for training language models cannot be prepared and used in recognition because of the diversity of topics, vocabularies, and speaking styles. To overcome difficulties in preparing such task-specific language models in advance, we propose collaborative training of language models on the basis of wisdom of crowds. On our public web service for LVCSR-based spoken document retrieval PodCastle, over half a million recognition errors were corrected by anonymous users. By leveraging such corrected transcriptions, component language models for various topics can be built and dynamically mixed to generate an appropriate language model for each podcast episode in an unsupervised manner. Experimental results with Japanese podcasts showed that the mixed languages models significantly reduced the word error rate. Index Terms: web service, LVCSR, language modeling, wisdom of crowds, error correction
Jun Ogata, Masataka Goto
INTERSPEECH2
2012 Integrating Additional Chord Information Into HMM-Based Lyrics-to-Audio Alignment
abstract
Aligning lyrics to audio has a wide range of applications such as the automatic generation of karaoke scores, song-browsing by lyrics, and the generation of audio thumbnails. Existing methods are restricted to using only lyrics and match them to phoneme features extracted from the audio (usually mel-frequency cepstral coefficients). Our novel idea is to integrate the textual chord information provided in the paired chords-lyrics format known from song books and Internet sites into the inference procedure. We propose two novel methods that implement this idea: First, assuming that all chords of a song are known, we extend a hidden Markov model (HMM) framework by including chord changes in the Markov chain and an additional audio feature (chroma) in the emission vector; second, for the more realistic case in which some chord information is missing, we present a method that recovers the missing chord information by exploiting repetition in the song. We conducted experiments with five changing parameters and show that with accuracies of 87.5% and 76.7%, respectively, both methods perform better than the baseline with statistical significance. We introduce the new accompaniment interface Song Prompter, which uses the automatically aligned lyrics to guide musicians through a song. It demonstrates that the automatic alignment is accurate enough to be used in a musical performance.
Matthias Mauch, Hiromasa Fujihara, Masataka Goto
IEEE Trans. Speech Audio Process.3
2012 A Nonparametric Bayesian Multipitch Analyzer Based on Infinite Latent Harmonic Allocation
abstract
The statistical multipitch analyzer described in this paper estimates multiple fundamental frequencies (F0s) in polyphonic music audio signals produced by pitched instruments. It is based on hierarchic4al nonparametric Bayesian models that can deal with uncertainty of unknown random variables such as model complexities (e.g., the number of F0s and the number of harmonic partials), model parameters (e.g., the values of F0s and the relative weights of harmonic partials), and hyperparameters (i.e., prior knowledge on complexities and parameters). Using these models, we propose a statistical method called infinite latent harmonic allocation (iLHA). To avoid model-complexity control, we allow the observed spectra to contain an unbounded number of sound sources (F0s), each of which is allowed to contain an unbounded number of harmonic partials. More specifically, to model a set of time-sliced spectra, we formulated nested infinite Gaussian mixture models based on hierarchical and generalized Dirichlet processes. To avoid manual tuning of influential hyperparameters, we put noninformative hyperprior distributions on them in a hierarchical manner. For efficient Bayesian inference, we used a modern technique called collapsed variational Bayes. In comparative experiments using audio recordings of piano and guitar solo performances, iLHA yielded promising results and we found that there would be room for improvement based on modeling of temporal continuity and spectral smoothness.
Kazuyoshi Yoshii, Masataka Goto
IEEE Trans. Speech Audio Process.2
2011 Social Infobox: collaborative knowledge construction by social property tagging
abstract
We propose a novel style of social tagging to construct knowledge collaboratively called Social Property Tagging and introduce the prototype system Social Infobox. Structured data is useful for computer system, however defining structure of knowledge for representing data semantics is usually a costly and time consuming task. In general, data structures are constructed by experts of knowledge engineering. Our method aims to construct not only structured data but also structure of data collaboratively by simple user input.
Masahiro Hamasaki, Masataka Goto, Hideaki Takeda 0001
CSCW2
2011 Concurrent estimation of singing voice F0 and phonemes by using spectral envelopes estimated from polyphonic music
abstract
The scarcity of available multi-track recordings constitutes a severe constraint on the training of probabilistic models for voice extraction from polyphonic music. We propose a novel training method to estimate a spectral envelope of a singing voice that makes it possible to train the models from a polyphonic music without segregating a singing voice. We implement this method as an extension to the existing W-PST method, which concurrently estimates singing voice fundamental frequency (F0) and phoneme from polyphonic music. The novel training method is based on random sampling from probabilistic distributions. We conducted experiments on concurrent F0 and phoneme estimation and confirm the effectiveness of our method.
Hiromasa Fujihara, Masataka Goto
ICASSP2
2011 Simultaneous processing of sound source separation and musical instrument identification using Bayesian spectral modeling
abstract
This paper presents a method of both separating audio mixtures into sound sources and identifying the musical instruments of the sources. A statistical tone model of the power spectrogram, called an integrated model, is defined and source separation and instrument identification are carried out on the basis of Bayesian inference. Since, the parameter distributions of the integrated model depend on each instrument, the instrument name is identified by selecting the one that has the maximum relative instrument weight. Experimental results showed correct instrument identification enables precise source separation even when many overtones overlap.
Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP2
2011 Polyphonic audio-to-score alignment based on Bayesian Latent Harmonic Allocation Hidden Markov Model
abstract
This paper presents a Bayesian method for temporally aligning a music score and an audio rendition. A critical problem in audio-to-score alignment is in dealing with the wide variety of timbre and volume of the audio rendition. In contrast with existing works that achieve this through ad-hoc feature design or careful training of tone models, we propose a Bayesian audio-to-score alignment method by modeling music performance as a Bayesian Hidden Markov Model, each state of which emits a Bayesian signal model based on Latent Harmonic Allocation. After attenuating reverberation, variational Bayes method is used to iteratively adapt the alignment, instrument tone model and the volume balance at each position of the score. The method is evaluated using sixty works of classical music of a variety of instrumentation ranging from solo piano to full orchestra. We verify that our method improves the alignment accuracy compared to dynamic time warping based on chroma vector for orchestral music, or our method employed in a maximum likelihood setting.
Akira Maezawa, Hiroshi G. Okuno, Tetsuya Ogata, Masataka Goto
ICASSP4
2011 Vocalistener2: A singing synthesis system able to mimic a user's singing in terms of voice timbre changes as well as pitch and dynamics
abstract
This paper presents a singing synthesis system, VocaListener2, that can automatically synthesize a singing voice by mimicking the timbre changes of a user's singing voice. The system is an extension of our previous VocaListener system which deals with only pitch and dynamics. Most previous techniques for manipulating voice timbre have focused on voice conversion and voice morphing, and they cannot deal with the timbre changes during singing. To develop VocaListener2, we constructed a voice timbre space on the basis of various singing voices that are synchronized under pitch, dynamics, and phoneme by using VocaListener. In this space, the timbre changes can be reflected in the synthesized singing voice. The system was evaluated by the Euclidean distance in the space between an estimated result and a ground-truth under closed/open conditions.
Tomoyasu Nakano, Masataka Goto
ICASSP2
2011 PodCastle: Recent Advances of a Spoken Document Retrieval Service Improved by Anonymous User Contributions
abstract
In this paper, we introduce recent advances of a speech retrieval web service, PodCastle, that collects and amplifies voluntary contributions by anonymous users. Our goal is to provide users with a public web service based on speech recognition and crowdsourcing so that they can experience state-of-the-art speech recognition performance through a useful service. Pod-Castle enables users to find speech data (such as podcasts and YouTube video clips) that include a search term, read full texts of their recognition results, and easily correct recognition errors by simply selecting from a list of candidates. The resulting corrections were used to improve both the speech retrieval and recognition performances. In our experiences from its practical use over the past four years (since December, 2006), over half a million recognition errors in about one hundred thousand speech data were corrected by anonymous users and we confirmed that the speech recognition performance of PodCastle was actually improved by those corrections. Index Terms: speech retrieval, spoken document retrieval, wisdom of crowds, crowdsourcing
Masataka Goto, Jun Ogata
INTERSPEECH1
2011 VocaWatcher: Natural singing motion generator for a humanoid robot
abstract
In this paper, we describe VocaWatcher, a novel robot motion generator that enables a humanoid robot to sing with realistic facial expressions and naturally synthesized singing voices. This robot singer is an important and attractive humanoid robot application for the entertainment scene; moreover, it promotes state-of-the-art integration of robot engineering, music processing, and image processing. To overcome the difficulties of generating natural facial expressions that are precisely synchronized with singing voices, VocaWatcher imitates a human singer by analyzing a video clip of a human singing, recorded by a single video camera. VocaWatcher can control mouth, eye, and neck motions by imitating the corresponding human movements, which are estimated without using any markers in the video. It can also synthesize singing voices by imitating the pitch and dynamics of the human singing in the same video.
Shuuji Kajita, Tomoyasu Nakano, Masataka Goto, Yosuke Matsusaka, Shinichiro Nakaoka, Kazuhito Yokoi
IROS3
2010 Singing information processing based on singing voice modeling
abstract
In this paper, we propose a novel area of research referred to as singing information processing. To shape the concept of this area, we first introduce singing understanding systems for synchronizing between vocal melody and corresponding lyrics, identifying the singer name, evaluating singing skills, creating hyperlinks between phrases in the lyrics of songs, and detecting breath sounds. We then introduce music information retrieval systems based on similarity of vocal melody timbre and vocal percussion, and singing synthesis systems. Common signal processing techniques for modeling singing voices that are used in these systems, such as techniques for extracting the vocal melody from polyphonic music recordings and modeling the lyrics by using phoneme HMMs for singing voices, are discussed.
Masataka Goto, Takeshi Saitou, Tomoyasu Nakano, Hiromasa Fujihara
ICASSP1
2010 PodCastle: A Spoken Document Retrieval Service Improved by Anonymous User Contributions
Masataka Goto, Jun Ogata
PACLIC1
2010 Editorial for the Special Issue on Signal Models and Representations of Musical and Environmental Sounds
abstract
The 23 papers in this special issue focus on signal models and representations of musical and environmental sounds.
Bertrand David 0001, Masataka Goto, Laurent Daudet, Paris Smaragdis
IEEE Trans. Speech Audio Process.2
2010 A Modeling of Singing Voice Robust to Accompaniment Sounds and Its Application to Singer Identification and Vocal-Timbre-Similarity-Based Music Information Retrieval
abstract
This paper describes a method of modeling the characteristics of a singing voice from polyphonic musical audio signals including sounds of various musical instruments. Because singing voices play an important role in musical pieces with vocals, such representation is useful for music information retrieval systems. The main problem in modeling the characteristics of a singing voice is the negative influences caused by accompaniment sounds. To solve this problem, we developed two methods,accompaniment sound reduction and reliable frame selection.The former makes it possible to calculate feature vectors that represent a spectral envelope of a singing voice after reducing accompaniment sounds. It first extracts the harmonic components of the predominant melody from sound mixtures and then resynthesizes the melody by using a sinusoidal model driven by these components. The latter method then estimates the reliability of frame of the obtained melody (i.e., the influence of accompaniment sound) by using two Gaussian mixture models (GMMs) for vocal and nonvocal frames to select the reliable vocal portions of musical pieces. Finally, each song is represented by its GMM consisting of the reliable frames. This new representation of the singing voice is demonstrated to improve the performance of an automatic singer identification system and to achieve an MIR system based on vocal timbre similarity.
Hiromasa Fujihara, Masataka Goto, Tetsuro Kitahara, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.2
2009 The use of acoustically detected filled and silent pauses in spontaneous speech recognition
abstract
In recognizing spontaneous speech, the performance of typical speech recognizers tends to be degraded by filled and silent pauses, which are hesitation phenomena frequently occurred in such speech. In this paper, we present a method for improving the performance of a speech recognizer by detecting and handling both filled pauses (lengthened vowels) and silent (unfilled) pauses. Our method automatically detects these pauses by using a bottom-up acoustical analysis in parallel with a typical speech decoding process, and then incorporates the detected results into the decoding process. From the results of experiments conducted using the CIAIR spontaneous speech corpus, the effectiveness of the proposed method was confirmed.
Jun Ogata, Masataka Goto, Katunobu Itou
ICASSP2
2009 Podcastle: collaborative training of acoustic models on the basis of wisdom of crowds for podcast transcription
abstract
This paper presents acoustic-model-training techniques for improving automatic transcription of podcasts. A typical approach for acoustic modeling is to create a task-specific corpus including hundreds (or even thousands) of hours of speech data and their accurate transcriptions. This approach, however, is impractical in podcast-transcription task because manual generation of the transcriptions of the large amounts of speech covering all the various types of podcast contents will be too costly and time consuming. To solve this problem, we introduce collaborative training of acoustic models on the basis of wisdom of crowds, i.e., the transcriptions of podcast-speech data are generated by anonymous users on our web service PodCastle. We then describe a podcast-dependent acoustic modeling system by using RSS metadata to deal with the differences of acoustic conditions in podcast speech data. From our experimental results on actual podcast speech data, the effectiveness of the proposed acoustic model training was confirmed. Index Terms: podcast, LVCSR, acoustic model training, wisdom of crowds, error correction
Jun Ogata, Masataka Goto
INTERSPEECH2
2009 Acoustic and perceptual effects of vocal training in amateur male singing
abstract
This paper reports our investigation of the acoustic effects of vocal training for amateur singers and of the contribution of those effects to perceived vocal quality. Recording singing voices before and after vocal training and then analyzing changes in acoustic parameters with a focus on features unique to singing voices, we found that two different F0 fluctuations (vibrato and overshoot) and singing formant were improved by the training. The results of psychoacoustic experiments showed that perceived voice quality was influenced more by the changes of F0 characteristics than by the changes of spectral characteristics and that acoustic features unique to singing voices contribute to perceived voice quality in the following order: vibrato, singing formant, overshoot, and preparation. Index Terms: singing voice, vocal training, psychoacoustic experiment 1.
Takeshi Saitou, Masataka Goto
INTERSPEECH2
2009 Acoustic event detection for spotting "hot spots" in podcasts
abstract
This paper presents a method to detect acoustic events that can be used to find “hot spots ” in podcast programs. We focus on meaningful non-verbal audible reactions which suggest hot spots such as laughter and reactive tokens. In order to detect this kind of short events and segment the counterpart utterances, we need accurate audio segmentation and classification, dealing with various recording environments and background music. Thus, we propose a method for automatically estimating and switching penalty weights for the BIC-based segmentation depending on background environments. Experimental results show significant improvement in detection accuracy by proposed method compared to when using a constant penalty weight. Index Terms: acoustic event detection, laughter detection, podcast, Bayesian Information Criterion
Kouhei Sumi, Tatsuya Kawahara, Jun Ogata, Masataka Goto
INTERSPEECH4
2009 MusicCommentator: Generating Comments Synchronized with Musical Audio Signals by a Joint Probabilistic Model of Acoustic and Textual Features
Kazuyoshi Yoshii, Masataka Goto
ICEC2
2008 Three techniques for improving automatic synchronization between music and lyrics: Fricative detection, filler model, and novel feature vectors for vocal activity detection
abstract
Three techniques are described that improve a previously developed system for automatically synchronizing lyrics with musical audio signals. Although this system achieves state-of-the-art accuracy by extracting vocal vowels from polyphonic sound mixtures and using forced alignment between those vowels and a phoneme network of the lyrics, there was still room for improvement. The first technique detects nonexistence regions in which fricative consonant sounds do not exist, which were not utilized in the previous system, and prohibits the alignment of the fricative phonemes to those regions. The second technique inserts a filler model between phrases of the phoneme network. This model improves the accuracy of the forced alignment by ignoring inter-phrase vowel utterances not included in the lyrics. The third technique introduces novel feature vectors for vocal activity detection that enable a distance calculation between two sets of the harmonic structure without estimating their spectral envelopes. Experimental results showed that all three techniques contribute to improved synchronization.
Hiromasa Fujihara, Masataka Goto
ICASSP2
2008 A similar content retrieval method for podcast episodes
abstract
Given podcasts (audio blogs) that are sets of speech files called episodes, this paper describes a method for retrieving episodes that have similar content. Although most previous retrieval methods were based on bibliographic information, tags, or users' playback behaviors without considering spoken content, our method can compute content-based similarity based on speech recognition results of podcast episodes even if the recognition results include some errors. To overcome those errors, it converts intermediate speech-recognition results to a confusion network containing competitive candidates, and then computes the similarity by using keywords extracted from the network. Experimental results with episodes that have different word accuracy and content showed that keywords obtained from competitive candidates were useful in retrieving similar episodes. To show relevant episodes, our method will be incorporated into PodCastle, a public web service that provides full-text searching of podcasts on the basis of speech recognition.
Junta Mizuno, Jun Ogata, Masataka Goto
SLT3
2008 Content-Based Music Information Retrieval: Current Directions and Future Challenges
abstract
The steep rise in music downloading over CD sales has created a major shift in the music industry away from physical media formats and towards online products and services. Music is one of the most popular types of online information and there are now hundreds of music streaming and download services operating on the World-Wide Web. Some of the music collections available are approaching the scale of ten million tracks and this has posed a major challenge for searching, retrieving, and organizing music content. Research efforts in music information retrieval have involved experts from music perception, cognition, musicology, engineering, and computer science engaged in truly interdisciplinary activity that has resulted in many proposed algorithmic and methodological solutions to music search using content-based methods. This paper outlines the problems of content-based music information retrieval and explores the state-of-the-art methods using audio cues (e.g., query by humming, audio fingerprinting, content-based music retrieval) and other cues (e.g., music notation and symbolic representation), and identifies some of the major challenges for the coming years.
Michael A. Casey, Remco C. Veltkamp, Masataka Goto, Marc Leman, Christophe Rhodes, Malcolm Slaney
Proc. IEEE3
2008 Computational Models of Similarity for Drum Samples
abstract
In this paper, we optimize and evaluate computational models of similarity for sounds from the same instrument class. We investigate four instrument classes: bass drums, snare drums, high-pitched toms, and low-pitched toms. We evaluate two similarity models: one is defined in the ISO/IEC MPEG-7 standard, and the other is based on auditory images. For the second model, we study the impact of various parameters. We use data from listening tests, and instrument class labels to evaluate the models. Our results show that the model based on auditory images yields a very high average correlation with human similarity ratings and clearly outperforms the MPEG-7 recommendation. The average correlations range from 0.89-0.96 depending on the instrument class. Furthermore, our results indicate that instrument class data can be used as alternative to data from listening tests to evaluate sound similarity models.
Elias Pampalk, Perfecto Herrera, Masataka Goto
IEEE Trans. Speech Audio Process.3
2008 An Efficient Hybrid Music Recommender System Using an Incrementally Trainable Probabilistic Generative Model
abstract
This paper presents a hybrid music recommender system that ranks musical pieces while efficiently maintaining collaborative and content-based data, i.e., rating scores given by users and acoustic features of audio signals. This hybrid approach overcomes the conventional tradeoff between recommendation accuracy and variety of recommended artists. Collaborative filtering, which is used on e-commerce sites, cannot recommend nonbrated pieces and provides a narrow variety of artists. Content-based filtering does not have satisfactory accuracy because it is based on the heuristics that the user's favorite pieces will have similar musical content despite there being exceptions. To attain a higher recommendation accuracy along with a wider variety of artists, we use a probabilistic generative model that unifies the collaborative and content-based data in a principled way. This model can explain the generative mechanism of the observed data in the probability theory. The probability distribution over users, pieces, and features is decomposed into three conditionally independent ones by introducing latent variables. This decomposition enables us to efficiently and incrementally adapt the model for increasing numbers of users and rating scores. We evaluated our system by using audio signals of commercial CDs and their corresponding rating scores obtained from an e-commerce site. The results revealed that our system accurately recommended pieces including nonrated ones from a wide variety of artists and maintained a high degree of accuracy even when new users and rating scores were added.
Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.2
2007 Active Music Listening Interfaces Based on Signal Processing
abstract
This paper introduces our research aimed at building "active music listening interfaces". This research approach is intended to enrich end-users' music listening experiences by applying music-understanding technologies based on signal processing. Active music listening is a way of listening to music through active interactions. We have developed seven interfaces for active music listening, such as interfaces for skipping sections of no interest within a musical piece while viewing a graphical overview of the entire song structure, for displaying virtual dancers or song lyrics synchronized with the music, for changing the timbre of instrument sounds in compact-disc recordings, and for browsing a large music collection to encounter interesting musical pieces or artists. These interfaces demonstrate the importance of music-understanding technologies and the benefit they offer to end users. Our hope is that this work will help change music listening into a more active, immersive experience.
Masataka Goto
ICASSP (4)1
2007 Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
abstract
This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-disc recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-disc recordings and achieved the instrument equalizer.
Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (1)2
2007 Presentation sensei: a presentation training system using speech and image processing
abstract
In this paper we present a presentation training system that observes a presentation rehearsal and provides the speaker with recommendations for improving the delivery of the presentation, such as to speak more slowly and to look at the audience. Our system "Presentation Sensei" is equipped with a microphone and camera to analyze a presentation by combining speech and image processing techniques. Based on the results of the analysis, the system gives the speaker instant feedback with respect to the speaking rate, eye contact with the audience, and timing. It also alerts the speaker when some of these indices exceed predefined warning thresholds. After the presentation, the system generates visual summaries of the analysis results for the speaker's self-examinations. Our goal is not to improve the content on a semantic level, but to improve the delivery of it by reducing inappropriate basic behavior patterns. We asked a few test users to try the system and they found it very useful for improving their presentations. We also compared the system's output with the observations of a human evaluator. The result shows that the system successfully detected some inappropriate behavior. The contribution of this work is to introduce a practical recognition-based human training system and to show its feasibility despite the limitations of state-of-the-art speech and video recognition technologies.
Kazutaka Kurihara, Masataka Goto, Jun Ogata, Yosuke Matsusaka, Takeo Igarashi
ICMI2
2007 Podcastle: a web 2.0 approach to speech recognition research
abstract
In this paper, we describe a public web service, “PodCastle”, that provides full-text searching of Japanese podcasts on the basis of automatic speech recognition. This is an instance of our research approach, “Speech Recognition Research 2.0”, which is aimed at providing users with a web service based on Web 2.0 so that they can experience state-of-the-art speech recognition performance, and at promoting speech recognition technologies in cooperation with anonymous users. PodCastle enables users to find podcasts that include a search term, read full texts of their recognition results, and easily correct recognition errors. The results of the error correction can then be used to improve the performance of both full-text search and speech recognition. Although we know of no state-of-the-art speech recognizer that can successfully transcribe all of the various kinds of podcasts, the mechanism we propose will gradually increase the usefulness and applicability of PodCastle. Index Terms: information retrieval, speech recognition, error correction, wisdom of crowds, Web 2.0
Masataka Goto, Jun Ogata, Kouichirou Eto
INTERSPEECH1
2007 Automatic transcription for a web 2.0 service to search podcasts
abstract
This paper describes speech recognition techniques that enable a Web 2.0 service “PodCastle ” where users can search and read transcribed texts of podcasts, and correct recognition errors in those texts. Most previous speech recognizers had difficulties transcribing podcasts because podcasts include various kinds of contents recorded in different conditions and cover recent topics that tend to have many out-of-vocabulary words. To overcome such difficulties, we continuously improve speech recognizers by using information aggregated on the basis of Web 2.0. For example, a language model is adapted to a topic of the target podcast on the fly, the pronunciations of out-of-vocabulary words are obtained from a Web 2.0 service, and an acoustic model is trained by using the results of the error correction by anonymous users. The experiments we report in this paper show that our techniques produce promising results for podcasts.
Jun Ogata, Masataka Goto, Kouichirou Eto
INTERSPEECH2
2007 Vocal conversion from speaking voice to singing voice using STRAIGHT
Takeshi Saitou, Masataka Goto, Masashi Unoki, Masato Akagi
INTERSPEECH2
2007 Drum Sound Recognition for Polyphonic Audio Signals by Adaptation and Matching of Spectrogram Templates With Harmonic Structure Suppression
abstract
This paper describes a system that detects onsets of the bass drum, snare drum, and hi-hat cymbals in polyphonic audio signals of popular songs. Our system is based on a template-matching method that uses power spectrograms of drum sounds as templates. This method calculates the distance between a template and each spectrogram segment extracted from a song spectrogram, using Goto's distance measure originally designed to detect the onsets in drums-only signals. However, there are two main problems. The first problem is that appropriate templates are unknown for each song. The second problem is that it is more difficult to detect drum-sound onsets in sound mixtures including various sounds other than drum sounds. To solve these problems, we propose template-adaptation and harmonic-structure-suppression methods. First of all, an initial template of each drum sound, called a seed template, is prepared. The former method adapts it to actual drum-sound spectrograms appearing in the song spectrogram. To make our system robust to the overlapping of harmonic sounds with drum sounds, the latter method suppresses harmonic components in the song spectrogram before the adaptation and matching. Experimental results with 70 popular songs showed that our template-adaptation and harmonic-structure-suppression methods improved the recognition accuracy and achieved 83%, 58%, and 46% in detecting onsets of the bass drum, snare drum, and hi-hat cymbals, respectively.
Kazuyoshi Yoshii, Masataka Goto, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.2
2006 Speech pen: predictive handwriting based on ambient multimodal recognition
abstract
It is tedious to handwrite long passages of text by hand. To make this process more efficient, we propose predictive handwriting that provides input predictions when the user writes by hand. A predictive handwriting system presents possible next words as a list and allows the user to select one to skip manual writing. Since it is not clear if people are willing to use prediction, we first run a user study to compare handwriting and selecting from the list. The result shows that, in Japanese, people prefer to select, especially when the expected performance gain from using selection is large. Based on these observations, we designed a multimodal input system, called speech-pen, that assists digital writing during lectures or presentations with background speech and handwriting recognition. The system recognizes speech and handwriting in the background and provides the instructor with predictions for further writing. The speech-pen system also allows the sharing of context information for predictions among the instructor and the audience; the result of the instructor's speech recognition is sent to the audience to support their own note-taking. Our preliminary study shows the effectiveness of this system and the implications for further improvements.
Kazutaka Kurihara, Masataka Goto, Jun Ogata, Takeo Igarashi
CHI2
2006 F0 Estimation Method for Singing Voice in Polyphonic Audio Signal Based on Statistical Vocal Model and Viterbi Search
abstract
This paper describes a method for estimating F0s of vocal from polyphonic audio signals. Because melody is sung by a singer in many musical pieces, the estimation of F0s of the vocal part is useful for many applications. Based on existing multiple-F0 estimation method, we evaluate the vocal probabilities of the harmonic structure of each F0 candidate. In order to calculate the vocal probabilities of the harmonic structure, we extract and resynthesize the harmonic structure by using a sinusoidal model and extract feature vectors. Then, we evaluate the vocal probability by using vocal and non-vocal Gaussian mixture models (GMMs). Finally, we track F0 trajectories using these probabilities based on Viterbi search. Experimental results show that our method improves estimation accuracy from 78.1% to 84.3%, which is 28.3% reduction of misestimation
Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (5)3
2006 Instrogram: A New Musical Instrument Recognition Technique Without Using Onset Detection NOR F0 Estimation
abstract
This paper describes a new technique for recognizing musical instruments in polyphonic music. Because the conventional framework for musical instrument recognition in polyphonic music had to estimate the onset time and fundamental frequency (F0) of each note, instrument recognition strictly suffered from errors of onset detection and F0 estimation. Unlike such a note-based processing framework, our technique calculates the temporal trajectory of instrument existence probabilities for every possible F0, and the results are visualized with a spectrogram-like graphical representation called instrogram. The instrument existence probability is defined as the product of a nonspecific instrument existence probability calculated using PreFEst and a conditional instrument existence probability calculated using the hidden Markov model. Experimental results show that the obtained instrograms reflect the actual instrumentations and facilitate instrument recognition
Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (5)2
2006 An Error Correction Framework Based on Drum Pattern Periodicity for Improving Drum Sound Detection
abstract
This paper presents a framework for correcting errors of automatic drum sound detection focusing on the periodicity of drum patterns. We define drum patterns as periodic structures found in onset sequences of bass and snare drum sounds. Our framework extracts periodic drum patterns from imperfect onset sequences of detected drum sounds (bottom-up processing) and corrects errors using the periodicity of the drum patterns (top-down processing). We implemented this framework on our drum-sound detection system. We first obtained onset sequences of the drum sounds with our system and extracted drum patterns. On the basis of our observation that the same drum patterns tend to be repeated, we detected time points which deviate from the periodicity as error candidates. Finally, we verified each error candidate to judge whether it is an actual onset or not. Experiments of drum sound detection for polyphonic audio signals of popular CD recordings showed that our correction framework improved the average detection accuracy from 77.4% to 80.7%
Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (5)2
2006 Speaker identification under noisy environments by using harmonic structure extraction and reliable frame weighting
abstract
We present methods for automatic speaker identification in noisy environments. To improve noise robustness of speaker identification, we developed two methods, theharmonic structure extraction method and the reliable frame weighting method. The harmonic structure extraction method enables the speaker of input speech signals to be identified after environmental noise has been reduced. This method first extracts harmonic components of the speech from the sound mixtures and then resynthesizes a clean speech signal by using a sinusoidal model driven by harmonic components. The reliable frame weighting method then determines how each frame of the resynthesized speech is reliable (i.e. little influenced by environmental noises) by using two Gaussian mixture models for the speech and noise. The speaker can be robustly identified by attaching importance to reliable frames. Experimental results with thirty speakers showed that our method was able to reduce the influences of environmental noise and achieved an error rate of 10.7%, while the error rate for a conventional method was 18.9%. Index Terms: speaker identification, noise robustness, voice extraction, voice reliability, Gaussian mixture model.
Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH3
2006 An automatic singing skill evaluation method for unknown melodies using pitch interval accuracy and vibrato features
abstract
This paper presents a method of evaluating singing skills that does not require score information of the sung melody. This requires an approach that is different from existing systems, such as those currently used for Karaoke systems. Previous research on singing evaluation has focused on analyzing the characteristics of singing voice, but were not aimed at developing an automatic evaluation method. The approach presented in this study uses pitch interval accuracy and vibrato as acoustic features which are independent from specific characteristics of the singer or melody. The approach was tested by a 2-class (good/poor) classification test with 600 song sequences, and achieved an average classification rate of
Tomoyasu Nakano, Masataka Goto, Yuzuru Hiraga
INTERSPEECH2
2006 Automatic Synchronization between Lyrics and Music CD Recordings Based on Viterbi Alignment of Segregated Vocal Signals
abstract
This paper describes a system that can automatically synchronize between polyphonic musical audio signals and corresponding lyrics. Although there were methods that can synchronize between monophonic speech signals and corresponding text transcriptions by using Viterbi alignment techniques, they cannot be applied to vocals in CD recordings because accompaniment sounds often overlap with vocals. To align lyrics with such vocals, we therefore developed three methods: a method for segregating vocals from polyphonic sound mixtures, a method for detecting vocal sections, and a method for adapting a speech-recognizer phone model to segregated vocal signals. Experimental results for 10 Japanese popular-music songs showed that our system can synchronize between music and lyrics with satisfactory accuracy for 8 songs
Hiromasa Fujihara, Masataka Goto, Jun Ogata, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ISM2
2006 Musical Instrument Recognizer "Instrogram" and Its Application to Music Retrieval Based on Instrumentation Similarity
abstract
Instrumentation is an important cue in retrieving musical content. Conventional methods for instrument recognition performing notewise require accurate estimation of the onset time and fundamental frequency (FO) for each note, which is not easy in polyphonic music. This paper presents a non-notewise method for instrument recognition in polyphonic musical audio signals. Instead of such note-wise estimation, our method calculates the temporal trajectory of instrument existence probabilities for every FO and visualizes it as a spectrogram-like graphical representation, called an instrogram. This method can avoid the influence by errors of onset detection and FO estimation because it does not use them. We also present methods for MPEG-7-based instrument annotation and music information retrieval based on the similarity between instrograms. Experimental results with realistic music show the average accuracy of 76.2% for the instrument annotation and that the instrogram-based similarity measure represents the actual instrumentation similarity better than an MFCC-based one
Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ISM2
2006 A chorus section detection method for musical audio signals and its application to a music listening station
abstract
This paper describes a method for obtaining a list of repeated chorus ("hook") sections in compact-disc recordings of popular music. The detection of chorus sections is essential for the computational modeling of music understanding and is useful in various applications, such as automatic chorus-preview/search functions in music listening stations, music browsers, or music retrieval systems. Most previous methods detected as a chorus a repeated section of a given length and had difficulty identifying both ends of a chorus section and dealing with modulations (key changes). By analyzing relationships between various repeated sections, our method, called RefraiD, can detect all the chorus sections in a song and estimate both ends of each section. It can also detect modulated chorus sections by introducing a perceptually motivated acoustic feature and a similarity that enable detection of a repeated chorus section even after modulation. Experimental results with a popular music database showed that this method correctly detected the chorus sections in 80 of 100 songs. This paper also describes an application of our method, a new music-playback interface for trial listening called SmartMusicKIOSK , which enables a listener to directly jump to and listen to the chorus section while viewing a graphical overview of the entire song structure. The results of implementing this application have demonstrated its usefulness
Masataka Goto
IEEE Trans. Speech Audio Process.1
2005 An Auto-Regressive, Non-Stationary Excited Signal Parameter Estimation Method and an Evaluation of a Singing-Voice Recognition
abstract
We have previously described an auto-regressive hidden Markov model (AR-HMM) and an accompanying parameter estimation method. The AR-HMM was obtained by combining an AR process with an HMM introduced as a non-stationary excitation model. We demonstrated that the AR-HMM can accurately estimate the characteristics of both articulatory systems and excitation signals from high-pitched speech. In this paper, we apply the AR-HMM to feature extraction from singing voices and evaluate the recognition accuracy of the AR-HMM-based approach.
Akira Sasou, Masataka Goto, Satoru Hayamizu, Kazuyo Tanaka
ICASSP (1)2
2005 Speech repair: quick error correction just by using selection operation for speech input interfaces
abstract
In this paper, we propose a novel speech input interface function, called “Speech Repair ” in which recognition errors can be easily corrected by selecting candidates. During the speech input, this function displays not only the typical speechrecognition result but also other competitive candidates. Each word in the result is separated by line segments and accompanied by other word candidates. A user who finds a recognition error can simply select the correct word from the candidates for that temporal region. In order to overcome the difficulty of generating appropriate candidates, we adopted a confusion network that can condense a huge internal word graph of a large vocabulary continuous speech recognition (LVCSR) system. In our experiments, almost all recognition errors were corrected and the effectiveness of speech repair was confirmed. 1.
Jun Ogata, Masataka Goto
INTERSPEECH2
2005 Discrimination between singing and speaking voices
abstract
Discriminating between singing and speaking voices by using the local and global characteristics of voice signals is discussed. From the results of subjective experiments, we show that human beings can discriminate singing and speaking voices with more than 70 % and 95 % accuracy from 300 ms and one second long signals, respectively. From the subjective experiment results, assuming that different features are effective for shortterm and long-term signals, we designed two measures using a spectral envelope (MFCC) and the fundamental frequency (F0, perceived as pitch) contour. Experimental results show that the F0 measure performs better than the spectral envelope measure when the input voice signals are longer than one second. Particularly, it can discriminate singing and speaking voices with more than 80 % accuracy with two-second signals. On the other hand, when the input signals are shorter than one second, the spectral envelope measure performs better than the F0 measure. Finally, by simply combining the two measures, more than 90% accuracy is obtained for two-second signals. 1.
Yasunori Ohishi, Masataka Goto, Katunobu Itou, Kazuya Takeda
INTERSPEECH2
2005 A wireless LAN architecture using PANA for secure network selection
abstract
In this paper, we propose a new network architecture that supports IEEE 802.11i and WPA and provides the network selection functionality and the functionality to physically separate AAA clients from access points by using PANA (protocol for carrying authentication for network access) as the network access authentication protocol; both functionalities have been difficult to support solely with IEEE 802.11i. We also show our prototype system in which the proposed architecture is realized using a newly developed WPA-capable access point which provides connectivity to multiple networks, namely, a guest network, a management network, and multiple user networks. We also show performance results of the prototype system and an example of usage in an enterprise environment.
Yoshimichi Tanizawa, Masataka Goto, Victor Fajardo, Yoshihiro Ohba
WiMob (2)2
2005 Pitch-Dependent Identification of Musical Instrument Sounds
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
Appl. Intell.2
2004 Category-level identification of non-registered musical instrument sounds
abstract
This paper describes a method that identifies sounds of non-registered musical instruments (i.e., musical instruments that are not contained in the training data) at a category level. Although the problem of how to deal with non-registered musical instruments is essential in musical instrument identification, it has not been dealt with in previous studies. Our method solves this problem by distinguishing between registered and non-registered instruments and identifying the category name of the non-registered instruments. When a given sound is registered, its instrument name, e.g. violin, is identified. Even if it is not registered, its category name, e.g. strings, can be identified. The important issue in achieving such identification is to adopt a musical instrument hierarchy reflecting the acoustical similarity. We present a method for acquiring such a hierarchy from a musical instrument sound database. Experimental results show that around 77% of non-registered instrument sounds, on average, were correctly identified at the category level.
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
ICASSP (4)2
2004 Speech spotter: on-demand speech recognition in human-human conversation on the telephone or in face-to-face situations
abstract
This paper describes a novel speech-interface function, called “speech spotter”,whichenablesausertoentervoicecommands into a speech recognizer in the midst of natural human-human conversation. In the past, it has been difficult to use automatic speech recognition in human-human conversation since it was not easy to judge, from only microphone input, whether a user was speaking to another person or a speech recognizer. We solve this problem by using two kinds of nonverbal speech information: a filled pause (a vowel-lengthening hesitation like “er...”) and voice pitch. Only when a user utters a voice command with a high pitch just after a filled pause is the voice command accepted by the speech recognizer. By using this speechspotter function, we have built two application systems: an ondemand information system for assisting human-human conversation and a music-playback system for enriching telephone conversation. The results from using these systems have shown thatthespeech-spotter functionisrobustandconvenientenough to be used in face-to-face or cellular-phone conversations.
Masataka Goto, Koji Kitayama, Katunobu Itou, Tetsunori Kobayashi
INTERSPEECH1
2004 A real-time music-scene-description system: predominant-F0 estimation for detecting melody and bass lines in real-world audio signals
Masataka Goto
Speech Commun.1
2003 A chorus-section detecting method for musical audio signals
abstract
This paper describes a method for obtaining a list of chorus (refrain) sections in compact-disc recordings of popular music. The detection of chorus sections is essential for the computational modeling of music understanding and is useful in various applications, such as automatic chorus-preview functions in music browsers or retrieval systems. Most previous methods detected as a chorus a repeated section of a given length and had difficulty in identifying both ends of a chorus section and in dealing with modulations (key changes). By analyzing relationships between various repeated sections, our method called RefraiD can detect all the chorus sections in a song and estimate both ends of each section. It can also detect modulated chorus sections by introducing a similarity that enables modulated repetition to be judged correctly. Experimental results with a popular-music database show that this method detects the correct chorus sections in 80 of 100 songs.
Masataka Goto
ICASSP (5)1
2003 Musical instrument identification based on F0-dependent multivariate normal distribution
abstract
The pitch dependency of timbres has not been fully exploited in musical instrument identification. In this paper, we present a method using an F0-dependent multivariate normal distribution of which mean is represented by a function of fundamental frequency (F0). This F0-dependent mean function represents the pitch dependency of each feature, while the F0-normalized covariance represents the nonpitch dependency. Musical instrument sounds are first analyzed by the F0-dependent multivariate normal distribution, and then identified by using the discriminant function based on the Bayes decision rule. Experimental results of identifying 6247 solo tones of 19 musical instruments by 10-fold cross validation showed that the proposed method improved the recognition rate at individual-instrument level from 75.73% to 79.73%, and the recognition rate at category level from 88.20% to 90.65%.
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
ICASSP (5)2
2003 Musical instrument identification based on F0-dependent multivariate normal distribution
abstract
The pitch dependency of timbres has not been fully exploited in musical instrument identification. In this paper, we present a method using an F0-dependent multivariate normal distribution of which mean is represented by a function of fundamental frequency (FO). This F0-dependent mean function represents the pitch dependency of each feature, while the F0-normalized covariance represents the non-pitch dependency. Musical instrument sounds are first analyzed by the F0-dependent multivariate normal distribution, and then identified by using the discriminant function based on the Bayes decision rule. Experimental results of identifying 6,247 solo tones of 19 musical instruments by 10-fold cross validation showed that the proposed method improved the recognition rate at individual-instrument level from 75.73% to 79.73%, and the recognition rate at category level from 88.20% to 90.65%.
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
ICME2
2003 Pitch-Dependent Musical Instrument Identification and Its Application to Musical Sound Ontology
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
IEA/AIE2
2003 A Learning-Based Jam Session System that Imitates a Player's Personality Model
Masatoshi Hamanaka, Masataka Goto, Hideki Asoh, Nobuyuki Otsu
IJCAI2
2003 Speech shift: direct speech-input-mode switching through intentional control of voice pitch
abstract
This paper describes a speech-input interface function, called speech shift, that enables a user to specify a speech-input mode by simply changing (shifting) voice pitch. While current speech-input interfaces have used only verbal information, we aimed at building a more user-friendly speech interface by making use of nonverbal information, the voice pitch. By intentionally controlling the pitch, a user can enter the same word with it having different meanings (functions) without explicitly changing the speech-input mode. Our speech-shift function implemented on a voice-enabled word processor, for example, can distinguish an utterance with a high pitch from one with a normal (low) pitch, and regard the former as voice-command-mode input(suchasfile-menuandedit-menucommands)andthelatter as regular dictation-mode text input. Our experimental results from twenty subjects showed that the speech-shift function is effective, easy to use, and a labor-saving input method.
Masataka Goto, Yukihiro Omoto, Katunobu Itou, Tetsunori Kobayashi
INTERSPEECH1
2003 Speech starter: noise-robust endpoint detection by using filled pauses
abstract
In this paper we propose a speech interface function, called speech starter, that enables noise-robust endpoint (utterance) detection for speech recognition. When current speech recognizers are used in a noisy environment, a typical recognition error is caused by incorrect endpoints because their automatic detection is likely to be disturbed by non-stationary noises. The speech starter function enables a user to specify the beginning of each utterance by uttering a filler with a filled pause, which is used as a trigger to start speech-recognition processes. Since filled pauses can be detected robustly in a noisy environment, practical endpoint detection is achieved. Speech starter also offers the advantage of providing a hands-free speech interface and it is user-friendly because a speaker tends to utter filled pauses (e.g., “er...”) at the beginning of utterances when hesitating in human-human communication. Experimental results from a 10-dB-SNR noisy environment show that the recognition error rate with speech starter was lower than with conventional endpoint-detection methods. 1.
Koji Kitayama, Masataka Goto, Katunobu Itou, Tetsunori Kobayashi
INTERSPEECH2
2003 SmartMusicKIOSK: music listening station with chorus-search function
abstract
This paper describes a new music-playback interface for trial listening, SmartMusicKIOSK. In music stores, short trial listening of CD music is not usually a passive experience -- customers often search out the chorus or "hook" of a song using the fast-forward button. Listening of this type, however, has not been traditionally supported. This research achieves a function for jumping to the chorus section and other key parts of a song plus a function for visualizing song structure. These functions make it easier for a listener to find desired parts of a song and thereby facilitate an active listening experience. The proposed functions are achieved by an automatic chorus-section detecting method, and the results of implementing them as a listening station have demonstrated their usefulness.
Masataka Goto
UIST1
2002 Speech completion: on-demand completion assistance using filled pauses for speech input interfaces
abstract
This paper describes a novel speech interface function, called speech completion, that helps a user enter a word or phrase by completing (filling in the rest of) a phrase fragment uttered by the user. Although the concept of completion is widely used in text-based interfaces, there have been no reports of completion being effectively applied to speech. By using a filled pause, we enable a user to effortlessly invoke the speech-completion function which helps the user recall uncertain phrases and saves labor when the input phrase is long. When a user hesitates by lengthening a vowel (a filled pause is uttered) during a phrase, our system immediately displays completion candidates whose beginnings acoustically resemble the uttered fragment so that the user can select the correct one. In our experiments with a system that included a filled-pause detector and a speech recognizer capable of listing candidates, the effectiveness of speech completion was confirmed.
Masataka Goto, Katunobu Itou, Satoru Hayamizu
INTERSPEECH1
2001 A predominant-F0 estimation method for CD recordings: MAP estimation using EM algorithm for adaptive tone models
abstract
This paper describes a predominant-F/sub 0/ (fundamental frequency) estimation method called PreFEst, which can detect melody and bass lines in monaural audio signals containing sounds of various instruments, While most previous methods premised mixtures of a few sounds and had difficulty dealing with such complex signals, our method can estimate the F/sub 0/ of the melody and bass lines without assuming the number of sound sources in compact-disc recordings. In this paper we propose the following three extensions to our previous PreFEst to make it more adaptive and flexible: introducing multiple harmonic-structure tone models, estimating the shape of tone models, and introducing a prior distribution of its shape and F/sub 0/ estimates These extensions were implemented by the MAP (maximum a posteriori probability) estimation by using the expectation-maximization algorithm. Experimental results with compact-disc recordings showed that our real-time system based on the extended PreFEst achieved performance improvement.
Masataka Goto
ICASSP1
2001 Real-time sound source localization and separation system and its application to automatic speech recognition
abstract
A real-time sound localization/separation system for near-field sound sources was constructed and evaluated in a real office environment. As for the sound localization, the experimental results showed that the direction of the two sources was estimated with high accuracy while the range of the sources was estimated with moderate accuracy. As for the sound separation, a recognition rate of 70 % for an on-line recognizer on a network and of 90% for an off-line recognizer were achieved, respectively.
Futoshi Asano, Masataka Goto, Katunobu Itou, Hideki Asoh
INTERSPEECH2
2000 A robust predominant-F0 estimation method for real-time detection of melody and bass lines in CD recordings
abstract
This paper describes a robust method for estimating the fundamental frequency (F0) of melody and bass lines in monaural real-world musical audio signals containing sounds of various instruments. Most previous F0-estimation methods had great difficulty dealing with such complex audio signals because they were designed to deal with mixtures of only a few sounds. To make it possible to estimate the F0 of the melody and bass lines, we propose a predominant-F0 estimation method called PreFEst that does not rely on the F0's unreliable frequency component and obtains the most predominant F0 supported by harmonics within an intentionally limited frequency range. It evaluates the relative dominance of every possible F0 by using the expectation-maximization algorithm and considers the temporal continuity of F0s by using a multiple-agent architecture. Experimental results show that our real-time system can detect the melody and bass lines in audio signals sampled from commercially distributed compact discs.
Masataka Goto
ICASSP1
1999 A real-time filled pause detection system for spontaneous speech recognition
abstract
This paper describes a method for automatically detecting filled (vocalized) pauses, which are one of the hesitation phenomena that current speech recognizers typically cannot handle. The detection of these pauses is important in spontaneous speech dialogue systems because they play valuable roles, such as helping a speaker keep a conversational turn, in oral communication. Although a few speech recognition systems have processed filled pauses within subword-based connected word recognition or word-spotting frameworks, they did not detect the pauses individually and consequently could not consider their roles. In this paper we propose a method that detects filled pauses and word lengthening on the basis of small fundamental frequency transition and small spectral envelope deformation under the assumption that speakers do not change articulator parameters during filled pauses. Experimental results for a Japanese spoken dialogue corpus show that our real-time filled-pause-detection system yielded a recall rate of 84.9 % and a precision rate of 91.5%.
Masataka Goto, Katunobu Itou, Satoru Hayamizu
EUROSPEECH1
1999 Real-time beat tracking for drumless audio signals: Chord change detection for musical decisions
Masataka Goto, Yoichi Muraoka
Speech Commun.1
1996 Localization by harmonic structure and its application to harmonic sound stream segregation
abstract
Sound stream segregation is essential for understanding auditory events in the real-world. In this paper, we present a new method for sound stream segregation using harmonic structure and localization, or direction, in the horizontal plane. The direction of the sound source is determined by using the harmonic structure extracted from binaural inputs. The fundamental frequency of each sound is then refined by using the direction of its source. This paper discusses how the effectiveness of the harmonic-based stream segregation system (HBSS) is improved by incorporating the new method and presents the binaural HBSS (Bi-HBSS). In particular, experimental results show that the Bi-HBSS reduces the spectrum distortions and the fundamental frequency errors, compared with the HBSS and with a direction-based stream segregation system.
Tomohiro Nakatani, Masataka Goto, Hiroshi G. Okuno
ICASSP2
1994 A Beat Tracking System for Acoustic Signals of Music
abstract
This paper presents a beat tracking system that processes acoustic signals of music and recognizes temporal positions of beats in time. Musical beat tracking is needed by various multimedia applications such as video editing, audio editing, and stage lighting control. Previous systems were not able to deal with acoustic signals that contained sounds of various instruments, especially drums. They dealt with either MIDI signals or acoustic signals played on a few instruments, and in the latter case, did not work in real time. Our system deals with popular music in which drums maintain the beat. Because our system examines multiple hypotheses in parallel, it can follow beats without losing track of them, even if some hypotheses become wrong. Our system has been implemented on a parallel computer, the Fujitsu AP1000. In our experiment, the system correctly tracked beats in 27 out of 30 commercially distributed popular songs.
Masataka Goto, Yoichi Muraoka
ACM Multimedia1