Satoru Fukayama

dblp:38/7252 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-6506-2796ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Benchmarking Prosody Encoding in Discrete Speech Tokens
abstract
Recently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks. However, these discrete tokens are typically learned in advance, separately from the training of language models or downstream tasks. As a result, choices related to discretization, such as the SSL model used or the number of clusters, must be made heuristically. In particular, speech language models are expected to understand and generate responses that reflect not only the semantic content but also prosodic features. Yet, there has been limited research on the ability of discrete tokens to capture prosodic information. To address this gap, this study conducts a comprehensive analysis focusing on prosodic encoding based on their sensitivity to the artificially modified prosody, aiming to provide practical guidelines for designing discrete tokens.
Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu
ASRU2
2025 Investigation of Spatial Self-Supervised Learning and Its Application to Target Speaker Speech Recognition
abstract
In this paper, we investigate spatial self-supervised learning for target speaker speech recognition. Neural separation models can be trained in a self-supervised manner by using only multichannel mixture signals. Such a framework is typically based on a physics-informed generative model, widely studied in blind source separation (BSS). Our study focuses on the application of spatial self-supervised learning for distant speech recognition (DSR). Multi-talker DSR systems often rely on a BSS method called guided source separation (GSS), which separates target speech utterances from multichannel mixtures and their speaker activities. Its performance, however, would be limited by its assumption that each time-frequency (TF) bin contains only one sound source or noise. To address this limitation, factor models have been studied by assuming each TF bin as a sum of all the sources. We investigate both the classic BSS and self-supervised methods based on factor models by evaluating them extensively on multiple different DSR scenarios in the CHiME-8 DASR challenge. We show that the classic BSS methods based on factor models suffer from initialization sensitivity, while self-supervised learning mitigated this problem and slightly outperformed GSS. The best-performing method and audio samples are available online at https://ybando.jp/projects/icassp2025/.
Yoshiaki Bando, Samuele Cornell, Satoru Fukayama, Shinji Watanabe 0001
ICASSP3
2025 Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu
INTERSPEECH3
2025 Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu
INTERSPEECH3
2025 Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
Hitoshi Suda, Shinnosuke Takamichi, Satoru Fukayama
INTERSPEECH3
2025 SingDistVis: interactive Overview+Detail visualization for F0 trajectories of numerous singers singing the same song
abstract
Abstract This paper describes SingDistVis, an information visualization technique for fundamental frequency (F0) trajectories of large-scale singing data where numerous singers sing the same song. SingDistVis allows to explore F0 trajectories interactively by combining two views: OverallView and DetailedView. OverallView visualizes a distribution of the F0 trajectories of the song in a time-frequency heatmap. When a user specifies an interesting part, DetailedView zooms in on the specified part and visualizes singing assessment (rating) results. Here, it displays high-rated singings in red and low-rated singings in blue. When the user clicks on a particular singing, the audio source is played and its F0 trajectory through the song is displayed in OverallView. We selected heatmap-based visualization for OverallView to provide an overview of a large-scale F0 dataset, and polyline-based visualization for DetailedView to provide a more precise representation of a small number of particular F0 trajectories. This paper introduces a subjective experiment using 1,000 singing voices to determine suitable visualization parameters. Then, this paper presents user evaluations where we asked participants to compare visualization results of four types of Overview+Detail designs and concluded that the presented design archived better evaluations than other designs in all the seven questions. Finally, this paper describes a user experiment in which eight participants compare SingDistVis with a baseline implementation in exploring interested singing voices and concludes that the proposed SingDistVis archived better evaluations in nine of the questions.
Takayuki Itoh, Tomoyasu Nakano, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto
Multim. Tools Appl.3
2024 Self-Supervised Speech Representations are More Phonetic than Semantic
Kwanghee Choi, Ankita Pasad, Tomohiko Nakamura, Satoru Fukayama, Karen Livescu, Shinji Watanabe 0001
INTERSPEECH4
2023 jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
abstract
We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese children’s songs and have six voice parts (lead vocal, soprano, alto, tenor, bass, and vocal percussion). They are divided into seven subsets, each of which features typical characteristics of a music genre such as jazz and enka. The variety in genre and voice part match vocal ensembles recently widespread in social media services such as YouTube, although the main targets of conventional vocal ensemble datasets are choral singing made up of soprano, alto, tenor, and bass. Experimental evaluation demonstrates that our corpus is a challenging resource for vocal ensemble separation. Our corpus is available on our project page.
Tomohiko Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama, Hiroshi Saruwatari
ICASSP4
2022 Exploiting Fine-tuning of Self-supervised Learning Models for Improving Bi-modal Sentiment Analysis and Emotion Recognition
Satoru Fukayama, Panikos Heracleous, Jun Ogata
INTERSPEECH2
2022 Singer Diarization for Polyphonic Music With Unison Singing
abstract
This paper introduces a new framework for singer diarization, which is a technique to reveal who sings when in songs with multiple singers. Although various techniques have been developed to analyze and extract features of singing voices in musical audio signals, most of them assume that a song is sung by a single singer, and singer diarization for multiple singers has not been well studied in the field of singing information processing. To deal with multiple speakers in speech analysis, speaker diarization has been explored to handle overlapped speech voices, but cannot handle singing voices well because of acoustic differences between singing and speech voices. This paper therefore proposes a new diarization framework specialized in singing voices. To achieve high accuracy in overlap detection, this paper proposes a novel acoustic feature named Cosacorr score, which is helpful in estimating whether a song is sung by more than one singer. After extracting singing voices from polyphonic music by using a singing voice separation technique, the framework adopts an existing ArcFace technique to extract discriminative singer representations from short segments of the separated singing voices. The framework is evaluated by using a new private dataset of unison singing voices, which is constructed using commercially available compact discs (CDs). The experimental results show that the proposed framework outperformed the baseline method for speaker diarization in terms of diarization error rate (DER).
Hitoshi Suda, Daisuke Saito, Satoru Fukayama, Tomoyasu Nakano, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Audio-visual object removal in 360-degree videos
abstract
Abstract We present a novel concept audio–visual object removal in 360-degree videos, in which a target object in a 360-degree video is removed in both the visual and auditory domains synchronously. Previous methods have solely focused on the visual aspect of object removal using video inpainting techniques, resulting in videos with unreasonable remaining sounds corresponding to the removed objects. We propose a solution which incorporates direction acquired during the video inpainting process into the audio removal process. More specifically, our method identifies the sound corresponding to the visually tracked target object and then synthesizes a three-dimensional sound field by subtracting the identified sound from the input 360-degree video. We conducted a user study showing that our multi-modal object removal supporting both visual and auditory domains could significantly improve the virtual reality experience, and our method could generate sufficiently synchronous, natural and satisfactory 360-degree videos.
Ryo Shimamura, Yuki Koyama 0001, Takayuki Nakatsuka, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto, Shigeo Morishima
Vis. Comput.5
2019 Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention Mechanism
abstract
This paper describes neural singing transcription that estimates a sequence of musical notes directly from the audio signal of singing voice in an end-to-end manner without time-aligned training data. A conventional approach to singing transcription is to perform vocal F0 estimation followed by musical note estimation. The performance of this approach, however, is severely limited because the F0 estimation errors propagate to the note estimation step and rich acoustic information cannot be used. In addition, it is difficult and time-consuming to split continuous signals of singing voices into segments corresponding to musical notes for making precise time-aligned transcriptions. To solve these problems, we use an encoder-decoder model with an attention mechanism that can automatically learn an input-output alignment and mapping, even from non-aligned training data. The main challenge of our study is to estimate temporal categories (note values) in addition to instantaneous categories (pitches). We thus propose a novel loss function for the attention weights of time-aligned notes for semi-supervised alignment training. By gradually reducing the weight of the loss function, a better input-output alignment can be learned much more quickly. We showed that our method performed well for isolated singing voice in popular music.
Ryo Nishikimi, Eita Nakamura, Satoru Fukayama, Masataka Goto, Kazuyoshi Yoshii
ICASSP3
2019 Transdrums: A Drum Pattern Transfer System Preserving Global Pattern Structure
abstract
This paper presents TransDrums, which is a system that transfers drum patterns from a drum-pattern-source song (D-song) to a base song (B-song) and synthesizes the audio with the substituted drum pattern. Typical drum parts consist of multiple drum patterns that are concatenated to form a structure by, for example, inserting fill-in patterns at structural boundaries. The previous system that replaced the drum parts was not able to form such a structure. Therefore, we propose TransDrums, which extracts and transfers multiple drum patterns to form the structure. It takes two songs as the input and extracts multiple typical drum patterns from each song. It then makes pairs of those patterns between B-song and D-song and replaces them using the counterpart drum patterns to synthesize audio with the altered drum pattern. To achieve the key idea of properly replacing a drum phrase in B-song with that in D-song, it is necessary to model the structure of the drum parts by analyzing the transition probabilities between the typical drum patterns. The appropriate pairs are determined so that the sum of the Jensen-Shannon divergence between the transition probabilities is minimized. Our experimental results show that TransDrums can generate audio to change by altering the drum patterns with the structure.
Shun Sawada, Satoru Fukayama, Masataka Goto, Keiji Hirata 0001
ICASSP2
2019 Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model
abstract
This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical language model that represents the characteristics of each part in order to solve the ambiguity of part assignment and estimate a musically-natural score. We propose a factorial hidden semi-Markov model that consists of three language models corresponding to the three guitar parts (three latent chains) and an acoustic model of a mixture spectrogram (emission model). The language model for rhythm guitar represents a homophonic sequence of musical notes (chord sequence) and those for lead and bass guitars represent a monophonic sequence of musical notes in a higher and lower frequency range respectively. The acoustic model represents a spectrogram as a sum of low-rank spectrograms of the three guitar parts approximated by NMF. Given a spectrogram, we estimate the note sequences using Gibbs sampling. We show that our model outperforms a state-of-the-art multi-pitch detection method in the accuracy and naturalness of the transcribed scores.
Kentaro Shibata, Ryo Nishikimi, Satoru Fukayama, Masataka Goto, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii
ICASSP3
2019 Audio-Based Automatic Generation of a Piano Reduction Score by Considering the Musical Structure
Hirofumi Takamori, Takayuki Nakatsuka, Satoru Fukayama, Masataka Goto, Shigeo Morishima
MMM (2)3
2019 Query-by-Dancing: A Dance Music Retrieval System Based on Body-Motion Similarity
Shuhei Tsuchida, Satoru Fukayama, Masataka Goto
MMM (1)2
2019 ABCPRec: Adaptively Bridging Consumer and Producer Roles for User-Generated Content Recommendation
abstract
In Web services dealing with user-generated content (UGC), a user can have two roles: a role of a consumer and that of a producer. Since most item recommendation models have only considered the role of a user as a consumer, how to leverage the two roles to improve UGC recommendation accuracy has been underexplored. In this paper, based on the state-of-the-art UGC recommendation method called CPRec (consumer and producer based recommendation), we propose ABCPRec (adaptively bridging CPRec). Unlike CPRec, which assumes that the two roles of a user are always related to each other, ABCPRec adaptively bridges the two roles according to the similarity between her nature as a consumer and that as a producer. This enables the model to learn each user's characteristics as both a consumer and a producer and to recommend items to each user more accurately. By using two real-world datasets, we showed that our proposed method significantly outperformed comparative methods in terms of AUC.
Kosetsu Tsukuda, Satoru Fukayama, Masataka Goto
SIGIR2
2018 A Melody-Conditioned Lyrics Language Model
abstract
Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano
NAACL-HLT3
2017 Automatic System for Editing Dance Videos Recorded Using Multiple Cameras
Shuhei Tsuchida, Satoru Fukayama, Masataka Goto
ACE2
2017 LyriSys: An Interactive Support System for Writing Lyrics Based on Topic Transition
abstract
This paper presents LyriSys, a novel lyric-writing support system. Previous systems for lyric writing can fully automatically only generate a single line of lyrics that satisfies given constraints on accent and syllable patterns or an entire lyric. In contrast to such systems, LyriSys allows users to create and revise their work incrementally in a trial-and-error manner. Through fine-grained interactions with the system, the user can create the specifications of the musical structure and the story of the lyrics in terms of the verse-bridge-chorus structure, the number of lines, words and syllables, and most importantly, the transition over semantic topics such as "scene", "dark" and "sweet love". This paper provides an overview of the design of the system and its user interface and describes how the writing process is guided by a state-of-the-art probabilistic generative topic model that is trained without supervision. The system works for both Japanese and English.
Kento Watanabe, Yuichiroh Matsubayashi, Kentaro Inui, Tomoyasu Nakano, Satoru Fukayama, Masataka Goto
IUI5
2016 Modeling Discourse Segments in Lyrics Using Repeated Patterns
abstract
This study proposes a computational model of the discourse segments in lyrics to understand and to model the structure of lyrics. To test our hypothesis that discourse segmentations in lyrics strongly correlate with repeated patterns, we conduct the first large-scale corpus study on discourse segments in lyrics. Next, we propose the task to automatically identify segment boundaries in lyrics and train a logistic regression model for the task with the repeated pattern and textual features. The results of our empirical experiments illustrate the significance of capturing repeated patterns in predicting the boundaries of discourse segments in lyrics.
Kento Watanabe, Yuichiroh Matsubayashi, Naho Orita, Naoaki Okazaki, Kentaro Inui, Satoru Fukayama, Tomoyasu Nakano, Jordan B. L. Smith, Masataka Goto
COLING6
2016 Music emotion recognition with adaptive aggregation of Gaussian process regressors
abstract
This paper describes a novel method for estimating the emotions elicited by a piece of music from its acoustic signals. Previous research in this field has centered on finding effective acoustic features and regression methods to relate features to emotions. The state-of-the-art method is based on a multi-stage regression, which aggregates the results from different regressors trained with training data. However, after training, the aggregation happens in a fixed way and cannot be adapted to acoustic signals with different musical properties. We propose a method that adapts the aggregation by taking into account new acoustic signal inputs. Since we cannot know the emotions elicited by new inputs beforehand, we need a way of adapting the aggregation weights. We do so by exploiting the deviation observed in the training data using Gaussian process regressions. We confirmed with an experiment comparing different aggregation approaches that our adaptive aggregation is effective in improving recognition accuracy.
Satoru Fukayama, Masataka Goto
ICASSP1
2015 AutoGuitarTab: Computer-Aided Composition of Rhythm and Lead Guitar Parts in the Tablature Space
abstract
We present AutoGuitarTab, a system for generating realistic guitar tablature given an input symbolic chord and key sequence. Our system consists of two modules: AutoRhythmGuitar and AutoLeadGuitar. The first of these generates rhythm guitar tablatures which outline the input chord sequence in a particular style (using Markov chains to ensure playability) and performs a structural analysis to produce a structurally consistent composition. AutoLeadGuitar generates lead guitar parts in distinct musical phrases, guiding the pitch classes towards chord tones and steering the evolution of the rhythmic and melodic intensity according to user preference. Experimentally, we uncover musician-specific trends in guitar playing style, and demonstrate our system's ability to produce playable, realistic and style-specific tablature using a combination of algorithmic, user-surveyed and expert evaluation techniques.
Matt McVicar, Satoru Fukayama, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Automated choreography synthesis using a Gaussian process leveraging consumer-generated dance motions
abstract
We propose a novel method of automatically generating dance choreography using machine learning. In a typical approach to automatic choreography, a dance is constructed by concatenating segments of existing dances which are maximally correlated to the target audio features with connectivity constraints. However, researchers using this approach are unable to produce dances with much variety, since the set of examples used in these experiments (usually motion-capture of existing choreographies) is limited and costly to produce. To solve this issue, we propose a probabilistic model which maps beat structures to dance movements using a Gaussian process, trained with a large amount of consumer-generated dance motion obtained from the web. The main contribution of our work is the combination of two approaches: the previously mentioned correlation based approach which seeks for relationships between music and dance, and a machine learning approach which is based on human motion modeling. Inspection of the generated dances proves that our method can generate choreographies with different characters by switching the training dataset, and highlights opportunities in training with further dance motions on the web to generate more expressive dance choreography.
Satoru Fukayama, Masataka Goto
Advances in Computer Entertainment1
2009 Orpheus: Automatic Composition System Considering Prosody of Japanese Lyrics
Satoru Fukayama, Kei Nakatsuma, Shinji Sako, Yuichiro Yonebayashi, Tae Hun Kim, Si Wei Qin, Takuho Nakano, Takuya Nishimoto, Shigeki Sagayama
ICEC1