EDBT 2026 Demo / reviewers in the wild / expert
Tomoyasu Nakano
dblp:49/5220
· DBLP profile ↗
26ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0001-8014-2209ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 88% Multimedia systems and quality of experience · 12% | |
| Human-computer interaction and pervasive computing
2 papers |
User interface design and tools · 52% Human-AI interaction · 48% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › music information retrieval
singing information processing |
0.6 | 1 | 2022 | Singer Diarization for Polyphonic Music With Unison Singing · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization |
0.2 | 1 | 2015 | TextAlive: Integrated Design Environment for Kinetic Typography · CHI 2015 |
User interface design and tools
authoring tools |
0.2 | 1 | 2015 | TextAlive: Integrated Design Environment for Kinetic Typography · CHI 2015 |
Audio and music processing › source separation › music source separation
singing voice separation |
0.2 | 1 | 2022 | Singer Diarization for Polyphonic Music With Unison Singing · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
User interface design and tools
designer-developer collaboration |
0.1 | 1 | 2015 | TextAlive: Integrated Design Environment for Kinetic Typography · CHI 2015 |
Methods — techniques the papers use, named apart from their topics
user study · 2.2text-to-speech · 1.7cosacorr score · 0.6arcface · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer VoiceabstractKawaii" is the Japanese concept of cute, which carries sociocultural connotations related to social identities and emotional responses.Yet, virtually all work to date has focused on the visual side of kawaii, including in studies of computer agents and social robots.In pursuit of formalizing the new science of kawaii vocalics, we explored what elements of voice relate to kawaii and how they might be manipulated, manually and automatically.We conducted a four-phase study (grand 𝑁 = 512) with two varieties of computer voices: text-to-speech (TTS) and game character voices.We found kawaii "sweet spots" through manipulation of fundamental and formant frequencies, but only for certain voices and to a certain extent.Findings also suggest a ceiling effect for the kawaii vocalics of certain voices.We offer empirical validation of the preliminary kawaii vocalics model and an elementary method for manipulating kawaii perceptions of computer voice. Yuto Mandai, Katie Seaborn, Tomoyasu Nakano, Xin Sun 0016, Yijia Wang 0001, Jun Kato 0001 |
CHI | 3 |
| 2025 | SingDistVis: interactive Overview+Detail visualization for F0 trajectories of numerous singers singing the same songabstractAbstract This paper describes SingDistVis, an information visualization technique for fundamental frequency (F0) trajectories of large-scale singing data where numerous singers sing the same song. SingDistVis allows to explore F0 trajectories interactively by combining two views: OverallView and DetailedView. OverallView visualizes a distribution of the F0 trajectories of the song in a time-frequency heatmap. When a user specifies an interesting part, DetailedView zooms in on the specified part and visualizes singing assessment (rating) results. Here, it displays high-rated singings in red and low-rated singings in blue. When the user clicks on a particular singing, the audio source is played and its F0 trajectory through the song is displayed in OverallView. We selected heatmap-based visualization for OverallView to provide an overview of a large-scale F0 dataset, and polyline-based visualization for DetailedView to provide a more precise representation of a small number of particular F0 trajectories. This paper introduces a subjective experiment using 1,000 singing voices to determine suitable visualization parameters. Then, this paper presents user evaluations where we asked participants to compare visualization results of four types of Overview+Detail designs and concluded that the presented design archived better evaluations than other designs in all the seven questions. Finally, this paper describes a user experiment in which eight participants compare SingDistVis with a baseline implementation in exploring interested singing voices and concludes that the proposed SingDistVis archived better evaluations in nine of the questions. Takayuki Itoh, Tomoyasu Nakano, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto |
Multim. Tools Appl. | 2 |
| 2022 | Singer Diarization for Polyphonic Music With Unison SingingabstractThis paper introduces a new framework for singer diarization, which is a technique to reveal who sings when in songs with multiple singers. Although various techniques have been developed to analyze and extract features of singing voices in musical audio signals, most of them assume that a song is sung by a single singer, and singer diarization for multiple singers has not been well studied in the field of singing information processing. To deal with multiple speakers in speech analysis, speaker diarization has been explored to handle overlapped speech voices, but cannot handle singing voices well because of acoustic differences between singing and speech voices. This paper therefore proposes a new diarization framework specialized in singing voices. To achieve high accuracy in overlap detection, this paper proposes a novel acoustic feature named Cosacorr score, which is helpful in estimating whether a song is sung by more than one singer. After extracting singing voices from polyphonic music by using a singing voice separation technique, the framework adopts an existing ArcFace technique to extract discriminative singer representations from short segments of the separated singing voices. The framework is evaluated by using a new private dataset of unison singing voices, which is constructed using commercially available compact discs (CDs). The experimental results show that the proposed framework outperformed the baseline method for speaker diarization in terms of diarization error rate (DER). Hitoshi Suda, Daisuke Saito, Satoru Fukayama, Tomoyasu Nakano, Masataka Goto |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | An automated system recommending background music to listen to while workingabstractAbstract Many people listen to music while working nowadays. However, conventional recommendation systems that are designed for playing songs matching user preferences cannot be applied for such a situation. This is because previous research showed that listeners’ concentration can be negatively affected not only by music that listeners strongly dislike but also by music that the listeners strongly like. Therefore, when we consider a recommendation system to be used while working, it is desirable to avoid both songs the user likes very much and songs the user dislikes very much. Given this background, we propose FocusMusicRecommender, a system designed specifically for recommending music to listen to while working. It summarizes songs automatically and plays them successively in order to enable users to give not only “dislike (very much)” feedback via a “skip” button but also “like (very much)” feedback via a “keep listening” button. The feedback is then combined with the users’ concentration level that is estimated from their behavioral history during the playback of the corresponding song, which allows the system to obtain preference information that distinguishes between “like” and “like very much” without burdening the user who is working. Based on the preference information, the system estimates the preference levels of unplayed songs and prioritizes the songs for subsequent playback by also considering the user’s current concentration level. Our experiments showed the validity and effectiveness of the proposed method, including the accuracy of the concentration level estimation. Moreover, our user study verified the suitability of the recommendation results from both the observed behavior and obtained comments of the participants. Hiromu Yakura, Tomoyasu Nakano, Masataka Goto |
User Model. User Adapt. Interact. | 2 |
| 2020 | Interactive deep singing-voice separation based on human-in-the-loop adaptationabstractThis paper presents a deep-learning-based interactive system separating the singing voice from input polyphonic music signals. Although deep neural networks have been successful for singing voice separation, no approach using them allows any user interaction for improving the separation quality. We present a framework that allows a user to interactively fine-tune the deep neural model at run time to adapt it to the target song. This is enabled by designing unified networks consisting of two U-Net architectures based on frequency spectrogram representations: one for estimating the spectrogram mask that can be used to extract the singing-voice spectrogram from the input polyphonic spectrogram; the other for estimating the fundamental frequency (F0) of the singing voice. Although it is not easy for the user to edit the mask, he or she can iteratively correct errors in part of the visualized F0 trajectory through simple interaction. Our unified networks leverage the user-corrected F0 to improve the rest of the F0 trajectory through the model adaptation, which results in better separation quality. We validated this approach in a simulation experiment showing that the F0 correction can improve the quality of singing-voice separation. We also conducted a pilot user study with an expert musician, who used our system to produce a high-quality singing-voice separation result. Tomoyasu Nakano, Yuki Koyama 0001, Masahiro Hamasaki, Masataka Goto |
IUI | 1 |
| 2018 | Instlistener: An Expressive Parameter Estimation System Imitating Human Performances of Monophonic Musical InstrumentsabstractWe present InstListener, a system that takes an expressive monophonic solo instrument performance by a human performer as the input and imitates its audio recordings by using an existing MIDI (Musical Instrument Digital Interface) synthesizer. It automatically analyzes the input and estimates, for each musical note, expressive performance parameters such as the timing, duration, discrete semitone-level pitch, amplitude, continuous pitch contour, and continuous amplitude contour. The system uses an iterative process to estimate and update those parameters by analyzing both the input and output of the system so that the output from the MIDI synthesizer can be similar enough to the input. Our evaluation results showed that the iterative parameter estimation improved the accuracy of imitating of the input performance and thus increased the naturalness and expressiveness of the output performance. Zhengshan Shi, Tomoyasu Nakano, Masataka Goto |
ICASSP | 2 |
| 2018 | FocusMusicRecommender: A System for Recommending Music to Listen to While WorkingabstractThis paper proposes FocusMusicRecommender, an automated system recommending background music to listen to while working. Recommendation systems matching user preferences have been widely researched even though research has shown that music that listeners strongly like is not suitable background music because it interferes with their concentration. FocusMusicRecommender plays songs that users may "neither like nor dislike" instead of "like very much." It is designed to by default summarize a song automatically so that users can give "like very much" feedback by pressing a "keep listening" button or "dislike very much" feedback by pressing a "skip" button. It uses this feedback, along with users» concentration levels estimated from their behavior history, to distinguish between the preference levels "like" and "like very much." It then estimates the preference levels of unplayed songs and selects the most suitable song by considering the user»s current concentration level. The effectiveness of the proposed feedback method and suitability of the recommendation results were verified experimentally and in user studies. Furthermore, it is confirmed that the proposed method can estimate the user»s concentration level more accurately than the previous methods. Hiromu Yakura, Tomoyasu Nakano, Masataka Goto |
IUI | 2 |
| 2018 | A Melody-Conditioned Lyrics Language ModelabstractKento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano |
NAACL-HLT | 6 |
| 2017 | LyriSys: An Interactive Support System for Writing Lyrics Based on Topic TransitionabstractThis paper presents LyriSys, a novel lyric-writing support system. Previous systems for lyric writing can fully automatically only generate a single line of lyrics that satisfies given constraints on accent and syllable patterns or an entire lyric. In contrast to such systems, LyriSys allows users to create and revise their work incrementally in a trial-and-error manner. Through fine-grained interactions with the system, the user can create the specifications of the musical structure and the story of the lyrics in terms of the verse-bridge-chorus structure, the number of lines, words and syllables, and most importantly, the transition over semantic topics such as "scene", "dark" and "sweet love". This paper provides an overview of the design of the system and its user interface and describes how the writing process is guided by a state-of-the-art probabilistic generative topic model that is trained without supervision. The system works for both Japanese and English. Kento Watanabe, Yuichiroh Matsubayashi, Kentaro Inui, Tomoyasu Nakano, Satoru Fukayama, Masataka Goto |
IUI | 4 |
| 2016 | Modeling Discourse Segments in Lyrics Using Repeated PatternsabstractThis study proposes a computational model of the discourse segments in lyrics to understand and to model the structure of lyrics. To test our hypothesis that discourse segmentations in lyrics strongly correlate with repeated patterns, we conduct the first large-scale corpus study on discourse segments in lyrics. Next, we propose the task to automatically identify segment boundaries in lyrics and train a logistic regression model for the task with the repeated pattern and textual features. The results of our empirical experiments illustrate the significance of capturing repeated patterns in predicting the boundaries of discourse segments in lyrics. Kento Watanabe, Yuichiroh Matsubayashi, Naho Orita, Naoaki Okazaki, Kentaro Inui, Satoru Fukayama, Tomoyasu Nakano, Jordan B. L. Smith, Masataka Goto |
COLING | 7 |
| 2016 | An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singerabstractThis paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice timbre expression words, such as "Age" and "Gender", and they usually need to be manually assigned to individual singers' singing voices through listening. To make it possible to automatically estimate them from given singer's singing voices, an acoustic feature to well capture only each singer's voice timbre is extracted with a Gaussian mixture model trained using parallel data between singing voices sung by many pre-stored target singers and same voices sung by a reference singer. Then, the voice timbre evaluation values are estimated from the extracted feature using regression models. The experimental results showed that the proposed method is capable of accurately estimating those values for some expression words, such as "Age" and "Gender", and nonlinear regression is effective for the expression words, "Powerfulness" and "Uniqueness." Soichi Yamane, Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2016 | A soundtrack generation system to synchronize the climax of a video clip with musicabstractIn this paper, we present a soundtrack generation system that can automatically add a soundtrack with the length and climax points aligned to those of a video clip. Adding a soundtrack to a video clip is an important process in video editing. Editors tend to add chorus sections to the climax points of the video clip by replacing and concatenating musical segments. However, this process is time-consuming. Our system automatically detects climaxes of both the video clips and music based on feature extraction and analysis. This enables the system to add a soundtrack in which the climax is synchronized to the climax of the video clip. We evaluated the generated soundtracks through a subjective evaluation. Haruki Sato, Tatsunori Hirai, Tomoyasu Nakano, Masataka Goto, Shigeo Morishima |
ICME | 3 |
| 2016 | PlaylistPlayer: An Interface Using Multiple Criteria to Change the Playback Order of a Music PlaylistabstractWe propose a novel interface that allows the user to interactively change the playback order of multiple songs by choosing one or more criteria. The criteria include not only the song's title and artist name but also its content automatically estimated by music/singing signal processing and artist-level social analysis. The artist-level social information is discovered from Wikipedia and DBpedia. With regard to manipulating playback order, existing interfaces typically allow the user to change it manually or automatically by choosing one of a few types of criteria. The proposed interface, on the other hand, deals with nine properties and multiple integrations of them (e.g., vocal gender and beats per minute). To realize the ordering by multiple criteria, a distance matrix is computed from the criteria vectors and is then used to estimate paths for ascending, descending, and random orders by applying principle component analysis or to estimate a path for a smooth order by solving the travelling salesman problem. Tomoyasu Nakano, Jun Kato 0001, Masahiro Hamasaki, Masataka Goto |
IUI | 1 |
| 2015 | TextAlive: Integrated Design Environment for Kinetic TypographyabstractThis paper presents TextAlive, a graphical tool that allows interactive editing of kinetic typography videos in which lyrics or transcripts are animated in synchrony with the corresponding music or speech. While existing systems have allowed the designer and casual user to create animations, most of them do not take into account synchronization with audio signals. They allow predefined motions to be applied to objects and parameters to be tweaked, but it is usually impossible to extend the predefined set of motion algorithms within these systems. We therefore propose an integrated design environment featuring (1) GUIs that designers can use to create and edit animations synchronized with audio signals, (2) integrated tools that programmers can use to implement animation algorithms, and (3) a framework for bridging the interfaces for designers and programmers. A preliminary user study with designers, programmers, and casual users demonstrated its capability in authoring various kinetic typography videos. Jun Kato 0001, Tomoyasu Nakano, Masataka Goto |
CHI | 2 |
| 2015 | Songle Widget: Making Animation and Physical Devices Synchronized with Music Videos on the WebabstractThis paper describes a web-based multimedia development framework, Songle Widget, that makes it possible to control computer-graphic animation and physical devices such as lighting devices and robots in synchronization with music publicly available on the web. To avoid the difficulty of time-consuming manual annotation, Songle Widget makes it easy to develop web-based applications with rigid music synchronization by leveraging music-understanding technologies. Four types of musical elements (music structure, hierarchical beat structure, melody line, and chords) have been automatically annotated for more than 920,000 songs on music-or video-sharing services and can readily be used by music-synchronized applications. Since errors are inevitable when elements are annotated automatically, Songle Widget takes advantage of a user-friendly crowdsourcing interface that enables users to correct them. This is effective when applications require error-free annotation. We made Songle Widget open to the public, and its capabilities and usefulness have been demonstrated in seven music-synchronized applications. Masataka Goto, Kazuyoshi Yoshii, Tomoyasu Nakano |
ISM | 3 |
| 2015 | Musical Similarity and Commonness Estimation Based on Probabilistic Generative ModelsabstractThis paper proposes a novel concept we call musical commonness, which is the similarity of a song to a set of songs, in other words, its typicality. This commonness can be used to retrieve representative songs from a song set (e.g., songs released in the 80s or 90s). Previous research on musical similarity has compared two songs but has not evaluated the similarity of a song to a set of songs. The methods presented here for estimating the similarity and commonness of polyphonic musical audio signals are based on a unified framework of probabilistic generative modeling of four musical elements (vocal timbre, musical timbre, rhythm, and chord progression). To estimate the commonness, we use a generative model trained from a song set instead of estimating musical similarities of all possible song-pairs by using a model trained from each song. In experimental evaluation, we used 3278 popular music songs. Estimated song-pair similarities are comparable to ratings by a musician at the 0.1% significance level for vocal and musical timbre, at the 1% level for rhythm, and the 5% level for chord progression. Results of commonness evaluation show that the higher the musical commonness is, the more similar a song is to songs of a song set. Tomoyasu Nakano, Kazuyoshi Yoshii, Masataka Goto |
ISM | 1 |
| 2014 | Regression approaches to perceptual age control in singing voice conversionabstractThe perceptual age of a singing voice is the age of the singer as perceived by the listener, and is one of the notable characteristics that determines perceptions of a song. In this paper, we describe a novel voice timbre control technique based on the perceptual age for singing voice conversion (SVC). Singers can sing expressively by controlling prosody and voice timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome the limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice timbre of an arbitrary source singer into that of an arbitrary target singer. However, it is still difficult to intuitively control singing voice characteristics by manipulating parameters corresponding to specific physical traits, such as gender and age. In this paper, we develop a technique for controlling the voice timbre based on perceptual age that maintains the singer's individuality. The experimental results show that the proposed voice timbre control method makes it possible to change the singer's perceptual age while not having an adverse effect on the perceived individuality. Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2014 | Vocal timbre analysis using latent Dirichlet allocation and cross-gender vocal timbre similarityabstractThis paper presents a vocal timbre analysis method based on topic modeling using latent Dirichlet allocation (LDA). Although many works have focused on analyzing characteristics of singing voices, none have dealt with “latent” characteristics (topics) of vocal timbre, which are shared by multiple singing voices. In the work described in this paper, we first automatically extracted vocal timbre features from polyphonic musical audio signals including vocal sounds. The extracted features were used as observed data, and mixing weights of multiple topics were estimated by LDA. Finally, the semantics of each topic were visualized by using a word-cloud-based approach. Experimental results for a singer identification task using 36 songs sung by 12 singers showed that our method achieved a mean reciprocal rank of 0.86. We also proposed a method for estimating cross-gender vocal timbre similarity by generating pitch-shifted (frequency-warped) signals of every singing voice. Experimental results for a cross-gender singer retrieval task showed that our method discovered interesting similar pitch-shifted singers. Tomoyasu Nakano, Kazuyoshi Yoshii, Masataka Goto |
ICASSP | 1 |
| 2014 | Cultivating vocal activity detection for music audio signals in a circulation-type crowdsourcing ecosystemabstractThis paper presents a crowdsourcing-based self-improvement framework of vocal activity detection (VAD) for music audio signals. A standard approach to VAD is to train a vocal-and-non-vocal classifier by using labeled audio signals (training set) and then use that classifier to label unseen signals. Using this technique, we have developed an online music-listening service called Songle that can help users better understand music by visualizing automatically estimated vocal regions and pitches of arbitrary songs existing on the Web. The accuracy of VAD is limited, however, because in general the acoustic characteristics of the training set are different from those of real songs on the Web. To overcome this limitation, we adapt a classifier by leveraging vocal regions and pitches corrected by volunteer users. UnlikeWikipedia-type crowdsourcing, our Songle-based framework can amplify user contributions: error corrections made for a limited number of songs improve VAD for all songs. This gives better music listening experiences to all users as non-monetary rewards. Kazuyoshi Yoshii, Hiromasa Fujihara, Tomoyasu Nakano, Masataka Goto |
ICASSP | 3 |
| 2013 | Evaluation of a singing voice conversion method based on many-to-many eigenvoice conversionabstractIn this paper, we evaluate our proposed singing voice conver-sion method from various perspectives. To enable singers to freely control their voice timbre of singing voice, we have pro-posed a singing voice conversion method based on many-to-many eigenvoice conversion (EVC) that enables to convert the voice timbre of an arbitrary source singer into that of another arbitrary target singer using a probabilistic model. Further-more, to easily develop training data consisting of multiple par-allel data sets between a single reference singer and many other singers, a technique for efficiently and effectively generating the parallel data sets from nonparallel singing voice data sets of many singers using a singing-to-singing synthesis system have been proposed. However, we have never conducted sufficient investigations into the effectiveness of these proposed methods. In this paper, we conduct both objective and subjective eval-uations to carefully investigate the effectiveness of proposed methods. Moreover, the differences between singing voice con-version and speaking voice conversion are also analyzed. Ex-perimental results show that our proposed method succeeds in enabling people to control their own voice timbre by using only an extremely small amount of the target singing voice. Index Terms: singing voice, voice conversion, eigenvoice con-version, singing-to-singing synthesis, performance evaluation Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2013 | An investigation of acoustic features for singing voice conversion based on perceptual ageabstractIn this paper, we investigate the acoustic features that can be modified to control the perceptual age of a singing voice. Singers can sing expressively by controlling prosody and vocal timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome this limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, it is still difficult to intu-itively control singing voice characteristics by manipulating pa-rameters corresponding to specific physical traits, such as gen-der and age. In this paper, we focus on controlling the perceived age of the singer and, as a first step, perform an investigation of the factors that play a part in the listener’s perception of the singer’s age. The experimental results demonstrate that 1) the perceptual age of singing voices corresponds relatively well to the actual age of the singer, 2) speech analysis/synthesis pro-cessing and statistical voice conversion processing don’t cause adverse effects on the perceptual age of singing voices, and 3) prosodic features have a larger effect on the perceptual age than spectral features. Kazuhiro Kobayashi, Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2012 | VocaListener and VocaWatcher: Imitating a human singer by using signal processingabstractIn this paper, we describe three singing information processing systems, VocaListener, VocaListener2, and VocaWatcher, that imitate singing expressions of the voice and face of a human singer. VocaListener can synthesize natural singing voices by analyzing and imitating the pitch and dynamics of the human singing. VocaListener2 imitates temporal timbre changes in addition to the pitch and dynamics. In synchronization with the synthesized singing voices, VocaWatcher can generate realistic facial motions of a humanoid robot, the HRP-4C, by analyzing and imitating facial motions of a human singing that are recorded by a single video camera. These systems that focus on “imitation” are not only promising for representing human-like naturalness, but also useful for providing intuitive control means. Masataka Goto, Tomoyasu Nakano, Shuuji Kajita, Yosuke Matsusaka, Shinichiro Nakaoka, Kazuhito Yokoi |
ICASSP | 2 |
| 2011 | Vocalistener2: A singing synthesis system able to mimic a user's singing in terms of voice timbre changes as well as pitch and dynamicsabstractThis paper presents a singing synthesis system, VocaListener2, that can automatically synthesize a singing voice by mimicking the timbre changes of a user's singing voice. The system is an extension of our previous VocaListener system which deals with only pitch and dynamics. Most previous techniques for manipulating voice timbre have focused on voice conversion and voice morphing, and they cannot deal with the timbre changes during singing. To develop VocaListener2, we constructed a voice timbre space on the basis of various singing voices that are synchronized under pitch, dynamics, and phoneme by using VocaListener. In this space, the timbre changes can be reflected in the synthesized singing voice. The system was evaluated by the Euclidean distance in the space between an estimated result and a ground-truth under closed/open conditions. Tomoyasu Nakano, Masataka Goto |
ICASSP | 1 |
| 2011 | VocaWatcher: Natural singing motion generator for a humanoid robotabstractIn this paper, we describe VocaWatcher, a novel robot motion generator that enables a humanoid robot to sing with realistic facial expressions and naturally synthesized singing voices. This robot singer is an important and attractive humanoid robot application for the entertainment scene; moreover, it promotes state-of-the-art integration of robot engineering, music processing, and image processing. To overcome the difficulties of generating natural facial expressions that are precisely synchronized with singing voices, VocaWatcher imitates a human singer by analyzing a video clip of a human singing, recorded by a single video camera. VocaWatcher can control mouth, eye, and neck motions by imitating the corresponding human movements, which are estimated without using any markers in the video. It can also synthesize singing voices by imitating the pitch and dynamics of the human singing in the same video. Shuuji Kajita, Tomoyasu Nakano, Masataka Goto, Yosuke Matsusaka, Shinichiro Nakaoka, Kazuhito Yokoi |
IROS | 2 |
| 2010 | Singing information processing based on singing voice modelingabstractIn this paper, we propose a novel area of research referred to as singing information processing. To shape the concept of this area, we first introduce singing understanding systems for synchronizing between vocal melody and corresponding lyrics, identifying the singer name, evaluating singing skills, creating hyperlinks between phrases in the lyrics of songs, and detecting breath sounds. We then introduce music information retrieval systems based on similarity of vocal melody timbre and vocal percussion, and singing synthesis systems. Common signal processing techniques for modeling singing voices that are used in these systems, such as techniques for extracting the vocal melody from polyphonic music recordings and modeling the lyrics by using phoneme HMMs for singing voices, are discussed. Masataka Goto, Takeshi Saitou, Tomoyasu Nakano, Hiromasa Fujihara |
ICASSP | 3 |
| 2006 | An automatic singing skill evaluation method for unknown melodies using pitch interval accuracy and vibrato featuresabstractThis paper presents a method of evaluating singing skills that does not require score information of the sung melody. This requires an approach that is different from existing systems, such as those currently used for Karaoke systems. Previous research on singing evaluation has focused on analyzing the characteristics of singing voice, but were not aimed at developing an automatic evaluation method. The approach presented in this study uses pitch interval accuracy and vibrato as acoustic features which are independent from specific characteristics of the singer or melody. The approach was tested by a 2-class (good/poor) classification test with 600 song sequences, and achieved an average classification rate of Tomoyasu Nakano, Masataka Goto, Yuzuru Hiraga |
INTERSPEECH | 1 |