EDBT 2026 Demo / reviewers in the wild / expert
Yuki Saito 0001
dblp:36/7818-1
· DBLP profile ↗
47ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-7967-2613ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 7 first-author · 26 since 2021Artificial intelligence and machine learning · 31 · 8 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki |
LREC | 2 |
| 2026 | J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language ModelingabstractSpoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications. Wataru Nakata, Kentaro Seki, Hitomi Yanaka, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
LREC | 4 |
| 2026 | DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue AudioabstractFull-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference. Wataru Nakata, Yuki Saito 0001, Kazuki Yamauchi, Emiru Tsunoo, Hiroshi Saruwatari |
SIGDIAL | 2 |
| 2026 | Speaker-conditioned phrase break prediction for text-to-speech with phoneme-level pre-trained language modelabstractThis paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We further demonstrate that these speaker embeddings can capture speaker-related characteristics solely from the phrasing task. Besides, we explore the potential of pre-trained speaker embeddings for unseen speakers through a few-shot adaptation method. Furthermore, we pioneer the application of phoneme-level pre-trained language models to this TTS front-end task, which significantly boosts the accuracy of the phrasing model. Our methods are rigorously assessed through both objective and subjective evaluations, demonstrating their effectiveness. • Speaker-conditioned phrasing model improves accuracy in multi-speaker phrasing tasks. • We explore various speaker embeddings in phrasing models. • We apply phoneme-level pre-trained language models to enhance phrasing accuracy. • We propose a speaker adaptation method for few-shot phrasing tasks. • We verify that speaker embeddings learn human-aligned features via phrasing tasks. Yuki Saito 0001, Takaaki Saeki, Tomoki Koriyama, Wataru Nakata, Detai Xin, Hiroshi Saruwatari |
Speech Commun. | 2 |
| 2025 | CAVIARES: Corpus for Audio-Visual Expressive Voice AgentabstractHigh-quality audio-visual corpora are essential for building voice agents capable of natural human-machine communication, but existing corpora commonly contain a limited amount of data per speaker, making personalized modeling difficult. We present CAVIARES, a new audio-visual corpus comprising 9.5 hours of expressive speech recorded by a single professional Japanese female speaker. CAVIARES consists of two subsets: acted dialogue and expressive reading, providing a diverse range of speaking styles for speech-to-facial motion modeling and multimodal learning tasks. In this paper, we describe the construction process of CAVIARES and the results of corpus analysis. CAVIARES will be released for research purposes only. Jinsheng Chen, Yuki Saito 0001, Naoko Tanji, Hironori Doi, Byeongseon Park, Yuma Shirahata, Kentaro Tachibana, Hiroshi Saruwatari |
ASRU | 2 |
| 2025 | Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent LayerabstractWe introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model. Go Nishikawa, Wataru Nakata, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari, Tomohiko Nakamura |
ASRU | 3 |
| 2025 | Analysing the Language of Neural Audio CodecsabstractThis study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to linguistic statistical laws such as Zipf’s law and Heaps’ law, as well as their entropy and redundancy. To assess how these token-level properties relate to semantic and acoustic preservation in synthesized speech, we evaluate intelligibility using error rates of automatic speech recognition, and quality using the UTMOS score. Our results reveal that NAC tokens, particularly 3-grams, exhibit language-like statistical patterns. Moreover, these properties, together with measures of information content, are found to correlate with improved performances in speech recognition and resynthesis tasks. These findings offer insights into the structure of NAC token sequences and inform the design of more effective generative speech models. Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito 0001, Hiroshi Saruwatari |
ASRU | 5 |
| 2025 | Measuring Time Delay Tolerance in Third-Person Live Commentary for Super Smash Bros. UltimateabstractThis study proposes a methodology for measuring the acceptable delay tolerance for third-person game commentary. Third-person game commentary refers to commentary delivered by someone other than the player, with the role of helping viewers better understand the game and enhancing the viewing experience. With the recent advancement of AI, there has been increasing interest in automating such commentary using video understanding and audio generation. However, automating this process using video understanding and audio generation introduces delays, potentially affecting the naturalness of the commentary. In this context, since the extent to which such delays are acceptable to viewers remains unclear, we address this issue. The tolerance is modeled using an unnormalized Gaussian function. Through experiments on Super Smash Bros. Ultimate with 727 participants, we found that the average acceptable delay for this game is 3.71 seconds, with variations depending on different viewer attributes and gameplay contexts. Ryosuke Matsushita, Ryosuke Sakai, Koki Fukuda, Shinnosuke Takamichi, Kota Iura, Yuki Saito 0001, Graham Neubig, Katsuhito Sudoh, Hiroya Takamura, Tatsuya Ishigaki |
CoG | 6 |
| 2025 | Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning FeaturesabstractReal-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features of self-supervised learning (SSL) models. This work is the first to incorporate SSL features with causality into an SE model. The causal SSL features are encoded and combined with spectrogram features using feature-wise linear modulation to estimate a mask for enhancing the noisy input speech. Simultaneously, we quantize the causal SSL features using vector quantization to represent phonetic characteristics as semantic tokens. The model not only encodes SSL features but also predicts the future semantic tokens in multi-task learning (MTL). The experimental results using VoiceBank + DEMAND dataset show that our proposed method achieves 2.88 in PESQ, especially with semantic prediction MTL, in which we confirm that the semantic prediction played an important role in causal SE. Emiru Tsunoo, Yuki Saito 0001, Wataru Nakata, Hiroshi Saruwatari |
ICASSP | 2 |
| 2025 | RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Yusuke Kanamori, Yuki Okamoto, Taisei Takano, Shinnosuke Takamichi, Yuki Saito 0001, Hiroshi Saruwatari |
INTERSPEECH | 5 |
| 2025 | Shallow Flow Matching for Coarse-to-Fine Text-to-Speech SynthesisabstractWe propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations from the weak generator as conditions, SFM constructs intermediate states along the FM paths from these representations. During training, we introduce an orthogonal projection method to adaptively determine the temporal position of these states, and apply a principled construction strategy based on a single-segment piecewise flow. The SFM inference starts from the intermediate state rather than pure noise, thereby focusing computation on the latter stages of the FM paths. We integrate SFM into multiple TTS models with a lightweight SFM head. Experiments demonstrate that SFM yields consistent gains in speech naturalness across both objective and subjective evaluations, and significantly accelerates inference when using adaptive-step ODE solvers. Demo and codes are available at https://ydqmkkx.github.io/SFMDemo/. Yiyi Cai, Yuki Saito 0001, Lixu Wang, Hiroshi Saruwatari |
NeurIPS | 3 |
| 2024 | STYLECAP: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-Supervised Learning ModelsabstractWe propose StyleCap, a method to generate natural language descriptions of speaking styles appearing in speech. Although most of conventional techniques for para-/non-linguistic information recognition focus on the category classification or the intensity estimation of pre-defined labels, they cannot provide the reasoning of the recognition result in an interpretable manner. StyleCap is a first step towards an end-to-end method for generating speaking-style prompts from speech, i.e., automatic speaking-style captioning. StyleCap is trained with paired data of speech and natural language descriptions. We train neural networks that convert a speech representation vector into prefix vectors that are fed into a large language model (LLM)-based text decoder. We explore an appropriate text decoder and speech feature representation suitable for this new task. The experimental results demonstrate that our StyleCap leveraging richer LLMs for the text decoder, speech self-supervised learning (SSL) features, and sentence rephrasing augmentation improves the accuracy and diversity of generated speaking-style captions. Samples of speaking-style captions generated by our StyleCap are publicly available1. Kazuki Yamauchi, Yusuke Ijima, Yuki Saito 0001 |
ICASSP | 3 |
| 2024 | Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
Takuto Igarashi, Yuki Saito 0001, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2024 | SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
Yuki Saito 0001, Takuto Igarashi, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 1 |
| 2024 | Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
Kentaro Seki, Shinnosuke Takamichi, Norihiro Takamune, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2024 | Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
Tomoki Koriyama, Yuki Saito 0001 |
INTERSPEECH | 3 |
| 2024 | The T05 System for the voicemos challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic SpeechabstractWe present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-quality synthetic speech. In addition to a pretrained self-supervised learning (SSL)-based speech feature extractor, our system incorporates a pretrained image feature extractor to capture the difference of synthetic speech observed in speech spectrograms. We first separately train two MOS predictors that use either of an SSL-based or spectrogram-based feature. Then, we fine-tune the two predictors for better MOS prediction using the fusion of two extracted features. In the VMC 2024 Track 1, our T05 system achieved first place in 7 out of 16 evaluation metrics and second place in the remaining 9 metrics, with a significant difference compared to those ranked third and below. We also report the results of our ablation study to investigate essential factors of our system. Kaito Baba, Wataru Nakata, Yuki Saito 0001, Hiroshi Saruwatari |
SLT | 3 |
| 2024 | Cross-Dialect Text-to-Speech In Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BertabstractWe explore cross-dialect text-to-speech(CD-TTS),a task to synthesize learned speakers’voices in non-native dialects,especially in pitch-accent languages.CD-TTS is important for developing voice agents that naturally communicate with people across regions.We present a novel TTS model comprising three sub-modules to perform competitively at this task.We first train a backbone TTS model to synthesize dialect speech from a text conditioned on phoneme-level accent latent variables(ALVs)extracted from speech by a reference encoder. Then,we train an ALV predictor to predict ALVs tailored to a target dialect from input text leveraging our novel multi-dialect phoneme-level BERT.We conduct multi-dialect TTS experiments and evaluate the effectiveness of our model by comparing it with a baseline derived from conventional dialect TTS methods.The results show that our model improves the dialectal naturalness of synthetic speech in CD-TTS. Kazuki Yamauchi, Yuki Saito 0001, Hiroshi Saruwatari |
SLT | 2 |
| 2023 | COCO-NUT: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-Based ControlabstractIn text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as text-to-image, free-form text instruction should be useful for intuitive and complicated control of voice characteristics. A sufficiently large corpus of high-quality and diverse voice samples with corresponding free-form descriptions can advance such control research. However, neither an open corpus nor a scalable method is currently available. To this end, we develop Coco-Nut, a new corpus including diverse Japanese utterances, along with text transcriptions and free-form voice characteristics descriptions. Our methodology to construct this corpus consists of 1) automatic collection of voice-related audio data from the Internet, 2) quality assurance, and 3) manual annotation using crowdsourcing. Additionally, we benchmark our corpus on the prompt embedding model trained by contrastive speech-text learning. Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Wataru Nakata, Detai Xin, Hiroshi Saruwatari |
ASRU | 3 |
| 2023 | MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture ModelsabstractIn this paper, we propose a method for intermediating multiple speakers’ attributes and diversifying their voice characteristics in “speaker generation,” an emerging task that aims to synthesize a nonexistent speaker’s naturally sounding voice. The conventional TacoSpawn-based speaker generation method represents the distributions of speaker embeddings by Gaussian mixture models (GMMs) conditioned with speaker attributes. Although this method enables the sampling of various speakers from the speaker-attribute-aware GMMs, it is not yet clear whether the learned distributions can represent speakers with an intermediate attribute (i.e., mid-attribute). To this end, we propose an optimal-transport-based method that interpolates the learned GMMs to generate nonexistent speakers with mid-attribute (e.g., gender-neutral) voices. We empirically validate our method and evaluate the naturalness of synthetic speech and the controllability of two speaker attributes: gender and language fluency. The evaluation results show that our method can control the generated speakers’ attributes by a continuous scalar value without statistically significant degradation of speech naturalness. Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Detai Xin, Hiroshi Saruwatari |
ICASSP | 3 |
| 2023 | Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-SpeechabstractPause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styles of inserting silent pauses, which can degrade the performance of the model trained on a multi-speaker speech corpus. To this end, we propose more powerful pause insertion frameworks based on a pre-trained language model. Our approach uses bidirectional encoder representations from transformers (BERT) pre-trained on a large-scale text corpus, injecting speaker embeddings to capture various speaker characteristics. We also leverage duration-aware pause insertion for more natural multi-speaker TTS. We develop and evaluate two types of models. The first improves conventional phrasing models on the position prediction of respiratory pauses (RPs), i.e., silent pauses at word transitions without punctuation. It performs speaker-conditioned RP prediction considering contextual information and is used to demonstrate the effect of speaker information on the prediction. The second model is further designed for phoneme-based TTS models and performs duration-aware pause insertion, predicting both RPs and punctuation-indicated pauses (PIPs) that are categorized by duration. The evaluation results show that our models improve the precision and recall of pause insertion and the rhythm of synthetic speech. Tomoki Koriyama, Yuki Saito 0001, Takaaki Saeki, Detai Xin, Hiroshi Saruwatari |
ICASSP | 3 |
| 2023 | CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
Yuki Saito 0001, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 1 |
| 2023 | ChatGPT-EDSS: Empathetic Dialogue Speech Synthesis Trained from ChatGPT-derived Context Word Embeddings
Yuki Saito 0001, Shinnosuke Takamichi, Eiji Iimori, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 1 |
| 2023 | HumanDiffusion: diffusion model using perceptual gradients
Yota Ueda, Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Hiroshi Saruwatari |
INTERSPEECH | 3 |
| 2022 | Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2022 | Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue HistoryabstractWe propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history.Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems.Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context.As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling.To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterancewise modeling.The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method. Yuto Nishimura, Yuki Saito 0001, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2022 | STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice AgentabstractWe present STUDIES, a new speech corpus for developing a voice agent that can speak in a friendly manner.Humans naturally control their speech prosody to empathize with each other.By incorporating this "empathetic dialogue" behavior into a spoken dialogue system, we can develop a voice agent that can respond to a user more naturally.We designed the STUDIES corpus to include a speaker who speaks with empathy for the interlocutor's emotion explicitly.We describe our methodology to construct an empathetic dialogue speech corpus and report the analysis results of the STUDIES corpus.We conducted a text-to-speech experiment to initially investigate how we can develop more natural voice agent that can tune its speaking style corresponding to the interlocutor's emotion.The results show that the use of interlocutor's emotion label and conversational context embedding can produce speech with the same degree of naturalness as that synthesized by using the agent's emotion label. Yuki Saito 0001, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 1 |
| 2022 | Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTSabstractThis paper proposes a human-in-the-loop speaker-adaptation method for multi-speaker text-to-speech.With a conventional speaker-adaptation method, a target speaker's embedding vector is extracted from his/her reference speech using a speaker encoder trained on a speaker-discriminative task.However, this method cannot obtain an embedding vector for the target speaker when the reference speech is unavailable.Our method is based on a human-in-the-loop optimization framework, which incorporates a user to explore the speakerembedding space to find the target speaker's embedding.The proposed method uses a sequential line search algorithm that repeatedly asks a user to select a point on a line segment in the embedding space.To efficiently choose the best speech sample from multiple stimuli, we also developed a system in which a user can switch between multiple speakers' voices for each phoneme while looping an utterance.Experimental results indicate that the proposed method can achieve comparable performance to the conventional one in objective and subjective evaluations even if reference speech is not used as the input of a speaker encoder directly. Kenta Udagawa, Yuki Saito 0001, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2021 | Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme PerceptionabstractWe propose a conditional generative adversarial network (GAN) incorporating humans’ perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a real-data distribution accurately but can never represent a human-acceptable distribution, which are ranges of data in which humans accept the naturalness regardless of whether the data are real or not. A Human-GAN was proposed to model the human-acceptable distribution. A DNN-based generator is trained using a human-based discriminator, i.e., humans’ perceptual evaluations, instead of the GAN’s DNN-based discriminator. However, the HumanGAN cannot represent conditional distributions. This paper proposes the HumanACGAN, a theoretical extension of the HumanGAN, to deal with conditional human-acceptable distributions. Our HumanACGAN trains a DNN-based conditional generator by regarding humans as not only a discriminator but also an auxiliary classifier. The generator is trained by deceiving the human-based discriminator that scores the unconditioned naturalness and the human-based classifier that scores the class-conditioned perceptual acceptability. The training can be executed using the backpropagation algorithm involving humans’ perceptual evaluations. Our experimental results in phoneme perception demonstrate that our HumanACGAN can successfully train this conditional generator. Yota Ueda, Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari |
ICASSP | 3 |
| 2021 | Cross-Lingual Speaker Adaptation Using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
Interspeech | 2 |
| 2021 | Perceptual-Similarity-Aware Deep Speaker Representation Learning for Multi-Speaker Generative ModelingabstractWe propose novel deep speaker representation learning that considers perceptual similarity among speakers for multi-speaker generative modeling. Following its success in accurate discriminative modeling of speaker individuality, knowledge of deep speaker representation learning (i.e., speaker representation learning using deep neural networks) has been introduced to multi-speaker generative modeling. However, the conventional discriminative algorithm does not necessarily learn speaker embeddings suitable for such generative modeling, which may result in lower quality and less controllability of synthetic speech. We propose three representation learning algorithms that utilize a perceptual speaker similarity matrix obtained by large-scale perceptual scoring of speaker-pair similarity. The algorithms train a speaker encoder to learn speaker embeddings with three different representations of the matrix: a set of vectors, the Gram matrix, and a graph. Furthermore, we propose an active learning algorithm that iterates the perceptual scoring and speaker encoder training. To obtain accurate embeddings while reducing costs of scoring and training, the algorithm selects unscored speaker-pairs to be scored next on the basis of the sequentially-trained speaker encoder's similarity prediction results. Experimental evaluation results show that 1) the proposed representation learning algorithms learn speaker embeddings strongly correlated with perceptual speaker-pair similarity, 2) the embeddings improve synthetic speech quality in speech autoencoding tasks better than conventional d-vectors learned by discriminative modeling, 3) the proposed active learning algorithm achieves higher synthetic speech quality while reducing costs of scoring and training, and 4) among the proposed similarity {vector, matrix, graph} embedding algorithms, the first achieves the best speaker similarity for synthetic speech and the third gives the most improvement in the synthetic speech naturalness. Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception ModelingabstractWe propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribution by fooling the discriminator that distinguishes real and generated data. Therefore, the basic GAN cannot represent the outside of a real-data distribution. In the case of speech perception, humans can recognize not only human voices but also processed (i.e., a non-existent human) voices as human voice. Such a human-acceptable distribution is typically wider than a real-data one and cannot be modeled by the basic GAN. To model the human-acceptable distribution, we formulate a backpropagation-based generator training algorithm by regarding human perception as a black-boxed discriminator. The training efficiently iterates generator training by using a computer and discrimination by human. We evaluate our HumanGAN in speech naturalness modeling and demonstrate that it can represent a human-acceptable distribution that is wider than a real-data distribution. Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari |
ICASSP | 2 |
| 2020 | Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral DifferentialsabstractIn this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials. The conventional method with a minimum-phase filter achieves high-quality conversion but requires heavy computation in filtering. This is because the minimum phase using a fixed lifter of the Hilbert transform often results in a long-tap filter. One of our methods is a data-driven method for lifter training. Since this method takes filter truncation into account in training, it can shorten the tap length of the filter while preserving conversion accuracy. Our other method is sub-band processing for extending the conventional method from narrow-band (16 kHz) to full-band (48 kHz) VC, which can convert a full-band waveform with higher converted-speech quality. Experimental results indicate that 1) the proposed lifter-training method for narrow-band VC can shorten the tap length to 1/16 without degrading the converted-speech quality and 2) the proposed sub-band-processing method for full-band VC can improve the converted-speech quality than the conventional method. Takaaki Saeki, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
ICASSP | 2 |
| 2020 | Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image
Shunsuke Goto, Kotaro Onishi, Yuki Saito 0001, Kentaro Tachibana, Koichiro Mori |
INTERSPEECH | 3 |
| 2020 | Real-Time, Full-Band, Online DNN-Based Voice Conversion System Using a Single CPU
Takaaki Saeki, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2020 | Cross-Lingual Text-To-Speech Synthesis via Domain Adaptation and Perceptual Similarity Regression in Speaker Space
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2020 | Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 3 |
| 2020 | SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on GameplayabstractDeveloping a spontaneous speech corpus would be beneficial for spoken language processing and understanding. We present a speech corpus named the SMASH corpus, which includes spontaneous speech of two Japanese male commentators that made third-person audio commentaries during the gameplay of a fighting game. Each commentator ad-libbed while watching the gameplay with various topics covering not only explanations of each moment to convey the information on the fight but also comments to entertain listeners. We made transcriptions and topic tags as annotations on the recorded commentaries with our two-step method. We first made automatic and manual transcriptions of the commentaries and then manually annotated the topic tags. This paper describes how we constructed the SMASH corpus and reports some results of the annotations. Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
LREC | 1 |
| 2020 | DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech CorpusabstractIn this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features called context instead of directly using text. In such frameworks, we can synthesize not only reading-style speech but also speech with paralinguistic and nonlinguistic features by adding such information to the context. However, it is not clear what kind of information is crucial for reproducing paralinguistic and nonlinguistic features. Therefore, we investigate the effectiveness of rich tags in DNN-based speech synthesis according to the Corpus of Spontaneous Japanese (CSJ), which has a large amount of annotations on paralinguistic features such as prosody, disfluency, and morphological features. Experimental evaluation results shows that the reproducibility of paralinguistic features of synthetic speech was enhanced by adding such information as context. Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
LREC | 3 |
| 2020 | Phase reconstruction from amplitude spectrograms based on directional-statistics deep neural networksabstractThis paper presents a deep neural network (DNN)-based phase reconstruction method from amplitude spectrograms. In speech processing, an amplitude spectrogram is often used for processing, and the corresponding phases are reconstructed from the amplitude spectrogram by using the Griffin-Lim method. However, the Griffin-Lim method causes unnatural artifacts in synthetic speech. To solve this problem, we propose the directional-statistics DNNs for predicting phases from the amplitude spectrograms. We first propose the von Mises distribution DNN, which is a generative model having the von Mises distribution and models histograms of a periodic variable. We extend it for modeling group delay that has a stronger connection to the amplitude spectrograms. Furthermore, we generalize the group-delay modeling and propose another DNN called the sine-skewed generalized cardioid distribution DNN for modeling asymmetric histograms such as a group delay. Results from objective and subjective evaluations indicate that (1) our von Mises distribution DNN can predict group delay more accurately than predicting phases, (2) our DNN works as better initialization of the Griffin-Lim method, (3) the phase reconstruction methods based on our von Mises distribution DNN achieve better speech quality than the conventional Griffin-Lim method, and (4) our sine-skewed generalized cardioid distribution DNN models the group delay more accurately than our von Mises distribution DNN. Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari |
Signal Process. | 2 |
| 2019 | Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-trackingabstractThis paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double-tracking, a recording method in which two performances of the same phrase are recorded and mixed to create a richer, layered sound. However, singing voices synthesized using conventional DNN-based methods never vary because the synthesis process is deterministic and only one waveform is synthesized from one musical score. To address this problem, we use a GMMN to model the variation of the modulation spectrum of the pitch contour of natural singing voices and add a randomized inter-utterance variation to the pitch contour generated by conventional DNN-based singing voice synthesis. Experimental evaluations suggest that 1) our approach can provide perceptible inter-utterance pitch variation while preserving speech quality. We extend our approach to double-tracking, and the evaluation demonstrates that 2) GMMN-based neural double-tracking is perceptually closer to natural double-tracking than conventional signal processing-based artificial double-tracking is. Hiroki Tamaru, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
ICASSP | 2 |
| 2019 | Vocoder-free text-to-speech synthesis incorporating generative adversarial networks using low-/multi-frequency STFT amplitude spectraabstractThis paper proposes novel training algorithms for vocoder-free text-to-speech (TTS) synthesis based on generative adversarial networks (GANs) that compensate for short-term Fourier transform (STFT) amplitude spectra in low/multi frequency resolution. Vocoder-free TTS using STFT amplitude spectra can avoid degradation of synthetic speech quality caused by the vocoder-based parameterization used in conventional TTS. Our previous work for the vocoder-based TTS proposed a method for incorporating the GAN-based distribution compensation into acoustic model training to improve synthetic speech quality. This paper extends the algorithm to the vocoder-free TTS and propose a GAN-based training algorithm using low-frequency-resolution amplitude spectra to overcome the difficulty in modeling complicated distribution of the high-dimensional spectra. In the proposed algorithm, amplitude spectra are transformed into low-frequency-resolution amplitude spectra by applying an average pooling function along with a frequency axis; then the GAN-based distribution compensation is performed in the low-frequency-resolution domain. Because the low-frequency-resolution amplitude spectra approximately emulate filter banks, the proposed algorithm is expected to improve synthetic speech quality by reducing differences in spectral envelopes of natural and synthetic speech. Furthermore, various frequency scales that are related to human speech perception (e.g., mel and inverse mel frequency scales) can be introduced to the proposed training algorithm by applying an frequency warping function to amplitude spectra. This paper also proposes a GAN-based training algorithm using multi-frequency-resolution amplitude spectra that uses both low- and original-frequency-resolution amplitude spectra to reduce the differences in not only spectral envelopes but also fine structures. Experimental results demonstrate that (1) GANs using low-frequency-resolution amplitude spectra improve speech quality and work robustly against the settings of the frequency resolution and hyperparameters, (2) in comparison among low-, original-, and multi-frequency-resolution amplitude spectra, the use of low-frequency-resolution ones work best improve the synthetic speech quality, and (3) the use of the inverse mel frequency scale for obtaining low-frequency-resolution amplitude spectra further improves synthetic speech quality. Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
Comput. Speech Lang. | 1 |
| 2018 | Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-VectorsabstractThis paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish because of an over-regularization issue often observed in latent variables of the VAEs. To overcome the issue, this paper proposes a VAE-based non-parallel VC conditioned by not only the speaker representations but also phonetic contents of speech represented as phonetic posteriorgrams (PPGs). Since the phonetic contents are given during the training, we can expect that the VC models effectively learn speaker-independent latent features of speech. Focusing on the point, this paper also extends the conventional VAE-based non-parallel VC to many-to-many VC that can convert arbitrary speakers' characteristics into another arbitrary speakers' ones. We investigate two methods to estimate speaker representations for speakers not included in speech corpora used for training VC models: 1) adapting conventional speaker codes, and 2) using d-vectors for the speaker representations. Experimental results demonstrate that 1) PPGs successfully improve both naturalness and speaker similarity of the converted speech, and 2) both speaker codes and d-vectors can be adopted to the VAE-based many-to-many non-parallel VC. Yuki Saito 0001, Yusuke Ijima, Kyosuke Nishida, Shinnosuke Takamichi |
ICASSP | 1 |
| 2018 | Text-to-Speech Synthesis Using STFT Spectra Based on Low-/Multi-Resolution Generative Adversarial NetworksabstractThis paper proposes novel training algorithms for vocoder-free statistical parametric speech synthesis (SPSS) using short-term Fourier transform (STFT) spectra. Recently, text-to-speech synthesis using STFT spectra has been investigated since it can avoid quality degradation caused by the vocoder-based parameterization in conventional SPSS using a vocoder. In conventional SPSS using a vocoder, we previously proposed a training algorithm for integrating generative adversarial network (GAN)-based distribution compensation. To extend the algorithm to vocoder-free SPSS, we propose low- and multi-resolution GAN-based training algorithms for vocoder-free SPSS. In our algorithm that uses the low-resolution GAN, acoustic models are trained to minimize the weighted sum of the mean squared error between natural and generated spectra in the original resolution and adversarial loss to deceive discriminative models in the lower resolution. Since the low-resolution spectra are close to filter banks and their distribution becomes simpler, GAN-based distribution compensation works well. Furthermore, we propose an algorithm using multi-resolution GANs, which uses both the low-resolution GAN and original-resolution GAN. Experimental results demonstrate that 1) the low-resolution GAN works robustly to the setting of its frequency resolution and hyperparameter, and 2) compared the low-, original-, and multi-resolution GANs, the low-resolution GAN works the best to improve synthetic speech quality. Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
ICASSP | 1 |
| 2018 | Statistical Parametric Speech Synthesis Incorporating Generative Adversarial NetworksabstractA method for statistical parametric speech synthesis incorporating generative adversarial networks (GANs) is proposed. Although powerful deep neural networks techniques can be applied to artificially synthesize speech waveform, the synthetic speech quality is low compared with that of natural speech. One of the issues causing the quality degradation is an oversmoothing effect often observed in the generated speech parameters. A GAN introduced in this paper consists of two neural networks: a discriminator to distinguish natural and generated samples, and a generator to deceive the discriminator. In the proposed framework incorporating the GANs, the discriminator is trained to distinguish natural and generated speech parameters, while the acoustic models are trained to minimize the weighted sum of the conventional minimum generation loss and an adversarial loss for deceiving the discriminator. Since the objective of the GANs is to minimize the divergence (i.e., distribution difference) between the natural and generated speech parameters, the proposed method effectively alleviates the oversmoothing effect on the generated speech parameters. We evaluated the effectiveness for text-to-speech and voice conversion, and found that the proposed method can generate more natural spectral parameters and F0than conventional minimum generation error training algorithm regardless of its hyperparameter settings. Furthermore, we investigated the effect of the divergence of various GANs, and found that a Wasserstein GAN minimizing the Earth-Mover's distance works the best in terms of improving the synthetic speech quality. Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesisabstractThis paper proposes a novel training algorithm for high-quality Deep Neural Network (DNN)-based speech synthesis. The parameters of synthetic speech tend to be over-smoothed, and this causes significant quality degradation in synthetic speech. The proposed algorithm takes into account an Anti-Spoofing Verification (ASV) as an additional constraint in the acoustic model training. The ASV is a discriminator trained to distinguish natural and synthetic speech. Since acoustic models for speech synthesis are trained so that the ASV recognizes the synthetic speech parameters as natural speech, the synthetic speech parameters are distributed in the same manner as natural speech parameters. Additionally, we find that the algorithm compensates not only the parameter distributions, but also the global variance and the correlations of synthetic speech parameters. The experimental results demonstrate that 1) the algorithm outperforms the conventional training algorithm in terms of speech quality, and 2) it is robust against the hyper-parameter settings. Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
ICASSP | 1 |
| 2017 | Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior ProbabilitiesabstractVoice conversion (VC) using sequence-to-sequence learning of context posterior probabilities is proposed.Conventional VC using shared context posterior probabilities predicts target speech parameters from the context posterior probabilities estimated from the source speech parameters.Although conventional VC can be built from non-parallel data, it is difficult to convert speaker individuality such as phonetic property and speaking rate contained in the posterior probabilities because the source posterior probabilities are directly used for predicting target speech parameters.In this work, we assume that the training data partly include parallel speech data and propose sequence-to-sequence learning between the source and target posterior probabilities.The conversion models perform non-linear and variable-length transformation from the source probability sequence to the target one.Further, we propose a joint training algorithm for the modules.In contrast to conventional VC, which separately trains the speech recognition that estimates posterior probabilities and the speech synthesis that predicts target speech parameters, our proposed method jointly trains these modules along with the proposed probability conversion modules.Experimental results demonstrate that our approach outperforms the conventional VC. Hiroyuki Miyoshi, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
INTERSPEECH | 2 |