Shinnosuke Takamichi

dblp:124/9221 · DBLP profile ↗
← Back
90ranked-venue papers
12as first author
55since 2021 · last 2026
0000-0003-0520-7847ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 10 first-author · 46 since 2021Artificial intelligence and machine learning · 61 · 7 first-author · 38 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Sign-to-Speech Prosody Transfer via Sign Reconstruction-Based GAN
Toranosuke Manabe, Yuto Shibata, Shinnosuke Takamichi, Yoshimitsu Aoki
ICPR (8)3
2026 Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki
LREC5
2026 J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
abstract
Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications.
Wataru Nakata, Kentaro Seki, Hitomi Yanaka, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
LREC5
2025 Analysing the Language of Neural Audio Codecs
abstract
This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to linguistic statistical laws such as Zipf’s law and Heaps’ law, as well as their entropy and redundancy. To assess how these token-level properties relate to semantic and acoustic preservation in synthesized speech, we evaluate intelligibility using error rates of automatic speech recognition, and quality using the UTMOS score. Our results reveal that NAC tokens, particularly 3-grams, exhibit language-like statistical patterns. Moreover, these properties, together with measures of information content, are found to correlate with improved performances in speech recognition and resynthesis tasks. These findings offer insights into the structure of NAC token sequences and inform the design of more effective generative speech models.
Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito 0001, Hiroshi Saruwatari
ASRU2
2025 Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification
abstract
The marmoset, a highly vocal primate, is a key model for studying social-communicative behavior. Unlike human speech, marmoset vocalizations are less structured, highly variable, and recorded in noisy, low-resource conditions. Learning marmoset communication requires joint call segmentation, classification, and caller identification-challenging domain tasks. Previous CNNs handle local patterns but struggle with long-range temporal structure. We applied Transformers using self-attention for global dependencies. However, Transformers show overfitting and instability on small, noisy annotated datasets. To address this, we pretrain Transformers with MAE-a self-supervised method reconstructing masked segments from hundreds of hours of unannotated marmoset recordings. The pretraining improved stability and generalization. Results show MAE-pretrained Transformers outperform CNNs, demonstrating modern self-supervised architectures effectively model low-resource non-human vocal communication.
Shinnosuke Takamichi, Sakriani Sakti, Satoshi Nakamura 0001
ASRU2
2025 Measuring Time Delay Tolerance in Third-Person Live Commentary for Super Smash Bros. Ultimate
abstract
This study proposes a methodology for measuring the acceptable delay tolerance for third-person game commentary. Third-person game commentary refers to commentary delivered by someone other than the player, with the role of helping viewers better understand the game and enhancing the viewing experience. With the recent advancement of AI, there has been increasing interest in automating such commentary using video understanding and audio generation. However, automating this process using video understanding and audio generation introduces delays, potentially affecting the naturalness of the commentary. In this context, since the extent to which such delays are acceptable to viewers remains unclear, we address this issue. The tolerance is modeled using an unnormalized Gaussian function. Through experiments on Super Smash Bros. Ultimate with 727 participants, we found that the average acceptable delay for this game is 3.71 seconds, with variations depending on different viewer attributes and gameplay contexts.
Ryosuke Matsushita, Ryosuke Sakai, Koki Fukuda, Shinnosuke Takamichi, Kota Iura, Yuki Saito 0001, Graham Neubig, Katsuhito Sudoh, Hiroya Takamura, Tatsuya Ishigaki
CoG4
2025 The Text-to-speech in the Wild (TITW) Database
Jee-Weon Jung, Wangyou Zhang, Soumi Maiti, Yihan Wu 0008, Xin Wang 0037, Yuta Matsunaga, Seyun Um, Jinchuan Tian, Hye-Jin Shim, Nicholas W. D. Evans, Joon Son Chung, Shinnosuke Takamichi, Shinji Watanabe 0001
INTERSPEECH13
2025 RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Yusuke Kanamori, Yuki Okamoto, Taisei Takano, Shinnosuke Takamichi, Yuki Saito 0001, Hiroshi Saruwatari
INTERSPEECH4
2025 Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
Shunsuke Kando, Yusuke Miyao, Shinnosuke Takamichi
INTERSPEECH3
2025 J-SPAW: Japanese speaker verification and spoofing attacks recorded in-the-wild dataset
Sayaka Shiota, Suzuka Horie, Kouta Kanno, Shinnosuke Takamichi
INTERSPEECH4
2025 Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
Hitoshi Suda, Shinnosuke Takamichi, Satoru Fukayama
INTERSPEECH2
2025 Analysis of the Correlation Between Theory of Mind and Dialogue Ability to Identify Essential ToM for Dialogue Systems
Haruhisa Iseno, Atsumoto Ohashi, Tetsuji Ogawa, Shinnosuke Takamichi, Ryuichiro Higashinaka
PACLIC4
2024 Environmental Sound Synthesis from Vocal Imitations and Sound Event Labels
abstract
One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as rhythm and pitch, using vocal imitations, which cannot be expressed by conventional input information, such as sound event labels, images, or texts, in an environmental sound synthesis model. In this paper, we propose a framework for environmental sound synthesis from vocal imitations and sound event labels based on a framework of a vector quantized encoder and the Tacotron2 decoder. Using vocal imitations is expected to control the pitch and rhythm of the synthesized sound, which only sound event labels cannot control. Our objective and subjective experimental results show that vocal imitations effectively control the pitch and rhythm of synthesized sounds.
Yuki Okamoto, Keisuke Imoto, Shinnosuke Takamichi, Ryotaro Nagase, Takahiro Fukumori, Yoichi Yamashita
ICASSP3
2024 Diversity-Based Core-Set Selection for Text-to-Speech with Linguistic and Acoustic Features
abstract
This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed for constructing large-scale TTS corpora by collecting diverse data from massive sources such as audiobooks and YouTube. Although these methods have gained significant attention for enhancing the expressive capabilities of TTS systems, they often prioritize collecting vast amounts of data without considering practical constraints like storage capacity and computation time in training, which limits the available data quantity. Consequently, the need arises to efficiently collect data within these volume constraints. To address this, we propose a method for selecting the core subset (known as core-set) from a TTS corpus on the basis of a diversity metric, which measures the degree to which a subset encompasses a wide range. Experimental results demonstrate that our proposed method performs significantly better than the baseline phoneme-balanced data selection across language and corpus size.
Kentaro Seki, Shinnosuke Takamichi, Takaaki Saeki, Hiroshi Saruwatari
ICASSP2
2024 Do Learned Speech Symbols Follow Zipf's Law?
abstract
In this study, we investigate whether speech symbols, learned through deep learning, follow Zipf’s law, akin to natural language symbols. Zipf’s law is an empirical law that delineates the frequency distribution of words, forming fundamentals for statistical analysis in natural language processing. Natural language symbols, which are invented by humans to symbolize speech content, are recognized to comply with this law. On the other hand, recent breakthroughs in spoken language processing have given rise to the development of learned speech symbols; these are data-driven symbolizations of speech content. Our objective is to ascertain whether these datadriven speech symbols follow Zipf’s law, as the same as natural language symbols. Through our investigation, we aim to forge new ways for the statistical analysis of spoken language processing.
Shinnosuke Takamichi, Hiroki Maeda, Joonyong Park, Daisuke Saito, Hiroshi Saruwatari
ICASSP1
2024 Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
Takuto Igarashi, Yuki Saito 0001, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH4
2024 Textless Dependency Parsing by Labeled Sequence Prediction
Shunsuke Kando, Yusuke Miyao, Jason Naradowsky, Shinnosuke Takamichi
INTERSPEECH4
2024 SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
Takaaki Saeki, Soumi Maiti, Shinnosuke Takamichi, Shinji Watanabe 0001, Hiroshi Saruwatari
INTERSPEECH3
2024 SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
Yuki Saito 0001, Takuto Igarashi, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH4
2024 Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
Kentaro Seki, Shinnosuke Takamichi, Norihiro Takamune, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari
INTERSPEECH2
2024 Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
Hitoshi Suda, Aya Watanabe, Shinnosuke Takamichi
INTERSPEECH3
2024 SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
Osamu Take, Shinnosuke Takamichi, Kentaro Seki, Yoshiaki Bando, Hiroshi Saruwatari
INTERSPEECH2
2024 DNN-Based Ensemble Singing Voice Synthesis With Interactions Between Singers
abstract
We propose a singing voice synthesis (SVS) method for a more unified ensemble singing voice by modeling interactions between singers. Most existing SVS methods aim to synthesize a solo voice, and do not consider interactions between singers, i.e., adjusting one’s own voice to the others’ voices. Since the production of ensemble voices from solo singing voices ignores the interactions, it can degrade the unity of the vocal ensemble. Therefore, we propose a SVS that reproduces the interactions. It is based on an architecture that uses musical scores of multiple voice parts, and loss functions that simulate the interactions’ effect to acoustic features. Experimental results show that our methods improve the unity of the vocal ensemble.
Hiroaki Hyodo, Shinnosuke Takamichi, Tomohiko Nakamura, Junya Koguchi, Hiroshi Saruwatari
SLT2
2024 JNV corpus: A corpus of Japanese nonverbal vocalizations with diverse phrases and emotions
abstract
We present JNV (Japanese Nonverbal Vocalizations) corpus, a corpus of Japanese nonverbal vocalizations (NVs) with diverse phrases and emotions. Existing Japanese NV corpora either lack phrase diversity or focus on a small number of emotions, which makes it difficult to analyze the characteristics of Japanese NVs and support downstream tasks like emotion recognition. We first propose a corpus-design method that contains two phases: (1) collecting NVs phrases based on crowd-sourcing; (2) recording NVs by stimulating speakers with emotional scenarios. We then collect 420 audio clips from 4 speakers that cover 6 emotions based on the proposed method. Results of comprehensive objective and subjective experiments demonstrate that (1) the emotions of the collected NVs can be recognized with high accuracy by both human evaluators and statistical models; (2) the collected NVs have a high authenticity comparable to previous corpora of English NVs. Additionally, we analyze the distributions of vowel types in Japanese and conduct feature importance analysis to show discriminative acoustic features between emotion categories in Japanese NVs. We publicate JNV to advance further development in this field.
Detai Xin, Shinnosuke Takamichi, Hiroshi Saruwatari
Speech Commun.2
2024 Text-Inductive Graphone-Based Language Adaptation for Low-Resource Speech Synthesis
abstract
Neural text-to-speech (TTS) systems have made significant progress in generating natural synthetic speech. However, neural TTS requires large amounts of paired training data, which limits its applicability to a small number of resource-rich languages. Previous work on low-resource TTS has addressed the data hungriness based on transfer learning from a multilingual model to low-resource languages, but it still relies heavily on the availability of paired data for the target languages. In this paper, we propose a text-inductive language adaptation framework for low-resource TTS to address the cost of collecting the paired data for low-resource languages. To inject textual knowledge during transfer learning, our framework employs a two-stage adaptation scheme that utilizes both text-only and supervised data for the target language. In the text-based adaptation stage, we update the language-aware embedding layer with a masked language model objective using text-only data for the target language. In the supervised adaptation stage, the entire TTS model is updated using paired data for the target language. We also propose a graphone-based multilingual training method that jointly uses graphemes and International Phonetic Alphabet symbols (referred to as graphones) for resource-rich languages, while using only graphemes for low-resource languages. This approach facilitates the transfer of pronunciation knowledge from resource-rich to low-resource languages. Through extensive evaluations, we demonstrate that 1) our framework with text-based adaptation outperforms the previous supervised transfer learning approach, 2) the proposed graphone-based training method further improves the performance of both multilingual TTS and low-resource language adaptation. With only 5 minutes of paired data for fine-tuning, our method achieved highly intelligible synthetic speech with the character error rates of around 6 % for a target language.
Takaaki Saeki, Soumi Maiti, Shinji Watanabe 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Yodas: Youtube-Oriented Dataset for Audio and Speech
abstract
In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled YouTube speech datasets. The labeled subsets, including manual or automatic subtitles, facilitate supervised model training. Conversely, the unlabeled subsets are apt for self-supervised learning applications. YODAS is distinctive as the first publicly available dataset of its scale, and it will be distributed under a Creative Commons license. We introduce the collection methodology utilized for YODAS, which contributes to the large-scale speech dataset construction. Subsequently, we provide a comprehensive analysis of speech, text contained within the dataset. Finally, we describe the speech recognition baselines over the top-15 languages.
Shinnosuke Takamichi, Takaaki Saeki, Sayaka Shiota, Shinji Watanabe 0001
ASRU2
2023 COCO-NUT: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-Based Control
abstract
In text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as text-to-image, free-form text instruction should be useful for intuitive and complicated control of voice characteristics. A sufficiently large corpus of high-quality and diverse voice samples with corresponding free-form descriptions can advance such control research. However, neither an open corpus nor a scalable method is currently available. To this end, we develop Coco-Nut, a new corpus including diverse Japanese utterances, along with text transcriptions and free-form voice characteristics descriptions. Our methodology to construct this corpus consists of 1) automatic collection of voice-related audio data from the Internet, 2) quality assurance, and 3) manual annotation using crowdsourcing. Additionally, we benchmark our corpus on the prompt embedding model trained by contrastive speech-text learning.
Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Wataru Nakata, Detai Xin, Hiroshi Saruwatari
ASRU2
2023 jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
abstract
We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese children’s songs and have six voice parts (lead vocal, soprano, alto, tenor, bass, and vocal percussion). They are divided into seven subsets, each of which features typical characteristics of a music genre such as jazz and enka. The variety in genre and voice part match vocal ensembles recently widespread in social media services such as YouTube, although the main targets of conventional vocal ensemble datasets are choral singing made up of soprano, alto, tenor, and bass. Experimental evaluation demonstrates that our corpus is a challenging resource for vocal ensemble separation. Our corpus is available on our project page.
Tomohiko Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama, Hiroshi Saruwatari
ICASSP2
2023 Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images
abstract
We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental sounds from the desired onomatopoeia texts. Onomatopoeias have another representation: visual-text representations of sounds in comics, advertisements, and virtual reality. A visual onomatopoeia (visual text of onomatopoeia) contains rich information that is not present in the text, such as a long-short duration of the image, so the use of this representation is expected to synthesize diverse sounds. Therefore, we propose visual onoma-to-wave for environmental sound synthesis from visual onomatopoeia. The method can transfer visual concepts of the visual text and sound-source image to the synthesized sound. We also propose a data augmentation method focusing on the repetition of onomatopoeias to enhance the performance of our method. An experimental evaluation shows that the methods can synthesize diverse environmental sounds from visual text and sound-source images.
Hien Ohnaka, Shinnosuke Takamichi, Keisuke Imoto, Yuki Okamoto, Kazuki Fujii, Hiroshi Saruwatari
ICASSP2
2023 MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture Models
abstract
In this paper, we propose a method for intermediating multiple speakers’ attributes and diversifying their voice characteristics in “speaker generation,” an emerging task that aims to synthesize a nonexistent speaker’s naturally sounding voice. The conventional TacoSpawn-based speaker generation method represents the distributions of speaker embeddings by Gaussian mixture models (GMMs) conditioned with speaker attributes. Although this method enables the sampling of various speakers from the speaker-attribute-aware GMMs, it is not yet clear whether the learned distributions can represent speakers with an intermediate attribute (i.e., mid-attribute). To this end, we propose an optimal-transport-based method that interpolates the learned GMMs to generate nonexistent speakers with mid-attribute (e.g., gender-neutral) voices. We empirically validate our method and evaluate the naturalness of synthetic speech and the controllability of two speaker attributes: gender and language fluency. The evaluation results show that our method can control the generated speakers’ attributes by a continuous scalar value without statistically significant degradation of speech naturalness.
Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Detai Xin, Hiroshi Saruwatari
ICASSP2
2023 Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts
abstract
We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does not fully represent the context information. The proposed method uses an acoustic context encoder and a textual context encoder to aggregate context information and feeds it to the TTS model, which enables the model to predict context-dependent prosody. We conducted comprehensive objective and subjective evaluations on a multi-speaker Japanese audiobook dataset. Experimental results demonstrate that the proposed method significantly outperforms two previous works. Additionally, we present insights about the different choices of context - modalities, lateral information and length - for audiobook TTS that have never been discussed in the literature before.
Detai Xin, Sharath Adavanne, Federico Ang, Ashish Kulkarni, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP5
2023 Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining
abstract
While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the target language. The use of text-only data allows the development of TTS systems for low-resource languages for which only textual resources are available, making TTS accessible to thousands of languages. Inspired by the strong cross-lingual transferability of multilingual language models, our framework first performs masked language model pretraining with multilingual text-only data. Then we train this model with a paired data in a supervised manner, while freezing a language-aware embedding layer. This allows inference even for languages not included in the paired data but present in the text-only data. Evaluation results demonstrate highly intelligible zero-shot TTS with a character error rate of less than 12% for an unseen language.
Takaaki Saeki, Soumi Maiti, Shinji Watanabe 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IJCAI5
2023 How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics
Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki, Detai Xin, Hiroshi Saruwatari
INTERSPEECH2
2023 CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
Yuki Saito 0001, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH3
2023 ChatGPT-EDSS: Empathetic Dialogue Speech Synthesis Trained from ChatGPT-derived Context Word Embeddings
Yuki Saito 0001, Shinnosuke Takamichi, Eiji Iimori, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH2
2023 HumanDiffusion: diffusion model using perceptual gradients
Yota Ueda, Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Hiroshi Saruwatari
INTERSPEECH2
2023 Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus
Detai Xin, Shinnosuke Takamichi, Ai Morimatsu, Hiroshi Saruwatari
INTERSPEECH2
2023 TimToShape: Supporting Practice of Musical Instruments by Visualizing Timbre with 2D Shapes based on Crossmodal Correspondences
abstract
Timbre is high-dimensional and sensuous, making it difficult for musical-instrument learners to improve their timbre. Although some systems exist to improve timbre, they require expert labeling for timbre evaluation; however, solely visualizing the results of unsupervised learning lacks the intuitiveness of feedback because human perception is not considered. Therefore, we employ crossmodal correspondences for intuitive visualization of the timbre. We designed TimToShape, a system that visualizes timbre with 2D shapes based on the user’s input of timbre–shape correspondences. TimToShape generates a shape morphed by linear interpolation according to the timbre’s position in the latent space, which is obtained by unsupervised learning with a variational autoencoder (VAE). We confirmed that people perceived shapes generated by TimToShape to correspond more to timbre than randomly generated shapes. Furthermore, a user study of six violin players revealed that TimToShape was well-received in terms of visual clarity and interpretability.
Kota Arai, Yutaro Hirao, Takuji Narumi, Tomohiko Nakamura, Shinnosuke Takamichi, Shigeo Yoshida
IUI5
2022 Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
INTERSPEECH3
2022 Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
abstract
We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history.Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems.Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context.As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling.To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterancewise modeling.The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method.
Yuto Nishimura, Yuki Saito 0001, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH3
2022 SelfRemaster: Self-Supervised Speech Restoration with Analysis-by-Synthesis Approach Using Channel Modeling
abstract
We present a self-supervised speech restoration method without paired speech corpora.Because the previous general speech restoration method uses artificial paired data created by applying various distortions to high-quality speech corpora, it cannot sufficiently represent acoustic distortions of real data, limiting the applicability.Our model consists of analysis, synthesis, and channel modules that simulate the recording process of degraded speech and is trained with real degraded speech data in a self-supervised manner.The analysis module extracts distortionless speech features and distortion features from degraded speech, while the synthesis module synthesizes the restored speech waveform, and the channel module adds distortions to the speech waveform.Our model also enables audio effect transfer, in which only acoustic distortions are extracted from degraded speech and added to arbitrary high-quality audio.Experimental evaluations with both simulated and real data show that our method achieves significantly higher-quality speech restoration than the previous supervised method, suggesting its applicability to real degraded speech materials.
Takaaki Saeki, Shinnosuke Takamichi, Tomohiko Nakamura, Naoko Tanji, Hiroshi Saruwatari
INTERSPEECH2
2022 UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
abstract
We present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022.The challenge is to predict the MOS values of speech samples collected from previous Blizzard Challenges and Voice Conversion Challenges for two tracks: a main track for in-domain prediction and an out-of-domain (OOD) track for which there is less labeled data from different listening tests.Our system is based on ensemble learning of strong and weak learners.Strong learners incorporate several improvements to the previous finetuning models of self-supervised learning (SSL) models, while weak learners use basic machine-learning methods to predict scores from SSL features.In the Challenge, our system had the highest score on several metrics for both the main and OOD tracks.In addition, we conducted ablation studies to investigate the effectiveness of our proposed methods.
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH5
2022 STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
abstract
We present STUDIES, a new speech corpus for developing a voice agent that can speak in a friendly manner.Humans naturally control their speech prosody to empathize with each other.By incorporating this "empathetic dialogue" behavior into a spoken dialogue system, we can develop a voice agent that can respond to a user more naturally.We designed the STUDIES corpus to include a speaker who speaks with empathy for the interlocutor's emotion explicitly.We describe our methodology to construct an empathetic dialogue speech corpus and report the analysis results of the STUDIES corpus.We conducted a text-to-speech experiment to initially investigate how we can develop more natural voice agent that can tune its speaking style corresponding to the interlocutor's emotion.The results show that the use of interlocutor's emotion label and conversational context embedding can produce speech with the same degree of naturalness as that synthesized by using the agent's emotion label.
Yuki Saito 0001, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH3
2022 J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis
abstract
In this paper, we construct a Japanese audiobook speech corpus called "J-MAC" for speech synthesis research.With the success of reading-style speech synthesis, the research target is shifting to tasks that use complicated contexts.Audiobook speech synthesis is a good example that requires cross-sentence, expressiveness, etc.Unlike reading-style speech, speaker-specific expressiveness in audiobook speech also becomes the context.To enhance this research, we propose a method of constructing a corpus from audiobooks read by professional speakers.From many audiobooks and their texts, our method can automatically extract and refine the data without any language dependency.Specifically, we use vocal-instrumental separation to extract clean data, connectionist temporal classification to roughly align text and audio, and voice activity detection to refine the alignment.J-MAC is open-sourced in our project page.We also conduct audiobook speech synthesis evaluations, and the results give insights into audiobook speech synthesis.
Shinnosuke Takamichi, Wataru Nakata, Naoko Tanji, Hiroshi Saruwatari
INTERSPEECH1
2022 Personalized Filled-pause Generation with Group-wise Prediction Models
abstract
In this paper, we propose a method to generate personalized filled pauses (FPs) with group-wise prediction models. Compared with fluent text generation, disfluent text generation has not been widely explored. To generate more human-like texts, we addressed disfluent text generation. The usage of disfluency, such as FPs, rephrases, and word fragments, differs from speaker to speaker, and thus, the generation of personalized FPs is required. However, it is difficult to predict them because of the sparsity of position and the frequency difference between more and less frequently used FPs. Moreover, it is sometimes difficult to adapt FP prediction models to each speaker because of the large variation of the tendency within each speaker. To address these issues, we propose a method to build group-dependent prediction models by grouping speakers on the basis of their tendency to use FPs. This method does not require a large amount of data and time to train each speaker model. We further introduce a loss function and a word embedding model suitable for FP prediction. Our experimental results demonstrate that group-dependent models can predict FPs with higher scores than a non-personalized one and the introduced loss function and word embedding model improve the prediction performance.
Yuta Matsunaga, Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
LREC3
2022 VTTS: Visual-Text To Speech
abstract
This paper proposes a visual-text to speech (vTTS) method, a method for synthesizing speech directly from visual text (i.e., text as an image). vTTS can use visual features in visual text that should be important for speech synthesis such as emphasis and radicals (components in Chinese characters), but they are not available in conventional TTS using discrete symbols as the input. The proposed vTTS method extracts visual features with a convolutional neural network and then generates acoustic features with a non-autoregressive model in an end-to-end manner. Experimental results show that 1) the vTTS method is capable of generating speech with naturalness comparable to or better than a conventional TTS, 2) it can transfer emphasis and emotion attributes in visual text to speech without additional labels and architectures, and 3) it can synthesize more natural and intelligible speech from unseen and rare characters than conventional TTS.
Yoshifumi Nakano, Takaaki Saeki, Shinnosuke Takamichi, Katsuhito Sudoh, Hiroshi Saruwatari
SLT3
2022 Lightweight and irreversible speech pseudonymization based on data-driven optimization of cascaded voice modification modules
abstract
In this paper, we propose a speech pseudonymization framework that utilizes cascaded and superposition-based voice modification modules. With increasing opportunities to use spoken dialogue systems nowadays, research regarding protecting the privacy of speaker information encapsulated in speech data is attracting attention. Pseudonymization, which is one method for voice privacy protection, aims to keep the intelligibility of speech while simultaneously suppressing speaker-specific information. One motivation of our framework is to achieve a reliable pseudonymization performance with light computation. To do this, we utilize the advantages of both machine learning-based and signal processing-based approaches. The advantages are (1) using signal processing-based methods parameterized with few hyperparameters and (2) using machine learning-based optimization to optimize all hyperparameters on the basis of black-box systems consisting of automatic speaker verification and automatic speech recognition. Our method of cascading signal processing modules, which are jointly optimized in a data-driven manner, can pseudonymize speech in a lightweight manner. Additionally, we discuss irreversible pseudonymization approaches and propose a superposition approach, yet another pseudonymization method that is more irreversible than the cascade method in terms of estimating the adequate parameters to recover the original signal. From the experimental results conducted under the VoicePrivacy 2020 protocols, we can demonstrate that (1) our cascade method succeeds in deteriorating the speaker recognition rate by over 24% while simultaneously improving the speech recognition rate by approximately 8% compared with a signal processing-based baseline system of VoicePrivacy 2020 and that (2) our superposition method works comparable to our cascade method in terms of pseudonymization performance.
Hiroto Kai, Shinnosuke Takamichi, Sayaka Shiota, Hitoshi Kiya
Comput. Speech Lang.2
2021 Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network
abstract
Incremental text-to-speech (TTS) synthesis generates utterances in small linguistic units for the sake of real-time and low-latency applications. We previously proposed an incremental TTS method that leverages a large pre-trained language model to take unobserved future context into account without waiting for the subsequent segment. Although this method achieves comparable speech quality to that of a method that waits for the future context, it entails a huge amount of processing for sampling from the language model at each time step. In this paper, we propose an incremental TTS method that directly predicts the unobserved future context with a lightweight model, instead of sampling words from the large-scale language model. We perform knowledge distillation from a GPT2-based context prediction network into a simple recurrent model by minimizing a teacher-student loss defined between the context embedding vectors of those models. Experimental results show that the proposed method requires about ten times less inference time to achieve comparable synthetic speech quality to that of our previous method, and it can perform incremental synthesis much faster than the average speaking speed of human English speakers, demonstrating the availability of our method to real-time applications.
Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
ASRU2
2021 Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception
abstract
We propose a conditional generative adversarial network (GAN) incorporating humans’ perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a real-data distribution accurately but can never represent a human-acceptable distribution, which are ranges of data in which humans accept the naturalness regardless of whether the data are real or not. A Human-GAN was proposed to model the human-acceptable distribution. A DNN-based generator is trained using a human-based discriminator, i.e., humans’ perceptual evaluations, instead of the GAN’s DNN-based discriminator. However, the HumanGAN cannot represent conditional distributions. This paper proposes the HumanACGAN, a theoretical extension of the HumanGAN, to deal with conditional human-acceptable distributions. Our HumanACGAN trains a DNN-based conditional generator by regarding humans as not only a discriminator but also an auxiliary classifier. The generator is trained by deceiving the human-based discriminator that scores the unconditioned naturalness and the human-based classifier that scores the class-conditioned perceptual acceptability. The training can be executed using the backpropagation algorithm involving humans’ perceptual evaluations. Our experimental results in phoneme perception demonstrate that our HumanACGAN can successfully train this conditional generator.
Yota Ueda, Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari
ICASSP4
2021 Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS
abstract
We propose a method for obtaining disentangled speaker and language representations via mutual information minimization and domain adaptation for cross-lingual text-to-speech (TTS) synthesis. The proposed method extracts speaker and language embeddings from acoustic features by a speaker encoder and a language encoder. Then the proposed method applies domain adaptation on the two embeddings to obtain language-invariant speaker embedding and speaker-invariant language embedding. To get more disentangled representations, the proposed method further uses mutual information minimization between the two embeddings to remove entangled information within each embedding. Disentangled representations of speaker and language are critical for cross-lingual TTS synthesis since entangled representations make it difficult to maintain speaker identity information when changing the language representation and consequently causes performance degradation. We evaluate the proposed method using English and Japanese multi-speaker datasets with a total of 207 speakers. Experimental result demonstrates that the proposed method significantly improves the naturalness and speaker similarity of both intra-lingual and cross-lingual TTS synthesis. Furthermore, we show that the proposed method has a good capability of maintaining the speaker identity between languages.
Detai Xin, Tatsuya Komatsu, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP3
2021 Digital Speech Makeup: Voice Conversion Based Altered Auditory Feedback for Transforming Self-Representation
abstract
Makeup (i.e., cosmetics) has long been used to transform not only one’s appearance but also their self-representation. Previous studies have demonstrated that visual transformations can induce a variety of effects on self-representation. Herein, we introduce Digital Speech Makeup (DSM), the novel concept of using voice conversion (VC) based auditory feedback to transform human self-representation. We implemented a proof-of-concept system that leverages a state-of-the-art algorithm for near real-time VC and bone-conduction headphones for resolving speech disruptions caused by delayed auditory feedback. Our user study confirmed that conversing for a few dozen minutes using the system influenced participants’ speech ownership and implicit bias. Furthermore, we reviewed the participants’ comments about the experience of DSM and gained additional qualitative insight into possible future directions for the concept. Our work represents the first step towards utilizing VC to design various interpersonal interactions, centered on influencing the users’ psychological state.
Riku Arakawa, Zendai Kashino, Shinnosuke Takamichi, Adrien Verhulst, Masahiko Inami
ICMI3
2021 Cross-Lingual Speaker Adaptation Using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
Interspeech3
2021 Lightweight Voice Anonymization Based on Data-Driven Optimization of Cascaded Voice Modification Modules
abstract
In this paper, we propose a voice anonymization framework based on data-driven optimization of cascaded voice modification modules. With increasing opportunities to use speech dialogue with machines nowadays, research regarding privacy protection of speaker information encapsulated in speech data is attracting attention. Anonymization, which is one of the methods for privacy protection, is based on signal processing manners, and the other one based on machine learning ones. Both approaches have a trade off between intelligibility of speech and degree of anonymization. The proposed voice anonymization framework utilizes advantages of machine learning and signal processing-based approaches to find the optimized trade off between the two. We use signal processing methods with training data for optimizing hyperparameters in a data-driven manner. The speech is modified using cascaded lightweight signal processing methods and then evaluated using black-box ASR and ASV, respectively. Our proposed method succeeded in deteriorating the speaker recognition rate by approximately 22% while simultaneously improved the speech recognition rate by over 3% compared to a signal processing-based conventional method.
Hiroto Kai, Shinnosuke Takamichi, Sayaka Shiota, Hitoshi Kiya
SLT2
2021 Incremental Text-to-Speech Synthesis Using Pseudo Lookahead With Large Pretrained Language Model
abstract
This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech. Incremental TTS is generally subject to a trade-off between latency and synthetic speech quality. It is challenging to produce high-quality speech with a low-latency setup that does not make much use of an unobserved future sentence (hereafter, “lookahead”). To resolve this issue, we propose an incremental TTS method that uses a pseudo lookahead generated with a language model to take the future contextual information into account without increasing latency. Our method can be regarded as imitating a human's incremental reading and uses pretrained GPT2, which accounts for the large-scale linguistic knowledge, for the lookahead generation. Evaluation results show that our method 1) achieves higher speech quality than the method taking only observed information into account and 2) achieves a speech quality equivalent to waiting for the future context observation.
Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE Signal Process. Lett.2
2021 Perceptual-Similarity-Aware Deep Speaker Representation Learning for Multi-Speaker Generative Modeling
abstract
We propose novel deep speaker representation learning that considers perceptual similarity among speakers for multi-speaker generative modeling. Following its success in accurate discriminative modeling of speaker individuality, knowledge of deep speaker representation learning (i.e., speaker representation learning using deep neural networks) has been introduced to multi-speaker generative modeling. However, the conventional discriminative algorithm does not necessarily learn speaker embeddings suitable for such generative modeling, which may result in lower quality and less controllability of synthetic speech. We propose three representation learning algorithms that utilize a perceptual speaker similarity matrix obtained by large-scale perceptual scoring of speaker-pair similarity. The algorithms train a speaker encoder to learn speaker embeddings with three different representations of the matrix: a set of vectors, the Gram matrix, and a graph. Furthermore, we propose an active learning algorithm that iterates the perceptual scoring and speaker encoder training. To obtain accurate embeddings while reducing costs of scoring and training, the algorithm selects unscored speaker-pairs to be scored next on the basis of the sequentially-trained speaker encoder's similarity prediction results. Experimental evaluation results show that 1) the proposed representation learning algorithms learn speaker embeddings strongly correlated with perceptual speaker-pair similarity, 2) the embeddings improve synthetic speech quality in speech autoencoding tasks better than conventional d-vectors learned by discriminative modeling, 3) the proposed active learning algorithm achieves higher synthetic speech quality while reducing costs of scoring and training, and 4) among the proposed similarity {vector, matrix, graph} embedding algorithms, the first achieves the best speaker similarity for synthetic speech and the third gives the most improvement in the synthetic speech naturalness.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception Modeling
abstract
We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribution by fooling the discriminator that distinguishes real and generated data. Therefore, the basic GAN cannot represent the outside of a real-data distribution. In the case of speech perception, humans can recognize not only human voices but also processed (i.e., a non-existent human) voices as human voice. Such a human-acceptable distribution is typically wider than a real-data one and cannot be modeled by the basic GAN. To model the human-acceptable distribution, we formulate a backpropagation-based generator training algorithm by regarding human perception as a black-boxed discriminator. The training efficiently iterates generator training by using a computer and discrimination by human. We evaluate our HumanGAN in speech naturalness modeling and demonstrate that it can represent a human-acceptable distribution that is wider than a real-data distribution.
Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari
ICASSP3
2020 Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials
abstract
In this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials. The conventional method with a minimum-phase filter achieves high-quality conversion but requires heavy computation in filtering. This is because the minimum phase using a fixed lifter of the Hilbert transform often results in a long-tap filter. One of our methods is a data-driven method for lifter training. Since this method takes filter truncation into account in training, it can shorten the tap length of the filter while preserving conversion accuracy. Our other method is sub-band processing for extending the conventional method from narrow-band (16 kHz) to full-band (48 kHz) VC, which can convert a full-band waveform with higher converted-speech quality. Experimental results indicate that 1) the proposed lifter-training method for narrow-band VC can shorten the tap length to 1/16 without degrading the converted-speech quality and 2) the proposed sub-band-processing method for full-band VC can improve the converted-speech quality than the conventional method.
Takaaki Saeki, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP3
2020 End-to-End Text-to-Speech Synthesis with Unaligned Multiple Language Units Based on Attention
Masashi Aso, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH2
2020 Real-Time, Full-Band, Online DNN-Based Voice Conversion System Using a Single CPU
Takaaki Saeki, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH3
2020 Cross-Lingual Text-To-Speech Synthesis via Domain Adaptation and Perceptual Similarity Regression in Speaker Space
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
INTERSPEECH3
2020 Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
INTERSPEECH4
2020 SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay
abstract
Developing a spontaneous speech corpus would be beneficial for spoken language processing and understanding. We present a speech corpus named the SMASH corpus, which includes spontaneous speech of two Japanese male commentators that made third-person audio commentaries during the gameplay of a fighting game. Each commentator ad-libbed while watching the gameplay with various topics covering not only explanations of each moment to convey the information on the fight but also comments to entertain listeners. We made transcriptions and topic tags as annotations on the recorded commentaries with our two-step method. We first made automatic and manual transcriptions of the commentaries and then manually annotated the topic tags. This paper describes how we constructed the SMASH corpus and reports some results of the annotations.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
LREC2
2020 DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech Corpus
abstract
In this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features called context instead of directly using text. In such frameworks, we can synthesize not only reading-style speech but also speech with paralinguistic and nonlinguistic features by adding such information to the context. However, it is not clear what kind of information is crucial for reproducing paralinguistic and nonlinguistic features. Therefore, we investigate the effectiveness of rich tags in DNN-based speech synthesis according to the Corpus of Spontaneous Japanese (CSJ), which has a large amount of annotations on paralinguistic features such as prosody, disfluency, and morphological features. Experimental evaluation results shows that the reproducibility of paralinguistic features of synthetic speech was enhanced by adding such information as context.
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
LREC4
2020 Phase reconstruction from amplitude spectrograms based on directional-statistics deep neural networks
abstract
This paper presents a deep neural network (DNN)-based phase reconstruction method from amplitude spectrograms. In speech processing, an amplitude spectrogram is often used for processing, and the corresponding phases are reconstructed from the amplitude spectrogram by using the Griffin-Lim method. However, the Griffin-Lim method causes unnatural artifacts in synthetic speech. To solve this problem, we propose the directional-statistics DNNs for predicting phases from the amplitude spectrograms. We first propose the von Mises distribution DNN, which is a generative model having the von Mises distribution and models histograms of a periodic variable. We extend it for modeling group delay that has a stronger connection to the amplitude spectrograms. Furthermore, we generalize the group-delay modeling and propose another DNN called the sine-skewed generalized cardioid distribution DNN for modeling asymmetric histograms such as a group delay. Results from objective and subjective evaluations indicate that (1) our von Mises distribution DNN can predict group delay more accurately than predicting phases, (2) our DNN works as better initialization of the Griffin-Lim method, (3) the phase reconstruction methods based on our von Mises distribution DNN achieve better speech quality than the conventional Griffin-Lim method, and (4) our sine-skewed generalized cardioid distribution DNN models the group delay more accurately than our von Mises distribution DNN.
Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari
Signal Process.1
2020 Acoustic model-based subword tokenization and prosodic-context extraction without language knowledge for text-to-speech synthesis
abstract
This paper presents text tokenization and context extraction without using language knowledge for text-to-speech (TTS) synthesis. To generate prosody, statistical parametric TTS synthesis typically requires the professional knowledge of the target language. Therefore, languages suitable for TTS synthesis are limited to only rich-resource languages. To achieve TTS synthesis without using language knowledge, we propose acoustic model-based subword tokenization and unsupervised extraction of prosodic contexts. The subword tokenization can determine language units suitable for prosody generation. The context extraction can retrieve contexts from pairs of subwords and prosody. The proposed methods function without language knowledge and can improve F0 prediction accuracy. Experimental evaluation demonstrates that 1) the training of proposed subword tokenization, which uses the expectation-maximization algorithm and deep neural networks, is empirically stable, 2) the proposed subword tokenization tokenizes text into subwords that are close to language-specific units, and 3) the proposed methods outperform the conventional methods using language model-based tokenization in terms of synthetic speech quality.
Masashi Aso, Shinnosuke Takamichi, Norihiro Takamune, Hiroshi Saruwatari
Speech Commun.2
2019 Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking
abstract
This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double-tracking, a recording method in which two performances of the same phrase are recorded and mixed to create a richer, layered sound. However, singing voices synthesized using conventional DNN-based methods never vary because the synthesis process is deterministic and only one waveform is synthesized from one musical score. To address this problem, we use a GMMN to model the variation of the modulation spectrum of the pitch contour of natural singing voices and add a randomized inter-utterance variation to the pitch contour generated by conventional DNN-based singing voice synthesis. Experimental evaluations suggest that 1) our approach can provide perceptible inter-utterance pitch variation while preserving speech quality. We extend our approach to double-tracking, and the evaluation demonstrates that 2) GMMN-based neural double-tracking is perceptually closer to natural double-tracking than conventional signal processing-based artificial double-tracking is.
Hiroki Tamaru, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
ICASSP3
2019 Speech Quality Evaluation of Synthesized Japanese Speech Using EEG
Ivan Halim Parmonangan, Hiroki Tanaka, Sakriani Sakti, Shinnosuke Takamichi, Satoshi Nakamura 0001
INTERSPEECH4
2019 Vocoder-free text-to-speech synthesis incorporating generative adversarial networks using low-/multi-frequency STFT amplitude spectra
abstract
This paper proposes novel training algorithms for vocoder-free text-to-speech (TTS) synthesis based on generative adversarial networks (GANs) that compensate for short-term Fourier transform (STFT) amplitude spectra in low/multi frequency resolution. Vocoder-free TTS using STFT amplitude spectra can avoid degradation of synthetic speech quality caused by the vocoder-based parameterization used in conventional TTS. Our previous work for the vocoder-based TTS proposed a method for incorporating the GAN-based distribution compensation into acoustic model training to improve synthetic speech quality. This paper extends the algorithm to the vocoder-free TTS and propose a GAN-based training algorithm using low-frequency-resolution amplitude spectra to overcome the difficulty in modeling complicated distribution of the high-dimensional spectra. In the proposed algorithm, amplitude spectra are transformed into low-frequency-resolution amplitude spectra by applying an average pooling function along with a frequency axis; then the GAN-based distribution compensation is performed in the low-frequency-resolution domain. Because the low-frequency-resolution amplitude spectra approximately emulate filter banks, the proposed algorithm is expected to improve synthetic speech quality by reducing differences in spectral envelopes of natural and synthetic speech. Furthermore, various frequency scales that are related to human speech perception (e.g., mel and inverse mel frequency scales) can be introduced to the proposed training algorithm by applying an frequency warping function to amplitude spectra. This paper also proposes a GAN-based training algorithm using multi-frequency-resolution amplitude spectra that uses both low- and original-frequency-resolution amplitude spectra to reduce the differences in not only spectral envelopes but also fine structures. Experimental results demonstrate that (1) GANs using low-frequency-resolution amplitude spectra improve speech quality and work robustly against the settings of the frequency resolution and hyperparameters, (2) in comparison among low-, original-, and multi-frequency-resolution amplitude spectra, the use of low-frequency-resolution ones work best improve the synthetic speech quality, and (3) the use of the inverse mel frequency scale for obtaining low-frequency-resolution amplitude spectra further improves synthetic speech quality.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
Comput. Speech Lang.2
2019 Independent Deeply Learned Matrix Analysis for Determined Audio Source Separation
abstract
In this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based multichannel audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy.
Naoki Makishima, Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hayato Sumino, Shinnosuke Takamichi, Hiroshi Saruwatari, Nobutaka Ono
IEEE ACM Trans. Audio Speech Lang. Process.6
2018 Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors
abstract
This paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish because of an over-regularization issue often observed in latent variables of the VAEs. To overcome the issue, this paper proposes a VAE-based non-parallel VC conditioned by not only the speaker representations but also phonetic contents of speech represented as phonetic posteriorgrams (PPGs). Since the phonetic contents are given during the training, we can expect that the VC models effectively learn speaker-independent latent features of speech. Focusing on the point, this paper also extends the conventional VAE-based non-parallel VC to many-to-many VC that can convert arbitrary speakers' characteristics into another arbitrary speakers' ones. We investigate two methods to estimate speaker representations for speakers not included in speech corpora used for training VC models: 1) adapting conventional speaker codes, and 2) using d-vectors for the speaker representations. Experimental results demonstrate that 1) PPGs successfully improve both naturalness and speaker similarity of the converted speech, and 2) both speaker codes and d-vectors can be adopted to the VAE-based many-to-many non-parallel VC.
Yuki Saito 0001, Yusuke Ijima, Kyosuke Nishida, Shinnosuke Takamichi
ICASSP4
2018 Text-to-Speech Synthesis Using STFT Spectra Based on Low-/Multi-Resolution Generative Adversarial Networks
abstract
This paper proposes novel training algorithms for vocoder-free statistical parametric speech synthesis (SPSS) using short-term Fourier transform (STFT) spectra. Recently, text-to-speech synthesis using STFT spectra has been investigated since it can avoid quality degradation caused by the vocoder-based parameterization in conventional SPSS using a vocoder. In conventional SPSS using a vocoder, we previously proposed a training algorithm for integrating generative adversarial network (GAN)-based distribution compensation. To extend the algorithm to vocoder-free SPSS, we propose low- and multi-resolution GAN-based training algorithms for vocoder-free SPSS. In our algorithm that uses the low-resolution GAN, acoustic models are trained to minimize the weighted sum of the mean squared error between natural and generated spectra in the original resolution and adversarial loss to deceive discriminative models in the lower resolution. Since the low-resolution spectra are close to filter banks and their distribution becomes simpler, GAN-based distribution compensation works well. Furthermore, we propose an algorithm using multi-resolution GANs, which uses both the low-resolution GAN and original-resolution GAN. Experimental results demonstrate that 1) the low-resolution GAN works robustly to the setting of its frequency resolution and hyperparameter, and 2) compared the low-, original-, and multi-resolution GANs, the low-resolution GAN works the best to improve synthetic speech quality.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP2
2018 CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects
Shinnosuke Takamichi, Hiroshi Saruwatari
LREC1
2018 An end-to-end model for cross-lingual transformation of paralinguistic information
Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001
Mach. Transl.2
2018 Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks
abstract
A method for statistical parametric speech synthesis incorporating generative adversarial networks (GANs) is proposed. Although powerful deep neural networks techniques can be applied to artificially synthesize speech waveform, the synthetic speech quality is low compared with that of natural speech. One of the issues causing the quality degradation is an oversmoothing effect often observed in the generated speech parameters. A GAN introduced in this paper consists of two neural networks: a discriminator to distinguish natural and generated samples, and a generator to deceive the discriminator. In the proposed framework incorporating the GANs, the discriminator is trained to distinguish natural and generated speech parameters, while the acoustic models are trained to minimize the weighted sum of the conventional minimum generation loss and an adversarial loss for deceiving the discriminator. Since the objective of the GANs is to minimize the divergence (i.e., distribution difference) between the natural and generated speech parameters, the proposed method effectively alleviates the oversmoothing effect on the generated speech parameters. We evaluated the effectiveness for text-to-speech and voice conversion, and found that the proposed method can generate more natural spectral parameters and F0than conventional minimum generation error training algorithm regardless of its hyperparameter settings. Furthermore, we investigated the effect of the divergence of various GANs, and found that a Wasserstein GAN minimizing the Earth-Mover's distance works the best in terms of improving the synthetic speech quality.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Blind source separation based on independent low-rank matrix analysis with sparse regularization for time-series activity
abstract
In this paper, we propose a new blind source separation (BSS) method based on independent low-rank matrix analysis (ILRMA) with novel sparse regularization. ILRMA is a recently proposed BSS algorithm that simultaneously estimates a demixing matrix and source spectrogram models based on nonnegative matrix factorization (NMF). To improve the separation accuracy and stability, an additional constraint such as sparseness is needed but there have been no studies on this so far. In this study, we introduce an a priori statistical model for time-series amplitudes of source spectrograms, employing a new frequency-wise sparse regularization using estimates from the Bayesian postfilter to enhance the modeling accuracy. This regularization results in a bilevel optimization problem that consists of the estimation of a sparsity-emphasized source model using NMF and the separation of sources by ILRMA. In this paper, we present two approximated optimization schemes and their combination for performing regularized ILRMA. The efficacy of the proposed method is confirmed in a BSS experiment.
Yoshiki Mitsui, Daichi Kitamura, Shinnosuke Takamichi, Nobutaka Ono, Hiroshi Saruwatari
ICASSP3
2017 Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesis
abstract
This paper proposes a novel training algorithm for high-quality Deep Neural Network (DNN)-based speech synthesis. The parameters of synthetic speech tend to be over-smoothed, and this causes significant quality degradation in synthetic speech. The proposed algorithm takes into account an Anti-Spoofing Verification (ASV) as an additional constraint in the acoustic model training. The ASV is a discriminator trained to distinguish natural and synthetic speech. Since acoustic models for speech synthesis are trained so that the ASV recognizes the synthetic speech parameters as natural speech, the synthetic speech parameters are distributed in the same manner as natural speech parameters. Additionally, we find that the algorithm compensates not only the parameter distributions, but also the global variance and the correlations of synthetic speech parameters. The experimental results demonstrate that 1) the algorithm outperforms the conventional training algorithm in terms of speech quality, and 2) it is robust against the hyper-parameter settings.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP2
2017 Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities
abstract
Voice conversion (VC) using sequence-to-sequence learning of context posterior probabilities is proposed.Conventional VC using shared context posterior probabilities predicts target speech parameters from the context posterior probabilities estimated from the source speech parameters.Although conventional VC can be built from non-parallel data, it is difficult to convert speaker individuality such as phonetic property and speaking rate contained in the posterior probabilities because the source posterior probabilities are directly used for predicting target speech parameters.In this work, we assume that the training data partly include parallel speech data and propose sequence-to-sequence learning between the source and target posterior probabilities.The conversion models perform non-linear and variable-length transformation from the source probability sequence to the target one.Further, we propose a joint training algorithm for the modules.In contrast to conventional VC, which separately trains the speech recognition that estimates posterior probabilities and the speech synthesis that predicts target speech parameters, our proposed method jointly trains these modules along with the proposed probability conversion modules.Experimental results demonstrate that our approach outperforms the conventional VC.
Hiroyuki Miyoshi, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH3
2017 Sampling-Based Speech Parameter Generation Using Moment-Matching Networks
abstract
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis.Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis produces completely the same speech, i.e., there is no inter-utterance variation in synthetic speech.To give synthetic speech natural inter-utterance variation, this paper builds DNN acoustic models that make it possible to randomly sample speech parameters.The DNNs are trained so that they make the moments of generated speech parameters close to those of natural speech parameters.Since the variation of speech parameters is compressed into a low-dimensional simple prior noise vector, our algorithm has lower computation cost than direct sampling of speech parameters.As the first step towards generating synthetic speech that has natural inter-utterance variation, this paper investigates whether or not the proposed sampling-based generation deteriorates synthetic speech quality.In evaluation, we compare speech quality of conventional maximum likelihood-based generation and proposed sampling-based generation.The result demonstrates the proposed generation causes no degradation in speech quality.
Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
INTERSPEECH1
2016 The NU-NAIST Voice Conversion System for the Voice Conversion Challenge 2016
Kazuhiro Kobayashi, Shinnosuke Takamichi, Satoshi Nakamura 0001, Tomoki Toda
INTERSPEECH2
2016 Postfilters to Modify the Modulation Spectrum for Statistical Parametric Speech Synthesis
abstract
This paper presents novel approaches based on modulation spectrum (MS) for high-quality statistical parametric speech synthesis, including text-to-speech (TTS) and voice conversion (VC). Although statistical parametric speech synthesis offers various advantages over concatenative speech synthesis, the synthetic speech quality is still not as good as that of concatenative speech synthesis or the quality of natural speech. One of the biggest issues causing the quality degradation is the over-smoothing effect often observed in the generated speech parameter trajectories. Global variance (GV) is known as a feature well correlated with the over-smoothing effect, and the effectiveness of keeping the GV of the generated speech parameter trajectories similar to those of natural speech has been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we propose using the MS of the generated speech parameter trajectories as a new feature to effectively quantify the over-smoothing effect. Moreover, we propose postfilters to modify the MS utterance by utterance or segment by segment to make the MS of synthetic speech close to that of natural speech. The proposed postfilters are applicable to various synthesizers based on statistical parametric speech synthesis. We first perform an evaluation of the proposed method in the framework of hidden Markov model (HMM)-based TTS, examining its properties from different perspectives. Furthermore, effectiveness of the proposed postfilters are also evaluated in Gaussian mixture model (GMM)-based VC and classification and regression trees (CART)-based TTS (a.k.a., CLUSTERGEN). The experimental results demonstrate that 1) the proposed utterance-level postfilter achieves quality comparable to the conventional generation algorithm considering the GV, and yields significant improvements by applying to the GV-based generation algorithm in HMM-based TTS, 2) the proposed segment-level postfilter capable of achieving low-delay synthesis also yields significant improvements in synthetic speech quality, and 3) the proposed postfilters are also effective in not only HMM-based TTS but also GMM-based VC and CLUSTERGEN.
Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesis
abstract
This paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global Variance (GV) is known as a feature well correlated with the over-smoothing effect and a metric on the GV of the generated parameters is effectively used as a penalty term in the conventional parameter generation. However, the quality of the synthetic speech is far from that of the natural speech. Recently, we have found that a Modulation Spectrum (MS) of the generated parameters, which is also regarded as an extension of the GV, is more sensitively correlated with the over-smoothing effect than the GV. This paper incorporates a metric on the MS as a new penalty term in the proposed parameter generation algorithm. The experimental results demonstrate that the proposed parameter generation algorithm considering the MS yields significant improvements in synthetic speech quality compared to the conventional parameter generation algorithm considering the GV.
Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001
ICASSP1
2015 Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice Conversion
abstract
This paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech is still significantly worse than that of natural speech. In order to address this problem while preserving the computationally efficient conversion processing, the proposed training method enables 1) to use a consistent optimization criterion between training and conversion and 2) to compensate a Modulation Spectrum (MS) of the converted parameter trajectory as a feature sensitively correlated with over-smoothing effects causing quality degradation of the converted speech. The experimental results demonstrate that the proposed algorithm yields significant improvements in term of both the converted speech quality and the conversion accuracy for speaker individuality compared to the basic training algorithm.
Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001
ICASSP1
2015 Preserving word-level emphasis in speech-to-speech translation using linear regression HSMMs
abstract
In speech, emphasis is an important type of paralinguistic information that helps convey the focus of an utterance, new information, and emotion. If emphasis can be incorporated into a speech-to-speech (S2S) translation system, it will be possible to convey this information across the language barrier. However, previous related work focuses only on the translation of particular prosodic features, such as F0, or works with emphasis but focuses on extremely small vocabularies, such as the 10 digits. In this paper, we describe a new S2S method that is able to translate the emphasis across languages and consider multiple features of emphasis such as power, F0, and duration over larger vocabularies. We do so by introducing two new components: word-level emphasis estimation using linear regression hidden semi-Markov models, and emphasis translation that translates the word-level emphasis to the target language with conditional random fields. The text-to-speech synthesis system is also modified to be able to synthesize emphasized speech. The result shows that our system can translate the emphasis correctly with 91.6% F -measure for objective test, and 87.8% for subjective test.
Quoc Truong Do, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001
INTERSPEECH2
2015 Non-native speech synthesis preserving speaker individuality based on partial correction of prosodic and phonetic characteristics
abstract
This paper presents a novel non-native speech synthesis technique that preserves the individuality of a non-native speaker. Cross-lingual speech synthesis based on voice conversion or HMM-based speech synthesis, which synthesizes foreign language speech of a specific non-native speaker reflecting the speaker-dependent acoustic characteristics extracted from the speaker’s natural speech in his/her mother tongue, tends to cause a degradation of speaker individuality in synthetic speech compared to intra-lingual speech synthesis. This paper proposes a new approach to cross-lingual speech synthesis that preserves speaker individuality by explicitly using non-native speech spoken by the target speaker. Although the use of nonnative speech makes it possible to preserve the speaker individuality in the synthesized target speech, naturalness is significantly degraded as the speech is directly affected by unnatural prosody and pronunciation often caused by differences in the linguistic systems of the source and target languages. To improve naturalness while preserving speaker individuality, we propose (1) a prosodic correction method based on model adaptation, and (2) a phonetic correction method based on spectrum replacement for unvoiced consonants. The experimental results demonstrate that these proposed methods are capable of significantly improving naturalness while preserving the speaker individuality in synthetic speech.
Yuji Oshima, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001
INTERSPEECH2
2015 Modulation spectrum-constrained trajectory training algorithm for HMM-based speech synthesis
Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001
INTERSPEECH1
2014 A postfilter to modify the modulation spectrum in HMM-based speech synthesis
abstract
In this paper, we propose a postfilter to compensate modulation spectrum in HMM-based speech synthesis. In order to alleviate over-smoothing effects which is a main cause of quality degradation in HMM-based speech synthesis, it is necessary to consider features that can capture over-smoothing. Global Variance (GV) is one well-known example of such a feature, and the effectiveness of parameter generation algorithm considering GV have been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we introduce the Modulation Spectrum (MS) of speech parameter trajectory as a new feature to effectively capture the over-smoothing effect, and we propose a postfilter based on the MS. The MS is represented as a power spectrum of the parameter trajectory. The generated speech parameter sequence is filtered to ensure that its MS has a pattern similar to natural speech. Experimental results show quality improvements when the proposed methods are applied to spectral and F0components, compared with conventional methods considering GV.
Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001
ICASSP1
2014 A hearing impairment simulation method using audiogram-based approximation of auditory charatecteristics
abstract
Hearing impairment simulation is an effective technique to educate normal-hearing people about auditory perception of the hearing-impaired. Because auditory characteristics of the hearing impaired vary greatly from person-to-person, personalization of the hearing impairment simulation systems is essential to accurately simulate these individual differences. However, measurement of auditory characteristics of individuals is time-consuming work. In this paper, we propose a hearing impairment simulation method that is easily applied to individual hearing-impaired persons. Auditory filter characteristics and gain characteristics are estimated from easily measurable audiograms of each individual. We also implement a method for manually adjusting the hearing impairment level to improve accuracy of the proposed hearing impairment simulation. An experimental evaluation is conducted to compare intelligibility between hearing-impaired and normal-hearing persons with the proposed hearing impairment simulation. The experimental results show that the proposed method effectively makes the word correct rate and phoneme confusion tendency of the normal hearing persons similar to those of the hearing impaired persons. Index Terms: hearing-impairment simulation, personalization, auditory filter characteristics, gain characteristics, audiogram,
Nozomi Jinbo, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001
INTERSPEECH2
2013 Generalizing continuous-space translation of paralinguistic information
abstract
In previous work, we proposed a model for speech-to-speech translation that is sensitive to paralinguistic information such as duration and power of spoken words [1]. This model uses linear regression to map source acoustic features to target acoustic features directly and in continuous space. However, while the model is effective, it faces scalability issues as a single model must be trained for every word, which makes it difficult to generalize to words for which we do not have parallel speech. In this work we first demonstrate that simply training a linear regression model on all words is not sufficient to express paralinguistic translation. We next describe a neural network model that has sufficient expressive power to perform paralinguistic translation with a single model. We evaluate the proposed method on a digit translation task and show that we achieve similar results with a single neural network-based model as previous work did using word-dependent models. Index Terms: speech translation, paralinguistic information, linear regression, neural network
Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001
INTERSPEECH2
2013 Improvements to HMM-based speech synthesis based on parameter generation with rich context models
abstract
In this paper, we improve parameter generation with rich context models by modifying an initialization method and further apply it to both spectral and F0 components in HMM-based speech synthesis. To alleviate over-smoothing effects caused by the traditional parameter generation methods, we have previously proposed an iterative parameter generation method with rich context models. It has been reported that this method yields quality improvements in synthetic speech but there are still limitations. This is because 1) this generation method still suffers from the over-smoothing effect, as it uses the parameters generated by the traditional method as an initial parameters, which strongly affect on the finally generated parameters and 2) it is applied to only the spectral component. To address these issues, we propose 1) an initialization method to generate less smoothed but more discontinuous initial parameters that tend to yield better generated parameters, and 2) a parameter generation method with rich context models for the F0 component. Experimental results show that the proposed methods yield significant improvements in quality of synthetic speech. Index Terms: HMM-based speech synthesis, rich context models, GMM, context clustering, over-smoothing, MSD-HMM
Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001
INTERSPEECH1
2012 An Evaluation of Parameter Generation Methods with Rich Context Models in HMM-Based Speech Synthesis
Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai, Sakriani Sakti, Satoshi Nakamura 0001
INTERSPEECH1