EDBT 2026 Demo / reviewers in the wild / expert
Kentaro Tachibana
dblp:35/8056
· DBLP profile ↗
25ranked-venue papers
3as first author
18since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAVIARES: Corpus for Audio-Visual Expressive Voice AgentabstractHigh-quality audio-visual corpora are essential for building voice agents capable of natural human-machine communication, but existing corpora commonly contain a limited amount of data per speaker, making personalized modeling difficult. We present CAVIARES, a new audio-visual corpus comprising 9.5 hours of expressive speech recorded by a single professional Japanese female speaker. CAVIARES consists of two subsets: acted dialogue and expressive reading, providing a diverse range of speaking styles for speech-to-facial motion modeling and multimodal learning tasks. In this paper, we describe the construction process of CAVIARES and the results of corpus analysis. CAVIARES will be released for research purposes only. Jinsheng Chen, Yuki Saito 0001, Naoko Tanji, Hironori Doi, Byeongseon Park, Yuma Shirahata, Kentaro Tachibana, Hiroshi Saruwatari |
ASRU | 8 |
| 2025 | Description-Based Controllable Text-to-Speech With Cross-Lingual Voice ControlabstractWe propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target language with a description control model trained on another language, which maps input text descriptions to the conditional features of the TTS model. These two models share disentangled timbre and style representations based on self-supervised learning (SSL), allowing for disentangled voice control, such as controlling speaking styles while retaining the original timbre. Furthermore, because the SSL-based timbre and style representations are language-agnostic, combining the TTS and description control models while sharing the same embedding space effectively enables cross-lingual control of voice characteristics. Experiments on English and Japanese TTS demonstrate that our method achieves high naturalness and controllability for both languages, even though no Japanese audio-description pairs are used. Ryuichi Yamamoto, Yuma Shirahata, Masaya Kawamura, Kentaro Tachibana |
ICASSP | 4 |
| 2024 | PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language DescriptionsabstractWe propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the concept of speaker prompt, which describes voice characteristics (e.g., gender-neutral, young, old, and muffled) designed to be approximately independent of speaking style. Since there is no large-scale dataset containing speaker prompts, we first construct a dataset based on the LibriTTS-R corpus with manually annotated speaker prompts. We then employ a diffusion-based acoustic model with mixture density networks to model diverse speaker factors in the training data. Unlike previous studies that rely on style prompts describing only a limited aspect of speaker individuality, such as pitch, speaking speed, and energy, our method utilizes an additional speaker prompt to effectively learn the mapping from natural language descriptions to the acoustic features of diverse speakers. Our subjective evaluation results show that the proposed method can better control speaker characteristics than the methods without the speaker prompt. Audio samples are available at https://reppy4620.github.io/demo.promptttspp/. Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, Kentaro Tachibana |
ICASSP | 7 |
| 2024 | Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
Takuto Igarashi, Yuki Saito 0001, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 6 |
| 2024 | LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi, Kentaro Tachibana |
INTERSPEECH | 5 |
| 2024 | SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
Yuki Saito 0001, Takuto Igarashi, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 6 |
| 2024 | Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
INTERSPEECH | 4 |
| 2023 | Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier TransformabstractWe propose a lightweight end-to-end text-to-speech model using multi-band generation and inverse short-time Fourier transform. Our model is based on VITS, a high-quality end-to-end text-to-speech model, but adopts two changes for more efficient inference: 1) the most computationally expensive component is partially replaced with a simple inverse short-time Fourier transform, and 2) multi-band generation, with fixed or trainable synthesis filters, is used to generate waveforms. Unlike conventional lightweight models, which employ optimization or knowledge distillation separately to train two cascaded components, our method enjoys the full benefits of end-to-end optimization. Experimental results show that our model synthesized speech as natural as that synthesized by VITS, while achieving a real-time factor of 0.066 on an Intel Core i7 CPU, 4.1 times faster than VITS. Moreover, a smaller version of the model significantly outperformed a lightweight baseline model with respect to both naturalness and inference speed. Code and audio samples are available from https://github.com/MasayaKawamura/MB-iSTFT-VITS. Masaya Kawamura, Yuma Shirahata, Ryuichi Yamamoto, Kentaro Tachibana |
ICASSP | 4 |
| 2023 | Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech SynthesisabstractSeveral fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate unstable pitch contour with audible artifacts when the dataset contains emotional attributes, i.e., large diversity of pronunciation and prosody. To address this problem, we propose Period VITS, a novel end-to-end TTS model that incorporates an explicit periodicity generator. In the proposed method, we introduce a frame pitch predictor that predicts prosodic features, such as pitch and voicing flags, from the input text. From these features, the proposed periodicity generator produces a sample-level sinusoidal source that enables the waveform decoder to accurately reproduce the pitch. Finally, the entire model is jointly optimized in an end-to-end manner with variational inference and adversarial objectives. As a result, the decoder becomes capable of generating more stable, expressive, and natural output waveforms. The experimental results showed that the proposed model significantly outperforms baseline models in terms of naturalness, with improved pitch stability in the generated samples. Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima, Jae-Min Kim, Kentaro Tachibana |
ICASSP | 6 |
| 2023 | Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANsabstractNeural audio super-resolution models are typically trained on low- and high-resolution audio signal pairs. Although these methods achieve highly accurate super-resolution if the acoustic characteristics of the input data are similar to those of the training data, challenges remain: the models suffer from quality degradation for out-of-domain data, and paired data are required for training. To address these problems, we propose Dual-CycleGAN, a high-quality audio super-resolution method that can utilize unpaired data based on two connected cycle consistent generative adversarial networks (CycleGAN). Our method decomposes the super-resolution method into domain adaptation and resampling processes to handle acoustic mismatch in the unpaired low- and high-resolution signals. The two processes are then jointly optimized within the CycleGAN framework. Experimental results verify that the proposed method significantly outperforms conventional methods when paired data are not available. Code and audio samples are available from https://chomeyama.github.io/DualCycleGAN-Demo/. Reo Yoneyama, Ryuichi Yamamoto, Kentaro Tachibana |
ICASSP | 3 |
| 2023 | CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
Yuki Saito 0001, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2023 | ChatGPT-EDSS: Empathetic Dialogue Speech Synthesis Trained from ChatGPT-derived Context Word Embeddings
Yuki Saito 0001, Shinnosuke Takamichi, Eiji Iimori, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2022 | Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue HistoryabstractWe propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history.Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems.Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context.As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling.To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterancewise modeling.The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method. Yuto Nishimura, Yuki Saito 0001, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2022 | A Unified Accent Estimation Method Based on Multi-Task Learning for Japanese Text-to-Speech
Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
INTERSPEECH | 3 |
| 2022 | DRSpeech: Degradation-Robust Text-to-Speech Synthesis with Frame-Level and Utterance-Level Acoustic Representation LearningabstractMost text-to-speech (TTS) methods use high-quality speech corpora recorded in a well-designed environment, incurring a high cost for data collection.To solve this problem, existing noise-robust TTS methods are intended to use noisy speech corpora as training data.However, they only address either timeinvariant or time-variant noises.We propose a degradationrobust TTS method, which can be trained on speech corpora that contain both additive noises and environmental distortions.It jointly represents the time-variant additive noises with a framelevel encoder and the time-invariant environmental distortions with an utterance-level encoder.We also propose a regularization method to attain clean environmental embedding that is disentangled from the utterance-dependent information such as linguistic contents and speaker characteristics.Evaluation results show that our method achieved significantly higher-quality synthetic speech than previous methods in the condition including both additive noise and reverberation. Takaaki Saeki, Kentaro Tachibana, Ryuichi Yamamoto |
INTERSPEECH | 2 |
| 2022 | STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice AgentabstractWe present STUDIES, a new speech corpus for developing a voice agent that can speak in a friendly manner.Humans naturally control their speech prosody to empathize with each other.By incorporating this "empathetic dialogue" behavior into a spoken dialogue system, we can develop a voice agent that can respond to a user more naturally.We designed the STUDIES corpus to include a speaker who speaks with empathy for the interlocutor's emotion explicitly.We describe our methodology to construct an empathetic dialogue speech corpus and report the analysis results of the STUDIES corpus.We conducted a text-to-speech experiment to initially investigate how we can develop more natural voice agent that can tune its speaking style corresponding to the interlocutor's emotion.The results show that the use of interlocutor's emotion label and conversational context embedding can produce speech with the same degree of naturalness as that synthesized by using the agent's emotion label. Yuki Saito 0001, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2022 | Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data AugmentationabstractData augmentation via voice conversion (VC) has been successfully applied to low-resource expressive text-to-speech (TTS) when only neutral data for the target speaker are available. Although the quality of VC is crucial for this approach, it is challenging to learn a stable VC model because the amount of data is limited in low-resource scenarios, and highly expressive speech has large acoustic variety. To address this issue, we propose a novel data augmentation method that combines pitch-shifting and VC techniques. Because pitch-shift data augmentation enables the coverage of a variety of pitch dynamics, it greatly stabilizes training for both VC and TTS models, even when only 1,000 utterances of the target speaker's neutral data are available. Subjective test results showed that a FastSpeech 2-based emotional TTS system with the proposed method improved naturalness and emotional similarity compared with conventional methods. Ryo Terashima, Ryuichi Yamamoto, Eunwoo Song, Yuma Shirahata, Hyun-Wook Yoon, Jae-Min Kim, Kentaro Tachibana |
INTERSPEECH | 7 |
| 2021 | Phrase Break Prediction with Bidirectional Encoder Representations in Japanese Text-to-Speech SynthesisabstractWe propose a novel phrase break prediction method that combines implicit features extracted from a pre-trained large language model, a.k.a BERT, and explicit features extracted from BiLSTM with linguistic features. In conventional BiLSTM based methods, word representations and/or sentence representations are used as independent components. The proposed method takes account of both representations to extract the latent semantics, which cannot be captured by previous methods. The objective evaluation results show that the proposed method obtains an absolute improvement of 3.2 points for the F1 score compared with BiLSTM-based conventional methods using linguistic features. Moreover, the perceptual listening test results verify that a TTS system that applied our proposed method achieved a mean opinion score of 4.39 in prosody naturalness, which is highly competitive with the score of 4.37 for synthesized speech with ground-truth phrase breaks. Kosuke Futamata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
Interspeech | 4 |
| 2020 | Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image
Shunsuke Goto, Kotaro Onishi, Yuki Saito 0001, Kentaro Tachibana, Koichiro Mori |
INTERSPEECH | 4 |
| 2018 | An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic FeaturesabstractAlthough a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis because the model size becomes too large to train with a consumer GPU. For a WaveNet vocoder with a sampling frequency of 48 kHz with a consumer GPU, this paper introduces a subband WaveNet architecture to a speaker-dependent WaveNet vocoder and proposes a subband WaveNet vocoder. In experiments, each conditional subband WaveNet with a sampling frequency of 8 kHz was well trained using a consumer GPU. The results of subjective evaluations with a Japanese male speech corpus indicate that the proposed subband WaveNet vocoder with 36-dimensional simple acoustic features significantly outperformed the conventional source-filter model-based vocoders including STRAIGHT with 86-dimensional features. Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 2 |
| 2018 | An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech GenerationabstractWe propose a noise shaping method to improve the sound quality of speech signals generated by WaveNet, which is a convolutional neural network (CNN) that predicts a waveform sample sequence as a discrete symbol sequence. Speech signals generated by WaveNet often suffer from noise signals caused by the quantization error generated by representing waveform samples as discrete symbols and the prediction error of the CNN. We analyze these noise signals and show that 1) since the prediction error is much larger than the quantization error, the effect of the quantization error on the noise signals is practically negligible, and 2) noise signals tend to cause large spectral distortion in a high-frequency band. To alleviate the adverse effect of these noise signals on the generated speech signals, the proposed noise shaping method applies a perceptual weighting filter to WaveNet, making it possible to use the frequency masking properties of the human auditory system. We conducted objective and subjective evaluations to investigate the effectiveness of the proposed method and demonstrated that it significantly improved the sound quality of the generated speech signals. Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 1 |
| 2017 | Subband wavenet with overlapped single-sideband filterbanksabstractCompared with conventional vocoders, deep neural network-based raw audio generative models, such as WaveNet and SampleRNN, can more naturally synthesize speech signals, although the synthesis speed is a problem, especially with high sampling frequency. This paper provides subband WaveNet based on multirate signal processing for high-speed and high-quality synthesis with raw audio generative models. In the training stage, speech waveforms are decomposed and decimated into subband short waveforms with a low sampling rate, and each subband WaveNet network is trained using each subband stream. In the synthesis stage, each generated signal is up-sampled and integrated into a fullband speech signal. The results of objective and subjective experiments for unconditional WaveNet with a sampling frequency of 32 kHz indicate that the proposed subband WaveNet with a square-root Hann window-based overlapped 9-channel single-sideband filterbank can realize about four times the synthesis speed and improve the synthesized speech quality more than the conventional fullband WaveNet. Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ASRU | 2 |
| 2016 | Model Integration for HMM- and DNN-Based Speech Synthesis Using Product-of-Experts Framework
Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 1 |
| 2009 | Source adaptive blind signal extraction using closed-form ICA for hands-free robot spoken dialogue systemabstractIn this paper, we propose a new ICA-based BSS algorithm including estimation of sources' probability density functions (PDFs) to adapt the nonlinear activation function to various noise conditions. In the proposed method, closed-form second-order ICA is introduced as a computational-cost-efficient preprocessing to extract sources' PDFs, which is beneficial for real-time application. Compared with various type of conventional ICAs, e.g., fixed activation-function type and ML-based type, our proposed algorithm can give a faster and higher convergence. Based on the proposed source-adaptive ICA, we show a real-time noise reduction results under diffuse noise environment. Also we can demonstrate our recently developed hands-free robot spoken dialogue system via real-time ICA. Yu Takahashi, Hiroshi Saruwatari, Yuki Fujihara, Kentaro Tachibana, Yoshimitsu Mori, Shigeki Miyabe, Kiyohiro Shikano, Akira Tanaka |
ICASSP | 4 |
| 2007 | Efficient Blind Source Separation Combining Closed-Form Second-Order ICA and Nonclosed-Form Higher-Order ICAabstractIn this paper, first, we propose a computational-cost efficient blind source separation combining closed-form 2nd-order independent component analysis (ICA) and nonclosed-form higher-order ICA. The closed-form solution of the 2nd-order ICA has been recently presented by one of the authors. This finding motivates us to combine the closed-form 2nd-order ICA and higher-order ICA, where the preceding closed-form ICA produces a good initial value and the following higher-order ICA updates the separation filters from the advantageous status. Secondly, we utilize the proposed architecture to address an essential question that which type of statistics is more beneficial to ICA among non-stationarity and non-Gaussianity. This can be conducted owing to the attractive property that the closed-form ICA can provide a good estimate of the theoretical upper limitation of the separation performance among 2nd-order ICAs without suffering from poor-convergence problems. Experimental results reveal that the non-Gaussianity-based ICA can outperform the non-stationarity-based ICA. Kentaro Tachibana, Hiroshi Saruwatari, Yoshimitsu Mori, Shigeki Miyabe, Kiyohiro Shikano, Akira Tanaka |
ICASSP (1) | 1 |