EDBT 2026 Demo / reviewers in the wild / expert
Wataru Nakata
dblp:312/5142
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-3953-6534ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language ModelingabstractSpoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications. Wataru Nakata, Kentaro Seki, Hitomi Yanaka, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
LREC | 1 |
| 2026 | DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue AudioabstractFull-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference. Wataru Nakata, Yuki Saito 0001, Kazuki Yamauchi, Emiru Tsunoo, Hiroshi Saruwatari |
SIGDIAL | 1 |
| 2026 | Speaker-conditioned phrase break prediction for text-to-speech with phoneme-level pre-trained language modelabstractThis paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We further demonstrate that these speaker embeddings can capture speaker-related characteristics solely from the phrasing task. Besides, we explore the potential of pre-trained speaker embeddings for unseen speakers through a few-shot adaptation method. Furthermore, we pioneer the application of phoneme-level pre-trained language models to this TTS front-end task, which significantly boosts the accuracy of the phrasing model. Our methods are rigorously assessed through both objective and subjective evaluations, demonstrating their effectiveness. • Speaker-conditioned phrasing model improves accuracy in multi-speaker phrasing tasks. • We explore various speaker embeddings in phrasing models. • We apply phoneme-level pre-trained language models to enhance phrasing accuracy. • We propose a speaker adaptation method for few-shot phrasing tasks. • We verify that speaker embeddings learn human-aligned features via phrasing tasks. Yuki Saito 0001, Takaaki Saeki, Tomoki Koriyama, Wataru Nakata, Detai Xin, Hiroshi Saruwatari |
Speech Commun. | 5 |
| 2025 | Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent LayerabstractWe introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model. Go Nishikawa, Wataru Nakata, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari, Tomohiko Nakamura |
ASRU | 2 |
| 2025 | Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning FeaturesabstractReal-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features of self-supervised learning (SSL) models. This work is the first to incorporate SSL features with causality into an SE model. The causal SSL features are encoded and combined with spectrogram features using feature-wise linear modulation to estimate a mask for enhancing the noisy input speech. Simultaneously, we quantize the causal SSL features using vector quantization to represent phonetic characteristics as semantic tokens. The model not only encodes SSL features but also predicts the future semantic tokens in multi-task learning (MTL). The experimental results using VoiceBank + DEMAND dataset show that our proposed method achieves 2.88 in PESQ, especially with semantic prediction MTL, in which we confirm that the semantic prediction played an important role in causal SE. Emiru Tsunoo, Yuki Saito 0001, Wataru Nakata, Hiroshi Saruwatari |
ICASSP | 3 |
| 2024 | The T05 System for the voicemos challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic SpeechabstractWe present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-quality synthetic speech. In addition to a pretrained self-supervised learning (SSL)-based speech feature extractor, our system incorporates a pretrained image feature extractor to capture the difference of synthetic speech observed in speech spectrograms. We first separately train two MOS predictors that use either of an SSL-based or spectrogram-based feature. Then, we fine-tune the two predictors for better MOS prediction using the fusion of two extracted features. In the VMC 2024 Track 1, our T05 system achieved first place in 7 out of 16 evaluation metrics and second place in the remaining 9 metrics, with a significant difference compared to those ranked third and below. We also report the results of our ablation study to investigate essential factors of our system. Kaito Baba, Wataru Nakata, Yuki Saito 0001, Hiroshi Saruwatari |
SLT | 2 |
| 2023 | COCO-NUT: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-Based ControlabstractIn text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as text-to-image, free-form text instruction should be useful for intuitive and complicated control of voice characteristics. A sufficiently large corpus of high-quality and diverse voice samples with corresponding free-form descriptions can advance such control research. However, neither an open corpus nor a scalable method is currently available. To this end, we develop Coco-Nut, a new corpus including diverse Japanese utterances, along with text transcriptions and free-form voice characteristics descriptions. Our methodology to construct this corpus consists of 1) automatic collection of voice-related audio data from the Internet, 2) quality assurance, and 3) manual annotation using crowdsourcing. Additionally, we benchmark our corpus on the prompt embedding model trained by contrastive speech-text learning. Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Wataru Nakata, Detai Xin, Hiroshi Saruwatari |
ASRU | 4 |
| 2022 | Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 1 |
| 2022 | UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022abstractWe present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022.The challenge is to predict the MOS values of speech samples collected from previous Blizzard Challenges and Voice Conversion Challenges for two tracks: a main track for in-domain prediction and an out-of-domain (OOD) track for which there is less labeled data from different listening tests.Our system is based on ensemble learning of strong and weak learners.Strong learners incorporate several improvements to the previous finetuning models of self-supervised learning (SSL) models, while weak learners use basic machine-learning methods to predict scores from SSL features.In the Challenge, our system had the highest score on several metrics for both the main and OOD tracks.In addition, we conducted ablation studies to investigate the effectiveness of our proposed methods. Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Hiroshi Saruwatari |
INTERSPEECH | 3 |
| 2022 | J-MAC: Japanese multi-speaker audiobook corpus for speech synthesisabstractIn this paper, we construct a Japanese audiobook speech corpus called "J-MAC" for speech synthesis research.With the success of reading-style speech synthesis, the research target is shifting to tasks that use complicated contexts.Audiobook speech synthesis is a good example that requires cross-sentence, expressiveness, etc.Unlike reading-style speech, speaker-specific expressiveness in audiobook speech also becomes the context.To enhance this research, we propose a method of constructing a corpus from audiobooks read by professional speakers.From many audiobooks and their texts, our method can automatically extract and refine the data without any language dependency.Specifically, we use vocal-instrumental separation to extract clean data, connectionist temporal classification to roughly align text and audio, and voice activity detection to refine the alignment.J-MAC is open-sourced in our project page.We also conduct audiobook speech synthesis evaluations, and the results give insights into audiobook speech synthesis. Shinnosuke Takamichi, Wataru Nakata, Naoko Tanji, Hiroshi Saruwatari |
INTERSPEECH | 2 |