VLDB 2026 Research / reviewers in the wild / expert
Ryo Terashima
dblp:67/8451
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
Ryo Terashima, Yuma Shirahata, Masaya Kawamura |
INTERSPEECH | 1 |
| 2025 | Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
Reo Yoneyama, Masaya Kawamura, Ryo Terashima, Ryuichi Yamamoto, Tomoki Toda |
INTERSPEECH | 3 |
| 2024 | Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior EmbeddingabstractThis paper proposes a multilingual, multi-speaker (MM) TTS system by using a voice conversion (VC)-based data augmentation method. Creating an MM-TTS model is challenging, owing to the difficulties of collecting polyglot data from multiple speakers. To address this problem, we adopt a cross-lingual, multi-speaker VC model trained with multiple speakers’ monolingual databases. As this model effectively transfers acoustic attributes while retaining the content information, it is possible to generate each speaker’s polyglot corpora. Subsequently, we design the MM-TTS model with variational autoencoder (VAE)-based posterior embeddings. It is to be noted that incorporating VC-augmented polyglot corpora into the TTS training process might degrade synthetic quality, since the corpora sometimes contain unwanted artifacts. To mitigate this issue, the VAE is trained to capture the acoustic dissimilarity between the recorded and VC-augmented datasets. Through the selective choice of the posterior embeddings obtained from the original recordings in the training set, the proposed model enables the generation of acoustically clearer voices. Hyun-Wook Yoon, Jin-Seob Kim, Ryuichi Yamamoto, Ryo Terashima, Chan-Ho Song, Jae-Min Kim, Eunwoo Song |
ICASSP | 4 |
| 2024 | Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis FrameworkabstractWe propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization.Our proposed model converts song data with choral singing commonly contained in popular music and unsuitable for generating a simulated dataset to the solo singing data.This cleansing is based on NANSY++, which is a framework trained to reconstruct an input non-overlapped audio signal.We exploit the pretrained NANSY++ to convert choral singing into clean, nonoverlapped audio.This cleansing process mitigates the mislabeling of choral singing to solo singing and helps the effective training of EEND models even when the majority of available song data contains choral singing sections.We experimentally evaluated the EEND model trained with a dataset using our proposed method using annotated popular duet songs.As a result, our proposed method improved 14.8 points in diarization error rate. Hokuto Munakata, Ryo Terashima, Yusuke Fujita |
INTERSPEECH | 2 |
| 2024 | Hibikino-Musashi@Home RoboCup@Home DSPL Champion 2024
Akinobu Mizutani, Kosei Isomoto, Kosei Yamao, Ryohei Kobayashi 0003, Soma Fumoto, Koshun Arimura, Naoki Yamaguchi, Tomoya Shiba, Kouki Kimizuka, Yuta Ohno, Ryo Terashima, Hiromasa Yamaguchi, Tomoaki Fujino, Ryoga Maruno, Wataru Yoshimura, Kazuhito Mine, Tang Phu Thien Nhan, Yuga Yano, Yuichiro Tanaka, Takeshi Nishida, Takashi Morie, Hakaru Tamukoh |
RoboCup | 11 |
| 2023 | Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech SynthesisabstractSeveral fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate unstable pitch contour with audible artifacts when the dataset contains emotional attributes, i.e., large diversity of pronunciation and prosody. To address this problem, we propose Period VITS, a novel end-to-end TTS model that incorporates an explicit periodicity generator. In the proposed method, we introduce a frame pitch predictor that predicts prosodic features, such as pitch and voicing flags, from the input text. From these features, the proposed periodicity generator produces a sample-level sinusoidal source that enables the waveform decoder to accurately reproduce the pitch. Finally, the entire model is jointly optimized in an end-to-end manner with variational inference and adversarial objectives. As a result, the decoder becomes capable of generating more stable, expressive, and natural output waveforms. The experimental results showed that the proposed model significantly outperforms baseline models in terms of naturalness, with improved pitch stability in the generated samples. Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima, Jae-Min Kim, Kentaro Tachibana |
ICASSP | 4 |
| 2022 | Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data AugmentationabstractData augmentation via voice conversion (VC) has been successfully applied to low-resource expressive text-to-speech (TTS) when only neutral data for the target speaker are available. Although the quality of VC is crucial for this approach, it is challenging to learn a stable VC model because the amount of data is limited in low-resource scenarios, and highly expressive speech has large acoustic variety. To address this issue, we propose a novel data augmentation method that combines pitch-shifting and VC techniques. Because pitch-shift data augmentation enables the coverage of a variety of pitch dynamics, it greatly stabilizes training for both VC and TTS models, even when only 1,000 utterances of the target speaker's neutral data are available. Subjective test results showed that a FastSpeech 2-based emotional TTS system with the proposed method improved naturalness and emotional similarity compared with conventional methods. Ryo Terashima, Ryuichi Yamamoto, Eunwoo Song, Yuma Shirahata, Hyun-Wook Yoon, Jae-Min Kim, Kentaro Tachibana |
INTERSPEECH | 1 |