EDBT 2026 Demo / reviewers in the wild / expert
Ye Jia
dblp:217/2520
· DBLP profile ↗
28ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-0457-8083ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SimulTron: On-Device Simultaneous Speech to Speech TranslationabstractSimultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST architecture designed to tackle this task. SimulTron is a lightweight direct S2ST model that uses the strengths of the Translatotron framework while incorporating key modifications for streaming operation, and an adjustable fixed delay. Our experiments show that SimulTron surpasses Translatotron 2 in offline evaluations. Furthermore, real-time evaluations reveal that SimulTron improves upon the performance achieved by Translatotron 1. Additionally, SimulTron achieves superior BLEU scores and latency compared to previous real-time S2ST method on the MuST-C dataset. Significantly, we have successfully deployed SimulTron on a Pixel 7 Pro device, show its potential for simultaneous S2ST on-device. Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding 0004, Ye Jia, Nadav Bar, Heiga Zen, Michelle Tadmor Ramanovich |
ICASSP | 5 |
| 2025 | HiBench: Benchmarking LLMs Capability on Hierarchical Structure ReasoningabstractStructure reasoning is a fundamental capability of large language models (LLMs), enabling them to reason about structured commonsense and answer multi-hop questions. However, existing benchmarks for structure reasoning mainly focus on horizontal and coordinate structures (e.g. graphs), overlooking the hierarchical relationships within them. Hierarchical structure reasoning is crucial for human cognition, particularly in memory organization and problem-solving. It also plays a key role in various real-world tasks, such as information extraction and decision-making. To address this gap, we propose HiBench, the first framework designed to systematically benchmark the hierarchical reasoning capabilities of LLMs from initial structure generation to final proficiency assessment. It encompasses six representative scenarios, covering both fundamental and practical aspects, and consists of 30 tasks with varying hierarchical complexity, totaling 39,519 queries. To evaluate LLMs comprehensively, we develop five capability dimensions that depict different facets of hierarchical structure understanding. Through extensive evaluation of 20 LLMs from 10 model families, we reveal key insights into their capabilities and limitations: 1) existing LLMs show proficiency in basic hierarchical reasoning tasks; 2) they still struggle with more complex structures and implicit hierarchical representations, especially in structural modification and textual reasoning. Based on these findings, we create a small yet well-designed instruction dataset, which enhances LLMs' performance on HiBench by an average of 88.84% (Llama-3.1-8B) and 31.38% (Qwen2.5-7B) across all tasks. The HiBench dataset and toolkit are available at https://github.com/jzzzzh/HiBench to encourage evaluation. Zhuohang Jiang, Pangjing Wu, Ziran Liang, Peter Q. Chen, Xu Yuan 0007, Ye Jia, Jiancheng Tu, Chen Li 0023, Peter Hiu Fung Ng, Qing Li 0001 |
KDD (2) | 6 |
| 2025 | Can You Move These Over There? Exploring an LLM-based VR Mover to Support Natural Multi-object Manipulation
Xiangzhi Eric Wang, Zackary P. T. Sin, Ye Jia, Dan Archer 0001, Wynonna H. Y. Fong, Qing Li 0001, Chen Li 0023 |
UIST | 3 |
| 2025 | Traceable teleportation: Improving spatial learning in virtual locomotion
Ye Jia, Zackary P. T. Sin, Chen Li 0023, Peter Hiu Fung Ng, Xiao Huang 0001, George Baciu, Jiannong Cao 0001, Qing Li 0001 |
Int. J. Hum. Comput. Stud. | 1 |
| 2024 | illumotion: An Optical-illusion-based VR Locomotion Technique for Long-Distance 3D MovementabstractLocomotion has a marked impact on user experience in VR, but currently, common to-go techniques such as steering and teleportation have their limitations. Particularly, steering is prone to cybersickness, while teleportation trades presence for mitigating cybersickness. Inspired by how we manipulate a picture on a mobile phone, we propose illumotion, an optical-illusion-based method that, we believe, can provide an alternative to these two typical techniques. Instead of zooming in a picture by pinching two fingers, we can move forward by “zooming” toward part of the 3D virtual scene with pinched hands. Not only is the proposed technique easy to use, it also seems to minimize cybersickness to some degree.illumotion relies on the manipulation of optics; as such, it requires solving motion parameters in screen space and a model of how we perceive depth. To evaluate it, a comprehensive user study with 66 users was conducted. Results show that, compared with either teleportation, steering or both, illumotion has better performance, presence, usability, user experience and cybersickness alleviation. We believe the result is a clear indication that our novel optically-driven method is a promising candidate for generalized locomotion. Zackary P. T. Sin, Ye Jia, Chen Li 0023, Hong Va Leong, Qing Li 0001, Peter Hiu Fung Ng |
VR | 2 |
| 2023 | Textless Direct Speech-to-Speech Translation with Discrete Speech RepresentationabstractResearch on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade systems, which are often composed of recognition, translation and synthesis sub-systems. However, most of end-to-end systems still rely on intermediate textual supervision during training, which makes it infeasible to work for languages without written forms. In this work, we propose a novel model, Textless Translatotron, which is based on Translatotron 2 [1], for training an end-to-end direct S2ST model without any textual supervision. Instead of jointly training with an auxiliary task predicting target phonemes as in Translatotron 2, the proposed model uses an auxiliary task predicting discrete speech representations which are obtained from learned or random speech quantizers. When a speech encoder pre-trained with unsupervised speech data is used for both models, the proposed model obtains translation quality nearly on-par with Translatotron 2 on the multilingual CVSS-C corpus [2] as well as the bilingual Fisher Spanish-English corpus [3]. On the latter, it outperforms the prior state-of-the-art textless model by +18.5 BLEU. Ye Jia, Chung-Cheng Chiu |
ICASSP | 2 |
| 2023 | Speech Aware Dialog System Technology Challenge (DSTC11)
Hagen Soltau, Izhak Shafran, Mingqiu Wang, Abhinav Rastogi, Jeffrey Zhao, Ye Jia, Wei Han 0002, Yuan Cao 0007, Aramys Miranda |
INTERSPEECH | 6 |
| 2022 | More than Words: In-the-Wild Visually-Driven Prosody for Text-to-SpeechabstractIn this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes ad-vantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate how this allows VDTTS to, unlike plain TTS models, generate speech that not only has prosodic variations like natural pauses and pitch, but is also synchronized to the input video. Experimentally, we show our model produces well-synchronized outputs, approaching the video-speech synchronization quality of the ground-truth, on several challenging benchmarks including “in-the-wild” content from VoxCeleb2. Supplementary demo videos demonstrating video-speech synchronization, robustness to speaker ID swapping, and prosody, presented at the project page.11Project page: http://google-research.github.io/lingvo-lab/vdtts Michael Hassid, Michelle Tadmor Ramanovich, Brendan Shillingford, Miaosen Wang, Ye Jia, Tal Remez |
CVPR | 5 |
| 2022 | Translatotron 2: High-quality direct speech-to-speech translation with voice preservationabstractWe present Translatotron 2, a neural direct speech-to-speech translation model that can be trained end-to-end. Translatotron 2 consists of a speech encoder, a linguistic decoder, an acoustic synthesizer, and a single attention module that connects them together. Experimental results on three datasets consistently show that Translatotron 2 outperforms the original Translatotron by a large margin on both translation quality (up to +15.5 BLEU) and speech generation quality, and approaches the same of cascade systems. In addition, we propose a simple method for preserving speakers’ voices from the source speech to the translation speech in a different language. Unlike existing approaches, the proposed method is able to preserve each speaker’s voice on speaker turns without requiring for speaker segmentation. Furthermore, compared to existing approaches, it better preserves speaker’s privacy and mitigates potential misuse of voice cloning for creating spoofing audio artifacts. Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, Roi Pomerantz |
ICML | 1 |
| 2022 | XTREME-S: Evaluating Cross-lingual Speech RepresentationsabstractWe introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages.XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval.Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning.This paper describes the new benchmark and establishes the first speech-only and speechtext baselines using XLS-R and mSLAM on all downstream tasks.We motivate the design choices and detail how to use the benchmark.Datasets and fine-tuning scripts are made easily accessible through the HuggingFace platform. 1 Alexis Conneau, Ankur Bapna, Yu Zhang 0033, Patrick von Platen, Anton Lozhkov, Colin Cherry, Ye Jia, Clara Rivera, Mihir Kale, Daan van Esch, Vera Axelrod, Simran Khanuja, Jonathan H. Clark, Orhan Firat, Michael Auli, Sebastian Ruder, Jason Riesa, Melvin Johnson |
INTERSPEECH | 8 |
| 2022 | Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer TasksabstractTransfer tasks in text-to-speech (TTS) synthesis -where one or more aspects of the speech of one set of speakers is transferred to another set of speakers that do not feature these aspects originally -remains a challenging task.One of the challenges is that models that have high-quality transfer capabilities can have issues in stability, making them impractical for user-facing critical tasks.This paper demonstrates that transfer can be obtained by training a robust TTS system on data generated by a less robust TTS system designed for a high-quality transfer task; in particular, a CHiVE-BERT monolingual TTS system is trained on the output of a Tacotron model designed for accent transfer.While some quality loss is inevitable with this approach, experimental results show that the models trained on synthetic data this way can produce high quality audio displaying accent transfer, while preserving speaker characteristics such as speaking style. Lev Finkelstein, Heiga Zen, Norman Casagrande, Chun-an Chan, Ye Jia, Tom Kenter, Alexey Petelin, Jonathan Shen, Vincent Wan, Yu Zhang 0033, Rob Clark |
INTERSPEECH | 5 |
| 2022 | Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translationabstractEnd-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research.Recent works have demonstrated that the performance of such direct S2ST systems is approaching that of conventional cascade S2ST when trained on comparable datasets.However, in practice, the performance of direct S2ST is bounded by the availability of paired S2ST training data.In this work, we explore multiple approaches for leveraging much more widely available unsupervised and weakly-supervised speech and text data to improve the performance of direct S2ST based on Translatotron 2. With our most effective approaches, the average translation quality of direct S2ST on 21 language pairs on the CVSS-C corpus is improved by +13.6 BLEU (or +113% relatively), as compared to the previous state-of-the-art trained without additional data.The improvements on low-resource language are even more significant (+398% relatively on average).Our comparative studies suggest future research directions for S2ST and speech representation learning. Ye Jia, Yifan Ding 0004, Ankur Bapna, Colin Cherry, Yu Zhang 0033, Alexis Conneau, Nobuyuki Morioka |
INTERSPEECH | 1 |
| 2022 | CVSS Corpus and Massively Multilingual Speech-to-Speech TranslationabstractWe introduce CVSS, a massively multilingual-to-English speech-to-speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English. CVSS is derived from the Common Voice speech corpus and the CoVoST 2 speech-to-text translation (ST) corpus, by synthesizing the translation text from CoVoST 2 into speech using state-of-the-art TTS systems. Two versions of translation speech in English are provided: 1) CVSS-C: All the translation speech is in a single high-quality canonical voice; 2) CVSS-T: The translation speech is in voices transferred from the corresponding source speech. In addition, CVSS provides normalized translation text which matches the pronunciation in the translation speech. On each version of CVSS, we built baseline multilingual direct S2ST models and cascade S2ST models, verifying the effectiveness of the corpus. To build strong cascade S2ST baselines, we trained an ST model on CoVoST 2, which outperforms the previous state-of-the-art trained on the corpus without extra data by 5.8 BLEU. Nevertheless, the performance of the direct S2ST models approaches the strong cascade baselines when trained from scratch, and with only 0.1 or 0.7 BLEU difference on ASR transcribed translation when initialized from matching ST models. Ye Jia, Michelle Tadmor Ramanovich, Heiga Zen |
LREC | 1 |
| 2021 | Textual Echo CancellationabstractIn this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo11TTS playback denotes the synthesized TTS voice, and TTS playback echo denotes the reverberated TTS voice captured by the microphone. from overlapping speech recordings. Such a system can largely im-prove speech recognition performance and user experience for in-telligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS play-back as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to en-hancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC). Shaojin Ding, Ye Jia |
ASRU | 2 |
| 2021 | Parallel Tacotron: Non-Autoregressive and Controllable TTSabstractAlthough neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness. This paper proposes a non-autoregressive neural text-to-speech model augmented with a variational autoencoder-based residual encoder. This model, called Parallel Tacotron, is highly parallelizable during both training and inference, allowing efficient synthesis on modern parallel hardware. The use of the variational autoencoder relaxes the one-to-many mapping nature of the text-to-speech problem and improves naturalness. To further improve the naturalness, we use lightweight convolutions, which can efficiently capture local contexts, and introduce an iterative spectrogram loss inspired by iterative refinement. Experimental results show that Parallel Tacotron matches a strong autoregressive baseline in subjective evaluations with significantly decreased inference time. Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang 0033, Ye Jia, Ron J. Weiss |
ICASSP | 5 |
| 2021 | Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration ModelingabstractThis paper introduces Parallel Tacotron 2, a non-autoregressive neural text-to-speech model with a fully differentiable duration model which does not require supervised duration signals.The duration model is based on a novel attention mechanism and an iterative reconstruction loss based on Soft Dynamic Time Warping, this model can learn token-frame alignments as well as token durations automatically.Experimental results show that Parallel Tacotron 2 outperforms baselines in subjective naturalness in several diverse multi speaker evaluations. Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang 0033, Ye Jia, R. J. Skerry-Ryan |
Interspeech | 5 |
| 2021 | PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTSabstractThis paper introduces PnG BERT, a new encoder model for neural TTS.This model is augmented from the original BERT model, by taking both phoneme and grapheme representations of text as input, as well as the word-level alignment between them.It can be pre-trained on a large text corpus in a selfsupervised manner, and fine-tuned in a TTS task.Experimental results show that a neural TTS model using a pre-trained PnG BERT as its encoder yields more natural prosody and more accurate pronunciation than a baseline model using only phoneme input with no pre-training.Subjective side-by-side preference evaluations show that raters have no statistically significant preference between the speech synthesized using a PnG BERT and ground truth recordings from professional speakers. Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang 0033 |
Interspeech | 1 |
| 2020 | Improved Noisy Student Training for Automatic Speech RecognitionabstractRecently, a semi-supervised learning method known as "noisy student training" has been shown to improve image classification performance of deep networks significantly. Noisy student training is an iterative self-training method that leverages augmentation to improve network performance. In this work, we adapt and improve noisy student training for automatic speech recognition, employing (adaptive) SpecAugment as the augmentation method. We find effective methods to filter, balance and augment the data generated in between self-training iterations. By doing so, we are able to obtain word error rates (WERs) 4.2%/8.6% on the clean/noisy LibriSpeech test sets by only using the clean 100h subset of LibriSpeech as the supervised set and the rest (860h) as the unlabeled set. Furthermore, we are able to achieve WERs 1.7%/3.4% on the clean/noisy LibriSpeech test sets by using the unlab-60k subset of LibriLight as the unlabeled set for LibriSpeech 960h. We are thus able to improve upon the previous state-of-the-art clean/noisy test WERs achieved on LibriSpeech 100h (4.74%/12.20%) and LibriSpeech (1.9%/4.1%). Daniel S. Park, Yu Zhang 0033, Ye Jia, Wei Han 0002, Chung-Cheng Chiu, Bo Li 0028, Quoc V. Le |
INTERSPEECH | 3 |
| 2019 | Speech Recognition with Augmented Synthesized SpeechabstractRecent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific, human speech that is used to train speech recognizers. The multi-speaker speech synthesis architecture can learn latent embedding spaces of prosody, speaker and style variations derived from input acoustic representations thereby allowing for manipulation of the synthesized speech. In this paper, we evaluate the feasibility of enhancing speech recognition performance using speech synthesis using two corpora from different domains. We explore algorithms to provide the necessary acoustic and lexical diversity needed for robust speech recognition. Finally, we demonstrate the feasibility of this approach as a data augmentation strategy for domain-transfer. We find that improvements to speech recognition performance is achievable by augmenting training data with synthesized material. However, there remains a substantial gap in performance between recognizers trained on human speech those trained on synthesized speech. Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Ye Jia, Pedro J. Moreno 0001, Zelin Wu |
ASRU | 4 |
| 2019 | Leveraging Weakly Supervised Data to Improve End-to-end Speech-to-text TranslationabstractEnd-to-end Speech Translation (ST) models have many potential advantages when compared to the cascade of Automatic Speech Recognition (ASR) and text Machine Translation (MT) models, including lowered inference latency and the avoidance of error compounding. However, the quality of end-to-end ST is often limited by a paucity of training data, since it is difficult to collect large parallel corpora of speech and translated transcript pairs. Previous studies have proposed the use of pre-trained components and multi-task learning in order to benefit from weakly supervised training data, such as speech-to-transcript or text-to-foreign-text pairs. In this paper, we demonstrate that using pre-trained MT or text-to-speech (TTS) synthesis models to convert weakly supervised data into speech-to-translation pairs for ST training can be more effective than multi-task learning. Furthermore, we demonstrate that a high quality end-to-end ST model can be trained using only weakly supervised datasets, and that synthetic data sourced from unlabeled monolingual text or speech can be used to improve performance. Finally, we discuss methods for avoiding overfitting to synthetic speech with a quantitative ablation study. Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao 0007, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo |
ICASSP | 1 |
| 2019 | Hierarchical Generative Modeling for Controllable Speech Synthesis
Wei-Ning Hsu, Yu Zhang 0033, Ron J. Weiss, Heiga Zen, Yuxuan Wang 0002, Yuan Cao 0007, Ye Jia, Jonathan Shen, Patrick Nguyen, Ruoming Pang |
ICLR (Poster) | 8 |
| 2019 | Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech SeparationabstractWe describe Parrotron, an end-to-end-trained speech-to-speech conversion model that maps an input spectrogram directly to another spectrogram, without utilizing any intermediate discrete representation.The network is composed of an encoder, spectrogram and phoneme decoders, followed by a vocoder to synthesize a time-domain waveform.We demonstrate that this model can be trained to normalize speech from any speaker regardless of accent, prosody, and background noise, into the voice of a single canonical target speaker with a fixed accent and consistent articulation and prosody.We further show that this normalization model can be adapted to normalize highly atypical speech from a deaf speaker, resulting in significant improvements in intelligibility and naturalness, measured via a speech recognizer and listening tests.Finally, demonstrating the utility of this model on other speech tasks, we show that the same model architecture can be trained to perform a speech separation task. Fadi Biadsy, Ron J. Weiss, Pedro J. Moreno 0001, Dimitri Kanvesky, Ye Jia |
INTERSPEECH | 5 |
| 2019 | Direct Speech-to-Speech Translation with a Sequence-to-Sequence ModelabstractWe present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an intermediate text representation.The network is trained end-to-end, learning to map speech spectrograms into target spectrograms in another language, corresponding to the translated content (in a different canonical voice).We further demonstrate the ability to synthesize translated speech using the voice of the source speaker.We conduct experiments on two Spanish-to-English speech translation datasets, and find that the proposed model slightly underperforms a baseline cascade of a direct speech-to-text translation model and a text-to-speech synthesis model, demonstrating the feasibility of the approach on this very challenging task. Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson |
INTERSPEECH | 1 |
| 2019 | VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram MaskingabstractIn this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker.We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings;(2) A spectrogram masking network that takes both noisy spectrogram and speaker embedding as input, and produces a mask.Our system significantly reduces the speech recognition WER on multi-speaker signals, with minimal WER degradation on single-speaker signals. Hannah Muckenhirn, Kevin W. Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, Ignacio López-Moreno |
INTERSPEECH | 9 |
| 2019 | LibriTTS: A Corpus Derived from LibriSpeech for Text-to-SpeechabstractThis paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use.It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic speech recognition systems.The new corpus inherits desired properties of the LibriSpeech corpus while addressing a number of issues which make LibriSpeech less than ideal for text-to-speech work.The released corpus consists of 585 hours of speech data at 24kHz sampling rate from 2,456 speakers and the corresponding texts.Experimental results show that neural end-to-end TTS models trained from the LibriTTS corpus achieved above 4.0 in mean opinion scores in naturalness in five out of six evaluation speakers. Heiga Zen, Viet Dang, Rob Clark, Yu Zhang 0033, Ron J. Weiss, Ye Jia |
INTERSPEECH | 6 |
| 2019 | Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice CloningabstractWe present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish speech using an English speaker's voice, without training on any bilingual or parallel examples.Such transfer works across distantly related languages, e.g.English and Mandarin.Critical to achieving this result are: 1. using a phonemic input representation to encourage sharing of model capacity across languages, and 2. incorporating an adversarial loss term to encourage the model to disentangle its representation of speaker identity (which is perfectly correlated with language in the training data) from the speech content.Further scaling up the model by training on multiple speakers of each language, and incorporating an autoencoding input to help stabilize attention during training, results in a model which can be used to consistently synthesize intelligible speech for training speakers in all languages seen during training, and in native or foreign accents. Yu Zhang 0033, Ron J. Weiss, Heiga Zen, R. J. Skerry-Ryan, Ye Jia, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 7 |
| 2018 | Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech SynthesisabstractIn this work, we propose “global style tokens” (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The embeddings are trained with no explicit labels, yet learn to model a large range of acoustic expressiveness. GSTs lead to a rich set of significant results. The soft interpretable “labels” they generate can be used to control synthesis in novel ways, such as varying speed and speaking style – independently of the text content. They can also be used for style transfer, replicating the speaking style of a single audio clip across an entire long-form text corpus. When trained on noisy, unlabeled found data, GSTs learn to factorize noise and speaker identity, providing a path towards highly scalable but robust speech synthesis. Yuxuan Wang 0002, Daisy Stanton, Yu Zhang 0033, R. J. Skerry-Ryan, Eric Battenberg, Joel Shor, Ye Jia, Rif A. Saurous |
ICML | 8 |
| 2018 | Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech SynthesisabstractWe describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently trained components: (1) a speaker encoder network, trained on a speaker verification task using an independent dataset of noisy speech from thousands of speakers without transcripts, to generate a fixed-dimensional embedding vector from seconds of reference speech from a target speaker; (2) a sequence-to-sequence synthesis network based on Tacotron 2, which generates a mel spectrogram from text, conditioned on the speaker embedding; (3) an auto-regressive WaveNet-based vocoder that converts the mel spectrogram into a sequence of time domain waveform samples. We demonstrate that the proposed model is able to transfer the knowledge of speaker variability learned by the discriminatively-trained speaker encoder to the new task, and is able to synthesize natural speech from speakers that were not seen during training. We quantify the importance of training the speaker encoder on a large and diverse speaker set in order to obtain the best generalization performance. Finally, we show that randomly sampled speaker embeddings can be used to synthesize speech in the voice of novel speakers dissimilar from those used in training, indicating that the model has learned a high quality speaker representation. Ye Jia, Yu Zhang 0033, Ron J. Weiss, Jonathan Shen, Patrick Nguyen, Ruoming Pang, Ignacio López-Moreno |
NeurIPS | 1 |