Nobuyuki Morioka

dblp:39/3539 · also Nobu Morioka · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
abstract
Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages.
Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov
ICASSP3
2023 E3 TTS: Easy End-to-End Diffusion-Based Text To Speech
abstract
We propose Easy End-to-End Diffusion-based Text to Speech, a simple and efficient end-to-end text-to-speech model based on diffusion. E3 TTS directly takes plain text as input and generates an audio waveform through an iterative refinement process. Unlike many prior work, E3 TTS does not rely on any intermediate representations like spectrogram features or alignment information. Instead, E3 TTS models the temporal structure of the waveform through the diffusion process. Without relying on additional conditioning information, E3 TTS could support flexible latent structure within the given audio. This enables E3 TTS to be easily adapted for zero-shot tasks such as editing without any additional training. Experiments show that E3 TTS can generate high-fidelity audio, approaching the performance of a state-of-the-art neural TTS system. Audio samples are available at https://e3tts.github.io.
Yuan Gao 0046, Nobuyuki Morioka, Yu Zhang 0033, Nanxin Chen
ASRU2
2023 Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-Speech
abstract
This paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty to scale multilingual TTS to hundreds of languages is collecting high-quality speech–text paired data in low-resource languages. This study extends Maestro, a speech–text joint pretraining framework for automatic speech recognition (ASR), to speech generation tasks. To train a TTS model from various types of speech and text data, different training schemes are designed to handle supervised (paired TTS and ASR data) and unsupervised (untranscribed speech and unspoken text) datasets. Experimental evaluation shows that 1) multilingual TTS models trained on Virtuoso can achieve significantly better naturalness and intelligibility than baseline ones in seen languages, and 2) they can synthesize reasonably intelligible and naturally sounding speech for unseen languages where no high-quality paired TTS data is available.
Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka, Gary Wang, Yu Zhang 0033, Ankur Bapna, Andrew Rosenberg, Bhuvana Ramabhadran
ICASSP4
2023 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna
INTERSPEECH6
2022 Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation
abstract
End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research.Recent works have demonstrated that the performance of such direct S2ST systems is approaching that of conventional cascade S2ST when trained on comparable datasets.However, in practice, the performance of direct S2ST is bounded by the availability of paired S2ST training data.In this work, we explore multiple approaches for leveraging much more widely available unsupervised and weakly-supervised speech and text data to improve the performance of direct S2ST based on Translatotron 2. With our most effective approaches, the average translation quality of direct S2ST on 21 language pairs on the CVSS-C corpus is improved by +13.6 BLEU (or +113% relatively), as compared to the previous state-of-the-art trained without additional data.The improvements on low-resource language are even more significant (+398% relatively on average).Our comparative studies suggest future research directions for S2ST and speech representation learning.
Ye Jia, Yifan Ding 0004, Ankur Bapna, Colin Cherry, Yu Zhang 0033, Alexis Conneau, Nobuyuki Morioka
INTERSPEECH7
2021 Monotonic Kronecker-Factored Lattice
William Taylor Bakst, Nobuyuki Morioka, Erez Louidor
ICLR2
2020 Multidimensional Shape Constraints
abstract
We propose new multi-input shape constraints across four intuitive categories: complements, diminishers, dominance, and unimodality constraints. We show these shape constraints can be checked and even enforced when training machine-learned models for linear models, generalized additive models, and the nonlinear function class of multi-layer lattice models. Real-world experiments illustrate how the different shape constraints can be used to increase explainability and improve regularization, especially for non-IID train-test distribution shift.
Maya R. Gupta, Erez Louidor, Alexander Mangylov, Nobuyuki Morioka, Taman Narayan
ICML4
2011 Compact correlation coding for visual object categorization
abstract
Spatial relationships between local features are thought to play a vital role in representing object categories. However, learning a compact set of higher-order spatial features based on visual words, e.g., doublets and triplets, remains a challenging problem as possible combinations of visual words grow exponentially. While the local pairwise codebook achieves a compact codebook of pairs of spatially close local features without feature selection, its formulation is not scale invariant and is only suitable for densely sampled local features. In contrast, the proximity distribution kernel is a scale-invariant and robust representation capturing rich spatial proximity information between local features, but its representation grows quadratically in the number of visual words. Inspired by the two abovementioned techniques, this paper presents the compact correlation coding that combines the strengths of the two. Our method achieves a compact representation that is scaleinvariant and robust against object deformation. In addition, we adopt sparse coding instead of k-means clustering during the codebook construction to increase the discriminative power of our method. We systematically evaluate our method against both the local pairwise codebook and proximity distribution kernel on several challenging object categorization datasets to show performance improvements.
Nobuyuki Morioka, Shin'ichi Satoh 0001
ICCV1
2011 Robust visual reranking via sparsity and ranking constraints
abstract
Visual reranking has become a widely-accepted method to improve traditional text-based image search engines. Its basic principle is that visually similar images should have similar ranking scores. While existing methods are different in specifics, almost all of them are based on explicit or implicit pseudo-relevance feedback (PRF). Explicit PRF-based approaches, including classification-based and clustering-based reranking, suffer from the difficulty of selecting reliable positive and negative samples. Implicit PRF-based approaches, such as graph-based and Bayesian visual reranking, deal with such unreliability by making use of the initial ranking in a soft manner, but have limited capability of promoting relevant images and lowering down irrelevant images.
Nobuyuki Morioka, Jingdong Wang 0001
ACM Multimedia1
2011 Generalized Lasso based Approximation of Sparse Coding for Visual Recognition
abstract
Sparse coding, a method of explaining sensory data with as few dictionary bases as possible, has attracted much attention in computer vision. For visual object category recognition, L1 regularized sparse coding is combined with spatial pyramid representation to obtain state-of-the-art performance. However, because of its iterative optimization, applying sparse coding onto every local feature descriptor extracted from an image database can become a major bottleneck. To overcome this computational challenge, this paper presents "Generalized Lasso based Approximation of Sparse coding" (GLAS). By representing the distribution of sparse coefficients with slice transform, we fit a piece-wise linear mapping function with generalized lasso. We also propose an efficient post-refinement procedure to perform mutual inhibition between bases which is essential for an overcomplete setting. The experiments show that GLAS obtains comparable performance to L1 regularized sparse coding, yet achieves significant speed up demonstrating its effectiveness for large-scale visual recognition problems.
Nobuyuki Morioka, Shin'ichi Satoh 0001
NIPS1
2010 Learning Directional Local Pairwise Bases with Sparse Coding
abstract
Recently, sparse coding has been receiving much attention in object and scene recognition tasks because of its superiority in learning an effective codebook over k-means clustering. However, empirically, such codebook requires a relatively large number of visual words, essentially bases, to achieve high recognition accuracy. Therefore, due to the combinatorial explosion of visual words, it is infeasible to use this codebook to represent higher-order spatial features which are equally important in capturing distinct properties of scenes and objects. Contrasted with many previous techniques that exploit higher-order spatial features, Local Pairwise Codebook (LPC) is a simple and effective method to learn a compact set of clusters representing pairs of spatially close descriptors with k-means. Based on LPC, this paper proposes Directional Local Pairwise Bases (DLPB) that applies sparse coding to learn a compact set of bases capturing correlation between these descriptors, so to avoid the combinatorial explosion. Furthermore, such bases are learned for each quantized direction thereby explicitly adding directional information to the representation. We have evaluated DLPB with several challenging object and scene category datasets. Our experimental results show that DLPB outperforms the baselines across all datasets and achieves the state-of-the-art performance on some datasets.
Nobuyuki Morioka, Shin'ichi Satoh 0001
BMVC1
2010 Building Compact Local Pairwise Codebook with Joint Feature Space Clustering
Nobuyuki Morioka, Shin'ichi Satoh 0001
ECCV (1)1
2006 Enhancing Information Retrieval Using Problem Specific Knowledge
Nobuyuki Morioka, Ashesh Mahidadia
PKAW1