VLDB 2026 Research / reviewers in the wild / expert
Nobuyuki Morioka
dblp:39/3539 · also Nobu Morioka
· DBLP profile ↗
13ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed DataabstractCollecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages. Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov |
ICASSP | 3 |
| 2023 | E3 TTS: Easy End-to-End Diffusion-Based Text To SpeechabstractWe propose Easy End-to-End Diffusion-based Text to Speech, a simple and efficient end-to-end text-to-speech model based on diffusion. E3 TTS directly takes plain text as input and generates an audio waveform through an iterative refinement process. Unlike many prior work, E3 TTS does not rely on any intermediate representations like spectrogram features or alignment information. Instead, E3 TTS models the temporal structure of the waveform through the diffusion process. Without relying on additional conditioning information, E3 TTS could support flexible latent structure within the given audio. This enables E3 TTS to be easily adapted for zero-shot tasks such as editing without any additional training. Experiments show that E3 TTS can generate high-fidelity audio, approaching the performance of a state-of-the-art neural TTS system. Audio samples are available at https://e3tts.github.io. Yuan Gao 0046, Nobuyuki Morioka, Yu Zhang 0033, Nanxin Chen |
ASRU | 2 |
| 2023 | Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-SpeechabstractThis paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty to scale multilingual TTS to hundreds of languages is collecting high-quality speech–text paired data in low-resource languages. This study extends Maestro, a speech–text joint pretraining framework for automatic speech recognition (ASR), to speech generation tasks. To train a TTS model from various types of speech and text data, different training schemes are designed to handle supervised (paired TTS and ASR data) and unsupervised (untranscribed speech and unspoken text) datasets. Experimental evaluation shows that 1) multilingual TTS models trained on Virtuoso can achieve significantly better naturalness and intelligibility than baseline ones in seen languages, and 2) they can synthesize reasonably intelligible and naturally sounding speech for unseen languages where no high-quality paired TTS data is available. Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka, Gary Wang, Yu Zhang 0033, Ankur Bapna, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2023 | LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna |
INTERSPEECH | 6 |
| 2022 | Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translationabstractEnd-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research.Recent works have demonstrated that the performance of such direct S2ST systems is approaching that of conventional cascade S2ST when trained on comparable datasets.However, in practice, the performance of direct S2ST is bounded by the availability of paired S2ST training data.In this work, we explore multiple approaches for leveraging much more widely available unsupervised and weakly-supervised speech and text data to improve the performance of direct S2ST based on Translatotron 2. With our most effective approaches, the average translation quality of direct S2ST on 21 language pairs on the CVSS-C corpus is improved by +13.6 BLEU (or +113% relatively), as compared to the previous state-of-the-art trained without additional data.The improvements on low-resource language are even more significant (+398% relatively on average).Our comparative studies suggest future research directions for S2ST and speech representation learning. Ye Jia, Yifan Ding 0004, Ankur Bapna, Colin Cherry, Yu Zhang 0033, Alexis Conneau, Nobuyuki Morioka |
INTERSPEECH | 7 |
| 2021 | Monotonic Kronecker-Factored Lattice
William Taylor Bakst, Nobuyuki Morioka, Erez Louidor |
ICLR | 2 |
| 2020 | Multidimensional Shape ConstraintsabstractWe propose new multi-input shape constraints across four intuitive categories: complements, diminishers, dominance, and unimodality constraints. We show these shape constraints can be checked and even enforced when training machine-learned models for linear models, generalized additive models, and the nonlinear function class of multi-layer lattice models. Real-world experiments illustrate how the different shape constraints can be used to increase explainability and improve regularization, especially for non-IID train-test distribution shift. Maya R. Gupta, Erez Louidor, Alexander Mangylov, Nobuyuki Morioka, Taman Narayan |
ICML | 4 |
| 2011 | Compact correlation coding for visual object categorizationabstractSpatial relationships between local features are thought to play a vital role in representing object categories. However, learning a compact set of higher-order spatial features based on visual words, e.g., doublets and triplets, remains a challenging problem as possible combinations of visual words grow exponentially. While the local pairwise codebook achieves a compact codebook of pairs of spatially close local features without feature selection, its formulation is not scale invariant and is only suitable for densely sampled local features. In contrast, the proximity distribution kernel is a scale-invariant and robust representation capturing rich spatial proximity information between local features, but its representation grows quadratically in the number of visual words. Inspired by the two abovementioned techniques, this paper presents the compact correlation coding that combines the strengths of the two. Our method achieves a compact representation that is scaleinvariant and robust against object deformation. In addition, we adopt sparse coding instead of k-means clustering during the codebook construction to increase the discriminative power of our method. We systematically evaluate our method against both the local pairwise codebook and proximity distribution kernel on several challenging object categorization datasets to show performance improvements. Nobuyuki Morioka, Shin'ichi Satoh 0001 |
ICCV | 1 |
| 2011 | Robust visual reranking via sparsity and ranking constraintsabstractVisual reranking has become a widely-accepted method to improve traditional text-based image search engines. Its basic principle is that visually similar images should have similar ranking scores. While existing methods are different in specifics, almost all of them are based on explicit or implicit pseudo-relevance feedback (PRF). Explicit PRF-based approaches, including classification-based and clustering-based reranking, suffer from the difficulty of selecting reliable positive and negative samples. Implicit PRF-based approaches, such as graph-based and Bayesian visual reranking, deal with such unreliability by making use of the initial ranking in a soft manner, but have limited capability of promoting relevant images and lowering down irrelevant images. Nobuyuki Morioka, Jingdong Wang 0001 |
ACM Multimedia | 1 |
| 2011 | Generalized Lasso based Approximation of Sparse Coding for Visual RecognitionabstractSparse coding, a method of explaining sensory data with as few dictionary bases as possible, has attracted much attention in computer vision. For visual object category recognition, L1 regularized sparse coding is combined with spatial pyramid representation to obtain state-of-the-art performance. However, because of its iterative optimization, applying sparse coding onto every local feature descriptor extracted from an image database can become a major bottleneck. To overcome this computational challenge, this paper presents "Generalized Lasso based Approximation of Sparse coding" (GLAS). By representing the distribution of sparse coefficients with slice transform, we fit a piece-wise linear mapping function with generalized lasso. We also propose an efficient post-refinement procedure to perform mutual inhibition between bases which is essential for an overcomplete setting. The experiments show that GLAS obtains comparable performance to L1 regularized sparse coding, yet achieves significant speed up demonstrating its effectiveness for large-scale visual recognition problems. Nobuyuki Morioka, Shin'ichi Satoh 0001 |
NIPS | 1 |
| 2010 | Learning Directional Local Pairwise Bases with Sparse CodingabstractRecently, sparse coding has been receiving much attention in object and scene recognition tasks because of its superiority in learning an effective codebook over k-means clustering. However, empirically, such codebook requires a relatively large number of visual words, essentially bases, to achieve high recognition accuracy. Therefore, due to the combinatorial explosion of visual words, it is infeasible to use this codebook to represent higher-order spatial features which are equally important in capturing distinct properties of scenes and objects. Contrasted with many previous techniques that exploit higher-order spatial features, Local Pairwise Codebook (LPC) is a simple and effective method to learn a compact set of clusters representing pairs of spatially close descriptors with k-means. Based on LPC, this paper proposes Directional Local Pairwise Bases (DLPB) that applies sparse coding to learn a compact set of bases capturing correlation between these descriptors, so to avoid the combinatorial explosion. Furthermore, such bases are learned for each quantized direction thereby explicitly adding directional information to the representation. We have evaluated DLPB with several challenging object and scene category datasets. Our experimental results show that DLPB outperforms the baselines across all datasets and achieves the state-of-the-art performance on some datasets. Nobuyuki Morioka, Shin'ichi Satoh 0001 |
BMVC | 1 |
| 2010 | Building Compact Local Pairwise Codebook with Joint Feature Space Clustering
Nobuyuki Morioka, Shin'ichi Satoh 0001 |
ECCV (1) | 1 |
| 2006 | Enhancing Information Retrieval Using Problem Specific Knowledge
Nobuyuki Morioka, Ashesh Mahidadia |
PKAW | 1 |