Yusuke Yasuda

dblp:228/9342 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Stable Normalization for Preference-based Subset Selection with an Unbounded External Archive
abstract
An unbounded external archive (UA) has recently been introduced into preference-based evolutionary multi-objective optimization (PBEMO), which enables postprocessing to select solutions in the region of interest (ROI) from all evaluated non-dominated solutions. However, in scaled multi-objective optimization problems (scaled MOPs), differences in the scales of objectives and the presence of dominance resistant solutions (DRSs) in the UA often destabilize normalization and nadir point estimation, leading to degraded ROI approximation performance even after normalization. To address this issue, we propose a novel reference point-based postprocessing method that achieves accurate normalization in scaled MOPs. The method combines: (a) α-weighted objective transformation to remove residual DRSs from the UA, (b) Pareto front modeling-based linearization to reduce shape dependency by linearizing the Pareto front (PF), and (c) sigma-based preprocessing to remove outliers for stable front linearization execution. Experiments on scaled DTLZ and WFG benchmarks demonstrate that the proposed method consistently outperforms existing postprocessing and DRSs removal approaches in both ROI solution quality and nadir point estimation accuracy. These results indicate that our method can stably provide the decision maker with preference-aligned solution sets, even for MOPs with varying objective scales and PF shapes.
Yuta Nakanishi, Mamoru Doi, Yusuke Yasuda, Hiroyuki Sato 0003
GECCO3
2026 Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment
Yusuke Yasuda, Tomoki Toda
Comput. Speech Lang.1
2025 Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
Cheng-Hung Hu, Yusuke Yasuda, Akifumi Yoshimoto, Tomoki Toda
INTERSPEECH2
2025 E2EPref: An end-to-end preference-based framework for speech quality assessment to alleviate bias in direct assessment scores
abstract
In speech quality assessment (SQA), direct assessment (DA) scores are frequently used as the objective of model training. However, because the DA scores themselves have listener-wise bias and equal range bias, the scores predicted by models trained with DA scores do not always reflect the true quality score. In this study, we utilize preference-based learning for SQA by transforming the DA score prediction framework into a preference prediction framework. Our proposed End-to-End Preference-based framework (E2EPref) for SQA is designed for predicting system-level quality scores directly. It contains four proposed components: pair generation, preference function, threshold selection, and preference aggregation. Through these functions of E2EPref, we aim to mitigate biases introduced by directly using DA scores for training. In experiments, we show that this framework helps the SQA model alleviate biases, resulting in higher system-level Spearman’s rank correlation coefficient and linear correlation coefficient. Additionally, we evaluate the quality prediction capability of the framework in a zero-shot out-of-domain scenario. Finally, we collect subjective preference scores on a dataset already containing DA scores and analyze the advantages and disadvantages of using DA scores versus subjective preference scores as the ground truth or for model training. • Proposed E2EPref: an end-to-end preference-based SQA framework. • Pair generation, preference function, and aggregation enhance model performance. • E2EPref outperforms baselines in in-domain, out-of-domain, and cross-dataset tests. • Collected preference score dataset for advancing preference-based SQA research.
Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda
Comput. Speech Lang.2
2024 Exploring the Robustness of Text-to-Speech Synthesis Based on Diffusion Probabilistic Models to Heavily Noisy Transcriptions
Jingyi Feng, Yusuke Yasuda, Tomoki Toda
INTERSPEECH2
2024 Embedding Learning for Preference-based Speech Quality Assessment
Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda
INTERSPEECH2
2024 A Bilayered Decomposition Technique for Handling Complex Constrained Multi-Objective Optimization Problems
abstract
Constrained multi-objective optimization problems (CMOPs) are prevalent in real-world applications, dealing with multi-objective and constraint functions. Multiobjective evolutionary algorithm based on decomposition with differential evolution (MOEA/D-DE) has proven effective for unconstrained multi-objective optimization problems with complex Pareto front (PF). However, in CMOPs, the conflicting nature between objectives and constraints often makes it challenging to appropriately manage constraints while ensuring convergence, diversity, and feasibility of the solution set towards the PF. To address this issue, this paper proposes a bilayered decomposition technique. The proposed algorithm treats constraints as an additional objective function, separately decomposing the objective space and the constraint violation space. This method allows for the simultaneous attainment of convergence, diversity, and feasibility in the solution set while efficiently exploring the PF. Experimental results demonstrate that, within the MOEA/D-DE framework, our algorithm efficiently navigates complex PFs on challenging problems such as LIR-CMOPs and DAS-CMOPs, matching the performance of state-of-the-art constrained multi-objective evolutionary algorithms.
Yusuke Yasuda, Kenichi Tamura, Keiichiro Yasuda
SMC1
2023 Text-To-Speech Synthesis Based on Latent Variable Conversion Using Diffusion Probabilistic Model and Variational Autoencoder
abstract
Text-to-speech synthesis (TTS) is a task to convert texts into speech. Two of the factors that have been driving TTS are the advancements of probabilistic models and latent representation learning. We propose a TTS method based on latent variable conversion using a diffusion probabilistic model and the variational autoencoder (VAE). In our TTS method, we use a waveform model based on VAE, a diffusion model that predicts the distribution of latent variables in the waveform model from texts, and an alignment model that learns alignments between the text and speech latent sequences. Our method integrates diffusion with VAE by modeling both mean and variance parameters with diffusion, where the target distribution is determined by approximation from VAE. This latent variable conversion framework potentially enables us to flexibly incorporate various latent feature extractors. Our experiments show that our method is robust to linguistic labels with ambiguous orthography and alignment errors.
Yusuke Yasuda, Tomoki Toda
ICASSP1
2023 Preference-based training framework for automatic speech quality assessment using deep neural network
abstract
One objective of Speech Quality Assessment (SQA) is to estimate the ranks of synthetic speech systems. However, recent SQA models are typically trained using low-precision direct scores such as mean opinion scores (MOS) as the training objective, which is not straightforward to estimate ranking. Although it is effective for predicting quality scores of individual sentences, this approach does not account for speech and system preferences when ranking multiple systems. We propose a training framework of SQA models that can be trained with only preference scores derived from pairs of MOS to improve ranking prediction. Our experiment reveals conditions where our framework works the best in terms of pair generation, aggregation functions to derive system score from utterance preferences, and threshold functions to determine preference from a pair of MOS. Our results demonstrate that our proposed method significantly outperforms the baseline model in Spearman's Rank Correlation Coefficient.
Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda
INTERSPEECH2
2023 Analysis of Mean Opinion Scores in Subjective Evaluation of Synthetic Speech Based on Tail Probabilities
Yusuke Yasuda, Tomoki Toda
INTERSPEECH1
2023 Differential Evolution Using Superior Infeasible Solutions for Constrained Optimization
abstract
Differential Evolution (DE) is one of the effective metaheuristics for solving unconstrained optimization problems. Constraint Handling Technique (CHT) is needed to extend DE to constrained optimization. Feasibility Rule (FR) is one of the typical CHT. FR addresses constraints by using a simple rule that considers the objective function and constraint violation when comparing search individuals. However, since the solutions with small constraint violation, i.e., feasible solutions, are preferentially selected, the set of search individuals may be biased toward feasible regions and the improvement of the objective function value may stagnate. This paper overcomes this challenge by proposing a DE that uses a superior infeasible solution in an external archive. The external archive in the proposed DE stores search individuals that are superior in both objective function value and constraint violation and utilize them to generate mutant individuals. Finally, we verify the effectiveness of the proposed method using a benchmark problem where the feasible region is a convex set.
Yuji Sato, Watatu Kumagai, Yusuke Yasuda, Kenichi Tamura, Keiichiro Yasuda
SMC3
2022 Spoken-Text-Style Transfer with Conditional Variational Autoencoder and Content Word Storage
Daiki Yoshioka, Yusuke Yasuda, Noriyuki Matsunaga, Yamato Ohtani, Tomoki Toda
INTERSPEECH2
2021 How Similar or Different is Rakugo Speech Synthesizer to Professional Performers?
abstract
We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authentically entertains audiences. In this paper, we propose a novel evaluation methodology using synthesized rakugo speech and real rakugo speech uttered by professional performers of three different ranks. The naturalness of the synthesized speech was comparable to that of the human speech, but the synthesized speech entertained listeners less than the performers of any rank. However, we obtained some interesting insights into challenges to be solved in order to achieve a truly entertaining rakugo synthesizer. For example, naturalness was not the most important factor, even though it has generally been emphasized as the most important point to be evaluated in the conventional speech synthesis field. More important factors were the understandability of the content and distinguishability of the characters in the rakugo story, both of which the synthesized rakugo speech was relatively inferior at as compared with the professional performers. We also found that fundamental frequency (fo) modeling should at least be further improved to better entertain audiences. These results show important steps to reaching authentically entertaining speech synthesis.
Shuhei Kato, Yusuke Yasuda, Xin Wang 0037, Erica Cooper, Junichi Yamagishi
ICASSP2
2021 End-to-End Text-to-Speech Using Latent Duration Based on VQ-VAE
abstract
Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling that incorporates duration as a discrete latent variable to TTS and enables joint optimization of whole modules from scratch. We formulate our method based on conditional VQ-VAE to handle discrete duration in a variational autoencoder and provide a theoretical explanation to justify our method. In our framework, a connectionist temporal classification (CTC) -based force aligner acts as the approximate posterior, and text-to-duration works as the prior in the variational autoencoder. We evaluated our proposed method with a listening test and compared it with other TTS methods based on soft-attention or explicit duration modeling. The results showed that our systems rated between soft-attention-based methods (Transformer-TTS, Tacotron2) and explicit duration modeling-based methods (Fastspeech).
Yusuke Yasuda, Xin Wang 0037, Junichi Yamagishi
ICASSP1
2021 Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis
Yusuke Yasuda, Xin Wang 0037, Junichi Yamagishi
Comput. Speech Lang.1
2020 Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings
abstract
While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate multi-speaker modeling for end-to-end text-to-speech synthesis and study the effects of different types of state-of-the-art neural speaker embeddings on speaker similarity for unseen speakers. Learnable dictionary encoding-based speaker embeddings with angular softmax loss can improve equal error rates over x-vectors in a speaker verification task; these embeddings also improve speaker similarity and naturalness for unseen speakers when used for zero-shot adaptation to new speakers in end-to-end speech synthesis.
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang 0037, Nanxin Chen, Junichi Yamagishi
ICASSP3
2020 Effect of Choice of Probability Distribution, Randomness, and Search Methods for Alignment Modeling in Sequence-to-Sequence Text-to-Speech Synthesis Using Hard Alignment
abstract
Sequence-to-sequence text-to-speech (TTS) is dominated by soft-attention-based methods. Recently, hard-attention-based methods have been proposed to prevent fatal alignment errors, but their sampling method of discrete alignment is poorly investigated. This research investigates various combinations of sampling methods and probability distributions for alignment transition modeling in a hard-alignment-based sequence-to-sequence TTS method called SSNT-TTS. We clarify the common sampling methods of discrete variables including greedy search, beam search, and random sampling from a Bernoulli distribution in a more general way. Furthermore, we introduce the binary Concrete distribution to model discrete variables more properly. The results of a listening test shows that deterministic search is more preferable than stochastic search, and the binary Concrete distribution is robust with stochastic search for natural alignment transition.
Yusuke Yasuda, Xin Wang 0037, Junichi Yamagishi
ICASSP1
2020 Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
abstract
Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation paradigms, speaker augmentation, by creating artificial speakers and by taking advantage of low-quality data. The base Tacotron2 model is modified to account for the channel and dialect factors inherent in these corpora. In addition, we describe a warm-start training strategy that we adopted for Tacotron2 training. A large-scale listening test is conducted, and a distance metric is adopted to evaluate synthesis of dialects. This is followed by an analysis on synthesis quality, speaker and dialect similarity, and a remark on the effectiveness of our speaker augmentation approach. Audio samples are available online.
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Junichi Yamagishi
INTERSPEECH3
2019 Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language
abstract
End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its applicability to other languages is still unknown. Japanese could be one of the most difficult languages for which to achieve end-to-end speech synthesis, largely due to its character diversity and pitch accents. Therefore, state-of-the-art systems are still based on a traditional pipeline framework that requires a separate text analyzer and duration model. Towards end-to-end Japanese speech synthesis, we extend Tacotron to systems with self-attention to capture long-term dependencies related to pitch accents and compare their audio quality with classical pipeline systems under various conditions to show their pros and cons. In a large-scale listening test, we investigated the impacts of the presence of accentual-type labels, the use of force or predicted alignments, and acoustic features used as local condition parameters of the Wavenet vocoder. Our results reveal that although the proposed systems still do not match the quality of a top-line pipeline system for Japanese, we show important stepping stones towards end-to-end Japanese speech synthesis.
Yusuke Yasuda, Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
ICASSP1