EDBT 2026 Demo / reviewers in the wild / expert
Yuto Kondo
dblp:165/3947
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking Mean Opinion Scores in Speech Quality Assessment: Score Aggregation through Quantized Distribution FittingabstractThis study addresses the task of speech quality assessment (SQA), which aims to automatically predict the subjective quality of a given speech. Recent efforts have focused on training neural-based models to predict the mean opinion score (MOS) of speech samples produced by text-to-speech or voice conversion systems. We aim to enhance the performance of the models by a score aggregation method instead of MOS. The proposed method mitigates the effects of some issues arising from constraints imposed by limited options. Our method assumes annotators internally consider continuous scores and pick the nearest discrete rating. By modeling this process, we approximate the rating distribution by quantizing the latent continuous distribution. We then use the peak of the latent distribution, estimated through the loss between the distribution and actual ratings, as the new value instead of MOS. Experimental results demonstrate that substituting MOSNet’s target with this proposed value improves prediction performance. Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
ICASSP | 1 |
| 2025 | FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
INTERSPEECH | 4 |
| 2025 | Vocoder-Projected Feature Discriminator
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
INTERSPEECH | 4 |
| 2025 | JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
INTERSPEECH | 1 |
| 2024 | Selecting N-Lowest Scores for Training MOS Prediction ModelsabstractThe automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models have been actively developed for speech samples produced by text-to-speech or voice conversion, with a primary focus on training mean opinion score (MOS) prediction models. The quality of each speech sample may not be consistent across the entire duration, and it remains unclear which segments of the speech receive the primary focus from humans when assigning subjective evaluation for MOS calculation. We hypothesize that when humans rate speech, they tend to assign more weight to low-quality speech segments, and the variance in ratings for each sample is mainly due to accidental assignment of higher scores when overlooking the poor quality speech segments. Motivated by the hypothesis, we analyze the VCC2018 and BVCC datasets. Based on the hypothesis, we propose the more reliable representative value Nlow-MOS, the mean of the N-lowest opinion scores. Our experiments show that LCC and SRCC improve compared to regular MOS when employing Nlow-MOS to MOSNet training. This result suggests that Nlow-MOS is a more intrinsic representative value of subjective speech quality and makes MOSNet a better comparator of VC models. Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
ICASSP | 1 |
| 2024 | FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
INTERSPEECH | 4 |
| 2024 | PRVAE-VC2: Non-Parallel Voice Conversion by Distillation of Speech Representations
Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Yuto Kondo |
INTERSPEECH | 4 |
| 2021 | Deficient Basis Estimation of Noise Spatial Covariance Matrix for Rank-Constrained Spatial Covariance Matrix Estimation Method in Blind Speech ExtractionabstractRank-constrained spatial covariance matrix estimation (RCSCME) is a state-of-the-art blind speech extraction method applied to cases where one directional target speech and diffuse noise are mixed. In this paper, we proposed a new algorithmic extension of RCSCME. RCSCME complements a deficient one rank of the diffuse noise spatial covariance matrix, which cannot be estimated via preprocessing such as independent low-rank matrix analysis, and estimates the source model parameters simultaneously. In the conventional RC- SCME, a direction of the deficient basis is fixed in advance and only the scale is estimated; however, the candidate of this deficient basis is not unique in general. In the proposed RCSCM model, the deficient basis itself can be accurately estimated as a vector variable by solving a vector optimization problem. Also, we derive new update rules based on the EM algorithm. We confirm that the proposed method outperforms conventional methods under several noise conditions. Yuto Kondo, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari |
ICASSP | 1 |