VLDB 2026 Research / reviewers in the wild / expert
Yoshinao Sato
dblp:235/7051
· DBLP profile ↗
11ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-0657-0269ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluation of Paralinguistic-Aware Spoken Dialogue Systems using Next-Utterance ClassificationabstractIn spoken dialogues, paralinguistic cues frequently convey crucial information not captured by linguistic content alone. However, conventional spoken dialogue systems (SDSs), which typically comprise a cascade of an automatic speech recognition model and a large language model (LLM), lack the ability to recognize paralinguistic cues. Relying solely on transcribed text, paralinguistic-agnostic SDSs often cause dialogue breakdowns. To address this difficulty, integrating a paralinguistic recognition model into SDSs is essential. Therefore, this study focuses on evaluating such paralinguistic-aware SDSs. To this end, we propose a corpus-based evaluation method utilizing next-utterance classification as an automated alternative to human evaluation. Specifically, an LLM is tasked with predicting the dialogue act of the subsequent utterance given a transcribed dialogue history, comparing scenarios with and without paralinguistic attitude classes. Our experiments demonstrate that incorporating a paralinguistic attitude recognition model improves prediction performance, as measured by the mean reciprocal rank. Furthermore, we assessed the alignment of our proposed corpus-based evaluation method with subjective human evaluations. The results confirmed the validity of the proposed method as a reliable proxy for human evaluation, demonstrating its correlation with human rankings and agreement with human preferences in pairwise comparisons. Taken together, this work highlights that integrating paralinguistic recognition is a crucial step toward realizing more robust and natural spoken dialogue systems. Kouki Miyazawa, Yoshinao Sato |
SIGDIAL | 2 |
| 2025 | Transition Relevance Point Detection for Spoken Dialogue Systems with Self-Attention TransformerabstractMost conventional spoken dialogue systems determine when to respond based on the elapsed time of silence following user speech utterances. This approach often results in failures of turn-taking, disrupting smooth communications with users. This study addresses the detection of when it is acceptable for the dialogue system to start speaking. Specifically, we aim to detect transition relevant points (TRPs) rather than predict whether the dialogue participants will actually start speaking. To achieve this, we employ a self-supervised speech representation using contrastive predictive coding and a self-attention transformer. The proposed model, TRPDformer, was trained and evaluated on the corpus of everyday Japanese conversation. TRPDformer outperformed a baseline model based on the elapsed time of silence. Furthermore, third-party listeners rated the timing of system responses determined using the proposed model as superior to that of the baseline in a preference test. Kouki Miyazawa, Yoshinao Sato |
SIGDIAL | 2 |
| 2023 | Shuffleaugment: A Data Augmentation Method Using Time ShufflingabstractWe present ShuffleAugment, a data augmentation method for speech processing that randomly shuffles data in the time direction. Every speech processing task has a characteristic time scale depending on the phenomenon it addresses. The proposed method randomizes the time order of an input sequence on the irrelevant time scales and obtains many variants without sacrificing the essential information on the proper time scale. The shuffling process can be implemented as a neural network layer and applied to low- and high-level features at an arbitrary depth. We evaluate the efficiency of the proposed method by applying it to two tasks: speaker recognition and speech emotion recognition. Our experiments demonstrated that long-term and short-term shuffles improved the performance of speaker recognition and speech emotion recognition, respectively. These results indicate that ShuffleAugment is an effective data augmentation method. Yoshinao Sato, Narumitsu Ikeda, Hirokazu Takahashi |
ICASSP | 1 |
| 2023 | Domain Adaptation without Catastrophic Forgetting on a Small-Scale Partially-Labeled Corpus for Speech Emotion RecognitionabstractHence, developing a corpus for speech emotion recognition (SER) in the target domain is significant; however, this is time-consuming and cost-intensive. In this study, we aim to fully use a partially-labeled corpus in the target domain (target corpus) with the help of an existing fully-labeled corpus (common corpus). To this end, we proposed a method that leverages domain adversarial multi-task learning to reconcile the definitions of emotion classes across domains and noisy student training to utilize unlabeled data. Our experimental results demonstrated that the proposed method improved the SER performance in the target domain when the target corpus was small in size and imbalanced in classes. Furthermore, performance on the common corpus was not deteriorated by the proposed method. Yoshinao Sato |
ICASSP | 2 |
| 2023 | A Voice-Activity-Aware Loss Function for Continuous Speech SeparationabstractContinuous speech separation (CSS) is gaining popularity as a realistic scenario in which speech utterances partially overlap. Most previous studies on CSS employed the conventional scale-invariant signal-to-noise ratio (SI-SNR) loss, which is ill-defined when the reference audio is silent and thus requires some modifications. Hence, we propose a voice-activity-aware loss function that combines SI-SNR loss for speech segments and logarithmic root mean square loss for nonspeech segments. We trained and evaluated the block-online temporal convolutional network models on synthetic single-channel speech mixtures in noisy and reverberant environments. The results demonstrated that the proposed loss function outperformed the conventional loss function. Furthermore, qualitative analyses indicated that the proposed loss function accurately detected speech onsets and offsets, yielding reduced residual noise in nonspeech segments. Conggui Liu, Yoshinao Sato |
MMSP | 2 |
| 2021 | Speech Emotion Recognition Using Semi-Supervised Learning with Efficient Labeling StrategiesabstractThe collection of large amounts of labeled data for speech emotion recognition requires considerable time and effort. As a result, the sizes of existing corpora are limited. One promising solution to this difficulty is semi-supervised learning, i.e., learning from both labeled and unlabeled data. In this study, we applied the noisy student training (NST) method to speech emotion recognition. We experimentally investigate the trade-off between the amount and reliability of labeled data. For this purpose, we prepared labeled and unlabeled data by lim-iting the available annotations in the CREMA-D dataset. The experimental results showed that a model trained using the NST method with some of the annotations achieved almost the same performance as the one trained using supervised learning with all the annotations if the amount and reliabil-ity of the available annotations were appropriate. Our findings are significant in identifying the most efficient labeling strategy when utilizing a large-scale dataset without labels for speech emotion recognition. Yoshinao Sato |
ASRU | 2 |
| 2021 | Enhancing Block-Online Speech Separation using Interblock Context FlowabstractDespite recent progress in speech separation, online processing is still challenging. One promising approach is the block-online structure, which has been examined in a few previous studies. However, in a blockwise model, the available context information is limited to the same block. To overcome this limitation, we investigate the enhancement of a block-online speech separation model using interblock context flow. Specifically, we propose a blockwise temporal convolution network with layers between adjacent blocks that allow the propagation of interblock context information. We evaluate this model on single-channel speech mixtures with different context widths, latencies, and intervals in noisy and reverberant environments generated by the image method using the Wall Street Journal 0 corpus. The experimental results indicate that the proposed model outperforms the baseline blockwise model under all the conditions. Conggui Liu, Yoshinao Sato |
MMSP | 2 |
| 2021 | Efficient corpus design for wake-word detectionabstractWake-word detection is an indispensable technology for preventing virtual voice agents from being unintentionally triggered. Although various neural networks were proposed for wake-word detection, less attention has been paid to efficient corpus design, which we address in this study. For this purpose, we collected speech data via a crowdsourcing platform and evaluated the performance of several neural networks when different subsets of the corpus were used for training. The results reveal the following requirements for efficient corpus design to produce a lower misdetection rate: (1) short segments of continuous speech can be used as negative samples, but they are not as effective as random words; (2) utterances of "adversarial" words, i.e., phonetically similar words to a wake-word, contribute to improving performance significantly when they are used as negative samples; (3) it is preferable for individual speakers to provide both positive and negative samples; (4) increasing the number of speakers is better than increasing the number of repetitions of a wake-word by each speaker. Delowar Hossain, Yoshinao Sato |
SLT | 2 |
| 2020 | Reconciliation of Multiple Corpora for Speech Emotion Recognition by Multiple Classifiers with an Adversarial Corpus Discriminator
Yoshinao Sato |
INTERSPEECH | 2 |
| 2020 | Quality Estimation for Partially Subjective Classification Tasks via CrowdsourcingabstractThe quality estimation of artifacts generated by creators via crowdsourcing has great significance for the construction of a large-scale data resource. A common approach to this problem is to ask multiple reviewers to evaluate the same artifacts. However, the commonly used majority voting method to aggregate reviewers’ evaluations does not work effectively for partially subjective or purely subjective tasks because reviewers’ sensitivity and bias of evaluation tend to have a wide variety. To overcome this difficulty, we propose a probabilistic model for subjective classification tasks that incorporates the qualities of artifacts as well as the abilities and biases of creators and reviewers as latent variables to be jointly inferred. We applied this method to the partially subjective task of speech classification into the following four attitudes: agreement, disagreement, stalling, and question. The result shows that the proposed method estimates the quality of speech more effectively than a vote aggregation, measured by correlation with a fine-grained classification by experts. Yoshinao Sato, Kouki Miyazawa |
LREC | 1 |
| 2018 | Short Utterance Speaker Recognition by Reservoir with Self-Organized MappingabstractShort utterances cause performance degradation in conventional speaker recognition systems based on i-vector, which relies on the statistics of spectral features. To overcome this difficulty, we propose a novel method that utilizes the dynamics of the spectral features as well as their distribution. Our model integrates echo state network (ESN), a type of reservoir computing architecture, and self-organizing map (SOM), a competitive learning network. The ESN consists of a single-hidden-layer recurrent neural network with randomly fixed weights, which extracts temporal patterns of the spectral features. The input weights of our model are trained using the unsupervised competitive learning algorithm of the SOM, before enrollment, to extract the intrinsic structure of the spectral features, whereas the input weights are fixed randomly in the original ESN. In enrollment, the output weights are trained in a supervised manner to recognize an individual in a group of speakers. Our experiment demonstrates that the proposed method outperforms or is comparable to a baseline i-vector system for text-independent speaker identification on short utterances. Narumitsu Ikeda, Yoshinao Sato, Hirokazu Takahashi |
SLT | 2 |