VLDB 2026 Research / reviewers in the wild / expert
Yui Sudo
dblp:257/3712
· DBLP profile ↗
25ranked-venue papers
11as first author
24since 2021 · last 2025
0000-0003-2094-6701ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 21 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language ModelsabstractDialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs. Tomoya Mizumoto, Yusuke Fujita, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 6 |
| 2025 | Unifying Diarization, Separation, and ASR with Multi-Speaker EncoderabstractThis paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent inter-dependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively. Muhammad Shakeel 0001, Yui Sudo, Yifan Peng 0003, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
ASRU | 2 |
| 2025 | Serialized Output Prompting for Large Language Model-based Multi-Talker Speech RecognitionabstractPrompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit prompts or rely on simple task-definition prompts, with no prior work exploring the design of prompts to enhance performance. In this paper, we propose extracting serialized output prompts (SOP) and explicitly guiding the LLM using structured prompts to improve system performance (SOP-MT-ASR). A Separator and serialized Connectionist Temporal Classification (CTC) layers are inserted after the speech encoder to separate and extract MT content from the mixed speech encoding in a first-speaking-first-out manner. Subsequently, the SOP, which serves as a prompt for LLMs, is obtained by decoding the serialized CTC outputs using greedy search. To train the model effectively, we design a threestage training strategy, consisting of serialized output training (SOT) fine-tuning, serialized speech information extraction, and SOP-based adaptation. Experimental results on the LibriMix dataset show that, although the LLM-based SOT model performs well in the two-talker scenario, it fails to fully leverage LLMs under more complex conditions, such as the three-talker scenario. The proposed SOP approach significantly improved performance under both two- and three-talker conditions. Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 6 |
| 2025 | OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
Yifan Peng 0003, Muhammad Shakeel 0001, Yui Sudo, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2025 | AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima, Lianbo Liu, Yui Sudo |
INTERSPEECH | 5 |
| 2025 | Joint Target-Speaker ASR and Activity Detection
Chikara Maeda, Muhammad Shakeel 0001, Yui Sudo |
INTERSPEECH | 3 |
| 2025 | Is Synthetic Data Truly Effective for Training Speech Language Models?
Tomoya Mizumoto, Atsushi Kojima, Yusuke Fujita, Lianbo Liu, Yui Sudo |
INTERSPEECH | 5 |
| 2025 | DYNAC: Dynamic Vocabulary-based Non-Autoregressive Contextualization for Speech Recognition
Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel 0001, Yifan Peng 0003, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2025 | OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
Yui Sudo, Yusuke Fujita, Atsushi Kojima, Tomoya Mizumoto, Lianbo Liu |
INTERSPEECH | 1 |
| 2024 | OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language IdentificationabstractThere has been an increasing interest in large speech models that can perform multiple tasks in a single model.Such models usually adopt an encoder-decoder or decoder-only architecture due to their popularity and good performance in many domains.However, autoregressive models can be slower during inference compared to non-autoregressive models and also have potential risks of hallucination.Though prior studies observed promising results of non-autoregressive models for certain tasks at small scales, it remains unclear if they can be scaled to speech-to-text generation in diverse languages and tasks.Inspired by the Open Whisper-style Speech Model (OWSM) project, we propose OWSM-CTC, a novel encoder-only speech foundation model based on Connectionist Temporal Classification (CTC).It is trained on 180k hours of public audio data for multilingual automatic speech recognition (ASR), speech translation (ST), and language identification (LID).Compared to encoder-decoder OWSM, our OWSM-CTC achieves competitive results on ASR and up to 24% relative improvement on ST, while it is more robust and 3 to 4 times faster for inference.OWSM-CTC also improves the longform ASR result with 20x speed-up.We will publicly release our code, pre-trained model, and training logs to promote open science in speech foundation models. 1 Yifan Peng 0003, Yui Sudo, Muhammad Shakeel 0001, Shinji Watanabe 0001 |
ACL (1) | 2 |
| 2024 | Contextualized Automatic Speech Recognition With Attention-Based Bias Phrase Boosted Beam SearchabstractEnd-to-end (E2E) automatic speech recognition (ASR) methods exhibit remarkable performance. However, since the performance of such methods is intrinsically linked to the context present in the training data, E2E-ASR methods do not perform as desired for unseen user contexts (e.g., technical terms, personal names, and playlists). Thus, E2E-ASR methods must be easily contextualized by the user or developer. This paper proposes an attention-based contextual biasing method that can be customized using an editable phrase list (referred to as a bias list). The proposed method can be trained effectively by combining a bias phrase index loss and special tokens to detect the bias phrases in the input speech data. In addition, to improve the contextualization performance during inference further, we propose a bias phrase boosted (BPB) beam search algorithm based on the bias phrase index probability. Experimental results demonstrate that the proposed method consistently improves the word error rate and the character error rate of the target phrases in the bias list on both the Librispeech-960 (English) and our in-house (Japanese) dataset, respectively. Yui Sudo, Muhammad Shakeel 0001, Yosuke Fukumoto, Yifan Peng 0003, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2024 | Improving Noise Robustness of Automatic Speech Recognition Based on a Parallel Adapter Model with Near-Identity Initialization
Takahiro Osaki, Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai |
IEA/AIE | 2 |
| 2024 | Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
Muhammad Shakeel 0001, Yui Sudo, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2024 | OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
Yifan Peng 0003, Jinchuan Tian, Siddhant Arora, Brian Yan, Yui Sudo, Muhammad Shakeel 0001, Kwanghee Choi, Jiatong Shi, Xuankai Chang, Jee-Weon Jung, Shinji Watanabe 0001 |
INTERSPEECH | 6 |
| 2024 | Contextualized Automatic Speech Recognition With Dynamic VocabularyabstractDeep biasing (DB) enhances the performance of end-to-end automatic speech recognition (E2E-ASR) models for rare words or contextual phrases using a bias list. However, most existing methods treat bias phrases as sequences of subwords in a predefined static vocabulary. This naive sequence decomposition produces unnatural token patterns, significantly lowering their occurrence probability. More advanced techniques address this problem by expanding the vocabulary with additional modules, including the external language model shallow fusion or rescoring. However, they result in increasing the workload due to the additional modules. This paper proposes a dynamic vocabulary where bias tokens can be added during inference. Each entry in a bias list is represented as a single token, unlike a sequence of existing subword tokens. This approach eliminates the need to learn subword dependencies within the bias phrases. This method is easily applied to various architectures because it only expands the embedding and output layers in common E2E-ASR architectures. Experimental results demonstrate that the proposed method improves the bias phrase WER on English and Japanese datasets by $3.1-4.9$ points compared with the conventional DB method. Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel 0001, Yifan Peng 0003, Shinji Watanabe 0001 |
SLT | 1 |
| 2023 | Reproducing Whisper-Style Training Using An Open-Source Toolkit And Publicly Available DataabstractPre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks even in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly accessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open Whisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available data. OWSM even supports more translation directions and can be more efficient to train. We will publicly release all scripts used for data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.11https://github.com/espnet/espnet Yifan Peng 0003, Jinchuan Tian, Brian Yan, Dan Berrebbi, Xuankai Chang, Jiatong Shi, Siddhant Arora, Roshan S. Sharma, Wangyou Zhang, Yui Sudo, Muhammad Shakeel 0001, Jee-Weon Jung, Soumi Maiti, Shinji Watanabe 0001 |
ASRU | 12 |
| 2023 | DPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech Models
Yifan Peng 0003, Yui Sudo, Muhammad Shakeel 0001, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2023 | Time-synchronous one-pass Beam Search for Parallel Online and Offline Transducers with Dynamic Block Training
Yui Sudo, Muhammad Shakeel 0001, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2023 | 4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders
Yui Sudo, Muhammad Shakeel 0001, Brian Yan, Jiatong Shi, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2023 | Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation
Yui Sudo, Kazuya Hata, Kazuhiro Nakadai |
INTERSPEECH | 1 |
| 2023 | Online Adaptation of Fourier Series Based Acoustic Transfer Function Model to Improve Sound Source Localization and SeparationabstractThis paper proposes an online adaptation method for Fourier series based acoustic transfer function (TF) models for robot audition systems based on microphone array signal processing. The TF represents the signal propagation characteristics from a sound source to a microphone, which is an essential component for real-world auditory scene analysis, including sound source localization and separation. The real-world applications of TF-based array signal processing requires two characteristics: 1) adaptability to changes in the acoustic environment (changes in the signal propagation characteristics between the sound source and the microphone), and 2) a lightweight TF set for use in embedded systems such as robots with limited memory and computational resources. This paper proposes an online adaptation method for lightweight TF models using the Fourier series expansion. This method has both above two characteristics. Experimental results showed that the use of TF set adapted online using the proposed method performs better sound source localization and separation performance than existing online TF adaptation methods. Yui Sudo, Masayuki Takigahira, Hideo Tsuru, Kazuhiro Nakadai, Hirofumi Nakajima |
RO-MAN | 1 |
| 2022 | Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated Voice Activity Detection
Yui Sudo, Muhammad Shakeel 0001, Kazuhiro Nakadai, Jiatong Shi, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2022 | Empirical Sampling from Latent Utterance-wise Evidence Model for Missing Data ASR based on Neural Encoder-Decoder Model
Ryu Takeda, Yui Sudo, Kazuhiro Nakadai, Kazunori Komatani |
INTERSPEECH | 2 |
| 2021 | Multichannel environmental sound segmentationabstractAbstract This paper proposes a multichannel environmental sound segmentation method. Environmental sound segmentation is an integrated method to achieve sound source localization, sound source separation and classification, simultaneously. When multiple microphones are available, spatial features can be used to improve the localization and separation accuracy of sounds from different directions; however, conventional methods have three drawbacks: (a) Sound source localization and sound source separation methods using spatial features and classification using spectral features trained in the same neural network, may overfit to the relationship between the direction of arrival and the class of a sound, thereby reducing their reliability to deal with novel events. (b) Although permutation invariant training used in autonomous speech recognition could be extended, it is impractical for environmental sounds that include an unlimited number of sound sources. (c) Various features, such as complex values of short time Fourier transform and interchannel phase differences have been used as spatial features, but no study has compared them. This paper proposes a multichannel environmental sound segmentation method comprising two discrete blocks, a sound source localization and separation block and a sound source separation and classification block. By separating the blocks, overfitting to the relationship between the direction of arrival and the class is avoided. Simulation experiments using created datasets including 75-class environmental sounds showed the root mean squared error of the proposed method was lower than that of conventional methods. Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai |
Appl. Intell. | 1 |
| 2019 | Environmental sound segmentation utilizing Mask U-NetabstractThis paper proposes an environmental sound segmentation method using Mask U-Net. Recent research in robot audition has analyzed noise reduction, section detection, and sound source separation for use in a real-world environment with many noises and overlaps. However, conventional methods apply respective functions in cascades. The biggest problem of cascade systems is the accumulation of errors generated at each function block. Although many methods of human voice separation have been proposed, robots operating in a real-world environment must be able to separate not only human voices but other environmental sounds. Unlike traditional sound source separation using spatial information, environmental sound segmentation must simultaneously detect sections and separate sound sources based on pre-trained features. One such method, U-Net, which was proposed for semantic segmentation of images, has been applied to the separation of singing voices. However, this method deals only with limited classes of sounds. The current study proposes an environmental sound segmentation method using Mask U-Net, which combines segmentation using U-Net with sound event detection using CNN to 75-classes of environmental sounds. Experimental application confirmed that this method improved learning speed and sound source separation compared with the conventional method. Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai |
IROS | 1 |