VLDB 2026 Research / reviewers in the wild / expert
Muhammad Shakeel 0001
dblp:217/1039-1
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0003-3822-0917ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unifying Diarization, Separation, and ASR with Multi-Speaker EncoderabstractThis paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent inter-dependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively. Muhammad Shakeel 0001, Yui Sudo, Yifan Peng 0003, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
ASRU | 1 |
| 2025 | OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
Yifan Peng 0003, Muhammad Shakeel 0001, Yui Sudo, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2025 | Joint Target-Speaker ASR and Activity Detection
Chikara Maeda, Muhammad Shakeel 0001, Yui Sudo |
INTERSPEECH | 2 |
| 2025 | DYNAC: Dynamic Vocabulary-based Non-Autoregressive Contextualization for Speech Recognition
Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel 0001, Yifan Peng 0003, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2024 | OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language IdentificationabstractThere has been an increasing interest in large speech models that can perform multiple tasks in a single model.Such models usually adopt an encoder-decoder or decoder-only architecture due to their popularity and good performance in many domains.However, autoregressive models can be slower during inference compared to non-autoregressive models and also have potential risks of hallucination.Though prior studies observed promising results of non-autoregressive models for certain tasks at small scales, it remains unclear if they can be scaled to speech-to-text generation in diverse languages and tasks.Inspired by the Open Whisper-style Speech Model (OWSM) project, we propose OWSM-CTC, a novel encoder-only speech foundation model based on Connectionist Temporal Classification (CTC).It is trained on 180k hours of public audio data for multilingual automatic speech recognition (ASR), speech translation (ST), and language identification (LID).Compared to encoder-decoder OWSM, our OWSM-CTC achieves competitive results on ASR and up to 24% relative improvement on ST, while it is more robust and 3 to 4 times faster for inference.OWSM-CTC also improves the longform ASR result with 20x speed-up.We will publicly release our code, pre-trained model, and training logs to promote open science in speech foundation models. 1 Yifan Peng 0003, Yui Sudo, Muhammad Shakeel 0001, Shinji Watanabe 0001 |
ACL (1) | 3 |
| 2024 | Contextualized Automatic Speech Recognition With Attention-Based Bias Phrase Boosted Beam SearchabstractEnd-to-end (E2E) automatic speech recognition (ASR) methods exhibit remarkable performance. However, since the performance of such methods is intrinsically linked to the context present in the training data, E2E-ASR methods do not perform as desired for unseen user contexts (e.g., technical terms, personal names, and playlists). Thus, E2E-ASR methods must be easily contextualized by the user or developer. This paper proposes an attention-based contextual biasing method that can be customized using an editable phrase list (referred to as a bias list). The proposed method can be trained effectively by combining a bias phrase index loss and special tokens to detect the bias phrases in the input speech data. In addition, to improve the contextualization performance during inference further, we propose a bias phrase boosted (BPB) beam search algorithm based on the bias phrase index probability. Experimental results demonstrate that the proposed method consistently improves the word error rate and the character error rate of the target phrases in the bias list on both the Librispeech-960 (English) and our in-house (Japanese) dataset, respectively. Yui Sudo, Muhammad Shakeel 0001, Yosuke Fukumoto, Yifan Peng 0003, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2024 | Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
Muhammad Shakeel 0001, Yui Sudo, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2024 | OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
Yifan Peng 0003, Jinchuan Tian, Siddhant Arora, Brian Yan, Yui Sudo, Muhammad Shakeel 0001, Kwanghee Choi, Jiatong Shi, Xuankai Chang, Jee-Weon Jung, Shinji Watanabe 0001 |
INTERSPEECH | 7 |
| 2024 | Contextualized Automatic Speech Recognition With Dynamic VocabularyabstractDeep biasing (DB) enhances the performance of end-to-end automatic speech recognition (E2E-ASR) models for rare words or contextual phrases using a bias list. However, most existing methods treat bias phrases as sequences of subwords in a predefined static vocabulary. This naive sequence decomposition produces unnatural token patterns, significantly lowering their occurrence probability. More advanced techniques address this problem by expanding the vocabulary with additional modules, including the external language model shallow fusion or rescoring. However, they result in increasing the workload due to the additional modules. This paper proposes a dynamic vocabulary where bias tokens can be added during inference. Each entry in a bias list is represented as a single token, unlike a sequence of existing subword tokens. This approach eliminates the need to learn subword dependencies within the bias phrases. This method is easily applied to various architectures because it only expands the embedding and output layers in common E2E-ASR architectures. Experimental results demonstrate that the proposed method improves the bias phrase WER on English and Japanese datasets by $3.1-4.9$ points compared with the conventional DB method. Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel 0001, Yifan Peng 0003, Shinji Watanabe 0001 |
SLT | 3 |
| 2023 | Reproducing Whisper-Style Training Using An Open-Source Toolkit And Publicly Available DataabstractPre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks even in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly accessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open Whisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available data. OWSM even supports more translation directions and can be more efficient to train. We will publicly release all scripts used for data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.11https://github.com/espnet/espnet Yifan Peng 0003, Jinchuan Tian, Brian Yan, Dan Berrebbi, Xuankai Chang, Jiatong Shi, Siddhant Arora, Roshan S. Sharma, Wangyou Zhang, Yui Sudo, Muhammad Shakeel 0001, Jee-Weon Jung, Soumi Maiti, Shinji Watanabe 0001 |
ASRU | 13 |
| 2023 | DPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech Models
Yifan Peng 0003, Yui Sudo, Muhammad Shakeel 0001, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2023 | Time-synchronous one-pass Beam Search for Parallel Online and Offline Transducers with Dynamic Block Training
Yui Sudo, Muhammad Shakeel 0001, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2023 | 4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders
Yui Sudo, Muhammad Shakeel 0001, Brian Yan, Jiatong Shi, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2022 | Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated Voice Activity Detection
Yui Sudo, Muhammad Shakeel 0001, Kazuhiro Nakadai, Jiatong Shi, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2021 | Detecting earthquakes: a novel deep learning-based approach for effective disaster response
Muhammad Shakeel 0001, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai |
Appl. Intell. | 1 |