Yiwen Shao

dblp:226/2078 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
abstract
Extending pre-trained text Large Language Models (LLMs)’s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention in the speech research community. However, building a unified speech understanding and generation model still faces the following challenges: (1) Due to the huge modality gap between speech and text tokens, extending text LLMs to unified speech LLMs relies on large-scale paired data for fine-tuning, and (2) Generation and understanding tasks prefer information at different levels, e.g., generation benefits from detailed acoustic features, while understanding favors high-level semantics. This divergence leads to difficult performance optimization in one unified model. To solve these challenges, in this paper, we present two key insights in speech tokenization and speech language modeling. Specifically, we first propose an Understanding-driven Speech Tokenizer (USTokenizer), which extracts high-level semantic information essential for accomplishing understanding tasks using text LLMs. In this way, USToken enjoys better modality commonality with text, which reduces the difficulty of modality alignment in adapting text LLMs to speech LLMs. Secondly, we present DualSpeechLM, a dual-token modeling framework that concurrently models USToken as input and acoustic token as output within a unified, end-to-end framework, seamlessly integrating speech understanding and generation capabilities. Furthermore, we propose a novel semantic supervision loss and a Chain-of-Condition (CoC) strategy to stabilize model training and enhance speech generation performance. Experimental results demonstrate that our proposed approach effectively fosters a complementary relationship between understanding and generation tasks, highlighting the promising strategy of mutually enhancing both tasks in one unified model.
Dongchao Yang, Yiwen Shao, Hangting Chen, Jiankun Zhao, Zhiyong Wu 0001, Helen M. Meng, Xixin Wu
AAAI3
2026 TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
abstract
We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization.The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and(2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking.Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models who spoke what and when in an end-to-end manner.Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps.Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost 1 . Recent Work on Multi-speaker ASR & Diarization
Mingyue Huo, Yiwen Shao
ACL (1)2
2026 Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
abstract
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored.Crucially, there is no consensus on whether audio-language models can build effective general-purpose audio encoders, nor a systematic understanding of how pretraining objectives behave across diverse tasks and scales.We identify three key barriers: limited scale of audio-text corpora, limited coverage of audio attributes in existing caption corpora, and lack of systematic exploration and evaluation.To fill this gap, we present the first principled empirical study of ALP.We first introduce Cap-tionStew, a 10.7M caption dataset aggregating open-source audio-text corpora across multiple domains and captioning focuses.We then conduct the first comprehensive evaluation comparing contrastive and captioning objectives for learning audio representation across speech, music, and environmental sound tasks.Our results not only demonstrate that ALP yields competitive, transferable representations, but reveal critical trade-offs: contrastive learning offers superior data efficiency, while captioning exhibits better scalability.Furthermore, we find that the benefits of supervised initialization often diminish at larger scales, challenging common practices.By grounding these claims in empirical evidence, we establish a viable pathway toward general-purpose audio representation learning, guiding future research.
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao, Hao Zhang 0112, Dong Yu 0001
ACL (1)4
2025 Efficient Scaling for LLM-based ASR
abstract
Large language model (LLM)-based automatic speech recognition (ASR) achieves strong performance but often incurs high computational costs. This work investigates how to obtain the best LLM-ASR performance efficiently. Through comprehensive and controlled experiments, we find that pretraining the speech encoder before integrating it with the LLM leads to significantly better scaling efficiency than the standard practice of joint post-training of LLM-ASR. Based on this insight, we propose a new multi-stage LLM-ASR training strategy, EFIN: Encoder First Integration. Among all training strategies evaluated, EFIN consistently delivers better performance (relative to 21.1 % CERR) with significantly lower computation budgets ($49.9 \%$ FLOPs). Furthermore, we derive a scaling law that approximates ASR error rates as a computation function, providing practical guidance for LLM-ASR scaling.
Bingshen Mu, Yiwen Shao, Dong Yu 0001, Lei Xie 0001
ASRU2
2025 Efficient Multilingual ASR Finetuning via LoRA Language Experts
Yiwen Shao, Jianheng Zhuo, Chenda Li, Liliang Tang, Dong Yu 0001, Yanmin Qian
INTERSPEECH2
2025 VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
Jianheng Zhuo, Yifan Yang 0005, Yiwen Shao, Yong Xu 0004, Dong Yu 0001, Kai Yu 0004, Xie Chen 0001
INTERSPEECH3
2024 UniX-Encoder: A Universal X-Channel Speech Encoder for AD-HOC Microphone Array Speech Processing
abstract
The speech field is evolving to solve more challenging scenarios, such as multi-channel recordings with multiple simultaneous talkers. In response to the diversity of microphone configurations in use, we introduce the UniX-Encoder, a universal encoder for multi-channel speech recordings. The UniX-Encoder is versatile, catering to a variety of speech tasks, and seamlessly integrates with any microphone array, whether in single-talker or multi-talker environments. Our research enhances previous multi-channel speech processing efforts in four aspects: 1) Adaptability: Contrasting traditional models constrained to certain microphone array configurations, our encoder is universally compatible. 2) Multi-Task Capability: Contrasting previous systems that were designed for single-task applications, the UniX-Encoder serves as a versatile upstream model, capable of extracting features for diverse speech tasks. 3) Self-Supervised Training: The UniX-Encoder is pretrained without the need for labeled multi-channel data. 4) End-to-End Integration: In contrast to models that first beamform then process single-channels, our encoder offers an end-to-end solution, bypassing explicit beamforming or separation. To validate its effectiveness, we tested the UniX-Encoder on a synthetic multi-channel dataset from the LibriSpeech corpus. Across various tasks, including ASR and speaker diarization, our encoder consistently outperformed combinations such as the WavLM model with the BeamformIt frontend.
Zili Huang, Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
ICASSP2
2024 RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios
Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH1
2024 Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Daniel Povey, Sanjeev Khudanpur
INTERSPEECH1
2024 Spatialemb: Extract and Encode Spatial Information for 1-Stage Multi-Channel Multi-Speaker ASR on Arbitrary Microphone Arrays
abstract
Spatial information is a critical clue for multi-channel multispeaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation stage, followed by standard single-channel ASR on the separated speech. This approach results in an inefficient, lengthy pipeline and sub-optimal ASR performance due to the accumulated errors from preprocessing modules. Furthermore, most spatial feature extraction methods depend on the knowledge of speaker positions and microphone topology, making the systems reliant on specific settings and challenging to adapt to new equipment. In this work, we propose a solution to these issues with a lightweight embedding module named SpatialEmb, which extracts and encodes spatial information directly for the ASR model, supporting both fixed and arbitrary microphone topology. We conduct comprehensive experiments on AliMeeting, a real meeting corpus, to determine the optimal model design for SpatialEmb in terms of both performance and efficiency. Our best model trained with 105 hours Train-Ali-far achieves 17.04% and 20.32% character error rates (CER) on the Eval and Test sets, establishing a new state-of-the-art result with the same training data.
Yiwen Shao, Yong Xu 0004, Sanjeev Khudanpur, Dong Yu 0001
SLT1
2024 Advancing Multi-Talker ASR Performance With Large Language Models
abstract
Recognizing overlapping speech from multiple speakers in conversational scenarios is one of the most challenging problem for automatic speech recognition (ASR). Serialized output training (SOT) is a classic method to address multi-talker ASR, with the idea of concatenating transcriptions from multiple speakers according to the emission times of their speech for training. However, SOT-style transcriptions, derived from concatenating multiple related utterances in a conversation, depend significantly on modeling long contexts. Therefore, compared to traditional methods that primarily emphasize encoder performance in attention-based encoderdecoder (AED) architectures, a novel approach utilizing large language models (LLMs) that leverages the capabilities of pre-trained decoders may be better suited for such complex and challenging scenarios. In this paper, we propose an LLM-based SOT approach for multi-talker ASR, leveraging pre-trained speech encoder and LLM, fine-tuning them on multi-talker dataset using appropriate strategies. Experimental results demonstrate that our approach surpasses traditional AED-based methods on the simulated dataset LibriMix and achieves state-of-the-art performance on the evaluation set of the real-world dataset AMI, outperforming the AED model trained with 1000 times more supervised data in previous works.
Mohan Shi, Zengrui Jin, Yaoxun Xu, Yong Xu 0004, Shixiong Zhang 0001, Yiwen Shao, Dong Yu 0001
SLT7
2022 Multi-Channel Multi-Speaker ASR Using 3D Spatial Feature
abstract
Automatic speech recognition (ASR) of multi-channel multi-speaker overlapped speech remains one of the most challenging tasks to the speech community. In this paper, we look into this challenge by utilizing the location information of target speakers in the 3D space for the first time. To explore the strength of proposed the 3D spatial feature, two paradigms are investigated. 1) a pipelined system with a multi-channel speech separation module followed by the state-of-the-art single-channel ASR module; 2) a "All-In-One" model where the 3D spatial feature is directly used as an input to ASR system without explicit separation modules. Both of them are fully differentiable and can be back-propagated end-to-end. We test them on simulated overlapped speech and real recordings. Experimental results show that 1) the proposed ALL-In-One model achieved a comparable error rate to the pipelined system while reducing the inference time by half; 2) the proposed 3D spatial feature significantly outperformed (31% CERR) all previous works of using the 1D directional information in both paradigms.
Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
ICASSP1
2022 Defense against Adversarial Attacks on Hybrid Speech Recognition System using Adversarial Fine-tuning with Denoiser
Sonal Joshi, Saurabh Kataria 0001, Yiwen Shao, Piotr Zelasko, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH3
2022 Chunking Defense for Adversarial Attacks on ASR
Yiwen Shao, Jesús Villalba 0001, Sonal Joshi, Saurabh Kataria 0001, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH1
2020 Speaker Diarization with Region Proposal Network
abstract
Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and cannot deal with the overlapped speech. In this paper, we propose a novel speaker diarization method: Region Proposal Network based Speaker Diarization (RPNSD). In this method, a neural network generates overlapped speech segment proposals, and compute their speaker embeddings at the same time. Compared with standard diarization systems, RPNSD has a shorter pipeline and can handle the overlapped speech. Experimental results on three diarization datasets reveal that RPNSD achieves remarkable improvements over the state-of-the-art x-vector baseline.
Zili Huang, Shinji Watanabe 0001, Yusuke Fujita, L. Paola García-Perera, Yiwen Shao, Daniel Povey, Sanjeev Khudanpur
ICASSP5
2020 PyChain: A Fully Parallelized PyTorch Implementation of LF-MMI for End-to-End ASR
abstract
We present PYCHAIN, a fully parallelized PyTorch implementation of end-to-end lattice-free maximum mutual information (LF-MMI) training for the so-called chain models in the Kaldi automatic speech recognition (ASR) toolkit.Unlike other Py-Torch and Kaldi based ASR toolkits, PYCHAIN is designed to be as flexible and light-weight as possible so that it can be easily plugged into new ASR projects, or other existing PyTorchbased ASR tools, as exemplified respectively by a new project PYCHAIN-EXAMPLE, and ESPRESSO, an existing end-to-end ASR toolkit.PYCHAIN's efficiency and flexibility is demonstrated through such novel features as full GPU training on numerator/denominator graphs, and support for unequal length sequences.Experiments on the WSJ dataset show that with simple neural networks and commonly used machine learning techniques, PYCHAIN can achieve competitive results that are comparable to Kaldi and better than other end-to-end ASR systems.
Yiwen Shao, Yiming Wang 0006, Daniel Povey, Sanjeev Khudanpur
INTERSPEECH1
2019 Espresso: A Fast End-to-End Neural Speech Recognition Toolkit
abstract
We present Espresso, an open-source, modular, extensible end-to-end neural automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch and the popular neural machine translation toolkit FAIRSEQ. ESRESSO supports distributed training across GPUs and computing nodes, and features various decoding approaches commonly employed in ASR, including look-ahead word-based language model fusion, for which a fast, parallelized decoder is implemented. Espresso achieves state-of-the-art ASR performance on the WSJ, LibriSpeech, and Switchboard data sets among other end-to-end systems without data augmentation, and is 4-11x faster for decoding than similar systems (e.g. ESPNET).
Yiming Wang 0006, Sanjeev Khudanpur, Tongfei Chen, Hainan Xu, Shuoyang Ding, Hang Lv 0001, Yiwen Shao, Nanyun Peng 0001, Lei Xie 0001, Shinji Watanabe 0001
ASRU7
2019 Using ASR Methods for OCR
abstract
Hybrid deep neural network hidden Markov models (DNN-HMM) have achieved impressive results on large vocabulary continuous speech recognition (LVCSR) tasks. However, the recent approaches using DNN-HMM models are not explored much for text recognition. Inspired by the current work in automatic speech recognition (ASR) and machine translation, we present an open vocabulary sub-word text recognition system. The sub-word lexicon and sub-word language model (LM) helps in overcoming the challenge of recognizing out of vocabulary (OOV) words, and a time delay neural network (TDNN) and convolution neural network (CNN) based DNN-HMM optical model (OM) efficiently models the sequence dependency in the line image. We present results on 12 datasets with training data varying from 6k lines to 600k lines. The system is built for 8 languages, i.e., English, French, Arabic, Chinese, Farsi, Tamil, Russian, and Korean. We report competitive results on several commonly used handwritten and printed text datasets.
Ashish Arora, L. Paola García-Perera, Shinji Watanabe 0001, Vimal Manohar, Yiwen Shao, Sanjeev Khudanpur, Chun-Chieh Chang, Babak Rekabdar, Bagher BabaAli, Daniel Povey, David Etter, Desh Raj, Hossein Hadian, Jan Trmal
ICDAR5
2019 Practices of backuping homomorphically encrypted databases
Sa Wang, Yiwen Shao, Yungang Bao
Frontiers Comput. Sci.2
2018 Use of Pitch Continuity for Robust Speech Activity Detection
abstract
Speech activity detection (SAD) is an important component for various speech processing applications and has been researched extensively recently. The pitch continuity, a significant characteristic of speech, however, has not successfully played a role in existing SAD methods. In this work, we propose a novel way to integrate the pitch continuity with pitch-related features. Practice is carried out through the Combo-SAD approach: We examine three consecutive frames and assume that they all have the same pitch as the center frame due to pitch continuity. Corresponding feature values are recomputed at the adjusted pitch location and then used in the final expression. The new combo feature is evaluated with various types of additive noise at different signal-to-noise ratios (SNR). The results show that the new feature leads to better SAD performance (with an up to 39.3% relative improvement on miss rate compared to Combo-SAD). We also introduce a novel variant of the underlying autocorrelation function and illustrate how it can improve the accuracy of pitch detection.
Yiwen Shao, Qiguang Lin
ICASSP1
2018 A Novel Normalization Method for Autocorrelation Function for Pitch Detection and for Speech Activity Detection
Qiguang Lin, Yiwen Shao
INTERSPEECH2