VLDB 2026 Research / reviewers in the wild / expert
Tao Wei 0003
dblp:64/5099-3
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2025
0009-0000-2134-7984ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC LossabstractContextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the alignment between the ASR and bias tasks has been largely neglected. This paper introduces a novel contextual biasing approach that employs the Multi-label Synchronous Output CTC (MCTC) algorithm to enhance the synchronization between ASR and bias task outputs. Furthermore, we propose an enhancement to phrase-level contextual representation. Our proposed method demonstrates significant improvements in Word Error Rate (WER). Specifically, our experiments reveal a 15.0% reduction in WER on the Librispeech-960 dataset compared to the CPPN method, with an impressive 29.0% reduction in WER for context phrases. Tao Wei 0003, Ziyang Zhuang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2025 | Token-Level Contextual Network with Ladder-Shaped Attention for End-to-End ASRabstractContextual automatic speech recognition (ASR) plays an increasingly important role in addressing the long-tail issues of general ASR. In the past, contextual ASR mainly focused on phrase-level discussions, providing a convenient way to handle biasing phrases. This paper introduces a new contextual network for extracting context tokens, with a focus on optimizing contextual knowledge using token-level information. We enhance the fusion of context information and ASR acoustic features, recognizing that textual knowledge inherently represents a distinct modality compared to acoustic information. More importantly, we propose a creative approach to address the challenge of the increasing size of expanding token list compared to phrase list. Our method is tested on the LibriSpeech and AISHELL-2 datasets, the results demonstrate a 31.0%/23.3% Word Error Rate (WER) reduction on LibriSpeech and a 26.9% reduction Character Error Rate (CER) on the named entity (NE) set from AISHELL-2. Tao Wei 0003, Ziyang Zhuang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 3 |
| 2025 | Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning DistillationabstractThe rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation can help small models acquire reasoning capabilities, but most existing methods focus primarily on improving teacher-generated reasoning paths. Our observations reveal that small models can generate high-quality reasoning paths during sampling, even without chain-of-thought prompting, though these paths are often latent due to their low probability under standard decoding strategies. To address this, we propose Self-Enhanced Reasoning Training (SERT), which activates and leverages latent reasoning capabilities in small models through self-training on filtered, self-generated reasoning paths under zero-shot conditions. Experiments using OpenAI’s GPT-3.5 as the teacher model and GPT-2 models as the student models demonstrate that SERT enhances the reasoning abilities of small models, improving their performance in reasoning distillation. Yong Zhang 0058, Zhitao Li 0002, Ming Li 0010, Ning Cheng 0001, Minchuan Chen, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 7 |
| 2025 | EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference SpeedabstractNon-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. In this paper, we propose a single-step NAR ASR architecture with high accuracy and inference speed, called EffectiveASR. It uses an Index Mapping Vector (IMV) based alignment generator to generate alignments during training, and an alignment predictor to learn the alignments for inference. It can be trained end-to-end (E2E) with cross-entropy loss combined with alignment loss. The proposed EffectiveASR achieves competitive results on the AISHELL-1 and AISHELL-2 Mandarin benchmarks compared to the leading models. Specifically, it achieves character error rates (CER) of 4.26%/4.62% on the AISHELL-1 dev/test dataset, which outperforms the AR Conformer with about 30x inference speedup. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 5 |
| 2025 | Enhancing Serialized Output Training for Multi-Talker ASR with Soft Monotonic Alignment and Utterance-level Timestamp
Fengyun Tan, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 2 |
| 2025 | Towards Efficiently Whisper Fine-tuning with Monotonic Alignments
Ziyang Zhuang, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 2 |
| 2024 | Improving Attention-Based End-to-End Speech Recognition by Monotonic Alignment Attention Matrix ReconstructionabstractIn automatic speech recognition (ASR) task, the output sequence should correspond to a linear transcription of the input sequence. Lots of works have been done to learn the monotonic alignment in end-to-end (E2E) ASR model, but their methods mainly focus on streaming propose and usually result in a decline in ASR performance. On the contrary, some studies have shown that for non-streaming attention-based models, monotonic alignment is beneficial to model performance. Based on this motivation, we propose the enhanced Gaussian Monotonic Alignment (e-GMA), which reduces the difficulty of learning monotonic alignment, and the reconstructed attention matrix leads to an improved accuracy in ASR tasks. Experiments on the LibriSpeech dataset demonstrate the effectiveness of the proposed approach. Comparing with a strong baseline obtained from WeNet, the proposed model yields 12.2% relative WER reduction on test-clean benchmark and 9.9% on test-other. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Jing Xiao 0006 |
ICASSP | 5 |
| 2024 | E-Paraformer: A Faster and Better Parallel Transformer for Non-autoregressive End-to-End Mandarin Speech Recognition
Fengyun Tan, Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Shaodan Zhai, Jing Xiao 0006 |
INTERSPEECH | 5 |
| 2023 | DASA: Difficulty-Aware Semantic Augmentation for Speaker VerificationabstractData augmentation is vital to the generalization ability and robustness of deep neural networks (DNNs) models. Existing augmentation methods for speaker verification manipulate the raw signal, which are time-consuming and the augmented samples lack diversity. In this paper, we present a novel difficulty-aware semantic augmentation (DASA) approach for speaker verification, which can generate diversified training samples in speaker embedding space with negligible extra computing cost. Firstly, we augment training samples by perturbing speaker embeddings along semantic directions, which are obtained from speaker-wise covariance matrices. Secondly, accurate covariance matrices are estimated from robust speaker embeddings during training, so we introduce difficulty-aware additive margin softmax (DAAM-Softmax) to obtain optimal speaker embeddings. Finally, we assume the number of augmented samples goes to infinity and derive a closed-form upper bound of the expected loss with DASA, which achieves compatibility and efficiency. Extensive experiments demonstrate the proposed approach can achieve a remarkable performance improvement. The best result achieves a 14.6% relative reduction in EER metric on CN-Celeb evaluation set. Yang Zhang 0025, Zhiyong Wu 0001, Tao Wei 0003, Helen M. Meng |
ICASSP | 5 |
| 2023 | Improving End-to-End Modeling For Mandarin-English Code-Switching Using Lightweight Switch-Routing Mixture-of-Experts
Fengyun Tan, Chaofeng Feng, Tao Wei 0003, Shuai Gong, Jinqiang Leng, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2022 | Towards Efficiently Learning Monotonic Alignments for Attention-based End-to-End Speech Recognition
Chenfeng Miao, Ziyang Zhuang, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2022 | FFM: A Frame Filtering Mechanism To Accelerate Inference Speed For Conformer In Speech Recognition
Zongfeng Quan, Nick J. C. Wang, Tao Wei 0003, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2021 | SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech RecognitionabstractInspired by the contrastive predictive coding (CPC), we propose a feature representation scheme for automatic speech recognition (ASR), which encodes sequential dependency information from raw audio signals. Following the original CPC, for a given frame, mutual information (MI) lower bound is maximized between historical context and future prediction. While computing the MI lower bound, based on original CPC, we develop the sequential CPC (SEQ-CPC), which takes the sequential information between frames into consideration. Since speech frames are not independent events, incorporating sequential information leads to better recognition performance. Experimental results on WSJ corpus show that SEQ-CPC achieves the best performance than CPC and NCE which is the contrastive objective used in wav2vec. Haimei Kang, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 7 |