VLDB 2026 Research / reviewers in the wild / expert
Ziyang Zhuang
dblp:330/8999
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC LossabstractContextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the alignment between the ASR and bias tasks has been largely neglected. This paper introduces a novel contextual biasing approach that employs the Multi-label Synchronous Output CTC (MCTC) algorithm to enhance the synchronization between ASR and bias task outputs. Furthermore, we propose an enhancement to phrase-level contextual representation. Our proposed method demonstrates significant improvements in Word Error Rate (WER). Specifically, our experiments reveal a 15.0% reduction in WER on the Librispeech-960 dataset compared to the CPPN method, with an impressive 29.0% reduction in WER for context phrases. Tao Wei 0003, Ziyang Zhuang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2025 | Token-Level Contextual Network with Ladder-Shaped Attention for End-to-End ASRabstractContextual automatic speech recognition (ASR) plays an increasingly important role in addressing the long-tail issues of general ASR. In the past, contextual ASR mainly focused on phrase-level discussions, providing a convenient way to handle biasing phrases. This paper introduces a new contextual network for extracting context tokens, with a focus on optimizing contextual knowledge using token-level information. We enhance the fusion of context information and ASR acoustic features, recognizing that textual knowledge inherently represents a distinct modality compared to acoustic information. More importantly, we propose a creative approach to address the challenge of the increasing size of expanding token list compared to phrase list. Our method is tested on the LibriSpeech and AISHELL-2 datasets, the results demonstrate a 31.0%/23.3% Word Error Rate (WER) reduction on LibriSpeech and a 26.9% reduction Character Error Rate (CER) on the named entity (NE) set from AISHELL-2. Tao Wei 0003, Ziyang Zhuang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2025 | EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference SpeedabstractNon-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. In this paper, we propose a single-step NAR ASR architecture with high accuracy and inference speed, called EffectiveASR. It uses an Index Mapping Vector (IMV) based alignment generator to generate alignments during training, and an alignment predictor to learn the alignments for inference. It can be trained end-to-end (E2E) with cross-entropy loss combined with alignment loss. The proposed EffectiveASR achieves competitive results on the AISHELL-1 and AISHELL-2 Mandarin benchmarks compared to the leading models. Specifically, it achieves character error rates (CER) of 4.26%/4.62% on the AISHELL-1 dev/test dataset, which outperforms the AR Conformer with about 30x inference speedup. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 1 |
| 2025 | Towards Efficiently Whisper Fine-tuning with Monotonic Alignments
Ziyang Zhuang, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 1 |
| 2024 | Improving Attention-Based End-to-End Speech Recognition by Monotonic Alignment Attention Matrix ReconstructionabstractIn automatic speech recognition (ASR) task, the output sequence should correspond to a linear transcription of the input sequence. Lots of works have been done to learn the monotonic alignment in end-to-end (E2E) ASR model, but their methods mainly focus on streaming propose and usually result in a decline in ASR performance. On the contrary, some studies have shown that for non-streaming attention-based models, monotonic alignment is beneficial to model performance. Based on this motivation, we propose the enhanced Gaussian Monotonic Alignment (e-GMA), which reduces the difficulty of learning monotonic alignment, and the reconstructed attention matrix leads to an improved accuracy in ASR tasks. Experiments on the LibriSpeech dataset demonstrate the effectiveness of the proposed approach. Comparing with a strong baseline obtained from WeNet, the proposed model yields 12.2% relative WER reduction on test-clean benchmark and 9.9% on test-other. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Jing Xiao 0006 |
ICASSP | 1 |
| 2024 | E-Paraformer: A Faster and Better Parallel Transformer for Non-autoregressive End-to-End Mandarin Speech Recognition
Fengyun Tan, Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Shaodan Zhai, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2022 | Towards Efficiently Learning Monotonic Alignments for Attention-based End-to-End Speech Recognition
Chenfeng Miao, Ziyang Zhuang, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 3 |