VLDB 2026 Research / reviewers in the wild / expert
Zhanheng Yang
dblp:277/9424
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Contextualized End-to-End Speech Recognition with Contextual Phrase Prediction Network
Kaixun Huang, Zhanheng Yang, Bingshen Mu, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2023 | Adaptive Contextual Biasing for Transducer Based Streaming Speech Recognition
Zhanheng Yang, Kaixun Huang, Changru Chen, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2023 | Two Stage Contextual Word Filtering for Context Bias in Unified Streaming and Non-streaming Transducer
Zhanheng Yang, Sining Sun, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2022 | Minimizing Sequential Confusion Error in Speech Command RecognitionabstractSpeech command recognition (SCR) has been commonly used on resource constrained devices to achieve hands-free user experience.However, in real applications, confusion among commands with similar pronunciations often happens due to the limited capacity of small models deployed on edge devices, which drastically affects the user experience.In this paper, inspired by the advances of discriminative training in speech recognition, we propose a novel minimize sequential confusion error (MSCE) training criterion particularly for SCR, aiming to alleviate the command confusion problem.Specifically, we aim to improve the ability of discriminating the target command from other commands on the basis of MCE discriminative criteria.We define the likelihood of different commands through connectionist temporal classification (CTC).During training, we propose several strategies to use prior knowledge creating a confusing sequence set for similar-sounding command instead of creating the whole non-target command set, which can better save the training resources and effectively reduce command confusion errors.Specifically, we design and compare three different strategies for confusing set construction.By using our proposed method, we can relatively reduce the False Reject Rate (FRR) by 33.7% at 0.01 False Alarm Rate (FAR) and confusion errors by 18.28% on our collected speech command set. Zhanheng Yang, Hang Lv 0001, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2022 | CaTT-KWS: A Multi-stage Customized Keyword Spotting Framework based on Cascaded Transducer-TransformerabstractCustomized keyword spotting (KWS) has great potential to be deployed on edge devices to achieve hands-free user experience. However, in real applications, false alarm (FA) would be a serious problem for spotting dozens or even hundreds of keywords, which drastically affects user experience. To solve this problem, in this paper, we leverage the recent advances in transducer and transformer based acoustic models and propose a new multi-stage customized KWS framework named Cascaded Transducer-Transformer KWS (CaTT-KWS), which includes a transducer based keyword detector, a frame-level phone predictor based force alignment module and a transformer based decoder. Specifically, the streaming transducer module is used to spot keyword candidates in audio stream. Then force alignment is implemented using the phone posteriors predicted by the phone predictor to finish the first stage keyword verification and refine the time boundaries of keyword. Finally, the transformer decoder further verifies the triggered keyword. Our proposed CaTT-KWS framework reduces FA rate effectively without obviously hurting keyword recognition accuracy. Specifically, we can get impressively 0.13 FA per hour on a challenging dataset, with over 90% relative reduction on FA comparing to the transducer based detection model, while keyword recognition accuracy only drops less than 2%. Zhanheng Yang, Sining Sun, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2021 | IEEE SLT 2021 Alpha-Mini Speech Challenge: Open Datasets, Tracks, Rules and BaselinesabstractThe IEEE Spoken Language Technology Workshop (SLT) 2021 Alpha-mini Speech Challenge (ASC) is intended to improve research on keyword spotting (KWS) and sound source location (SSL) on humanoid robots. Many publications report significant improvements in deep learning based KWS and SSL on open source datasets in recent years. For deep learning model training, it is necessary to expand the data coverage to improve the model robustness. Thus, simulating multi-channel noisy and reverberant data from single-channel speech, noise, echo and room impulsive response (RIR) is widely adopted. However, this approach may generate mismatch between simulated data and recorded data in real application scenarios, especially echo data. In this challenge, we open source a sizable speech, keyword, echo and noise corpus for promoting data-driven methods, particularly deep-learning approaches on KWS and SSL. We also choose Alpha-mini, a humanoid robot produced by UBTECH equipped with a built-in four-microphone array on its head, to record development and evaluation sets under the actual Alpha-mini robot application scenario, including environ-mental noise as well as echo and mechanical noise generated by the robot itself for model evaluation. Furthermore, we illustrate the rules, evaluation methods and baselines for re-searchers to quickly assess their achievements and optimize their models. Yihui Fu, Zhuoyuan Yao, Weipeng He, Jian Wu 0027, Zhanheng Yang, Lei Xie 0001, Dong-Yan Huang, Hui Bu, Petr Motlícek, Jean-Marc Odobez |
SLT | 6 |