Zelin Wu

dblp:195/0252 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DiffAC-Seg: A Diffusion Model Enhanced with Attention and Contextual Features for Stroke Lesion Segmentation
Fenglian Li, Lixia Huang, Guijun Chen, Zelin Wu
PRCV (13)5
2024 Stroke-CVAE-CGAN: A Medical Imaging Data Augmentation Network Incorporating Stroke Lesion Distribution
abstract
In recent years, the incidence of stroke has significantly increased, posing a serious threat to public health. Accurate stroke lesion segmentation techniques can assist physicians in promptly formulating appropriate treatment plans based on specific patient conditions, significantly reducing the risk of disability and mortality, thereby improving patient outcomes. Against this backdrop, leveraging its powerful representation and reasoning capabilities, deep learning has emerged as a key research direction in the field of medical image processing. However, deep learning-based stroke lesion segmentation methods rely on a large amount of precisely labeled medical imaging data, the acquisition of which often faces challenges such as high costs, insufficient quantities, and time-intensive efforts. Traditional data augmentation methods provide some relief but are still limited by data distribution and diversity constraints. Addressing these issues, this paper introduces a novel data generation model, Stroke-CVAE-CGAN, which generates “synthetic lesion masks” based on the spatial distribution characteristics of stroke lesions, serving as constraints for Conditional Generative Adversarial Networks (CGANs), thus enabling the augmented training data that aligns with the distribution patterns of real stroke lesions. Experiments conducted on the ATLAS stroke lesion segmentation dataset show that the augmented data generated by Stroke-CVAE-CGAN closely matches the training data in terms of distribution and exhibits superior Frechet Inception Distance (FID) quality. Utilizing this augmented data to train the U-Net segmentation model significantly enhances the accuracy of stroke lesion segmentation.
Haisheng Hui, Fenglian Li, Zelin Wu
DSAA4
2024 Optimizing Large-Scale Context Retrieval for End-to-End ASR
Diamantino Caseiro, Kandarp Joshi, Christopher Li, Pat Rondon, Zelin Wu, Petr Zadrazil, Lillian Zhou
INTERSPEECH6
2024 Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, Rohit Prabhavalkar, Cal Peyser, Nanxin Chen, Tara N. Sainath, Bhuvana Ramabhadran
INTERSPEECH2
2024 Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro J. Moreno 0001
INTERSPEECH2
2023 Contextual Spelling Correction with Large Language Models
abstract
Contextual Spelling Correction (CSC) models are used to improve automatic speech recognition (ASR) quality given userspecific context. Typically, context is modeled as a large set of text spans to compare against a given ASR hypothesis using some distance measure (text, phonetic, or neural embedding). In this work we propose a CSC system based on a single Large Language Model (LLM) adapted with prompt tuning. Our approach is shown to be data efficient, and does not require dedicated serving. Our system exhibits advanced contextualization capabilities, such as support for phonetic spellings, cross-lingual scripts, and context specified as topics, with little to no data engineering. On voice assistant datasets, our system achieves $7.8 \%$ absolute word error rate reduction from a reference ASR system with relevant context and improving upon other contextualization solutions. Finally, we test our system in a prompt-injection attack scenario and report vulnerabilities and mitigations.
Gan Song, Zelin Wu, Golan Pundak, Angad Chandorkar, Kandarp Joshi, Xavier Velez, Diamantino Caseiro, Ben Haynor, Nikhil Siddhartha, Pat Rondon, Khe Chai Sim
ASRU2
2023 SLM: Bridge the Thin Gap Between Speech and Text Foundation Models
abstract
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter with just 1% (156M) of the foundation models’ parameters. This adaptation not only leads SLM to achieve strong performance on conventional tasks such as automatic speech recognition (ASR) and automatic speech translation (AST), but also unlocks the novel capability of zero-shot instruction-following for more diverse tasks. Given a speech input and a text instruction, SLM is able to perform unseen generation tasks including contextual biasing ASR using real-time context, dialog generation, speech continuation, and question answering. Our approach demonstrates that the representational gap between pretrained speech and language models is narrower than one would expect, and can be bridged by a simple adaptation mechanism. As a result, SLM is not only efficient to train, but also inherits strong capabilities already present in foundation models of different modalities.
Mingqiu Wang, Wei Han 0002, Izhak Shafran, Zelin Wu, Chung-Cheng Chiu, Yuan Cao 0007, Nanxin Chen, Yu Zhang 0033, Hagen Soltau, Paul K. Rubenstein, Lukas Zilka, Golan Pundak, Nikhil Siddhartha, Johan Schalkwyk
ASRU4
2023 Dual-Mode NAM: Effective Top-K Context Injection for End-to-End ASR
Zelin Wu, Tsendsuren Munkhdalai, Pat Rondon, Golan Pundak, Khe Chai Sim, Christopher Li
INTERSPEECH1
2023 W-Net: A boundary-enhanced segmentation network for stroke lesions
Zelin Wu, Fenglian Li, Suzhe Wang, Lixia Huang
Expert Syst. Appl.1
2022 Streaming Intended Query Detection using E2E Modeling for Continued Conversation
abstract
In voice-enabled applications, a predetermined hotword is usually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each time introduces a cognitive burden in continued conversations.To avoid repeating a hotword, we propose a streaming end-to-end (E2E) intended query detector that identifies the utterances directed towards the device and filters out other utterances not directed towards device.The proposed approach incorporates the intended query detector into the E2E model that already folds different components of the speech recognition pipeline into one neural network.The E2E modeling on speech decoding and intended query detection also allows us to declare a quick intended query detection based on early partial recognition result, which is important to decrease latency and make the system responsive.We demonstrate that the proposed E2E approach yields a 22% relative improvement on equal error rate (EER) for the detection accuracy and 600 ms latency improvement compared with an independent intended query detector.In our experiment, the proposed model detects whether the user is talking to the device with a 8.7% EER within 1.4 seconds of median latency after user starts speaking.
Shuo-Yiin Chang, Guru Prakash Arumugam, Zelin Wu, Tara N. Sainath, Bo Li 0028, Qiao Liang 0001, Adam Stambler, Shyam Upadhyay, Manaal Faruqui, Trevor Strohman
INTERSPEECH3
2022 NAM+: Towards Scalable End-to-End Contextual Biasing for Adaptive ASR
abstract
Attention-based biasing techniques for end-to-end ASR systems are able to achieve large accuracy gains without requiring the inference algorithm adjustments and parameter tuning common to fusion approaches. However, it is challenging to simultaneously scale up attention-based biasing to realistic numbers of biased phrases; maintain in-domain WER gains, while minimizing out-of-domain losses; and run in real time. We present NAM+, an attention-based biasing approach which achieves a 16X inference speedup per acoustic frame over prior work when run with 3,000 biasing entities, as measured on a typical mobile CPU. NAM+ achieves these run-time gains through a combination of Two-Pass Hierarchical Attention and Dilated Context Update. Compared to the adapted baseline, NAM+ further decreases the in-domain WER by up to 12.6% relative, while incurring an out-of-domain WER regression of 20% relative. Compared to the non-adapted baseline, the out-of-domain WER regression is 7.1 % relative.
Tsendsuren Munkhdalai, Zelin Wu, Golan Pundak, Khe Chai Sim, Pat Rondon, Tara N. Sainath
SLT2
2021 A Deliberation-Based Joint Acoustic and Text Decoder
abstract
We propose a new two-pass E2E speech recognition model that improves ASR performance by training on a combination of paired data and unpaired text data.Previously, the joint acoustic and text decoder (JATD) has shown promising results through the use of text data during model training and the recently introduced deliberation architecture has reduced recognition errors by leveraging first-pass decoding results.Our method, dubbed Deliberation-JATD, combines the spelling correcting abilities of deliberation with JATD's use of unpaired text data to further improve performance.The proposed model produces substantial gains across multiple test sets, especially those focused on rare words, where it reduces word error rate (WER) by between 12% and 22.5% relative.This is done without increasing model size or requiring multi-stage training, making Deliberation-JATD an efficient candidate for on-device applications.
Sepand Mavandadi, Tara N. Sainath, Zelin Wu
Interspeech4
2020 Multistate Encoding with End-To-End Speech RNN Transducer Network
abstract
Recurrent Neural Network Transducer (RNN-T) models [1] for automatic speech recognition (ASR) provide high accuracy speech recognition. Such end-to-end (E2E) models combine acoustic, pronunciation and language models (AM, PM, LM) of a conventional ASR system into a single neural network, dramatically reducing complexity and model size.In this paper, we propose a technique for incorporating contextual signals, such as intelligent assistant device state or dialog state, directly into RNN-T models. We explore different encoding methods and demonstrate that RNN-T models can effectively utilize such context. Our technique results in reduction in Word Error Rate (WER) of up to 10.4% relative on a variety of contextual recognition tasks. We also demonstrate that proper regularization can be used to model context independently for improved overall quality.
Zelin Wu, Bo Li 0028, Yu Zhang 0033, Petar S. Aleksic, Tara N. Sainath
ICASSP1
2019 Speech Recognition with Augmented Synthesized Speech
abstract
Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific, human speech that is used to train speech recognizers. The multi-speaker speech synthesis architecture can learn latent embedding spaces of prosody, speaker and style variations derived from input acoustic representations thereby allowing for manipulation of the synthesized speech. In this paper, we evaluate the feasibility of enhancing speech recognition performance using speech synthesis using two corpora from different domains. We explore algorithms to provide the necessary acoustic and lexical diversity needed for robust speech recognition. Finally, we demonstrate the feasibility of this approach as a data augmentation strategy for domain-transfer. We find that improvements to speech recognition performance is achievable by augmenting training data with synthesized material. However, there remains a substantial gap in performance between recognizers trained on human speech those trained on synthesized speech.
Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Ye Jia, Pedro J. Moreno 0001, Zelin Wu
ASRU7
2019 Semi-supervised Training for End-to-end Models via Weak Distillation
abstract
End-to-end (E2E) models are a promising research direction in speech recognition, as the single all-neural E2E system offers a much simpler and more compact solution compared to a conventional model, which has a separate acoustic (AM), pronunciation (PM) and language model (LM). However, it has been noted that E2E models perform poorly on tail words and proper nouns, likely because the end-to-end optimization requires joint audio-text pairs, and does not take advantage of additional lexicons and large amounts of text-only data used to train the LMs in conventional models. There has been numerous efforts in training an RNN-LM on text-only data and fusing it into the end-to-end model. In this work, we contrast this approach to training the E2E model with audio-text pairs generated from unsupervised speech data. To target the proper noun issue specifically, we adopt a Part-of-Speech (POS) tagger to filter the unsupervised data to use only those with proper nouns. We show that training with filtered unsupervised-data provides up to a 13% relative reduction in word-error-rate (WER), and when used in conjunction with a cold-fusion RNN-LM, up to a 17% relative improvement.
Bo Li 0028, Tara N. Sainath, Ruoming Pang, Zelin Wu
ICASSP4
2019 Improving Performance of End-to-End ASR on Numeric Sequences
abstract
Recognizing written domain numeric utterances (e.g., I need $1.25.) can be challenging for ASR systems, particularly when numeric sequences are not seen during training.This out-ofvocabulary (OOV) issue is addressed in conventional ASR systems by training part of the model on spoken domain utterances (e.g., I need one dollar and twenty five cents.),for which numeric sequences are composed of in-vocabulary numbers, and then using an FST verbalizer to denormalize the result.Unfortunately, conventional ASR models are not suitable for the low memory setting of on-device speech recognition.E2E models such as RNN-T are attractive for on-device ASR, as they fold the AM, PM and LM of a conventional model into one neural network.However, in the on-device setting the large memory footprint of an FST denormer makes spoken domain training more difficult.In this paper, we investigate techniques to improve E2E model performance on numeric data.We find that using a text-to-speech system to generate additional numeric training data, as well as using a small-footprint neural network to perform spoken-to-written domain denorming, yields improvement in several numeric classes.In the case of the longest numeric sequences, we see reduction of WER by up to a factor of 8.
Cal Peyser, Hao Zhang 0010, Tara N. Sainath, Zelin Wu
INTERSPEECH4
2019 VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
abstract
In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker.We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings;(2) A spectrogram masking network that takes both noisy spectrogram and speaker embedding as input, and produces a mask.Our system significantly reduces the speech recognition WER on multi-speaker signals, with minimal WER degradation on single-speaker signals.
Hannah Muckenhirn, Kevin W. Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, Ignacio López-Moreno
INTERSPEECH5
2016 Unsupervised context learning for speech recognition
abstract
It has been shown in the literature that automatic speech recognition systems can greatly benefit from contextual information [1, 2, 3, 4, 5]. Contextual information can be used to simplify the beam search and improve recognition accuracy. Types of useful contextual information can include the name of the application the user is in, the contents of the user's phone screen, the user's location, a certain dialog state, etc. Building a separate language model for each of these types of context is not feasible due to limited resources or limited amounts of training data. In this paper we describe an approach for unsupervised learning of contextual information and automatic building of contextual biasing models. Our approach can be used to build a large number of small contextual models from a limited amount of available unsupervised training data. We describe how n-grams relevant for a particular context are automatically selected as well as how an optimal size of a final contextual model is chosen. Our experimental results show great accuracy improvements for several types of context.
Assaf Hurwitz Michaely, Mohammadreza Ghodsi, Zelin Wu, Justin Scheiner, Petar S. Aleksic
SLT3