VLDB 2026 Research / reviewers in the wild / expert
Yang Zhang 0089
dblp:06/6785-89
· DBLP profile ↗
8ranked-venue papers
4as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Chat about Boring Problems: Studying GPT-Based Text NormalizationabstractText normalization - the conversion of text from written to spoken form - is traditionally assumed to be an ill-formed task for language modeling. In this work, we argue otherwise. We empirically show the capacity of Large-Language Models (LLM) for text normalization in few-shot scenarios. Combining self-consistency reasoning with linguistic-informed prompt engineering, we find LLM-based text normalization to achieve error rates approximately 40% lower than production-level normalization systems. Further, upon error analysis, we note key limitations in the conventional design of text normalization tasks. We create a new taxonomy of text normalization errors and apply it to results from GPT-3.5-Turbo and GPT-4.0. Through this new framework, we identify strengths and weaknesses of LLM-based TN, opening opportunities for future work. Yang Zhang 0089, Travis M. Bartley, Mariana Graterol-Fuenmayor, Vitaly Lavrukhin, Evelina Bakhturina, Boris Ginsburg |
ICASSP | 1 |
| 2023 | Conformer-Based Target-Speaker Automatic Speech Recognition For Single-Channel AudioabstractWe propose CONF-TSASR, a non-autoregressive end-to-end time-frequency domain architecture for single-channel target-speaker automatic speech recognition (TS-ASR). The model consists of a TitaNet based speaker embedding module, a Conformer based masking as well as ASR modules. These modules are jointly optimized to transcribe a target-speaker, while ignoring speech from other speakers. For training we use Connectionist Temporal Classification (CTC) loss and introduce a scale-invariant spectrogram reconstruction loss to encourage the model better separate the target-speaker’s spectrogram from mixture. We obtain state-of-the-art target-speaker word error rate (TS-WER) on WSJ0-2mix-extr (4.2%). Further, we report for the first time TS-WER on WSJ0-3mix-extr (12.4%), LibriSpeech2Mix (4.2%) and LibriSpeech3Mix (7.6%) datasets, establishing new benchmarks for TS-ASR. The proposed model will be open-sourced through NVIDIA NeMo toolkit. Yang Zhang 0089, Krishna C. Puvvada, Vitaly Lavrukhin, Boris Ginsburg |
ICASSP | 1 |
| 2022 | Shallow Fusion of Weighted Finite-State Transducer and Language Model for Text NormalizationabstractText normalization (TN) systems in production are largely rule-based using weighted finite-state transducers (WFST).However, WFST-based systems struggle with ambiguous input when the normalized form is context-dependent.On the other hand, neural text normalization systems can take context into account but they suffer from unrecoverable errors and require labeled normalization datasets, which are hard to collect.We propose a new hybrid approach that combines the benefits of rule-based and neural systems.First, a non-deterministic WFST outputs all normalization candidates, and then a neural language model picks the best one -similar to shallow fusion for automatic speech recognition.While the WFST prevents unrecoverable errors, the language model resolves contextual ambiguity.The approach is easy to extend and we show it is effective.It achieves comparable or better results than existing state-of-theart TN models. Evelina Bakhturina, Yang Zhang 0089, Boris Ginsburg |
INTERSPEECH | 2 |
| 2021 | Hi-Fi Multi-Speaker English TTS DatasetabstractThis paper introduces a new multi-speaker English dataset for training text-to-speech models.The dataset is based on Lib-riVox audiobooks and Project Gutenberg texts, both in the public domain.The new dataset contains about 292 hours of speech from 10 speakers with at least 17 hours per speaker sampled at 44.1 kHz.To select speech samples with high quality, we considered audio recordings with a signal bandwidth of at least 13 kHz and a signal-to-noise ratio (SNR) of at least 32 dB.The dataset is publicly released at "http://www.openslr.org/109/". Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, Yang Zhang 0089 |
Interspeech | 4 |
| 2021 | NeMo (Inverse) Text Normalization: From Development to Production
Yang Zhang 0089, Evelina Bakhturina, Boris Ginsburg |
Interspeech | 1 |
| 2021 | NeMo Inverse Text Normalization: From Development to ProductionabstractInverse text normalization (ITN) converts spoken-domain automatic speech recognition (ASR) output into written-domain text to improve the readability of the ASR output. Many state-of-the-art ITN systems use hand-written weighted finite-state transducer(WFST) grammars since this task has extremely low tolerance to unrecoverable errors. We introduce an open-source Python WFST-based library for ITN which enables a seamless path from development to production. We describe the specification of ITN grammar rules for English, but the library can be adapted for other languages. It can also be used for written-to-spoken text normalization. We evaluate the NeMo ITN library using a modified version of the Google Text normalization dataset. Yang Zhang 0089, Evelina Bakhturina, Kyle Gorman, Boris Ginsburg |
Interspeech | 1 |
| 2020 | BioMegatron: Larger Biomedical Domain Language ModelabstractHoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, Raghav Mani. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Hoo-Chang Shin, Yang Zhang 0089, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, Raghav Mani |
EMNLP (1) | 2 |
| 2020 | Quartznet: Deep Automatic Speech Recognition with 1D Time-Channel Separable ConvolutionsabstractWe propose a new end-to-end neural acoustic model for automatic speech recognition. The model is composed of multiple blocks with residual connections between them. Each block consists of one or more modules with 1D time-channel separable convolutional layers, batch normalization, and ReLU layers. It is trained with CTC loss. The proposed network achieves near state-of-the-art accuracy on LibriSpeech and Wall Street Journal, while having fewer parameters than all competing models. We also demonstrate that this model can be effectively fine-tuned on new datasets. Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li 0007, Yang Zhang 0089 |
ICASSP | 9 |