VLDB 2026 Research / reviewers in the wild / expert
Yutong Pang
dblp:254/1854
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-4829-4741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic dataabstractWe introduce DAS (Domain Adaptation with Synthetic data), a novel domain adaptation framework for pre-trained ASR model, designed to efficiently adapt to various language-defined domains without requiring any real data. In particular, DAS first prompts large language models (LLMs) to generate domain-specific texts before converting these texts to speech via text-to-speech technology. The synthetic data is used to finetune Whisper with Low-Rank Adapters (LoRAs) for targeted domains such as music, weather, and sports. We introduce a novel one-pass decoding strategy that merges predictions from multiple LoRA adapters efficiently during the auto-regressive text generation process. Experimental results show significant improvements, reducing the Word Error Rate (WER) by 10% to 17% across all target domains compared to the original model, with minimal performance regression in out-of-domain settings −1% on Librispeech test sets. We also demonstrate that DAS operates efficiently during inference, introducing an additional 9% increase in Real Time Factor (RTF) compared to the original model when inferring with three LoRA adapters. Yutong Pang, Debjyoti Paul, Laxmi Pandey, Kevin Jiang, Jinxi Guo, Ke Li 0023 |
ICASSP | 2 |
| 2025 | R2S: Real-to-Synthetic Representation Learning for Training Speech Recognition Models on Synthetic Data
Debjyoti Paul, Yutong Pang, Laxmi Pandey, Jinxi Guo, Ke Li 0023 |
INTERSPEECH | 3 |
| 2024 | Contextual Biasing of Named-Entities with Large Language ModelsabstractWe explore contextual biasing with Large Language Models (LLMs) to enhance Automatic Speech Recognition (ASR) in second-pass rescoring. Our approach introduces the utilization of prompts for LLMs during rescoring without the need for fine-tuning. These prompts incorporate a biasing list and a set of few-shot examples, serving as supplementary sources of information when evaluating the hypothesis score. Furthermore, we introduce multi-task training for LLMs to predict entity class and the subsequent token. To address sequence length constraints and improve the efficiency of contextual biasing, we propose dynamic prompting based on class tag predictions. Through dynamic prompting, we leverage the class tag predictions to identify the most probable entity class and subsequently utilize entities within this class as biasing context for the next token prediction. We evaluate the performance of proposed methods in terms of Word Error Rate (WER) on an internal entity-heavy and the SLUE-Voxpopuli datasets. Our results show significant improvements: biasing lists and few-shot examples achieved a relative improvement of 17.8% and 9.6%, while multitask training and dynamic prompting achieved 20.0% and 11.3% relative WER improvement, respectively. Chuanneng Sun, Yingyi Ma, Zhe Liu 0011, Lucas Kabela, Yutong Pang, Ozlem Kalinli |
ICASSP | 6 |
| 2024 | Recovering from Privacy-Preserving Masking with Large Language ModelsabstractModel adaptation is crucial to handle the discrepancy between proxy training data and actual users’ data received. To effectively perform adaptation, textual data of users is typically stored on servers or their local devices, where downstream natural language processing (NLP) models can be directly trained using such in-domain data. However, this might raise privacy and security concerns due to the extra risks of exposing user information to adversaries. Replacing identifying information in textual data with a generic marker has been recently explored. In this work, we leverage large language models (LLMs) to suggest substitutes of masked tokens and have their effectiveness evaluated on downstream language modeling tasks. Specifically, we propose multiple pre-trained and fine-tuned LLM-based approaches and perform empirical studies on various datasets for the comparison of these methods. Experimental results show that models trained on the obfuscation corpora are able to achieve comparable performance with the ones trained on the original data without privacy-preserving token masking. Arpita Vats, Zhe Liu 0011, Debjyoti Paul, Yingyi Ma, Yutong Pang, Ozlem Kalinli |
ICASSP | 6 |
| 2023 | Language Agnostic Data-Driven Inverse Text Normalization
Szu-Jui Chen, Debjyoti Paul, Yutong Pang |
INTERSPEECH | 3 |
| 2022 | Improving Data Driven Inverse Text Normalization using Data Augmentation and Machine Translation
Debjyoti Paul, Yutong Pang, Szu-Jui Chen |
INTERSPEECH | 2 |