VLDB 2026 Research / reviewers in the wild / expert
Yingyi Ma
dblp:242/3346
· DBLP profile ↗
13ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Dataset Distillation via Generative PruningabstractDataset distillation (DD) has demonstrated the promise of synthesizing smaller datasets that enable competitive performance. While most of DD methods operate in pixel space, they often suffer from poor scalability on high-resolution datasets. Recent works have thus shifted towards parameterizing synthetic data using deep generative priors. However, these approaches apply latent updates to all generator layers, leading to substantial computational overhead. To address this limitation, we propose Generative Lightweight Distillation (GLiD), a unified framework that jointly compresses the generator and accelerates latent optimization for efficient distillation. GLiD introduces two key components: (1) We introduce a Classification-Diversity Sensitivity Pruning mechanism that quantifies both discriminative utility and semantic diversity of each output channel to guide structural pruning and selective layer-wise optimization. (2) We also present a Layer-Adaptive Scheduling strategy that dynamically allocates latent update steps across generator stages based on convergence behavior. Extensive experiments on CIFAR-10 and ImageNet-1K and its subsets demonstrate that our GLiD achieves up to 10× acceleration across different datasets, while maintaining performance competitive with state-of-the-art methods. Yingyi Ma, Muquan Li, Guiduo Duan, Ke Qin, Shuang Liang 0002, Dongyang Zhang 0001 |
IEEE Trans. Big Data | 1 |
| 2025 | Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech RecognitionabstractWhile large language models (LLMs) have been applied to automatic speech recognition (ASR), the task of making the model streamable remains a challenge. This paper proposes a novel model architecture, Transducer-Llama, that integrates LLMs into a Factorized Transducer (FT) model, naturally enabling streaming capabilities. Furthermore, given that the large vocabulary of LLMs can cause data sparsity issue and increased training costs for spoken language systems, this paper introduces an efficient vocabulary adaptation technique to align LLMs with speech system vocabularies. The results show that directly optimizing the FT model with a strong pre-trained LLM-based predictor using the RNN-T loss yields some but limited improvements over a smaller pre-trained LM predictor. Therefore, this paper proposes a weak-to-strong LM swap strategy, using a weak LM predictor during RNN-T loss training and then replacing it with a strong LLM. After LM replacement, the minimum word error rate (MWER) loss is employed to finetune the integration of the LLM predictor with the Transducer-Llama model. Experiments on the LibriSpeech and large-scale multi-lingual LibriSpeech corpora show that the proposed streaming Transducer-Llama approach gave a 17% relative WER reduction (WERR) over a strong FT baseline and a 32% WERR over an RNN-T baseline. Keqi Deng, Jinxi Guo, Yingyi Ma, Niko Moritz, Philip C. Woodland, Ozlem Kalinli, Mike Seltzer |
ICASSP | 3 |
| 2024 | DOC-RAG: ASR Language Model Personalization with Domain-Distributed Co-occurrence Retrieval AugmentationabstractWe propose DOC-RAG - Domain-distributed Co-occurrence Retrieval Augmentation for ASR language model personalization aiming to improve the automatic speech recognition of rare word patterns in unseen domains. Our approach involves contrastively training a document retrieval module to rank external knowledge domains based on their semantic similarity with respect to the input query. We further use n-gram co-occurrence distribution to recognize rare word patterns associated with specific domains. We aggregate the next word probability distribution based on the relative importance of different domains. Extensive experiments on three user-specific speech-to-text tasks for meetings, TED talks, and financial earnings calls show that DOC-RAG significantly outperforms strong baselines with an 8-15% improvement in terms of perplexity and a 4-7% reduction in terms of Word Error Rates in various settings. Puneet Mathur, Zhe Liu 0011, Ke Li 0018, Yingyi Ma, Gil Keren, Dinesh Manocha |
LREC/COLING | 4 |
| 2024 | Effective Internal Language Model Training and Fusion for Factorized Transducer ModelabstractThe internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracted during inference to facilitate improved integration with external language models. Recently, various of factorized transducer models have been proposed, which explicitly embrace a standalone internal language model for non-blank token prediction. However, even with the adoption of factorized transducer models, limited improvement has been observed compared to shallow fusion. In this paper, we propose a novel ILM training and decoding strategy for factorized transducer models, which effectively combines the blank, acoustic and ILM scores. Our experiments show a 17% relative improvement over the standard decoding method when utilizing a well-trained ILM and the proposed decoding strategy on LibriSpeech datasets. Furthermore, when compared to a strong RNN-T baseline enhanced with external LM fusion, the proposed model yields a 5.5% relative improvement on general-sets and an 8.9% WER reduction for rare words. The proposed model can achieve superior performance without relying on external language models, rendering it highly efficient for production use-cases. To further improve the performance, we propose a novel and memory-efficient ILM-fusion-aware minimum word error rate (MWER) training method which improves ILM integration significantly. Jinxi Guo, Niko Moritz, Yingyi Ma, Frank Seide, Chunyang Wu, Jay Mahadeokar, Ozlem Kalinli, Christian Fügen, Mike Seltzer |
ICASSP | 3 |
| 2024 | Correction Focused Language Model Training For Speech RecognitionabstractLanguage models (LMs) have been commonly adopted to boost the performance of automatic speech recognition (ASR) particularly in domain adaptation tasks. Conventional way of LM training treats all the words in corpora equally, resulting in suboptimal improvements in ASR performance. In this work, we introduce a novel correction focused LM training approach which aims to prioritize ASR fallible words. The word-level ASR fallibility score, representing the likelihood of ASR mis-recognition, is defined and shaped as a prior word distribution to guide the LM training. To enable correction focused training with text-only corpora, large language models (LLMs) are employed as fallibility score predictors and text generators through multi-task fine-tuning. Experimental results for domain adaptation tasks demonstrate the effectiveness of our proposed method. Compared with conventional LMs, correction focused training achieves up to relatively 5.5% word error rate (WER) reduction in sufficient text scenarios. In insufficient text scenarios, LM training with LLMgenerated text achieves up to relatively 13% WER reduction, while correction focused training further obtains up to relatively 6% WER reduction. Yingyi Ma, Zhe Liu 0011, Ozlem Kalinli |
ICASSP | 1 |
| 2024 | Contextual Biasing of Named-Entities with Large Language ModelsabstractWe explore contextual biasing with Large Language Models (LLMs) to enhance Automatic Speech Recognition (ASR) in second-pass rescoring. Our approach introduces the utilization of prompts for LLMs during rescoring without the need for fine-tuning. These prompts incorporate a biasing list and a set of few-shot examples, serving as supplementary sources of information when evaluating the hypothesis score. Furthermore, we introduce multi-task training for LLMs to predict entity class and the subsequent token. To address sequence length constraints and improve the efficiency of contextual biasing, we propose dynamic prompting based on class tag predictions. Through dynamic prompting, we leverage the class tag predictions to identify the most probable entity class and subsequently utilize entities within this class as biasing context for the next token prediction. We evaluate the performance of proposed methods in terms of Word Error Rate (WER) on an internal entity-heavy and the SLUE-Voxpopuli datasets. Our results show significant improvements: biasing lists and few-shot examples achieved a relative improvement of 17.8% and 9.6%, while multitask training and dynamic prompting achieved 20.0% and 11.3% relative WER improvement, respectively. Chuanneng Sun, Yingyi Ma, Zhe Liu 0011, Lucas Kabela, Yutong Pang, Ozlem Kalinli |
ICASSP | 3 |
| 2024 | Recovering from Privacy-Preserving Masking with Large Language ModelsabstractModel adaptation is crucial to handle the discrepancy between proxy training data and actual users’ data received. To effectively perform adaptation, textual data of users is typically stored on servers or their local devices, where downstream natural language processing (NLP) models can be directly trained using such in-domain data. However, this might raise privacy and security concerns due to the extra risks of exposing user information to adversaries. Replacing identifying information in textual data with a generic marker has been recently explored. In this work, we leverage large language models (LLMs) to suggest substitutes of masked tokens and have their effectiveness evaluated on downstream language modeling tasks. Specifically, we propose multiple pre-trained and fine-tuned LLM-based approaches and perform empirical studies on various datasets for the comparison of these methods. Experimental results show that models trained on the obfuscation corpora are able to achieve comparable performance with the ones trained on the original data without privacy-preserving token masking. Arpita Vats, Zhe Liu 0011, Debjyoti Paul, Yingyi Ma, Yutong Pang, Ozlem Kalinli |
ICASSP | 5 |
| 2024 | Effective Text Adaptation For LLM-Based ASR Through Soft Prompt Fine-TuningabstractThe advent of Large Language Models (LLM) has reformed the Automatic Speech Recognition (ASR). Prompting LLM with audio embeddings to generate transcriptions becomes the new state-of-the-art ASR. Despite LLMs being trained with an extensive amount of text corpora, high-quality domain-specific text data can still significantly enhance ASR performance on domain adaptation tasks. Although LLM-based ASR can naturally incorporate more text corpora by fine-tuning the LLM decoder, fine-tuning such ASR on text-only data without paired prompts may diminish the effectiveness of domain-specific knowledge. To mitigate this issue, we propose a two-step soft prompt fine-tuning strategy that enhances domain-specific text adaptation. Experimental results show that text adaptation with our proposed method achieved a relative up to 9% Word Error Rate (WER) reduction and up to 18% Entity Error Rate (EER) reduction on the target domain compared to the baseline ASR. Combining this with domain-specific Language Model (LM) fusion can further improve the EER by a relative 2-5%. Yingyi Ma, Zhe Liu 0011, Ozlem Kalinli |
SLT | 1 |
| 2023 | Adaptive Multi-Corpora Language Model Training for Speech RecognitionabstractNeural network language model (NNLM) plays an essential role in automatic speech recognition (ASR) systems, especially in adaptation tasks when text-only data is available. In practice, an NNLM is typically trained on a combination of data sampled from multiple corpora. Thus, the data sampling strategy is important to the adaptation performance. Most existing works focus on designing static sampling strategies. However, each corpus may show varying impacts at different NNLM training stages. In this paper, we introduce a novel adaptive multi-corpora training algorithm that dynamically learns and adjusts the sampling probability of each corpus along the training process. The algorithm is robust to corpora sizes and domain relevance. Compared with static sampling strategy baselines, the proposed approach yields remarkable improvement by achieving up to relative 7% and 9% word error rate (WER) reductions on in-domain and out-of-domain adaptation tasks, respectively. Yingyi Ma, Zhe Liu 0011 |
ICASSP | 1 |
| 2022 | Warping Layer: Representation Learning for Label Structures in Weakly Supervised LearningabstractMany learning tasks only receive weak supervision, such as semi-supervised learning and few-shot learning. With limited labeled data, prior structures become especially important, and prominent examples include hierarchies and mutual exclusions in the class space. However, most existing approaches only learn the representations separately in the feature space and the label space, and do not explicitly enforce the logical relationships. In this paper, we propose a novel warping layer that jointly learns representations in both spaces, and thanks to the modularity and differentiability, it can be directly embedded into generative models to leverage the prior hierarchical structure and unlabeled data. The effectiveness of the warping layer is demonstrated on both few-shot and semi-supervised learning, outperforming the state of the art in practice. Yingyi Ma |
AISTATS | 1 |
| 2020 | Convex Representation Learning for Generalized Invariance in Semi-Inner-Product SpaceabstractInvariance (defined in a general sense) has been one of the most effective priors for representation learning. Direct factorization of parametric models is feasible only for a small range of invariances, while regularization approaches, despite improved generality, lead to nonconvex optimization. In this work, we develop a \emph{convex} representation learning algorithm for a variety of generalized invariances that can be modeled as semi-norms. Novel Euclidean embeddings are introduced for kernel representers in a semi-inner-product space, and approximation bounds are established. This allows invariant representations to be learned efficiently and effectively as confirmed in our experiments, along with accurate predictions. Yingyi Ma, Vignesh Ganapathiraman, Yaoliang Yu |
ICML | 1 |
| 2020 | Proximal Mapping for Deep RegularizationabstractUnderpinning the success of deep learning is effective regularizations that allow a variety of priors in data to be modeled. For example, robustness to adversarial perturbations, and correlations between multiple modalities. However, most regularizers are specified in terms of hidden layer outputs, which are not themselves optimization variables. In contrast to prevalent methods that optimize them indirectly through model weights, we propose inserting proximal mapping as a new layer to the deep network, which directly and explicitly produces well regularized hidden layer outputs. The resulting technique is shown well connected to kernel warping and dropout, and novel algorithms were developed for robust temporal learning and multiview modeling, both outperforming state-of-the-art methods. Yingyi Ma |
NeurIPS | 2 |
| 2019 | Learning Invariant Representations with Kernel WarpingabstractInvariance is an effective prior that has been extensively used to bias supervised learning with a \emph{given} representation of data. In order to learn invariant representations, wavelet and scattering based methods “hard code” invariance over the \emph{entire} sample space, hence restricted to a limited range of transformations. Kernels based on Haar integration also work only on a \emph{group} of transformations. In this work, we break this limitation by designing a new representation learning algorithm that incorporates invariances \emph{beyond transformation}. Our approach, which is based on warping the kernel in a data-dependent fashion, is computationally efficient using random features, and leads to a deep kernel through multiple layers. We apply it to convolutional kernel networks and demonstrate its stability. Yingyi Ma, Vignesh Ganapathiraman |
AISTATS | 1 |