EDBT 2026 Demo / reviewers in the wild / expert
Yuang Li
dblp:283/6035
· DBLP profile ↗
20ranked-venue papers
8as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMsabstractLarge language models (LLMs) have advanced to encompass extensive knowledge across diverse domains. Yet controlling what a LLMs should not know is important for ensuring alignment and thus safe use. However, effective unlearning in LLMs is difficult due to the fuzzy boundary between knowledge retention and forgetting. This challenge is exacerbated by entangled parameter spaces from continuous multi-domain training, often resulting in collateral damage, especially under aggressive unlearning strategies. Furthermore, the computational overhead required to optimize State-of-the-Art (SOTA) models with billions of parameters poses an additional barrier. In this work, we present ALTER, a lightweight unlearning framework for LLMs to address both the challenges of knowledge entanglement and unlearning efficiency. ALTER operates through two phases: (I) high entropy tokens are captured and learned via the shared A matrix in LoRA, followed by (II) an asymmetric LoRA architecture that achieves a specified forgetting objective by parameter isolation and unlearning tokens within the target subdomains. Serving as a new research direction for achieving unlearning via token-level isolation in the asymmetric framework. ALTER achieves SOTA performance on TOFU, WMDP, and MUSE benchmarks with over 95% forget quality and shows minimal side effects through preserving foundational tokens. By decoupling unlearning from LLMs' billion-scale parameters, this framework delivers excellent efficiency while preserving over 90% of model utility, exceeding baseline preservation rates of 47.8-83.6%. Xunlei Chen, Jinyu Guo, Yuang Li, Zhaokun Wang, Jie Zou 0001, Jiwei Wei, Wenhong Tian |
AAAI | 3 |
| 2026 | AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache ReuseabstractJie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Ruiqi Wu, Zhaokun Wang, Wenyi Li, Wenhong Tian. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Zhaokun Wang, Wenhong Tian |
ACL (1) | 4 |
| 2026 | Why not transform chat large language models to non-English?
Xiang Geng, Ming Zhu 0010, Jiahuan Li, Zhejian Lai, Shuaijie She, Yinglu Li, Yuang Li, Chang Su 0001, Xinglin Lyu, Min Zhang 0042, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang |
Frontiers Comput. Sci. | 10 |
| 2026 | Seeking commonality while preserving diversity: A differential dual-path MoE solution for Multilingual Neural Machine Translation
Jinyu Guo, Yuang Li, Wenxian Liu, Zhaokun Wang, Wenhong Tian |
Knowl. Based Syst. | 3 |
| 2025 | Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio EncodersabstractWeiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo, Chen Xu, Xiaofeng Zhao, Long Meng, Yunfei Lu, Min Zhang, Hao Yang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weiqiao Shan, Yuang Li, Yingfeng Luo, Chen Xu 0008, Long Meng, Yunfei Lu, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001 |
EMNLP | 2 |
| 2025 | Investigating Numerical Translation with Large Language ModelsabstractThe inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explored. This study focuses on evaluating the reliability of LLM-based machine translation systems when handling numerical data. In order to systematically test the numerical translation capabilities of currently open source LLMs, we have constructed a numerical translation dataset between Chinese and English based on real business data, encompassing ten types of numerical translation. Experiments on the dataset indicate that errors in numerical translation are a common issue, with most open-source LLMs faltering when faced with our test scenarios. Especially when it comes to numerical types involving large units like "million", "billion", and "亿" , even the latest llama3.1 8b model can have error rates as high as 20%. Finally, we introduce three potential strategies to mitigate the numerical mistranslations for large units. Wei Tang 0013, Yuang Li, Min Zhang 0042, Hao Yang 0006 |
ICASSP | 3 |
| 2025 | Large Language Model Should Understand Pinyin for Chinese ASR Error CorrectionabstractLarge language models (LLMs) can enhance automatic speech recognition (ASR) systems through generative error correction (GEC). In this paper, we propose Pinyin-enhanced GEC (PY-GEC), which leverages Pinyin—the phonetic representation of Mandarin Chinese—as supplementary information to improve Chinese ASR error correction. Our approach only utilizes synthetic errors for training and employs the one-best hypothesis during inference. Additionally, we introduce a multitask training approach involving conversion tasks between Pinyin and text to align their feature spaces. Experiments on the Aishell-1 and the Common Voice datasets demonstrate that our approach consistently outperforms GEC with text-only input. More importantly, we provide intuitive explanations for the effectiveness of PY-GEC and multitask training from two aspects: 1) increased attention weight on Pinyin features; and 2) aligned feature space between Pinyin and text hidden states. Yuang Li, Xiaosong Qiao, Wei Tang 0013, Min Zhang 0042, Hao Yang 0006 |
ICASSP | 1 |
| 2025 | Optimizing Speech Multi-View Feature Fusion through Conditional ComputationabstractRecent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset. Weiqiao Shan, Yuchen Han 0001, Yuang Li, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001 |
ICASSP | 6 |
| 2025 | "I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen EntitiesabstractSpoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we demonstrate that existing Spoken NER systems perform poorly when dealing with previously unseen named entities. To tackle this challenge, we propose a method for generating Spoken NER data based on a named entity dictionary (NED) to reduce costs. Specifically, we first use a large language model (LLM) to generate sentences from the sampled named entities and then use a text-to-speech (TTS) system to generate the speech. Furthermore, we introduce a noise metric to filter out noisy data. To evaluate our approach, we release a novel Spoken NER benchmark along with a corresponding NED containing 8,853 entities. Experiment results show that our method achieves state-of-the-art (SOTA) performance in the in-domain, zero-shot domain adaptation, and fully zero-shot settings. Our data will be available at https://github.com/DeepLearnXMU/HeardU. Xiang Geng, Yuang Li, Mengxin Ren, Wei Tang 0013, Jiahuan Li, Zhibin Lan, Min Zhang 0042, Hao Yang 0006, Shujian Huang, Jinsong Su |
ICASSP | 3 |
| 2025 | Graph Alignment Using Seed-Oriented Subgraph MatchingabstractThis paper addresses the challenge of unsupervised plain graph alignment, specifically in scenarios where auxiliary information, such as node attributes, is unavailable. Existing alignment algorithms primarily fall into two categories: spectral methods and representation learning-based methods. Spectral methods typically leverage alignment consistency principles, employing heuristic strategies to iteratively infer the alignment matrix. In contrast, representation learning methods focus on encoding the geometric structural features of nodes to generate node representations, thereby transforming the node matching task into a similarity computation based on these representations. While both approaches demonstrate robust performance in the graph alignment domain, their time complexity poses significant concerns. To mitigate this issue, we propose a novel, efficient algorithm grounded in seed-oriented subgraph matching. Our method begins by extracting a limited number of reliable pseudo alignment seeds derived from graph geometric features. Subsequently, we extract the corresponding K-hop seed-oriented subgraphs, allowing us to reformulate the graph alignment problem into a series of subgraph matching tasks. The final alignment matrix is then constructed by aggregating the results of these subgraph matches. Experimental evaluations conducted on public datasets reveal that our method not only improves efficiency but also outperforms current state-of-the-art techniques in terms of accuracy. Wei Tang 0013, Xinglin Lv, Yuang Li, Min Zhang 0042, Hao Yang 0006 |
ICMR | 3 |
| 2025 | Efficient adaptive defense scheme for differential privacy in federated learning
Fangfang Shan, Yanlong Lu, Shuaifeng Li, Shiqi Mao, Yuang Li |
J. Inf. Secur. Appl. | 5 |
| 2024 | CB-Whisper: Contextual Biasing Whisper Using Open-Vocabulary Keyword-SpottingabstractEnd-to-end automatic speech recognition (ASR) systems often struggle to recognize rare name entities, such as personal names, organizations and terminologies that are not frequently encountered in the training data. This paper presents Contextual Biasing Whisper (CB-Whisper), a novel ASR system based on OpenAI’s Whisper model that can recognize user-defined name entities by performing open-vocabulary keyword-spotting (KWS) before the decoder. The KWS module leverages text-to-speech (TTS) techniques and a convolutional neural network (CNN) classifier to match the features between the entities and the utterances. To integrate the recognized entities into the Whipser decoder and avoid hallucinations, we carefully crafted multiple prompts with spoken form hints. Experiments show that the KWS module based on Whisper encoder’s features can recognize unseen user-defined keywords effectively. More importantly, the proposed CB-Whisper substantially improves the mixed-error-rate (MER) and entity recall compared to the original Whisper model on three internal datasets and two publicly available datasets including Aishell and ACL datasets that cover English-only, Chinese-only, and code-switching scenarios. Yuang Li, Yinglu Li, Min Zhang 0042, Chang Su 0001, Mengyao Piao, Xiaosong Qiao, Miaomiao Ma, Hao Yang 0006 |
LREC/COLING | 1 |
| 2024 | Cross-Domain Audio Deepfake Detection: Dataset and AnalysisabstractAudio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy.Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance.However, the existing ADD datasets are outdated, leading to suboptimal generalization of detection models.In this paper, we construct a new cross-domain ADD dataset comprising over 300 hours of speech data that is generated by five advanced zeroshot TTS models.To simulate real-world scenarios, we employ diverse attack methods and audio prompts from different datasets.Experiments show that, through novel attackaugmented training, the Wav2Vec2-large and Whisper-medium models achieve equal error rates of 4.1% and 6.5% respectively.Additionally, we demonstrate our models' outstanding few-shot ADD ability by fine-tuning with just one minute of target-domain data.Nonetheless, neural codec compressors greatly affect the detection accuracy, necessitating further research.Our dataset is publicly available 1 . Yuang Li, Min Zhang 0042, Mengxin Ren, Xiaosong Qiao, Miaomiao Ma, Daimeng Wei, Hao Yang 0006 |
EMNLP | 1 |
| 2024 | Using Large Language Model for End-to-End Chinese ASR and NER
Yuang Li, Min Zhang 0042, Mengxin Ren, Shimin Tao, Jinsong Su, Hao Yang 0006 |
INTERSPEECH | 1 |
| 2024 | A Multitask Training Approach to Enhance Whisper with Open-Vocabulary Keyword SpottingabstractThe recognition of rare named entities, such as personal names and terminologies, is challenging for automatic speech recognition (ASR) systems, especially when they are not frequently observed in the training data.In this paper, we introduce keyword spotting enhanced Whisper (KWS-Whisper), a novel ASR system that leverages the Whisper model and performs openvocabulary keyword spotting (OV-KWS) on the hidden states of the Whisper encoder to recognize user-defined named entities.These entities serve as prompts for the Whisper decoder.To optimize the model, we propose a multitask training approach that learns OV-KWS and contextual-ASR tasks.We evaluate our approach on Chinese Aishell hot word subsets and two internal code-switching test sets and show that it significantly improves the entity recall compared to the original Whisper model.Moreover, we demonstrate that the OV-KWS can be a plug-andplay module to enhance the ASR error correction methods and frozen Whisper models. Yuang Li, Min Zhang 0042, Chang Su 0001, Yinglu Li, Xiaosong Qiao, Mengxin Ren, Miaomiao Ma, Daimeng Wei, Shimin Tao, Hao Yang 0006 |
INTERSPEECH | 1 |
| 2023 | Prompting Large Language Models for Zero-Shot Domain Adaptation in Speech RecognitionabstractThe integration of Language Models (LMs) has proven to be an effective way to address domain shifts in speech recognition. However, these approaches usually require a significant amount of target domain text data for the training of LMs. Different from these methods, in this work, with only a domain-specific text prompt, we propose two zero-shot ASR domain adaptation methods using LLaMA, a 7-billionparameter large language model (LLM). LLM is used in two ways: 1) second-pass rescoring: reranking N-best hypotheses of a given ASR system with LLaMA; 2) deep LLM-fusion: incorporating LLM into the decoder of an encoder-decoder based ASR system. Experiments show that, with only one domain prompt, both methods can effectively reduce word error rates (WER) on out-of-domain TedLium-2 and SPGISpeech datasets. Especially, the deep LLM-fusion has the advantage of better recall of entity and out-of-vocabulary words. Yuang Li, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001 |
ASRU | 1 |
| 2023 | Knowledge Prompt for Whisper: An ASR Entity Correction Approach with Knowledge BaseabstractEntity correction is crucial in Automatic Speech TABLE I Recognition (ASR), since erroneous entities seriously affect our understanding of ASR results. In this paper, in order to correct entity errors, we propose a knowledge prompt approach for Whisper (a recent ASR model trained with a corpus containing 680k hours of labeled speech recorded in various conditions). For a given audio, our approach consists of three steps: (1) obtaining its ASR result by Whisper; (2) fuzzy matching the ASR result with a knowledge base to obtain candidate entities; (3) using the candidate entities as a prompt to obtain the final ASR result by Whisper again. We conduct experiments on the test dataset of open-source Chinese speech corpus AISHELLNER. Experimental results show that our approach not only significantly improves the entity recall rate in ASR results (from 70.97% to 84.82%), but also reduces the overall Character Error Rate (CER). Min Zhang 0042, Xiaosong Qiao, Chang Su 0001, Yinglu Li, Yuang Li, Ming Zhu 0010, Mengyao Piao, Shimin Tao, Hao Yang 0006, Yanfei Jiang |
IEEE Big Data | 6 |
| 2023 | Self-Supervised Learning-Based Source Separation for Meeting DataabstractSource separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel source separation, most studies have focused on simulated setups. In this paper, seven SSL models were compared on both simulated and real-world corpora. Then, we propose to integrate the best-performing model WavLM into an automatic transcription system through a novel iterative source selection method. To improve real-world performance, time-domain unsupervised mixture invariant training was adapted to the time-frequency domain. Experiments showed that in the transcription system when source separation was inserted before an ASR model fine-tuned on separated speech, absolute reductions of 1.9% and 1.5% in concatenated minimum-permutation word error rate for an unknown number of speakers (cpWER-us) were observed on the AMI dev and test sets. Yuang Li, Xianrui Zheng, Philip C. Woodland |
ICASSP | 1 |
| 2023 | Accurate and Structured Pruning for Efficient Automatic Speech Recognition
Huiqiang Jiang, Li Lyna Zhang, Yuang Li, Shijie Cao, Yuqing Yang 0001, Mao Yang 0004, Lili Qiu |
INTERSPEECH | 3 |
| 2023 | Accelerating Transducers through Adjacent Token Merging
Yuang Li, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001 |
INTERSPEECH | 1 |