VLDB 2026 Research / reviewers in the wild / expert
Yiming Chen 0010
dblp:51/5612-10
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0001-8557-1077ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsabstractLLM serving is limited by provider-side resources: longer generations consume more GPU time, increase latency, and reduce throughput in multi-tenant systems.This creates a denial-of-service (DoS) risk, where attackers degrade service by inducing excessive generation.Prior work on LLM DoS primarily relies on adversarial perturbations that delay end-of-sequence termination.We show perturbations are often unnecessary: natural, benignlooking instructions that specify impractical and meaningless tasks can already trigger excessive generation.To study this overlooked vulnerability, we introduce NaturalSloth, an adversarial dataset of natural, instruction-based DoS prompts.Starting from a human-curated seed set spanning diverse attack categories, we design a multi-agent synthesis framework to scale the dataset while preserving malicious intent and increasing semantic diversity.Experiments across a wide range of proprietary and open-source LLMs show that NaturalSloth consistently induces excessive generation, with attack effectiveness further amplified when combined with jailbreak techniques.Our analysis also reveals significant limitations of existing defenses, highlighting the need for dedicated protections against natural DoS attacks. 1 Yiming Chen 0010, Zexin Li 0001, Xianghu Yue, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 1 |
| 2026 | CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksabstractLarge Language Models (LLMs) are increasingly used not only to generate code, but also to judge it: comparing, ranking, or scoring competing solutions.However, their reliability in this evaluative role remains poorly understood.Inconsistent or flawed judgments can undermine benchmarks and distort training signals.This paper investigates the performance and robustness of LLMs when used as code judges.We introduce CodeJudgeBench, a benchmark explicitly designed to evaluate LLM-as-a-Judge models across three critical coding tasks: code generation, code repair, and unit test generation.We comprehensively benchmark the performance of 26 LLM-as-a-Judge models, encompassing general-purpose, code-tuned, and reasoning models.Our empirical findings reveal that relatively small reasoning models (e.g., Qwen3-8B) can outperform much larger non-reasoning models up to 70B.We further stress-test robustness by applying both general and code-specific perturbations.All models show significant instability and are sensitive to changes such as response ordering, variable naming, and misleading comments.These findings highlight serious concerns about the consistency and robustness of LLM-based judges for coding tasks. Hongchao Jiang, Yiming Chen 0010, Yushi Cao, Hung-yi Lee, Robby T. Tan |
ACL (1) | 2 |
| 2026 | PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001 |
Pattern Recognit. | 2 |
| 2026 | VoiceBench: Benchmarking LLM-Based Voice AssistantsabstractAbstract Recent advancements in large language models (LLMs) like GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering an improved user experience over text-based interactions. However, a suitable benchmark to rigorously evaluate such speech interactions systems is currently lacking. To bridge this gap, we introduce VoiceBench, the first benchmark specifically designed to assess LLM-based voice assistants. VoiceBench comprises 6,783 synthetic and real spoken instructions recorded from diverse speakers across eight distinct tasks. These instructions are meticulously crafted to assess three crucial capability areas: general knowledge, instruction-following, and safety compliance. Furthermore, VoiceBench systematically incorporates realistic variations common in spoken interactions, including differences in speaker characteristics (e.g., accents), heterogeneous environmental conditions (e.g., reverberation), and content complexities such as mispronunciations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.1 Yiming Chen 0010, Xianghu Yue, Chen Zhang 0020, Xiaoxue Gao, Robby T. Tan, Haizhou Li 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationabstractSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang, Zhongwei Wan, Yixin He, Dezhi Ran, Tianle Gu, Haizhou Li, Tao Xie, Baishakhi Ray. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yiming Chen 0010, Zexin Li 0001, Zhongwei Wan, Yixin He 0002, Dezhi Ran, Tianle Gu, Haizhou Li 0001, Tao Xie 0001, Baishakhi Ray |
EMNLP | 2 |
| 2025 | Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference OptimizationabstractCurrent emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs’ in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines. Xiaoxue Gao, Chen Zhang 0020, Yiming Chen 0010, Huayun Zhang, Nancy F. Chen |
ICASSP | 3 |
| 2024 | A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsabstractAutomatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human evaluations. Notably among them, large language models (LLMs), particularly the instruction-tuned variants like ChatGPT, are shown to be promising substitutes for human judges. Yet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc. Hence, it remains inconclusive how effective these LLMs are. To this end, we conduct a comprehensive study on the application of LLMs for automatic dialogue evaluation. Specifically, we analyze the multi-dimensional evaluation capability of 30 recently emerged LLMs at both turn and dialogue levels, using a comprehensive set of 12 meta-evaluation datasets. Additionally, we probe the robustness of the LLMs in handling various adversarial perturbations at both turn and dialogue levels. Finally, we explore how model-level and dimension-level ensembles impact the evaluation performance. All resources are available at https://github.com/e0397123/comp-analysis. Chen Zhang 0020, Luis Fernando D'Haro, Yiming Chen 0010, Malu Zhang, Haizhou Li 0001 |
AAAI | 3 |
| 2024 | MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental TasksabstractClass-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy. Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 3 |
| 2024 | Transferable Adversarial Attacks Against ASRabstractGiven the extensive research and real-world applications of automatic speech recognition (ASR), ensuring the robustness of ASR models against minor input perturbations becomes a crucial consideration for maintaining their effectiveness in real-time scenarios. Previous explorations into ASR model robustness have predominantly revolved around evaluating accuracy on white-box settings with full access to ASR models. Nevertheless, full ASR model details are often not available in real-world applications. Therefore, evaluating the robustness of black-box ASR models is essential for a comprehensive understanding of ASR model resilience. In this regard, we thoroughly study the vulnerability of practical black-box attacks in cutting-edge ASR models and propose to employ two advanced time-domain-based transferable attacks alongside our differentiable feature extractor. We also propose a speech-aware gradient optimization approach (SAGO) for ASR, which forces mistranscription with minimal impact on human imperceptibility through voice activity detection rule and a speech-aware gradient-oriented optimizer. Our comprehensive experimental results reveal performance enhancements compared to baseline approaches across five models on two databases. Xiaoxue Gao, Zexin Li 0001, Yiming Chen 0010, Cong Liu 0005, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | Dynamic Transformers Provide a False Sense of EfficiencyabstractDespite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference.Multi-exit is a mainstream approach to address this issue by making a tradeoff between efficiency and accuracy, where the saving of computation comes from an early exit.However, whether such saving from earlyexiting is robust remains unknown.Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency.To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models.By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction.Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability. 1 Yiming Chen 0010, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 1 |
| 2022 | Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkabstractMost sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals.Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks.In this work, we propose a semi-supervised sentence embedding framework, GenSE, that effectively leverages large-scale unlabeled data.Our method include three parts: 1) Generate: A generator/discriminator model is jointly trained to synthesize sentence pairs from open-domain unlabeled corpus; 2) Discriminate: Noisy sentence pairs are filtered out by the discriminator to acquire high-quality positive and negative sentence pairs; 3) Contrast: A prompt-based contrastive approach is presented for sentence representation learning with both annotated and synthesized data.Comprehensive experiments show that GenSE achieves an average correlation score of 85.19 on the STS datasets and consistent performance improvement on four domain adaptation tasks, significantly surpassing the state-of-the-art methods and convincingly corroborating its effectiveness and generalization ability. 1 Yiming Chen 0010, Yan Zhang 0004, Bin Wang 0040, Zuozhu Liu, Haizhou Li 0001 |
EMNLP | 1 |
| 2022 | Analyzing and Evaluating Faithfulness in Dialogue SummarizationabstractDialogue summarization is abstractive in nature, making it suffer from factual errors.The factual correctness of summaries has the highest priority before practical applications.Many efforts have been made to improve faithfulness in text summarization.However, there is a lack of systematic study on dialogue summarization systems.In this work, we first perform the fine-grained human analysis on the faithfulness of dialogue summaries and observe that over 35% of generated summaries are faithfully inconsistent respective the source dialogues.Furthermore, we present a new model-level faithfulness evaluation method.It examines generation models with multi-choice questions created by rule-based transformations.Experimental results show that our evaluation schema is a strong proxy for the factual correctness of summarization models.The humanannotated faithfulness samples and the evaluation toolkit are released to facilitate future research toward faithful dialogue summarization.Code Bin Wang 0040, Chen Zhang 0020, Yan Zhang 0004, Yiming Chen 0010, Haizhou Li 0001 |
EMNLP | 4 |
| 2021 | DynaEval: Unifying Turn and Dialogue Level EvaluationabstractChen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, Haizhou Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Zhang 0055, Yiming Chen 0010, Luis Fernando D'Haro, Yan Zhang 0004, Thomas Friedrichs, Grandee Lee, Haizhou Li 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Revisiting Self-training for Few-shot Learning of Language ModelabstractAs unlabeled data carry rich task-relevant information, they are proven useful for fewshot learning of language model.The question is how to effectively make use of such data.In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-based fewshot learner, SFLM.Given two views of a text sample via weak and strong augmentation techniques, SFLM generates a pseudo label on the weakly augmented version.Then, the model predicts the same pseudo label when fine-tuned with the strongly augmented version.This simple approach is shown to outperform other state-of-the-art supervised and semi-supervised counterparts on six sentence classification and six sentence-pair classification benchmarking tasks.In addition, SFLM only relies on a few in-domain unlabeled data.We conduct a comprehensive analysis to demonstrate the robustness of our proposed approach under various settings, including augmentation techniques, model scale, and fewshot knowledge transfer across tasks. Yiming Chen 0010, Yan Zhang 0004, Chen Zhang 0020, Grandee Lee, Haizhou Li 0001 |
EMNLP (1) | 1 |