VLDB 2026 Research / reviewers in the wild / expert
Yue Xing 0002
dblp:185/5744-2
· DBLP profile ↗
24ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0001-7723-0048ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 8 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Retrieval Heads are DynamicabstractYuping Lin, Zitao Li, Yue Xing, Pengfei He, Yingqian Cui, Yaliang Li, Bolin Ding, Jingren Zhou, Jiliang Tang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuping Lin, Zitao Li, Yue Xing 0002, Yingqian Cui, Yaliang Li, Bolin Ding, Jingren Zhou 0001, Jiliang Tang |
ACL (1) | 3 |
| 2025 | Unveiling Privacy Risks in LLM Agent MemoryabstractLarge Language Model (LLM) agents have become increasingly prevalent across various realworld applications.They enhance decisionmaking by storing private user-agent interactions in the memory module for demonstrations, introducing new privacy risks for LLM agents.In this work, we systematically investigate the vulnerability of LLM agents to our proposed Memory EXTRaction Attack (MEXTRA) under a black-box setting.To extract private information from memory, we propose an effective attacking prompt design and an automated prompt generation method based on different levels of knowledge about the LLM agent.Experiments on two representative agents demonstrate the effectiveness of MEXTRA.Moreover, we explore key factors influencing memory leakage from both the agent designer's and the attacker's perspectives.Our findings highlight the urgent need for effective memory safeguards in LLM agent design and deployment. Bo Wang 0069, Weiyi He, Shenglai Zeng, Zhen Xiang, Yue Xing 0002, Jiliang Tang |
ACL (1) | 5 |
| 2025 | Towards Context-Robust LLMs: A Gated Representation Fine-tuning ApproachabstractLarge Language Models (LLMs) enhanced with external contexts, such as through retrieval-augmented generation (RAG), often face challenges in handling imperfect evidence.They tend to over-rely on external knowledge, making them vulnerable to misleading and unhelpful contexts.To address this, we propose the concept of context-robust LLMs, which can effectively balance internal knowledge with external context, similar to human cognitive processes.Specifically, context-robust LLMs should rely on external context only when lacking internal knowledge, identify contradictions between internal and external knowledge, and disregard unhelpful contexts.To achieve this goal, we introduce Grft, a lightweight and plug-and-play gated representation fine-tuning approach.Grft consists of two key components: a gating mechanism to detect and filter problematic inputs, and low-rank representation adapters to adjust hidden representations.By training a lightweight intervention function with only 0.0004% of model size on fewer than 200 examples, Grft can effectively adapts LLMs towards context-robust behaviors. Shenglai Zeng, Kai Guo 0003, Hanqing Lu, Yue Xing 0002, Hui Liu 0031 |
ACL (1) | 6 |
| 2025 | A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware DemonstrationabstractFew-shot Chain-of-Thought (CoT) prompting has demonstrated strong performance in improving the reasoning capabilities of large language models (LLMs). While theoretical investigations have been conducted to understand CoT, the underlying transformer used in these studies isolates the CoT reasoning process into separated in-context learning steps (Stepwise ICL). In this work, we theoretically show that, compared to Stepwise ICL, the transformer gains better error correction ability and more accurate predictions if the reasoning from earlier steps (Coherent CoT) is integrated. Given that this coherent reasoning changes the behavior of the transformer, we further investigate the sensitivity of the transformer with Coherent CoT when the demonstration examples are corrupted at the inference stage. Our theoretical results indicate that the transformer is more sensitive to errors in intermediate reasoning steps than the final outcome. Building upon this observation, we propose an improvement on CoT by incorporating both correct and incorrect reasoning paths in the demonstration. Our experiments validate the effectiveness of the proposed approach. Yingqian Cui, Xianfeng Tang, Qi He 0002, Chen Luo 0003, Jiliang Tang, Yue Xing 0002 |
AISTATS | 7 |
| 2025 | Superiority of Multi-Head Attention: A Theoretical Study in Shallow Transformers in In-Context Linear RegressionabstractWe present a theoretical analysis of the performance of transformer with softmax attention in in-context learning with linear regression tasks. While the existing theoretical literature predominantly focuses on providing convergence upper bounds to show that trained transformers with single-/multi-head attention can obtain a good in-context learning performance, our research centers on comparing the exact convergence of single- and multi-head attention more rigorously. We conduct an exact theoretical analysis to demonstrate that multi-head attention with a substantial embedding dimension performs better than single-head attention. When the number of in-context examples $D$ increases, the prediction loss using single-/multi-head attention is in $O(1/D)$, and the one for multi-head attention has a smaller multiplicative constant. In addition to the simplest data distribution setting, our technical framework in calculating the exact convergence further facilitates studying more scenarios, e.g., noisy labels, local examples, correlated features, and prior knowledge. We observe that, in general, multi-head attention is preferred over single-head attention. Our results verify the effectiveness of the design of multi-head attention in the transformer architecture. Yingqian Cui, Jie Ren 0019, Hui Liu 0031, Jiliang Tang, Yue Xing 0002 |
AISTATS | 6 |
| 2025 | Six-CD: Benchmarking Concept Removals for Text-to-image Diffusion ModelsabstractText-to-image (T2I) diffusion models have shown exceptional capabilities in generating images that closely correspond to textual prompts. However, the advancement of T2I diffusion models presents significant risks, as the models could be exploited for malicious purposes, such as generating images with violence or nudity, or creating unauthorized portraits of public figures in inappropriate contexts. To mitigate these risks, concept removal methods have been proposed. These methods aim to modify diffusion models to prevent the generation of malicious and unwanted concepts. Despite these efforts, existing research faces several challenges: (1) a lack of consistent comparisons on a comprehensive dataset, (2) ineffective prompts in harmful and nudity concepts, (3) overlooked evaluation of the ability to generate the benign part within prompts containing malicious concepts. To address these gaps, we propose to benchmark the concept removal methods by introducing a new dataset, Six-CD, along with a novel evaluation metric. In this benchmark, we conduct a thorough evaluation of concept removals, with the experimental observations and discussions offering valuable insights in the field. Jie Ren 0019, Kangrui Chen, Yingqian Cui, Shenglai Zeng, Hui Liu 0003, Yue Xing 0002, Jiliang Tang, Lingjuan Lyu |
CVPR | 6 |
| 2025 | Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic DataabstractShenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, Jiliang Tang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shenglai Zeng, Jiankun Zhang 0001, Jie Ren 0019, Hanqing Lu, Han Xu 0002, Hui Liu 0003, Yue Xing 0002, Jiliang Tang |
EMNLP | 9 |
| 2025 | Towards Knowledge Checking in Retrieval-augmented Generation: A Representation PerspectiveabstractShenglai Zeng, Jiankun Zhang, Bingheng Li, Yuping Lin, Tianqi Zheng, Dante Everaert, Hanqing Lu, Hui Liu, Hui Liu, Yue Xing, Monica Xiao Cheng, Jiliang Tang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shenglai Zeng, Jiankun Zhang 0001, Bingheng Li, Yuping Lin, Dante Everaert, Hanqing Lu, Hui Liu 0033, Hui Liu 0031, Yue Xing 0002, Monica Xiao Cheng, Jiliang Tang |
NAACL (Long Papers) | 10 |
| 2025 | LLM Safety Alignment is Divergence Estimation in DisguiseabstractWe present a theoretical framework showing that popular LLM alignment methods—including RLHF and its variants—can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less-preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance–refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety. Rajdeep Haldar, Guang Lin 0001, Yue Xing 0002, Qifan Song |
NeurIPS | 4 |
| 2025 | Keeping an Eye on LLM Unlearning: The Hidden Risk and RemedyabstractAlthough Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks, growing concerns have emerged over the misuse of sensitive, copyrighted, or harmful data during training. To address these concerns, unlearning techniques have been developed to remove the influence of specific data without retraining from scratch. However, this paper reveals a critical vulnerability in fine-tuning-based unlearning: a malicious user can craft a manipulated forgetting request that stealthily degrades the model’s utility for benign users. We demonstrate this risk through a red-teaming Stealthy Attack (SA), which is inspired by two key limitations of existing unlearning—the inability to constrain the scope of unlearning effect and the failure to distinguish benign tokens from unlearning signals. Prior work has shown that unlearned models tend to memorize forgetting data as unlearning signals, and respond with hallucinations or feigned ignorance when unlearning signals appear in the input. By subtly increasing the presence of common benign tokens in the forgetting data, SA enhances the connection between benign tokens and unlearning signals. As a result, when normal users include such tokens in their prompts, the model exhibits unlearning behaviors, leading to unintended utility degradation. To address this vulnerability, we propose Scope-aware Unlearning (SU), a lightweight enhancement that introduces a scope term into the unlearning objective, encouraging the model to localize the forgetting effect. Our method requires no additional data processing, integrates seamlessly with existing fine-tuning frameworks, and significantly improves robustness against SA. Extensive experiments validate the effectiveness of both SA and SU. Jie Ren 0019, Zhenwei Dai, Xianfeng Tang, Yue Xing 0002, Shenglai Zeng, Jingying Zeng, Qiankun Peng, Samarth Varshney, Suhang Wang, Qi He 0002, Charu C. Aggarwal, Hui Liu 0003 |
NeurIPS | 4 |
| 2025 | Self-Comparison for Dataset-Level Membership Inference in Large (Vision-)Language ModelabstractLarge Language Models (LLMs) and Vision-Language Models (VLMs) have made significant advancements in a wide range of natural language processing and vision-language tasks. Access to large web-scale datasets has been a key factor in their success. However, concerns have been raised about the unauthorized use of copyrighted materials and potential copyright infringement. Existing methods, such as sample-level Membership Inference Attacks (MIA) and distribution-based dataset, inference distinguish member and non-member data by leveraging the common observation that models tend to memorize and show greater confidence in member data. Nevertheless, these methods face challenges when applied to LLMs and VLMs, such as the requirement for ground-truth member data or non-member data that shares the same distribution as the test data. In this paper, we propose a novel dataset-level membership inference method based on Self-Comparison. We find that a member prefix followed by a non-member suffix (paraphrased from a member suffix) can further trigger the model's memorization on training data. Instead of directly comparing member and non-member data, we introduce paraphrasing to the second half of the sequence and evaluate how the likelihood changes before and after paraphrasing. Unlike prior approaches, our method does not require access to ground-truth member data or non-member data in identical distribution, making it more practical. Extensive experiments demonstrate that our proposed method outperforms traditional MIA and dataset inference techniques across various datasets and models, including GPT-4o. Jie Ren 0019, Kangrui Chen, Chen Chen 0043, Vikash Sehwag, Yue Xing 0002, Jiliang Tang, Lingjuan Lyu |
WWW | 5 |
| 2024 | Exploring Memorization in Fine-tuned Language ModelsabstractShenglai Zeng, Yaxin Li, Jie Ren, Yiding Liu, Han Xu, Pengfei He, Yue Xing, Shuaiqiang Wang, Jiliang Tang, Dawei Yin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shenglai Zeng, Yaxin Li 0001, Jie Ren 0019, Han Xu 0002, Yue Xing 0002, Shuaiqiang Wang, Jiliang Tang, Dawei Yin 0001 |
ACL (1) | 7 |
| 2024 | Better Representations via Adversarial Training in Pre-Training: A Theoretical PerspectiveabstractPre-training is known to generate universal representations for downstream tasks in large-scale deep learning such as large language models. Existing literature, e.g., Kim et al. (2020), empirically observe that the downstream tasks can inherit the adversarial robustness of the pre-trained model. We provide theoretical justifications for this robustness inheritance phenomenon. Our theoretical results reveal that feature purification plays an important role in connecting the adversarial robustness of the pre-trained model and the downstream tasks in two-layer neural networks. Specifically, we show that (i) with adversarial training, each hidden node tends to pick only one (or a few) feature; (ii) without adversarial training, the hidden nodes can be vulnerable to attacks. This observation is valid for both supervised pre-training and contrastive learning. With purified nodes, it turns out that clean training is enough to achieve adversarial robustness in downstream tasks. Yue Xing 0002, Xiaofeng Lin 0005, Qifan Song, Yi Xu 0011, Belinda Zeng, Guang Cheng 0003 |
AISTATS | 1 |
| 2024 | Effect of Ambient-Intrinsic Dimension Gap on Adversarial VulnerabilityabstractThe existence of adversarial attacks on machine learning models imperceptible to a human is still quite a mystery from a theoretical perspective. In this work, we introduce two notions of adversarial attacks: natural or on-manifold attacks, which are perceptible by a human/oracle, and unnatural or off-manifold attacks, which are not. We argue that the existence of the off-manifold attacks is a natural consequence of the dimension gap between the intrinsic and ambient dimensions of the data. For 2-layer ReLU networks, we prove that even though the dimension gap does not affect generalization performance on samples drawn from the observed data space, it makes the clean-trained model more vulnerable to adversarial perturbations in the off-manifold direction of the data space. Our main results provide an explicit relationship between the $\ell_2,\ell_{\infty}$ attack strength of the on/off-manifold attack and the dimension gap. Rajdeep Haldar, Yue Xing 0002, Qifan Song |
AISTATS | 2 |
| 2024 | Unveiling and Mitigating Memorization in Text-to-Image Diffusion Models Through Cross Attention
Jie Ren 0019, Yaxin Li 0001, Shenglai Zeng, Han Xu 0002, Lingjuan Lyu, Yue Xing 0002, Jiliang Tang |
ECCV (77) | 6 |
| 2024 | Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisabstractLarge language models (LLMs) are susceptible to a type of attack known as jailbreaking, which misleads LLMs to output harmful contents.Although there are diverse jailbreak attack strategies, there is no unified understanding on why some methods succeed and others fail.This paper explores the behavior of harmful and harmless prompts in the LLM's representation space to investigate the intrinsic properties of successful jailbreak attacks.We hypothesize that successful attacks share some similar properties: They are effective in moving the representation of the harmful prompt towards the direction to the harmless prompts.We leverage hidden representations into the objective of existing jailbreak attacks to move the attacks along the acceptance direction, and conduct experiments to validate the above hypothesis using the proposed objective.We hope this study provides new insights into understanding how LLMs understand harmfulness information.1 * These authors contributed equally to this work.1 Our code is available at https://github.com/ yuplin2333/representation-space-jailbreak. Yuping Lin, Han Xu 0002, Yue Xing 0002, Makoto Yamada, Hui Liu 0031, Jiliang Tang |
EMNLP | 4 |
| 2022 | Unlabeled Data Help: Minimax Analysis and Adversarial RobustnessabstractThe recent proposed self-supervised learning (SSL) approaches successfully demonstrate the great potential of supplementing learning algorithms with additional unlabeled data. However, it is still unclear whether the existing SSL algorithms can fully utilize the information of both labelled and unlabeled data. This paper gives an affirmative answer for the reconstruction-based SSL algorithm (Lee et al., 2020) under several statistical models. While existing literature only focuses on establishing the upper bound of the convergence rate, we provide a rigorous minimax analysis, and successfully justify the rate-optimality of the reconstruction-based SSL algorithm under different data generation models. Furthermore, we incorporate the reconstruction-based SSL into the exist- ing adversarial training algorithms and show that learning from unlabeled data helps improve the robustness. Yue Xing 0002, Qifan Song, Guang Cheng 0003 |
AISTATS | 1 |
| 2022 | Why Do Artificially Generated Data Help Adversarial RobustnessabstractIn the adversarial training framework of \cite{carmon2019unlabeled,gowal2021improving}, people use generated/real unlabeled data with pseudolabels to improve adversarial robustness. We provide statistical insights to explain why the artificially generated data improve adversarial training. In particular, we study how the attack strength and the quality of the unlabeled data affect adversarial robustness in this framework. Our results show that with a high-quality unlabeled data generator, adversarial training can benefit greatly from this framework under large attack strength, while a poor generator can still help to some extent. To make adaptions concerning the quality of generated data, we propose an algorithm that performs online adjustment to the weight between the labeled real data and the generated data, aiming to optimize the adversarial risk. Numerical studies are conducted to verify our theories and show the effectiveness of the proposed algorithm. Yue Xing 0002, Qifan Song, Guang Cheng 0003 |
NeurIPS | 1 |
| 2022 | Phase Transition from Clean Training to Adversarial TrainingabstractAdversarial training is one important algorithm to achieve robust machine learning models. However, numerous empirical results show a great performance degradation from clean training to adversarial training (e.g., 90+\% vs 67\% testing accuracy on CIFAR-10 dataset), which does not match the theoretical guarantee delivered by the existing studies. Such a gap inspires us to explore the existence of an (asymptotic) phase transition phenomenon with respect to the attack strength: adversarial training is as well behaved as clean training in the small-attack regime, but there is a sharp transition from clean training to adversarial training in the large-attack regime. We validate this conjecture in linear regression models, and conduct comprehensive experiments in deep neural networks. Yue Xing 0002, Qifan Song, Guang Cheng 0003 |
NeurIPS | 1 |
| 2021 | Predictive Power of Nearest Neighbors Algorithm under Random PerturbationabstractThis work investigates the predictive performance of the classical $k$ Nearest Neighbors ($k$-NN) algorithm when the testing data are corrupted by random perturbation. The impact of corruption level on the asymptotic regret is carefully characterized and we reveal a phase-transition phenomenon that, when the corruption level of the random perturbation $\omega$ is below a critical order (i.e., small-$\omega$ regime), the asymptotic regret remains the same; when it is beyond that order (i.e., large-$\omega$ regime), the asymptotic regret deteriorates polynomially. More importantly, the regret of $k$-NN classifier heuristically matches the rate of minimax regret for randomly perturbed testing data, thus implies the strong robustness of $k$-NN against random perturbation on testing data. In fact, we show that the classical $k$-NN can achieve no worse predictive performance, compared to the NN classifiers trained via the popular noise-injection strategy. Our numerical experiment also illustrates that combining $k$-NN component with modern learning algorithms will inherit the strong robustness of $k$-NN. As a technical by-product, we prove that under different model assumptions, the pre-processed 1-NN proposed in \cite{xue2017achieving} will at most achieve a sub-optimal rate when the data dimension $d>4$ even if $k$ is chosen optimally in the pre-processing step. Yue Xing 0002, Qifan Song, Guang Cheng 0003 |
AISTATS | 1 |
| 2021 | On the Generalization Properties of Adversarial TrainingabstractModern machine learning and deep learning models are shown to be vulnerable when testing data are slightly perturbed. Theoretical studies of adversarial training algorithms mostly focus on their adversarial training losses or local convergence properties. In contrast, this paper studies the generalization performance of a generic adversarial training algorithm. Specifically, we consider linear regression models and two-layer neural networks (with lazy training) using squared loss under low-dimensional regime and high-dimensional regime. In the former regime, after overcoming the non-smoothness of adversarial training, the adversarial risk of the trained models will converge to the minimal adversarial risk. In the latter regime, we discover that data interpolation prevents the adversarial robust estimator from being consistent (i.e. converge in probability). Therefore, inspired by successes of the least absolute shrinkage and selection operator (LASSO), we incorporate the $\mathcal{L}_1$ penalty in the high dimensional adversarial learning, and show that it leads to consistent adversarial robust estimation. A series of numerical studies are conducted to demonstrate that how the smoothness and $\mathcal{L}_1$ penalization help to improve the adversarial robustness of DNN models. Yue Xing 0002, Qifan Song, Guang Cheng 0003 |
AISTATS | 1 |
| 2021 | Adversarially Robust Estimate and Risk Analysis in Linear RegressionabstractAdversarial robust learning aims to design algorithms that are robust to small adversarial perturbations on input variables. Beyond the existing studies on the predictive performance to adversarial samples, our goal is to understand statistical properties of adversarial robust estimates and analyze adversarial risk in the setup of linear regression models. By discovering the statistical minimax rate of convergence of adversarial robust estimators, we emphasize the importance of incorporating model information, e.g., sparsity, in adversarial robust learning. Further, we reveal an explicit connection of adversarial and standard estimates, and propose a straightforward two-stage adversarial training framework, which facilitates to utilize model structure information to improve adversarial robustness. In theory, the consistency of the adversarial robust estimator is proven and its Bahadur representation is also developed for the statistical inference purpose. The proposed estimator converges in a sharp rate under either low-dimensional or sparse scenario. Moreover, our theory confirms two phenomena in adversarial robust learning: adversarial robustness hurts generalization, and unlabeled data help improve the generalization. In the end, we conduct numerical simulations to verify our theory. Yue Xing 0002, Guang Cheng 0003 |
AISTATS | 1 |
| 2021 | On the Algorithmic Stability of Adversarial TrainingabstractThe adversarial training is a popular tool to remedy the vulnerability of deep learning models against adversarial attacks, and there is rich theoretical literature on the training loss of adversarial training algorithms. In contrast, this paper studies the algorithmic stability of a generic adversarial training algorithm, which can further help to establish an upper bound for generalization error. By figuring out the stability upper bound and lower bound, we argue that the non-differentiability issue of adversarial training causes worse algorithmic stability than their natural counterparts. To tackle this problem, we consider a noise injection method. While the non-differentiability problem seriously affects the stability of adversarial training, injecting noise enables the training trajectory to avoid the occurrence of non-differentiability with dominating probability, hence enhancing the stability performance of adversarial training. Our analysis also studies the relation between the algorithm stability and numerical approximation error of adversarial attacks. Yue Xing 0002, Qifan Song, Guang Cheng 0003 |
NeurIPS | 1 |
| 2020 | Directional Pruning of Deep Neural NetworksabstractIn the light of the fact that the stochastic gradient descent (SGD) often finds a flat minimum valley in the training loss, we propose a novel directional pruning method which searches for a sparse minimizer in or close to that flat region. The proposed pruning method does not require retraining or the expert knowledge on the sparsity level. To overcome the computational formidability of estimating the flat directions, we propose to use a carefully tuned $\ell_1$ proximal gradient algorithm which can provably achieve the directional pruning with a small learning rate after sufficient training. The empirical results demonstrate the promising results of our solution in highly sparse regime (92% sparsity) among many existing pruning methods on the ResNet50 with the ImageNet, while using only a slightly higher wall time and memory footprint than the SGD. Using the VGG16 and the wide ResNet 28x10 on the CIFAR-10 and CIFAR-100, we demonstrate that our solution reaches the same minima valley as the SGD, and the minima found by our solution and the SGD do not deviate in directions that impact the training loss. The code that reproduces the results of this paper is available at https://github.com/donlan2710/gRDA-Optimizer/tree/master/directional_pruning. Shih-Kang Chao, Zhanyu Wang, Yue Xing 0002, Guang Cheng 0003 |
NeurIPS | 3 |