Tianshuo Cong

dblp:233/3780 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-3189-8223ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Covert Knowledge Poisoning Attacks in Retrieval-Augmented Code Generation
abstract
Retrieval-Augmented Code Generation (RACG) systems enhance code generation by dynamically integrating examples retrieved from open knowledge bases, but their reliance on external sources exposes them to knowledge poisoning. While prior attacks inject explicit vulnerable code, such payloads are easily detected, limiting their real-world impact and failing to probe the full attack surface. To overcome this limitation, we propose Arachne, a covert knowledge poisoning attack that, for the first time, achieves fine-grained control over vulnerability types while evading detection. Arachne relies on two mechanisms: (1) Benign-Appearing Fragment Construction, where vulnerable code is decomposed into benign-appearing fragments that pre serve compilability and contextual cues, rendering them effectively undetectable by vulnerability analyzers; and (2) Retrieval Driven Knowledge Completion, in which retrieved fragments activate the large language model's (LLM) contextual reasoning to autonomously generate complete vulnerable code. We evaluate Arachne on eight mainstream LLMs and four retrievers across four common vulnerability types. On GPT-4, Qwen2.5, Gemini 3-flash, Claude-4-sonnet and DeepSeekCoder, Arachne achieves attack success rates exceeding 70% in most settings for each vulnerability type across CWE-78, CWE-295, CWE-367, and CWE-614, peaking at 100% for CWE-367. Arachne demonstrates up to a 37% improvement in attack success rate over existing poisoning attacks. In two real-world applications, the attack achieves 87% success rate with 70% benign utility, indicating that the generated code often satisfies user requirements while introducing vulnerable code. Critically, our poisoned samples, which are designed without explicit malicious patterns, fully bypass rule-based analyzers and challenge state-of-the-art LLM based detectors, exemplifying their stealth among multiple failed defense paradigms. Our findings expose critical limitations in existing defense frameworks, highlighting the urgent need for novel defense mechanisms specifically designed for RACG systems.
Xinlei He 0001, Tianshuo Cong, Ke Xu 0002, Qi Li 0002
IEEE Trans. Dependable Secur. Comput.4
2026 Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models
abstract
Large Language Models (LLMs) undergo continuous updates to improve user experience. However, prior research on the security and safety implications of LLMs has primarily focused on their specific versions, overlooking the impact of successive LLM updates. This prompts the need for a holistic understanding of the risks in these different versions of LLMs. To fill this gap, in this paper, we conduct a longitudinal study to examine the adversarial robustness – specifically misclassification, jailbreak, and hallucination – of three prominent LLM families: GPT, Llama, and Qwen. Our study reveals that LLM updates do not consistently improve adversarial robustness as expected. For instance, a later version of GPT-3.5 degrades regarding misclassification and hallucination despite its improved resilience against jailbreaks. GPT-4 and GPT-4o demonstrate (incrementally) higher robustness overall. Larger Llama and Qwen models do not uniformly exhibit improved robustness across all three aspects studied. In addition, larger model sizes do not necessarily yield improved robustness. Minor updates lacking substantial robustness improvements can exacerbate existing issues rather than resolve them. We hope our study can offer valuable insights into navigating model updates and informed decisions in model development and usage.
Yugeng Liu, Tianshuo Cong, Zhengyu Zhao 0001, Michael Backes 0001, Yang Zhang 0016
IEEE Trans. Inf. Forensics Secur.2
2025 FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
abstract
Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating additional modalities (e.g., images). Despite this advancement, the safety of LVLMs remains adequately underexplored, with a potential overreliance on the safety assurances purported by their underlying LLMs. In this paper, we propose FigStep, a straightforward yet effective black-box jailbreak algorithm against LVLMs. Instead of feeding textual harmful instructions directly, FigStep converts the prohibited content into images through typography to bypass the safety alignment. The experimental results indicate that FigStep can achieve an average attack success rate of 82.50% on six promising open-source LVLMs. Not merely to demonstrate the efficacy of FigStep, we conduct comprehensive ablation studies and analyze the distribution of the semantic embeddings to uncover that the reason behind the success of FigStep is the deficiency of safety alignment for visual embeddings. Moreover, we compare FigStep with five text-only jailbreaks and four image-based jailbreaks to demonstrate the superiority of FigStep, i.e., negligible attack costs and better attack performance. Above all, our work reveals that current LVLMs are vulnerable to jailbreak attacks, which highlights the necessity of novel cross-modality safety alignment techniques.
Yichen Gong, Delong Ran, Conglei Wang, Tianshuo Cong, Anyu Wang 0001, Sisi Duan, Xiaoyun Wang 0001
AAAI5
2025 CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
abstract
Backdoor attacks significantly compromise the security of large language models by triggering them to output specific and controlled content. Currently, triggers for textual backdoor attacks fall into two categories: fixed-token triggers and sentence-pattern triggers. However, the former are typically easy to identify and filter, while the latter, such as syntax and style, do not apply to all original samples and may lead to semantic shifts. In this paper, inspired by cross-lingual (CL) prompts of LLMs in real-world scenarios, we propose a higher-dimensional trigger method at the paragraph level, namely CL-Attack. CL-Attack injects the backdoor by using texts with specific structures that incorporate multiple languages, thereby offering greater stealthiness and universality compared to existing backdoor attack techniques. Extensive experiments on different tasks and model architectures demonstrate that CL-Attack can achieve nearly 100 percents attack success rate with a low poisoning rate in both classification and generation tasks. We also empirically show that CL-Attack is more robust against current major defense methods compared to baseline backdoor attacks. Additionally, in response to CL-Attack, we further develop a new defense called TranslateDefense, which can partially mitigate the impact of CL-Attack.
Jingyi Zheng, Tianshuo Cong, Xinlei He 0001
AAAI3
2025 Safety Misalignment Against Large Language Models
Yichen Gong, Delong Ran, Xinlei He 0001, Tianshuo Cong, Anyu Wang 0001, Xiaoyun Wang 0001
NDSS4
2025 ErrorTrace: A Black-Box Traceability Mechanism Based on Model Family Error Space
abstract
The open-source release of large language models (LLMs) enables malicious users to create unauthorized derivative models at low cost, posing significant threats to intellectual property (IP) and market stability. Existing IP protection methods either require access to model parameters or are vulnerable to fine-tuning attacks. To fill this gap, we propose ErrorTrace, a robust and black-box traceability mechanism for protecting LLM IP. Specifically, ErrorTrace leverages the unique error patterns of model families by mapping and analyzing their distinct error spaces, enabling robust and efficient IP protection without relying on internal parameters or specific query responses. Experimental results show that ErrorTrace achieves a traceability accuracy of 0.8518 for 27 base models when the suspect model is not included in ErrorTrace's training set, outperforming the baseline by 0.2593. Additionally,ErrorTrace successfully tracks 34 fine-tuned, pruned and merged models across various scenarios, demonstrating its broad applicability and robustness. In addition, ErrorTrace shows a certain level of resilience when subjected to adversarial attacks. Our code is available at: https://github.com/csdatazcc/ErrorTrace.
Chuanchao Zang, Xiangtao Meng, Tianshuo Cong, Yaxing Zha, Dong Qi, Zheng Li 0023, Shanqing Guo
NeurIPS4
2025 PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
abstract
Fine-tuning is an essential process to improve the performance of Large Language Models (LLMs) in specific domains, with Parameter-Efficient Fine-Tuning (PEFT) gaining popularity due to its capacity to reduce computational demands through the integration of low-rank adapters. These lightweight adapters, such as LoRA, can be shared and utilized on open-source platforms. However, adversaries could exploit this mechanism to inject backdoors into these adapters, resulting in malicious behaviors like incorrect or harmful outputs, which pose serious security risks to the community. Unfortunately, few current efforts concentrate on analyzing the backdoor patterns or detecting the backdoors in the adapters. To fill this gap, we first construct and release PADBench, a comprehensive benchmark that contains 13, 300 benign and backdoored adapters fine-tuned with various datasets, attack strategies, PEFT methods, and LLMs. Moreover, we propose PEFTGuard, the first backdoor detection framework against PEFT-based adapters. Extensive evaluation upon PADBench shows that PEFTGuard outperforms existing detection methods, achieving nearly perfect detection accuracy (100%) in most cases. Notably, PEFTGuard exhibits zero-shot transferability on three aspects, including different attacks, PEFT methods, and adapter ranks. In addition, we consider various adaptive attacks to demonstrate the high robustness of PEFTGuard. We further explore several possible backdoor mitigation defenses, finding fine-mixing to be the most effective method. We envision that our benchmark and method can shed light on future LLM backdoor detection research.11Our code and dataset are available at: https://github.com/Vincent-HKUSTGZ/PEFTGuard.
Zhen Sun 0001, Tianshuo Cong, Yule Liu, Chenhao Lin, Xinlei He 0001, Rongmao Chen, Xingshuo Han, Xinyi Huang 0001
SP2
2025 From Purity to Peril: Backdooring Merged Models From "Harmless" Benign Components
Lijin Wang, Tianshuo Cong, Xinlei He 0001, Zhan Qin, Xinyi Huang 0001
USENIX Security Symposium3
2024 Test-Time Poisoning Attacks Against Test-Time Adaptation Models
abstract
Deploying machine learning (ML) models in the wild is challenging as it suffers from distribution shifts, where the model trained on an original domain cannot generalize well to unforeseen diverse transfer domains. To address this challenge, several test-time adaptation (TTA) methods have been proposed to improve the generalization ability of the target pre-trained models under test data to cope with the shifted distribution. The success of TTA can be credited to the continuous fine-tuning of the target model according to the distributional hint from the test samples during test time. Despite being powerful, it also opens a new attack surface, i.e., test-time poisoning attacks, which are substantially different from previous poisoning attacks that occur during the training time of ML models (i.e., adversaries cannot intervene in the training process). In this paper, we perform the first test-time poisoning attack against four mainstream TTA methods, including TTT, DUA, TENT, and RPL. Concretely, we generate poisoned samples based on the surrogate models and feed them to the target TTA models. Experimental results show that the TTA methods are generally vulnerable to test-time poisoning attacks. For instance, the adversary can feed as few as 10 poisoned samples to degrade the performance of the target model from 76.20% to 41.83%. Our results demonstrate that TTA algorithms lacking a rigorous security assessment are unsuitable for deployment in real-life scenarios. As such, we advocate for the integration of defenses against test-time poisoning attacks into the design of TTA methods.1
Tianshuo Cong, Xinlei He 0001, Yang Zhang 0016
SP1
2022 SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained Encoders
abstract
Self-supervised learning is an emerging machine learning (ML) paradigm. Compared to supervised learning which leverages high-quality labeled datasets, self-supervised learning relies on unlabeled datasets to pre-train powerful encoders which can then be treated as feature extractors for various downstream tasks. The huge amount of data and computational resources consumption makes the encoders themselves become the valuable intellectual property of the model owner. Recent research has shown that the ML model's copyright is threatened by model stealing attacks, which aim to train a surrogate model to mimic the behavior of a given model. We empirically show that pre-trained encoders are highly vulnerable to model stealing attacks. However, most of the current efforts of copyright protection algorithms such as watermarking concentrate on classifiers. Meanwhile, the intrinsic challenges of pre-trained encoder's copyright protection remain largely unstudied. We fill the gap by proposing SSLGuard, the first watermarking scheme for pre-trained encoders. Given a clean pre-trained encoder, SSLGuard injects a watermark into it and outputs a watermarked version. The shadow training technique is also applied to preserve the watermark under potential model stealing attacks. Our extensive evaluation shows that SSLGuard is effective in watermark injection and verification, and it is robust against model stealing and other watermark removal attacks such as input noising, output perturbing, overwriting, model pruning, and fine-tuning.
Tianshuo Cong, Xinlei He 0001, Yang Zhang 0016
CCS1