VLDB 2026 Research / reviewers in the wild / expert
Shang Wang 0004
dblp:53/448-4
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-5114-4659ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unshaken by Weak Embedding: Robust Probabilistic Watermarking for Dataset Copyright Protection
Shang Wang 0004, Tianqing Zhu, Dayong Ye, Bo Liu 0001, Ming Ding 0001, Shengfang Zhai, Yansong Gao 0001 |
NDSS | 1 |
| 2026 | When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?abstractThe deployment of large language models (LLMs) like ChatGPT and Gemini has shown their powerful natural language generation capabilities. However, these models can inadvertently learn and retain sensitive information and harmful content during training, raising significant ethical and legal concerns. To address these issues, machine unlearning has been introduced as a potential solution. While existing unlearning methods take into account the specific characteristics of LLMs, they often suffer from high computational demands, limited applicability, or the risk of catastrophic forgetting. To address these limitations, we propose a lightweight behavioral unlearning framework based on Retrieval-Augmented Generation (RAG) technology. By modifying the external knowledge base of RAG, we simulate the effects of forgetting without directly interacting with the unlearned LLM. We approach the construction of unlearned knowledge as a constrained optimization problem, deriving two key components that underpin the effectiveness of RAG-based unlearning. This RAG-based approach is particularly effective for closed-source LLMs, where existing unlearning methods often fail. We evaluate our framework through extensive experiments on both open-source and closed-source models, including ChatGPT, Gemini, Llama-2-7b-chat, and PaLM 2. The results demonstrate that our approach meets five key unlearning criteria: effectiveness, universality, harmlessness, simplicity, and robustness. Meanwhile, this approach can extend to multimodal large language models and LLM-based agents. Shang Wang 0004, Tianqing Zhu, Dayong Ye, Wanlei Zhou 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Parameter-Agnostic Privacy-Preserving Machine Unlearning for Large Language ModelsabstractIn recent years, advancements in large language models have led to significant innovation and critical progress in AI. However, some of these innovations are raising privacy and security concerns. Machine unlearning has therefore emerged as a potential solution to mitigate such risks. Yet, while erasing data records from traditional models is relatively straightforward, making a large language model “forget” what it has learned is often very challenging. This is not just because they include so many parameters, it is also because the knowledge they possess is intricately entangled. Further, the privacy risk of unlearned data remains neglected in most unlearning solutions. To overcome these limitations, we took advantage of information retrieval and developed an efficient privacy-preserving unlearning mechanism. Our solution eliminates the impact of targeted information by removing high-risk semantic meanings from the model’s output. It also incorporates differentially-private randomization to make the unlearned information statistically indiscernible. Most importantly, the algorithm requires neither parametric fine-tuning nor in-context prompt calibration. A theoretical analysis demonstrates that this method satisfies rigorous privacy and unlearning guarantees. Additionally, experiments on real-world datasets prove that the method is both effective and has the capacity to handle practical unlearning tasks for large language model applications. Lefeng Zhang, Tianqing Zhu, Zihan Xie, Shang Wang 0004, Binxing Fang, Wanlei Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Data-Free Model-Related Attacks: Unleashing the Potential of Generative AI
Dayong Ye, Tianqing Zhu, Shang Wang 0004, Bo Liu 0001, Leo Yu Zhang, Wanlei Zhou 0001, Yang Zhang 0016 |
USENIX Security Symposium | 3 |
| 2024 | Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade DefenseabstractAll current backdoor attacks on deep learning (DL) models fall under the category of a vertical class backdoor (VCB).In VCB attacks, any sample from a class activates the implanted backdoor when the secret trigger is present, regardless of whether it is a sub-type source-class-agnostic backdoor or a source-class-specific backdoor. For example, a trigger of sunglasses could mislead a facial recognition model when either an arbitrary (source-class-agnostic) or a specific (source-class-specific) person wears sunglasses. Existing defense strategiesoverwhelmingly focus on countering VCB attacks, especially those that are source-class-agnostic. This narrow focus neglects the potential threat of other simpler yet general backdoor types, leading to false security implications. It is, therefore, crucial to discover and elucidate unknown backdoor types, particularly those that can be easily implemented, as a mandatory step before developing countermeasures. Shang Wang 0004, Yansong Gao 0001, Zhi Zhang 0001, Huming Qiu, Minhui Xue 0001, Alsharif Abuadbba, Anmin Fu, Surya Nepal, Derek Abbott |
CCS | 2 |
| 2024 | CareFL: Contribution Guided Byzantine-Robust Federated LearningabstractByzantine-robust federated learning (FL) endeavors to empower service providers in acquiring a precise global model, even in the presence of potentially malicious FL clients. While considerable strides have been taken in the development of robust aggregation algorithms for FL in recent years, their efficacy is confined to addressing particular forms of Byzantine attacks, and they exhibit vulnerabilities when confronted with a spectrum of attack vectors. Notably, a prevailing issue lies in the heavy reliance of these algorithms on the examination of local model gradients. It is worth noting that an attacker possesses the ability to manipulate a carefully chosen small gradient of a model within a context where there could be millions of gradients available, thereby facilitating adaptive attacks. Drawing inspiration from the foundational Shapley value methodology in game theory, we introduce an effective FL scheme namedCareFL. This scheme is designed to provide robustness against a spectrum of state-of-the-art Byzantine attacks. Unlike approaches that rely on the examination of gradients,CareFLemploys a universal metric, the loss of the local model—independent of specific gradients, to identify potentially malicious clients. Specifically, in each aggregation round, the FL server trains a reference model using a small auxiliary dataset— the auxiliary dataset can be removed with a slight defense degradation trade-off. It employs the Shapley value to assess the contribution of each client-submitted model in minimizing the global model loss. Subsequently, the server selects client models closer to the reference model in terms of Shapley values for the global model update. To reduce the computational overhead ofCareFLwhen the number of clients is relatively scaled-up, we construct its variant, namelyCareFL+ generally by grouping clients. Extensive experimentation conducted on well-established MNIST and CIFAR-10 datasets, encompassing diverse model architectures, including AlexNet, demonstrates thatCareFLconsistently achieves accuracy levels comparable to those attained under attack-free conditions when faced with five formidable attacks.CareFLand CareFL+ outperform six existing state-of-the-art Byzantine-robust FL aggregation methods, includingFLTrust, across both IID and non-IID data distribution settings. Qihao Dong, Shengyuan Yang, Zhiyang Dai, Yansong Gao 0001, Shang Wang 0004, Yuan Cao 0003, Anmin Fu, Willy Susilo |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | CASSOCK: Viable Backdoor Attacks against DNN in the Wall of Source-Specific Backdoor DefensesabstractAs a critical threat to deep neural networks (DNNs), backdoor attacks can be categorized into two types, i.e., source-agnostic backdoor attacks (SABAs) and source-specific backdoor attacks (SSBAs). Compared to traditional SABAs, SSBAs are more advanced in that they have superior stealthier in bypassing mainstream countermeasures that are effective against SABAs. Nonetheless, existing SSBAs suffer from two major limitations. First, they can hardly achieve a good trade-off between ASR (attack success rate) and FPR (false positive rate). Besides, they can be effectively detected by the state-of-the-art (SOTA) countermeasures (e.g., SCAn [40]). Shang Wang 0004, Yansong Gao 0001, Anmin Fu, Zhi Zhang 0001, Yuqing Zhang 0001, Willy Susilo, Dongxi Liu |
AsiaCCS | 1 |
| 2022 | LinkBreaker: Breaking the Backdoor-Trigger Link in DNNs via Neurons Consistency CheckabstractBackdoor attacks cause model misbehaving by first implanting backdoors in deep neural networks (DNNs) during training and then activating the backdoor via samples with triggers during inference. The compromised models could pose serious security risks to artificial intelligence systems, such as misidentifying ‘stop’ traffic sign into ‘80km/h’. In this paper, we investigate the connection characteristic between the backdoor and the trigger in DNNs and observe the fact that the backdoor is implanted via establishing a link between a cluster of neurons, representing the backdoor, and the triggers. Based on this observation, we design LinkBreaker, a new generic scheme for defending against backdoor attacks. In particular, LinkBreaker deploys a neuron consistency check mechanism for identifying compromised neuron set related to the trigger. Then, the LinkBreaker regulates the model to make predictions based on benign neuron set only and thus breaks the link between the backdoor and the trigger. Compared to previous defenses, LinkBreaker offers a more general backdoor countermeasure that is not only effective against input-agnostic backdoors but also source-specific backdoors, which the later can not be defeated by majority of state-of-the-arts. Besides, LinkBreaker is robust against adversarial examples, which, to a large extent, provides a holistic defense against adversarial example attacks on DNNs, while almost all current backdoor defenses do not have such consideration and capability. Extensive experimental evaluations on real datasets demonstrate that LinkBreaker is with high efficacy of suppressing trigger inputs while incurring no noticeable accuracy deterioration on benign inputs. Zhenzhu Chen, Shang Wang 0004, Anmin Fu, Yansong Gao 0001, Shui Yu 0001, Robert H. Deng |
IEEE Trans. Inf. Forensics Secur. | 2 |