VLDB 2026 Research / reviewers in the wild / expert
Ganghua Wang
dblp:200/9632
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 43% Efficient and distributed learning · 19% Learning paradigms · 9% | |
| Network and information security
5 papers |
Security and privacy of machine learning · 75% Digital forensics and information hiding · 25% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 20 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
1.6 | 3 | 2024 | Demystifying Poisoning Backdoor Attacks from a Statistical Perspective · ICLR 2024 Understanding Backdoor Attacks through the Adaptability Hypothesis · ICML 2023 A Unified Detection Framework for Inference-Stage Backdoor Defenses · NeurIPS 2023 |
Security and privacy of machine learning › adversarial attack
backdoor attack |
1.4 | 2 | 2024 | Demystifying Poisoning Backdoor Attacks from a Statistical Perspective · ICLR 2024 Understanding Backdoor Attacks through the Adaptability Hypothesis · ICML 2023 |
Information retrieval
retrieval-augmented generation |
0.9 | 1 | 2025 | On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains · ICML 2025 |
Information retrieval
retrieval models |
0.9 | 1 | 2025 | On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains · ICML 2025 |
Security and privacy of machine learning
retrieval-augmented generation security |
0.9 | 1 | 2025 | On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains · ICML 2025 |
Security and privacy of machine learning › poisoning attack
retrieval poisoning attack |
0.9 | 1 | 2025 | On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness › data poisoning
backdoor attack |
0.8 | 1 | 2024 | Demystifying Poisoning Backdoor Attacks from a Statistical Perspective · ICLR 2024 |
Digital forensics and information hiding › watermarking
robust watermarking |
0.8 | 1 | 2024 | RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable Guarantees · NeurIPS 2024 |
Digital forensics and information hiding
watermarking |
0.8 | 1 | 2024 | RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable Guarantees · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › neural network security
backdoor vulnerability |
0.7 | 1 | 2023 | Understanding Backdoor Attacks through the Adaptability Hypothesis · ICML 2023 |
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability |
0.7 | 1 | 2023 | Provable Identifiability of Two-Layer ReLU Neural Networks via LASSO Regularization · IEEE Trans. Inf. Theory 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Pruning Deep Neural Networks from a Sparsity Perspective · ICLR 2023 |
Machine learning › Learning paradigms › supervised learning
neural network regression |
0.7 | 1 | 2023 | Provable Identifiability of Two-Layer ReLU Neural Networks via LASSO Regularization · IEEE Trans. Inf. Theory 2023 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.7 | 1 | 2023 | Pruning Deep Neural Networks from a Sparsity Perspective · ICLR 2023 |
Machine learning › Deep learning architectures and training › ReLU networks
two-layer ReLU network |
0.7 | 1 | 2023 | Provable Identifiability of Two-Layer ReLU Neural Networks via LASSO Regularization · IEEE Trans. Inf. Theory 2023 |
Security and privacy of machine learning › adversarial attack › backdoor attack › backdoor defense
backdoor detection |
0.7 | 1 | 2023 | A Unified Detection Framework for Inference-Stage Backdoor Defenses · NeurIPS 2023 |
Security and privacy of machine learning › adversarial attack › backdoor attack
backdoor poisoning attacks |
0.7 | 1 | 2023 | Understanding Backdoor Attacks through the Adaptability Hypothesis · ICML 2023 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable Guarantees · NeurIPS 2024 |
Machine learning › Learning theory › model selection
variable selection |
0.2 | 1 | 2023 | Provable Identifiability of Two-Layer ReLU Neural Networks via LASSO Regularization · IEEE Trans. Inf. Theory 2023 |
Methods — techniques the papers use, named apart from their topics
conformal prediction · 2.8embedding similarity analysis · 2.6detection-based defense · 2.6statistical bounds · 1.5randomized smoothing · 1.5lower and upper bounds · 1.5adversarial training · 1.5kernel-based learning · 1.3adaptability hypothesis · 1.3sparsity analysis · 0.7latent representation analysis · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsabstractRetrieval-Augmented Generation (RAG) has been empirically shown to enhance the performance of large language models (LLMs) in knowledge-intensive domains such as healthcare, finance, and legal contexts. Given a query, RAG retrieves relevant documents from a corpus and integrates them into the LLMs’ generation process. In this study, we investigate the adversarial robustness of RAG, focusing specifically on examining the retrieval system. First, across 225 different setup combinations of corpus, retriever, query, and targeted information, we show that retrieval systems are vulnerable to universal poisoning attacks in medical Q&A. In such attacks, adversaries generate poisoned documents containing a broad spectrum of targeted information, such as personally identifiable information. When these poisoned documents are inserted into a corpus, they can be accurately retrieved by any users, as long as attacker-specified queries are used. To understand this vulnerability, we discovered that the deviation from the query’s embedding to that of the poisoned document tends to follow a pattern in which the high similarity between the poisoned document and the query is retained, thereby enabling precise retrieval. Based on these findings, we develop a new detection-based defense to ensure the safe use of RAG. Through extensive experiments spanning various Q&A domains, we observed that our proposed method consistently achieves excellent detection rates in nearly all cases. Xun Xian, Ganghua Wang, Xuan Bi, Rui Zhang 0028, Jayanth Srinivasa, Ashish Kundu, Charles Fleming, Mingyi Hong 0001, Jie Ding 0002 |
ICML | 2 |
| 2024 | Demystifying Poisoning Backdoor Attacks from a Statistical PerspectiveabstractBackdoor attacks pose a significant security risk to machine learning applications due to their stealthy nature and potentially serious consequences. Such attacks involve embedding triggers within a learning model with the intention of causing malicious behavior when an active trigger is present while maintaining regular functionality without it. This paper derives a fundamental understanding of backdoor attacks that applies to both discriminative and generative models, including diffusion models and large language models. We evaluate the effectiveness of any backdoor attack incorporating a constant trigger, by establishing tight lower and upper boundaries for the performance of the compromised model on both clean and backdoor test data. The developed theory answers a series of fundamental but previously underexplored problems, including (1) what are the determining factors for a backdoor attack's success, (2) what is the direction of the most effective backdoor attack, and (3) when will a human-imperceptible trigger succeed. We demonstrate the theory by conducting experiments using benchmark datasets and state-of-the-art backdoor attack scenarios. Our code is available \href{https://github.com/KeyWgh/DemystifyBackdoor}{here}. Ganghua Wang, Xun Xian, Ashish Kundu, Jayanth Srinivasa, Xuan Bi, Mingyi Hong 0001, Jie Ding 0002 |
ICLR | 1 |
| 2024 | RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable GuaranteesabstractSafeguarding intellectual property and preventing potential misuse of AI-generated images are of paramount importance. This paper introduces a robust and agile plug-and-play watermark detection framework, referred to as RAW.
As a departure from existing encoder-decoder methods, which incorporate fixed binary codes as watermarks within latent representations, our approach introduces learnable watermarks directly into the original image data. Subsequently, we employ a classifier that is jointly trained with the watermark to detect the presence of the watermark.
The proposed framework is compatible with various generative architectures and supports on-the-fly watermark injection after training. By incorporating state-of-the-art smoothing techniques, we show that the framework also provides provable guarantees regarding the false positive rate for misclassifying a watermarked image, even in the presence of adversarial attacks targeting watermark removal.
Experiments on a diverse range of images generated by state-of-the-art diffusion models demonstrate substantially improved watermark encoding speed and watermark detection performance, under adversarial attacks, while maintaining image quality. Our code is publicly available [here](https://github.com/jeremyxianx/RAWatermark). Xun Xian, Ganghua Wang, Xuan Bi, Jayanth Srinivasa, Ashish Kundu, Mingyi Hong 0001, Jie Ding 0002 |
NeurIPS | 2 |
| 2023 | Pruning Deep Neural Networks from a Sparsity Perspective
Enmao Diao, Ganghua Wang, Jiawei Zhang 0007, Yuhong Yang 0002, Jie Ding 0002, Vahid Tarokh |
ICLR | 2 |
| 2023 | Understanding Backdoor Attacks through the Adaptability HypothesisabstractA poisoning backdoor attack is a rising security concern for deep learning. This type of attack can result in the backdoored model functioning normally most of the time but exhibiting abnormal behavior when presented with inputs containing the backdoor trigger, making it difficult to detect and prevent. In this work, we propose the adaptability hypothesis to understand when and why a backdoor attack works for general learning models, including deep neural networks, based on the theoretical investigation of classical kernel-based learning models. The adaptability hypothesis postulates that for an effective attack, the effect of incorporating a new dataset on the predictions of the original data points will be small, provided that the original data points are distant from the new dataset. Experiments on benchmark image datasets and state-of-the-art backdoor attacks for deep neural networks are conducted to corroborate the hypothesis. Our finding provides insight into the factors that affect the attack’s effectiveness and has implications for the design of future attacks and defenses. Xun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu, Xuan Bi, Mingyi Hong 0001, Jie Ding 0002 |
ICML | 2 |
| 2023 | A Unified Detection Framework for Inference-Stage Backdoor DefensesabstractBackdoor attacks involve inserting poisoned samples during training, resulting in a model containing a hidden backdoor that can trigger specific behaviors without impacting performance on normal samples. These attacks are challenging to detect, as the backdoored model appears normal until activated by the backdoor trigger, rendering them particularly stealthy. In this study, we devise a unified inference-stage detection framework to defend against backdoor attacks. We first rigorously formulate the inference-stage backdoor detection problem, encompassing various existing methods, and discuss several challenges and limitations. We then propose a framework with provable guarantees on the false positive rate or the probability of misclassifying a clean sample. Further, we derive the most powerful detection rule to maximize the detection power, namely the rate of accurately identifying a backdoor sample, given a false positive rate under classical learning scenarios. Based on the theoretically optimal detection rule, we suggest a practical and effective approach for real-world applications based on the latent representations of backdoored deep nets. We extensively evaluate our method on 14 different backdoor attacks using Computer Vision (CV) and Natural Language Processing (NLP) benchmark datasets. The experimental findings align with our theoretical results. We significantly surpass the state-of-the-art methods, e.g., up to 300\% improvement on the detection power as evaluated by AUCROC, over the state-of-the-art defense against advanced adaptive backdoor attacks. Xun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu, Xuan Bi, Mingyi Hong 0001, Jie Ding 0002 |
NeurIPS | 2 |
| 2023 | Provable Identifiability of Two-Layer ReLU Neural Networks via LASSO RegularizationabstractLASSO regularization is a popular regression tool to enhance the prediction accuracy of statistical models by performing variable selection through the$\ell _{1}$penalty, initially formulated for the linear model and its variants. In this paper, the territory of LASSO is extended to two-layer ReLU neural networks, a fashionable and powerful nonlinear regression model. Specifically, given a neural network whose output$y$depends only on a small subset of input$\boldsymbol {x}$, denoted by$\mathcal {S}^{\star }$, we prove that the LASSO estimator can stably reconstruct the neural network and identify$\mathcal {S}^{\star }$when the number of samples scales logarithmically with the input dimension. This challenging regime has been well understood for linear models while barely studied for neural networks. Our theory lies in an extended Restricted Isometry Property (RIP)-based analysis framework for two-layer ReLU neural networks, which may be of independent interest to other LASSO or neural network settings. Based on the result, we advocate a neural network-based variable selection method. Experiments on simulated and real-world datasets show promising performance of the variable selection approach compared with existing techniques. Gen Li 0005, Ganghua Wang, Jie Ding 0002 |
IEEE Trans. Inf. Theory | 2 |