Mengyuan Sun 0001

dblp:209/5193-1 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0001-0002-0680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 InverTune: A Backdoor Defense Method for Multimodal Contrastive Learning via Backdoor-Adversarial Correlation Analysis
Mengyuan Sun 0001, Yu Li 0006, Yunjie Ge, Bo Du 0001, Qian Wang 0002
NDSS1
2026 Armor: Shielding Unlearnable Examples Against Data Augmentation
abstract
Private data, when published online, may be collected by unauthorized parties to train deep neural networks (DNNs). To protect privacy, defensive noises can be added to original samples to degrade their learnability by DNNs. Recently, unlearnable examples (Huang et al., 2021) are proposed to minimize the training loss such that the model learns almost nothing. However, raw data are often pre-processed before being used for training, which may restore the private information of protected data. In this paper, we reveal the data privacy violation induced by data augmentation, a commonly used data pre-processing technique to improve model generalization capability, which is the first of its kind as far as we are concerned. We demonstrate that data augmentation can significantly raise the accuracy of the model trained on unlearnable examples from 21.3% to 66.1%. To address this issue, we propose a defense framework, dubbed Armor, to protect data privacy from potential breaches of data augmentation. To overcome the difficulty of having no access to the model training process, we design a non-local module-assisted surrogate model that better captures the effect of data augmentation. In addition, we design a surrogate augmentation selection strategy that maximizes distribution alignment between augmented and non-augmented samples, to choose the optimal augmentation strategy for each class. We also use a dynamic step size adjustment algorithm to enhance the defensive noise generation process. Extensive experiments are conducted on 4 datasets and 5 data augmentation methods to verify the performance of Armor. Comparisons with 6 state-of-the-art defense methods have demonstrated that Armor can preserve the unlearnability of protected private data under data augmentation. Armor reduces the test accuracy of the model trained on augmented protected samples by as much as 60% more than baselines. We also show that Armor is robust to adversarial training. We will open-source our codes upon publication.
Xueluan Gong, Yuji Wang, Yanjiao Chen, Haocheng Dong, Yiming Li 0004, Mengyuan Sun 0001, Shuaike Li, Qian Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Sleight: Hidden Data Privacy Breaches in Federated Learning
abstract
Federated Learning (FL) has emerged as a paradigm for conducting machine learning across broad and decentralized datasets, promising enhanced privacy by obviating the need for direct data sharing. However, recent studies show that attackers can steal private data through model manipulation or gradient analysis. Existing attacks are constrained by low theft quantity or low-resolution data, and they are often easily detected through anomaly monitoring in gradients or weights. In this paper, we propose Sleight, a novel data-reconstruction attack, supported by two key techniques, i.e., distinctive and sparse encoding design and block partitioning. Unlike conventional methods that require detectable changes to the model, Sleight stealthily embeds a hidden model using parameter sharing to systematically extract sensitive data. The Fibonacci-based index design ensures efficient, structured retrieval of memorized data, while the block partitioning method enhances Sleight's capability to handle high-resolution images by dividing them into smaller, manageable units. Extensive experiments on 4 datasets confirmed that Sleight is superior to 5 state-of-the-art data-reconstruction attacks under 5 respective detection methods. Sleight can handle large-scale and high-resolution data without being detected or mitigated by state-of-the-art data reconstruction defense methods. In contrast to baselines, Sleight can be directly applied to both FedAvg and FedSGD scenarios, underscoring the need for developers to devise new defenses against such vulnerabilities. We will open-source our code upon acceptance.
Xueluan Gong, Yuji Wang, Shuike Li, Mengyuan Sun 0001, Chen Chen 0115, Qian Wang 0002, Kwok-Yan Lam
IEEE Trans. Dependable Secur. Comput.4