EDBT 2026 Demo / reviewers in the wild / expert
Yinggui Wang
dblp:136/1775 · also Ying-Gui Wang
· DBLP profile ↗
23ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-6686-6603ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
Wei Huang 0039, Anda Cheng, Yinggui Wang, Lei Wang 0251, Tao Wei 0002 |
Proc. VLDB Endow. | 3 |
| 2025 | GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language ModelsabstractKai Yao, Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao, Yixin Ji, Jianke Zhu, Wei Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao 0015, Yixin Ji, Jianke Zhu, Wei Wang 0002 |
ACL (1) | 6 |
| 2025 | Transferable Adversarial Examples with Bayesian Approach
Mingyuan Fan 0003, Cen Chen 0001, Wenmeng Zhou, Yinggui Wang |
AsiaCCS | 4 |
| 2025 | A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and OptimizationabstractLarge Language Model Unlearning (LLMU) is a promising way to remove private or sensitive information from large language models.However, the comprehensive evaluation of LLMU remains underexplored.The dominant deterministic evaluation can yield overly optimistic assessments of unlearning efficacy.To mitigate this, we propose a Fully Probabilistic Evaluation (FPE) framework that incorporates input and output distributions in LLMU evaluation.FPE obtains a probabilistic evaluation result by querying unlearned models with various semantically similar inputs and multiple sampling attempts.We introduce an Input Distribution Sampling method in FPE to select high-quality inputs, enabling a stricter measure of information leakage risks.Furthermore, we introduce a Contrastive Embedding Loss (CEL) to advance the performance of LLMU.CEL employs contrastive learning to distance latent representations of unlearned samples from adaptively clustered contrast samples while aligning them with random vectors, leading to improved efficacy and robustness for LLMU.Our experiments show that FPE uncovers more unlearned information leakage risks than prior evaluation methods, and CEL improves unlearning effectiveness by at least 50.1% and robustness by at least 37.2% on Llama-2-7B while retaining high model utility. Anda Cheng, Wei Huang 0039, Yinggui Wang |
EMNLP | 3 |
| 2025 | Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware PruningabstractRecent advancements in large language models (LLMs) have shown impressive capabilities in various downstream tasks but typically face Catastrophic Forgetting (CF) during fine-tuning.In this paper, we propose the Forgetting-Aware Pruning Metric (FAPM), a novel pruning-based approach to balance CF and downstream task performance.Our investigation reveals that the degree to which task vectors (i.e., the subtraction of pre-trained weights from the weights fine-tuned on downstream tasks) overlap with pre-trained model parameters is a critical factor for CF.Based on this finding, FAPM employs the ratio of the task vector to pre-trained model parameters as a metric to quantify CF, integrating this measure into the pruning criteria.Importantly, FAPM does not necessitate modifications to the training process or model architecture, nor does it require any auxiliary data.We conducted extensive experiments across eight datasets, covering natural language inference, General Q&A, Medical Q&A, Math Q&A, reading comprehension, and cloze tests.The results demonstrate that FAPM limits CF to just 0.25% while maintaining 99.67% accuracy on downstream tasks.We provide the code to reproduce our results.1 . Wei Huang 0039, Anda Cheng, Yinggui Wang |
EMNLP | 3 |
| 2025 | Fine-grained Prompt Screening: Defending Against Backdoor Attack on Text-to-Image Diffusion ModelsabstractText-to-image (T2I) diffusion models exhibit impressive generation capabilities in recently studies. However, they are vulnerable to backdoor attacks, where model outputs are manipulated by malicious triggers. In this paper, we propose a novel input-level defense method, called Fine-grained Prompt Screening (GrainPS). Our method is motivated by the phenomenon, i.e., Semantics Misalignment, where the backdoor trigger causes the inconsistency between the cross-attention projections of object words (the key words to determine the main content of the generated image) and their true semantics. In particular, we divide each prompt into pieces and conduct fine-grained analysis by examining the impact of the trigger on object words in the cross-attention layers rather than their global influence on the entire generated image. To assess the impact of each word on object words, we formulate "semantics alignment score'' as the metric with a carefully crafted detection strategy to identify the trigger. Therefore, our implementation can detect backdoor input prompts and localize of triggers simultaneously. Evaluations across four advanced backdoor attack scenarios demonstrate the effectiveness of our proposed defense method. Nan Zhong, Guobiao Li, Anda Cheng, Yinggui Wang, Zhenxing Qian, Xinpeng Zhang 0001 |
IJCAI | 5 |
| 2025 | AnchorSync: Global Consistency Optimization for Long Video Editing
Zichi Liu, Yinggui Wang, Tao Wei 0002, Chao Ma 0004 |
ACM Multimedia | 2 |
| 2025 | AegisGuard: RL-Guided Adapter Tuning for TEE-Based Efficient & Secure On-Device InferenceabstractOn-device large models (LMs) reduce cloud dependency but expose proprietary model weights to the end-user, making them vulnerable to white-box model stealing (MS) attacks. A common defense is TEE-Shielded DNN Partition (TSDP), which places all trainable LoRA adapters (fine tuned on private data) inside a trusted execution environment (TEE). However, this design suffers from excessive host-to-TEE communication latency. We propose AegisGuard, a fine tuning and deployment framework that selectively shields the MS sensitive adapters while offloading the rest to the GPU, balancing security and efficiency. AegisGuard integrates two key components: i) RL-based Sensitivity Measurement (RSM), which injects Gaussian noise during training and applies a lightweight reinforcement learning to rank adapters based on their impact on model stealing; and (ii) Shielded-Adapter Compression (SAC), which structurally prunes the selected adapters to reduce both parameter size and intermediate feature maps, further lowering TEE computation and data transfer costs. Extensive experiments demonstrate that AegisGuard achieves black-box level MS resilience (surrogate accuracy around 39%, matching fully shielded baselines), while reducing end-to-end inference latency by 2–3× and cutting TEE memory usage by 4× compared to state-of-the-art TSDP methods. Ziqi Zhang 0017, Yinggui Wang, Tiantong Wang, Yurong Hao, Tao Wei 0002, Yang Cao 0011, Wei Yang Bryan Lim |
NeurIPS | 3 |
| 2025 | PoiSAFL: Scalable Poisoning Attack Framework to Byzantine-resilient Semi-asynchronous Federated Learning
Xiaoyi Pang, Zhibo Wang 0001, Jiahui Hu 0001, Yinggui Wang, Lei Wang 0251, Tao Wei 0002, Kui Ren 0001, Chun Chen 0001 |
USENIX Security Symposium | 5 |
| 2025 | StegGuard: Secrets Encoder and Decoder Act as Fingerprint of Self-Supervised Pretrained ModelabstractIn this work, we propose StegGuard, a novel fingerprinting mechanism to verify the ownership of a suspect pretrained model using steganography, where the pre-trained model is obtained via self-supervised learning. A critical perspective in StegGuard is that the unique characteristic of the transformation from an image to an embedding, conducted by the pre-trained model, can be equivalently captured by how an encoder embeds secrets into images and how a decoder extracts them from the embeddings with tolerable error. While each independently trained pre-trained model has a distinct transformation, a piracy model exhibits a transformation similar to that of the victim. Based on these observations, StegGuard learns a pair of secrets encoder and decoder as the fingerprint of the victim model. Additionally, a frequency-domain channel attention embedding block is introduced into the encoder to adaptively embed secrets into suitable frequency bands. During verification, if the secrets embedded into the query images can be extracted with an acceptable error from the embeddings of the query images, the suspect model is determined to be piracy; otherwise, it is deemed independent. Extensive experiments demonstrate that with as few as 100 query images, StegGuard achieves high piracy detection accuracy and robustness against model stealing attacks including model extraction, fine-tuning, pruning, embedding noising and shuffle. Compared to existing methods, StegGuard consistently achieves lower p-values for piracy models (as low as 1e-14) and higher p-values for independent models (up to 0.99), confirming its effectiveness and reliability. Xingdong Ren, Hanzhou Wu, Yinggui Wang, Guangling Sun |
IEEE Internet Things J. | 3 |
| 2024 | TaiChi: Improving the Robustness of NLP Models by Seeking Common Ground While Reserving DifferencesabstractRecent studies have shown that Pre-trained Language Models (PLMs) are vulnerable to adversarial examples, crafted by introducing human-imperceptible perturbations to clean examples to deceive the models. This vulnerability stems from the divergence in the data distributions of clean and adversarial examples. Therefore, addressing this issue involves teaching the model to diminish the differences between the two types of samples and to focus more on their similarities. To this end, we propose a novel approach named TaiChi that employs a Siamese network architecture. Specifically, it consists of two sub-networks sharing the same structure but trained on clean and adversarial samples, respectively, and uses a contrastive learning strategy to encourage the generation of similar language representations for both kinds of samples. Furthermore, it utilizes the Kullback-Leibler (KL) divergence loss to enhance the consistency in the predictive behavior of the two sub-networks. Extensive experiments across three widely used datasets demonstrate that TaiChi achieves superior trade-offs between robustness to adversarial attacks at token and character levels and accuracy on clean examples compared to previous defense methods. Our code and data are publicly available at https://github.com/sai4july/TaiChi. Chengyu Wang 0001, Yanhao Wang 0001, Cen Chen 0001, Yinggui Wang |
LREC/COLING | 5 |
| 2024 | A Fast, Performant, Secure Distributed Training Framework For LLMabstractThe distributed (federated) LLM is an important method for co-training the domain-specific LLM using siloed data. However, maliciously stealing model parameters and data from the server or client side has become an urgent problem to be solved. In this paper, we propose a secure distributed LLM based on model slicing. In this case, we deploy the Trusted Execution Environment (TEE) on both the client and server side, and put the fine-tuned structure (LoRA or embedding of P-tuning v2) into the TEE. Then, secure communication is executed in the TEE and general environments through lightweight encryption. In order to further reduce the equipment cost as well as increase the model performance and accuracy, we propose a split fine-tuning scheme. In particular, we split the LLM by layers and place the latter layers in a server-side TEE (the client does not need a TEE). We then combine the proposed Sparsification Parameter Fine-tuning (SPF) with the LoRA part to improve the accuracy of the downstream task. Numerous experiments have shown that our method guarantees accuracy while maintaining security. Wei Huang 0039, Yinggui Wang, Anda Cheng, Aihui Zhou, Chaofan Yu, Lei Wang 0251 |
ICASSP | 2 |
| 2024 | Enhanced Face Recognition using Intra-class Incoherence ConstraintabstractThe current face recognition (FR) algorithms has achieved a high level of accuracy, making further improvements increasingly challenging. While existing FR algorithms primarily focus on optimizing margins and loss functions, limited attention has been given to exploring the feature representation space. Therefore, this paper endeavors to improve FR performance in the view of feature representation space. Firstly, we consider two FR models that exhibit distinct performance discrepancies, where one model exhibits superior recognition accuracy compared to the other. We implement orthogonal decomposition on the features from the superior model along those from the inferior model and obtain two sub-features. Surprisingly, we find the sub-feature perpendicular to the inferior still possesses a certain level of face distinguishability. We adjust the modulus of the sub-features and recombine them through vector addition. Experiments demonstrate this recombination is likely to contribute to an improved facial feature representation, even better than features from the original superior model. Motivated by this discovery, we further consider how to improve FR accuracy when there is only one FR model available. Inspired by knowledge distillation, we incorporate the intra-class incoherence constraint (IIC) to solve the problem. Experiments on various FR benchmarks show the existing state-of-the-art method with IIC can be further improved, highlighting its potential to further enhance FR performance. Yuanqing Huang 0002, Yinggui Wang, Le Yang 0001, Lei Wang 0251 |
ICLR | 2 |
| 2024 | UPFL: Unsupervised Personalized Federated Learning towards New ClientsabstractPersonalized federated learning (pFL) has gained significant attention as a promising approach to address the challenge of data heterogeneity. In this paper, we address a relatively unexplored problem in federated learning. When a federated model has been trained and deployed, and an unla-beled new client joins, providing a personalized model for the new client becomes a highly challenging task. To address this challenge, we extend the adaptive risk minimization technique into the unsupervised pFL setting and propose our method, FedTTA. We further improve FedTTA with two simple yet highly effective optimization strategies: enhancing the training of the adaptation model with proxy regularization and early-stopping the adaptation through entropy. Moreover, we propose a knowledge distillation loss specifically designed for FedTTA to address the device heterogeneity. Extensive experiments on five datasets against eleven baselines demonstrate the effectiveness of our proposed FedTTA and its variants. The code is available at: https://github.com/anonymous-federated-learning/code. Tiandi Ye, Cen Chen 0001, Yinggui Wang, Xiang Li 0067, Ming Gao 0001 |
SDM | 3 |
| 2024 | Federated Submodular Maximization With Differential PrivacyabstractSubmodular maximization is a fundamental problem in many Internet of Things applications, such as sensor placement, resource allocation, and mobile crowdsourcing. Despite being intensively studied over the last two decades, the problem of submodular maximization has not yet been considered in an emerging federated computation setting. In this article, we first comprehensively study federated submodular maximization, where a set of clients aims to cooperate in finding a set of items to maximize a monotone submodular function under the orchestration of a central server while providing strong privacy guarantees for their sensitive data. We consider the problem in a client-level differential privacy (DP) setting: the server is not necessarily trusted and the clients should perturb their results locally before sending them to the server. Specifically, we propose a novel approximation algorithm for federated submodular maximization by incorporating client-level DP mechanisms and decomposed function evaluations into the greedy algorithm, along with two heuristics to further reduce the privacy budget, computational cost, and communication overhead. Finally, we perform extensive experiments to demonstrate the effectiveness and efficiency of our proposed algorithms. Yanhao Wang 0001, Cen Chen 0001, Yinggui Wang |
IEEE Internet Things J. | 4 |
| 2024 | Fast-Convergent Wireless Federated Learning: A Voting-Based TopK Model Compression ApproachabstractFederated learning (FL) has been extensively exploited in the training of machine learning models to preserve data privacy. In particular, wireless FL enables multiple clients to collaboratively train models by sharing model updates via wireless communication without exposing raw data. The state-of-the-art wireless FL advocates efficient aggregation of model updates from multiple clients by over-the-air computing. However, a significant deficiency of over-the-air aggregation lies in the infeasibility of TopK model compression given that top model updates cannot be aggregated directly before they are aligned according to their indices. In view of the fact that TopK can greatly accelerate FL, we design a novel wireless FL with voting based TopK algorithm, namely WFL-VTopK, so that top model updates can be aggregated by over-the-air computing directly. Specifically, there are two phases in WFL-VTopK. In Phase 1, clients vote their top model updates, based on which global top model updates can be efficiently identified. In Phase 2, clients formally upload global top model updates so that they can be directly aggregated by over-the-air computing. Furthermore, the convergence of WFL-VTopK is theoretically guaranteed under non-convex loss. Based on the convergence of WFL-VTopK, we optimize model utility subjecting to training time and energy constraints. To validate the superiority of WFL-VTopK, we extensively conduct experiments with real datasets under wireless communication. The experimental results demonstrate that WFL-VTopK can effectively aggregate models by only communicating 1%-2% top models updates, and hence significantly outperforms the state-of-the-art baselines. By significantly reducing the wireless communication traffic, our work paves the road to train large models in wireless FL. Xiaoxin Su 0001, Yipeng Zhou, Laizhong Cui, Quan Z. Sheng, Yinggui Wang, Song Guo 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2024 | BapFL: You can Backdoor Personalized Federated LearningabstractIn federated learning (FL), malicious clients could manipulate the predictions of the trained model through backdoor attacks, posing a significant threat to the security of FL systems. Existing research primarily focuses on backdoor attacks and defenses within the generic federated learning scenario, where all clients collaborate to train a single global model. A recent study conducted by Qin et al. [ 24 ] marks the initial exploration of backdoor attacks within the personalized federated learning (pFL) scenario, where each client constructs a personalized model based on its local data. Notably, the study demonstrates that pFL methods with parameter decoupling can significantly enhance robustness against backdoor attacks. However, in this article, we whistleblow that pFL methods with parameter decoupling are still vulnerable to backdoor attacks. The resistance of pFL methods with parameter decoupling is attributed to the heterogeneous classifiers between malicious clients and benign counterparts. We analyze two direct causes of the heterogeneous classifiers: (1) data heterogeneity inherently exists among clients and (2) poisoning by malicious clients further exacerbates the data heterogeneity. To address these issues, we propose a two-pronged attack method, BapFL, which comprises two simple yet effective strategies: (1) poisoning only the feature encoder while keeping the classifier fixed and (2) diversifying the classifier through noise introduction to simulate that of the benign clients. Extensive experiments on three benchmark datasets under varying conditions demonstrate the effectiveness of our proposed attack. Additionally, we evaluate the effectiveness of six widely used defense methods and find that BapFL still poses a significant threat even in the presence of the best defense, Multi-Krum. We hope to inspire further research on attack and defense strategies in pFL scenarios. The code is available at: https://github.com/BapFL/code Tiandi Ye, Cen Chen 0001, Yinggui Wang, Xiang Li 0067, Ming Gao 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Privacy-Preserving End-to-End Spoken Language UnderstandingabstractSpoken language understanding (SLU), one of the key enabling technologies for human-computer interaction in IoT devices, provides an easy-to-use user interface. Human speech can contain a lot of user-sensitive information, such as gender, identity, and sensitive content. New types of security and privacy breaches have thus emerged. Users do not want to expose their personal sensitive information to malicious attacks by untrusted third parties. Thus, the SLU system needs to ensure that a potential malicious attacker cannot deduce the sensitive attributes of the users, while it should avoid greatly compromising the SLU accuracy. To address the above challenge, this paper proposes a novel SLU multi-task privacy-preserving model to prevent both the speech recognition (ASR) and identity recognition (IR) attacks. The model uses the hidden layer separation technique so that SLU information is distributed only in a specific portion of the hidden layer, and the other two types of information are removed to obtain a privacy-secure hidden layer. In order to achieve good balance between efficiency and privacy, we introduce a new mechanism of model pre-training, namely joint adversarial training, to further enhance the user privacy. Experiments over two SLU datasets show that the proposed method can reduce the accuracy of both the ASR and IR attacks close to that of a random guess, while leaving the SLU performance largely unaffected. Yinggui Wang, Wei Huang 0039, Le Yang 0001 |
IJCAI | 1 |
| 2022 | Privacy-Preserving Face Recognition in the Frequency DomainabstractSome applications may require performing face recognition (FR) on third-party servers, which could be accessed by attackers with malicious intents to compromise the privacy of users’ face information. This paper advocates a practical privacy-preserving FR scheme without key management realized in the frequency domain. The new scheme first collects the components of the same frequency from different blocks of a face image to form component channels. Only part of the channels are retained and fed into the analysis network that performs an interpretable privacy-accuracy trade-off analysis to identify channels important for face image visualization but not crucial for maintaining high FR accuracy. For this purpose, the loss function of the analysis network consists of the empirical FR error loss and a face visualization penalty term, and the network is trained in an end-to-end manner. We find that with the developed analysis network, more than 94% of the image energy can be dropped while the face recognition accuracy stays almost undegraded. In order to further protect the remaining frequency components, we propose a fast masking method. Effectiveness of the new scheme in removing the visual information of face images while maintaining their distinguishability is validated over several large face datasets. Results show that the proposed scheme achieves a recognition performance and inference time comparable to ArcFace operating on original face images directly. Yinggui Wang, Le Yang 0001 |
AAAI | 1 |
| 2015 | TOA-based joint synchronization and source localization with random errors in sensor positions and sensor clock biases
Yinggui Wang, Le Yang 0001, Yanbo Xue |
Ad Hoc Networks | 1 |
| 2014 | Compressive detection of stochastic signals with the measurement matrix not necessarily orthonormal
Yinggui Wang, Le Yang 0001, Zheng Liu 0012, Fucheng Guo 0001, Wenli Jiang |
FUSION | 1 |
| 2014 | Block-sparse signal recovery with synthesized multitask compressive sensingabstractThe paper considers the problem of reconstructing blocks-sparse signals. A new algorithm, called synthesized multitask compressive sensing (SMCS), is proposed. In contrast to existing methods that rely on the availability of the sparsity structure information, the SMCS algorithm resorts to the multitask compressive sensing (MCS) technique for signal recovery. The SMCS algorithm synthesizes new compressive sensing (CS) tasks via circular-shifting operations and utilizes the minimum description length (MDL) principle to determine the proper set of the synthesized CS tasks for signal reconstruction. An outstanding advantage of SMCS is that it can achieve good signal reconstruction performance without using prior information on the block-sparsity structure. Simulations corroborate the theoretical developments. Yinggui Wang, Zheng Liu 0012, Wenli Jiang, Le Yang 0001 |
ICASSP | 1 |
| 2013 | An MDL-based multi-task classification and reconstruction algorithm
Yinggui Wang, Zheng Liu 0012, Dao-Wang Feng, Wenli Jiang |
FUSION | 1 |