EDBT 2026 Demo / reviewers in the wild / expert
Huiwen Wu
dblp:90/3516
· DBLP profile ↗
14ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synergizing Multigrid Algorithms with Vision Transformer: A Novel Approach to Enhance the Seismic Foundation ModelabstractDue to the rapid advancement and homogenization of Artificial Intelligence (AI) technology development, transformer-based foundation models have revolutionized scientific applications, such as drug discovery, materials research, and astronomy. However, seismic data presents unique characteristics that require specialized processing techniques for pretraining foundation models in seismic contexts with high- and low-frequency features playing crucial roles. Existing Vision Transformer (ViT) with sequential image tokenization fails to efficiently and effectively capture both high- and low-frequency seismic information because they ignore the intrinsic structural patterns of seismograms. This work introduces ADATG, a novel adaptive two-grid training strategy with Hilbert encoding, explicitly tailored for seismogram data and leveraging the hierarchical structures inherent in seismic data. Specifically, our approach employs spectrum decomposition to separate high- and low-frequency components, and hierarchical Hilbert encoding to represent the data effectively. Moreover, inspired by the frequency principle, we propose an adaptive training strategy that initially emphasizes coarse-level information and then progressively refines the model's focus on fine-level features. Extensive experiments demonstrate the effectiveness and efficiency of our method. This research highlights the importance of data encoding and training strategies informed by the distinct characteristics of high- and low-frequency features in seismic images, ultimately enhancing the pretraining of visual seismic foundation models. Huiwen Wu, Hongbin Ye |
AAAI | 1 |
| 2026 | PDFL: A Privacy-Enhancing and Robust Poisoning Defense Federated Learning SchemeabstractThis paper addresses the security and privacy issues of the global models in Federated Learning by proposing a new approach, called PDFL, which tackles the challenges of poisoning attacks and privacy leakage in FL rounds. PDFL is based on secure multi-party computation and performs privacy-preserving cluster analysis on encrypted data from participants in order to identify malicious poisoning attackers. This approach involves a two-server mechanism and integrates four privacy-preserving protocols based on two-party computation (2PC): SecJudge for normalizing gradients, SecCosine for computing the cosine similarity values among gradients, SecClu for countering poisoning attacks, and SecAgg for secure aggregation by the server. These protocols are designed to achieve low computational costs, preserve client data privacy, and mitigate poisoning attacks from the potentially malicious clients. We provide a theoretical proof that our four sub-protocols and the PDFL scheme are both safe and reliable, demonstrating that PDFL can ensure the privacy and security of the participating data. Additionally, we conduct extensive simulation experiments to evaluate the accuracy, efficiency, computational overhead, and communication overhead associated with the PDFL scheme. Experimental results show the potential of the PDFL scheme in significantly enhancing the ability to identify malicious poisoning attackers in federated learning systems accurately and efficiently, hence making PDFL a promising solution for addressing privacy and security concerns in this domain. Huiwen Wu, Qingming Li, Ziyao Liu, Jun Zhao 0007, Kwok-Yan Lam, Qingkuan Dong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially PrivatelyabstractThe emergence of the large language model (LLM) has shown its superiority in a wide range of disciplines, including language understanding and translation, relational logic reasoning, and even partial differential equations solving. The transformer is the pervasive backbone architecture for the foundation model construction. It is vital to research how to adjust the Transformer architecture to achieve an end-to-end privacy guarantee in LLM fine-tuning. This paper investigates three potential information leaks during a federated fine-tuning procedure for LLM (FedLLM). Based on the potential information leakage, we insert two-stage randomness into FedLLM to provide an end-to-end privacy guarantee solution. The first stage is to train a gradient auto-encoder with a Gaussian random prior based on the statistical information of the gradients generated by local clients. The second stage is fine-tuning the overall LLM with a differential privacy guarantee by adopting appropriate Gaussian noises. We show our proposed method's efficiency and accuracy gains with several foundation models and two popular evaluation benchmarks. Furthermore, we present a comprehensive privacy analysis with Gaussian Differential Privacy (GDP) and Renyi Differential Privacy (RDP). Huiwen Wu, Deyi Zhang, Xiaogang Xu 0002, Jiafei Wu, Zhe Liu 0001 |
AAAI | 1 |
| 2025 | Comparing and Improving Frequency Estimation Perturbation Mechanisms Under Local Differential Privacy
She Sun, Jiafei Wu, Huiwen Wu |
ACISP (3) | 5 |
| 2025 | DPFedSub: A Differentially Private Federated Learning with Randomized Subspace Descend
Huiwen Wu, Chuan Ma 0001, Xueran Li, Deyi Zhang, She Sun |
ACISP (3) | 1 |
| 2025 | Comparing and Improving Perturbation Mechanisms Under Local Differential Privacy
She Sun, Xiaoran Yan, Huiwen Wu |
Inscrypt (3) | 5 |
| 2025 | CG-FedLLM: How to Compress Gradients in Federated Fine-Tuning for Large Language ModelsabstractThe success of current Large-Language Models (LLMs) hinges on extensive training data that are collected and stored centrally, called Centralized Learning (CL). However, such a collection manner poses a privacy threat, and one potential solution is Federated Learning (FL), which transfers gradients, not raw data, among clients. Unlike traditional networks, FL for LLMs incurs significant communication costs due to their tremendous parameters. In this study, we introduce an innovative approach to compress gradients to improve communication efficiency during LLM FL, formulating the new FL pipeline named CG-FedLLM. This approach integrates an encoder on the client side to acquire the compressed gradient features and a decoder on the server side to reconstruct the gradients. We also develop a novel training strategy that comprises Temporal-ensemble Gradient-Aware Pre-training (TGAP) to identify characteristic gradients of the target model and Federated AutoEncoder-Involved Fine-tuning (FAF) to compress gradients adaptively. Extensive experiments confirm that our approach reduces communication costs and improves performance (e.g., average 3 points increment compared with traditional CL- and FL-based fine-tuning with several foundation models on well-recognized benchmarks, MMLU and C-Eval). This is because our encoder-decoder, trained via TGAP and FAF, can filter gradients while selectively preserving critical features. Furthermore, we present a series of experimental analyses that focus on the communication efficiency, accuracy, and generalization ability within this privacy-centric framework, providing insights into the development of more efficient and private LLMs fine-tuning. Huiwen Wu, Xiaogang Xu 0002, Deyi Zhang, Jiafei Wu, Zhe Liu 0001 |
ECAI | 1 |
| 2025 | VFGCN: A Vertical Federated Learning Framework With Privacy Preserving for Graph Convolutional NetworkabstractDue to the robust representational capabilities of graph data, employing graph neural networks for its processing has demonstrated superior performance over conventional deep learning algorithms. Graph data encompasses abundant features and structural information; however, its large-scale collection is often challenging in practice. This difficulty arises because data predominantly exists in isolated compartments, making it arduous to harmonize information across various organizations or to enable multiple organizations to collaborate effectively while safeguarding local data privacy. In light of an extreme data distribution scenario, where each client possesses distinct nodes with partially overlapping segments yet divergent data features, we introduce a dual-cloud server architecture. This framework encompasses the design of four secure subprotocols: ReEnc (secure re-encryption), SecPSI (secure outsourcing of PSI), SecWeight (secure weight calculation), and SecAgg (secure aggregation). Together, these components facilitate a vertical federated learning framework for graph convolutional networks, ensuring privacy preservation. We provide a security proof for the entire system and extensive evaluation on three benchmark datasets (Cora, Citeseer, and Pubmed) illustrates that our Vertical Federated Graph Convolutional Network (VFGCN) surpasses existing privacy-preserving methodologies. Qingming Li, Ximeng Liu, Xiaoran Yan, Qingkuan Dong, Huiwen Wu, Xiangjie Kong 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Fast and Robust Differential Private Stochastic Gradient Descent with Preconditioner
Huiwen Wu |
PKAW | 1 |
| 2024 | A Variational Approach to Personalized Federated Learning and Its Improvement
Huiwen Wu, Shuo Zhang 0004 |
PKAW | 1 |
| 2024 | FedScale: A Federated Unlearning Method Mimicking Human Forgetting Processes
Wenshu Huang, Huiwen Wu, Liming Fang 0001, Lu Zhou 0002 |
WASA (1) | 2 |
| 2022 | Vertically Federated Graph Neural Network for Privacy-Preserving Node ClassificationabstractRecently, Graph Neural Network (GNN) has achieved remarkable progresses in various real-world tasks on graph data, consisting of node features and the adjacent information between different nodes. High-performance GNN models always depend on both rich features and complete edge information in graph. However, such information could possibly be isolated by different data holders in practice, which is the so-called data isolation problem. To solve this problem, in this paper, we propose VFGNN, a federated GNN learning paradigm for privacy-preserving node classification task under data vertically partitioned setting, which can be generalized to existing GNN models. Specifically, we split the computation graph into two parts. We leave the private data (i.e., features, edges, and labels) related computations on data holders, and delegate the rest of computations to a semi-honest server. We also propose to apply differential privacy to prevent potential information leakage from the server. We conduct experiments on three benchmarks and the results demonstrate the effectiveness of VFGNN. Chaochao Chen 0001, Jun Zhou 0011, Longfei Zheng, Huiwen Wu, Lingjuan Lyu, Jia Wu 0001, Bingzhe Wu, Li Wang 0056 |
IJCAI | 4 |
| 2022 | Differential Private Knowledge Transfer for Privacy-Preserving Cross-Domain RecommendationabstractCross Domain Recommendation (CDR) has been popularly studied to alleviate the cold-start and data sparsity problem commonly existed in recommender systems. CDR models can improve the recommendation performance of a target domain by leveraging the data of other source domains. However, most existing CDR models assume information can directly ‘transfer across the bridge’, ignoring the privacy issues. To solve this problem, we propose a novel two stage based privacy-preserving CDR framework (PriCDR). In the first stage, we propose two methods, i.e., Johnson-Lindenstrauss Transform (JLT) and Sparse-aware JLT (SJLT), to publish the rating matrix of the source domain using Differential Privacy (DP). We theoretically analyze the privacy and utility of our proposed DP based rating publishing methods. In the second stage, we propose a novel heterogeneous CDR model (HeteroCDR), which uses deep auto-encoder and deep neural network to model the published source rating matrix and target rating matrix respectively. To this end, PriCDR can not only protect the data privacy of the source domain, but also alleviate the data sparsity of the source domain. We conduct experiments on two benchmark datasets and the results demonstrate the effectiveness of PriCDR and HeteroCDR. Chaochao Chen 0001, Huiwen Wu, Jiajie Su, Lingjuan Lyu, Li Wang 0056 |
WWW | 2 |
| 2002 | Evaluating the Utility of Statistical Phrases and Latent Semantic Indexing for Text ClassificationabstractThe term-based vector space model is a prominent technique for retrieving textual information. In this paper we examine the usefulness of phrases as terms in vector-based document classification. We focus on statistical techniques to extract both adjacent and window phrases from documents. We discover that the positive effect of adding phrase terms is very limited, if we have already achieved good performance using single-word terms, even when SVD/LSI is used as the dimensionality reduction method. Huiwen Wu, Dimitrios Gunopulos |
ICDM | 1 |