VLDB 2026 Research / reviewers in the wild / expert
Rui Hu 0005
dblp:08/3712-5
· DBLP profile ↗
18ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-3317-1765ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XMark: Reliable Multi-Bit Watermarking for LLM-Generated TextsabstractMulti-bit watermarking has emerged as a promising solution for embedding imperceptible binary messages into Large Language Model (LLM)-generated text, enabling reliable attribution and tracing of malicious usage of LLMs.Despite recent progress, existing methods still face key limitations: some become computationally infeasible for large messages, while others suffer from a poor trade-off between text quality and decoding accuracy.Moreover, the decoding accuracy of existing methods drops significantly when the number of tokens in the generated text is limited, a condition that frequently arises in practical usage.To address these challenges, we propose XMARK, a novel method for encoding and decoding binary messages in LLM-generated texts.The unique design of XMARK's encoder produces a less distorted logit distribution for watermarked token generation, preserving text quality, and also enables its tailored decoder to reliably recover the encoded message with limited tokens.Extensive experiments across diverse downstream tasks show that XMARK significantly improves decoding accuracy while preserving the quality of watermarked text, outperforming prior methods. Rui Hu 0005, Olivera Kotevska, Zikai Zhang 0003 |
ACL (1) | 2 |
| 2026 | Optimal Client Sampling in Federated Learning With Client-Level Heterogeneous Differential Privacy
Rui Hu 0005, Olivera Kotevska |
IEEE Internet Things J. | 2 |
| 2025 | Detecting Backdoor Attacks in Federated Learning via Direction Alignment InspectionabstractThe distributed nature of training makes Federated Learning (FL) vulnerable to backdoor attacks, where malicious model updates aim to compromise the global model’s performance on specific tasks. Existing defense methods show limited efficacy as they overlook the inconsistency between benign and malicious model updates regarding both general and fine-grained directions. To fill this gap, we introduce AlignIns, a novel defense method designed to safeguard FL systems against backdoor attacks. AlignIns looks into the direction of each model update through a direction alignment inspection process. Specifically, it examines the alignment of model updates with the overall update direction and analyzes the distribution of the signs of their significant parameters, comparing them with the principle sign across all model updates. Model updates that exhibit an unusual degree of alignment are considered malicious and thus be filtered out. We provide the theoretical analysis of the robustness of AlignIns and its propagation error in FL. Our empirical results on both independent and identically distributed (IID) and non-IID datasets demonstrate that AlignIns achieves higher robustness compared to the state-of-the-art defense methods. The code is available at https://github.com/JiiahaoXU/AlignIns. Zikai Zhang 0003, Rui Hu 0005 |
CVPR | 3 |
| 2025 | On the Out-of-Distribution Backdoor Attack for Federated LearningabstractTraditional backdoor attacks in federated learning (FL) operate within constrained attack scenarios, as they depend on visible triggers and require physical modifications to the target object, which limits their practicality. To address this limitation, we introduce a novel backdoor attack prototype for FL called the out-of-distribution (OOD) backdoor attack (OBA), which uses OOD data as both poisoned samples and triggers simultaneously. Our approach significantly broadens the scope of backdoor attack scenarios in FL. To improve the stealthiness of OBA, we propose SoDa, which regularizes both the magnitude and direction of malicious local models during local training, aligning them closely with their benign versions to evade detection. Empirical results demonstrate that OBA effectively circumvents state-of-the-art defenses while maintaining high accuracy on the main task. Zikai Zhang 0003, Rui Hu 0005 |
MobiHoc | 3 |
| 2025 | FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language ModelsabstractLarge Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications. Yan Gao 0016, Massimo Roberto Scamarcia, Javier Fernández-Marqués, Mohammad Naseri, Chong Shen Ng, Dimitris Stripelis, Zexi Li 0001, Tao Shen 0002, Jiamu Bai, Daoyuan Chen, Zikai Zhang 0003, Rui Hu 0005, Inseo Song, Kangyoon Lee, Hong Jia, Ting Dang, Zheyuan Liu 0002, Daniel J. Beutel, Lingjuan Lyu, Nicholas D. Lane |
NeurIPS | 12 |
| 2025 | Achieving Byzantine-Resilient Federated Learning via Layer-Adaptive Sparsified Model AggregationabstractFederated Learning (FL) enables multiple clients to train a model collaboratively without sharing their local data. Yet the FL system is vulnerable to well-designed Byzan-tine attacks, which aim to disrupt the model training pro-cess by uploading malicious model updates. Existing ro-bust aggregation rule-based defense methods overlook the diversity of magnitude and direction across different lay-ers of the model updates, resulting in limited robustness performance, particularly in non-IID settings. To address these challenges, we propose the Layer-Adaptive Sparsified Model Aggregation (LASA) approach, which combines pre-aggregation sparsification with layer-wise adaptive aggre-gation to improve robustness. Specifically, LASA includes a pre-aggregation sparsification module that sparsifies up-dates from each client before aggregation, reducing the im-pact of malicious parameters and minimizing the interfer-ence from less important parameters for the subsequent fil-tering process. Based on sparsified updates, a layer-wise adaptive filter then adaptively selects benign layers using both magnitude and direction metrics across all clients for aggregation. We provide a detailed theoretical robustness analysis of LASA and a resilience analysis of the FL inte-grated with LASA. Extensive experiments are conducted on various IID and non-IID datasets. The numerical results demonstrate the effectiveness of LASA. Code is available at https://github.com/JiiahaoXU/LASA. Zikai Zhang 0003, Rui Hu 0005 |
WACV | 3 |
| 2025 | Identify Backdoored Model in Federated Learning via Individual UnlearningabstractBackdoor attacks present a significant threat to the robustness of Federated Learning (FL) due to their stealth and effectiveness. They maintain both the main task of the FL system and the backdoor task simultaneously, causing malicious models to appear statistically similar to benign ones, which enables them to evade detection by existing defense methods. We find that malicious parameters in backdoored models are inactive on the main task, resulting in a significantly large empirical loss during the machine unlearning process on clean inputs. Inspired by this, we propose MASA, a method that utilizes individual unlearning on local models to identify malicious models in FL. To improve the performance of MASA in challenging non-independent and identically distributed (non-IID) settings, we design pre-unlearning model fusion that integrates local models with knowledge learned from other datasets to mitigate the divergence in their unlearning behaviors caused by the non-IID data distributions of clients. Additionally, we propose a new anomaly detection metric with minimal hyperparameters to filter out malicious models efficiently. Extensive experiments on IID and non-IID datasets across six different attacks validate the effectiveness of MASA. To the best of our knowledge, this is the first work to leverage machine unlearning to identify malicious models in FL. Code is available at https://github.com/JiiahaoXU/MASA. Zikai Zhang 0003, Rui Hu 0005 |
WACV | 3 |
| 2025 | Fed-HeLLo: Efficient Federated Foundation Model Fine-Tuning With Heterogeneous LoRA AllocationabstractFederated learning (FL) has recently been used to collaboratively fine-tune foundation models (FMs) across multiple clients. Notably, federated low-rank adaptation (LoRA)-based fine-tuning methods have recently gained attention, which allows clients to fine-tune FMs with a small portion of trainable parameters locally. However, most existing methods do not account for the heterogeneous resources of clients or lack an effective local training strategy to maximize global fine-tuning performance under limited resources. In this work, we propose federated LoRA-based fine-tuning framework with heterogeneous LoRA allocation (Fed-HeLLo), a novel federated LoRA-based fine-tuning framework that enables clients to collaboratively fine-tune an FM with different local trainable LoRA layers. To ensure its effectiveness, we develop several heterogeneous LoRA allocation (HLA) strategies that adaptively allocate local trainable LoRA layers based on clients' resource capabilities and the layer importance. Specifically, based on the dynamic layer importance, we design a Fisher information matrix score-based HLA (FIM-HLA) that leverages dynamic gradient norm information. To better stabilize the training process, we consider the intrinsic importance of LoRA layers and design a geometrically defined HLA (GD-HLA) strategy. It shapes the collective distribution of trainable LoRA layers into specific geometric patterns, such as triangle, inverted triangle, bottleneck, and uniform. Moreover, we extend GD-HLA into a randomized version, named randomized GD-HLA (RGD-HLA), for enhanced model accuracy with randomness. By codesigning the proposed HLA strategies, we incorporate both the dynamic and intrinsic layer importance into the design of our HLA strategy. To thoroughly evaluate our approach, we simulate various complex federated LoRA-based fine-tuning settings using five datasets and three levels of data distributions ranging from independent identically distributed (i.i.d.) to extreme non-i.i.d. The experimental results demonstrate the effectiveness and efficiency of Fed-HeLLo with the proposed HLA strategies. The code is available at https://github.com/ TNI-playground/Fed_HeLLo. Zikai Zhang 0003, Rui Hu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Federated Learning With Sparsified Model Perturbation: Improving Accuracy Under Client-Level Differential PrivacyabstractFederated learning (FL) that enables edge devices to collaboratively learn a shared model while keeping their training data locally has received great attention recently and can protect privacy in comparison with the traditional centralized learning paradigm. However, sensitive information about the training data can still be inferred from model parameters shared in FL. Differential privacy (DP) is the state-of-the-art technique to defend against those attacks. The key challenge to achieving DP in FL lies in the adverse impact of DP noise on model accuracy, particularly for deep learning models with large numbers of parameters. This paper develops a novel differentially-private FL scheme named Fed-SMP that provides a client-level DP guarantee while maintaining high model accuracy. To mitigate the impact of privacy protection on model accuracy, Fed-SMP leverages a new technique called Sparsified Model Perturbation (SMP) where local models are sparsified first before being perturbed by Gaussian noise. We provide a tight end-to-end privacy analysis for Fed-SMP using Rényi DP and prove the convergence of Fed-SMP with both unbiased and biased sparsifications. Extensive experiments on real-world datasets are conducted to demonstrate the effectiveness of Fed-SMP in improving model accuracy with the same DP guarantee and saving communication cost simultaneously. Rui Hu 0005, Yuanxiong Guo, Yanmin Gong 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Energy-Efficient Distributed Machine Learning at Wireless Edge with Device-to-Device CommunicationabstractThis paper considers a federated edge learning (FEL) system where a base station (BS) coordinates a set of edge devices to train a shared machine learning model collaboratively. One of the fundamental issues in such systems is maintaining the learning performance with the limited and heterogeneous resource capabilities of edge devices. Our goal is to improve the energy efficiency of edge devices in FEL by mitigating the temporal and spatial heterogeneity of their energy resources. Specifically, to balance the heterogeneous energy levels among edge devices, energy-hungry devices can offload their data to nearby devices that have sufficient energy via device-to-device (D2D) communication links at low transmission overheads. Be-sides, to mitigate the impact of the time-varying energy level of a device, data collected by edge devices can be queued to be processed when sufficient energy is available. To compute the optimal offloading and queuing strategies, we propose an online control algorithm based on Lyapunov optimization to determine the amount of data to be offloaded, queued, and processed at each time slot. Our simulation results on the real-world dataset demonstrate that our approach achieves a better overall energy efficiency than baselines. Rui Hu 0005, Yuanxiong Guo, Yanmin Gong 0001 |
ICC | 1 |
| 2022 | Hybrid Local SGD for Federated Learning with Heterogeneous Communications
Yuanxiong Guo, Ying Sun 0003, Rui Hu 0005, Yanmin Gong 0001 |
ICLR | 3 |
| 2021 | Federated Learning with Sparsification-Amplified Privacy and Adaptive OptimizationabstractFederated learning (FL) enables distributed agents to collaboratively learn a centralized model without sharing their raw data with each other. However, data locality does not provide sufficient privacy protection, and it is desirable to facilitate FL with rigorous differential privacy (DP) guarantee. Existing DP mechanisms would introduce random noise with magnitude proportional to the model size, which can be quite large in deep neural networks. In this paper, we propose a new FL framework with sparsification-amplified privacy. Our approach integrates random sparsification with gradient perturbation on each agent to amplify privacy guarantee. Since sparsification would increase the number of communication rounds required to achieve a certain target accuracy, which is unfavorable for DP guarantee, we further introduce acceleration techniques to help reduce the privacy cost. We rigorously analyze the convergence of our approach and utilize Renyi DP to tightly account the end-to-end DP guarantee. Extensive experiments on benchmark datasets validate that our approach outperforms previous differentially-private FL approaches in both privacy guarantee and communication efficiency. Rui Hu 0005, Yanmin Gong 0001, Yuanxiong Guo |
IJCAI | 1 |
| 2020 | Certified Robustness of Graph Classification against Topology Attack with Randomized SmoothingabstractGraph classification has practical applications in diverse fields. Recent studies show that graph-based machine learning models are especially vulnerable to adversarial perturbations due to the non i.i. d nature of graph data. By adding or deleting a small number of edges in the graph, adversaries could greatly change the graph label predicted by a graph classification model. In this work, we propose to build a smoothed graph classification model with certified robustness guarantee. We have proven that the resulting graph classification model would output the same prediction for a graph under l0bounded adversarial perturbation. We also evaluate the effectiveness of our approach under graph convolutional network (GCN) based multi-class graph classification model. Zhidong Gao, Rui Hu 0005, Yanmin Gong 0001 |
GLOBECOM | 2 |
| 2020 | Trading Data For Learning: Incentive Mechanism For On-Device Federated LearningabstractFederated Learning rests on the notion of training a global model distributedly on various devices. Under this setting, users' devices perform computations on their own data and then share the results with the cloud server to update the global model. A fundamental issue in such systems is to effectively incentivize user participation. The users suffer from privacy leakage of their local data during the federated model training process. Without well-designed incentives, self-interested users will be unwilling to participate in federated learning tasks and contribute their private data. To bridge this gap, in this paper, we adopt the game theory to design an effective incentive mechanism, which selects users that are most likely to provide reliable data and compensates for their costs of privacy leakage. We formulate our problem as a two-stage Stackelberg game and solve the game's equilibrium. Effectiveness of the proposed mechanism is demonstrated by extensive simulations. Rui Hu 0005, Yanmin Gong 0001 |
GLOBECOM | 1 |
| 2020 | Privacy-Preserving Personalized Federated LearningabstractTo provide intelligent and personalized services on smart devices, machine learning techniques have been widely used to learn from data, identify patterns, and make automated decisions. Machine learning processes typically require a large amount of representative data that are often collected through crowdsourcing from end users. However, user data could be sensitive in nature, and learning machine learning models on these data may expose sensitive information of users, violating their privacy. Moreover, to meet the increasing demand of personalized services, these learned models should capture their individual characteristics. This paper proposes a privacy-preserving approach for learning effective personalized models on distributed user data while guaranteeing the differential privacy of user data. Practical issues in a distributed learning system such as user heterogeneity are considered in the proposed approach. Moreover, the convergence property and privacy guarantee of the proposed approach are rigorously analyzed. Experiments on realistic mobile sensing data demonstrate that the proposed approach is robust to high user heterogeneity and offer a trade-off between accuracy and privacy. Rui Hu 0005, Yuanxiong Guo, Hongning Li, Qingqi Pei, Yanmin Gong 0001 |
ICC | 1 |
| 2020 | Personalized Federated Learning With Differential PrivacyabstractTo provide intelligent and personalized services on smart devices, machine learning techniques have been widely used to learn from data, identify patterns, and make automated decisions. Machine learning processes typically require a large amount of representative data that are often collected through crowdsourcing from end users. However, user data could be sensitive in nature, and training machine learning models on these data may expose sensitive information of users, violating their privacy. Moreover, to meet the increasing demand of personalized services, these learned models should capture their individual characteristics. This article proposes a privacy-preserving approach for learning effective personalized models on distributed user data while guaranteeing the differential privacy of user data. Practical issues in a distributed learning system such as user heterogeneity are considered in the proposed approach. In addition, the convergence property and privacy guarantee of the proposed approach are rigorously analyzed. The experimental results on realistic mobile sensing data demonstrate that the proposed approach is robust to user heterogeneity and offers a good tradeoff between accuracy and privacy. Rui Hu 0005, Yuanxiong Guo, Hongning Li, Qingqi Pei, Yanmin Gong 0001 |
IEEE Internet Things J. | 1 |
| 2020 | DP-ADMM: ADMM-Based Distributed Learning With Differential PrivacyabstractAlternating direction method of multipliers (ADMM) is a widely used tool for machine learning in distributed settings where a machine learning model is trained over distributed data sources through an interactive process of local computation and message passing. Such an iterative process could cause privacy concerns of data owners. The goal of this paper is to provide differential privacy for ADMM-based distributed machine learning. Prior approaches on differentially private ADMM exhibit low utility under high privacy guarantee and assume the objective functions of the learning problems to be smooth and strongly convex. To address these concerns, we propose a novel differentially private ADMM-based distributed learning algorithm called DP-ADMM, which combines an approximate augmented Lagrangian function with time-varying Gaussian noise addition in the iterative process to achieve higher utility for general objective functions under the same differential privacy guarantee. We also apply the moments accountant method to analyze the end-to-end privacy loss. The theoretical analysis shows that the DP-ADMM can be applied to a wider class of distributed learning problems, is provably convergent, and offers an explicit utility-privacy tradeoff. To our knowledge, this is the first paper to provide explicit convergence and utility properties for differentially private ADMM-based distributed learning algorithms. The evaluation results demonstrate that our approach can achieve good convergence and model accuracy under high end-to-end differential privacy guarantee. Zonghao Huang, Rui Hu 0005, Yuanxiong Guo, Eric Chan-Tin, Yanmin Gong 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Targeted Poisoning Attacks on Social Recommender SystemsabstractWith the popularity of online social networks, social recommendations that rely on ones social connections to make personalized recommendations have become possible. This introduces vulnerabilities for an adversarial party to compromise the recommendations for users by utilizing their social connections. In this paper, we propose the targeted poisoning attack on the factorization-based social recommender system in which the attacker aims to promote an item to a group of target users by injecting fake ratings and social connections. We formulate the optimal poisoning attack as a bi-level program and develop an efficient algorithm to find the optimal attacking strategy. We then evaluate the proposed attacking strategy on real-world dataset and demonstrate that the social recommender system is sensitive to the targeted poisoning attack. We find that users in the social recommender system can be attacked even if they do not have direct social connections with the attacker. Rui Hu 0005, Yuanxiong Guo, Miao Pan, Yanmin Gong 0001 |
GLOBECOM | 1 |