Zibo Wang 0001

dblp:133/7314-1 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-9554-6695ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LAFA: Agentic LLM-Driven Federated Analytics Over Decentralized Data Sources
abstract
Large Language Models (LLMs) have shown great promise in automating data analytics tasks by interpreting natural language queries and generating multi-operation execution plans. However, existing LLM-agent-based analytics frameworks operate under the assumption of centralized data access, offering little to no privacy protection. In contrast, federated analytics (FA) enables privacy-preserving computation across distributed data sources, but lacks support for natural language input and requires structured, machine-readable queries. In this work, we present LAFA, the first system that integrates LLM-agent-based data analytics with FA. LAFA introduces a hierarchical multi-agent architecture that accepts natural language queries and transforms them into optimized, executable FA workflows. A coarse-grained planner first decomposes complex queries into sub-queries, while a fine-grained planner maps each sub-query into a Directed Acyclic Graph of FA operations using prior structural knowledge. To improve execution efficiency, an optimizer agent rewrites and merges multiple DAGs, eliminating redundant operations and minimizing computational and communicational overhead. Our experiments demonstrate that LAFA consistently outperforms baseline prompting strategies by achieving higher execution plan success rates and reducing resource-intensive FA operations by a substantial margin. This work establishes a practical foundation for privacy-preserving, LLM-driven analytics that supports natural language input in the FA setting.
Haichao Ji, Zibo Wang 0001, Yifei Zhu 0001, Dan Wang 0002, Zhu Han 0001
CloudCom2
2025 Towards Fair and Scalable Trial Assignment in Federated Bandits: A Shapley Value Approach
abstract
Federated multi-armed bandits extend the multi-armed bandits framework to the federated learning setting where multiple clients in the same exploration space collaboratively identify the optimal arm. While previous studies demonstrated its efficiency like classical federated learning, in this paper, we present the first work that reveals the serious fairness problem in federated multi-armed bandits when clients have heterogeneous and overlapping armsets. The fairness problem happens because clients with different trial requirements should conduct different numbers of trials summing up to a global requirement, but they wish to minimize their exploration effort. To address the novel fairness concern, we formally formulate the fairness-aware trial assignment as a coalitional game. Based on the theoretically derived trial requirements, we devise a Shapley value-based trial assignment mechanism to guarantee fairness. Regardless of the #P-hard complexity when deriving the general Shapley value, we achieve an accurate computation of trial assignment with polynomial complexity by exploiting its unique characteristic. We further carefully control the numbers of trials in each iteration to resolve the communication bottleneck and minimize the wasted trials. Experiment results show that, compared to the naïve federated scheme, our design outperforms with both high fairness metrics and high efficiency in total trials and communication.
Zibo Wang 0001, Yifei Zhu 0001, Dan Wang 0002, Zhu Han 0001
IEEE Trans. Big Data1
2024 DPBalance: Efficient and Fair Privacy Budget Scheduling for Federated Learning as a Service
abstract
Federated learning (FL) has emerged as a prevalent distributed machine learning scheme that enables collaborative model training without aggregating raw data. Cloud service providers further embrace Federated Learning as a Service (FLaaS), allowing data analysts to execute their FL training pipelines over differentially-protected data. Due to the intrinsic properties of differential privacy, the enforced privacy level on data blocks can be viewed as a privacy budget that requires careful scheduling to cater to diverse training pipelines. Existing privacy budget scheduling studies prioritize either efficiency or fairness individually. In this paper, we propose DPBalance, a novel privacy budget scheduling mechanism that jointly optimizes both efficiency and fairness. We first develop a comprehensive utility function incorporating data analyst-level dominant shares and FL-specific performance metrics. A sequential allocation mechanism is then designed using the Lagrange multiplier method and effective greedy heuristics. We theoretically prove that DPBalance satisfies Pareto Efficiency, Sharing Incentive, Envy-Freeness, and Weak Strategy Proofness. We also theoretically prove the existence of a fairness-efficiency tradeoff in privacy budgeting. Extensive experiments demonstrate that DPBalance outperforms state-of-the-art solutions, achieving an average efficiency improvement of 1.44× ~ 3.49×, and an average fairness improvement of 1.37×~24.32×.
Zibo Wang 0001, Yifei Zhu 0001, Chen Chen 0067
INFOCOM2
2024 Federated Analytics-Empowered Frequent Pattern Mining for Decentralized Web 3.0 Applications
abstract
The emerging Web 3.0 paradigm aims to decentralize existing web services, enabling desirable properties such as transparency, incentives, and privacy preservation. However, current Web 3.0 applications supported by blockchain infrastructure still cannot support complex data analytics tasks in a scalable and privacy-preserving way. This paper introduces the emerging federated analytics (FA) paradigm into the realm of Web 3.0 services, enabling data to stay local while still contributing to complex web analytics tasks in a privacy-preserving way. We propose FedWeb, a tailored FA design for important frequent pattern mining tasks in Web 3.0. FedWeb remarkably reduces the number of required participating data owners to support privacy-preserving Web 3.0 data analytics based on a novel distributed differential privacy technique. The correctness of mining results is guaranteed by a theoretically rigid candidate filtering scheme based on Hoeffding’s inequality and Chebychev’s inequality. Two response budget saving solutions are proposed to further reduce participating data owners. Experiments on three representative Web 3.0 scenarios show that FedWeb can improve data utility by ∼25.3% and reduce the participating data owners by ∼98.4%.
Zibo Wang 0001, Yifei Zhu 0001, Dan Wang 0002, Zhu Han 0001
INFOCOM1
2023 Secure Trajectory Publication in Untrusted Environments: A Federated Analytics Approach
abstract
The increasing awareness of privacy and the adoption of data regulations challenge the traditional trajectory publication framework in which a trusted server has access to the raw data from mobile clients. In the new untrusted environment, the clients call for much stronger data privacy preservation locally without sharing their raw data. Based on the emerging paradigm of federated analytics, we propose a Federated Analytics-based Secure Trajectory PUBlication (FASTPub) mechanism to operate in such untrusted environments. Compared with existing local differential privacy (LDP) methods, FASTPub guarantees LDP and loss-bounded$k$-anonymity simultaneously with greatly improved data utility. Specifically, FASTPub works interactively between the server and clients and iteratively builds up the trajectory without exposing raw data. Sampled clients only respond to selected trajectory fragments with randomized answers to preserve privacy as much as possible. The server then intelligently aggregates these randomized responses leveraging the intrinsic Apriori property and a Markov independent assumption of trajectory data to guide further iterations. Extensive experiments on synthetic and real-world datasets on two downstream tasks demonstrate that FASTPub gains a remarkably improved data utility compared to the existing state-of-the-art solutions.
Zibo Wang 0001, Yifei Zhu 0001, Dan Wang 0002, Zhu Han 0001
IEEE Trans. Mob. Comput.1
2022 FedFPM: A Unified Federated Analytics Framework for Collaborative Frequent Pattern Mining
abstract
Frequent pattern mining is an important class of knowledge discovery problems. It aims at finding out high-frequency items or structures (e.g., itemset, sequence) in a database, and plays an essential role in deriving other interesting patterns, like association rules. The traditional approach of gathering data to a central server and analyze is no longer viable due to the increasing awareness of user privacy and newly established laws on data protection. Previous privacy-preserving frequent pattern mining approaches only target a particular problem with great utility loss when handling complex structures. In this paper, we take the first initiative to propose a unified federated analytics framework (FedFPM) for a variety of frequent pattern mining problems, including item, itemset, and sequence mining. FedFPM achieves high data utility and guarantees local differential privacy without uploading raw data. Specifically, FedFPM adopts an interactive query-response approach between clients and a server. The server meticulously employs the Apriori property and the Hoeffding’s inequality to generates informed queries. The clients randomize their responses in the reduced space to realize local differential privacy. Experiments on three different frequent pattern mining tasks demonstrate that FedFPM achieves better performances than the state-of-the-art specialized benchmarks, with a much smaller computation overhead.
Zibo Wang 0001, Yifei Zhu 0001, Dan Wang 0002, Zhu Han 0001
INFOCOM1
2021 FedACS: Federated Skewness Analytics in Heterogeneous Decentralized Data Environments
abstract
The emerging federated optimization paradigm performs data mining or artificial intelligence techniques locally on the edge devices, enabling scientists and engineers to utilize the blooming edge data with privacy protection. In such a paradigm, since data cannot be shared or gathered, data heterogeneity naturally emerges, which significantly degrades the performance of federated optimization, ultimately leading to poor quality of federated services. In this paper, we present the first work on characterizing the data heterogeneity in the framework of federated analytics, i.e., to collectively carry out analytics tasks without raw data sharing, and use the information to create a desirable data environment via intelligent client selection. Our proposed Analytics-driven Client Selection framework, named FedACS, tackles the data heterogeneity problem in three steps. First, clients are in charge of generating insights about local data without disclosure of sensitive information. Then, the server uses these insights to infer the situation of clients’ data heterogeneity based on the Hoeffding’s inequality. Finally, a dueling bandit is formulated to intelligently select clients with slighter data heterogeneity to form a desirable client pool. FedACS can be universally applied to all kinds of federated optimization tasks, and gains benefits including privacy protection, infrastructure reuse, and client load reduction. To test its efficiency, we further customize it to assist federated learning, a popular scenario of federated optimization. According to experiment results, FedACS reduces the accuracy degrading by up to 65.6%, and speeds up the convergence for up to 2.4 times.
Zibo Wang 0001, Yifei Zhu 0001, Dan Wang 0002, Zhu Han 0001
IWQoS1