EDBT 2026 Demo / reviewers in the wild / expert
Xuebin Ren
dblp:136/3467
· DBLP profile ↗
42ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0002-6498-0250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 16 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BIQ: Bisection Interval Quantization for Communication-efficient Federated LearningabstractQuantization is a pivotal technique for enhancing communication efficiency in Federated Learning (FL). Traditional quantization methods often set uniform intervals, may fail to adequately characterize non-uniform data distributions, thus leading to substantial estimation errors and degrated model performance. Non-uniform quantization can better solve the problem. However, when applied to FL, it would bring additional communication overheads for the alignment of parameter distributions among distributed models. To address this issue, we propose Bisection Interval Quantization (BIQ), a novel non-uniform quantization framework for FL with great communication efficiency. In particular, BIQ works by optimizing the interval selection through recursive bisection among distributed clients without extra parameter communication. For scenarios involving amounts of boundary inputs, we further design Weighted Bisection Interval Quantization (WBIQ), which incorporates maximum likelihood estimation to refine boundary value reconstruction to enhance the estimation quality of boundary inputs. Our theoretical analysis rigorously establishes, for the first time under biased quantization conditions, that both BIQ and WBIQ achieve tighter error bounds and enhanced stability. Extensive experiments validate that both BIQ and WBIQ significantly accelerate the convergence of FL model training when compared to the state-of-the-art quantizers under both convex and non-convex settings. Luyang Gai, Shusen Yang, Xuebin Ren |
AAAI | 3 |
| 2026 | Fine-Grained Manipulation Attacks to Local Differential Privacy Protocols for Range Query
Wenda Chen, Xuebin Ren |
ICDE | 3 |
| 2026 | Fine-Grained Manipulation Attacks to Local Differential Privacy Protocols for Data StreamsabstractLocal Differential Privacy (LDP) enables massive data collection and analysis while protecting end users' privacy against untrusted aggregators. It has been applied to various data types (e.g., categorical, numerical, and graph data) and application settings (e.g., static and streaming). Recent findings indicate that LDP protocols can be easily disrupted by poisoning or manipulation attacks, where an attacker can leverage injected/corrupted fake users to send crafted data to the aggregator in order to manipulate the final estimate of the aggregator. However, current attacks primarily target static protocols, neglecting the security of LDP protocols in the streaming settings. Our research fills the gap by developing novel fine-grained manipulation attacks to LDP protocols for data streams. By reviewing the attack surfaces in existing algorithms, we introduce a unified attack framework with composable modules, which can manipulate the LDP estimated stream toward a target stream. Our attack framework can adapt to state-of-the-art streaming LDP algorithms with different analytic tasks (e.g., frequency and mean) and LDP models (event-level, user-level, <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$w$</tex-math></inline-formula>-event level). We verify our attacks theoretically and validate them through extensive experiments on real-world datasets. Finally, we explore a possible defense mechanism for mitigating our attacks. Xuebin Ren, Shusen Yang, Chia-Mu Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | FairGFL: Privacy-Preserving Fairness-Aware Federated Learning With Overlapping SubgraphsabstractGraph federated learning enables the collaborative extraction of high-order information from distributed subgraphs while preserving the privacy of raw data. However, graph data often exhibits overlap among different clients. Previous research has demonstrated certain benefits of overlapping data in mitigating data heterogeneity. However, the negative effects have not been explored, particularly in cases where the overlaps are imbalanced across clients. In this paper, we uncover the unfairness issue arising from imbalanced overlapping subgraphs through both empirical observations and theoretical reasoning. To address this issue, we propose FairGFL (FAIRness-aware subGraph Federated Learning), a novel algorithm that enhances cross-client fairness while maintaining model utility in a privacy-preserving manner. Specifically, FairGFL incorporates an interpretable weighted aggregation approach to enhance fairness across clients, leveraging privacy-preserving estimation of their overlapping ratios. Furthermore, FairGFL improves the tradeoff between model utility and fairness by integrating a carefully crafted regularizer into the federated composite loss function. Through extensive experiments on four benchmark graph datasets, we demonstrate that FairGFL outperforms four representative baseline algorithms in terms of both model utility and fairness. Shusen Yang, Fangyuan Zhao, Xuebin Ren |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Differentially Private Fine-Tuning of Diffusion ModelsabstractThe integration of Differential Privacy (DP) with diffusion models (DMs) presents a promising yet challenging frontier, particularly due to the substantial memorization capabilities of DMs that pose significant privacy risks. Differential privacy offers a rigorous framework for safeguarding individual data points during model training, with Differential Privacy Stochastic Gradient Descent (DP-SGD) being a prominent implementation. Diffusion method decomposes image generation into iterative steps, theoretically aligning well with DP's incremental noise addition. Despite the natural fit, the unique architecture of DMs necessitates tailored approaches to effectively balance privacy-utility trade-off. Recent developments in this field have highlighted the potential for generating high-quality synthetic data by pre-training on public data (i.e., ImageNet) and fine-tuning on private data, however, there is a pronounced gap in research on optimizing the trade-offs involved in DP settings, particularly concerning parameter efficiency and model scalability. Our work addresses this by proposing a parameter-efficient fine-tuning strategy optimized for private diffusion models, which minimizes the number of trainable parameters to enhance the privacy-utility trade-off. We empirically demonstrate that our method achieves state-of-the-art performance in DP synthesis, significantly surpassing previous benchmarks on widely studied datasets (e.g., with only 0.47M trainable parameters, achieving a more than 35% improvement over the previous state-of-the-art with a small privacy budget on the CelebA-64 dataset). Anonymous codes available at https://anonymous.4open.science/r/DP-LORA-F02F. Yu-Lin Tsai, Chia-Mu Yu, Xuebin Ren, Francois Buet-Golfouse |
ICCV | 4 |
| 2025 | FedImpute: Personalized federated learning for data imputation with clusterer and auxiliary classifier
Yanan Li 0004, Shaocong Guo, Xinyuan Guo, Peng Zhao 0001, Xuebin Ren, Hui Wang 0071 |
Expert Syst. Appl. | 5 |
| 2025 | FedGen: Personalized federated learning with data generation for enhanced model customization and class imbalance
Peng Zhao 0001, Shaocong Guo, Yanan Li 0004, Shusen Yang, Xuebin Ren |
Future Gener. Comput. Syst. | 5 |
| 2025 | T³Planner: Multi-Phase Planning Across Structure-Constrained Optical, IP, and Routing TopologiesabstractNetwork topology planning is an essential multi-phase process to build and jointly optimize the multi-layer network topologies in wide-area networks (WANs). Most existing practices target single-phase/layer planning, and are incapable of satisfying all rigorous topological structure constraints (e.g., dual-homing rings) defined by network standards and operators, especially in large-scale networks. These significantly limit their usability and performance in production networks. We consider a general topology planning problem with typical structure constraints over three essential phases (greenfield, reconfiguration, and site expansion) and topological layers (optical, IP, and routing topologies). We present, T3Planner, a novel practical solver to this problem in production. Specifically, we develop a structure-driven encoder based on graph neural network (GNN) for concise structure encoding, and design a new learning framework with optical-centric layer compression/reconstruction and rule-aided reinforcement learning (RL) for fast convergence and high performance. Extensive experiments on nine real topologies demonstrate that T3Planner scales to large optical networks with hundreds of sites, saves 46.6% cost, and supports$3.12\times $more demand when compared to related existing approaches. Yijun Hao, Shusen Yang, Cong Zhao 0001, Xuebin Ren, Peng Zhao 0001, Chenren Xu, Shibo Wang 0002 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Learning Adaptive Multi-Timescale Scheduling for Mobile Edge ComputingabstractIn mobile edge computing (MEC), resource scheduling is crucial to task requests’ performance and service providers’ cost, involving multi-layer heterogeneous scheduling decisions. Existing MEC schedulers typically adopt static-timescale scheduling, where scheduling decisions are updated regularly at fixed intervals for all layers. The inflexible updating timescales lead to poor performance in the production networks. In this paper, we propose EdgeTimer, an unprecedented approach that automatically and adaptively determines respective updating timescales of multiple scheduling layers to achieve a better trade-off between the operation cost and service performance. Specifically, we design (i) a three-layer hierarchical deep reinforcement learning (DRL) framework for efficient learning of tightly coupled policies, (ii) a tailored multi-agent DRL algorithm for decentralized scheduling, with the convergence strictly proved, and (iii) a lightweight system defender for deterministic reliability assurance. Furthermore, we apply EdgeTimer to a wide range of Kubernetes scheduling rules, and evaluate it using production traces with different workload patterns. Through extensive trace-driven experiments, we demonstrate that EdgeTimer can significantly decrease the operation cost for service providers without sacrificing the delay performance, thereby improving overall profits, compared with the state-of-the-art approaches. Yijun Hao, Shusen Yang, Shibo Wang 0002, Xuebin Ren |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | FedDSV: Shapley Value-Based Contribution Estimation in Federated Learning With Dynamic ParticipationabstractFederated Learning (FL) succeeds in collaborative and privacy-preserving ML model training among multiple distributed data owners. To maintain a healthy FL ecosystem, it is crucial to estimate the contributions of all participants fairly. Due to provable fairness, Shapley value (SV) is widely used for contribution estimation in FL. However, current studies focus on static scenarios with fixed participants and neglect the dynamic settings with the random joining or leaving of participants in practice. This paper fills the gap by proposing FedDSV, a novel contribution estimation framework for FL with dynamic participation. FedDSV supports flexible weighting mechanisms and is compatible with the SV fairness properties in dynamic scenarios. To reduce the computational complexity, we propose a Monte Carlo variant sampling method (SMC), which can adapt well to dynamic scenarios and approximate the true SVs. To evaluate the effectiveness and efficiency of our proposed approaches, extensive experiments under different settings (e.g., frequency switching, low-quality detection, etc.) are conducted on both i.i.d and non-i.i.d. distributions. Experimental results demonstrate that FedDSV can reflect the real utility contribution of data sources for dynamic FL, and SMC can approximate the exact dynamic SVs with larger similarities in a much shorter time than the state-of-the-art methods. Kaijia Lei, Xuebin Ren, Shusen Yang, Fangyuan Zhao |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement LearningabstractIn mobile edge computing (MEC), resource scheduling is crucial to task requests’ performance and service providers’ cost, involving multi-layer heterogeneous scheduling decisions. Existing schedulers typically adopt static timescales to regularly update scheduling decisions of each layer, without adaptive adjustment of timescales for different layers, resulting in potentially poor performance in practice.We notice that the adaptive timescales would significantly improve the trade-off between the operation cost and delay performance. Based on this insight, we propose EdgeTimer, the first work to automatically generate adaptive timescales to update multi-layer scheduling decisions using deep reinforcement learning (DRL). First, EdgeTimer uses a three-layer hierarchical DRL framework to decouple the multi-layer decision-making task into a hierarchy of independent sub-tasks for improving learning efficiency. Second, to cope with each sub-task, EdgeTimer adopts a safe multi-agent DRL algorithm for decentralized scheduling while ensuring system reliability. We apply EdgeTimer to a wide range of Kubernetes scheduling rules, and evaluate it using production traces with different workload patterns. Extensive trace-driven experiments demonstrate that EdgeTimer can learn adaptive timescales, irrespective of workload patterns and built-in scheduling rules. It obtains up to 9:1 more profit than existing approaches without sacrificing the delay performance. Yijun Hao, Shusen Yang, Shibo Wang 0002, Xuebin Ren |
INFOCOM | 6 |
| 2024 | VertiMRF: Differentially Private Vertical Federated Data SynthesisabstractData synthesis is a promising solution to share data for various downstream analytic tasks without exposing raw data. However, without a theoretical privacy guarantee, a synthetic dataset would still leak some sensitive information in raw data. As a countermeasure, differential privacy is widely adopted to safeguard data synthesis by strictly limiting the released information. This technique is advantageous yet presents significant challenges in the vertical federated setting, where data attributes are distributed among different data parties. The main challenge lies in maintaining privacy while efficiently and precisely reconstructing the correlation between attributes. In this paper, we propose a novel algorithm called VertiMRF, designed explicitly for generating synthetic data in the vertical setting and providing differential privacy protection for all information shared from data parties. We introduce techniques based on the Flajolet-Martin (FM) sketch for encoding local data satisfying differential privacy and estimating cross-party marginals. We provide theoretical privacy and utility proof for encoding in this multi-attribute data. Collecting the locally generated private Markov Random Field (MRF) and the sketches, a central server can reconstruct a global MRF, maintaining the most useful information. Two critical techniques introduced in our VertiMRF are dimension reduction and consistency enforcement, preventing the noise of FM sketch from overwhelming the information of attributes with large domain sizes when building the global MRF. These two techniques allow flexible and inconsistent binning strategies of local private MRF and the data sketching module, which can preserve information to the greatest extent. We conduct extensive experiments on four real-world datasets to evaluate the effectiveness of VertiMRF. End-to-end comparisons demonstrate the superiority of VertiMRF. Fangyuan Zhao, Zitao Li, Xuebin Ren, Bolin Ding, Shusen Yang, Yaliang Li |
KDD | 3 |
| 2024 | FtlSPG: A Federated Transfer Learning Framework for Personalized Safety Protective Gear Detection in Electric Power IndustryabstractSafety protective gear (SPG) detection based on the machine learning model plays an important role in improving outdoor personnel safety in the electric power industry. However, the detection method of transmitting video to the cloud faces a series of challenges, such as privacy disclosure and high latency. To solve this problem, we present FtlSPG, a federated transfer learning framework for SPG detection. In particular, under the three-layer pyramid architecture of “site-companyCloud-globalServer,” we propose a federated personalized model based on local batch normalization and dynamical weighting for the source domain with labeled video. Moreover, a federated domain adaptation model based on a federated deep adversarial network and model self-training is presented for the target domain with unlabeled video. Finally, we verify the effectiveness of FtlSPG in real-world power companies. Extensive experiments demonstrate that FtlSPG can significantly outperform existing schemes, in terms of privacy protection, detection precision, and response latency. Shusen Yang, Cong Zhao 0001, Peng Zhao 0001, Xuebin Ren |
IEEE Internet Things J. | 7 |
| 2024 | Knowledge and Data Dual-Driven Fault Diagnosis in Industrial Scenarios: A SurveyabstractKnowledge and data dual-driven (KDDD) represents a novel paradigm that leverages the strengths of data-driven methods in feature representation and knowledge transfer, while also incorporating expertise accumulated by domain experts. This integration allows KDDD methods to enhance the interpretability, reliability, and robustness of fault diagnosis (FD) approaches, making them widely studied in the field of industrial equipment (IE) FD. Despite the existence of systematic and valuable reviews on IE FD, there remains a gap in the literature regarding the review of KDDD IE FD methods. Therefore, conducting a comprehensive investigation into KDDD IE FD methods is of utmost importance and necessity. Such an investigation will facilitate readers’ understanding of advanced technologies and enable the rapid design of effective solutions for real-world IE FD problems. In this survey, we first outline the limitations of data-driven and knowledge-based FD methods, highlighting the need for KDDD methods. Subsequently, we delve into the details of how domain knowledge can be effectively integrated with deep learning models. Additionally, we analyze challenges of KDDD methods in real-world IE FD applications, while also discussing novel solutions for prospective research directions. Finally, we conclude this survey, emphasizing the inspiration it offers to researchers interested in advancing IE FD, and its potential to stimulate practical IE FD research. Shusen Yang, Cong Zhao 0001, Peng Zhao 0001, Xuebin Ren |
IEEE Internet Things J. | 7 |
| 2024 | A personalized cross-domain recommendation with federated meta learning
Peng Zhao 0001, Yuanyang Jin, Xuebin Ren, Yanan Li 0004 |
Multim. Tools Appl. | 3 |
| 2024 | Multi-Stage Asynchronous Federated Learning With Adaptive Differential PrivacyabstractThe fusion of federated learning and differential privacy can provide more comprehensive and rigorous privacy protection, thus attracting extensive interests from both academia and industry. However, facing the system-level challenge of device heterogeneity, most current synchronous FL paradigms exhibit low efficiency due to the straggler effect, which can be significantly reduced by Asynchronous FL (AFL). However, AFL has never been comprehensively studied, which imposes a major challenge in the utility optimization of DP-enhanced AFL. Here, theoretically motivated multi-stage adaptive private algorithms are proposed to improve the trade-off between model utility and privacy for DP-enhanced AFL. In particular, we first build two DP-enhanced AFL frameworks with consideration of universal factors for different adversary models. Then, we give a solid analysis on the model convergence of AFL, based on which, DP can be adaptively achieved with high utility. Through extensive experiments on different training models and benchmark datasets, we demonstrate that the proposed algorithms achieve the overall best performances and improve up to 24% test accuracy with the same privacy loss and have faster convergence compared with the state-of-the-art algorithms. Our frameworks provide an analytical way for private AFL and adapt to more complex FL application scenarios. Yanan Li 0004, Shusen Yang, Xuebin Ren, Cong Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Reducing Traffic Wastage in Video Streaming via Bandwidth-Efficient Bitrate AdaptationabstractBitrate adaptation (also known as ABR) is a crucial technique to improve the quality of experience (QoE) for video streaming applications. However, existing ABR algorithms suffer from severe traffic wastage, which refers to the traffic cost of downloading the video segments that users do not finally consume, for example, due to early departure or video skipping. In this paper, we carefully formulate the dynamics of buffered data volume (BDV), a strongly correlated indicator of traffic wastage, which, to the best of our knowledge, is the first time to rigorously clarify the effect of downloading plans on potential wastage. To reduce wastage while keeping a high QoE, we present a bandwidth-efficient bitrate adaptation algorithm (named BE-ABR), achieving consistently low BDV without distinct QoE losses. Specifically, we design a precise, time-aware transmission delay prediction model over the Transformer architecture, and develop a fine-grained buffer control scheme. Through extensive experiments conducted on emulated and real network environments including WiFi, 4G, and 5G, we demonstrate that BE-ABR performs well in both QoE and bandwidth savings, enabling a 60.87% wastage reduction and a comparable, or even better, QoE, compared to the state-of-the-art methods. Hairong Su, Shibo Wang 0002, Shusen Yang, Tianchi Huang, Xuebin Ren |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Exploring the Benefits of Visual Prompting in Differential PrivacyabstractVisual Prompting (VP) is an emerging and powerful technique that allows sample-efficient adaptation to downstream tasks by engineering a well-trained frozen source model. In this work, we explore the benefits of VP in constructing compelling neural network classifiers with differential privacy (DP). We explore and integrate VP into canonical DP training methods and demonstrate its simplicity and efficiency. In particular, we discover that VP in tandem with PATE, a state-of-the-art DP training method that leverages the knowledge transfer from an ensemble of teachers, achieves the state-of-the-art privacy-utility tradeoff with minimum expenditure of privacy budget. Moreover, we conduct additional experiments on cross-domain image classification with a sufficient domain gap to further unveil the advantage of VP in DP. Lastly, we also conduct extensive ablation studies to validate the effectiveness and contribution of VP under DP consideration. Our code is available at https://github.com/EzzzLi/Prompt-PATE. Yu-Lin Tsai, Chia-Mu Yu, Xuebin Ren |
ICCV | 5 |
| 2023 | MPDM: A Multi-Paradigm Deployment Model for Large-Scale Edge-Cloud IntelligenceabstractThe development of cloud and edge computing has enabled the easy access of artificial intelligence (AI) services for massive heterogeneous and resource-constrained devices. Particularly, computation-intensive AI services can be orchestrated and deployed in the cloud or edge according to varying performance and cost requirements. Nonetheless, the improved accessibility of deep learning (DL) model variants and the evolving of computational intelligence paradigms pose great challenges for orchestrating large-scale DL inference services in the cloud-edge continuum. Focusing on cloud or edge-based deployment, existing work on multi-variant service orchestration often has a limited solution space of deployment plans. To address this limitation, we first propose a novel multi-paradigm deployment model (MPDM) for service orchestration, which not only considers the model variants but also allows the co-existence of multiple paradigms for large-scale inference service deployment. The service deployment in the MPDM model is then formulated as a multiobjective optimization problem of seeking a better tradeoff among the system accuracy, service scale, and deployment cost. To solve the multiobjective optimization, we further propose a weighted metric-based constructive heuristic algorithm (WCH), which can efficiently obtain an approximately optimal Pareto frontier. Extensive experimental results have validated the effectiveness and efficiency of WCH, and revealed the impacts of both multi-paradigm deployment and edge-cloud collaborative intelligence (ECCI) paradigm on large-scale DL serving systems. Luhui Wang, Xuebin Ren, Cong Zhao 0001, Fangyuan Zhao, Shusen Yang |
IEEE Internet Things J. | 2 |
| 2023 | Federated multi-objective reinforcement learning
Fangyuan Zhao, Xuebin Ren, Shusen Yang, Peng Zhao 0001 |
Inf. Sci. | 2 |
| 2023 | RTCoInfer: Real-Time Collaborative CNN Inference for Stream Analytics on Ubiquitous ImagesabstractEmerging intelligent applications based on accurate and timely stream analytics require real-time CNN inference of massive data continuously generated at the pervasive end devices. Due to the resource constraints, neither computing locally at end devices nor transmitting to remote servers is competent for computation-intensive CNN inference on large-volume images in real-time. Therefore, Collaborative Inference (CI), which conducts inference sequentially from the local device to the remote server with compressed intermediate inference data, is rapidly promoted. Due to the essential communication in collaboration, the CI efficiency is sensitive to network conditions, and will degrade under the unpredictable network fluctuations in practice, which may cause a severe delay in CI and degrade the responsiveness of stream analytics. For accurate and timely stream analytics in practical fluctuating networks, we present RTCoInfer, the real-time CI framework with run-time transmission adaption considering the network conditions. Specifically, we propose a novel Switchable CNN integrating CNNs with different compression rates on the partition layer for the run-time transmission adjustment, and construct a real-time controller determining the compression rate to maintain the real-time CI for stream analytics. Extensive experiments show that, compared with state-of-the-art methods, RTCoInfer achieves better efficiency and unprecedented resilience in real-time stream analytics. Zhanhua Zhang, Shusen Yang, Cong Zhao 0001, Xuebin Ren, Hanqiao Yu, Siyan Guo |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | LDP-IDS: Local Differential Privacy for Infinite Data StreamsabstractLocal differential privacy (LDP) is promising for private streaming data collection and analysis. However, existing few LDP studies over streams either apply to finite streams only or may suffer from insufficient protection. This paper investigates this problem by proposing LDP-IDS, a novel w-event LDP paradigm to provide practical privacy guarantee for infinite streams. By constructing a unified error analysis, we adapt the existing budget division framework in centralized differential privacy (CDP) for LDP-IDS, which however incurs prohibitive noise and expensive communication cost. To this end, we propose a novel and extensible framework of population division and recycling, as well as online adaptive population division algorithms for LDP-IDS. We provide theoretical guarantees and demonstrate, through extensive discussions, that our proposed framework not only achieves significant reduction in utility loss and communication overhead, but also enjoys great compatibility for varied analytic tasks and flexibility of incorporating ideas of many existing stream algorithms. Extensive experiments on synthetic and real-world datasets validate the high effectiveness, efficiency, and flexibility of our proposed framework and methods. Xuebin Ren, Weiren Yu, Shusen Yang, Cong Zhao 0001, Zongben Xu |
SIGMOD Conference | 1 |
| 2022 | PCFed: Privacy-Enhanced and Communication-Efficient Federated Learning for Industrial IoTsabstractFederated learning (FL) is capable of analyzing tremendous data from smart edge devices in Industrial Internet of Things (IIoTs), empowering numerous industrial applications. However, the increasing privacy concerns and deployment costs of IIoT environment have been posing new challenges for FL. This article proposes PCFed, a novel privacy-enhanced and communication-efficient FL framework to provide higher model accuracy with rigorous privacy guarantees and great communication efficiency. In particular, we develop a sampling-based intermittent communication strategy via a PID (proportional, integral, and derivative) controller on the cloud server to adaptively reduce the communication frequency. In addition, we design a budget allocation mechanism to balance the tradeoff between model accuracy and privacy loss. Then, we develop PCFed+, an enhanced variant for PCFed, with further consideration of infinite data streams on edge servers. Extensive experiments demonstrate that both PCFed and PCFed+ can significantly outperform existing schemes, in terms of communication efficiency, privacy protection, and model accuracy. Shusen Yang, Xuebin Ren, Peng Zhao 0001, Cong Zhao 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | IndustEdge: A Time-Sensitive Networking Enabled Edge-Cloud Collaborative Intelligent Platform for Smart IndustryabstractAn edge-cloud collaborative intelligent (ECCI) platform is of great significance for the agile development and rapid deployment of ECCI applications, which are essential for realizing smart industry in the era of Industry 4.0. However, the existing platforms lack considering the high real-time latency demand of industrial operations, which severely hinders the development of smart industry and may even lead to severe industrial accidents. To effectively reduce the response latency of industrial applications, in this article, we propose an ECCI platform IndustEdge. It takes time-sensitive networking as the deterministic transport for the link layer, and provides an extensible ECCI orchestration component to reduce the system level latency. Furthermore, IndustEdge has an ECCI algorithm library for different collaborative modes and provides the complete life cycle management for ECCI applications. We implement platforms for both the real-world prototype and emulated-world emulation, and conduct two case studies to evaluate the effectiveness of IndustEdge. Shusen Yang, Xuebin Ren, Peng Zhao 0001, Cong Zhao 0001, Xinyu Yang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Towards Efficient and Stable K-Asynchronous Federated Learning With Unbounded Stale Gradients on Non-IID DataabstractFederated learning (FL) is an emerging privacy-preserving paradigm that enables multiple participants collaboratively to train a global model without uploading raw data. Considering heterogeneous computing and communication capabilities of different participants, asynchronous FL can avoid the stragglers effect in synchronous FL and adapts to scenarios with vast participants. Both staleness and non-IID data in asynchronous FL would reduce the model utility. However, there exists an inherent contradiction between the solutions to the two problems. That is, mitigating the staleness requires to select less but consistent gradients while coping with non-IID data demands more comprehensive gradients. To address the dilemma, this paper proposes a two-stage weighted$K$asynchronous FL with adaptive learning rate (WKAFL). By selecting consistent gradients and adjusting learning rate adaptively, WKAFL utilizes stale gradients and mitigates the impact of non-IID data, which can achieve multifaceted enhancement in training speed, prediction accuracy and training stability. We also present the convergence analysis for WKAFL under the assumption of unbounded staleness to understand the impact of staleness and non-IID data. Experiments implemented on both benchmark and synthetic FL datasets show that WKAFL has better overall performance compared to existing algorithms. Yanan Li 0004, Xuebin Ren, Shusen Yang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Locally Private High-Dimensional Crowdsourced Data Release Based on Copula FunctionsabstractWith the increasing popularity of crowdsourcing services, high-dimensional crowdsourced data provides a wealth of knowledge. Nonetheless, unprecedented privacy threats to participants have emerged, due to complex correlations among multiple attributes and the vulnerabilities of untrusted crowdsourcing servers. Differential privacy-based paradigms have been proposed to release privacy-preserving datasets with statistical approximation. Nonetheless, most existing schemes are limited when facing highly correlated attributes, and cannot prevent privacy threats from untrusted crowdsourcing servers. To address this issue, we propose two novel solutions, namelyLoCopandDR_LoCop, which guarantee local differential privacy based on the randomized response technique while synthesizing and releasing high-dimensional crowdsourced data with high data utility. Particularly,LoCopleverages copula theory to synthesize high-dimensional crowdsourced data via univariate marginal distribution and attribute dependence. Univariate marginal distribution is estimated by the Lasso-based regression algorithm from aggregated privacy-preserving bit strings. Dependencies among attributes are modeled as multivariate Gaussian copula. Based onLoCop, the enhanced solutionDR_LoCopnot only takes advantage of C-vine copula to reflect conditional dependencies among high-dimensional attributes, but also achieves dimension reduction. Extensive experiments on real-world datasets demonstrate that our solutions substantially outperform the state-of-the-art techniques in terms of both data utility and computational overhead. Xinyu Yang 0001, Xuebin Ren, Wei Yu 0002, Shusen Yang |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Local Differential Privacy for data collection and analysis
Jun Zhao 0007, Xinyu Yang 0001, Xuebin Ren, Kwok-Yan Lam |
Neurocomputing | 5 |
| 2021 | DPCrowd: Privacy-Preserving and Communication-Efficient Decentralized Statistical Estimation for Real-Time Crowdsourced DataabstractIn Internet-of-Things (IoT)-driven smart-world systems, real-time crowdsourced databases from multiple distributed servers can be aggregated to extract dynamic statistics from a larger population, thus providing more reliable knowledge for our society. Particularly, multiple distributed servers in a decentralized network can realize real-time collaborative statistical estimation by disseminating statistics from their separate databases. Despite no raw data sharing, the real-time statistics could still expose the data privacy of crowdsourcing participants. For mitigating the privacy concern, while the traditional differential privacy (DP) mechanism can be simply implemented to perturb the statistics in each timestamp and independently for each dimension, this may suffer a great utility loss from the real-time and multidimensional crowdsourced data. Also, the real-time broadcasting would bring significant overheads in the whole network. To tackle the issues, we propose a novel privacy preserving and communication-efficient decentralized statistical estimation algorithm (DPCrowd), which only requires intermittently sharing the DP protected parameters with one-hop neighbors by exploiting the temporal correlations in real-time crowdsourced data. Then, with further consideration of spatial correlations, we develop an enhanced algorithm, DPCrowd+, to deal with multidimensional infinite crowd-data streams. Extensive experiments on several data sets demonstrate that our proposed schemes DPCrowd and DPCrowd+ can significantly outperform existing schemes in providing accurate and consensus estimation with rigorous privacy protection and great communication efficiency. Xuebin Ren, Chia-Mu Yu, Wei Yu 0002, Xinyu Yang 0001, Jun Zhao 0007, Shusen Yang |
IEEE Internet Things J. | 1 |
| 2021 | Survey on Improving Data Utility in Differentially Private Sequential Data PublishingabstractThe massive generation, extensive sharing, and deep exploitation of data in the big data era have raised unprecedented privacy threats. To address privacy concerns, various privacy paradigms have been proposed to achieve a good tradeoff between privacy and data utility. Particularly, differential privacy has been well accepted as one of the de facto standard for privacy preservation, and numerous schemes guaranteeing differential privacy have been proposed. Nonetheless, most of the existing works claiming a superior utility-privacy tradeoff only present specific methods, with distinct perspectives, and a complete comparative analysis and evaluation study has not been fully investigated. To this end, in this paper we review and investigate existing schemes on providing differential privacy from a broad and encompassing perspective to provide a comprehensive survey with respect to both the privacy guarantee and the effectiveness and efficiency in utility improvement. We categorize the existing schemes into distribution optimization, sensitivity calibration, transformation, decomposition, and correlations exploitation, based on their mechanisms in improving data utility. We also conduct some analysis and comparison of their various concepts and principles, focusing on improvements to data utility. Finally, we outline some challenges and provide future research directions. Xinyu Yang 0001, Xuebin Ren, Wei Yu 0002 |
IEEE Trans. Big Data | 3 |
| 2021 | Latent Dirichlet Allocation Model Training With Differential PrivacyabstractLatent Dirichlet Allocation (LDA) is a popular topic modeling technique for hidden semantic discovery of text data and serves as a fundamental tool for text analysis in various applications. However, the LDA model as well as the training process of LDA may expose the text information in the training data, thus bringing significant privacy concerns. To address the privacy issue in LDA, we systematically investigate the privacy protection of the main-stream LDA training algorithm based on Collapsed Gibbs Sampling (CGS) and propose several differentially private LDA algorithms for typical training scenarios. In particular, we present the first theoretical analysis on the inherent differential privacy guarantee of CGS based LDA training and further propose a centralized privacy-preserving algorithm (HDP-LDA) that can prevent data inference from the intermediate statistics in the CGS training. Also, we propose a locally private LDA training algorithm (LP-LDA) on crowdsourced data to provide local differential privacy for individual data contributors. Furthermore, we extend LP-LDA to an online version as OLP-LDA to achieve LDA training on locally private mini-batches in a streaming setting. Extensive analysis and experiment results validate both the effectiveness and efficiency of our proposed privacy-preserving LDA training algorithms. Fangyuan Zhao, Xuebin Ren, Shusen Yang, Peng Zhao 0001, Xinyu Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | MPCSM: Microservice Placement for Edge-Cloud Collaborative Smart ManufacturingabstractLatency-aware service placement is promising in reducing the overall service response latency of proliferating edge-cloud collaborative smart manufacturing systems. However, intuitive latency estimators used by existing service placement approaches cannot accurately depict the nonlinear end-to-end (E2E) latency of multihop microservices with complex dependencies, which is severely hindering the effectiveness of latency-aware service placement. To address this issue, in this article, we present a microservice placement mechanism for edge-cloud collaborative smart manufacturing (MPCSM), where a microservice placement algorithm latency-aware edge-cloud collaborative placement supported by an accurate data-driven E2E latency estimation method is proposed. We build a real-world collaborative prototype, and conduct a case study on semiconductor manufacturing to elaborate the construction of our latency estimator. Results of extensive experiments demonstrate that the error of our E2E latency estimator is up to 10× less than that of existing ones, and the overall service latency with MPCSM is up to 10× less than that with existing service placement approaches. Cong Zhao 0001, Shusen Yang, Xuebin Ren, Luhui Wang, Peng Zhao 0001, Xinyu Yang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Privacy-preserving Crowd-guided AI Decision-making in Ethical DilemmasabstractWith the rapid development of artificial intelligence (AI), ethical issues surrounding AI have attracted increasing attention. In particular, autonomous vehicles may face moral dilemmas in accident scenarios, such as staying the course resulting in hurting pedestrians or swerving leading to hurting passengers. To investigate such ethical dilemmas, recent studies have adopted preference aggregation, in which each voter expresses her/his preferences over decisions for the possible ethical dilemma scenarios, and a centralized system aggregates these preferences to obtain the winning decision. Although a useful methodology for building ethical AI systems, such an approach can potentially violate the privacy of voters since moral preferences are sensitive information and their disclosure can be exploited by malicious parties resulting in negative consequences. In this paper, we report a first-of-its-kind privacy-preserving crowd-guided AI decision-making approach in ethical dilemmas. We adopt the formal and popular notion of differential privacy to quantify privacy, and consider four granularities of privacy protection by taking voter-/record-level privacy protection and centralized/distributed perturbation into account, resulting in four approaches VLCP, RLCP, VLDP, and RLDP, respectively. Moreover, we propose different algorithms to achieve these privacy protection granularities, while retaining the accuracy of the learned moral preference model. Specifically, VLCP and RLCP are implemented with the data aggregator setting a universal privacy parameter and perturbing the averaged moral preference to protect the privacy of voters' data. VLDP and RLDP are implemented in such a way that each voter perturbs her/his local moral preference with a personalized privacy parameter. Extensive experiments based on both synthetic data and real-world data of voters' moral decisions demonstrate that the proposed approaches achieve high accuracy of preference aggregation while protecting individual voter's privacy. Jun Zhao 0007, Han Yu 0001, Xinyu Yang 0001, Xuebin Ren, Shuyu Shi |
CIKM | 6 |
| 2019 | On Privacy Protection of Latent Dirichlet Allocation Model TrainingabstractLatent Dirichlet Allocation (LDA) is a popular topic modeling technique for discovery of hidden semantic architecture of text datasets, and plays a fundamental role in many machine learning applications. However, like many other machine learning algorithms, the process of training a LDA model may leak the sensitive information of the training datasets and bring significant privacy risks. To mitigate the privacy issues in LDA, we focus on studying privacy-preserving algorithms of LDA model training in this paper. In particular, we first develop a privacy monitoring algorithm to investigate the privacy guarantee obtained from the inherent randomness of the Collapsed Gibbs Sampling (CGS) process in a typical LDA training algorithm on centralized curated datasets. Then, we further propose a locally private LDA training algorithm on crowdsourced data to provide local differential privacy for individual data contributors. The experimental results on real-world datasets demonstrate the effectiveness of our proposed algorithms. Fangyuan Zhao, Xuebin Ren, Shusen Yang, Xinyu Yang 0001 |
IJCAI | 2 |
| 2019 | Adaptive Differentially Private Data Stream Publishing in Spatio-temporal Monitoring of IoTabstractSpatio-temporal monitoring of the Internet of Things (IoT) has enabled the development and proliferation of third-party computing services by extensively exploiting the massive amount of sensing data. In particular, continuously generated data stream are monitored in real-time and exploited to facilitate people's daily lives, such as traffic monitoring and epidemic prevention. In its simplest way of deployment, the direct publishing of various streams could seriously compromise the privacy of participating users. Hence, a more sophisticated scheme is needed to regulate the privately publishing of data streams, which may possibly require control to be applied dynamically. However, most existing solutions are non-adaptive to dynamic changes of the streams due to constraints of predefined parameters, thus are vulnerable to low data utility. In this paper, we present AdaPub, a data-adaptive framework for infinite multidimensional stream real-time publishing with ω-event differential privacy while ensuring high data utility. Without predefining the parameters, AdaPub could learn and update the parameters that reflect the spatio-temporal correlations of the stream in a data-adaptive manner. Specifically, we propose two modules DimParti and AdaCluster which are seamlessly incorporated into AdaPub to simultaneously learn dimension correlations and time correlations in a data-adaptive way, thus greatly improving the data utility of the sanitized streams. Extensive experiments on real-world datasets demonstrate that our solution substantially outperforms state-of-the-art solutions with much lower errors while achieving strong privacy guarantees. Xinyu Yang 0001, Xuebin Ren, Jun Zhao 0007, Kwok-Yan Lam |
IPCCC | 3 |
| 2019 | Differentially Private Event Sequences over Infinite Streams with Relaxed Privacy Guarantee
Xuebin Ren, Xianghua Yao, Chia-Mu Yu, Wei Yu 0002, Xinyu Yang 0001 |
WASA | 1 |
| 2019 | Impact of Prior Knowledge and Data Correlation on Privacy Leakage: A Unified AnalysisabstractIt has been widely understood that differential privacy can guarantee rigorous privacy against adversaries with arbitrary prior knowledge. However, recent studies demonstrate that this may not be true for correlated data, and indicate that three factors could influence privacy leakage: the data correlation pattern, prior knowledge of adversaries, and sensitivity of the query function. This poses a fundamental problem: what is the mathematical relationship between the three factors and privacy leakage? In this paper, we present a unified analysis of this problem. A new privacy definition, named prior differential privacy (PDP), is proposed to evaluate privacy leakage considering the exact prior knowledge possessed by the adversary. We use two models, the weighted hierarchical graph and the multivariate Gaussian model, to analyze discrete and continuous data, respectively. We demonstrate that positive, negative, and hybrid correlations have distinct impacts on privacy leakage. Considering general correlations, a closed-form expression of privacy leakage is derived for continuous data, and a chain rule is presented for discrete data. Our results are valid for general linear queries, including count, sum, mean, and histogram. Numerical experiments are presented to verify our theoretical analysis. Yanan Li 0004, Xuebin Ren, Shusen Yang, Xinyu Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | LoPub: High-Dimensional Crowdsourced Data Publication With Local Differential PrivacyabstractHigh-dimensional crowdsourced data collected from numerous users produces rich knowledge about our society; however, it also brings unprecedented privacy threats to the participants. Local differential privacy (LDP), a variant of differential privacy, is recently proposed as a state-of-the-art privacy notion. Unfortunately, achieving LDP on high-dimensional crowdsourced data publication raises great challenges in terms of both computational efficiency and data utility. To this end, based on the expectation maximization (EM) algorithm and Lasso regression, we first propose efficient multi-dimensional joint distribution estimation algorithms with LDP. Then, we develop a local differentially private high-dimensional data publication algorithm (LoPub) by taking advantage of our distribution estimation techniques. In particular, correlations among multiple attributes are identified to reduce the dimensionality of crowdsourced data, thus speeding up the distribution learning process and achieving high data utility. Extensive experiments on real-world datasets demonstrate that our multivariate distribution estimation scheme significantly outperforms existing estimation schemes in terms of both communication overhead and estimation speed. Moreover, LoPub can keep, on average, 80% and 60% accuracy over the released datasets in terms of support vector machine and random forest classification, respectively. Xuebin Ren, Chia-Mu Yu, Weiren Yu, Shusen Yang, Xinyu Yang 0001, Julie A. McCann, Philip S. Yu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Copula-Based Multi-Dimensional Crowdsourced Data Synthesis and Release with Local PrivacyabstractVarious paradigms, based on differential privacy, have been proposed to release a privacy-preserving dataset with statistical approximation. Nonetheless, most existing schemes are limited when facing highly correlated attributes, and cannot prevent privacy threats from untrusted servers. In this paper, we propose a novel Copula- based scheme to efficiently synthesize and release multi-dimensional crowdsourced data with local differential privacy. In our scheme, each participant's (or user's) data is locally transformed into bit strings based on a randomized response technique, which guarantees a participant's privacy on the participant (user) side. Then, Copula theory is leveraged to synthesize multi-dimensional crowdsourced data based on univariate marginal distribution and attribute dependence. Univariate marginal distribution is estimated by the Lasso-based regression algorithm from the aggregated privacy- preserving bit strings. Dependencies among attributes are modeled as multivariate Gaussian Copula, of which parameter is estimated by Pearson correlation coefficients. We conduct experiments to validate the effectiveness of our scheme. Our experimental results demonstrate that our scheme is effective for the release of multi-dimensional data with local differential privacy guaranteed to distributed participants. Xinyu Yang 0001, Xuebin Ren, Wei Yu 0002 |
GLOBECOM | 3 |
| 2016 | On Binary Decomposition Based Privacy-Preserving Aggregation Schemes in Real-Time Monitoring SystemsabstractIn real-time monitoring systems, fine-grained measurements would pose great privacy threats to the participants as real-time measurements could disclose accurate people-centric activities. Differential privacy has been proposed to formalize and guide the design of privacy-preserving schemes. Nonetheless, due to the correlations and high fluctuations in time-series data, it is hard to achieve an effective privacy and utility tradeoff by differential privacy mechanisms. To address this issue, in this paper, we first proposed novel multi-dimensional decomposition based schemes to compress the noise and enhance the utility in differential privacy. The key idea is to decompose the measurements into multi-dimensional records and to achieve differential privacy in bounded dimensions so that the error caused by unbounded measurements can be significantly reduced. We then extended our developed scheme and developed a binary decomposition scheme for privacy-preserving time-series aggregation in real-time monitoring systems. Through a combination of extensive theoretical analysis and experiments, our data shows that our proposed schemes can effectively improve usability while achieving the same level of differential privacy than existing schemes. Xinyu Yang 0001, Xuebin Ren, Jie Lin 0002, Wei Yu 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | On binary decomposition based privacy-preserving aggregation schemes in real-time monitoring systemsabstractReal-time monitoring systems can introduce numerous benefits to the participants in terms of performing data mining and analysis. Nonetheless, due to the correlations in time-series data, it is hard to achieve an effective privacy and utility tradeoff through a normal differential privacy mechanism. To address this issue, we propose novel multi-dimensional decomposition based schemes, which can greatly improve the utility in differential privacy. After extending the developed scheme, we then develop a binary decomposition scheme for time-series aggregation in real-time monitoring systems. Through both extensive theoretical analysis and experiments, our data shows that our proposed schemes can effectively improve usability while achieving the same level of differential privacy than existing schemes. Xuebin Ren, Xinyu Yang 0001, Jie Lin 0002, Wei Yu 0002 |
ICC | 1 |
| 2015 | A novel temporal perturbation based privacy-preserving scheme for real-time monitoring systems
Xinyu Yang 0001, Xuebin Ren, Shusen Yang, Julie A. McCann |
Comput. Networks | 2 |
| 2013 | On Scaling Perturbation Based Privacy-Preserving Schemes in Smart Metering SystemsabstractThe smart grid poses great concern about the exposure of consumers' privacy as the fine-grained measurements in the smart metering system can expose consumer's privacy through the disclosure of accurate load profiles of home energy usage. To address this issue, in this paper we propose novel scaling perturbation based privacy-preserving schemes that can achieve great utility for fine-grained measurements in a privacy-friendly and cost-effective manner. Our schemes adopt the measurement-based scaling perturbation to hide original measurements with low cost. Through a combination of both extensive theoretical analysis and experiments, our results show that the proposed schemes can preserve consumers' privacy through fine-grained measurements and achieve a better utility-privacy tradeoff in comparison with the existing schemes. Xuebin Ren, Xinyu Yang 0001, Jie Lin 0002, Qingyu Yang 0003, Wei Yu 0002 |
ICCCN | 1 |