EDBT 2026 Demo / reviewers in the wild / expert
Peng Sun 0003
dblp:88/619-3
· DBLP profile ↗
58ranked-venue papers
10as first author
51since 2021 · last 2026
0000-0001-6221-8142ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 30 · 8 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Security and privacy · 7 · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SatCLA: A Collaborative LLM-Driven Framework for Annotation of LEO Satellite ImageryabstractReal-world applications such as on-orbit disaster assessment, precision agriculture, and maritime surveillance demand immediate semantic understanding of satellite imagery. Unfortunately, the lack of efficient, high-quality annotation for large-scale satellite imagery limits the responsiveness of low Earth orbit (LEO) constellations in dynamic observation scenarios. We propose SatCLA, a lightweight framework for collaborative LLM-driven annotation of LEO satellite imagery under resource-constrained edge environments. SatCLA enables lightweight large language models (LLMs) to perform zero-shot semantic labeling onboard by retrieving relevant visual-text features and generating structured prompts to guide annotation. However, challenges such as label inconsistency, LLM hallucination, and limited semantic generalization arise under constrained compute and data settings. To address these, SatCLA introduces three key designs: (1) a multimodal retrieval pipeline for context-aware prompting, (2) a structured prompt template to guide task-specific semantic constraints, and (3) an inter-satellite re-annotation mechanism that allows peer satellites to collaboratively verify and refine low-confidence outputs. We conduct experiments within an emulated heterogeneous LEO satellite-cluster testbed and on multiple remote sensing benchmarks, demonstrating that SatCLA achieves significant gains in annotation accuracy and consistency compared to baseline methods, while maintaining low latency and high throughput. Guogen Zeng, Juan Luo, Peng Sun 0003, Anping Liu |
ICMR | 3 |
| 2026 | THUS: A Two-Phase Cross-Platform Hybrid User Recruitment Strategy in Mobile CrowdsensingabstractIn recent years, the mobile crowdsensing (MCS) paradigm has enabled a diverse array of emerging sensing applications by harnessing the collective efforts of ubiquitous mobile users, who collaborate to carry out specific sensing tasks using smart devices. However, the majority of existing works concentrate on a single MCS platform, which struggles to accommodate diverse service requirements. Moreover, these existing researches either consider opportunistic users (OUs) or participatory users (PUs) for task execution, which leads to low task coverage or high recruitment costs, while reducing the sensing quality of tasks. Therefore, in this paper, we introduce a multi-platform scenario where OUs and PUs are combined to complement each other. Then, we formulate a multi-platform hybrid user recruitment (MPHUR) problem within the limited platform budget and user time budget and decompose it into two NP-hard subproblems. To maximize the total sensing quality of tasks, we propose a Two-phase cross-platform Hybrid User recruitment Strategy called THUS. In the first phase, we present a greedy-based opportunistic user recruitment algorithm to match the user-task pair iteratively with maximum sensing quality according to the shortage degree of PUs. In the second phase, the MCS platforms assign PUs to complete the tasks that OUs fail to cover based on their residual budget. We propose a multi-task minimum-cost flow algorithm to recruit PUs for the remaining tasks. The extensive experiments are conducted on two real-world datasets to demonstrate the effectiveness of our proposed THUS. Honglong Chen, Zhishuai Li, Ning Chen 0012, Peng Sun 0003, Liantao Wu |
IEEE Internet Things J. | 6 |
| 2026 | Robust Client-Server Watermarking for Split Federated LearningabstractSplit Federated Learning (SFL) is renowned for its privacy-preserving nature and low computational overhead among decentralized machine learning paradigms. In this framework, clients employ lightweight models to process private data locally and transmit intermediate outputs to a powerful server for further computation. However, SFL is a double-edged sword: while it enables edge computing and enhances privacy, it also introduces intellectual property ambiguity as both clients and the server jointly contribute to training. Existing watermarking techniques fail to protect both sides since no single participant possesses the complete model. To address this, we propose RISE, a Robust model Intellectual property protection scheme using client-Server watermark Embedding for SFL. Specifically, RISE adopts an asymmetric client-server watermarking design: the server embeds feature-based watermarks through a loss regularization term, while clients embed backdoor-based watermarks by injecting predefined trigger samples into private datasets. This co-embedding strategy enables both clients and the server to verify model ownership. Experimental results on standard datasets and multiple network architectures show that RISE achieves over $95\%$ watermark detection rate ($p-value \lt 0.03$) across most settings. It exhibits no mutual interference between client- and server-side watermarks and remains robust against common removal attacks. Jiaxiong Tang, Zhengchunmin Dai, Liantao Wu, Peng Sun 0003, Honglong Chen |
IEEE Internet Things J. | 4 |
| 2026 | Unifying Gradient Leakage Attacks Against Privacy-Protected Federated Learning in IoT NetworksabstractFederated learning (FL) is a transformative paradigm for the Internet of Things (IoT), enabling decentralized model training across distributed IoT devices while reducing reliance on centralized data collection. Crucially, FL cuts communication overhead, an essential benefit in bandwidth-limited IoT environments. However, repeated gradient exchanges between edge clients (e.g., sensors, mobile devices) and the central server expose vulnerabilities to gradient leakage attacks (GLAs), allowing adversaries to reconstruct private training data from shared gradients. While various gradient protection strategies, such as differential privacy, sparsification, and clipping, have been introduced to mitigate this risk, most existing GLAs are designed for specific protection schemes and fail under heterogeneous deployments. In this work, we propose a unified GLA framework that targets diverse gradient protection techniques in FL systems. Our approach tackles two core challenges: (i) aligning protected gradients with their raw counterparts to enable robust feature extraction, and (ii) identifying critical features for efficient and accurate data reconstruction.We introduce a Taylor-based gradient approximation method for alignment and design a feature reconstructor that enhances both performance and computational efficiency. Extensive experiments across various FL scenarios demonstrate the framework’s superior reconstruction capability under different protection schemes, emphasizing the need for robust privacy-preserving mechanisms in IoT networks. Hui Zhou 0014, Zheng Qin 0001, Peng Sun 0003, Yipeng Zou, Xiaoshuai Wu |
IEEE Internet Things J. | 3 |
| 2026 | MATE: A D2D-Enhanced Multi-Bitrate Video Caching Strategy for Cloud-Edge-Device Collaborative NetworksabstractEdge caching alleviates backhaul pressure and enhances video service quality by deploying video content near user devices. However, the limited storage capacity of edge servers struggles to cope with the exponential growth of video data, challenging the delivery of high-quality video services. While both Device-to-Device (D2D) caching and multi-bitrate video technology are promising solutions to relieve the pressure on edge servers, existing research suffers from a key limitation: studies on multi-bitrate caching are predominantly focused on the edge layer, while D2D caching is often limited to single-bitrate scenarios. This isolation neglects the significant benefits of integrating D2D caching with multi-bitrate technology and fails to develop a cross-layer caching strategy for multi-bitrate videos. To address this limitation, we propose a D2D-enhanced Multi-bitrate video cAching straTEgy (MATE) for cloud-edge-device collaborative networks. We formulate a joint service latency and caching replacement cost optimization problem, which can be modeled as a mixed-integer programming problem. To overcome the coupling between caching strategies at the edge layer and device layer, we employ an alternating iterative optimization approach to decouple the original problem into two subproblems. We design an edge-device double-layer joint caching strategy, i.e., a device-layer caching strategy based on greedy algorithm and Lagrange multipliers, and an edge-layer caching strategy based on multi-agent twin delayed deep deterministic policy gradient algorithm. Extensive simulations are conducted to demonstrate the effectiveness of the proposed MATE. Honglong Chen, Xinglong Fan, Zhichen Ni, Liantao Wu, Peng Sun 0003, Weifeng Liu 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | A Socially Optimal Marketplace for Splittable Task Offloading in Multi-User Multi-Server Edge Computing NetworksabstractMobile users can offload their tasks to adjacent edge servers to enhance service quality. These servers require suitable reimbursements to cover the operational and energy consumption costs incurred while assisting with offloaded tasks. Although previous studies have examined market mechanisms for multiple users offloading tasks to multiple servers, most of them have not investigated the market mechanism for splittable task offloading, where tasks can be divided into multiple subtasks and offloaded to multiple servers. In this work, we propose a novel edge computing marketplace that focuses on splittable task offloading in multi-user multi-server scenarios with the aim of maximizing social welfare. Designing such a marketplace presents several challenges. First, the problem of task and computing resource division introduced in this context results in a complex solution space, and the division decisions are interdependent. Second, the users and edge servers have conflicting objectives and hidden utility/cost information. To overcome these challenges and achieve socially optimal market operation, we devise an Iterative DoublE Auction (IDEA) mechanism.IDEAemploys a broker to facilitate the interactions between users and edge servers and induces truthful reporting of hidden information through iterative updates to the allocation and pricing rules. Rigorous theoretical analysis and extensive simulations demonstrate the effectiveness of the proposedIDEAmechanism in achieving optimal social performance. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Honglong Chen, Juan Luo, Yong Zuo, Yang Yang 0001 |
IEEE Trans. Netw. | 2 |
| 2026 | Cooperative and Competitive Pricing in Collaborative Edge ComputingabstractA user with limited computation resources can address his delay-sensitive and computation-intensive tasks through task offloading to nearby edge servers, by purchasing both network and computation resources from profitseeking providers. We identify a substitutability property of computation and network resources for realizing the delay requirement. That is, to reduce task delay, the user can purchase more network resources to reduce transmission delay or more computation resources to reduce computation delay. This property significantly affects the user's purchase behavior and leads to strategic interactions between the computation service provider (CSP) and the network service provider (NSP), which have not been systematically studied yet. To this end, we formulate a two-stage Stackelberg game. In Stage I, one CSP and one NSP set their prices. In Stage II, each user decides offloading ratio and the amount of resources to purchase. By deriving the closed-form solutions in Stage II, we analytically conclude that the substitutability affects the user's decision through the network price to computation price ratio. We then incorporate the solution in Stage II into Stage I and analyze the service providers' pricing under two market structures. In the cooperative setting, where two service providers are integrated and jointly maximize their total profit, they would flexibly adjust the price ratio based on computation and network costs. In the competitive setting, where they are separate firms and aim to maximize their own profit, we formulate a pricing game and characterize a counter-intuitive equilibrium: the service providers would set high prices instead of low prices. Experimental results show that users benefit from service providers' competitive interactions. Guocheng Liao, Peng Sun 0003, Qian Ma 0002, Jianguo Chen 0001, Xu Chen 0004 |
IEEE Trans. Serv. Comput. | 2 |
| 2026 | Pricing and Trading of Data Options on Data-as-a-Service PlatformsabstractIn current Data-as-a-Service (DaaS) platforms, data marketplaces are the primary mechanism by which data owners deliver data products and services to users. However, data owners incur substantial costs to collect or produce data before selling it, exposing them to economic risks due to market price fluctuations. Options trading, used in traditional commodity markets, can alleviate sellers' financial burdens with forward contracts at fixed option prices and pre-paid option premiums. To our knowledge, the application of option trading in data marketplaces has not been thoroughly studied or implemented. In this work, we advocate for the first data options marketplace on a DaaS platform that mitigates price fluctuation risks and reduces market entry barriers. Designing such a marketplace involves several key challenges, including the easy replication of data and the difficulty in assessing data quality before production. To address these challenges, we first design a data quality prediction method based on sellers' reputation. Following this, we model and quantify the competitive relationships among buyers. On this basis, we model the interactions between sellers and buyers in the data options market as a two-stage Stackelberg game, focusing on maximizing sellers' profit. We formally derive the perfect subgame equilibrium for option trading and derive each seller's optimal pricing strategy. Numerical experiments demonstrate the superiority of data options trading over conventional data trading methods in DaaS platforms. Cheng Zhang 0035, Yang Xu 0013, Runyu Kang, Hangfan Li, Shihao Xiao, Peng Sun 0003 |
IEEE Trans. Serv. Comput. | 7 |
| 2026 | Broker-Assisted Computation Offloading and Resource Pricing in MEC Networks: A Two-Stage Stackelberg Game ApproachabstractMobile Edge Computing (MEC) significantly enhances service response speeds and improves the Quality of Service (QoS) by facilitating the offloading of computation-intensive tasks from Mobile Users (MUs) to nearby Edge Servers (ESs). However, due to the inherent selfishness of involved entities, MUs may be unwilling to offload tasks without reasonable resource pricing, and ESs may lack motivation to provide computation resources without appropriate compensation. Furthermore, improving the utilization of ESs' computation resources and achieving efficient task scheduling remains a major challenge. To ad dress these issues, we introduce a profitable broker between ESs and MUs, and propose TORP, a Two-stage Stackelberg Game based Computation Offloading and Resource Pricing mechanism, to maximize the utility of each entity. Specifically, we model the interactions among three entities (i.e., the broker, MUs, and ESs) as a two-stage Stackelberg game, where the interactions between the broker and ESs is defined as Stage I, while the interactions between the broker and MUs is defined as Stage II. By using the backward induction method, we theoretically prove the Stackelberg Equilibrium (SE) for each stage of the two-stage Stackelberg, and the SE of the whole game. Then, recognizing that the optimization problem is a Mixed-Integer Nonlinear Programming (MINLP) problem, an Alternating Iteration-Based Resource Pricing and Task Offloading Algorithm (AIPOA) is proposed to obtain the optimal solution. Finally, we perform extensive simulations comparing TORP against multiple base lines. Experimental results show that TORP achieves substantial improvements, enhancing the utilities of three entities by about 2.00%-52.69% under different scenarios. Huan Zhou 0002, Deng Meng, Jianmeng Guo, Peng Sun 0003, Liang Zhao 0014, Bin Guo 0001, Zhiwen Yu 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | FLdetox: Detoxify Persistent Backdoors in Federated LearningabstractRecent advanced persistent backdoor attacks pose a serious security challenge for federated learning, allowing backdoors to be injected into the global model with lifespans of nearly a thousand rounds after the attack ceases. Existing anomaly detection-based defenses can effectively detect abnormal client models, but are limited in actively eliminating already implanted persistent backdoors. Other defenses such as norm clipping and differential privacy could shorten the lifespan of backdoors, but they would affect model performance. In this work, we propose FLdetox, a multi-layered defensive framework to detoxify persistent backdoors without relying on auxiliary datasets while guaranteeing the desirable performance of the main task. Specifically, FLdetox slows backdoor implantation speed through fine-grained clipping and weight freezing, eliminates already implanted backdoors by retraining stable weights, and hinders attackers from optimizing their attacks through decoy global models. Extensive experimental results validate that, FLdetox effectively prevents and rapidly eliminates the implantation of state-of-the-art persistent backdoors, including Neurotoxin, DBA, F3BA, and A3FL while maintaining high performance of the main task. Wenjuan Tang, Chenkai Liu, Jianan Zhao 0015, Peng Sun 0003 |
ICDCS | 5 |
| 2025 | FFCBA: Feature-based Full-target Clean-label Backdoor Attacks
Yangxu Yin, Honglong Chen, Yudong Gao, Peng Sun 0003, Liantao Wu, Zhe Li 0026, Weifeng Liu 0001 |
ACM Multimedia | 4 |
| 2025 | SoK: On Gradient Leakage in Federated Learning
Jiacheng Du, Jiahui Hu 0001, Zhibo Wang 0001, Peng Sun 0003, Neil Zhenqiang Gong, Kui Ren 0001, Chun Chen 0001 |
USENIX Security Symposium | 4 |
| 2025 | Energy-Efficient UAV-Based Data Collection 3-D Trajectory Optimization With Wireless Power Transfer for Forest MonitoringabstractForest environment monitoring is crucial for detecting and predicting natural disasters, such as forest fires. Uncrewed aerial vehicle (UAV) is frequently used to acquire environmental data in forest ecosystems for monitoring purposes. However, the complex geographical terrain of forests, limited UAV battery capacity, and power constraints of ground IoT devices pose significant challenges to data collection. In this article, we design a 3-D trajectory optimization framework for a UAV to collect forest environmental data, considering the altitude limitations of the UAV in complex forest environments and the need to recharge IoT devices using wireless power transfer (WPT) technology, aiming to extend the operational lifetime of the network. Specifically, we aim to achieve a desirable balance between the amount of data collected by the UAV and its energy consumption incurred by data collection and sensor recharging. To this end, we first formulate the trajectory design of the UAV for data collection in forest areas as a nonlinear optimization problem, aiming to maximize the amount of sensor data collected while minimizing the energy consumption of the UAV. To solve this problem, we propose a converging-trajectory design and data collection (C-TDDC) method, which includes two subalgorithms. The first is an ant colony optimization-based traveling salesman problem (ACO-TSP) algorithm to generate optimal UAV trajectories. The second is a proximal policy optimization-based reinforcement learning algorithm, which balances data collection and energy consumption for the UAV. The simulation results demonstrate that the proposed C-TDDC algorithm exhibits more stable convergence and performs better in completing data collection tasks, outperforming both the state-of-the-art algorithm and baseline schemes. Fan Li 0030, Juan Luo, Peng Sun 0003, Shuyang Teng |
IEEE Internet Things J. | 3 |
| 2025 | On-Orbit DNN Distributed Inference for Remote Sensing Images in Satellite Internet of ThingsabstractIn satellite Internet of Things (IoT), the remote sensing satellites capture images and then transmit them to a ground station through low Earth orbit (LEO) communication satellites for model inference. However, this process results in significant transmission latency and communication overhead. In response, researchers have proposed various satellite on-orbit model inference methods. Nonetheless, the limited computation capacity and memory space of a single remote sensing satellite impose processing delays when dealing with large quantities of high-resolution images, thereby making it difficult to ensure real-time service. To tackle this issue, we propose an on-orbit deep neural network (DNN) distributed inference framework for remote sensing images in satellite IoT, leveraging the availability of numerous LEO computing satellites. Designing such a framework involves two crucial questions: first, determining which LEO satellites should participate in distributed DNN inference, and second, how to partition the images among the selected LEO satellites. To address these questions, we formulate the distributed inference process as a mixed integer nonlinear optimization problem, which is known to be NP-hard. The objective is to minimize overall energy consumption while ensuring that the distributed inference is accomplished when the satellite dynamic network remains unchanged. We initially propose a dynamic optimization algorithm that derives the optimal solution with rigorous theoretical guarantees. Subsequently, to reduce computational complexity, we introduce an approximate solution based on an improved simulated annealing algorithm. We demonstrate that the approximate algorithm performs within a limited range of the optimal algorithm. Finally, we build a heterogeneous testbed based on Kubernetes and conduct extensive experiments to validate that our proposed algorithms reduce energy consumption by an average of 24.63% and 25.98% on the Faster-RCNN inference model, 47.09% and 47.51% on the RetinaNet inference model, and 53.36% and 48.08% on the Yolov5 inference model on the two datasets compared to the baselines. Shuyang Teng, Juan Luo, Peng Sun 0003, Fan Li 0030, Fengxiao Tang |
IEEE Internet Things J. | 4 |
| 2025 | Socially Optimal Mechanism Design for Relay-Assisted Asynchronous Federated LearningabstractFederated learning (FL) has been extensively applied in industrial cyber-physical systems (ICPSs) to develop powerful models for complex industrial tasks (e.g., fault diagnosis), while safeguarding industrial data confidentiality. Asynchronous federated learning (AFL) effectively mitigates the straggler issue in the synchronous paradigm by aggregating client models in a first-come-first-served manner. Proper client selection is crucial for achieving efficient model training in AFL. A widely adopted model for implementing client selection in AFL is multi-armed bandit (MAB), which models client selection as arm pulling. Existing MAB-based client selection schemes overlook practical scenarios where direct client-server communications are unfavorable or unavailable (for example, in ICPSs such as mines, where communication infrastructure is underdeveloped, direct client-server communication is often unreliable or even unfeasible). In such cases, the server needs to incentivize self-interested relays to perform arm-pulling actions, including selecting the right client and relaying the communication from the selected client to the server. This paper proposes the first framework of incentivized online client selection for AFL. The design and optimization of such a framework involve significant challenges due to the tight coupling between unknown client behavior and private relay cost. To circumvent this challenge, we adopt the dual-based method and construct a special Lagrangian function that incorporates client behavior learning and relay cost revelation, and utilize it to design a socially-optimal mechanism for the framework. Our mechanism satisfies several desirable properties, including voluntary participation, incentive compatibility, relay utilization fairness, and client participation fairness. The proposed mechanism achieves the same asymptotic performance as the state-of-the-art benchmark that requires additional information. Furthermore, our analysis reveals that more available relays bring our mechanism closer to the theoretical upper bound of social performance. Numerical results demonstrate that our proposed mechanism achieves up to 85% and 99% of the social welfare obtained by the benchmarks. Peng Sun 0003, Guocheng Liao, Jianwei Huang 0001, Xiang Li 0148, Yuwei Wang 0001, Xu Chen 0004 |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | A Triple Stealthy Backdoor: Hidden in Spatial, Frequency, and Feature DomainsabstractBackdoor attacks pose significant security risks to deep neural networks (DNNs). These attacks involve models that make intentionally incorrect (and potentially targeted) predictions on poisoned inputs containing carefully crafted triggers, while operating normally with clean inputs. Prior studies have investigated the invisibility of backdoor triggers to improve attack stealthiness. However, they primarily concentrate on achieving invisibility solely in the spatial domain, ignoring the generation of invisible triggers in the frequency and feature domains. This constraint makes the poisoned images vulnerable to detection by recent defense mechanisms. To tackle this problem, we introduce a Triple stealthy BAckdoor attack approach, termed TriBA, which simultaneously ensures the invisibility of triggers in all the spatial, frequency, and feature domains, to achieve desirable attack performance, while ensuring strong stealthiness. Specifically, we initially utilize Wavelet Transform to embed the high-frequency information from the trigger image into the clean image to ensure effective attack performance. Then, to achieve strong stealthiness across both spatial and frequency domains, we integrate Fourier Transform and Cosine Transform to blend the poisoned image and clean image in the frequency domain. Furthermore, TriBA adopts an attack strategy to make the backdoor features similar to clean features in the feature space, which guarantees trigger invisibility in the feature domain while maintaining attack effectiveness. We theoretically prove the effectiveness of this strategy. Finally, TriBA has been comprehensively evaluated on four datasets against popular image classifiers, demonstrating a marked improvement over existing state-of-the-art backdoor attacks in terms of both attack success rate and stealthiness. Yudong Gao, Honglong Chen, Peng Sun 0003, Junjian Li, Yangxu Yin, Zhibo Wang 0001, Weifeng Liu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Towards Fair Federated Learning via Unbiased Feature AggregationabstractFederated learning (FL) is a distributed machine learning framework that enables multiple clients to collaboratively train models without raw data exchange. Prior studies on FL mainly focus on optimizing learning performance, enhancing privacy preservation, and improving attack resilience. However, little work studies how to mitigate the unfairness of federated trained models while unfair models would make discriminatory decisions toward certain groups or populations (e.g., favoring males over females), leading to serious ethical concerns. Thus, it is crucial to mitigate model unfairness in FL, yet challenging as this requires centralized access to each data point's fairness-sensitive information (e.g., race, gender), which is prohibited in FL. In this work, we propose a novel fair FL framework FedUFA, where the server can aggregate clients’ learned knowledge in an unbiased manner, to obtain fair and high-usability federated trained models. Specifically, to unearth the bias in clients’ local data and account for potentially heterogeneous local models, we propose a knowledge distillation-based FL scheme, where clients’ knowledge of learned features on a public dataset is amalgamated to the server for aggregation. We train an unbiased feature mapper at the server to remove fairness-sensitive latent features and extract fair representations from clients’ submitted raw features. In particular, we design an adversarial training method to train the mapper, which involves apredictoraiming to maximize the prediction accuracy on the FL task and adiscriminatorintending to help identify fairness-sensitive features. Extensive experiments on real-world datasets demonstrate the effectiveness of FedUFA. Zeqing He, Zhibo Wang 0001, Xiaowei Dong, Peng Sun 0003, Ju Ren 0001, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Poisoning Attacks to Knowledge Distillation-Based Federated Learning Under Robust Aggregation RulesabstractFederated learning (FL) is susceptible to poisoning attacks. To defend against such threats, robust aggregation rules (AGRs) are typically deployed on the server to identify or filter clients’ potentially malicious submissions based on statistical similarity. Recently, knowledge distillation (KD) has been widely used in FL to facilitate collaborative learning among clients that have heterogeneous model architectures by aggregating and distilling architecture-independent model outputs (i.e., logits). However, the KD process introduces a novel poisoning attack surface, where adversaries can manipulate local model output logits to ruin the global model performance. To fully reveal and explore such a new security vulnerability and effectively poison the global model in the existence of robust AGRs, in this paper, we propose the first untargeted poisoning attack scheme to KD-based FL under robust AGRs, named ManipulatingKD. It manipulates compromised clients to send well-designed malicious logits during the KD process. To ensure attack effectiveness and stealthiness, ManipulatingKD models attacks as constrained optimization problems. This allows for crafting satisfactory malicious logits that are statistically similar to benign logits but can generate poisoned aggregated logits to provide deviated supervision and mislead the global model. Extensive experiments demonstrate the effectiveness of ManipulatingKD under both non-robust and robust AGRs. Particularly, under robust AGRs, the global model accuracy degradation caused by our attacks can exceed 2× that of state-of-the-art attacks. Xiaoyi Pang, Zhibo Wang 0001, Defang Liu, Jiahui Hu 0001, Peng Sun 0003, Meng Luo 0002, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Toward Defending Adversarial Patch Attacks With Mask-Reconstruction-Assisted Adversarial TrainingabstractAdversarial patch attacks pose a significant threat to deep learning models in real-world applications, such as autonomous driving, due to their physical feasibility and ease of deployment. Although several defenses exist, they often have limitations, such as requiring prior knowledge of adversarial patches, being tied to specific model architectures, or struggling to balance robustness with accuracy on clean inputs. In this work, we propose a novel Mask-Reconstruction-assisted Adversarial Training (MRAT) framework to enhance model robustness against adversarial patches while preserving accuracy. To achieve this, we first develop an adversarial example generation method that applies masks comprising randomly positioned square blocks with adversarial perturbations to clean images. This technique generates diverse adversarial examples for training, preventing the model from overfitting to a specific type of adversarial patch. To maintain accuracy on clean inputs while enhancing adversarial robustness, we incorporate a mask-reconstruction task into standard adversarial training paradigms. This task utilizes feature representations extracted from the backbone to reconstruct clean versions of adversarial examples. By jointly optimizing the reconstruction network and the classification model, the feature representations of adversarial examples retain essential features of their original images, leading to robust feature extraction and high classification accuracy. Extensive experimental results in both digital and physical domains demonstrate that MRAT significantly enhances the model’s defensive performance against adversarial patch attacks while maintaining high accuracy on clean inputs. Our code is available at: https://github.com/wanggroupAI/AI_Security/tree/main/Adversarial%20Example/MRAT. Hongshan Yang, Zhibo Wang 0001, Peng Sun 0003, Zhixuan Chu, Feng Lin 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | H-STEP: Heuristic Stable Edge Service Entity Placement for Mobile Virtual Reality SystemsabstractVirtual reality (VR) technology, as a latency-sensitive application, can achieve real-time response to enhance the user’s quality of experience (QoE) on edge devices. However, edge servers, unlike internally managed cloud servers, are prone to hardware failures, software abnormalities, and network attacks. Most prior studies have focused on reducing service delay and improving user coverage through service entity (SE) placement, often neglecting the critical impact of edge server malfunctions on user QoE. In this work, we design a stable service entity placement framework that connects users on faulty servers to collaborative edge servers, ensuring seamless task completion. This framework presents two primary challenges: determining the grouping of collaborative edge services and the placement of SEs. To address these challenges, we introduce a heuristic stable service entity placement (H-STEP) scheme. This scheme first determines the grouping of collaborative edge servers using an iterative search algorithm and then places SEs on suitable edge servers via a fast non-dominated sorting genetic placement algorithm. This approach balances stability benefits with total cost, enhancing the system’s economic benefits. We theoretically analyze the performance of H-STEP and derive the performance gap between H-STEP and the optimal scheme. Extensive real-data-driven simulations demonstrate that H-STEP’s performance closely approximates that of the optimal scheme and surpasses existing schemes. Xuejian Chi, Honglong Chen, Zhichen Ni, Peng Sun 0003, Dongxiao Yu |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | An Incentive Framework for Task Offloading in Edge Computing Marketplaces Under Price CompetitionabstractTo efficiently execute tasks, computation resource requesters (CRRs) with limited resources can offload their tasks to nearby computation resource providers (CRPs) with spare computing capacity. These CRPs require appropriate incentives to compensate for their incurred costs when helping process the offloaded tasks. Although several mechanisms have been designed to incentivize CRPs, none of them have investigated the incentive mechanism considering price-setting and price-taking CRPs simultaneously. In this work, we propose an incentive framework for task offloading in the edge computing marketplace that includes both price-setting and price-taking CRPs. We model the CRR's interactions with both types of CRPs as a three-stage Stackelberg game to maximize the profit for both the CRR and CRPs. We prove the existence of a unique subgame perfect equilibrium (SPE) of the formulated game and further develop iterative algorithms for the CRR and price-setting CRPs to achieve the equilibrium. Through the designed algorithms, each CRP does not require complete information about the CRR and other CRPs. Extensive simulations demonstrate that offloading tasks to both price-setting and price-taking CRPs achieves higher profits for the CRR and price-setting CRPs compared to offloading tasks solely to price-setting CRPs. Additionally, the obtained SPE can achieve near-optimal social welfare. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Xiaoyi Pang, Jiahui Hu 0001, Honglong Chen, Yang Yang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Towards Privacy-Enhanced and Robust Clustered Federated LearningabstractClustered federated learning (CFL) leverages data distribution similarities to cluster clients, facilitating personalized model training under data heterogeneity. However, most existing CFL schemes pose potential privacy risks for clients (e.g., gradient inversion attacks) as they rely on individual gradients for clustering. This also renders them incompatible with secure aggregation mechanisms that are widely employed in federated learning for privacy protection. Moreover, CFL introduces the risk of malicious clients dominating several clusters and conducting poisoning attacks therein, thereby threatening secure model training. To address these issues, we propose ProCFL, a Privacy-Enhanced and Robust CFL framework incorporating gradient-free clustering and peer validation. Specifically, we first design a new protocol for measuring data distribution similarity among clients without using their gradient information. Then, we transform the client clustering process into a weighted set covering problem and introduce a diversity-optimized clustering algorithm to achieve near-optimal clustering results while eliminating any need for prior knowledge. Furthermore, we develop a post-hoc detection mechanism that employs peer validation to identify and discard malicious client models. Extensive experimental evaluation of ProCFL validates its superior model robustness and accuracy performance compared to existing schemes. Yang Xu 0013, Yunlin Tan, Cheng Zhang 0035, Peng Sun 0003, Yibang Zhang, Ju Ren 0001, Hongbo Jiang 0001, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | A Dual Stealthy Backdoor: From Both Spatial and Frequency PerspectivesabstractBackdoor attacks pose serious security threats to deep neural networks (DNNs). Backdoored models make arbitrarily (targeted) incorrect predictions on inputs containing well-designed triggers, while behaving normally on clean inputs. Prior researches have explored the invisibility of backdoor triggers to enhance attack stealthiness. However, most of them only focus on the invisibility in the spatial domain, neglecting the generation of invisible triggers in the frequency domain. This limitation renders the generated poisoned images easily detectable by recent defense methods. To address this issue, we propose a DUal stealthy BAckdoor attack method named DUBA, which simultaneously considers the invisibility of triggers in both the spatial and frequency domains, to achieve desirable attack performance, while ensuring strong stealthiness. Specifically, we first use Wavelet Transform to embed the high-frequency information of the trigger image into the clean image to ensure attack effectiveness. Then, to attain strong stealthiness, we incorporate Fourier Transform and Cosine Transform to mix the poisoned image and clean image in the frequency domain. Moreover, DUBA adopts a novel attack strategy, training the model with weak triggers and attacking with strong triggers to further enhance attack performance and stealthiness. DUBA is evaluated extensively on four datasets against popular image classifiers, showing significant superiority over state-of-the-art backdoor attacks in attack success rate and stealthiness. Yudong Gao, Honglong Chen, Peng Sun 0003, Junjian Li, Anqing Zhang, Zhibo Wang 0001, Weifeng Liu 0001 |
AAAI | 3 |
| 2024 | Byzantine-robust Decentralized Federated Learning via Dual-domain Clustering and Trust BootstrappingabstractDecentralized federated learning (DFL) facilitates collaborative model training across multiple connected clients without a central coordination server, thereby avoiding the single point of failure in traditional centralized federated learning (CFL). However, DFL exhibits increased susceptibility to Byzantine attacks owing to the lack of a responsible central server. Furthermore, a benign client in DFL may be dominated by Byzantine clients (more than half of its neighbors are malicious), posing significant challenges for robust model training. In this work, we propose DFL-Dual, a novel Byzantine-robust DFL method through dual-domain client clustering and trust bootstrapping. Specifically, we first propose to leverage both data-domain and model-domain distance metrics to identify client discrepancies. Then, we design a trust evaluation mechanism centered on benign clients, which enables them to evaluate their neighbors. Building upon the dual-domain distance metric and trust evaluation mechanism, we further develop a two-stage clustering and trust bootstrapping technique to exclude Byzantine clients from local model aggregation. We extensively evaluate the proposed DFL-Dual method through rigorous experimentation, demonstrating its remarkable performance superiority over existing robust CFL and DFL schemes. Peng Sun 0003, Zhibo Wang 0001, Bo Liu 0034 |
CVPR | 1 |
| 2024 | An Adaptive Model Difference Clipping Method for Differentially Private Federated LearningabstractFederated learning (FL) is a privacy-preserving distributed machine learning framework that allows collaborative model training among multiple clients without disclosing their raw training data. Despite data staying localized in FL, clients are still susceptible to privacy threats as their private or sensitive information can be inferred from their shared model updates or gradients. To address this issue, differential privacy (DP) techniques have been widely employed in FL by introducing random noises to obfuscate clients’ shared model parameters, thus alleviating privacy leakage. However, the introduction of DP to FL often significantly compromises the accuracy of model training. In this work, we propose a novel differentially private federated learning method to strike a desirable balance between model training accuracy and client privacy. The core idea of our method is that in each round, each client first clips the model difference between their updated model and the received global model and introduces random noises to the clipped model difference, which is then uploaded to the server. In particular, we design an adaptive clipping strategy where the clipping bound is dynamically adjusted. We conduct extensive experimental evaluations of our proposed method, and the results show that it achieves superior model accuracy under the same level of privacy protection. Juan Luo, Peng Sun 0003, Bojun Jiang |
HPCC | 3 |
| 2024 | Learning from the History: Accurately and Efficiently Aggregating Geospatial Data Under Local Differential PrivacyabstractAggregating geospatial data plays a crucial role in location-based services. However, collecting such sensitive data raises concerns about location privacy leakage. Local Differential Privacy (LDP), as a de facto privacy paradigm, has been widely employed to ensure individual location privacy. Nonetheless, existing approaches for aggregating geospatial data under LDP either suffer from compromised accuracy or involve complex computations. In this work, we propose a history-aware geospatial data aggregation framework to enhance both accuracy and efficiency while guaranteeing LDP. To this end, we first investigate an efficient aggregation method, namely General Randomized Response (GRR), and find that its variance of aggregation error follows the sum of two zero-mean binomial distributions. This reveals that multiple aggregations can boost the accuracy of GRR. To obtain multiple aggregations without compromising privacy, we adopt a Markov transition model to complement current aggregations from historical ones. However, learning the Markov transition matrix on perturbed data is challenging. Accordingly, we propose a privacy-aware Markov Transition Matrix Estimation (MTME) algorithm. Finally, we introduce a truth discovery-based refinement algorithm to iteratively derive an accurate aggregated result from multiple inaccurate aggregations. We evaluate our proposed method on two real-world trajectory datasets, and thorough experiments demonstrate its superior accuracy and very low time overhead compared to competitors. Hongbo Jiang 0001, Jie Li 0058, Peng Sun 0003, Jiangchuan Liu |
ICDCS | 5 |
| 2024 | Energy-based Backdoor Defense without Task-Specific Samples and Model RetrainingabstractBackdoor defense is crucial to ensure the safety and robustness of machine learning models when under attack. However, most existing methods specialize in either the detection or removal of backdoors, but seldom both. While few works have addressed both, these methods rely on strong assumptions or entail significant overhead costs, such as the need of task-specific samples for detection and model retraining for removal. Hence, the key challenge is how to reduce overhead and relax unrealistic assumptions. In this work, we propose two Energy-Based BAckdoor defense methods, called EBBA and EBBA+, that can achieve both backdoored model detection and backdoor removal with low overhead. Our contributions are twofold: First, we offer theoretical analysis for our observation that a predefined target label is more likely to occur among the top results for various samples. Inspired by this, we develop an enhanced energy-based technique, called EBBA, to detect backdoored models without task-specific samples (i.e., samples from any tasks). Secondly, we theoretically analyze that after data corruption, the original clean label of a poisoned sample is more likely to be predicted as a top output by the model, a sharp contrast to clean samples. Accordingly, we extend EBBA to develop EBBA+, a new transferred energy approach to efficiently detect poisoned images and remove backdoors without model retraining. Extensive experiments on multiple benchmark datasets demonstrate the superior performance of our methods over baselines in both backdoor detection and removal. Notably, the proposed methods can effectively detect backdoored model and poisoned images as well as remove backdoors at the same time. Yudong Gao, Honglong Chen, Peng Sun 0003, Zhe Li 0026, Junjian Li, Huajie Shao |
ICML | 3 |
| 2024 | Towards Efficient Asynchronous Federated Learning in Heterogeneous Edge EnvironmentsabstractFederated learning (FL) is widely used in edge environments as a privacy-preserving collaborative learning paradigm. However, edge devices often have heterogeneous computation capabilities and data distributions, hampering the efficiency of co-training. Existing works develop staleness-aware semi-asynchronous FL that reduces the contribution of slow devices to the global model to mitigate their negative impacts. But this makes data on slow devices unable to be fully leveraged in global model updating, exacerbating the effects of data heterogeneity. In this paper, to cope with both system and data heterogeneity, we propose a clustering and two-stage aggregation-based Efficient Asynchronous Federated Learning (EAFL) framework, which can achieve better learning performance with higher efficiency in heterogeneous edge environments. In EAFL, we first propose a gradient similarity-based dynamic clustering mechanism to cluster devices with similar system and data characteristics together dynamically during the training process. Then, we develop a novel two-stage aggregation strategy consisting of staleness-aware semi-asynchronous intra-cluster aggregation and data size-aware synchronous inter-cluster aggregation to efficiently and comprehensively aggregate training updates across heterogeneous clusters. With that, the negative impacts of slow devices and Non-IID data can be simultaneously alleviated, thus achieving efficient collaborative learning. Extensive experiments demonstrate that EAFL is superior to state-of-the-art methods. Xiaoyi Pang, Zhibo Wang 0001, Jiahui Hu 0001, Peng Sun 0003, Kui Ren 0001 |
INFOCOM | 5 |
| 2024 | BABE: Backdoor attack with bokeh effects via latent separation suppression
Junjian Li, Honglong Chen, Yudong Gao, Shaozhong Guo, Peng Sun 0003 |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | A Profit-Maximizing Data Marketplace with Differentially Private Federated Learning under Price CompetitionabstractThe proliferation of machine learning (ML) applications has given rise to a new and popular data marketplace paradigm. These marketplaces facilitate ML model requesters in obtaining data from data owners to train their desired models. To mitigate the privacy concerns of data owners, federated learning (FL) has been introduced, enabling collaborative model training without raw data trading. Furthermore, researchers have incorporated differential privacy (DP) techniques into FL, resulting in differentially private federated learning (DPFL) to enhance privacy preservation. However, existing designs of DPFL-based data marketplaces consider a simplified but unrealistic scenario where the model requester holds dominant market power, and data owners cannot set their own prices. In this work, we propose a novel DPFL-based data marketplace that accommodates both price-taking and price-setting data owners. We model the interactions among the model requester and these two types of data owners as a three-stage Stackelberg game, focusing on maximizing the model requester's profit. We rigorously establish that the formulated game is a convex game with a unique subgame perfect equilibrium. Moreover, we devise iterative algorithms to determine the equilibrium strategies for the model requester and price-setting data owners. Notably, our algorithms allow data owners to operate without requiring complete information about the model requester or other data owners. Numerical experiments demonstrate the superiority of our proposed three-stage framework in terms of the model requester's profitability compared to scenarios where only price-taking data owners are involved. Furthermore, we reveal that price competition among price-setting data owners reduces equilibrium market prices. Peng Sun 0003, Liantao Wu, Zhibo Wang 0001, Jinfei Liu, Juan Luo, Wenqiang Jin |
Proc. ACM Manag. Data | 1 |
| 2024 | Investigating the Backdoor on DNNs Based on Recolorization and Reconstruction: From a Multi-Channel PerspectiveabstractRecently, backdoor attacks have become a serious security threat to Deep Neural Networks (DNNs). Backdoor attacks involve embedding a hidden backdoor into a DNN model, compelling it to correctly classify benign images while erroneously classifying images with backdoor triggers as the target label. However, both current backdoor attacks and defenses have their limitations. In backdoor attacks, they are either non-stealthy or vulnerable to well-designed backdoor defense strategies. As for backdoor defenses, they often rely heavily on additional assumptions (such as determined extra clean images) and are not universally applicable, which may become impractical in the face of the latest backdoor attacks. To address the above problems, in this paper, we investigate the backdoor attack and defense strategies from a multi-channel perspective. Specifically, in terms of attacks, we propose a recolorization based attack method (RC-Attack) to generate triggers in color ab channels, which is more stealthy and effective. In terms of defenses, we propose a reconstruction-based defense method (RC-Defense) to reconstruct the color AB channels and lightness channel respectively, thus making the triggers in the reconstructed images ineffective, which is a more practical solution. Extensive experiments are conducted to demonstrate the superior performance of the proposed RC-Attack in terms of effectiveness, stealthiness and defense-resistance, and also to validate the effectiveness of the proposed RC-Defense. Honglong Chen, Yudong Gao, Anqing Zhang, Peng Sun 0003, Nan Jiang 0013, Weifeng Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Call White Black: Enhanced Image-Scaling Attack in Industrial Artificial Intelligence SystemsabstractThe increasing prevalence of deep neural networks (DNNs) in industrial artificial intelligence systems (IAISs) promotes the development of industrial automation. However, the growing employment of DNNs also exposes them to various attacks. Recent studies have shown that the data preprocessing process of DNNs is vulnerable to image-scaling attack. Such attacks can craft an attack image, which looks like a given source image but becomes a different target image after being scaled to the target size. The attack images generated by existing image-scaling attacks are easily perceivable to the human visual system, significantly degrading the attack's stealthiness. In this paper, we investigate image-scaling attack from the perspective of signal processing. We unearth that the root cause of the weak deceiving effects of existing image-scaling attack images lies in the introduction of additional high-frequency signals during their construction. Thus, we propose an enhanced image-scaling attack (EIS), which employs adversarial images crafted based on the source (“clean”) images as the target images. Those adversarial images preserve the “clean” pixel information of source images, thereby significantly mitigating the emergence of additional high-frequency signals in the attack images. Specifically, we consider three realistic threat models covering deep models' training and inference phases. Correspondingly, we design three strategies tailored to generate adversarial images with vicious patterns. These patterns are subsequently integrated into the attack images, which can mislead a model with target input size after the necessary scaling operation. Extensive experiments validate the superior performance of the proposed image-scaling attack compared to the original one. Junjian Li, Honglong Chen, Peng Sun 0003, Zhibo Wang 0001, Zhichen Ni, Weifeng Liu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Does Differential Privacy Really Protect Federated Learning From Gradient Leakage Attacks?abstractFederated Learning (FL) is susceptible to the gradient leakage attack (GLA), which can recover local private training data from the shared gradients or model updates. To ensure privacy, differential privacy is applied in FL by clipping and adding noise to local gradients (i.e., Local Differential Privacy (LDP)) or the global model update (i.e., Central Differential Privacy (CDP)). However, the effectiveness of DP in defending GLAs needs to be thoroughly investigated since some works briefly verify that DP can guard FL against GLAs while others question its defense capability. In this paper, we empirically evaluate CDP and LDP on the resistance of GLAs, and pay close attention to the trade-offs between privacy and utility in FL. Our findings reveal that: 1) existing GLAs can be defended by CDP using a per-layer clipping strategy and LDP with a reasonable privacy guarantee and 2) both CDP and LDP ensure the trade-off between privacy and utility in training shallow model, but cannot guarantee this trade-off in deeper model training (e.g., ResNets). Triggered by the crucial role of clipping operation for DP, we propose an improved attack that incorporates the clipping operation into existing GLAs without requiring additional information. The experimental results show our attack can destruct the protection of CDP and weaken the effectiveness of LDP. Overall, our work validates the effectiveness as well as reveals the vulnerability of DP under GLAs. We hope this work can provide guidance on utilizing DP for defending against GLA in FL and inspire the design of future privacy-preserving FL. Jiahui Hu 0001, Jiacheng Du, Zhibo Wang 0001, Xiaoyi Pang, Peng Sun 0003, Kui Ren 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Computation Offloading in Multi-Cell Networks With Collaborative Edge-Cloud Computing: A Game Theoretic ApproachabstractWith the widespread application of 5G and the Internet of things (IoT), edge computing and cloud computing have been collaboratively utilized for task offloading and processing. However, though the massive devices (e.g., smartphones) are organized into multi-cells, most of the existing works do not explore the computation offloading for edge-cloud computing under inter-cell interference. Thus, the offloading decisions may be inappropriate as the transmission rate is overestimated. To address this issue, we propose COMEC, a novel Computation Offloading scheme in Multi-cell networks with Edge-Cloud collaboration, which could minimize the total cost in terms of delay and energy consumption. Specifically, we first formulate COMEC as an optimization problem taking into account inter-cell interference. Then, considering the offloading decisions of all users are coupled, a non-cooperative game is formulated to minimize the total cost of each user in a distributed manner. We prove that this game is a general (ordinal) potential game and possesses a pure strategy Nash equilibrium (NE). Based on the finite improvement property of the potential game, we develop the corresponding computation offloading algorithm to achieve the NE. Finally, simulation results show that the proposed scheme can achieve superior performance in overall system cost compared with other baselines. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Yanjun Li 0004, Yang Yang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Accelerating Federated Learning via Parameter Selection and Pre-Synchronization in Mobile Edge-Cloud NetworksabstractFederated learning (FL) is a distributed machine learning methodology that can achieve collaborative model training among clients without collecting their private training data. Despite the great benefits in privacy protection, FL still faces challenges like limited computation capabilities of clients (e.g., end devices) and significant communication overheads when applied to mobile edge-cloud networks. To address these issues, this paper proposes a novel three-layer FL framework with Parameter Selection and Pre-synchronization (PSPFL) to achieve fast and accurate model training in mobile edge-cloud networks. The basic idea of PSPFL is that clients can select partial model parameters for transmission. Then, base stations aggregate these model parameters cooperatively (i.e., pre-synchronization) and send the aggregated results to the server for global model update periodically. However, there is an intrinsic trade-off between parameter transmission overhead and model training loss. To strike a desirable balance between them, we investigate the optimal parameter pre-synchronization round and local training round under PSPFL. Specifically, we propose an Alternating Minimization (AM) algorithm to obtain the initial local training round and parameter pre-synchronization round. Moreover, we integrate Deep Q-network with AM (namely DQNAM) to explore and update the optimal solution. Finally, extensive simulations are conducted to evaluate the performance of the proposed method on commonly used datasets. The results show that the proposed method can reduce the sum of FL completion time and training loss by an average of 20.72% - 69.25% compared to benchmark methods. Huan Zhou 0002, Peng Sun 0003, Bin Guo 0001, Zhiwen Yu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | AdvST: Generating Unrestricted Adversarial Images via Style TransferabstractRecent years have witnessed extensive applications of Deep Neural Networks (DNNs) in various vision tasks. However, DNNs are vulnerable to adversarial images crafted by introducing perturbations into inputs to induce incorrect predictions. Unlike$L_{p}$-norm restricted adversarial attacks, many unrestricted attacks have been proposed by modifying attributes of the image (e.g., edge, color), while the critical components of the image are preserved. However, most existing unrestricted attacks easily introduce unnatural distortions, colors, stains and schemes, in the generated adversarial images. This paper proposes a novel unrestricted attack (named AdvST) to create stylized, natural-looking, and high-transferability adversarial images. The basic idea of AdvST is to embed adversarial perturbations when transferring the style from the reference image onto the original image (i.e., rendering the original image's semantic contents into the reference image's style). To further improve the image quality of generated adversarial images, we refine two kinds of reference images (i.e., photographs and artworks) based on different attractive styles and design two attacks accordingly. For photorealistic attack, we incorporate semantic information obtained from segmentation maps to improve the photo realism of adversarial images. For artistic attack, we propose integrating edge information extracted by the Laplace operator to preserve the structural integrity of the original image. Extensive experimental results validate the superior performance of AdvST in terms of adversarial image quality and black-box transferability compared to benchmark methods. Honglong Chen, Peng Sun 0003, Junjian Li, Anqing Zhang, Weifeng Liu 0001, Nan Jiang 0013 |
IEEE Trans. Multim. | 3 |
| 2024 | Shield Against Gradient Leakage Attacks: Adaptive Privacy-Preserving Federated LearningabstractFederated learning (FL) requires frequent uploading and updating of model parameters, which is naturally vulnerable to gradient leakage attacks (GLAs) that reconstruct private training data through gradients. Although some works incorporate differential privacy (DP) into FL to mitigate such privacy issues, their performance is not satisfactory since they did not notice that GLA incurs heterogeneous risks of privacy leakage (RoPL) with respect to gradients from different communication rounds and clients. In this paper, we propose an Adaptive Privacy-Preserving Federated Learning (Adp-PPFL) framework to achieve satisfactory privacy protection against GLA, while ensuring good performance in terms of model accuracy and convergence speed. Specifically, a leakage risk-aware privacy decomposition mechanism is proposed to provide adaptive privacy protection to different communication rounds and clients by dynamically allocating the privacy budget according to the quantified RoPL. In particular, we exploratively design a round-level and a client-level RoPL quantification method to measure the possible risks of GLA breaking privacy from gradients in different communication rounds and clients respectively, which only employ the limited information in general FL settings. Furthermore, to improve the FL model training performance (i.e., convergence speed and global model accuracy), we propose an adaptive privacy-preserving local training mechanism that dynamically clips the gradients and decays the noises added to the clipped gradients during the local training process. Extensive experiments show that our framework outperforms the existing differentially private FL schemes on model accuracy, convergence, and attack resistance. Jiahui Hu 0001, Zhibo Wang 0001, Yongsheng Shen, Bohan Lin, Peng Sun 0003, Xiaoyi Pang, Jian Liu 0012, Kui Ren 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | A Socially Optimal Data Marketplace With Differentially Private Federated LearningabstractFederated learning (FL) enables multiple data owners to collaboratively train machine learning (ML) models for different model requesters while keeping data localized. Thus, FL can mitigate privacy leakage in conventional data marketplaces for ML applications requiring raw data trading for centralized model training. Nevertheless, data owners involved in FL may still suffer potential privacy leakage from gradient exposure to the model requesters. In this work, we advocate a novel data marketplace with differentially private federated learning (DPFL) to reduce such threats and maximize the social welfare. Designing such a marketplace involves several challenges. First, it is difficult to determine the privacy budget that a data owner should choose for a model requester, since they have conflicting objectives and private utility/cost information. Second, each data owner sustains privacy costs from his friends’ participation in DPFL due to data correlations, which introduces a negative externality to the market. We design a social-aware iterative double auction (SARDA) mechanism to resolve these challenges and achieve socially optimal market operation. SARDA employs a broker to coordinate the interactions between data owners and model requesters and induces them to truthfully report by iteratively updating the allocation and pricing rules. Moreover, SARDA accounts for the negative externality by incorporating others’ bids to reimburse each data owner. We show that SARDA achieves the optimal social performance and creates up to$60\%$higher social welfare than the social-agnostic benchmark. Peng Sun 0003, Guocheng Liao, Xu Chen 0004, Jianwei Huang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Collaboration in Federated Learning With Differential Privacy: A Stackelberg Game AnalysisabstractAs a privacy-preserving distributed learning paradigm, federated learning (FL) enables multiple client devices to train a shared model without uploading their local data. To further enhance the privacy protection performance of FL, differential privacy (DP) has been successfully incorporated into FL systems to defend against privacy attacks from adversaries. In FL with DP, how to stimulate efficient client collaboration is vital for the FL server due to the privacy-preserving nature of DP and the heterogeneity of various costs (e.g., computation cost) of the participating clients. However, this kind of collaboration remains largely unexplored in existing works. To fill in this gap, we propose a novel analytical framework based on Stackelberg game to model the collaboration behaviors among clients and the server with reward allocation as incentive in FL with DP. We first conduct rigorous convergence analysis of FL with DP and reveal how clients’ multidimensional attributes would affect the convergence performance of FL model. Accordingly, we solve the Stackelberg game and derive the collaboration strategies for both clients and the server. We further devise an approximately optimal algorithm for the server to efficiently conduct the joint optimization of the client set selection, the number of global iterations, and the reward payment for the clients. Numerical evaluations using real-world datasets validate our theoretical analysis and corroborate the superior performance of the proposed solution. Guangjing Huang, Qiong Wu 0009, Peng Sun 0003, Qian Ma 0002, Xu Chen 0004 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Server-Initiated Federated Unlearning to Eliminate Impacts of Low-Quality DataabstractFederated unlearning (FUL) is an emerging distributed machine learning paradigm which enables the removal or unlearning of specific training data effects from trained Federated Learning (FL) models. While current studies mostly focus on client-side FUL to address the “right to be forgotten”, and ignore the server's right to remove local models from the global model, particularly when clients are trained with low-quality data. In this paper, we introduce the Server-Initiated Federated Unlearning (SIFU) algorithm, devised to eliminate low-quality data from the global model. SIFU consists of two main components: (i) Identifying low-quality data: we develop a category-based method for quantifying low-quality data for each client and filter out clients containing such data. Datasets are then divided accordingly. (ii) Unlearning low-quality data: we employ gradient ascent training to counteract the adverse effects of low-quality data on local models. To minimize any bias introduced, we concurrently perform several batches of boosting training with good-quality data. SIFU could identify and promptly eliminate the impact of low-quality data on the FL global model while still preserving the benefits of good-quality data. Finally, extensive evaluations are conducted to verify the performance of SIFU with four different kinds of datasets and models. Results show that, compared to retraining from scratch, SIFU accelerates the speed of unlearning by 15× for small datasets (i.e., MNIST and FMNIST) and 20× for large datasets (i.e., CIFAR-10 and CelebA) without any degradation in accuracy, which also outperforms the state of the arts. Pengfei Wang 0013, Heng Qi, Changjun Zhou, Fuliang Li, Yong Wang 0046, Peng Sun 0003, Qiang Zhang 0008 |
IEEE Trans. Serv. Comput. | 7 |
| 2023 | Privacy-preserving Adversarial Facial FeaturesabstractFace recognition service providers protect face privacy by extracting compact and discriminative facial features (representations) from images, and storing the facial features for real-time recognition. However, such features can still be exploited to recover the appearance of the original face by building a reconstruction network. Although sev-eral privacy-preserving methods have been proposed, the enhancement offace privacy protection is at the expense of accuracy degradation. In this paper, we propose an adver-sarial features-based face privacy protection (AdvFace) approach to generate privacy-preserving adversarial features, which can disrupt the mapping from adversarial features to facial images to defend against reconstruction attacks. To this end, we design a shadow model which simulates the attackers' behavior to capture the mapping function from facial features to images and generate adversarial la-tent noise to disrupt the mapping. The adversarial features rather than the original features are stored in the server's database to prevent leaked features from exposing facial information. Moreover, the AdvFace requires no changes to the face recognition network and can be implemented as a privacy-enhancing plugin in deployed face recognition systems. Extensive experimental results demonstrate that Adv Face outperforms the state-of-the-art face privacy-preserving methods in defending against reconstruction at-tacks while maintaining face recognition accuracy. Zhibo Wang 0001, He Wang 0005, Shuaifan Jin, Jiahui Hu 0001, Yan Wang 0002, Peng Sun 0003, Kui Ren 0001 |
CVPR | 7 |
| 2023 | Towards Transferable Targeted Adversarial ExamplesabstractTransferability of adversarial examples is critical for black-box deep learning model attacks. While most existing studies focus on enhancing the transferability of untargeted adversarial attacks, few of them studied how to generate transferable targeted adversarial examples that can mislead models into predicting a specific class. Moreover, existing transferable targeted adversarial attacks usually fail to sufficiently characterize the target class distribution, thus suffering from limited transferability. In this paper, we propose the Transferable Targeted Adversarial Attack (TTAA), which can capture the distribution information of the target class from both label-wise and feature-wise perspectives, to generate highly transferable targeted adversarial examples. To this end, we design a generative adversarial training framework consisting of a generator to produce targeted adversarial examples, and feature-label dual discriminators to distinguish the generated adversarial examples from the target class images. Specifically, we design the label discriminator to guide the adversarial examples to learn label-related distribution information about the target class. Meanwhile, we design a feature discriminator, which extracts the feature-wise information with strong cross-model consistency, to enable the adversarial examples to learn the transferable distribution information. Furthermore, we introduce the random perturbation dropping to further enhance the transferability by augmenting the diversity of adversarial examples used in the training process. Experiments demonstrate that our method achieves excellent performance on the transferability of targeted adversarial examples. The targeted fooling rate reaches 95.13% when transferred from VGG-19 to DenseNet-121, which significantly outperforms the state-of-the-art methods. Zhibo Wang 0001, Hongshan Yang, Yunhe Feng, Peng Sun 0003, Hengchang Guo, Kui Ren 0001 |
CVPR | 4 |
| 2023 | Towards Class-Balanced Privacy Preserving Heterogeneous Model AggregationabstractHeterogeneous model aggregation (HMA) is an effective paradigm that integrates on-device trained models heterogeneous in architecture and target task into a comprehensive model. Recent works adopt knowledge distillation to amalgamate the knowledge of learned features and predictions from heterogeneous on-device models to realize HMA. However, most of them ignore that the disclosure of learned features exposes on-device models to privacy attacks. Moreover, the aggregated model may suffer from the imbalanced supervision caused by the uneven distribution of amalgamated knowledge about each class and show class bias. In this article, to address these issues, we propose a response-based class-balanced heterogeneous model aggregation mechanism, called CBHMA. It can effectively achieve HMA in a privacy-preserving manner and alleviate class bias in the aggregated model. Specifically, CBHMA aggregates on-device models by using only their response information to reduce their privacy leakage risk. To mitigate the impact of imbalanced supervision, CBHMA quantitatively measures the imbalanced supervision level for each class. Based on that, CBHMA customizes fine-grained misclassification costs for each class and utilizes such costs to adjust the importance of each class (more importance to classes with weaker supervision) in the response-based HMA algorithm. Extensive experiments on two real-world datasets demonstrate the effectiveness of CBHMA. Xiaoyi Pang, Zhibo Wang 0001, Zeqing He, Peng Sun 0003, Meng Luo 0010, Ju Ren 0001, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Temporal Correlation Enhanced Multiuser Detection for Uplink Grant-Free NOMAabstractCompressed sensing (CS) has been identified as a good candidate for user detection in grant-free non-orthogonal multiple access (NOMA) by exploiting the inherent sparsity of user activity. However, most of the existing CS-based user detection schemes do not fully utilize the temporal correlation of user activity in NOMA and rely heavily on the unrealistic assumption that the number of active users is known in advance. To address these issues, we propose a temporal correlation enhanced multiuser detection scheme to achieve efficient and pragmatic multiuser detection. First, using 1-bit memory to piggyback the information on whether the active users still have data to transmit, the base station can realize that the active users in the current time slot will turn to be silent or remain active. Then, to make explicit use of the temporal correlation of active user sets, a cross validation based adaptive subspace pursuit (CVASP) algorithm is developed by utilizing the reported information on prior active users. The proposed CVASP is a highly practical algorithm that does not require any prior knowledge of the number of active users or the noise level, as the cross validation technique could properly determine the stopping condition. Extensive simulation results demonstrate that the proposed mechanism could achieve almost the same performance as compared to the existing state of art CS-based multiuser detection algorithms while eliminating the need for any prior knowledge. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Yang Yang 0001, Zhi Wang 0003 |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | FedInv: Byzantine-Robust Federated Learning by Inversing Local Model UpdatesabstractFederated learning (FL) is a privacy-preserving distributed machine learning paradigm that enables multiple clients to collaboratively train statistical models without disclosing raw training data. However, the inaccessible local training data and uninspectable local training process make FL susceptible to various Byzantine attacks (e.g., data poisoning and model poisoning attacks), aiming to manipulate the FL model training process and degrade the model performance. Most of the existing Byzantine-robust FL schemes cannot effectively defend against stealthy poisoning attacks that craft poisoned models statistically similar to benign models. Things worsen when many clients are compromised or data among clients are highly non-independent and identically distributed (non-IID). In this work, to address these issues, we propose FedInv, a novel Byzantine-robust FL framework by inversing local model updates. Specifically, in each round of local model aggregation in FedInv, the parameter server first inverses the local model updates submitted by each client to generate a corresponding dummy dataset. Then, the server identifies those dummy datasets with exceptional Wasserstein distances from others and excludes the related local model updates from model aggregation. We conduct an exhaustive experimental evaluation of FedInv. The results demonstrate that FedInv significantly outperforms the existing robust FL schemes in defending against stealthy poisoning attacks under highly non-IID data partitions. Peng Sun 0003, Keyu Jiang |
AAAI | 2 |
| 2022 | A Profit-Maximizing Model Marketplace with Differentially Private Federated LearningabstractExisting machine learning (ML) model marketplaces generally require data owners to share their raw data, leading to serious privacy concerns. Federated learning (FL) can partially alleviate this issue by enabling model training without raw data exchange. However, data owners are still susceptible to privacy leakage from gradient exposure in FL, which discourages their participation. In this work, we advocate a novel differentially private FL (DPFL)-based ML model marketplace. We focus on the broker-centric design. Specifically, the broker first incentivizes data owners to participate in model training via DPFL by offering privacy protection as per their privacy budgets and explicitly accounting for their privacy costs. Then, it conducts optimal model versioning and pricing to sell the obtained model versions to model buyers. In particular, we focus on the broker’s profit maximization, which is challenging due to the significant difficulties in the revenue characterization of model trading and the cost estimation of DPFL model training. We propose a two-layer optimization framework to address it, i.e., revenue maximization and cost minimization under model quality constraints. The latter is still challenging due to its non-convexity and integer constraints. We hence propose efficient algorithms, and their performances are both theoretically guaranteed and empirically validated. Peng Sun 0003, Xu Chen 0004, Guocheng Liao, Jianwei Huang 0001 |
INFOCOM | 1 |
| 2022 | DOT: Decentralized Offloading of Tasks in OFDMA-Based Heterogeneous Computing NetworksabstractA fundamental issue in multiaccess edge computing (MEC) is efficiently offloading multiple tasks to multiple helper nodes (MTMH), i.e., MEC servers. However, most of the existing decentralized schemes do not consider interuser interference or merely adopt time division multiple access (TDMA) as the multiple access scheme for MTMH in the heterogeneous scenario, leading to a large latency. To address these issues, we propose DOT, a novel Decentralized Offloading of Tasks scheme in orthogonal frequency division multiple access (OFDMA)-based heterogeneous MEC, to minimize the sum cost in terms of energy consumption and delay. Specifically, we first formulate DOT as an optimization problem considering the interuser interference and dynamics in communication and computation resource allocation. Then, considering the huge dimension of potential offloading decisions and conflicting objectives of different users, the total cost of each user is minimized in a distributed manner by modeling the offloading problem as a potential game. The formulated potential game is proved to be an ordinal potential game and thus admits a Nash equilibrium (NE). Further, we develop an offloading algorithm to achieve the NE by exploiting the finite improvement property. Finally, simulation results demonstrate that DOT can achieve a lower cost compared with other baselines. Liantao Wu, Zening Liu, Peng Sun 0003, Honglong Chen, Kunlun Wang 0001, Yong Zuo, Yang Yang 0001 |
IEEE Internet Things J. | 3 |
| 2022 | Towards Personalized Privacy-Preserving Incentive for Truth Discovery in Mobile Crowdsensing SystemsabstractIncentive mechanisms are essential for stimulating adequate worker participation to achieve good truth discovery performance in mobile crowdsensing (MCS) systems. However, most of existing incentive mechanisms only consider compensating workers’ sensing cost, while the cost incurred by potential privacy leakage has been largely neglected. Moreover, none of existing privacy-preserving incentive mechanisms has incorporated workers’ different privacy preferences to provide personalized payments for them. In this paper, we propose a contract-based personalized privacy-preserving incentive mechanism for truth discovery in MCS systems, named Paris-TD, which provides personalized payments for workers as a compensation for privacy cost while achieving accurate truth discovery. The basic idea is that the platform offers a set of different contracts to workers with different privacy preferences, and each worker chooses to sign a contract which specifies a privacy-preserving degree (PPD) and the corresponding payment the worker will receive if she submits perturbed data with that PPD. Specifically, we respectively design a set of optimal contracts analytically under both full and incomplete information models, which maximize the truth discovery accuracy under a given budget, while satisfying the individual rationality and incentive compatibility properties. The feasibility and effectiveness of Paris-TD are validated through experiments on both synthetic and real-world datasets. Peng Sun 0003, Zhibo Wang 0001, Liantao Wu, Yunhe Feng, Xiaoyi Pang, Hairong Qi 0001, Zhi Wang 0003 |
IEEE Trans. Mob. Comput. | 1 |
| 2021 | Micromobility in Smart Cities: A Closer Look at Shared Dockless E-Scooters via Big Social DataabstractThe micromobility is shaping first- and last-mile travels in urban areas. Recently, shared dockless electric scooters (e-scooters) have emerged as a daily alternative to driving for short-distance commuters in large cities due to the affordability, easy accessibility via an app, and zero emissions. Meanwhile, e-scooters come with challenges in city management, such as traffic rules, public safety, parking regulations, and liability issues. In this paper, we collected and investigated 5.8 million scooter-tagged tweets and 144,197 images, generated by 2.7 million users from October 2018 to March 2020, to take a closer look at shared e-scooters via crowdsourcing data analytics. We profiled e-scooter usages from spatial-temporal perspectives, explored different business roles (i.e., riders, gig workers, and ridesharing companies), examined operation patterns (e.g., injury types, and parking behaviors), and conducted sentiment analysis. To our best knowledge, this paper is the first large-scale systematic study on shared e-scooters using big social data. Yunhe Feng, Dong Zhong, Peng Sun 0003, Weijian Zheng, Qinglei Cao, Zheng Lu 0005 |
ICC | 3 |
| 2021 | Joint User Activity Identification and Channel Estimation for Grant-Free NOMA: A Spatial-Temporal Structure-Enhanced ApproachabstractExploiting the sparse nature of user activity, compressed sensing (CS) has been a powerful technique to realize efficient user detection in grant-free nonorthogonal multiple access (NOMA). However, most of the existing CS-based multiuser detection schemes merely independently incorporate the temporal correlation in frame-based transmission or spatial correlation induced by multiantenna reception, leading to unsatisfactory user detection performance. Driven by the observation in the CS theory that the signal recovery performance could be enhanced by an increased number of sparse vectors with a common support set, in this article, we propose a novel joint user activity identification and channel estimation (JUICE) framework by integrating the temporal correlation of active user sets with multiantenna reception, which could achieve superior user detection performance. Specifically, we first formulate the JUICE as a Kronecker CS (KCS) problem to model the CS measurement process, by fully extracting the spatial-temporal structure of user activity. Then, based on the mined spatial-temporal structure of user activity, an adaptive subspace pursuit algorithm is developed, i.e., spatial-temporal structure enhanced adaptive subspace pursuit (STS-ASP), which could realize efficient multiuser detection. A distinct advantage of the proposed algorithm is that it does not require any prior knowledge (e.g., the number of active users and the noise level), by adaptively acquiring the number of active users and employing the cross-validation technique to appropriately terminate the iterative procedures. Extensive experimental evaluation is conducted, and the results corroborate the superiority of the proposed framework compared with the existing CS-based multiuser detection methods. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Yang Yang 0001 |
IEEE Internet Things J. | 2 |
| 2021 | Pain-FL: Personalized Privacy-Preserving Incentive for Federated LearningabstractFederated learning (FL) is a privacy-preserving distributed machine learning framework, which involves training statistical models over a number of mobile users (i.e., workers) while keeping data localized. However, recent works have demonstrated that workers engaged in FL are still susceptible to advanced inference attacks when sharing model updates or gradients, which would discourage them from participating. Most of the existing incentive mechanisms for FL mainly account for workers’ resource cost, while the cost incurred by potential privacy leakage resulting from inference attacks has rarely been incorporated. To address these issues, in this paper, we propose a contract-based personalized privacy-preserving incentive for FL, named Pain-FL, to provide customized payments for workers with different privacy preferences as compensation for privacy leakage cost while ensuring satisfactory convergence performance of FL models. The core idea of Pain-FL is that each worker agrees on a customized contract, which specifies a kind of privacy-preserving level (PPL) and the corresponding payment, with the server in each round of FL. Then, the worker perturbs her calculated stochastic gradients to be uploaded with that PPL in exchange for that payment. In particular, we respectively derive a set of optimal contracts analytically under both complete and incomplete information models, which could optimize the convergence performance of the finally learned global model, while bearing some desired economic properties, i.e., budget feasibility, individual rationality, and incentive compatibility. An exhaustive experimental evaluation of Pain-FL is conducted, and the results corroborate its practicability and effectiveness. Peng Sun 0003, Haoxuan Che, Zhibo Wang 0001, Yuwei Wang 0001, Liantao Wu, Huajie Shao |
IEEE J. Sel. Areas Commun. | 1 |
| 2020 | Attention-over-Attention Field-Aware Factorization MachineabstractFactorization Machine (FM) has been a popular approach in supervised predictive tasks, such as click-through rate prediction and recommender systems, due to its great performance and efficiency. Recently, several variants of FM have been proposed to improve its performance. However, most of the state-of-the-art prediction algorithms neglected the field information of features, and they also failed to discriminate the importance of feature interactions due to the problem of redundant features. In this paper, we present a novel algorithm called Attention-over-Attention Field-aware Factorization Machine (AoAFFM) for better capturing the characteristics of feature interactions. Specifically, we propose the field-aware embedding layer to exploit the field information of features, and combine it with the attention-over-attention mechanism to learn both feature-level and interaction-level attention to estimate the weight of feature interactions. Experimental results show that the proposed AoAFFM improves FM and FFM with large margin, and outperforms state-of-the-art algorithms on three public benchmark datasets. Zhibo Wang 0001, Jinxin Ma, Qian Wang 0002, Ju Ren 0001, Peng Sun 0003 |
AAAI | 6 |
| 2020 | Towards Personalized Privacy-Preserving Incentive for Truth Discovery in Crowdsourced Binary-Choice Question AnsweringabstractTruth discovery is an effective tool to unearth truthful answers in crowdsourced question answering systems. Incentive mechanisms are necessary in such systems to stimulate worker participation. However, most of existing incentive mechanisms only consider compensating workers' resource cost, while the cost incurred by potential privacy leakage has been rarely incorporated. More importantly, to the best of our knowledge, how to provide personalized payments for workers with different privacy demands remains uninvestigated thus far. In this paper, we propose a contract-based personalized privacy-preserving incentive mechanism for truth discovery in crowdsourced question answering systems, named PINTION, which provides personalized payments for workers with different privacy demands as a compensation for privacy cost, while ensuring accurate truth discovery. The basic idea is that each worker chooses to sign a contract with the platform, which specifies a privacy-preserving level (PPL) and a payment, and then submits perturbed answers with that PPL in return for that payment. Specifically, we respectively design a set of optimal contracts under both complete and incomplete information models, which could maximize the truth discovery accuracy, while satisfying the budget feasibility, individual rationality and incentive compatibility properties. Experiments on both synthetic and real-world datasets validate the feasibility and effectiveness of PINTION. Peng Sun 0003, Zhibo Wang 0001, Yunhe Feng, Liantao Wu, Yanjun Li 0004, Hairong Qi 0001, Zhi Wang 0003 |
INFOCOM | 1 |
| 2020 | SCRA: Structured Compressive Random Access for Efficient Information Collection in IoTabstractIt is a fundamental issue to achieve efficient information collection in Internet of Things (IoT), where random (channel) access plays an indispensable role, especially when coordination among IoT end nodes is unachievable. Compressive sensing (CS) has been widely used in random access to facilitate energy efficient and accurate data collection. However, a joint sparsity structure, which commonly exists among signals acquired by different end nodes, has been long ignored by existing CS-based random access schemes, leading to insufficient energy efficiency and accuracy. In this article, capitalizing on this joint sparsity structure, we propose a structured compressive random access (SCRA) mechanism in order to achieve maximum energy efficiency with accuracy guarantee for data collection. Specifically, we first model the data loss induced by packet collisions during random access as an independent CS measurement process for each node, where the corresponding CS projection matrix is determined by the data loss pattern. Furthermore, in order to control the amount of data transmitted in the channel and alleviate the packet collisions, we employ the concept of sensing probability to perform random subsampling at each end node before transmission, where the optimal sensing probability is derived. Finally, we propose to jointly recover the set of original signals at all nodes based on the concept of group sparsity by formulating the data collection process as a single-measurement-vector problem in CS. The evaluation results validate the effectiveness of SCRA in utilizing the joint sparsity structure to obtain superior performance compared to the benchmark methods. Peng Sun 0003, Liantao Wu, Zhibo Wang 0001, Yunhe Feng, Zhi Wang 0003 |
IEEE Internet Things J. | 1 |
| 2020 | Toward Efficient Compressed-Sensing-Based RFID Identification: A Sparsity-Controlled ApproachabstractRadio-frequency identification (RFID) has pervasive applications in building ultralow-power ubiquitous networks, where backscatter communication during tag identification is neither reliable nor efficient. Inspired by the sparsity that only a few RFID tags communicate with the reader simultaneously, many compressed sensing (CS)-based schemes have been proposed to exploit the colliding tag responses to facilitate efficient tag identification. However, most of them suffer from huge ID search space and signature collision during the CS recovery process. To address these issues, we propose SCRIC, a novel sparsity-controlled RFID identification scheme using a random signature assignment, which achieves a faster and more robust identification performance. Specifically, to tackle identification failure caused by severe signature collision, we assign each active tag an access probability to control the sparsity and reduce signature collision. Theoretical analysis is given to prove that the signature collision probability of the proposed scheme is reduced compared with the existing random signature scheme and an optimal access probability is derived. Moreover, considering that conventional CS recovery algorithm relies heavily on the unpractical assumption that the active tag number is known precisely in advance, we integrate the cross validation (CV) into CS recovery algorithms and propose a greedy algorithm called the CV-based orthogonal matching pursuit (OMP-CV), which can reduce the tag identification false alarm rate without any prior knowledge. Extensive experimental results show that the proposed mechanism significantly outperforms the existing CS-based tag identification methods in terms of identification speed and robustness to noise. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Yang Yang 0001, Zhi Wang 0003 |
IEEE Internet Things J. | 2 |
| 2019 | Chasing Total Solar Eclipses on Twitter: Big Social Data Analytics for Once-in-a-Lifetime EventsabstractWith the popularity of social networking services, big social data analytics emerged in various applications, such as discovering trending topics, monitoring public sentiment, and identifying human mobility patterns. In this paper, we take the opportunity of The 2017 Great American Eclipse, a once-in-a-lifetime event, to look into its potential social, emotional, and human movement impacts at the national level. Specifically, we collected more than five million English eclipse- mentioning tweets in a real-time manner using Twitter Streaming APIs. Then we profiled spatio- temporal distributions of the data, extracted both hashtagged and latent topics, analyzed emotions using polarized words, emojis and emoticons, and revealed both interstate and intrastate eclipse- chasing travel patterns. Our study provides a comprehensive example of understanding big social data and its associated influence from diverse perspectives. Yunhe Feng, Zheng Lu 0005, Zhonghua Zheng, Peng Sun 0003, Wenjun Zhou 0001, Qing Cao 0001 |
GLOBECOM | 4 |
| 2019 | A 3-D Self-Calibration Method for Multiple Base Stations in Large Complex Indoor EnvironmentabstractBase stations deployed for localization are required to be location aware. Traditional techniques for calibrating base stations need to be done manually, which is time-consuming and inconvenient. To address this problem, Semidefinite Programming (SDP) -based approaches have been developed, which can achieve high precision. However, existing approaches mostly focus on the 2-dimensional (2-D) scenarios where base stations must be restricted to be of the same height. In this paper, we propose a 3-dimensional (3-D) SDP-based method to determine the positions of base stations given inaccurate even incomplete distance information and apply it to large indoor environments. The corresponding Cramer-Rao lower bounds (CRLB) are derived to ensure localization quality. We investigate the performance of this method on problems with sensing range. Simulation results verify that our method can achieve superior performance in terms of accuracy and robustness compared with traditional noncooperative localization techniques. Peng Sun 0003, Zhi Wang 0003 |
WCNC | 2 |
| 2019 | PRSS: A Prejudiced Random Sensing Strategy for Energy-Efficient Information Collection in the Internet of ThingsabstractCompressive sensing (CS) has been widely used in the Internet of Things (IoT) to achieve efficient information collection. However, existing works have mainly focused on utilizing CS to lower the sampling rate or reduce the number of transmissions, without explicitly accounting for the heterogeneity of energy consumption in IoT environments. In this paper, we propose a CS-based prejudiced random sensing strategy (PRSS) that explicitly considers the heterogeneous energy consumption of IoT sensor nodes at different locations, in order to accurately attain a desired tradeoff between the overall energy consumption and the sensing accuracy. Specifically, each sensor node participates in sensing via distributed random access based on an assigned sensing probability, which is determined by its energy consumption in sending the sensed data, data collision rate and its contribution to recovery accuracy. We employ the statistical restricted isometry property as a practical indicator of the recovery accuracy and derive a sufficiently good recovery error bound based on it. Accordingly, we devise a novel convex optimization framework to find the most energy-efficient sensing probability assignment strategy with accuracy guarantee. We evaluate the PRSS using real-world sea surface temperature data traces. Comparative simulations corroborate that the PRSS can significantly reduce energy consumption and prolong network lifetime without sacrificing sensing accuracy. Peng Sun 0003, Zhi Tian, Zhibo Wang 0001, Zhi Wang 0003 |
IEEE Internet Things J. | 1 |