VLDB 2026 Research / reviewers in the wild / expert
Siping Shi
dblp:203/1967
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-5555-417XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Personalized Federated Learning via Overlapping Coalition Formation GameabstractTo tackle the challenge of data heterogeneity in federated learning (FL), personalized FL has been proposed to maximize individual utility (model performance) by customizing personalized models for clients. Considering the significance ofindividual rationality, existing works have formulated clients' participation decisions problem ashedonicgames. However, they assume that clients can participate in only one collaborative coalition, constraining players' attempts to join multiple coalitions. Different from prior works, we approach personalized FL from the perspective of hedonicoverlapping coalition formation(OCF) games where rational clients can join multiple coalitions and generate their personalized model by weighting the local and coalition models. Nevertheless, the key challenge in analyzing the game is how to achieve a stable coalition structure where no clients would deviate from the current structure. This leads to our main question:what does a stable OCF structure look like?To address this problem, we first investigate the linear FL models for theoretical insights. Then, we design a heuristic algorithm for achieving anindividually stableOCF structure. Experimental results demonstrate the feasibility of our algorithm for both linear and non-linear models, and show that our mechanism can improve the personalized model performance by up to 19% over existing methods. Bing Luo 0002, Jiawei Jiang 0001, Siping Shi, Chuang Hu, Dazhao Cheng |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | A Privacy-preserving Edge-cloud Video Analytics System via Policy-based Frame TransformationabstractIn real-time edge-cloud video analytics systems, the edge conducts initial analytics on the video frames to a split layer of a trained neural network model. Then, it sends intermediate results to the cloud for follow-up analytics. In this article, we first show that malicious attackers can perform reconstruction attacks and attribute inference attacks on those intermediate results. We present Preva, a new P rivacy-preserving R eal-time E dge-cloud V ideo A nalytics system that defends against both attacks while respecting latency constraints. Preva first applies a lightweight, policy-based video frames transformation scheme generated on the fly by PrevaNetv2, with a lightweight backbone and an early-exit mechanism that adapts the computation resources of users. To guarantee end-to-end latency, a contextual multi-armed-bandit algorithm (BBSplit) dynamically chooses the optimal split layer, and a frame-similarity filter (PPReuse) amortizes policy generation across video frames. We present a formal privacy analysis and show that Preva can guarantee privacy leakage under reconstruction attacks and attribute inference attacks. We evaluate Preva through three video analytics applications and show that Preva outperforms existing systems by 41.6% in analytics accuracy and 52.7% in privacy leakage. Siping Shi, Chuang Hu, Dan Wang 0002 |
ACM Trans. Internet Techn. | 3 |
| 2025 | Accelerating Long Video Understanding via Compressed Scene Graph-Enabled Chain-of-Thought
Tao Ling, Siping Shi, Dan Wang 0002 |
ACM Multimedia | 2 |
| 2025 | PriFairFed: A Local Differentially Private Federated Learning Algorithm for Client-Level FairnessabstractLocal Differential Privacy (LDP) is a mechanism used to protect training privacy in Federated Learning (FL) systems, typically by introducing noise to data and local models. However, in real-world distributed edge systems, the non-independent and identically distributed nature of data means that clients in FL systems experience varying sensitivities to LDP-introduced noise. This disparity leads to fairness issues, potentially discouraging marginal clients from contributing further. In this paper, we explore how to enhance client-level performance fairness under LDP conditions. We model an FL system with LDP and formulate the problem PriFair using regularization, which assigns varied noise amplitudes to clients based on federated analytics. Additionally, we develop PriFairFed, a Tikhonov regularization-based algorithm that eliminates variable dependencies and optimizes variables alternately, while also offering a theoretical privacy guarantee. We further experimented with the algorithm on a real-world system with 20 Raspberry Pi clients, showing up to a 73.2% improvement in client-level fairness compared to existing state-of-the-art approaches, while maintaining a comparable level of privacy. Chuang Hu, Nanxi Wu, Siping Shi, Bing Luo 0002, Kanye Ye Wang, Jiawei Jiang 0001, Dazhao Cheng |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | An Efficient On-Device Federated Learning System Through the Interplay of Client Selection and Batch Size With Watermarked DataabstractFederated Learning (FL) enables edge devices to collaboratively train a global model using local data. However, the increasing prevalence of watermarks in datasets presents a new challenge to efficient FL. While watermarks assert data ownership and copyright, they introduce complexities that can lead to shortcut learning problems and mislead utility measurements for client selection. These issues are further exacerbated by batch size variations in efficient FL frameworks, ultimately undermining their time-to-accuracy performance. We introduceLotusFL, an FL system designed to address the challenges posed by watermarked datasets in efficient FL. Specifically, it tackles the increased time-to-accuracy due to erroneous client selection and the accuracy degradation observed with larger batch sizes.LotusFLfirst estimates the characteristics of watermarks through statistical estimation and then adjusts the batch size using this estimated watermark information to balance the negative impact of the watermark against device idle waiting time. Additionally, its client selection mechanism, based on historical information, avoids the misleading utility signals from watermarks. This mechanism, working in conjunction with batch size adjustment, aims to accurately predict device runtime and identify potentially valuable devices. We evaluatedLotusFLthrough a real-world deployment on 40 edge devices. Compared to state-of-the-art efficient FL frameworks,LotusFLachieves superior performance, enhancing accuracy by up to 8.2% and reducing training time by 1.97×. Tao Ling, Siping Shi, Hao Wang 0022, Chuang Hu, Dan Wang 0002 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Federated Morozov Regularization for Shortcut Learning in Privacy Preserving Learning with Watermarked Image DataabstractFederated learning is a promising privacy-preserving learning paradigm in which multiple clients can collaboratively learn a model with their image data kept local. For protecting data ownership, personalized watermarks are usually added to the image data by each client. However, the introduced watermarks can lead to a shortcut learning problem, where the learned model performs predictions over-rely on the simple watermark-related features and represents a low accuracy on real-world data. Existing works assume the central server can directly access the predefined shortcut features during the training process. However, these may fail in the federated learning setting as the shortcut features of the heterogeneous watermarked data are difficult to obtain. In this paper, we propose a federated Morozov regularization technique, where the regularization parameter can be adaptively determined based on the watermark knowledge of all the clients in a privacy-preserving way, to eliminate the shortcut learning problem caused by the watermarked data. Specifically, federated Morozov regularization firstly performs lightweight local watermark mask estimation in each client to obtain the locations and intensities knowledge of local watermarks. Then, it aggregates the estimated local watermark masks to generate the global watermark knowledge with a weighted averaging. Finally, federated Morozov regularization determines the regularization parameter for each client by combining the local and global watermark knowledge. With the regularization parameter determined, the model is trained as normal federated learning. We implement and evaluate federated Morozov regularization based on a real-world deployment of federated learning on 40 Jetson devices with real-world datasets. The results show that federated Morozov regularization improves model accuracy by 11.22% compared to existing baselines. Tao Ling, Siping Shi, Hao Wang 0022, Chuang Hu, Dan Wang 0002 |
ACM Multimedia | 2 |
| 2024 | Distributionally Robust Federated Learning for Network Traffic Classification With Noisy LabelsabstractNetwork traffic classifiers of mobile devices are widely learned with federated learning(FL) for privacy preservation. Noisy labels commonly occur in each device and deteriorate the accuracy of the learned network traffic classifier. Existing noise elimination approaches attempt to solve this by detecting and removing noisy labeled data before training. However, they may lead to poor performance of the learned classifier, as the remaining traffic data in each device is few after noise removal. Motivated by the observation that the data feature of the noisy labeled traffic data is clean and the underlying true distribution of the noisy labeled data is statistically close to the clean traffic data, we propose to utilize the noisy labeled data by normalizing it to be close to the clean traffic data distribution. Specifically, we first formulate a distributionally robust federated network traffic classifier learning problem (DR-NTC) to jointly take the normalized traffic data and clean data into training. Then we specify the normalization function under Wasserstein distance to transform the noisy labeled traffic data into a certified robust region around the clean data distribution, and we reformulate the DR-NTC problem into an equivalent DR-NTC-W problem. Finally, we design a robust federated network traffic classifier learning algorithm, RFNTC, to solve the DR-NTC-W problem. Theoretical analysis shows the robustness guarantee of RFNTC. We evaluate the algorithm by training classifiers on a real-world dataset. Our experimental results show that RFNTC significantly improves the accuracy of the learned classifier by up to 1.05 times. Siping Shi, Yingya Guo, Dan Wang 0002, Yifei Zhu 0001, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Federated HD Map Updating Through Overlapping Coalition Formation GameabstractHigh Definition (HD) maps have become core supporting components for autonomous driving. To date, their updates heavily depend on the vehicle fleets of the map vendors, which cannot scale and timely reflect the highly dynamic environment. To ensure the HD map quality, it is advocated social vehicles should be used. Nevertheless, there are privacy concerns and a lack of incentives for social vehicles to contribute data. In this paper, we leverage federated analytics (FA), a newly developed collaborative data analytics paradigm, where raw data are kept local and only the insights generated from local analytics are sent to a server for aggregation. We present a new Federated Analytics based HD map Updating model (FAUMap) to protect the privacy of social vehicles. To motivate social vehicles to contribute data and improve the HD map quality, we formulate an overlapping coalition formation game, OCFUMap, and develop an algorithm to find feasible coalitions. Simulations show that our approach can improve the quality of the updated HD map by 1.56 times. To study an end-to-end operation of the FAUMap model and OCFUMap game, we present a case of HD map updates of the Powell street in San Francisco using the autonomous driving simulator CarLA. Siping Shi, Chuang Hu, Dan Wang 0002, Yifei Zhu 0001, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Distributionally Robust Federated Learning for Differentially Private DataabstractLocal differential privacy (LDP) is a prominent approach and widely adopted in federated learning (FL) to preserve the privacy of local training data. It also nicely provides a rigorous privacy guarantee with computational efficiency in theory. However, a strong privacy guarantee with local differential privacy can degrade the adversarial robustness of the learned global model. To date, very few studies focus on the interplay between LDP and the adversarial robustness of federated learning. In this paper, we observe that LDP adds random noise to the data to achieve privacy guarantee of local data, and thus introduces uncertainty to the training dataset of federated learning. This leads to decreased robustness. To solve this robustness problem caused by uncertainty, we propose to leverage the promising distributionally robust optimization (DRO) modeling approach. Specifically, we first formulate a distributionally robust and private federated learning problem (DRPri). While our formulation successfully captures the uncertainty generated by the LDP, we show that it is not easily tractable. We thus transform our DRPri problem to another equivalent problem, under the Wasserstein distance-based uncertainty set, which is named the DRPri-W problem. We then design a robust and private federated learning algorithm, RPFL, to solve the DRPri-W problem. We analyze RPFL and theoretically show it satisfies differential privacy with a robustness guarantee. We evaluate algorithm RPFL by training classifiers on real-world datasets under a set of well-known attacks. Our experimental results show our algorithm RPFL can significantly improve the robustness of the trained global model under differentially private data by up to 4.33 times. Siping Shi, Chuang Hu, Dan Wang 0002, Yifei Zhu 0001, Zhu Han 0001 |
ICDCS | 1 |
| 2022 | Preva: Protecting Inference Privacy through Policy-based Video-frame TransformationabstractReal-time edge-cloud video analytics systems have been widely used to support such applications as traffic counting, surveillance, autonomous driving, Metaverse, etc. In such a system, the edge and the cloud cooperatively conduct model inference of the video frames captured by the camera of the edge, using a trained DNN model of the video analytics application. The edge conducts initial analytics on the video frames to a split layer of the DNN model; and then sends intermediate results to the cloud for follow-up analytics. In this paper, we show that an attacker can perform reconstruction attacks to the intermediate results; and private information of the raw video frames, e.g., a plate number of a car, can be leaked. In this paper, we present Preva, a new Privacy preserving Real-time Edge-cloud Video Analytics system. The core idea of Preva is to conduct image transformation on the video frames, as preprocessing, prior to the video frames starting the edge-cloud video analytics process, so that during edge-cloud video analytics, the intermediate results will not leak private information under attack. We design a policy-based video-frame transformation scheme. Given the resource constraints of the edge, Preva ensures high accuracy in the final video analytics results and minimizes privacy leakage in any split layer. We present a formal privacy analysis and we show that Preva can guarantee privacy leakage under the reconstruction attacks of both outsider attackers and insider attackers. We evaluate Preva through three video analytics applications and we show that Preva outperforms existing systems for 64.4% in analytics accuracy and 59.2% in privacy leakage. Siping Shi, Dan Wang 0002, Chuang Hu, Bihai Zhang |
SEC | 2 |
| 2022 | Federated Anomaly Analytics for Local Model Poisoning AttackabstractThe local model poisoning attack is an attack to manipulate the shared local models during the process of distributed learning. Existing defense methods are passive in the sense that they try to mitigate the negative impact of the poisoned local models instead of eliminating them. In this paper, we leverage the new federated analytics paradigm, to develop a proactive defense method. More specifically, federated analytics is to collectively carry out analytics tasks without disclosing local data of the edge devices. We propose a Federated Anomaly Analytics enhanced Distributed Learning (FAA-DL) framework, where the clients and the server collaboratively analyze the anomalies. FAA-DL firstly detects all the uploaded local models and splits out the potential malicious ones. Then, it verifies each potential malicious local model with functional encryption. Finally, it removes the verified anomalies and aggregates the remaining to produce the global model. We analyze the FAA-DL framework and show that it is accurate, robust, and efficient. We evaluate FAA-DL by training classifiers on MNIST and Fashion-MNIST under various local model poisoning attacks. Our experiment results show FAA-DL improves the accuracy of the learned global model under strong attacks up to 6.90 times and outperforms the state-of-the-art defense methods with a robustness guarantee. Siping Shi, Chuang Hu, Dan Wang 0002, Yifei Zhu 0001, Zhu Han 0001 |
IEEE J. Sel. Areas Commun. | 1 |