Shuya Feng

dblp:298/9381 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-8139-7366ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Harmonizing Differential Privacy Mechanisms for Federated Learning: Boosting Accuracy and Convergence
abstract
Differentially private federated learning (DP-FL) offers a compelling approach to collaborative model training by ensuring robust privacy for clients. Despite its potential, current methods face challenges in effectively balancing privacy, utility, and performance across diverse federated learning scenarios. Addressing these challenges, we introduce UDP-FL, to our knowledge the first DP-FL framework that universally harmonizes any randomization mechanism, including those considered optimal, by employing the Gaussian Moments Accountant (viz. DP-SGD). Central to UDP-FL is the 'Harmonizer,' a dynamic module engineered to intelligently select and apply the most suitable DP mechanism tailored to each client's specific privacy requirements, data sensitivities, and computational capacities. This selection process is driven by the principle of Rényi Differential Privacy, which serves as a crucial mediator for aligning privacy budgets effectively. Our comprehensive evaluation of UDP-FL, benchmarked against established baseline methods, demonstrates superior performance in upholding privacy guarantees and enhancing model functionality. The framework's robustness has been rigorously tested against a broad spectrum of privacy attacks, making it one of the most thorough validations of a DP-FL framework to date.
Shuya Feng, Meisam Mohammady, Hanbin Hong, Shenao Yan, Ashish Kundu, Binghui Wang, Yuan Hong 0001
CODASPY1
2025 DPED: Multi-Layer Noise Distillation for Privacy-Preserving Text Embeddings
abstract
Training text embedding models under differential privacy constraints is challenging due to the high dimensionality of language data and the presence of rare, identifying linguistic features.We propose DPED (Differentially Private Embedding Distillation), a framework that leverages teacher-student distillation with multi-layer noise injection to learn highquality embeddings while providing differential privacy guarantees.DPED trains an ensemble of teacher models on disjoint subsets of sensitive text data, then transfers their knowledge to a student model through noisy aggregation at multiple layers.A rare-word-aware strategy adaptively handles infrequent words, improving privacy-utility trade-offs.Experiments on benchmark datasets demonstrate that DPED outperforms standard differentially private training methods, achieving substantially higher utility at the same privacy budget.Our approach protects individual word usage patterns in training documents, preventing models from memorizing unique linguistic fingerprints while maintaining practical utility
Shuya Feng, Yuan Hong 0001
EMNLP1
2025 Delay-allowed Differentially Private Data Stream Release
Zhan Qin, Kui Ren 0001, Chen Gong 0005, Shuya Feng, Yuan Hong 0001, Tianhao Wang 0001
NDSS5
2024 Towards Accurate and Stronger Local Differential Privacy for Federated Learning with Staircase Randomized Response
abstract
Federated Learning (FL), a privacy-preserving training approach, has proven to be effective, yet its vulnerability to attacks that extract information from model weights is widely recognized. To address such privacy concerns, Local Differential Privacy (LDP) has been applied to FL: perturbing the weights trained for the local model by each client. However, besides high utility loss on the randomized model weights, we identify a new inference attack to the existing LDP method, that can reconstruct the original value from the noisy values with high confidence. To mitigate these issues, in this paper, we propose the Staircase Randomized Response (SRR)-FL framework, which assigns higher probabilities to weights closer to the true weight, reducing the distance between the true and perturbed data. This minimizes the noise for maintaining the same LDP guarantee, leading to better utility. Compared to existing LDP mechanisms (e.g., Generalized Randomized Response) on the FL, SRR-FL can further provide a more accurate privacy-preserving training model, and enhance the robustness against the inference attack while ensuring the same LDP guarantee. Furthermore, we also use the parameter shuffling method for privacy amplification. The efficacy of SRR-FL has been validated on widely used datasets MNIST, Medical-MNIST and CIFAR-10, demonstrating remarkable performance. Code is available at https://github.com/matta-varun/SRR-FL.
Matta Varun, Shuya Feng, Han Wang 0021, Shamik Sural, Yuan Hong 0001
CODASPY2
2024 DPI: Ensuring Strict Differential Privacy for Infinite Data Streaming
abstract
Streaming data, crucial for applications like crowd-sourcing analytics, behavior studies, and real-time monitoring, faces significant privacy risks due to the large and diverse data linked to individuals. In particular, recent efforts to release data streams, using the rigorous privacy notion of differential privacy (DP), have encountered issues with unbounded privacy leakage. This challenge limits their applicability to only a finite number of time slots ("finite data stream") or relaxation to protecting the events ("event or w-event DP") rather than all the records of users. A persistent challenge is managing the sensitivity of outputs to inputs in situations where users contribute many activities and data distributions evolve over time. In this paper, we present a novel technique for Differentially Private data streaming over Infinite disclosure (DPI) that effectively bounds the total privacy leakage of each user in infinite data streams while enabling accurate data collection and analysis. Furthermore, we also maximize the accuracy of DPI via a novel boosting mechanism. Finally, extensive experiments across various streaming applications and real datasets (e.g., COVID-19, Network Traffic, and USDA Production), show that DPI maintains high utility for infinite data streams in diverse settings. Code for DPI is available at https://github.com/ShuyaFeng/DPI.
Shuya Feng, Meisam Mohammady, Han Wang 0021, Zhan Qin, Yuan Hong 0001
SP1
2022 A Model-Agnostic Approach to Differentially Private Topic Mining
abstract
Topic mining extracts patterns and insights from text data (e.g., documents, emails and product reviews), which can be used in various applications such as intent detection. However, topic mining can result in severe privacy threats to the users who have contributed to the text corpus since they can be re-identified from the text data with certain background knowledge. To our best knowledge, we propose the first differentially private topic mining technique (namely TopicDP) which injects well-calibrated Gaussian noise into the matrix output of any topic mining algorithm to ensure differential privacy and good utility. Specifically, we smoothen the sensitivity for the Gaussian mechanism via sensitivity sampling, which addresses the major challenges resulted from the high sensitivity in topic mining for differential privacy. Furthermore, we theoretically prove the differential privacy guarantee under the Rényi differential privacy mechanism and the utility error bounds of TopicDP. Finally, we conduct extensive experiments on two real-word text datasets (Enron email and Amazon Reviews), and the experimental results demonstrate that TopicDP is a model-agnostic framework that can generate better privacy preserving performance for topic mining as compared against other differential privacy mechanisms.
Han Wang 0021, Jayashree Sharma, Shuya Feng, Kai Shu, Yuan Hong 0001
KDD3
2021 Security Analysis of Block Withholding Attacks in Blockchain
abstract
Blockchain technology has gained growing popularity in recent years. While the technology works well for most applications, the vulnerability of the blockchain-based system is not well understood. The core part of a blockchain-based system is the consensus algorithm. It can be regarded as a protocol used by all network users to make decisions on the growth of the chain. Malicious users can exploit the vulnerabilities of the consensus algorithm to launch block withholding attacks to earn more profit by increasing their winning possibilities, which are significant threats to the security of blockchain-based applications. In this paper, we describe some scenarios of block withholding attacks and suggest effective mitigation methods accordingly. We propose two hybrid consensus algorithms against the block withholding attacks and evaluate the effectiveness of these algorithms, both analytically and through simulation. Simulation results from the SimBlock simulator verified the effectiveness of the proposed algorithms.
Shuya Feng, Jia He 0006, Maggie Cheng 0001
ICC1