Caijun Sun

dblp:216/9100 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-1529-3179ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Split Unlearning
abstract
We introduce Split Unlearning, a novel machine unlearning technology designed for Split Learning (SL), enabling the first-ever implementation of Sharded, Isolated, Sliced, and Aggregated (SISA) unlearning in SL frameworks. Particularly, the tight coupling between clients and the server in existing SL frameworks results in frequent bidirectional data flows and iterative training across all clients, violating the ''Isolated'' principle and making them struggle to implement SISA for independent and efficient unlearning. To address this, we propose SplitWiper with a new one-way-one-off propagation scheme, which leverages the inherently ''Sharded'' structure of SL and decouples neural signal propagation between clients and the server, enabling effective SISA unlearning even in scenarios with absent clients. We further design SplitWiper+ to enhance client label privacy, which integrates differential privacy and label expansion strategy to defend the privacy of client labels against the server and other potential adversaries. Experiments across diverse data distributions and tasks demonstrate that SplitWiper achieves 0% accuracy for unlearned labels, and 8% better accuracy for retained labels than non-SISA unlearning in SL. Moreover, the one-way-one-off propagation maintains constant overhead, reducing computational and communication costs by 99%. SplitWiper+ preserves 90% of label privacy when sharing masked labels with the server.
Yanna Jiang, Guangsheng Yu, Qin Wang 0008, Xu Wang 0004, Baihe Ma, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001
CCS6
2025 Exploiting attribute correlation for reconstruction attacks on differentially private multi-attributed data
Yanna Jiang, Baihe Ma, Xu Wang 0004, Guangsheng Yu, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001
J. Inf. Secur. Appl.5
2025 Understanding DAOs: An Empirical Study on Governance Dynamics
abstract
As a typical instance of human–computer interaction, the notion of decentralized autonomous organization (DAO) represents an organization constructed by automatically executed rules, such as via smart contracts, incorporating features of the permissionless committee, transparent proposals, and fair contributions by stakeholders. As of May 2023, DAO has impacted over $24.3B market caps. However, there are limited studies focused on this emerging field. To fill the gap, we start from the ground truth by empirically studying the breadth and depth of the DAO markets in mainstream public chain ecosystems in this article. We dive into the most widely adoptable DAO launchpad,Snapshot, which covers 95% of the wild DAO projects for data collection and analysis. By integrating extensively enrolled DAOs and corresponding data measurements, we explore statistical resources from Snapshot and analyze data from 581 DAO projects, encompassing 16 246 proposals over the course of 3+ years. Our empirical research has uncovered a multitude of previously unknown facts about DAOs, spanning topics such as their status, features, performance, threats, and ways of improvement. We have distilled these findings into a series of key insights and takeaway messages, emphasizing their significance. Notably, our study is the first of its kind to comprehensively examine the DAO ecosystem with a focus on scale and scope of data, real-time relevance, practical implementations, and comprehensive metrics, addressing critical gaps in the current literature.
Qin Wang 0008, Guangsheng Yu, Yilin Sai, Caijun Sun, Lam Duc Nguyen, Shiping Chen 0001
IEEE Trans. Comput. Soc. Syst.4
2025 IronForge: An Open, Secure, Fair, Decentralized Federated Learning
abstract
Federated learning (FL) offers an effective learning architecture to protect data privacy in a distributed manner. However, the inevitable network asynchrony, overdependence on a central coordinator, and lack of an open and fair incentive mechanism collectively hinder FL's further development. We propose IronForge, a new generation of FL framework, that features a directed acyclic graph (DAG)-based structure, where nodes represent uploaded models, and referencing relationships between models form the DAG that guides the aggregation process. This design eliminates the need for central coordinators to achieve fully decentralized operations. IronForge runs in a public and open network and launches a fair incentive mechanism by enabling state consistency in the DAG. Hence, the system fits in networks where training resources are unevenly distributed. In addition, dedicated defense strategies against prevalent FL attacks on incentive fairness and data privacy are presented to ensure the security of IronForge. Experimental results based on a newly developed test bed FLSim highlight the superiority of IronForge to the existing prevalent FL frameworks under various specifications in performance, fairness, and security. To the best of our knowledge, IronForge is the first secure and fully decentralized FL (DFL) framework that can be applied in open networks with realistic network and training settings.
Guangsheng Yu, Xu Wang 0004, Caijun Sun, Qin Wang 0008, Wei Ni 0001, Ren Ping Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 TbDd: A new trust-based, DRL-driven framework for blockchain sharding in IoT
abstract
Integrating sharded blockchain with IoT presents a solution for trust issues and optimized data flow. Sharding boosts blockchain scalability by dividing its nodes into parallel shards, yet it is vulnerable to the 1% attacks where dishonest nodes target a shard to corrupt the entire blockchain. Balancing security with scalability is pivotal for such systems. Deep Reinforcement Learning (DRL) adeptly handles dynamic, complex systems and multi-dimensional optimization. This paper introduces a Trust-based and DRL-driven (TbDd) framework, crafted to counter collusion attack risks and dynamically adjust node allocation, enhancing throughput while maintaining network security. With a comprehensive trust evaluation mechanism, TbDd discerns node types and performs targeted resharding against potential threats. The TbDd framework maximizes the tolerance for dishonest nodes, optimizes node movement frequency, ensures even node distribution in shards, and balances sharding risks. Extensive evaluations validate TbDd’s superiority over conventional random-, community-, and trust-based sharding methods in shard risk equilibrium and reducing cross-shard transactions.
Zixu Zhang, Guangsheng Yu, Caijun Sun, Xu Wang 0004, Ying Wang 0096, Wei Ni 0001, Ren Ping Liu 0001, Andrew Reeves, Nektarios Georgalas
Comput. Networks3
2024 Preventing harm to the rare in combating the malicious: A filtering-and-voting framework with adaptive aggregation in federated learning
abstract
The distributed nature of Federated Learning (FL) introduces security vulnerabilities and issues related to the heterogeneous distribution of data. Traditional FL aggregation algorithms often mitigate security risks by excluding outliers, which compromises the diversity of shared information. In this paper, we introduce a novel filtering-and-voting framework that adeptly navigates the challenges posed by non-iid training data and malicious attacks on FL. The proposed framework integrates a filtering layer for defensive measures against the intrusion of malicious models and a voting layer to harness valuable contributions from diverse participants. Moreover, by employing Deep Reinforcement Learning (DRL) for dynamic aggregation weight adjustment, we ensure the optimized aggregation of participant data, enhancing the diversity of information used for aggregation and improving the performance of the global model. Experimental results demonstrate that the proposed framework presents superior accuracy over traditional and contemporary FL aggregation methods as diverse models are utilized. It also shows robust resistance against malicious poisoning attacks.
Yanna Jiang, Baihe Ma, Xu Wang 0004, Guangsheng Yu, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001
Neurocomputing5
2023 A First Look into Blockchain DAOs
abstract
Decentralized autonomous organizations (DAOs) are critical to the blockchain ecosystem as they enable decentralized decision-making and governance, and facilitate the creation of decentralized applications (DApps) and organizations. However, despite significant importance, there is currently a lack of a comprehensive overview and detailed understanding of DAOs. To address the gap, this work presents a primary investigation of DAOs (35+). We category, examine and evaluate existing DAOs regarding their operational features, (non-)functionalities and real-world performance. In addition, we provide a consolidated exploration of DAOs by conducting a literature review [1] and an empirical study on mainstream projects, particularly Snapshot [2]. Our research contributes to a better understanding of DAOs and their potential impact on the blockchain ecosystem.
Qin Wang 0008, Guangsheng Yu, Yilin Sai, Caijun Sun, Lam Duc Nguyen, Xiwei Xu 0001, Shiping Chen 0001
ICBC4
2023 Obfuscating the Dataset: Impacts and Applications
abstract
Obfuscating a dataset by adding random noises to protect the privacy of sensitive samples in the training dataset is crucial to prevent data leakage to untrusted parties when dataset sharing is essential. We conduct comprehensive experiments to investigate how the dataset obfuscation can affect the resultant model weights —in terms of the model accuracy, ℓ 2 -distance-based model distance, and level of data privacy—and discuss the potential applications with the proposed Privacy, Utility, and Distinguishability (PUD)-triangle diagram to visualize the requirement preferences. Our experiments are based on the popular MNIST and CIFAR-10 datasets under both independent and identically distributed (IID) and non-IID settings. Significant results include a tradeoff between the model accuracy and privacy level and a tradeoff between the model difference and privacy level. The results indicate broad application prospects for training outsourcing and guarding against attacks in federated learning both of which have been increasingly attractive in many areas, particularly learning in edge computing.
Guangsheng Yu, Xu Wang 0004, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001
ACM Trans. Intell. Syst. Technol.3
2018 Secure multi-keyword ranked search over encrypted cloud data for multiple data owners
Ziqing Guo, Hua Zhang 0001, Caijun Sun, Qiaoyan Wen, Wenmin Li 0001
J. Syst. Softw.3