EDBT 2026 Demo / reviewers in the wild / expert
Chengjun Cai
dblp:198/7220
· DBLP profile ↗
29ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-8045-4226ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 12 · 5 first-author · 12 since 2021Computer networks · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TimeProtect: Time-Bound Policy Enforcement for Metadata-Hiding Encrypted Analytics
Qingyuan Xie, Zhao Bai, Chengjun Cai, Xiaohua Jia |
ICDCS | 3 |
| 2026 | PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud Collaboration
Yi Liu 0057, Weixiang Han, Chengjun Cai, Xingliang Yuan, Cong Wang 0001 |
INFOCOM | 3 |
| 2026 | ParaVul: A Parallel Large Language Model and Retrieval-Augmented Framework for Smart Contract Vulnerability DetectionabstractSmart contracts play a significant role in automating blockchain services. Nevertheless, vulnerabilities in smart contracts pose serious threats to blockchain security. Currently, traditional detection methods primarily rely on static analysis and formal verification, which can result in high false-positive rates and poor scalability. Large Language Models (LLMs) have recently made significant progress in smart contract vulnerability detection. However, they still face challenges such as high inference costs and substantial computational overhead. In this paper, we propose ParaVul, a parallel LLM and retrievalaugmented framework to improve the reliability and accuracy of smart contract vulnerability detection. Specifically, we first develop Sparse Low-Rank Adaptation (SLoRA), a technique for efficient LLM fine-tuning tailored to smart contract vulnerability detection. Distinct from existing LoRA methods, SLoRA inserts parallel sparse and low-rank branches after the attention projection and the feed-forward block, enabling LLMs to capture both global code semantics and localized vulnerability patterns while maintaining low training overhead. We then construct a vulnerability contract knowledge base and develop a hybrid Retrieval-Augmented Generation (RAG) system that integrates Okapi BM25 with dense retrieval to provide complementary lexical and semantic evidence for smart contract vulnerability verification. Furthermore, we propose a meta-learner-based gated verification module to fuse the outputs of the SLoRA detector and the two RAG-based detectors, thereby generating the final detection results. After completing vulnerability detection, we design chain-of-thought prompts to guide LLMs to generate comprehensive vulnerability detection reports. Simulation results demonstrate the superiority of ParaVul, especially in terms of F1 scores, achieving 0.9398 for single-label detection and 0.9930 for multi-label detection. Tenghui Huang, Jinbo Wen, Jiawen Kang 0001, Siyong Chen, Zhengtao Li, Tao Zhang 0063, Dongning Liu, Jiacheng Wang 0001, Chengjun Cai, Yinqiu Liu |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2026 | Adapting Large Language Models for Encrypted Traffic Analysis Services: An Efficient Realization With Mixture of LoRA ExpertsabstractAs encrypted traffic grows, traditional rule-based and deep learning methods struggle with engineering costs and encryption complexity. While Large Language Models (LLMs) offer promise for traffic analysis via pre-trained feature learning, they face challenges in handling diverse tasks, retaining pre-training knowledge, and adapting efficiently. To address these issues, we propose a new traffic representation learning method and a new Parameter-Efficient Fine-Tuning (PEFT) method for multi-task encrypted traffic analysis services, calledTrafficLLM.TrafficLLMalleviates task heterogeneity by utilizing a universal multi-task prompt template and addresses pre-training knowledge forgetting by integrating Singular Value Decomposition based Low-Rank Adaptation (SVD-LoRA). To further reduce the cost of adapting to multiple tasks, we combine the strengths of the Mixture of Experts (MoE) for multi-task learning with SVD-LoRA for PEFT, enabling efficient multi-task traffic analysis. Additionally, we introduce task-aware gating functions to dynamically assign different weights to experts, facilitating the efficient fusion of expert knowledge. Comprehensive experiments on 7 datasets across 5 downstream tasks demonstrate thatTrafficLLMdelivers superior analysis performance and resource efficiency compared to state-of-the-art models, including DeepSeek, NetGPT, ET-BERT, and TFE-GNN. Detailed analysis of throughput, memory usage, and latency further highlights the practical advantages ofTrafficLLM. Our data and code are available athttps://github.com/yiliucs/TrafficLLMhttps://github.com/yiliucs/TrafficLLM. Yi Liu 0057, Chengjun Cai, Xingliang Yuan, Cong Wang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Decentralized and Fair Trading Via Blockchain: The Journey So Far and the Road AheadabstractCentralized trading platforms have long been the preferred choice for users, despite growing concerns regarding data privacy. Users have to place their trust in these platforms and provide sensitive personal information, like identities and financial accounts. However, these centralized platforms often lack transparency, making it challenging to ensure fairness, privacy, and security against both external and internal risks. In contrast, a decentralized fair trading paradigm, harnessing the potential of blockchain technology, is rapidly emerging. It empowers individuals to engage in the exchange of digital assets with others while guaranteeing fairness, efficiency, and privacy. In this paper, we conduct a comprehensive survey of decentralized fair trading. We commence by providing fundamental definitions of fair trading and tracing its evolution over time. We then delve into the essential framework of on-chain and off-chain trading and highlight key improvements that enhance the efficiency of decentralized fair trading within various application scenarios. Furthermore, we undertake a thorough analysis of privacy and security enhancements within the scope, summarizing defenses against known attacks. Finally, we outline the challenges and offer insights into the future prospects of decentralized fair trading, with the aim of inspiring the development of more innovative and promising designs in this evolving trend. Hao Zeng 0006, Helei Cui, Bo Zhang 0119, Chengjun Cai, Zhiwen Yu 0001, Bin Guo 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Private Sample Alignment for Vertical Federated Learning: An Efficient and Reliable RealizationabstractSample alignment is recognized as a vital component of vertical federated learning, which facilitates the integration of differential samples and high-quality model training. In this trend, providing Private Sample Alignment (PSA) among multi-clients becomes naturally necessary for preventing unauthorized sample access and client privacy exposure. However, exiting PSA protocols mainly focus on two-party scenarios and cannot be directly adapted to the multi-client delegated computing scenarios required for vertical federated learning. Besides, these studies fail to address the need for protocol robustness in practical federated Learning network environments. Therefore, we aim to design an efficient and reliable PSA protocol in multi-client vertical federated learning. In this work, we present the first practical PSA protocol for vertical federated learning, allowing multi-clients to efficiently identify common samples without revealing additional information. Toward this direction, our PSA protocol first explores the Learning With Errors (LWE) problem to create a lightweight delegated Private Set Intersection (PSI) scheme, enabling efficient sample intersection among multiple clients. To achieve the reliability of the PSA protocol, we devise a multi-client vector aggregation algorithm that securely delegates the server to calculate the sample intersection. Building on this foundation, we develop an efficient Threshold-based Private Sample Alignment (T-PSA) protocol that allows multiple clients to determine the intersection of their input samples only if the intersection size surpasses a specific threshold. We implement a prototype and conduct a thorough security analysis. Comprehensive evaluation results confirm the efficiency and practicality of our design. Yuxin Xi, Yu Guo 0003, Shiyuan Xu, Chengjun Cai, Xiaohua Jia |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Anti-Confounding Hashing: Enhancing Radiological Image Retrieval via Debiased Weighting and Counterfactual ReasoningabstractContent-based medical image retrieval (CBMIR) enables physicians to make evidence-based diagnoses by retrieving similar medical images and recalling previous cases stored in databases. However, existing CBMIR models are prone to capturing superficial correlations due to confounding factors such as complex host organs and lesions, imaging discrepancies, artifacts, and inconsistent protocols. To address this issue, we propose a plug-and-play anti-confounding hashing (ACH) method, which uses debiased sample weighting and lesion counterfactual reasoning (LCR) to directly capture the natural direct effect (NDE) of lesions on query medical images without bias. The devised debiased weighting (DBW) loss adopts a backdoor adjustment to separate lesions from confounders. To effectively locate salient areas of lesions, we present a coarse-to-fine lesion positioning (C2F-LP) module by counterfactual reasoning. On two real-world radiological image datasets, ACH achieves 0.2%-9% improvement in mean average precision (mAP) over the six state-of-the-art methods, when using code lengths ranging from 8-bit to 32-bit. Its robustness to confounding factors is demonstrated through explainable visual analysis. Yao Hu 0001, Chengjun Cai, Zhi-an Huang, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Combating Abusive Information in Encrypted Messaging Services: A Secure and Efficient RealizationabstractEnd-to-end encrypted messaging services (EEMSs) empower private communication through encrypting messages, yet also make content moderation for combating the spread of abusive messages challenging. There is an urgent call for supporting content moderation in EEMSs while ensuring user privacy. In this paper, we present a new system design for privacy-assured content moderation in EEMSs. At a high level, users in our system can privately report abusive messages, and the EEMS traces the source if a message has an aggregated report count exceeding a predefined threshold and is audited to be abusive. Our system mainly departs from prior works in that it allows flexible and adaptable thresholds, offers robustness against dishonest reporters providing malformed reports, and better ensures the privacy of all users during the moderation process. We also take a step further and propose a privacy-aware detection mechanism that relies on a blocklist built with transparency to mitigate the further spread of identified abusive messages from forwarders. Formal security analysis is provided and extensive experiments demonstrate the practical efficiency of our system. Rui Lian, Yifeng Zheng 0001, Yulong Ming, Chengjun Cai, Cong Wang 0001, Xiaohua Jia |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | SmartUpdater: Enabling Transparent, Automated, and Secure Maintenance of Stateful Smart ContractsabstractSmart contracts in the Ethereum system are stored tamper-resistant, complicating necessary maintenance for offering new functionalities or fixing security vulnerabilities. Previous contract maintenance approaches mainly focus on logic modification using delegatecall-based patterns. While popular, they fail to handle data state updates (like storage layout changes), leading to impracticality and security risks in real-world applications. To address these challenges, this paper introduces SmartUpdater, a novel toolchain designed for transparent, automated, and secure maintenance of stateful smart contracts. SmartUpdater employs a hyperproxy-based contract maintenance pattern, where the hyperproxy serves as a constant entry and ensures that any state/logic modifications remain transparent to end users. SmartUpdater automates the maintenance process in terms of development streamlining, gas cost efficiency, and state migration verifiability. In extensive evaluations, we show that SmartUpdater can reduce gas consumption in contract maintenance compared with actual maintenance approaches. The evaluations point out the potential of SmartUpdater to significantly simplify the maintenance process for developers. Xiaoli Zhang 0003, Yiqiao Song, Yuefeng Du 0001, Chengjun Cai, Hongbing Cheng, Ke Xu 0002, Qi Li 0002 |
IEEE Trans. Software Eng. | 4 |
| 2024 | Nemesis: Combating Abusive Information in Encrypted Messaging with Private Reporting
Rui Lian, Yulong Ming, Chengjun Cai, Yifeng Zheng 0001, Cong Wang 0001, Xiaohua Jia |
ESORICS (2) | 3 |
| 2024 | HiddenTor: Toward a User-Centric and Private Query System for Tor BridgeDBabstractTor bridges are crucial, unlisted relays designed to enhance system accessibility and circumvent censorship in the Tor network. Currently, Tor BridgeDB will randomly distribute 1–3 bridge relays to the user per request. Yet, those randomly selected bridges may not meet users' specific needs, e.g., adequate bandwidth for large-file sharing in a certain region. Also, a user's usage metadata (e.g., bridge choices) collected by Tor BridgeDB would inevitably reveal sensitive information about the user, discouraging the use of this censorship-circumvention service. In light of them, we introduce HiddenTor, a user-centric and privacy-focused bridge distribution system that allows Tor users to retrieve bridges privately and precisely (i.e., based on a range of specific criteria). At its core, HiddenTor designs a condition-based private information retrieval (PIR) protocol by building atop a suite of lightweight cryptographic primitives (i.e., function secret sharing). Besides, HiddenTor also crafts several optimization designs to balance the trade-offs between query efficiency and service reliability. The extensive experimental results have confirmed the feasibility and practicality of HiddenTor. For example, our prototype can efficiently handle private queries over 3000 bridges in approximately 2 seconds, which can further be reduced to 0.21 seconds using parallel computing techniques. Yichen Zang, Chengjun Cai, Lei Xu 0019, Cong Wang 0001 |
ICDCS | 2 |
| 2024 | SecMdp: Towards Privacy-Preserving Multimodal Deep Learning in End-Edge-CloudabstractMultimodal deep learning technologies have advanced significantly, which brings extensive applications in diverse fields. The substantial computational demands of training and prediction in multimodal deep learning have made the End-Edge-Cloud (EEC) framework popular. It is essential to protect multimodal data and model privacy in such a framework. However, traditional cryptographic methods, though secure for data and models at edge nodes, cause efficiency limitations. In this paper, we propose SecMdp, an SGX-assisted secure computational framework for multimodal data in the EEC architecture. Edge nodes are equipped with the trusted execution environment (e.g., Intel SGX) to run multimodal algorithms. Additionally, to address the side-channel attacks of SGX, we present an enhanced PathORAM algorithm, MM_PathORAM, for the multimodal training and prediction processes, which are tailored for multimodal deep learning scenarios. It accelerates multimodal data access while protecting data privacy and model security. Experimental evaluation supports the effectiveness of our design in preserving edge computing efficiency. It demonstrates negligible impact on the speed of multimodal data loading, the configuration of model parameters during training, or the accuracy of predictions. Zhao Bai, Fangda Guo, Yu Guo 0003, Chengjun Cai, Rongfang Bie, Xiaohua Jia |
ICDE | 5 |
| 2024 | ERL-MR: Harnessing the Power of Euler Feature Representations for Balanced Multi-modal LearningabstractMulti-modal learning leverages data from diverse perceptual media to obtain enriched representations, thereby empowering machine learning models to complete more complex tasks. However, recent research results indicate that multi-modal learning still suffers from " modality imbalance '': Certain modalities' contributions are suppressed by dominant ones, consequently constraining the overall performance enhancement of multimodal learning. To tackle this issue, current approaches attempt to mitigate modality competition in various ways, but their effectiveness is still limited. To this end, we propose an Euler Representation Learning-based Modality Rebalance (ERL-MR) strategy, which reshapes the underlying competitive relationships between modalities into mutually reinforcing win-win situations while maintaining stable feature optimization directions. Specifically, ERL-MR employs Euler's formula to map original features to complex space, constructing cooperatively enhanced non-redundant features for each modality, which helps reverse the situation of modality competition. Moreover, to counteract the performance degradation resulting from optimization drift among modalities, we propose a Multi-Modal Constrained (MMC) loss based on cosine similarity of complex feature phase and cross-entropy loss of individual modalities, guiding the optimization direction of the fusion network. Extensive experiments conducted on four multi-modal multimedia datasets and two task-specific multi-modal multimedia datasets demonstrate the superiority of our ERL-MR strategy over state-of-the-art baselines, achieving modality rebalancing and further performance improvements. Weixiang Han, Chengjun Cai, Yu Guo 0003, Jialiang Peng |
ACM Multimedia | 2 |
| 2024 | Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak PromptsabstractLarge Vision Language Models (VLMs) extend and enhance the perceptual abilities of Large Language Models (LLMs).Despite offering new possibilities for LLM applications, these advancements raise significant security and ethical concerns, particularly regarding the generation of harmful content.While LLMs have undergone extensive security evaluations with the aid of red teaming frameworks, VLMs currently lack a well-developed one.To fill this gap, we introduce Arondight, a standardized red team framework tailored specifically for VLMs.Arondight is dedicated to resolving issues related to the absence of visual modality and inadequate diversity encountered when transitioning existing red teaming methodologies from LLMs to VLMs.Our framework features an automated multi-modal jailbreak attack, wherein visual jailbreak prompts are produced by a red team VLM, and textual prompts are generated by a red team LLM guided by a reinforcement learning agent.To enhance the comprehensiveness of VLM security evaluation, we integrate entropy bonuses and novelty reward metrics.These elements incentivize the RL agent to guide the red team LLM in creating a wider array of diverse and previously unseen test cases.Our evaluation of ten cutting-edge VLMs exposes significant security vulnerabilities, particularly in generating toxic images and aligning multi-modal prompts.In particular, our Arondight achieves an average attack success rate of 84.5% on GPT-4 in all fourteen prohibited scenarios defined by OpenAI in terms of generating toxic text.For a clearer comparison, we also categorize existing VLMs based on their safety levels and provide corresponding reinforcement recommendations.Our multimodal prompt dataset and red team code will be released after ethics committee approval. Yi Liu 0057, Chengjun Cai, Xiaoli Zhang 0003, Xingliang Yuan, Cong Wang 0001 |
ACM Multimedia | 2 |
| 2024 | DPoSt: Dynamic Proof of Storage-timeabstractProof of storage-time (PoSt) enforces highly reliable data storage services to data owners by providing lightweight and continuous data availability or possession checks to the uploaded data. Despite being promising, current PoSt schemes mainly focus on static data that will not change after a PoSt process starts. This limitation, however, would largely undermine the practicality of PoSt schemes as they cannot be effectively used for important and emerging cloud storage services like document sharing, code collaboration, and website management that will modify the uploaded data whenever needed. In this paper, we propose DPoSt, a new PoSt system that provides continuous auditing protection to data owners while supporting efficient data updates. DPoSt is built with a tailored suite of cryptographic primitives to update the intermediate auditing proofs while reducing the changes we need to conduct for subsequent challenge and proof pairs in the PoSt process. We develop a prototype of DPoSt, and the evaluation results demonstrate the efficiency of DPoSt for supporting dynamic PoSt operations. Qingyuan Xie, Chengjun Cai, Zhi-an Huang, Xiaohua Jia |
MSN | 2 |
| 2024 | VizardFL: Enabling Private Participation in Federated Learning Systems
Yichen Zang, Chengjun Cai, Cong Wang 0001 |
WISE (2) | 2 |
| 2024 | EVM-Shield: In-Contract State Access Control for Fast Vulnerability Detection and PreventionabstractRecently, smart contracts have been widely applied in security-sensitive fields yet are fragile to various vulnerabilities and attacks. Regarding this, existing research efforts either statically scrutinize smart contracts’ code or detect suspicious transaction execution flows. However, they either fail to timely protect contracts or only handle a small subset of well-known vulnerabilities. In the paper, we propose$\mathtt {EVM}$-$\mathtt {Shield}$that secures vulnerable smart contracts in real-time via fine-grained access control over sensitive states. The behind rationale is most of attacks aim to manipulate money-related states (e.g., tokens) for profits. Specifically, transaction-level state access control policies are first defined by developers and then translated into EVM-level policies with contract-aware function-level state access permissions. In policy enforcement,$\mathtt {EVM}$-$\mathtt {Shield}$introduces a hybrid storage analyzer to accurately identify (dynamic-allocated) storage locations for policy-involved states and a multi-stage cache based filter to fast revert bad transactions with unexpected state access behaviors. Finally, we conduct thorough experiments using 12 types of real-world contract vulnerabilities and all open-source smart contracts on the first$8M$blocks of Ethereum. The results demonstrate that$\mathtt {EVM}$-$\mathtt {Shield}$outperforms two state-of-the-art runtime analysis tools in terms of attack detection. Extensive performance evaluations with$185M$real-world transactions show that$\mathtt {EVM}$-$\mathtt {Shield}$can block 100% unexpected state accesses at the cost of 8% throughput degradation (compared with the native EVM). Xiaoli Zhang 0003, Wenxiang Sun, Hongbing Cheng, Chengjun Cai, Helei Cui, Qi Li 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Vizard: A Metadata-hiding Data Analytic System with End-to-End Policy ControlsabstractOwner-centric control is a widely adopted method for easing owners' concerns over data abuses and motivating them to share their data out to gain collective knowledge. However, while many control enforcement techniques have been proposed, privacy threats due to the metadata leakage therein are largely neglected in existing works. Unfortunately, a sophisticated attacker can infer very sensitive information based on either owners' data control policies or their analytic task participation histories (e.g., participating in a mental illness or cancer study can reveal their health conditions). To address this problem, we introduce Vizard, a metadata-hiding analytic system that enables privacy-hardened and enforceable control for owners. Vizard is built with a tailored suite of lightweight cryptographic tools and designs that help us efficiently handle analytic queries over encrypted data streams coming in real-time (like heart rates). We propose extension designs to further enable advanced owner-centric controls (with AND, OR, NOT operators) and provide owners with release control to additionally regulate how the result should be protected before deliveries. We develop a prototype of Vizard that is interfaced with Apache Kafka, and the evaluation results demonstrate the practicality of Vizard for large-scale and metadata-hiding analytics over data streams. Chengjun Cai, Yichen Zang, Cong Wang 0001, Xiaohua Jia, Qian Wang 0002 |
CCS | 1 |
| 2022 | Toward a Secure, Rich, and Fair Query Service for Light Clients on Public BlockchainsabstractThe rapid growth of storage overhead on public blockchains has urged the use of light clients that only store a small fraction of blockchain data and rely on other bootstrapped full nodes for data retrievals. Unfortunately, current blockchain light client designs are far from satisfactory. First, outsourcing retrieval requests could raise severe concerns about result correctness and privacy threats. Second, current light clients do not support rich query and enforce fee payments to full nodes. Given that blockchain storage increases day by day, enabling effective rich blockchain queries and fairly compensating full nodes’ ever growing costs has become extremely necessary. In this article, we propose a general and secure paid query framework to simultaneously meet those demands above. Specifically, we leverage the integration of trusted hardware (e.g., Intel SGX) and smart contract as a starting point for building efficient yet secure query processing with fair payments. Then, we further craft several crucial performance and security refinement designs to boost query efficiency and enforce result correctness, and also explore an enclave-facilitated fair settlement mechanism for on-chain cost optimizations. We implement a prototype of our paid query framework and the experimental result has demonstrated its practically affordable cost. Chengjun Cai, Lei Xu 0019, Anxin Zhou, Cong Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Golden Grain: Building a Secure and Decentralized Model Marketplace for MLaaSabstractML-as-a-service (MLaaS) becomes increasingly popular and revolutionizes the lives of people. A natural requirement for MLaaS is, however, to provide highly accurate prediction services. To achieve this, current MLaaS systems integrate and combine multiple well-trained models in their services. Yet, in reality, there is no easy way for MLaaS providers, especially for startups, to collect sufficiently well-trained models from individual developers, due to the lack of incentives. In this article, we aim to fill this gap by building up a model marketplace, called as Golden Grain, to facilitate model sharing, which enforces the fair model-money swapping process between individual developers and MLaaS providers. Specifically, we deploy the swapping process on the blockchain, and further introduce a blockchain-empowered model benchmarking process for transparently determining the model prices according to their authentic performances, so as to motivate the faithful contributions of well-trained models. Especially, to ease the blockchain overhead for model benchmarking, our marketplace carefully offloads the heavy computation and designs a secure off-chain on-chain interaction protocol based on a trusted execution environment (TEE), for ensuring both the integrity and authenticity of benchmarking. We implement a prototype of our Golden Grain on the Ethereum blockchain, and conduct extensive experiments using standard benchmark datasets to demonstrate the practically affordable performance of our design. Jia-Si Weng 0001, Jian Weng 0001, Chengjun Cai, Cong Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | FedServing: A Federated Prediction Serving Framework Based on Incentive MechanismabstractData holders, such as mobile apps, hospitals and banks, are capable of training machine learning (ML) models and enjoy many intelligence services. To benefit more individuals lacking data and models, a convenient approach is needed which enables the trained models from various sources for prediction serving, but it has yet to truly take off considering three issues: (i) incentivizing prediction truthfulness; (ii) boosting prediction accuracy; (iii) protecting model privacy.We design FedServing, a federated prediction serving framework, achieving the three issues. First, we customize an incentive mechanism based on Bayesian game theory which ensures that joining providers at a Bayesian Nash Equilibrium will provide truthful (not meaningless) predictions. Second, working jointly with the incentive mechanism, we employ truth discovery algorithms to aggregate truthful but possibly inaccurate predictions for boosting prediction accuracy. Third, providers can locally deploy their models and their predictions are securely aggregated inside TEEs. Attractively, our design supports popular prediction formats, including top-1 label, ranked labels and posterior probability. Besides, blockchain is employed as a complementary component to enforce exchange fairness. By conducting extensive experiments, we validate the expected properties of our design. We also empirically demonstrate that FedServing reduces the risk of certain membership inference attack. Jia-Si Weng 0001, Jian Weng 0001, Chengjun Cai, Cong Wang 0001 |
INFOCOM | 4 |
| 2021 | Enabling Reliable Keyword Search in Encrypted Decentralized Storage with FairnessabstractBlockchain has led the trend of decentralized applications and shown great use beyond cryptocurrencies. Decentralized storage such as Storj and Sia leverages blockchain to establish an open platform for sharing economy, which provides private and reliable file-outsourcing services. However, the ubiquitous keyword search function over encrypted files is yet to be supported. To enable this function, we first apply searchable encryption techniques to the decentralized setting. But this primitive can hardly ensure the service integrity. The reason is that decentralized storage commonly faces severe threats from both clients and service peers. Service peers may return partial or incorrect results, while clients may intentionally slander the service peers to avoid payments. To address these threats, we utilize the smart contract to record the logs of encrypted search (aka evidence) on the blockchain, and devise a fair protocol to handle disputes and issue fair payments. Using a dynamic-efficient searchable encryption scheme as an instantiation, we craft a concrete scheme that preserves encrypted search capability and enforces ecosystem healthiness, so that service peers are incentivized to make real efforts and jointly guarantee service reliability. We implement our scheme in Python and Solidity, and test its search performance and transaction costs on Ethereum. Chengjun Cai, Jian Weng 0001, Xingliang Yuan, Cong Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Towards Private, Robust, and Verifiable Crowdsensing Systems via Public BlockchainsabstractPublic blockchains have emerged as a promising direction in revolutionizing existing data-driven systems relying on centralized service providers. Among others, one kind of such systems is the popular crowdsensing systems which promise convenient data collection and aggregation. Although promising, leveraging public blockchains to build crowdsensing systems is non-trivial and has to overcome several barriers. First, public blockchains are transparent and lack support for data privacy. Second, participants from the open blockchain environment may misbehave in serving crowdsensing applications, like providing invalid data or doing aggregation incorrectly. Further, on-chain processing incurs monetary cost, so simply putting all workload on-chain is highly uneconomical and a delicate joint on-chain and off-chain design is required. In this paper, we take the first research attempt and explore a new design point to bridge public blockchains with crowdsensing systems. We propose a framework for building private, robust, and verifiable blockchain-empowered crowdsensing systems. It features an open service paradigm where blockchain nodes can rent out their computing resources to serve crowdsensing applications, with custom and full-fledged mechanisms to foster a healthy and economical ecosystem and to simultaneously tackle the challenges of data privacy, robustness against misbehaving participants, and service correctness assurance. Extensive experiments demonstrate our designs practicality. Chengjun Cai, Yifeng Zheng 0001, Yuefeng Du 0001, Zhan Qin, Cong Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Building a Secure Knowledge Marketplace Over Crowdsensed Data StreamsabstractIt is increasingly popular to leverage the wisdom of crowd for knowledge discovery and monetization. Among others, crowdsensing with truth discovery has emerged as a promising way for leveraging the crowd wisdom, which can mine reliable knowledge from the generally unreliable sensory data contributed collected from diverse sources. Building a knowledge marketplace based on crowdsensing with truth discovery for knowledge discovery and monetization, however, is non-trivial and has to overcome several challenges. First, the sensory data should be protected as they may carry sensitive information. Second, many real crowdsensing applications usually yield sensory data in a streaming fashion, posing the demand that truth discovery should be conducted over data streams to continuously mine reliable knowledge in each data collection epoch. Third, knowledge monetization should be well treated, fully addressing the practical needs of parties in the monetization ecosystem. In this article, we take the first research attempt and propose a new full-fledged framework for building a secure knowledge marketplace over crowdsensed data streams. Our marketplace supports secure monetization of reliable knowledge mined privately from data streams in crowdsensing applications. Our framework leverages lightweight cryptographic techniques like additive secret sharing to enable privacy-preserving streaming truth discovery, continuously producing reliable knowledge over data streams. For monetization of the learned truth, i.e., knowledge, we resort to the emerging blockchain technology and deliver a tailored and full-fledged design, which promises monetization fairness, knowledge confidentiality, and streamlined processing. Extensive experiments on Amazon cloud and Ethereum blockchain demonstrate the practically affordable performance of our design. Chengjun Cai, Yifeng Zheng 0001, Anxin Zhou, Cong Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Augmenting Encrypted Search: A Decentralized Service Realization with Enforced ExecutionabstractSearchable symmetric encryption (SSE) allows the data owner to outsource an encrypted database to a remote server in a private manner while maintaining the ability for selectively search. So far, most existing solutions focus on an honest-but-curious server, while security designs against a malicious server have not drawn enough attention. A few recent works have attempted to construct verifiable SSE that enables the data owner to verify the integrity of search results. Nevertheless, these verification mechanisms are highly dependent on specific SSE schemes, and fail to support complex queries. A general verification mechanism is desired that can be applied to all SSE schemes. In this work, instead of concentrating on a central server, we explore the potential of the smart contract, an emerging blockchain-based decentralized technology, and construct decentralized SSE schemes where the data owner can receive correct search results with assurance without worrying about potential wrongdoings of a malicious server. We study both public and private blockchain environments and propose two designs with a trade-off between security and efficiency. To better support practical applications, the multi-user setting of SSE is further investigated where the data owner allows authenticated users to search keywords in shared documents. We implement prototypes of our two designs and present experiments and evaluations to demonstrate the practicability of our decentralized SSE schemes. Shengshan Hu, Chengjun Cai, Qian Wang 0002, Cong Wang 0001, Zhibo Wang 0001, Dengpan Ye |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | EncELC: Hardening and Enriching Ethereum Light Clients with Trusted EnclavesabstractThe rapid growth of Ethereum blockchain has brought extremely heavy overhead for coin owners or developers to bootstrap and access transactions on Ethereum. To address this, light client is enabled, which only stores a small fraction of blockchain data and relies on bootstrapped full nodes for transaction retrievals. However, because the retrieval requests are outsourced, it raises several severe concerns about the integrity of returned results and the leakage of sensitive blockchain access histories, largely hindering the wider adoption of this important lightweight design. In addition to security issues, the continuously increasing blockchain storage also urges for more effective query functionalities for the Ethereum blockchain, so as to enable more flexible and precise transaction retrievals.In this paper, we propose EncELC, a new Ethereum light client design that enforces full-fledged protections for clients and enables rich queries over the Ethereum blockchain. EncELC leverages trusted hardware (e.g., Intel SGX) as a starting point for building efficient yet secure processing, and further crafts several crucial performance and security refinement designs to boost query efficiency and conceal leakages inside and outside SGX enclave. We implement a prototype of EncELC and test its performance in several real settings, and the results have confirmed the practicality of EncELC. Chengjun Cai, Lei Xu 0019, Anxin Zhou, Cong Wang 0001, Qian Wang 0002 |
INFOCOM | 1 |
| 2018 | Leveraging Crowdsensed Data Streams to Discover and Sell Knowledge: A Secure and Efficient RealizationabstractLeveraging the wisdom of crowd for knowledge discovery and monetization is increasingly popular nowadays. Among others, one popular way of leveraging the crowd wisdom is crowdsensing with truth discovery, which is able to discover truthful knowledge from the unreliable sensory data harvested from mobile clients. In order to become truly successful, however, a number of challenges are yet to be addressed. First, safeguarding clients' sensory data is demanded for privacy protection. Second, in many real crowdsensing applications, data are usually collected in a streaming manner, so truth discovery is naturally required to be efficiently conducted in a streaming fashion. Thirdly, knowledge monetization should be made full-fledged, endowed with features of transparency and streamlined processing while fully addressing the practical needs of parties in the monetization ecosystem. In this paper, we present our initial effort on a crowdsensing framework that enables privacy-preserving knowledge discovery and full-fledged blockchain-based knowledge monetization. Our framework enables privacy-preserving and efficient truth discovery over encrypted crowdsensed data streams for truthful knowledge discovery. Meanwhile, with careful integration of the newly emerging blockchain-based smart contract technology, our framework allows full-fledged knowledge monetization. Tackling the challenges of monetization fairness and (on-chain) knowledge confidentiality, our customized knowledge monetization design well respects the interests of knowledge seller and requester, with full support of transparency, streamlined processing, and automatic quality-aware rewards for clients. Extensive experiments on Microsoft Azure cloud and Ethereum blockchain demonstrate the practically affordable performance of our design. Chengjun Cai, Yifeng Zheng 0001, Cong Wang 0001 |
ICDCS | 1 |
| 2018 | Searching an Encrypted Cloud Meets Blockchain: A Decentralized, Reliable and Fair RealizationabstractEnabling search directly over encrypted data is a desirable technique to allow users to effectively utilize encrypted data outsourced to a remote server like cloud service provider. So far, most existing solutions focus on an honest-but-curious server, while security designs against a malicious server have not drawn enough attention. It is not until recently that a few works address the issue of verifiable designs that enable the data owner to verify the integrity of search results. Unfortunately, these verification mechanisms are highly dependent on the specific encrypted search index structures, and fail to support complex queries. There is a lack of a general verification mechanism that can be applied to all search schemes. Moreover, no effective countermeasures (e.g., punishing the cheater) are available when an unfaithful server is detected. In this work, we explore the potential of smart contract in Ethereum, an emerging blockchain-based decentralized technology that provides a new paradigm for trusted and transparent computing. By replacing the central server with a carefully-designed smart contract, we construct a decentralized privacy-preserving search scheme where the data owner can receive correct search results with assurance and without worrying about potential wrongdoings of a malicious server. To better support practical applications, we introduce fairness to our scheme by designing a new smart contract for a financially-fair search construction, in which every participant (especially in the multiuser setting) is treated equally and incentivized to conform to correct computations. In this way, an honest party can always gain what he deserves while a malicious one gets nothing. Finally, we implement a prototype of our construction and deploy it to a locally simulated network and an official Ethereum test network, respectively. The extensive experiments and evaluations demonstrate the practicability of our decentralized search scheme over encrypted data. Shengshan Hu, Chengjun Cai, Qian Wang 0002, Cong Wang 0001, Xiangyang Luo 0001, Kui Ren 0001 |
INFOCOM | 2 |
| 2017 | Towards trustworthy and private keyword search in encrypted decentralized storageabstractEmerging decentralized storage services such as Storj and Filecoin show promise as a new paradigm for data outsourcing. These services tie cryptocurrency to personal storage resources and leverage blockchain technology to ensure data integrity in distributed networks. Compared to current cloud storage, they are expected to be more scalable, cost effective, and secure. In addition to the features above, strong guarantees of data privacy are seriously desired due to today's prevalent data leak and abuse incidents. However, simply using end-to-end encryption limits the search capability and thus will degrade the user experience. In this paper, we propose an encrypted decentralized storage architecture that can support trustworthy and private keyword search functions. We start from searchable encryption to achieve search on encrypted data. Yet, only adopting this primitive is not sufficient to address particular threats in our target decentralized service model. Service peers would maliciously return incorrect results, while user peers would fraudulently refuse to pay service fees. To resolve those threats, we devise specific secure data addition and keyword search protocols to enable client-side verifiability and blockchain based fair judgments on the search results. For practical considerations, we integrate an efficient dynamic searchable encryption scheme to our protocols as an instantiation to lower the blockchain overhead. Our security and performance analysis indicates the advance of the proposed architecture. Chengjun Cai, Xingliang Yuan, Cong Wang 0001 |
ICC | 1 |