Youwen Zhu

dblp:98/3177 · DBLP profile ↗
← Back
64ranked-venue papers
14as first author
38since 2021 · last 2026
0000-0003-4365-9713ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 22 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 8 since 2021Computer networks · 10 · 4 first-author · 4 since 2021Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 (ϵ, δ)-local differential privacy mechanisms for set-valued data analysis
Youwen Zhu
Frontiers Comput. Sci.2
2026 PK-Free, Blind and Collusion-Resistant Synthetic Tabular Fingerprinting With Diffusion Models
abstract
High-quality synthetic tabular data raises severe issues about potential misuse. Existing post-processing database fingerprinting schemes relying on either Primary Key (PK) or original database ( non- blind) cannot work well in PK-free table, and easily suffer from regeneration attack using Diffusion Models (DMs). Integrating watermarking into the generation process of tabular data is an effective solution, such as Tree-Ring (TR) and Tabular Watermarking (TabWak) for latent tabular DMs. But, these schemes only address copyright protection and lack strong liability guarantees (fingerprinting) in case of unauthorized redistribution against collusion attack. Thus, we propose a framework ofPK-free,Blind andCollusion-resistant syntheticTabularFingerprinting (PBC-TabFip) with DMs by readily incorporating with symmetric Tardos codes of arbitrary alphabet sizes. PBC-TabFip solves the problem of deleted or modified PK, and maintains a blind property which does not use original data to extract fingerprint. Based on PBC-TabFip, we actually propose binary TabFip and TabFip+, quaternary TabFip* and TabFip+* schemes, where “+” means using self-cloning. To identify malicious user, we use Bit Matching (BM) and Valid Bit Matching (VBM) mechanisms for our schemes. We theoretically deduce that TabFip with BM obtains the highest expected accuracy of fingerprint bit matching against random noises. In the case of binary fingerprint, we also demonstrate the fragility of TabFip without Tardos codes, resulting in high probability of detecting innocent, and stronger robustness of TabFip with Tardos codes against inversion-based collusion attack using different strategies. Experimentally, comparing our schemes with baselines on two datasets, we show that the quality of synthetic table achieved by TabFip is higher than that achieved by other methods (TR under proportion of same rows$\lesssim$10%); and TabFip with BM achieves higher average detecting rate of correct user against five single-handed post-editing attacks. In a practical scenario, we exhibit that TabFip with Tardos codes identifies at least one of the colluders with 100% probability and without detecting innocent against two types of collusion attack.
Shunsheng Zhang, Youwen Zhu, Yonglong Luo
IEEE Trans. Dependable Secur. Comput.2
2026 Decoupled and Privacy-Preserving Key Generation in ABE Under the Minimal Disclosure Principle
abstract
Attribute-Based Encryption (ABE) enables fine-grained access control over outsourced data, but its key generation process typically requires users to disclose their complete attribute sets, introducing significant privacy risks. Existing privacy-preserving approaches—such as those based on zero-knowledge proofs or tightly coupled interactive protocols—suffer from limited scalability, high communication costs, and insufficient support for selective attribute disclosure. To address these limitations, we propose a privacy-enhancing key generation protocol guided by the principle ofMinimal Disclosure, which ensures that users disclose only the minimally necessary subset of attributes required for authorization. Our protocol decouples attribute verification from key issuance: users first obtain cryptographically verifiable attribute tokens, and later issue blinded key requests over selectively chosen attributes. This design enables selective disclosure, supports reusable attribute credentials, and enhances user autonomy. To improve scalability, we introduce a lightweight batch verification mechanism that reduces computation and communication overhead for the attribute authority. We prove that our protocol achieves thebindingandhidingproperties under standard cryptographic assumptions, and we formally verify these guarantees in the symbolic model using the ProVerif tool. In addition, we propose two privacy metrics—AttributeInference Gain (AIG) andPrivacy Gain (PG)—alongside an entropy-based analysis to quantify resistance against attribute inference attacks. Experimental results show that our scheme effectively mitigates inference leakage while offering substantial efficiency gains compared to existing schemes.
Youwen Zhu, Xiaodong Yang 0006, Changhee Hahn, Jian Wang 0038, Junbeom Hur
IEEE Trans. Inf. Forensics Secur.2
2026 Model Inversion Attack Against Federated Unlearning
abstract
In response to emerging regulations on the “right to be forgotten”, federated unlearning (FU) has been proposed to ensure privacy compliance by efficiently eliminating the influence of specific data from federated learning (FL) models. However, existing FU studies primarily focus on improving unlearning efficiency, with little attention given to the potential privacy risks introduced by FU itself. To bridge this research gap, we propose a novel federated unlearning inversion attack (FUIA) to expose potential privacy leakage in FU. This work represents the first systematic study on the privacy vulnerabilities inherent in FU. FUIA can apply to three major FU scenarios: sample unlearning, client unlearning, and class unlearning, demonstrating broad applicability and threat potential. Specifically, the server, acting as an honest-but-curious attacker, continuously records model parameter changes throughout the unlearning process and analyzes the differences before and after unlearning to infer the gradient information of forgotten data, enabling the reconstruction of its features or labels. FUIA directly undermines the goal of FU to eliminate the influence of specific data, exploiting vulnerabilities in the FU process to reconstruct forgotten data, thereby revealing flaws in privacy protection. Moreover, we explore two potential defense strategies that introduce a trade-off between privacy protection and model performance. Extensive experiments on multiple benchmark datasets and various FU methods demonstrate that FUIA effectively reveals private information of forgotten data.
Lei Zhou 0039, Youwen Zhu, Rongke Liu
IEEE Trans. Inf. Forensics Secur.2
2026 Locally Differentially Private Truth Discovery for Sparse Crowdsensing
abstract
Truth discovery has emerged as an effective tool to mitigate data inconsistency in crowdsensing by prioritizing data from high-quality responders. While local differential privacy (LDP) has emerged as a crucial privacy-preserving paradigm, existing studies under LDP rarely explore a worker's participation in specific tasks for sparse scenarios, which may also reveal sensitive information such as individual preferences and behaviors. Existing LDP mechanisms, when applied to truth discovery in sparse settings, may create undesirable dense distributions, provide insufficient privacy protection, and introduce excessive noise, compromising the efficacy of subsequent non-private truth discovery. Additionally, the interplay between noise injection and truth discovery remains insufficiently explored in the current literature. To address these issues, we propose a lOcally differentially private truth diSCovery approach for spArse cRowdsensing, namely OSCAR. The main idea is to use advanced optimization techniques to reconstruct the sparse data distribution and re-formalize truth discovery by considering the statistical characteristics of injected Laplacian noise while protecting the privacy of both the tasks being completed and the corresponding sensory data. Specifically, to address the data density concerns while alleviating noise, we design a randomized response based Bernoulli matrix factorization method BerRR. To recover the sparse structures from densified, perturbed data, we formalize a 0-1 integer programming problem and develop a sparse recovery solving method SpaIE based on implicit enumeration. We further devise a Laplacian-sensitive truth discovery method LapCRH that leverages maximum likelihood estimation to re-formalize truth discovery by measuring differences between noisy values and truths based on the statistical characteristic of Laplacian noise. Our comprehensive theoretical analysis establishes OSCAR's privacy guarantees, utility bounds, and computational complexity. Experimental results show that OSCAR surpasses the state-of-the-arts by at least 30% in accuracy improvement.
Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Youwen Zhu, Zhiquan Liu 0001, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.5
2026 GBC-UG: An Advanced Location Data Distribution Estimation Mechanism Under Geo-Indistinguishability
abstract
The statistical distribution of user geographic location data is widely used in various mobile applications. Although geo-indistinguishability (GI) has emerged as an effective privacy-preserving framework for processing location data, due to the lack of robust perturbation probability calculation and post-processing for eliminating statistical errors caused by random perturbation on user-side data, GI exhibits low accuracy when directly applied to two-dimensional continuous location data distribution estimation. To overcome this, we propose a novel and efficient location data distribution estimation mechanism by improving GI, termed gamma-based circle and uniform grids (GBC-UG). The GBC-UG mechanism consists of two key algorithms: i) the gamma-based circle (GBC) algorithm, which perturbs users' location data and ensures the calculability of perturbation probabilities on the server side, and ii) the uniform grids (UG) algorithm, which post-processes the perturbed data to accurately estimate the original distribution. We provide a theoretical analyses of the upper and lower bounds of the statistical error in distribution estimation and identify optimal parameter values to minimize this error. Experimental results on multiple real-world datasets demonstrate that the proposed GBC-UG mechanism can significantly improve the accuracy of distribution estimation, as well as the prediction accuracy of both the top-k and popularity ranking while guaranteeing user privacy, outperforming existing GI-based approaches.
Cong Tang, Youwen Zhu, Ruoyang Chen, Changyan Yi, Jian Wang 0038
IEEE Trans. Mob. Comput.2
2025 Anomaly Detection for ADS-B Data Based on KAN-LSTM
Lixia Xie, Yazhou Ning, Hongyu Yang 0003, Youwen Zhu, Huiling Hu, Xiang Cheng 0004
Inscrypt (2)4
2025 Fine-Filter: An Effective Defense Against Poisoning Attacks on Frequency Estimation Under LDP
Yuxia Zhou, Qiao Xue, Youwen Zhu
ICICS (1)3
2025 Addressing Sensitivity Distinction in Local Differential Privacy: A General Utility-Optimized Framework
Youwen Zhu, Rongke Liu, Changyu Dong
USENIX Security Symposium2
2025 PrivRS: Differentially private synthetic data generation via role similarity
Xinxin Ye, Youwen Zhu, Hai Deng
Comput. Secur.2
2025 A Verifiable Deletion Protocol for Enhancing Constraints on Public Clouds
abstract
Public cloud environments enable multiparty collaboration and data sharing but are also prone to significant security challenges, particularly regarding verifiable data deletion. Existing tightly coupled protocols fall short in ensuring that deletion behaviors are executed honestly, leaving users unable to verify if their data has been truly erased. In this article, we address this critical issue by proposing a novel protocol paradigm to verify deletion behavior in public clouds. Our approach introduces the concepts of uncertainty requests and uncertainty roles, which obfuscate the cloud’s attack perspective by decoupling the relationship between deletion requests and credential responses. This decoupling prevents the cloud from identifying the requester’s identity or linking credentials to specific deletion requests, thereby imposing strict constraints on deletion behavior and enhancing resistance to unauthorized data retention. We formally define the security properties of our paradigm and provide a concrete instantiation of the protocol. Our security proofs demonstrate that the proposed protocol not only verifies deletion behavior but also resists backup attacks, targeting users, data blocks, and deletion requests. For performance, our experiments demonstrate from both computational efficiency and effectiveness evaluation. The results reveal lower computational overhead and superior security, making it highly suitable for practical deployment in public cloud environments.
Youwen Zhu, Jian Wang 0038, Yan Jiang 0002
IEEE Internet Things J.2
2025 Frequency estimation under relaxed input-discriminative local differential privacy
Xiqi Kuang, Youwen Zhu, Rongke Liu, Shunsheng Zhang
Inf. Sci.2
2025 Non-interactive K-mode clustering of high-dimensional categorical data under local differential privacy
Xinxin Ye, Youwen Zhu, Shunsheng Zhang, Hai Deng
Inf. Sci.2
2025 A two-phase approach for locally differentially private fuzzy co-clustering
Jianlong An, Youwen Zhu
Inf. Sci.4
2025 A Lightweight and Generic Access Rights Update Mechanism for Attribute-Based Encryption in cloud storage
Youwen Zhu, Jian Wang 0038, Junbeom Hur
J. Syst. Archit.2
2025 GFD: An Effective Defense Against Targeted Poisoning Attacks for Local Differential Privacy Frequency Estimation
abstract
Local Differential Privacy (LDP) enables an untrusted server to collect and analyze sensitive data while preserving user privacy. Recent studies reveal that LDP protocols are vulnerable to poisoning attacks, in which an adversary can manipulate aggregated frequencies by controlling malicious users to send forged data to the server. Some countermeasures have been proposed to mitigate poisoning attacks, but they have limitations: 1) requiring prior knowledge of the attack type; 2) exhibiting poor resistance to the adaptive maximal gain attack, i.e., MGA-A. To address the two limitations, in this paper, we propose a novel detection scheme named Group Filter Detection (GFD) to defend against poisoning attacks on LDP frequency estimation. GFD is a universal defense scheme, which can be applied to any LDP frequency estimation protocol without the prior knowledge of attack types, and exhibits high robustness against various poisoning attacks. GFD can first identify the adversary’s target itemset and then filters the suspicious perturbed data (from malicious users). In this way, GFD can exclude malicious data with high confidence, thereby improving the accuracy of LDP frequency estimation. Compared with the existing solutions, experimental results demonstrate the highest effectiveness of GFD.
Youwen Zhu, Shaowei Wang 0003, Qiao Xue, Jian Wang 0038
IEEE Trans. Inf. Forensics Secur.2
2025 Maximizing Area Coverage in Privacy-Preserving Worker Recruitment: A Prior Knowledge-Enhanced Geo-Indistinguishable Approach
abstract
Worker recruitment for area coverage maximization, typically requires participants to upload location information, which can deter potential participation without proper protection. While existing studies resort to geo-indistinguishability to address this concern, they primarily focus on either specific task locations (Target Coverage) or operate under pre-defined recruitment quotas for an interested region (Area Coverage). These focuses not only yield suboptimal area coverage when scaled but also fail to leverage valuable prior knowledge in the form of participants’ noisy historical registered locations, to enhance both location obfuscation and worker identification processes. To address these limitations, we presentWILTON, which optimizes area coverage under geo-indistinguishability by recruiting the minimum number of participants through the strategic utilization of noisy prior knowledge. InWILTON, to generate obfuscated locations, we propose a probabilistic and weight-aware input perturbation mechanism, which groups and weights prior locations rather than using only personal prior locations. To privately identify the recruited workers, we design a grid-based worker identification method, which provides a worst-case performance guarantee of ratio 1 − 1/eto the optimum. We provide a theoretical analysis of the privacy, utility, and complexity guarantees ofWILTON. Experimental results over two real-world datasets and one synthetic dataset show thatWILTONsurpasses the state-of-the-arts by at least 8% in area coverage improvement.
Pengfei Zhang 0010, Xiang Cheng 0003, Zhikun Zhang 0001, Youwen Zhu, Ji Zhang 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Privacy-Preserving Recommendations With Mixture Model-Based Matrix Factorization Under Local Differential Privacy
abstract
Matrix factorization-based recommendations have emerged as a cornerstone for many industrial recommendation algorithms. However, studies employing differential privacy often rely on a trusted server, while those implementing local differential privacy (LDP) frequently encounter substantial accuracy degradation and excessive communication overhead due to noise injection and frequent user interactions. Moreover, the presence of injected noise for privacy protection and inherent Gaussian noise within these perturbed values compounds the accuracy issues and may create a cascading effect under LDP, exacerbating the complexity of the problem at hand. To address these challenges, we proposeMENTOR, which isMixture model-based rEcommeNdations approach with matrix FacTORization under LDP. Its main idea lies in the adoption of a bounded input perturbation mechanism that closely approximates the Laplace distribution while providing rigorous LDP, coupled with accounting for various noise types present in the perturbed data. This approach allows us to reformulate the matrix factorization process with minimal user interaction, requiring only a single round of communication. In particular, to add noise, we design a bidirectional bounded LDP input perturbation mechanismBBVwhile minimizing variance. To generate recommendation results, we devise a matrix factorization techniqueGLMFbased on a Gaussian–Laplacian mixture model. Comprehensive experiments reveal thatMENTORoutperforms the state of the art by at least 15% inRMSEand 12% inF-Measure, showcasing its effectiveness in balancing privacy and utility in recommendation systems.
Pengfei Zhang 0010, Zhikun Zhang 0001, Xiang Cheng 0003, Youwen Zhu, Ji Zhang 0001
IEEE Trans. Ind. Informatics5
2025 Numerical Data Collection Under Input-Discriminative Local Differential Privacy
abstract
Input-discriminative local differential privacy (ID-LDP) protects user data with a different range of values, which improves the utility of the estimated data compared to traditional LDP. However, the existing ID-LDP methods are used for categorical data and cannot be directly applied to numerical data. In this paper, we propose a numerical data collection (NDC) framework with ID-LDP to provide discriminative protection for the data with different inputs. This framework uses a piecewise mechanism to divide the numerical data into several segments and designs two perturbation methods to minimize the mean value of numerical data based on values submitted by users. We first create an NDC-UE method that encodes the raw data into a binary vector. This method sets the uploaded data bit as 1 and the rest as zero and perturbs each bit with a given probability. We further propose an NDC-GRR algorithm to perturb the numerical data with an optimal privacy budget. To reduce the complexity of NDC-GRR, we apply a greedy algorithm-based spanner to shorten the computation time and improve the accuracy. Theoretical analysis proves that our schemes satisfy the definition of ID-LDP. Experimental results based on two real-world datasets and a synthetic dataset show that the proposed schemes have less mean square error compared with the benchmarks.
Youwen Zhu, Shibo Dai, Pengfei Zhang 0010, Xiqi Kuang
IEEE Trans. Knowl. Data Eng.1
2025 LDGI: Location-Discriminative Geo-Indistinguishability for Location Privacy
abstract
Geo-Indistinguishability (GI) is a powerful privacy model that can effectively protect location information by limiting the ability of an attacker to infer a user's true location. In real life, locations usually have different sensitive levels in terms of privacy; for example, shopping malls might be low-sensitive while home addresses might be high-sensitive for users. But the GI model does not consider the various sensitive levels of locations, and implements the same perturbation on all locations to meet the highest privacy requirement. This would cause overprotection of low-sensitive locations and reduce data utility. To strike a good balance between privacy and utility, in this paper, we propose a novel privacy notion, termedLocation-DiscriminativeGeo-Indistinguishability (LDGI), which takes into account different sensitive levels of location privacy. With LDGI model, we then develop a perturbation scheme called EM-LDGI based on the exponential mechanism, and an advance scheme MinQL to further enhance data utility. To improve the efficiency of the proposed schemes, we design a scheme MinQL-S with the assistance of the spanner graph, at the cost of a slight utility degradation. We theoretically analyze that the proposed schemes satisfy LDGI and evaluate their performance by extensive experiments on both synthetic and real datasets. The comparison with GI mechanisms demonstrates the advantages of the LDGI model.
Youwen Zhu, Yuanyuan Hong, Qiao Xue, Xiao Lan, Yushu Zhang 0001, Yong Xiang 0001
IEEE Trans. Knowl. Data Eng.1
2024 Indexing dynamic encrypted database in cloud for efficient secure k-nearest neighbor query
Xingxin Li, Youwen Zhu, Jian Wang 0038, Yushu Zhang 0001
Frontiers Comput. Sci.2
2024 Utility-Enhanced Image Obfuscation With Block Differential Privacy
abstract
With the popularity of cloud servers, an increasing number of people utilize cloud platforms to manage their images. Images uploaded to the cloud are no longer directly controlled by the owner, which has raised concerns among users about privacy leakage. In response, some thumbnail-preserving encryption methods have been proposed to protect image privacy. However, since privacy lacks a clear definition, these methods face challenges in quantifying privacy, making it difficult for users to assess the effectiveness of privacy protection. To address this issue, we propose an image obfuscation method based on novel differential privacy mechanism, which provides provable privacy protection while preserving thumbnails. Additionally, our method can maintain the consistency of image distributions to enhance the utility of protected images. Extensive experimental results demonstrate that our method significantly outperforms existing work in terms of privacy and utility.
Junhao Ji, Tao Wang 0084, Youwen Zhu
IEEE Signal Process. Lett.4
2024 A Consumer-Oriented Image Transformation Scheme With a Secret Key for Privacy Protection
abstract
Images in electronic devices may pose privacy threats since they can capture sensitive information about consumers. Meanwhile, face recognition (FR) systems are widely used, exacerbating these concerns. Some learning-based schemes have been proposed to protect the privacy of facial images. However, consumers might encounter challenges in implementing them due to specific requirements related to computing power and professional background. The non-learning semantic adversarial perturbation schemes address the aforementioned issues, but they are either irreversible or necessitate additional storage space. To this end, we propose a consumer-oriented image transformation scheme that can prevent the recognition of images by FR systems. It offers a consumer-friendly method compared to schemes that need high-performance equipment and specialized knowledge. The proposed scheme is implemented by reversible embedding the key to coefficients during the JPEG compression. The proposed scheme can be operated in general computing devices by ordinary consumers. The experiments demonstrated that facial images can be protected by the proposed scheme.
Wenying Wen, Youwen Zhu, Rushi Lan
IEEE Signal Process. Lett.3
2024 Understanding Visual Privacy Protection: A Generalized Framework With an Instance on Facial Privacy
abstract
With the widespread application of computer vision, the scenarios in terms of visual privacy have become increasingly diverse and meanwhile numerous studies have been conducted to address privacy concerns in these scenarios. However, these studies are individually tailored for specific scenarios, making their layouts challenging to be drawn upon easily. When encountering a new scenario, it takes significant additional efforts to redesign a scheme due to the low referability of previous works. To tackle this issue, we explore commonalities among existing works and propose a generalized framework to meet the demand for visual privacy protection in various scenarios. Our framework is elaborately organized into several crucial steps, including privacy definition, scenario abstraction, algorithm design, and effect evaluation. It serves as a guide for researchers to efficiently design visual privacy protection schemes. In our framework, we establish a unified standard for quantifying privacy and introduce a novel constrained optimization theory to balance privacy and usability, which contributes to a broader understanding of visual privacy protection. Furthermore, we present an instance under the guidance of the framework that can support identity protection and attribute control scenarios through a diffusion-based model. Extensive experimental results demonstrate the effectiveness of our framework.
Yushu Zhang 0001, Junhao Ji, Wenying Wen, Youwen Zhu, Zhihua Xia, Jian Weng 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Collusion-Resilient Privacy-Preserving Database Fingerprinting
abstract
Database sharing may bring about privacy disclosure and illegal redistribution. A previously proposed entry-level Differential Privacy FingerPrinting mechanism (DPFP) for relational database achieves privacy and liability guarantees simultaneously. However, it is only robust against common attacks from a vicious Data Analyzer (DA) and lacks robustness against logical AND or OR collusion attack even if Anti-Collusion Code (ACC) is used to trace who the colluders are. In this work, we propose a Collusion-Resilient entry-level DP FingerPrinting mechanism (CRDPFP) for uniquely identifying colluders by directly using ACCs. Specifically, we firstly theoretically and experimentally demonstrate the vulnerabilities of existing fingerprinting schemes by identification of logical AND/OR collusion attack. To survive 5 types of collusion attacks and identify colluders, a Group-oriented Concatenated (GC) ACC based on I-code and Cover Free Family code is constructed and a catch-all detector is designed. By leveraging the randomization nature of fingerprint, we transform GC code into provable entry-level DP guarantees on the entire database. We also show that CRDPFP inherits the same connection properties between privacy, fingerprint robustness, and database utility from DPFP. Via experiments on two real-world relational databases, we exhibit that our mechanism supplies stronger robustness against 50% random flipping attack from a vicious DA, achieves higher and lower detecting rates of at least one colluder and innocent, uniquely traces all colluders for logical AND or OR collusion attack and obtains near-optimal utility with fingerprint parameter being close to 2 compared to existing schemes.
Shunsheng Zhang, Youwen Zhu, Ao Zeng
IEEE Trans. Inf. Forensics Secur.2
2024 Heavy Hitter Identification Over Large-Domain Set-Valued Data With Local Differential Privacy
abstract
Set-valued data are widely used to represent information in the real word, such as individual daily behaviors, items in shopping carts and web browsing history. By collecting set-valued data and identifying heavy hitters, service providers (i.e., the collector) can learn usage preferences of costumers (i.e., users), and improve the quality of their services by the learned information. However, the collection of raw data would bring privacy risks to users. Recently, local differential privacy (LDP) has emerged as a rigorous privacy framework for user private data collection. At the same time, many LDP schemes have been designed to achieve heavy hitters, but most of them are limited by the large data domain due to the huge computation cost. In this paper, we propose an LDP framework: PemSet, to efficiently identify heavy hitters from set-valued data with a large domain. In PemSet, users mainly focus on the prefix of each item (i.e., the first few bits of the binary expression of each item), and only perturb and report prefixes to reduce computation cost. Sometimes the prefixes of different items are the same, so the reported set-valued data could be a multiset, i.e., a set including multiple same items. As such, we design four LDP protocols MOLH, MOLH-S, MPCKV, MWheel to estimate frequencies of items in the multiset setting, and compare their performance under PemSet framework by experiments. Experimental results demonstrate that MOLH can perform the best in a high privacy region, i.e.,$\epsilon < 1$, while MWheel can obtain the highest utility when privacy budget is large, i.e.,$\epsilon \geqslant 1$.
Youwen Zhu, Yiran Cao, Qiao Xue, Qihui Wu 0001, Yushu Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2023 DIVRS: Data integrity verification based on ring signature in cloud storage
Yushu Zhang 0001, Youwen Zhu, Liangmin Wang 0001, Yong Xiang 0001
Comput. Secur.3
2023 Fully distributed identity-based threshold signatures with identifiable aborts
Yan Jiang 0002, Youwen Zhu, Jian Wang 0038, Xingxin Li
Frontiers Comput. Sci.2
2023 Differential Privacy Data Release Scheme Using Microaggregation With Conditional Feature Selection
abstract
Differential privacy (DP) has achieved great progress in addressing the user privacy preservation issues related to data analysis in the Internet of Things (IoT) services and applications. However, the existing DP models tend to overlook the effect of the correlation of data features on the utility of the smart IoT data. To mitigate this gap, we propose a Microaggregation-based DP method using Conditional Mutual Information (M-DPCMI) for prerelease data processing and feature selection. With the new method, we leverage the anonymous microaggregation approach to improve data utility while preventing potential sensitive IoT user information leakage. In addition, M-DPCMI is theoretically proved to satisfy the definition of DP and experimentally validated over real data sets. It is shown that the new model achieves better data utility than the state-of-the-art DP methods.
Xinxin Ye, Youwen Zhu, Hai Deng
IEEE Internet Things J.2
2023 Avalon: A Scalable and Secure Distributed Transaction Ledger Based on Proof-of-Market
abstract
Blockchain technology has gained widespread use. However, it faces several challenges including throughput, transaction delay, security, and decentralization. This paper presents the Avalon protocol based on a novel Proof-of-Market (PoM) consensus mechanism to address these issues. PoM is a type of Proof-of-Work (PoW) consensus that incorporates market-driven leader election and shifts PoW from mining pools to consumers based on transactions. The matching incentive mechanism makes PoM incentive compatible. PoM decouples the scalability and security of Bitcoin, which means that Avalon can optimize the capacity and interval of blocks without compromising other performance goals. Our analysis shows that Avalon can tolerate malicious nodes possessing up to$\bf{1/3}$of the network's total computational power. Furthermore, the implementation of Avalon is similar to Bitcoin and is highly concise. We evaluate the performance of Avalon through a simulated network of over$\bf{1,000}$nodes. Experimental results demonstrate that Avalon can achieve a throughput of$\bf{4,000}$TPS (transactions per second), which is significantly better than state-of-the-art schemes ($\bf{10\boldsymbol{\times}}$Bitcoin-NG,$\bf{5\boldsymbol{\times}}$ByzCoin, and$\bf{4\boldsymbol{\times}}$Algorand). Additionally, it has a transaction confirmation delay of up to$\bf{40}$s, which is twice better than Bitcoin-NG and ByzCoin while experiencing only minimal blockchain splits and maintaining excellent decentralization.
Weilin Chen 0002, Wei Yang 0011, Lide Xue, Bingren Chen, Youwen Zhu, Liusheng Huang
IEEE Trans. Computers5
2023 RAPP: Reversible Privacy Preservation for Various Face Attributes
abstract
The tremendous progress in deep learning has enabled to extract soft-biometric attributes from faces, which raises privacy concerns over images collected for face recognition. Advances toward attribute privacy have been able to conceal multiple attributes while preserving identity information but suffer from limitations: they 1) only consider a few soft-biometric attributes and 2) fail to support reversibility for attribute privacy preservation. To break these limitations, we design a reversible privacy-preserving scheme for various face attributes, called reversible attribute privacy preservation (RAPP). RAPP benefits from two modules: 1) The attribute obfuscator introduces a stream cipher to determine that special attributes have to be concealed with the user-defined password, which also supports recovering original attributes. 2) The attribute adversarial network is proposed to generate perturbed images that conceal various attributes while retaining the utility of face verification. In addition, when a wrong password is provided, the returned image with wrong attribute classification results still keeps realistic, which confuses an attacker to know whether the recovery is correct. Extensive experiments demonstrate that RAPP enables to conceal various attributes and recover original images while facilitating face verification.
Yushu Zhang 0001, Tao Wang 0084, Wenying Wen, Youwen Zhu
IEEE Trans. Inf. Forensics Secur.5
2023 DDRM: A Continual Frequency Estimation Mechanism With Local Differential Privacy
abstract
Many applications rely on continual data collection to provide real-time information services, e.g., real-time road traffic forecasts. However, the collection of original data brings risks to user privacy. Recently, local differential privacy (LDP) has emerged as a private data collection framework for mass population. However, for continual data collection, existing LDP schemes, e.g., those employing the memoization technique, are known to have privacy leakage on data change points over time. In this paper, we propose a new scheme with stronger privacy guarantee for continual frequency estimation under LDP, namely, Dynamic Difference Report Mechanism (DDRM). In DDRM, we introduce difference trees to capture the data changes over time, which well addresses possible privacy leakage on data change points. As for the utility enhancement, DDRM exploits the common case of no data change in time series and thereby suppresses the consumption of privacy budget in such cases. Meanwhile, an optimal privacy budget allocation scheme is proposed to encourage users to report more data for better estimation accuracy. By both theoretical analysis and experimental evaluations, we show DDRM achieves highly accurate frequency estimation in real time.
Qiao Xue, Qingqing Ye 0001, Haibo Hu 0001, Youwen Zhu, Jian Wang 0038
IEEE Trans. Knowl. Data Eng.4
2023 FingerChain: Copyrighted Multi-Owner Media Sharing by Introducing Asymmetric Fingerprinting Into Blockchain
abstract
Nowadays, more and more people are engaged in media creation and sharing for income. A common way to earn income is to upload media to an intermediary platform and then let the platform distribute some profits. However, intermediary platforms generally not only extract most of the profits, but also lack transparency in their operation, where owners lose direct control over the media. Nevertheless, individual sharing is not feasible for owners because each owner holds too little media to attract enough users independently. Blockchain is a solution to the above problem by gathering media from multiple owners without intermediaries. Though a lot of works have studied the use of blockchain for decentralized management of media data, many of them either did not consider sharing needs or tracing the illegal redistribution by malicious users. As for other works, most of them adopted symmetric digital watermarking in their blockchain networks, and thus fail to protect the rights of users who may be framed by malicious owners. Although asymmetric watermarking has been used by two existing works, the owner-side embedding pattern results in low owner-side efficiency. In view of this, we design a media sharing blockchain network in which the asymmetric fingerprinting (i.e., watermarking) with user-side embedding is introduced. Besides superior owner-side efficiency, our scheme also outperforms the above two ones in terms of TTP-free. Moreover, our scheme is designed to offer a user-friendly experience and support record traceability. The performance of our scheme is verified by both theoretical and experimental evaluations.
Xiangli Xiao, Yushu Zhang 0001, Youwen Zhu, Pengfei Hu 0001, Xiaochun Cao
IEEE Trans. Netw. Serv. Manag.3
2023 Shuffle Differential Private Data Aggregation for Random Population
abstract
Bridging the advantages of differential privacy in both centralized model (i.e., high accuracy) and local model (i.e., minimum trust), the shuffle privacy model has potential applications in many privacy-sensitive scenarios, such as mobile user data aggregation and federated learning. Since messages from users are anonymized by semi-trusted shufflers (e.g., anonymous channels, edge servers), every user could hide message among other users’ messages and inject only part of noises (a.k.a. privacy amplification). However, existing works assume that the participating user population is known in advance, which is unrealistic for dynamic environments (e.g., mobile computing, vehicular networks). In this work, we study the shuffle privacy model with a random participating population, and give privacy amplification bounds for population size with commonly encountered binomial, Poisson, sub-Gaussian distribution and etc. For further improving accuracy, we formulate and derive optimal dummy sizes for both non-adaptive and adaptive dummies. Finally, to break the error barrier due to the constraint of sending one single message per user, we design a multi-message shuffle private protocol supporting random population. Experiment results show that our approaches reduce more than 60% error when compared to the local model and naive approaches. We hope this work provides tailored solutions of shuffle privacy for dynamic mobile/distributed computing.
Shaowei Wang 0003, Xuandi Luo, Yuqiu Qian, Youwen Zhu, Kongyang Chen, Qi Chen 0024, Bangzhou Xin, Wei Yang 0011
IEEE Trans. Parallel Distributed Syst.4
2022 Mean estimation over numeric data with personalized local differential privacy
Qiao Xue, Youwen Zhu, Jian Wang 0038
Frontiers Comput. Sci.2
2021 VAGA: Towards Accurate and Interpretable Outlier Detection Based on Variational Auto-Encoder and Genetic Algorithm for High-Dimensional Data
abstract
The curse of dimensionality in high-dimensional data makes it difficult to capture the abnormality of data points in full data space. To deal with this problem, we propose an outlier detection model based on Variational Autoencoder and Genetic Algorithm for subspace outlier analysis of high-dimensional data (VAGA). The proposed VAGA model constructs a variational autoencoder (VAE) to preliminarily detect outliers. Then the genetic algorithm (GA) is used to search the abnormal subspace of the outliers obtained by the VAE layer to provide a basis for subspace outlier analysis. The subsequent clustering of the abnormal subspaces help filter out the false positives which are fed back to the VAE layer to adjust network weights. The comparative experiments performed on three public benchmark datasets show that the outlier detection results of the proposed VAGA model are highly interpretable and have better accuracy performance than the state-of-the-art outlier detection methods.
Jiamu Li, Ji Zhang 0001, Jian Wang 0038, Youwen Zhu, Mohamed Jaward Bah, Gaoming Yang, Yuquan Gan
IEEE BigData4
2021 Locally differentially private distributed algorithms for set intersection and union
Qiao Xue, Youwen Zhu, Jian Wang 0038, Xingxin Li, Ji Zhang 0001
Sci. China Inf. Sci.2
2021 Privacy-Assured FogCS: Chaotic Compressive Sensing for Secure Industrial Big Image Data Processing in Fog Computing
abstract
In the age of the industrial big data, there are several significant problems such as high-overhead data acquisition, data privacy leakage, and data tampering. Fog computing capability is rapidly expanding to address not only network congestion issues but data security issues. This article presents a chaotic compressive sensing (CS) scheme for securely processing industrial big image data in the fog computing paradigm, called privacy-assured FogCS. Specially, the sine logistic modulation map is used to drive the privacy-assured, authenticated, and block CS for secure image data collection in the sensor nodes. After sampling, the measurements are normalized in the fog nodes. The normalized measurements can achieve the perfect secrecy and their energy values are further masked through the proposed permutation-diffusion architecture. Finally, these relevant data are transmitted to the clouds (data centers) for storage, reconstruction, and authentication if required. In addition, a hardware implementation reference on a field programmable gate array is designed. Simulation analyses show the feasibility and efficiency of the privacy-assured FogCS scheme.
Yushu Zhang 0001, Ping Wang 0029, Hui Huang 0008, Youwen Zhu, Di Xiao 0001, Yong Xiang 0001
IEEE Trans. Ind. Informatics4
2020 A Block-Level RNN Model for Resume Block Classification
abstract
Resume block classification is the most significant step in resume information extraction. However, the existing algorithms applied to resume block classification are all the general text classification algorithms, which failed to consider the contextual order of each block within a resume. In order to improve the performance of resume block classification, we propose in this paper a block-level bidirectional recurrent neural network model that makes full use of the contextual order relationship among different resume blocks. The experimental results show that the average F1-score value of our model on three 1,400 real resume datasets is 6% to 9% higher than the existing methods.
Qiqiang Xu, Ji Zhang 0001, Youwen Zhu, Bohan Li 0001, Donghai Guan, Xin Wang 0030
IEEE BigData3
2020 A survey of authenticated key agreement protocols for multi-server architecture
Inam ul Haq, Jian Wang 0038, Youwen Zhu, Saad Maqbool
J. Inf. Secur. Appl.3
2020 Secure two-factor lightweight authentication protocol using self-certified public key cryptography for multi-server 5G networks
Inam ul Haq, Jian Wang 0038, Youwen Zhu
J. Netw. Comput. Appl.3
2020 Efficient authentication protocol with anonymity and key protection for mobile Internet users
Yan Jiang 0002, Youwen Zhu, Jian Wang 0038, Yong Xiang 0001
J. Parallel Distributed Comput.2
2020 Privacy-preserving k-means clustering with local synchronization in peer-to-peer networks
Youwen Zhu, Xingxin Li
Peer-to-Peer Netw. Appl.1
2020 Cloud-assisted secure biometric identification with sub-linear search efficiency
Youwen Zhu, Xingxin Li, Jian Wang 0038
Soft Comput.1
2019 Improved collusion-resisting secure nearest neighbor query over encrypted data in cloud
abstract
Summary Securely performing nearest neighbor query over encrypted data in cloud is an important topic in the area of cloud computing, for which Wang et al recently put forward a scheme (ie, CloudBI‐II) to address the challenging security problem: resisting the collusion of cloud server and query users. In this paper, we propose an efficient attack method that indicates CloudBI‐II will reveal the difference vectors under the collusion attack. Furthermore, we show that the difference vector disclosure will result in serious privacy breach and, thus, attain an efficient attack method to break CloudBI‐II. Namely, CloudBI‐II cannot achieve their declared security. Through theoretical analysis and experiment evaluation, we confirm that our proposed attack approach can fast recover the original data from the encrypted data set in CloudBI‐II. Finally, we provide an enhanced scheme that can efficiently resist the collusion attack.
Youwen Zhu, Xingxin Li, Hongyang Yan, Jing Li 0045
Concurr. Comput. Pract. Exp.1
2019 Efficient and secure multi-dimensional geometric range query over encrypted data in cloud
Xingxin Li, Youwen Zhu, Jian Wang 0038, Ji Zhang 0001
J. Parallel Distributed Comput.2
2018 A Genetic Algorithm Based Technique for Outlier Detection with Fast Convergence
Ji Zhang 0001, Zewen Hu, Hongzhou Li, Liang Chang 0003, Youwen Zhu, Jerry Chun-Wei Lin, Yongrui Qin
ADMA6
2018 Efficiently and securely harnessing cloud to solve linear regression and other matrix operations
Lu Zhou 0002, Youwen Zhu, Kim-Kwang Raymond Choo
Future Gener. Comput. Syst.2
2018 Secure multi-label data classification in cloud by additionally homomorphic encryption
Yu Luo 0004, Youwen Zhu, Xingxin Li
Inf. Sci.3
2018 FTP: An Approximate Fast Privacy-Preserving Equality Test Protocol for Authentication in Internet of Things
abstract
Privacy-preserving string equality test is a fundamental operation of many algorithms, including privacy-preserving authentication in Internet of Things (IoT). Existing secure equality test schemes can theoretically achieve string equality comparison and preserve the private strings. However, they suffer from heavy computation and communication cost, especially while the strings are of hundreds of bits or longer, which is not suitable for IoT applications. In this paper, we propose an approximate Fast privacy-preserving equality Test Protocol (FTP), which can securely complete string equality test and achieve high running efficiency at the cost of little accuracy loss. We strictly analyze the accuracy of our proposed scheme and formally prove its security. Additionally, we leverage extensive simulation experiments to evaluate the running cost, which confirms our high efficiency; for instance, our proposed FTP can securely compare two 256 -bit strings within 0.7 seconds on ordinary laptops.
Youwen Zhu, Jiabin Yuan, Xianmin Wang
Secur. Commun. Networks1
2018 On the Soundness and Security of Privacy-Preserving SVM for Outsourcing Data Classification
abstract
Recently, Rahulamathavan et al. propose a privacy preserving scheme for outsourcing SVM classification. Their core contribution is a secure protocol to attain the sign of numbers in encrypted form. In this paper, we observe that Rahulamathavan et al.'s protocol will suffer from some soundness and security problems. Then, we propose a new scheme to securely obtain the encrypted numbers' sign. Theoretical analysis and experiment results show our proposed scheme can not only fix the soundness and security problems, but also achieve higher efficiency.
Xingxin Li, Youwen Zhu, Jian Wang 0038, Zhe Liu 0001, Yining Liu 0001, Mingwu Zhang
IEEE Trans. Dependable Secur. Comput.2
2017 Distributed Set Intersection and Union with Local Differential Privacy
abstract
Privacy-preserving distributed set intersection and union have been widely applied in many scenarios and lots of work has paid attention to the problem. Existing solutions to privacy-preserving set intersection and union are built on secure multiparty computation protocols, which can theoretically solve it, but result in heavy computation and communication overhead. Worse still, most of the existing schemes cannot work once some participant fails. In this paper, we propose two differentially private approaches for distributed set intersection and union, respectively. In our schemes, each data contributor possesses a secret data set and perturbs it by randomized response technique to satisfy local differential privacy. Then the collector gathers all contributors' perturbed data sets and utilizes maximum likelihood estimation to gain an accurate estimation of intersection and union. Compared to existing schemes, the proposed schemes can dramatically reduce computation and communication overhead, and tolerate participant's failure. We formally prove that the proposed schemes satisfy local differential privacy, and leverage extensive experiments to evaluate the proposed approaches. The results indicate that our schemes have low computation and communication complexity, strong robustness and good utility.
Qiao Xue, Youwen Zhu, Jian Wang 0038, Xingxin Li
ICPADS2
2017 Secure Multi-label Classification over Encrypted Data in Cloud
Xingxin Li, Youwen Zhu, Jian Wang 0038, Zhe Liu 0001
ProvSec3
2017 Efficient k-NN query over encrypted data in cloud with limited key-disclosure and offline data owner
Youwen Zhu, Aniello Castiglione
Comput. Secur.2
2016 On efficiently harnessing cloud to securely solve linear regression and other matrix operations
abstract
In this paper, we propose a new efficient solution for securely outsourcing linear regression to a public cloud with robust answer verification. Additionally, we show our construction can be utilized to efficiently and securely outsource other large-scale matrix operations, such as determinant computation.
Youwen Zhu, Zhikuan Wang, Jian Wang 0038
IWQoS1
2016 Collusion-resisting secure nearest neighbor query over encrypted data in cloud, revisited
abstract
It is a challenging problem to securely resist the collusion of cloud server and query users while implementing nearest neighbor query over encrypted data in cloud. Recently, CloudBI-II is put forward to support nearest neighbor query on encrypted cloud data, and declared to be secure while cloud server colludes with some untrusted query users. In this paper, we propose an efficient attack method which indicates CloudBI-II will reveal the difference vectors under the collusion attack. Further, we show that the difference vector disclosure will result in serious privacy breach, and thus attain an efficient attack method to break CloudBI-II. Namely, CloudBI-II cannot achieve their declared security. Through theoretical analysis and experiment evaluation, we confirm our proposed attack approach can fast recover the original data from the encrypted data set in CloudBI-II. Finally, we provide an enhanced scheme which can efficiently resist the collusion attack.
Youwen Zhu, Zhikuan Wang, Jian Wang 0038
IWQoS1
2016 Secure k-NN Query on Encrypted Cloud Data with Limited Key-Disclosure and Offline Data Owner
Youwen Zhu, Zhikuan Wang
PAKDD (2)1
2016 Secure Naïve Bayesian Classification over Encrypted Data in Cloud
Xingxin Li, Youwen Zhu, Jian Wang 0038
ProvSec2
2016 Secure and controllable k-NN query over encrypted cloud data with key confidentiality
Youwen Zhu, Tsuyoshi Takagi
J. Parallel Distributed Comput.1
2015 Fast Secure Scalar Product Protocol with (almost) Optimal Efficiency
Youwen Zhu, Zhikuan Wang, Bilal Hassan, Jian Wang 0038
CollaborateCom1
2015 Hamburger attack: A collusion attack against privacy-preserving data aggregation schemes
abstract
Performing efficient data aggregation while keeping the property of privacy preservation of user-related data is of high concern. Extensive research has been conducted to address this problem in multiple areas. As a fact, the security of a number of privacy-preserving protocols are threatened by collusion between participants. Therefore, security analysis plays a fundamental role in privacy-preserving data aggregation protocols. In this paper, we present a new kind of collusion attack strategy called Hamburger Attack. It is of logical simplicity, but has the advantage of being effective and efficient. We employ it to check the security of several existing privacy-preserving data aggregation schemes. We show that under hamburger attack, some prior privacy-preserving data aggregation protocols will disclose part, even all, of the private data that they intended to protect. On the other hand, the hamburger attack is beneficial to designing new privacy-preserving data aggregation schemes. It can assist them in avoiding this kind of collusion attack.
Wei Yang 0011, Liusheng Huang, Mingjun Xiao, Xiaorong Lu, Youwen Zhu
IWQoS7
2015 A cheater identifiable multi-secret sharing scheme based on the Chinese remainder theorem
abstract
Abstract There are many researches on the polynomial‐based verifiable (k,n) multi‐secret sharing scheme (VMSSS), but none of them focuses on the Chinese remainder theorem (CRT)‐based VMSSS so far. For the first time, we provide a cheater identifiable multi‐secret sharing scheme based on CRT as an alternative method for VMSSS, which is unconditionally secure when the number of cheaters t≤(k − 1)/3. We adopt an encoding method, which makes multiple secrets to be transferred as a single one. In addition, we utilize a single keyed message authenticated code (MAC) to detect and identify cheaters in the reconstruction phase. Then, combine these two methods with a CRT‐based Asmuth‐Bloom's SSS to achieve our design goals. In our scheme, all participants share a single key of MAC rather than each participant possesses an independent key to check the validity of shares, and the size of share is independent in any of n,k, and t. Analyses show that our scheme is more efficient and secure than existing ones. Finally, as an example of the practical impact of our work, we present how our techniques can be applied to secure sum computation. Copyright © 2015 John Wiley & Sons, Ltd.
Zhenhua Chen 0001, Youwen Zhu, Xinli Xu
Secur. Commun. Networks3
2015 On the Security of A Privacy-Preserving Product Calculation Scheme
abstract
Recently, Jung and Li [1] propose a highly efficient privacy-preserving product calculation scheme without requiring secure communication channels. Then, they present secure approaches to solve several application problems using the product calculation protocol. In this work, we observe several security flawsin their privacy-preserving product calculation scheme and some application protocols. We show the security vulnerabilities will result in the disclosure of private data. In two application protocols, almost all the private numbers will be revealed. We further suggest solutions to fix the security problems.
Youwen Zhu, Liusheng Huang, Tsuyoshi Takagi
IEEE Trans. Dependable Secur. Comput.1
2010 Relation of PPAtMP and scalar product protocol and their applications
abstract
Scalar product protocol and privacy preserving add to multiply protocol (PPAtMP) are two significant basic secure multiparty computation protocols. In this paper, we claim that the two protocols are equivalent to each other and we can achieve one based on the other with the same communication and computation complexity. Then, we propose Secure Two-party Mean Protocol, Secure Shared x ln x Protocol and Secure Shared Generic Polynomial Protocol based on scalar product protocol and PPAtMP. Additionally, we analyze the correctness, security, communication overheads and computation complexity of each protocol proposed in this paper.
Youwen Zhu, Liusheng Huang, Wei Yang 0011
ISCC1