Xiaojie Zhu

dblp:148/4491 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 21 · 6 first-author · 20 since 2021Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Privacy-preserving for user-uploaded images and text in Vision-Language Models
Zixiang Liu, Chi Chen 0001, Shuguang Yuan 0003, Weilong Huang, Xiaojie Zhu, Peizhuo Lv
Comput. Secur.5
2026 Detecting photo-taking actions in surveillance videos based on CPU-only devices
abstract
Abstract Taking photos of sensitive facilities and sensitive information in no photography area may cause sensitive information leakage if not discovered in time. Employing action recognition models to detect instances of photography can effectively prevent information leakage. Current action recognition models have shown unsatisfactory performance in detecting photo-taking actions in surveillance videos, and their reliance on GPU devices hinder their practicality. This paper presents a novel approach to address the detection of photo-taking actions. The method utilizes object detection to filter out background data and incorporates human pose estimation to extract human skeleton data. By combining these AI techniques, the method enables accurate recognition of photo-taking actions. We introduce a novel technique called self-annotation that enables the model to focus on the crucial elements associated with photo-taking actions. Additionally, we introduce a new alarm mechanism that leads to a 69 $$\%$$ % reduction in false positives while maintaining the same level of recall by integrating the labels over a period to recognize actions. Compared with traditional action recognition approaches, our method is more flexible and lightweight in actual engineering applications. Moreover, our model is capable of running on CPU-only devices. Experimental results show that our model achieves a precision of 91 $$\%$$ % on our dataset.
Zixiang Liu, Peisong Shen, Chi Chen 0001, Shuguang Yuan 0003, Xiaojie Zhu, Houzhe Wang
Cybersecur.5
2026 Multi-dimensional fine-grained modeling and assessment for non-stationary fatigue processes
Ruimin Hu, Xiaojie Zhu, Dongliang Zhu 0001
Pattern Recognit.3
2026 BinEnhance-Pro: Enhancing Binary Code Search by Distinguishing Similar but Non-Homologous Functions
Yongpan Wang, Siyuan Li 0014, Xiaojie Zhu, Xiaodong Gu 0002, Yongle Chen
IEEE Trans. Dependable Secur. Comput.5
2026 Minos:Bringing Accountability and Traceability to Attribute-Based Keyword Search
abstract
The rapid adoption of cloud computing has intensified concerns over data privacy, particularly in the context of outsourced datasets. Attribute-Based Keyword Search (ABKS) has emerged as a promising cryptographic solution for enabling secure and efficient search over encrypted data. However, most existing ABKS schemes typically assume that all clients behave honestly, which weakens their applicability in real-world scenarios. In this work, we challenge this assumption by introducing ABKSUT, a new framework that incorporates user accountability and traceable ownership into ABKS. We formalize ABKSUT and present a concrete instantiation, Minos, which properly embeds the user's identity in the trapdoor and the data owner's identity in the ciphertext, enabling secure search, fine-grained access control, and post-hoc accountability. We formally prove the security and correctness of Minos and evaluate its performance using real-world datasets. Results show that Minos achieves strong security guarantees with practical efficiency, making it suitable for real-world deployment.
Xiaojie Zhu, Paulo Veríssimo, Willy Susilo
IEEE Trans. Dependable Secur. Comput.1
2026 An Unbiased and Robust Privacy-Preserving Fingerprinting Scheme for Relational Databases
abstract
Sharing relational databases is essential in today’s data-driven world for fostering collaboration, enhancing efficiency, and enabling real-time data access. However, privacy and copyright issues arise when sharing privacy-sensitive or valuable data. Additionally, high utility is required in shared data to enable accurate data mining and analysis. Entry-level differentially private fingerprinting schemes (DPFS) could address these concerns. In a DPFS, data can be securely shared without leaking original values while still supporting accurate analysis. Moreover, detectable fingerprints can deter unauthorized redistribution. However, existing DPFSs often lack utility—due to format changes and entry-wise bias—or robustness, as fingerprints can be removed undetected. In this paper, we propose an unbiased and robust differential privacy-based fingerprinting scheme (DPFS), which ensures that the fingerprinted copy remains an unbiased estimate of the original data. By incorporating differential privacy noise, our scheme effectively mitigates alteration, collusion, and hybrid attacks. Our DPFS satisfies ϵ-entry-level differential privacy, enabling clients to conduct unbiased analysis. To improve robustness, we design group-based fingerprint detection, which estimates the mean of injected noise per group with error tolerance. We provide a theoretical robustness analysis and propose a method for achieving optimal robustness. Experiments on four real-world databases show that our scheme consistently detects fingerprints and improves accuracy by up to 20% on machine learning tasks compared to existing DPFSs.
Shujie Cui, Hui Cui 0001, Jiabao Qiu, Shuguang Yuan 0003, Xiaojie Zhu, Jing Yu 0007, Chi Chen 0001, Xun Yi
IEEE Trans. Inf. Forensics Secur.6
2025 MDFG: Multi-Dimensional Fine-Grained Modeling for Fatigue Detection
abstract
Fatigue is a critical factor contributing to accidents in industries such as safety monitoring and engineering construction. Fatigue exhibits dynamic complexity and non-stationary characteristics, so there are many intermediate states of short-term variation between alert and fatigue. Capturing and learning the signs of these intermediate states is essential for accurate fatigue assessment. However, current fatigue detection methods primarily rely on coarse-grained labels, typically spanning minutes to hours, and commonly treat alert and fatigue as two distinctly separate distributions, overlooking the expression of intermediate states and oversimplifying the rich distribution information of fatigue types and levels, thereby limiting detection effectiveness. To address these, this paper explores a refined representation of fatigue in terms of three dimensions: time, type, and level, and proposes a Multi-Dimensional Fine-Grained Modeling for Fatigue Detection (MDFG). This introduces the SmallLoss to extract trustworthy samples, utilizes clustering to identify diverse subtypes under alert and fatigued states, and establishes base class sets in each state. Subsequently, a complete base class set containing intermediate state bases is constructed using the base class synthesis method, which achieves the expression of intermediate fatigue states from absence to presence. Finally, fatigue levels are quantified based on the matching between samples and the complete base class set. Moreover, to cope with the complex variability of fatigue states, MDFG employs meta-learning for training. MDFG achieves an Average accuracy improvement of 10.0% and 12.1% on two real datasets compared to methods that do not consider fine-grained information. Extensive experiments demonstrate that the MDFG exhibits superior robustness and stability among current fatigue detection methods.
Xiaojie Zhu, Ruimin Hu, Dongliang Zhu 0001, Mang Ye
AAAI2
2025 PSDMQ: A Parallel Method for Shortest Distances Multi-Querying on Encrypted Graph
abstract
Searchable symmetric encryption (SSE) enables a client to outsource a set of encrypted data in the cloud and retain the ability to perform retrieval without revealing client's private information. Graph as an import data structure and shortest distance querying as a significant fundamental operations on graph, has been applied a lot in real-life. Although efficient SSE shortest distance querying on graph are known, previous solutions are highly sequential. This is mainly due to the fact that, currently, the method for querying shortest distance on encrypted is usually designed based on 2HCL, which requires the search algorithm to access a sequence of memory locations, each of which is unpredictable and stored at the previous location in the sequence. Motivated by advances in multi-core architectures, we construct a new parallized method for batch shortest distance multi-querying(PSDMQ) on encrypted graph-structured data. Our approach is highly parallel thus can significantly improve efficiency. Both theoretical analysis and experiments are provided. Experiments on 3 real-world data sets demonstrates the effectiveness of our proposed scheme.
Xiaotong Dong, Xiaojie Zhu
CSCWD2
2025 Dynamic Ring Signature: Towards Provable Anonymity in Blockchain-Based E-Voting
Shan Jiang 0005, Zhehao Huang, Shichang Xuan, Jiaxing Shen, Huakun Huang, Xiaojie Zhu
ICA3PP (8)6
2025 Exploring Backdoor Attacks in Federated Learning Under Parameter-Efficient Fine-Tuning
Xiaojie Zhu
ISC2
2025 An Efficient White-box LLM Watermarking for IP Protection on Online Market Platforms
abstract
Online market platforms serve as a central hub for sharing and deploying AI models among researchers, developers, and companies. In this context, watermarking techniques are essential to protect intellectual property (IP), preventing unauthorized use and duplication of large language models (LLMs). Two key challenges arise: (i) These platforms host diverse LLMs, yet current watermarking techniques are only tailored to specific models, such as fine-tuned or quantized LLMs. (ii) Efficient watermarking is critical. However, traditional methods require substantial data and costly hardware, which limits their feasibility. In this paper, we propose an efficient white-box LLM watermarking technique called ELLMark. This method treats LLMs as multi-layered matrices while embedding watermarks only relies on modifying the model's weights. To preserve LLMs' performance, it filters weights by correlations with the activation magnitudes and downstream tasks, then modifies weights as minimal as possible via histogram modulation. Notably, all phases are training-free with low hardware resources, making it efficient for online platforms. We conduct extensive experiments to evaluate the effectiveness of ELLMark on LLaMA-3, OPT, and Phi-3 LLMs. The results demonstrate that it achieves 100% success in watermark detection while preserving model performance. Moreover, the preprocessing, encoding, and decoding processes remain efficient, taking less than 7 minutes, 12 minutes, and 18 seconds, respectively, for models with 80B parameters. Lastly, it exhibits robustness against parameter overwriting, re-watermarking, forging, fine-tuning, and pruning attacks.
Shuguang Yuan 0003, Xingyu Su, Peizhuo Lv, Weiji Xue, Jing Yu 0007, Xiaojie Zhu, Chi Chen 0001
KDD (2)6
2025 BinEnhance: An Enhancement Framework Based on External Environment Semantics for Binary Code Search
Yongpan Wang, Hong Li 0004, Xiaojie Zhu, Siyuan Li 0014, Chaopeng Dong, Shouguo Yang, Kangyuan Qin
NDSS3
2025 Re-examine Federated Rank Learning: Analyzing Its Robustness Against Poisoning Attacks
abstract
Federated learning decentralizes the training process across clients, allowing clients that do not trust each other to collaboratively train machine learning models without sharing private local data. However, this decentralized approach makes federated learning vulnerable to Poisoning Attacks. Recently, some studies have proposed a new paradigm called Federated Rank Learning (FRL), which uses ranking updates instead of parameter or gradient updates in federated learning. This change transforms the space of updates from continuous to discrete, making the federated learning framework more robust to poisoning attacks.However, we found that some simple and direct poisoning attack methods can easily cause FRL to fail to converge, contrary to the intuition that reducing the update space should limit attackers. This led us to re-examine the security of FRL. We first analyzed the robustness guarantees of FRL and confirmed its vulnerability to poisoning attacks from both theoretical and experimental perspectives. Next, we revisited the impact of changing the update space from continuous to discrete in the framework and found that the advantage of this change does not lie in directly defending against poisoning attacks, but in greatly limiting the attacker’s ability to implement stealthy poisoning attacks. Based on this, we added appropriate defense strategies to FRL, further shrinking the discrete update space into a secure range, and limiting the effectiveness and stealthiness of attacks. Experiments show that this approach significantly improves the ability of FRL to resist poisoning attacks.
Xiaojie Zhu, Paulo Veríssimo
RAID2
2025 TSDAs: Enhanced Statistics-Based Query Recovery in Dynamic Searchable Symmetric Encryption via Probabilistic Modeling
abstract
Dynamic Searchable Symmetric Encryption (DSSE) schemes typically permit certain predefined leakages at runtime to ensure efficiency, such as the response/update volume and the response file volume. Query Recovery Attacks (QRAs) exploit such leakages, combined with auxiliary knowledge, to recover the user’s query contents. Powerful QRAs typically rely on the ground-truth knowledge of the dataset or queries. While QRA based on the weaker statistical knowledge, termed statistics-based QRA, is also possible, it yields lower accuracy and consequently is not considered a significant threat. In this work, we revisit statistics-based QRAs within the DSSE context and introduce a unified probabilistic optimization attack model. Building on this model, toward volumetric leakages, we develop Two-Stage Decode Attacks (TSDAs), which formulate the QRA as a Conditional Random Field (CRF) decoding problem and solve it through a two-stage optimization process. TSDAs not only generalize the state-of-the-art counterpart (i.e., LVIA, Xu et al., CCS 2023), but also significantly outperform it, even under more realistic and constrained leakage assumptions. Against traditional DSSE schemes lacking protection on response volume leakages, our attack achieves nearly 100% query recovery accuracy on the Lucene dataset under the same evaluation settings as Xu et al. In contrast, LVIA achieves only 77% accuracy. Even when applied to ShieldDB (Vo et al., TKDE 2023), which incorporates response volume obfuscation, our method maintains an accuracy at 50%, whereas LVIA drops sharply to just 0.2%. Our findings suggest that statistics-based QRAs pose a serious threat to DSSE schemes.
Yikang Fan, Xiaojie Zhu
TrustCom2
2025 SafeMLLM: Extending Safety Alignment from Single-Modal LLMs to Multimodal LLMs
abstract
Multimodal large language models (MLLMs) are capable of processing diverse multimodal inputs to generate informative and contextually relevant responses. In this paper, alignment involves safeguarding the model against producing harmful or inappropriate content and ensuring its outputs are aligned with ethical guidelines. Existing alignment works have primarily focused on making the outputs of large language models(LLMs) harmless. However, the integration of multimodal information into MLLM inputs can lead to unintended or undesirable responses, thereby undermining the original alignment mechanisms of LLMs. When image, audio, or video information is contained in queries, MLLMs cannot guarantee the safe alignment of responses. In this study, we first collect VAdvBench to demonstrate that the additional multimodal inputs can easily break the safeguards of LLMs in MLLMs, highlighting the security vulnerabilities in these MLLMs. To address this issue, we propose SafeMLLM, a safety alignment framework designed for MLLMs. SafeMLLM aligns multimodal content by transforming multimodal inputs into a unified space, enabling MLLM to effectively block multimodal harmful information. In addition, since the current multimodal safety alignment works only consider images as additional modality inputs and lack benchmarks for other modalities like audio and video, we propose a pipeline to collect and synthesize a multimodal safety alignment benchmark, MLGuard, to evaluate the safety alignment performance on multimodal queries. The experiment results on MLGuard demonstrate that our SafeMLLM effectively rejects harmful instructions containing multimodal queries of image, audio, and video, ensuring robust safety alignment across diverse multimodal inputs. The code is available at https://github.com/jiangdi841/SafeMLLM/tree/main.
Xiaojie Zhu, Chi Chen 0001
TrustCom2
2025 A review of privacy-preserving biometric identification and authentication protocols
Peisong Shen, Xiaojie Zhu, Xue Tian, Chi Chen 0001
Comput. Secur.3
2025 Efficient shortest distance approximate query on large scale encrypted graph data
abstract
Abstract The problem of querying shortest distance on a graph has attracted significant research attention due to the widespread applicability of graphs and the ability of graph shortest path queries to address numerous application problems. Given the limited capabilities of clients and the ongoing advancements in cloud computing, people would like to outsource their graph data. Outsourcing data, however, poses the problem of privacy breaches. We should enable clients to encrypt their data before outsource it to cloud servers while retaining the capability of querying the data. The major challenge lies in designing a scheme computing the shortest distance on encrypted graph is how to strike a balance between security, efficiency and accuracy. Moreover, this challenge becomes even more pronounced as the scale of the graph increases. In this article, we propose an efficient scheme called Encrypted Shortest Distance Approximate Query (ESDAQ). We design a new algorithm k-level BFS and make use of cryptographic primitive AES to fulfill the scheme where k is an optional parameter selected by user. The total time cost can be O(N) at best to finish setup and query, which is superior to SOTA solutions of O(NlogN). Theoretical security analysis shows that ESDAQ can reach CQA2-security . Theoretical analysis on security, performance and accuracy are provided of our proposed scheme. Meanwhile, experiments on 12 representative real-world datasets and comprehensive comparison with top and latest schemes are provided. The experiments results demonstrates that our scheme is highly efficient and can be effectively applied to large-scale graphs which comprising over 3 million nodes and 1 billion edges.
Xiaotong Dong, Xiaojie Zhu
Cybersecur.3
2025 A Verifiable and Efficient Symmetric Searchable Encryption Scheme for Dynamic Dataset With Forward and Backward Privacy
abstract
The adoption of symmetric searchable encryption (SSE) has become increasingly common. However, many current SSE schemes assume an honest-but-curious cloud service provider (CSP) or necessitate significant overhead to manage a malicious CSP. Furthermore, most of these schemes are tailored for static datasets. Our paper presents an efficient SSE scheme that aims to address these challenges. To the best of our knowledge, this is the first scheme that supports dynamic datasets with forward and backward privacy, integrity verification of non-empty and empty search results, efficient search, non-interactive, light client, and both forward and inverted indexes simultaneously. In this paper, we present two novel approaches, Hexie and Jianding. Hexie implements secret sharing to conceal index entries, enabling dynamic updates, non-interactive interactions, and lightweight clients. To enhance the reliability of search results and address the problem of empty, incomplete, or inaccurate outcomes, we introduce the Jianding scheme as an extension of Hexie. It combines a chained MAC structure with a secret sharing scheme, which enables a client to verify the data integrity of the search result efficiently. Moreover, we propose graph-based dictionary sharding to enhance search efficiency. Finally, we conduct comprehensive experiments to validate the effectiveness of the proposed schemes.
Xiaojie Zhu, Jiancong Zhou, Yueyue Dai, Peisong Shen, Shabnam Kasra Kermanshahi, Jiankun Hu
IEEE Trans. Dependable Secur. Comput.1
2024 PagPassGPT: Pattern Guided Password Guessing via Generative Pretrained Transformer
abstract
Amidst the surge in deep learning-based password guessing models, challenges of generating high-quality passwords and reducing duplicate passwords persist. To address these challenges, we present PagPassGPT, a password guessing model constructed on a Generative Pretrained Transformer (GPT). It can perform pattern guided guessing by incorporating pattern structure information as background knowledge, resulting in a significant increase in the hit rate. Furthermore, we propose D&C-GEN to reduce the repeat rate of generated passwords, which adopts the concept of a divide-and-conquer approach. The primary task of guessing passwords is recursively divided into non-overlapping subtasks. Each subtask inherits the knowledge from the parent task and predicts succeeding tokens. In comparison to the state-of-the-art model, our proposed scheme exhibits the capability to correctly guess 12% more passwords while producing 25% fewer duplicates.
Xingyu Su, Xiaojie Zhu, Paulo Veríssimo
DSN2
2024 Goldfish: An Efficient Federated Unlearning Framework
abstract
With recent legislation on the right to be forgotten, machine unlearning has emerged as a crucial research area. It facilitates the removal of a user's data from federated trained machine learning models without the necessity for retraining from scratch. However, current machine unlearning algorithms are confronted with challenges of efficiency and validity. To address the above issues, we propose a new framework, named Goldfish. It comprises four modules: basic model, loss function, optimization, and extension. To address the challenge of low validity in existing machine unlearning algorithms, we propose a novel loss function. It takes into account the loss arising from the discrepancy between predictions and actual labels in the remaining dataset. Simultaneously, it takes into consideration the bias of predicted results on the removed dataset. Moreover, it accounts for the confidence level of predicted results. Additionally, to enhance efficiency, we adopt knowledge a distillation technique in the basic model and introduce an optimization module that encompasses the early termination mechanism guided by empirical risk and the data partition mechanism. Furthermore, to bolster the robustness of the aggregated model, we propose an extension module that incorporates a mechanism using adaptive distillation temperature to address the heterogeneity of user local data and a mechanism using adaptive weight to handle the variety in the quality of uploaded models. Finally, we conduct comprehensive experiments to illustrate the effectiveness of proposed approach.
Houzhe Wang, Xiaojie Zhu, Paulo Veríssimo
DSN2
2024 Heterogeneous Spatio-Temporal Series Forecasting Using Dynamic Graph Neural Networks for Flood Prediction
abstract
Accurate flood prediction is essential for disaster mitigation, life protection, and minimizing community and infrastructure impact. However, current flood prediction models often struggle with challenges such as capturing intricate spatiotemporal dynamics, adapting to changing environmental conditions, and integrating diverse variables effectively. In this paper, we propose a novel approach for flood forecasting, say a Heterogeneous Dynamic Temporal Graph Convolutional Network (HD-TGCN). The Dynamic Temporal Graph Convolution Module (D-TGCM) adapts to evolving temporal graph structures, utilizing a Multi-Head Self-Attention mechanism to generate adjacency matrices dynamically. This enhances the adaptability of the model to changing temporal graph structures and captures intricate dependencies among features at different time steps. Additionally, HD-TGCN incorporates parallel DTGCMs to address the heterogeneity in flood spatiotemporal data. This module effectively captures intricate relationships among diverse variable types, improving the model's ability to handle complex multivariable interactions and enabling the model to process heterogeneous graphs effectively. Experiments on real-world datasets demonstrate that our model outperforms state-of-the-art models. At a prediction horizon of 30 minutes, the model exhibits improvements of 78.02%, 66.30%, and 0.15% in terms of MAE, RMSE, and NSE, respectively.
Jiange Jiang, Hailong Hou, Congjian Deng, Xiaojie Zhu
ICC7
2024 GPSSE: A GPU-Accelerated Dynamic SSE Scheme with Efficient Batch Updating
abstract
Dynamic Searchable Symmetric Encryption (DSSE) allows cloud users to securely retrieve and update their data while outsourcing it to untrusted cloud service providers. Although extensive research efforts in recent years have notably improved the retrieval efficiency of DSSE, there remains potential for enhancing update efficiency, particularly in large-scale dataset updating scenarios. Therefore, we proposed a pioneering scheme called GPSSE (GPU-Accelerated Dynamic Searchable Symmetric Encryption Scheme) to explore accelerating batch updates of DSSE through GPU. In our design, we break the traditional chain-based data structure and build independent label-based data blocks to store entries. It facilitates the parallel updating of entries. Moreover, GPSSE achieves forward privacy and Type II backward privacy. In addition, we formally prove the security of the proposed GPSSE and show its practicality by conducting experiments using the publicly well-known Enron Email dataset. Experimental results demonstrate that GPSSE outperforms 40× to 194156× than other state-of-the-art schemes in updating efficiency.
Jiancong Zhou, Xiaojie Zhu, Shuguang Yuan 0003, Chi Chen 0001
ISPA2
2024 UAV-RIS-Aided Energy-Efficient and QoS-Aware Emergency Communications Based on DRL
abstract
Ensuring reliable communication can be incredibly challenging in emergencies due to the breakdown of conventional infrastructure. However, a promising solution is on the horizon: the integration of reconfigurable intelligent surfaces (RIS) onto unmanned aerial vehicles (UAV), known as UAV-RIS. This innovative approach holds the potential to offer agile and adaptable communication services during crises, overcoming the limitations of traditional systems. This paper establishes an innovative UAV-RIS system with an active RIS to enhance the uplink communication between ground devices (GDs) and the air base station (ABS). We present an advanced communication strategy utilizing deep reinforcement learning (DRL) for UAV-RIS-supported uplink communication in dynamic emergencies. This scheme is designed to optimize the energy efficiency of the UAV-RIS communication system while adhering to quality of service (QoS) constraints for all GDs. It achieves this by jointly optimizing the trajectory of the UAV-RIS and the phase of the active RIS, ensuring efficient and reliable communication in challenging environments. To optimize the performance of the system, we propose a hierarchical Proximal Policy Optimization (H-PPO) algorithm and the upper and lower layers of H-PPO optimize the trajectory and phase control, respectively. Simulation results demonstrate that our scheme can effectively learn the communication strategy to enhance the performance of dynamic emergency communication networks.
Ying Ju 0001, Haoyu Wang 0015, Lei Liu 0031, Qingqi Pei, Yu Gang Shee, Xiaojie Zhu, Celimuge Wu
VTC Fall7
2024 Learning with noisy labels for robust fatigue detection
Ruimin Hu, Xiaojie Zhu, Dongliang Zhu 0001, Xiaochen Wang 0001
Knowl. Based Syst.3
2024 Privacy-Preserving and Trusted Keyword Search for Multi-Tenancy Cloud
abstract
Cloud service models intrinsically cater to multiple tenants. In current multi-tenancy model, cloud service providers isolate data within a single tenant boundary with no or minimum cross-tenant interaction. With the booming of cloud applications, allowing a user to search across tenants is crucial to utilize stored data more effectively. However, conducting such a search operation is inherently risky, primarily due to privacy concerns. Moreover, existing schemes typically focus on a single tenant and are not well suited to extend support to a multi-tenancy cloud, where each tenant operates independently. In this article, to address the above issue, we provide a privacy-preserving, verifiable, accountable, and parallelizable solution for “privacy-preserving keyword search problem" among multiple independent data owners. We consider a scenario in which each tenant is a data owner and a user’s goal is to efficiently search for granted documents that contain the target keyword among all the data owners. We first propose a verifiable yet accountable keyword searchable encryption (VAKSE) scheme through symmetric bilinear mapping. For verifiability, a message authentication code (MAC) is computed for each associated piece of data. To maintain a consistent size of MAC, the computed MACs undergo an exclusive OR operation. For accountability, we propose a keyword-based accountable token mechanism where the client’s identity is seamlessly embedded without compromising privacy. Furthermore, we introduce the parallel VAKSE scheme, in which the inverted index is partitioned into small segments and all of them can be processed synchronously. We also conduct formal security analysis and comprehensive experiments to demonstrate the data privacy preservation and efficiency of the proposed schemes, respectively.
Xiaojie Zhu, Peisong Shen, Yueyue Dai, Lei Xu 0019, Jiankun Hu
IEEE Trans. Inf. Forensics Secur.1
2023 Cancelable biometric schemes for Euclidean metric and Cosine metric
abstract
Abstract The handy biometric data is a double-edged sword, paving the way of the prosperity of biometric authentication systems but bringing the personal privacy concern. To alleviate the concern, various biometric template protection schemes are proposed to protect the biometric template from information leakage. The preponderance of existing proposals is based on Hamming metric, which ignores the fact that predominantly deployed biometric recognition systems (e.g. face, voice, gait) generate real-valued templates, more applicable to Euclidean metric and Cosine metric. Moreover, since the emergence of similarity-based attacks, those schemes are not secure under a stolen-token setting. In this paper, we propose a succinct biometric template protection scheme to address such a challenge. The proposed scheme is designed for Euclidean metric and Cosine metric instead of Hamming distance. Mainly, the succinct biometric template protection scheme consists of distance-preserving, one-way, and obfuscation modules. To be specific, we adopt location sensitive hash function to realize the distance-preserving and one-way properties simultaneously and use the modulo operation to implement many-to-one mapping. We also thoroughly analyze the proposed scheme in three aspects: irreversibility, unlinkability and revocability. Moreover, comprehensive experiments are conducted on publicly known face databases. All the results show the effectiveness of the proposed scheme.
Yubing Jiang, Peisong Shen, Xiaojie Zhu, Chi Chen 0001
Cybersecur.4
2023 Swarm Learning-based Secure and Fair Model Sharing for Metaverse Healthcare
Yueyue Dai, Xiaojie Zhu
Mob. Networks Appl.4
2023 A Privacy-Preserving Framework for Conducting Genome-Wide Association Studies Over Outsourced Patient Data
abstract
Due to the sheer volume of data, data owners (e.g., hospitals or other data collectors) tend to outsource their data to cloud service providers (CSPs) for the purpose of storage and analytics. However, privacy concerns about genomic and phenotype data significantly limit the data owners’ choice. In this work, we propose the first solution, to the best of our knowledge, that allows a CSP to perform efficient and privacy-preserving search and analysis over encrypted genomic and phenotype data that is multi-tenant, i.e. owned by multiple hospitals. We first propose an encryption mechanism for phenotype data, where each data owner is allowed to encrypt its data with a unique secret key. Moreover, the ciphertext supports privacy-preserving search and, consequently, enables the identification of the case and control groups for a genome-wide association study (GWAS) without any privacy violations. Furthermore, we provide a per-query based authorization mechanism for a client to access and operate on the data stored at the CSP. Additionally, we apply multi-key fully homomorphic encryption to encrypt genomic data and show how to compute GWAS statistics (e.g., chi-square distribution test) over the ciphertext of individuals in the identified case and control groups. Thus, for the first time, the proposed scheme provides privacy-preserving computation for the entire GWAS pipeline. Finally, we implement the proposed scheme and run experiments over a real-life genomic dataset to show its effectiveness. The result shows that the proposed solution is capable to efficiently identify the case/control groups and subsequently conduct GWAS on the identified case/control groups.
Xiaojie Zhu, Erman Ayday, Roman Vitenberg
IEEE Trans. Dependable Secur. Comput.1
2022 FAV-BFT: An Efficient File Authenticity Verification Protocol for Blockchain-Based File-Sharing System
Shuai Su, Xiaojie Zhu
CollaborateCom (1)3
2022 Privacy-Preserving Search for a Similar Genomic Makeup in the Cloud
abstract
Increasing affordability of genome sequencing and, as a consequence, widespread availability of genomic data opens up new opportunities for the field of medicine, as also evident from the emergence of popular cloud-based offerings in this area, such as Google Genomics [1]. To utilize this data more efficiently, it is crucial that different entities share their data with each other. However, such data sharing is risky mainly due to privacy concerns. In this article, we attempt to provide a privacy-preserving and efficient solution for the “similar patient search” problem among several parties (e.g., hospitals) by addressing the shortcomings of previous attempts. We consider a scenario in which each hospital has its own genomic dataset and the goal of a physician (or researcher) is to search for a patient similar to a given one (based on a genomic makeup) among all the hospitals in the system. To enable this search, we propose a hierarchical index structure to index each hospital’s dataset with low memory requirement. Furthermore, we develop a novel privacy-preserving index merging mechanism that generates a common search index from individual indices of each hospital to significantly improve the search efficiency. We also consider the storage of medical information associated with genomic data of a patient (e.g., diagnosis and treatment). We allow access to this information via a fine-grained access control policy that we develop through the combination of standard symmetric encryption and ciphertext policy attribute-based encryption. Using this mechanism, a physician can search for similar patients and obtain medical information about the matching records if the access policy holds. We conduct experiments on large-scale genomic data and show the high efficiency of the proposed scheme.
Xiaojie Zhu, Erman Ayday, Roman Vitenberg, Narasimha Raghavan
IEEE Trans. Dependable Secur. Comput.1
2021 A Privacy-Preserving Framework for Outsourcing Location-Based Services to the Cloud
abstract
Thanks to the popularity of mobile devices numerous location-based services (LBS) have emerged. While several privacy-preserving solutions for LBS have been proposed, most of these solutions do not consider the fact that LBS are typically cloud-based nowadays. Outsourcing data and computation to the cloud raises a number of significant challenges related to data confidentiality, user identity and query privacy, fine-grained access control, and query expressiveness. In this work, we propose a privacy-preserving framework for outsourcing LBS to the cloud. The framework supports multi-location queries with fine-grained access control, and search by location attributes, while providing semantic security. In particular, the framework implements a new model that allows the user to govern the trade-off between precision and privacy on a dynamic per-query basis. We also provide a security analysis to show that the proposed scheme preserves privacy in the presence of different threats. We also show the viability of our proposed solution and scalability with the number of locations through an experimental evaluation, using a real-life OpenStreetMap dataset.
Xiaojie Zhu, Erman Ayday, Roman Vitenberg
IEEE Trans. Dependable Secur. Comput.1
2017 Privacy-Preserving Relevance Ranking Scheme and Its Application in Multi-keyword Searchable Encryption
Peisong Shen, Chi Chen 0001, Xiaojie Zhu
SecureComm3
2016 Multidomain Subspace Classification for Hyperspectral Images
abstract
Hyperspectral imaging offers new opportunities for pattern recognition tasks in the remote sensing community through its improved discrimination in the spectral domain. However, such advanced image processing also brings new challenges due to the high data dimensionality in both the spatial and spectral domains. To relieve this issue, in this paper, we present a novel multidomain subspace (MDS) feature representation and classification method for hyperspectral images. The proposed method is based on a patch alignment framework. In order to optimally combine the feature representations from the various domains and simultaneously enhance the subspace discriminability, we incorporate the supervised label information into each domain and further generalize the framework to a multidomain version. Furthermore, we develop an iterative approach to alternately optimize the MDS objective function by considering it as two subconvex optimizations. The classification performance on three standard hyperspectral remote sensing images confirms the superiority of the proposed MDS algorithm over the state-of-the-art subspace learning methods.
Liangpei Zhang 0001, Xiaojie Zhu, Lefei Zhang, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 An Efficient Privacy-Preserving Ranked Keyword Search Method
abstract
Cloud data owners prefer to outsource documents in an encrypted form for the purpose of privacy preserving. Therefore it is essential to develop efficient and reliable ciphertext search techniques. One challenge is that the relationship between documents will be normally concealed in the process of encryption, which will lead to significant search accuracy performance degradation. Also the volume of data in data centers has experienced a dramatic growth. This will make it even more challenging to design ciphertext search schemes that can provide efficient and reliable online information retrieval on large volume of encrypted data. In this paper, a hierarchical clustering method is proposed to support more search semantics and also to meet the demand for fast ciphertext search within a big data environment. The proposed hierarchical approach clusters the documents based on the minimum relevance threshold, and then partitions the resulting clusters into sub-clusters until the constraint on the maximum size of cluster is reached. In the search phase, this approach can reach a linear computational complexity against an exponential size increase of document collection. In order to verify the authenticity of search results, a structure called minimum hash sub-tree is designed in this paper. Experiments have been conducted using the collection set built from the IEEE Xplore. The results show that with a sharp increase of documents in the dataset the search time of the proposed method increases linearly whereas the search time of the traditional method increases exponentially. Furthermore, the proposed method has an advantage over the traditional method in the rank privacy and relevance of retrieved documents.
Chi Chen 0001, Xiaojie Zhu, Peisong Shen, Jiankun Hu, Song Guo 0001, Zahir Tari, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.2
2015 SOLS: A scheme for outsourced location based service
Chi Chen 0001, Xiaojie Zhu, Peisong Shen, Jing Yu 0007, Hong Zou, Jiankun Hu
J. Netw. Comput. Appl.2