EDBT 2026 Demo / reviewers in the wild / expert
Lin Liu 0018
dblp:61/2115-18
· DBLP profile ↗
46ranked-venue papers
7as first author
42since 2021 · last 2026
0000-0002-5930-8881ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 23 · 3 first-author · 20 since 2021Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Computer networks · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Efficient Membership Inference Attacks Against Federated Large Language Models: A Projection Residual Approach
Guilin Deng, Silong Chen, Yuchuan Luo, Yi Liu 0057, Songlei Wang, Zhiping Cai, Lin Liu 0018, Xiaohua Jia, Shaojing Fu |
SP | 7 |
| 2026 | Euston: Efficient and User-Friendly Secure Transformer Inference with Non-Interactivity
Xinwen Gao, Shaojing Fu, Lin Liu 0018, Zhuotao Liu, Yuchuan Luo |
SP | 3 |
| 2026 | PriSEG: A Privacy-Preserving Scheme for Image Segmentation SystemabstractImage segmentation, a crucial technology in computer vision, is being applied in an increasingly wide range of fields. However, as data processing involves privacy, corresponding security issues have also emerged. The straightforward application of conventional security protocols in extensive image segmentation networks tends to yield suboptimal results, primarily due to the accumulation of errors. To protect sensitive or personal data, we introduce PriSEG — the first practical privacy-preserving image segmentation scheme based on Secure Multi-Party Computation (SMPC) technology. This study employs additive secret sharing and develops high-precision secure computation protocols for critical operations such as division and nonlinear activation functions (e.g., Sigmoid and ReLU) within neural networks. Our novel protocols facilitate secure inference in large and complex image segmentation networks. Additionally, our system design incorporates a tripartite service provider model, enhancing the system's robustness by ensuring it remains operational despite the dropout of any service provider. We have also validated the security of PriSEG under a semi-honest model. The experimental results on real-world datasets demonstrate that our scheme achieves performance comparable to plaintext implementations, with only a 0.005 degradation on average Mean Absolute Error (aveMAE). This proves the effectiveness and practicality of our scheme. Yuchuan Luo, Lin Liu 0018, Shaojing Fu |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | FLUTE: FSS-Based Secure Two-Party LLM Inference Using Partial Transformer EncryptionabstractRecently, transformer-based large language models (LLMs) have become mainstream, particularly when used as agents. However, users with exceedingly sensitive data cannot benefit due to a lack of high-performance devices to train LLMs or limited access to large companies' LLM APIs. Secure two-party computing (2PC), especially the recently popular function secret sharing (FSS), enables secure inference of LLMs by protecting both users' inputs and LLM parameters from leakage. In this paper, we propose${\sf FLUTE}$, the first FSS-based 2PC inference framework for LLMs with partial transformer encryption. We first identify a subset of transformer core blocks by simulating an adversary$\mathcal {A}$attempting to recover model parameters layer by layer, and then design GPU-friendly FSS protocols for each module within the core blocks, optimizing communication using the matrix multiplication protocol of${\sf ABY2.0}$. Security analysis shows that our partial encryption scheme provides security comparable to encrypting the entire LLM, thereby enhancing the performance-security trade-off of the entire end-to-end secure inference. The experimental results in the latest LLM, Llama 3.1-8B, show that${\sf FLUTE}$outperforms the SOTA (SIGMA) by$3\times$in both latency and communication, and it achieves even greater advantages in larger LLMs, such as Llama 3.1-70B. Yujie Xue, Lin Liu 0018, Yuchuan Luo, Shaojing Fu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | MM4flow: A Pre-trained Multi-modal Model for Versatile Network Traffic AnalysisabstractNetwork traffic analysis is a critical research area, playing an essential role in enhancing network security and ensuring high-quality network services. Existing methods, which primarily rely on a single modality, face two significant limitations. First, while existing approaches may achieve strong performance in specific tasks, they often lack sufficient adaptability for diverse tasks. Second, existing pre-trained models are only trained with GB-scale traffic, with which increases the risk of over-fitting and limiting the models' overall performance. To address these challenges, we propose MM4flow, a pre-trained multi-modal model designed for versatile network traffic analysis. We divide network flows into two modalities: raw byte streams and transmission patterns, which encapsulate the content and behavior information, respectively. MM4flow is composed of two key stages: uni-modal pre-training and multi-modal fine-tuning. We develop an efficient data collection scheme enabling TB-scale traffic pre-training. Leveraging a real-world traffic that exceeds 70 TB, MM4flow conducts uni-modal pre-training on each modality with a modified BERT architecture tailored for network flows. For specific downstream tasks, we introduce a modal fusion module based on cross-attention mechanisms. The fusion module facilitates effective integration of multi-modal information, enabling MM4flow to fully utilize both content and behavior cues during fine-tuning with minimal labeled dataset. We evaluate MM4flow on six public datasets covering six various tasks. Extensive experiments demonstrate that MM4flow achieves superior accuracy than baselines. Especially, compared to existing pre-trained models, MM4flow achieves an 84% improvement in accuracy for website identification under encrypted tunnels. Moreover, the pre-trained MM4flow significantly reduces the reliance on high-quality labeled training data for downstream tasks. Luming Yang, Lin Liu 0018, Junjie Huang 0001, Zhuotao Liu, Shiyu Liang, Shaojing Fu |
CCS | 2 |
| 2025 | MSFS: Maliciously Secure 3-Party Feature Selection via Mutual Information
Peijun Zhao, Lin Liu 0018, Shaojing Fu, Yuchuan Luo |
Inscrypt (2) | 2 |
| 2025 | A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model ConfidentialityabstractPrivacy-sensitive users require deploying large language models (LLMs) within their own infrastructure (on-premises) to safeguard private data and enable customization.However, vulnerabilities in local environments can lead to unauthorized access and potential model theft.To address this, prior research on small models has explored securing only the output layer within hardware-secured devices to balance model confidentiality and customization.Yet this approach fails to protect LLMs effectively.In this paper, we discover that (1) query-based distillation attacks targeting the secured top layer can produce a functionally equivalent replica of the victim model; (2) securing the same number of layers, bottom layers before a transition layer provide stronger protection against distillation attacks than top layers, with comparable effects on customization performance; and (3) the number of secured layers creates a trade-off between protection and customization flexibility.Based on these insights, we propose SOLID, a novel deployment framework that secures a few bottom layers in a secure environment and introduces an efficient metric to optimize the trade-off by determining the ideal number of hidden layers.Extensive experiments on five models (1.3B to 70B parameters) demonstrate that SOLID outperforms baselines, achieving a better balance between protection and downstream customization.Our code can be found at: https://github.com/ OTTO-OTO/SOLID-OnPremiseDeployment. Hanbo Huang, Lin Liu 0018, Zhuotao Liu, Ruoyu Sun 0001, Shiyu Liang |
EMNLP | 5 |
| 2025 | Concretely Efficient Three-party Oblivious SelectionabstractWith the increasing demand for access to and query of multimedia data, many applications rely on outsourcing data storage and queries to the cloud. Although efficient, this approach poses privacy risks to data owners. Researchers have sought secure methods for processing outsourced data, with the prerequisite of securely retrieving target data from datasets. In this paper, we address the performance limitations of existing three-party oblivious selection algorithms by proposing a novel online-efficient design. Furthermore, we optimize the proposed approach through algorithmic and system-level enhancements, achieving performance improvements across various settings. Shang Song, Lin Liu 0018, Rongmao Chen, Wei Peng 0005 |
ICME | 2 |
| 2025 | PrivMLLM: Efficient Three-Party Multimodal Large Language Model Secure Inference Supported Prompt PrivacyabstractMultimodal Large Language Models (MLLMs) represent the next frontier in artificial intelligence (AI), capable of processing and integrating diverse data types (text, images, audio and video) for richer understanding and generation. However, their potential is limited by privacy constraints: sensitive data resides with different owners who must keep it local, while MLLM parameters themselves require protection. We address this challenge by pioneering complete end-to-end privacy protection that simultaneously secures prompts, multimedia data, and MLLM parameters.In this work, we present PrivMLLM, the first general-purpose secure three-party (3PC) inference framework for MLLMs that serves two data owners while protecting both model parameters and inputs. Our approach is based on three key insights: (1) domain knowledge-aware optimization of fixed-point arithmetic for global performance gains, (2) hybrid secret-sharing protocols that reduce communication and rounds for linear/non-linear operations, and (3) constant-round function secret sharing (FSS)-based protocols for private embedding and maximum.We formally prove PrivMLLM’s security under the rigorous Universal Composability (UC) framework and demonstrate the first end-to-end secure inference system for MLLMs. Experimental results show that our solution enables secure inference for Llama3.2-11B-vision in under 0:5-minute per token, achieving 16 and 8 speedups compared to our implementations built with the SOTA 2PC (CrypTen+) and 3PC (ABY3+) privacypreserving machine learning (PPML) frameworks respectively. Moreover, we benchmark the submodules ViT and LLM against the SOTA schemes, SHAFT (NDSS 2025) and SIGMA (PETS 2024), achieving 1:3~3× performance gains. Yujie Xue, Lin Liu 0018, Yuchuan Luo, Shaojing Fu |
ICNP | 2 |
| 2025 | ENSI: Efficient Non-Interactive Secure Inference for Large Language ModelsabstractSecure inference enables privacy-preserving machine learning by leveraging cryptographic protocols that support computations on sensitive user data without exposing it. However, integrating cryptographic protocols with large language models (LLMs) presents significant challenges, as the inherent complexity of these protocols, together with LLMs' massive parameter scale and sophisticated architectures, severely limits practical usability. In this work, we propose ENSI, a novel non-interactive secure inference framework for LLMs, based on the principle of codesigning the cryptographic protocols and LLM architecture. ENSI employs an optimized encoding strategy that seamlessly integrates CKKS scheme with a lightweight LLM variant, BitNet, significantly reducing the computational complexity of encrypted matrix multiplications. In response to the prohibitive computational demands of softmax under homomorphic encryption (HE), we pioneer the integration of the sigmoid attention mechanism with HE as a seamless, retraining-free alternative. Furthermore, by embedding the Bootstrapping operation within the RMSNorm process, we efficiently refresh ciphertexts while markedly decreasing the frequency of costly bootstrapping invocations. Experimental evaluations demonstrate that ENSI achieves approximately an$8 \times$acceleration in matrix multiplications and a$2.6 \times$speedup in softmax inference on CPU compared to state-of-the-art method, with the proportion of bootstrapping is reduced to just 1 %. Maojiang Wang, Xinwen Gao, Yuchuan Luo, Lin Liu 0018, Shaojing Fu |
SRDS | 5 |
| 2025 | LEA: Label Enumeration Attack in Vertical Federated LearningabstractA typical Vertical Federated Learning (VFL) scenario involves several participants collaboratively training a machine learning model, where each party has different features for the same samples, with labels held exclusively by one party. Since labels contain sensitive information, VFL must ensure the privacy of labels. However, existing VFL-targeted label inference attacks are either limited to specific scenarios or require auxiliary data, rendering them impractical in real-world applications. We introduce a novel Label Enumeration Attack (LEA) that, for the first time, achieves applicability across multiple VFL scenarios and eschews the need for auxiliary data. Our intuition is that an adversary, employing clustering to enumerate mappings between samples and labels, ascertains the accurate label mappings by evaluating the similarity between the benign model and the simulated models trained under each mapping. To achieve that, the first challenge is how to measure model similarity, as models trained on the same data can have different weights. Drawing from our findings, we propose an efficient approach for assessing congruence based on the cosine similarity of the first-round loss gradients, which offers superior efficiency and precision compared to the comparison of parameter similarities. However, the computational cost may be prohibitive due to the necessity of training and comparing the vast number of simulated models generated through enumeration. To overcome this challenge, we propose Binary-LEA from the perspective of reducing the number of models and eliminating futile training, which lowers the number of enumerations from$\mathcal{O}(n!)$to$\mathcal{O}\left(n^{3}\right)$. Experiments across different VFL settings confirm LEA's effectiveness. In the absence of auxiliary datasets, LEA demonstrates a$\mathbf{5 0 \%}$to$\mathbf{9 0 \%}$enhancement in attack accuracy over the existing state-of-the-art label inference attacks in VFL. Moreover, LEA is resilient against common defense mechanisms such as gradient noise and gradient compression. Shaojing Fu, Yuchuan Luo, Lin Liu 0018 |
SRDS | 4 |
| 2025 | CertTA: Certified Robustness Made Practical for Learning-Based Traffic Analysis
Jinzhu Yan, Zhuotao Liu, Shiyu Liang, Lin Liu 0018, Ke Xu 0002 |
USENIX Security Symposium | 5 |
| 2025 | The Analysis of Encrypted Video Stream Based on Low-Dimensional Embedding MethodabstractIn recent years, encrypted video streaming takes up an increasing proportion of mobile network traffic, with encrypted video streams playing a significant role in illegal video detection. However, there are challenges in performing content analysis of encrypted video streams, including label limitations and complex calculations. In this paper, we proposed a low-dimensional embedding method based on Byte Rate Sequences (BRS), named EVS2vec (Encrypted Video Stream to Vector), to solve these problems effectively. It can represent the content of encrypted video streams with low-dimensional vectors by mapping the indefinite-length sequence into a low-dimensional Euclidean space. EVS2vec can thereby be applied for not only supervised analysis but also unsupervised analysis. Furthermore, using BRS can also save the time overhead on fine-grained network traffic parsing. In order to ensure the content-related distinguishability of the embedding result, inspired by contrastive learning, we designed a network structure based on Recurrent Neural Network (RNN) with self-attention mechanism in EVS2vec and trained it using triplet network. The experiments on a public dataset show that EVS2vec saves storage overhead while containing enough video content information. EVS2vec can achieve a high accuracy of similarity threshold, reaching 96.89%. An 8-dimensional fingerprint for each video is constructed. Moreover, classification and clustering analysis can also be performed with acceptable results. Luming Yang, Shaojing Fu, Lin Liu 0018, Yuchuan Luo |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Robustness Matters: Pre-Training Can Enhance the Performance of Encrypted Traffic AnalysisabstractModels with large-scale parameters and pre-training have been leveraged for encrypted traffic analysis. However, existing researches primarily focused on accuracy, often overlooking the role of large-scale pre-trained parameters in enhancing robustness. While machine learning (ML) and deep learning (DL) models trained from scratch can achieve high accuracy, they exhibit limited robustness. When subjected to network noise in real-world, their identification results can fluctuate significantly, which is unacceptable. Unfortunately, current robustness evaluation methods neglect samples diversity and employ unreasonable noise settings. This field still lacks a reasonable quantitative description of models robustness. In this paper, we propose the PA-curve to display the distribution of sample’s correct-decision stability, which can simultaneously reflect the model’s accuracy and robustness. By calculating the area under the PA-curve, called PA-area, we enable the quantitative assessment of robustness for encrypted traffic analysis. Furthermore, we design a pre-trained model based on packet length sequence, and pre-trained it on TB-scale traffic. By fine-tuning on limited labeled training data, it can achieve downstream analysis tasks. We conduct experiments on five encrypted traffic datasets with different tasks. Besides accuracy, we analyzed the robustness of the pre-trained model and existing methods under common network disturbances, including packet loss, retransmission, and disorder. Experimental results demonstrated that, compared to ML-based and DL-based models trained from scratch, the pre-trained model can not only achieve high accuracy, but also exhibit greater resilience to network noise. The source code is available at https://anonymous.4open.science/r/BERT-ps-4630. Luming Yang, Lin Liu 0018, Junjie Huang 0001, Jiangyong Shi, Shaojing Fu, Jinshu Su |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | unFlowS: An Unsupervised Construction Scheme of Flow Spectrum for Network Traffic DetectionabstractIn recent years, the construction of behavior-based analysis models is hindered by issues such as insufficient data, difficulty in labeling, and the complexity of behavior types. In reality, specific cyber threats often require manual analysis of raw network traffic, which is a complex and inefficient process. Flow spectrum can simplify the complex analysis process of raw network flow by mapping it from a high-dimensional space to a one-dimensional spectral space. However, the existing flow spectrum cannot adapt to the open-world scenarios and behavior-based detection for unknown cyber threats. To address these challenges, we propose a new flow spectrum construction scheme, named unFlowS, to effectively represent network flows and assist analysts to understand the behaviors of network traffic. unFlowS-Net, an unsupervised flow-based detection model we designed as the core of our scheme, can transform network flows into spectral lines. It makes unFlowS possible to detect unknown cyber threats. We further build spectral vectors for spectral lines generated by network flow sets, enabling the visualization of network behaviors within a period of time and automatic behavior-based detection. Experimental results demonstrated that unFlowS-Net can achieve better performance than state-of-the-art methods on unsupervised flow-based detection. Based on spectral vectors, not only can it intuitively display the network behavior characteristic of the target host, but also automatically detect suspicious network behaviors. Luming Yang, Lin Liu 0018, Junjie Huang 0001, Jiangyong Shi, Shaojing Fu, Shize Guo |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Diffusion-based Adversarial Attack to Automatic Speech Recognition
Yuchuan Luo, Shaojing Fu, Zhenyu Qiu, Lin Liu 0018 |
ACML | 5 |
| 2024 | Honest-Majority Maliciously Secure Skyline Queries on Outsourced DataabstractThe application of skyline queries on outsourced databases significantly aids online analysis, yet efficiently handling encrypted queries remains a formidable obstacle. Moreover, query outcomes are vulnerable to potential malicious cloud services. To circumvent these limitations, this work presents the Honest-Majority and Maliciously Skyline Query scheme (HMMSQ), which facilitates efficient skyline queries while safeguarding the privacy of datasets, queries, and skylines, as well as detecting malevolent activities. The core of HMMSQ is an optimized skyline diagram constructed by a novel skyline region-splitting algorithm for accurate skyline queries. Furthermore, it mitigates the frequency of dataset accesses by leveraging a multi-path R-tree for secure skyline retrieval. Notably, the majority of malicious behavior detection is focused on the servers, thereby minimizing user authentication overhead. The complexity and security are thoroughly analyzed, and experimental evaluations on various datasets demonstrate its efficiency and practicality in terms of computational cost and communication overhead. Remarkably, HMMSQ outperforms existing methods in query latency, achieving up to an order of magnitude improvement. Yu Chen 0113, Lin Liu 0018, Rongmao Chen, Shaojing Fu, Yuexiang Yang |
CIKM | 2 |
| 2024 | Privacy-Preserving Byzantine-Robust Federated Learning via Multiparty Homomorphic Encryption
Songwei Luo, Shaojing Fu, Lin Liu 0018, Yuchuan Luo |
COCOON (2) | 4 |
| 2024 | Speedy Privacy-Preserving Skyline Queries on Outsourced Data
Yu Chen 0113, Lin Liu 0018, Rongmao Chen, Shaojing Fu, Yuexiang Yang, Jiangyong Shi, Liangzhong He |
ESORICS (2) | 2 |
| 2024 | FedSV: A Privacy-Preserving Byzantine-Robust Federated Learning Scheme with Self-validation
Shaojing Fu, Yuchuan Luo, Lin Liu 0018 |
ICA3PP (1) | 4 |
| 2024 | PrivARM: Privacy-Preserving Association Rule Mining in the Cloud
Guangliang Sun, Wei Wu 0015, Lin Liu 0018 |
ICA3PP (1) | 5 |
| 2024 | KP-WF: Cross-Domain Few-Shot Website Fingerprinting
Lin Liu 0018, Ziling Wei, Shuhui Chen, Jinshu Su |
ICDF2C (2) | 1 |
| 2024 | Defend from Scratch: A Diffusion-Based Proactive Defense Method for Unauthorized Speech Synthesis
Yuchuan Luo, Zhenyu Qiu, Lin Liu 0018, Shaojing Fu |
ICONIP (1) | 4 |
| 2024 | FingerMamba: Mamba-based Efficient Multi-tab Website FingerprintingabstractNowadays, protecting user privacy on the Internet is paramount, especially with the increasing use of the Tor network to anonymize online activities. However, Tor is vulnerable to website fingerprinting (WF), where patterns in encrypted traffic are analyzed to infer visited websites. It can be utilized to monitor and investigate illegal activities on the dark web. Existing website fingerprinting techniques typically assume single-tab browsing, which is unrealistic as users often open multiple tabs consecutively or within a short period due to Tor’s slow loading speeds and typical user habits. Moreover, current multi-tab approaches face challenges in classification speed, which is crucial for high-throughput networks. FingerMamba, our proposed model, addresses these gaps by efficiently extracting local information and establishing long-range dependencies using a Mamba-based structured state-space model. It significantly enhances the accuracy and speed of multi-tab website fingerprinting. Extensive experiments on the largest real-world multi-tab dataset demonstrate that FingerMamba effectively improves classification accuracy in both closed-world and open-world settings. Furthermore, with maintaining similar accuracy performance, FingerMamba can increase inference speed by up to four times compared to the existing methods. To our knowledge, FingerMamba is the first model to tailor the Mamba architecture for website fingerprinting. Lin Liu 0018, Ziling Wei, Shuhui Chen, Zixuan Dong, Jinshu Su |
IPCCC | 1 |
| 2024 | AST-Trans: Detecting Web Tracking using Transformer-based Deep Learning with Abstract Syntax TreeabstractWeb tracking has become a key tool for service providers to collect online data and analyze user behaviors, raising concerns about the privacy of Internet users. In this paper, we propose a new web tracking detection method, namely AST-Trans, which detects and removes web tracking behavior using Transformer-based deep learning with abstract syntax trees. In the method, the abstract syntax tree is built for the detected website codes. Then, a sequence generation algorithm is proposed to convert tree-like code structures into one-dimensional sequences for deep models. To enhance training efficiency, we devise a reduction strategy to simplify the code tree structure by introducing equivalent nodes. After that, a Transformer-based deep learning algorithm is introduced to realize web tracking detection. By the proposed method, the exact tracking code blocks can be identified, and thus, we can implement the tracking code removal with minimum website breakage. To verify the effectiveness of the proposed method, an HTTPS proxy with AST-Trans on it is implemented to detect and remove tracking codes. We evaluate AST-Trans with the TrackSign-labeled dataset. The results show that the proposed method can detect the tracking behavior with high precision. In addition, we validate the feasibility of the method by measuring the website page breakage. Ziling Wei, Lin Liu 0018, Shuhui Chen, Jinshu Su |
IPCCC | 3 |
| 2024 | ExpMD: an Explainable Framework for Traffic Identification Based on Multi-Domain FeaturesabstractNetwork traffic identification has a significant impact on the QoS (Quality of Service) of network. However, there are challenges in performing network traffic identification, including underutilization of information and weak of interpretability. To address these issues, in this paper, we propose an Explainable framework based on Multi-Domain features for network traffic identification, named ExpMD. The network information is divided into four feature domains: metadata domain, byte-distribution (BD) domain, textual domain, and temporal domain. On these feature domains, we construct four identification models independently, thereby achieving the extraction and utilization of multi-domain information of network flows. Subsequently, efficient identification of network traffic can be achieved through the ensemble of models from four domains. In addition, we employ post-hoc explainable approaches to attribute features across multiple feature domains, thus mitigating the black-box issue of network traffic identification. We run a set of experiments in the identification tasks of application network traffic. The extensive evaluation demonstrates that our framework not only achieves an accuracy of 97. 5%, but also generates high-fidelity, sparse, complete, and stable explanation results for network flows. Furthermore, each sub-model performed at a state-of-the-art effect within their feature domains, respectively. Luming Yang, Lin Liu 0018, Shaojing Fu |
IWQoS | 3 |
| 2024 | BVDFed: Byzantine-resilient and verifiable aggregation for differentially private federated learning
Xinwen Gao, Shaojing Fu, Lin Liu 0018, Yuchuan Luo |
Frontiers Comput. Sci. | 3 |
| 2024 | SecDM: A Secure and Lossless Human Mobility Prediction SystemabstractWith the rapid development of deep neural network research, many deep neural network prediction services are provided for cloud service users. However, due to the untrusted nature of cloud computing, there are risks associated with this. To secure user data and cloud servers at the same time, we designed a secure prediction system, named SecDM (SecureDeepMove), which focuses on an attentional recurrent network for human mobility prediction. A neural network inference system that allows two parties to work together securely and efficiently without revealing any data is presented in this work. To design it, with the help of secret sharing technology, we first propose several secure, efficient, and lossless two-party protocols, which can securely calculate non-linear functions in the model, such as sigmoid, tanh, and softmax. In addition, a secure and effective strategy is introduced to maintain the accuracy of the calculation. Moreover, we also prove the security of our scheme in the semi-honest model. Finally, experimental results validate that the prediction result of our SecDM is not only as accurate as the non-privacy-preserving scheme but also highly efficient. Lin Liu 0018, Shaojing Fu, XueLun Huang, Yuchuan Luo, Xuyun Zhang, Kim-Kwang Raymond Choo |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Secure Approximate Nearest Neighbor Search with Locality-Sensitive Hashing
Shang Song, Lin Liu 0018, Rongmao Chen, Wei Peng 0005, Yi Wang 0055 |
ESORICS (3) | 2 |
| 2023 | Credit-based Differential Privacy Stochastic Model Aggregation Algorithm for Robust Federated Learning via BlockchainabstractBy encapsulating model parameters in blocks when training machine learning models collaboratively, blockchain is recognized as a promising enabling technology to facilitate reliable federated learning under a distributed and untrusted environment. However, storing updated models of each worker in blockchain induces potential privacy risks, such as membership inference attacks. Besides, the volatile network conditions in the distributed environment may cause the deterioration of system robustness. This paper aims at addressing the privacy and robustness issues mentioned above. Specifically, a Credit-based Differential Privacy stochastic model aggregation algorithm combined with SIGN operation (Cre-DPSIGN) is adopted in our peer-to-peer network, which can realize the trade-off between privacy and accuracy. Furthermore, leveraging the transparency and tamper-proofing of blockchain, we design practical and reliable smart contracts for unbiased sampling based on the credit of workers to improve system robustness against Byzantine workers. In addition, we have demonstrated that the use of biased differential privacy mechanisms can lead to performance degradation. Therefore, we have introduced two unbiased differential privacy mechanisms and have proven their convergence and privacy guarantee. Extensive experiments conducted on MNIST datasets show that our algorithm can achieve byzantine fault tolerance rate with a private loss ϵ = 0.4. Compared with the state-of-the-art, aka DP-RSA (IJCAI-22), Cre-DPSIGN shows lower privacy loss consumption and better system robustness. Mengyao Du, Miao Zhang 0037, Lin Liu 0018, Kai Xu 0014, Quanjun Yin |
ICPP | 3 |
| 2023 | Privacy-Preserving Federated Learning with Hierarchical Clustering to Improve Training on Non-IID Data
Songwei Luo, Shaojing Fu, Yuchuan Luo, Lin Liu 0018, Yanxiang Deng |
NSS | 4 |
| 2023 | Privacy Preserving Outsourced K-means Clustering Using Kd-tree
Yanxiang Deng, Lin Liu 0018, Shaojing Fu, Yuchuan Luo, Wei Wu 0015 |
ProvSec | 2 |
| 2023 | EVS2vec: A Low-dimensional Embedding Method for Encrypted Video Stream AnalysisabstractThe rise in video streaming has led to an increase in network traffic, with encrypted video streams playing a significant role in illegal video detection. However, there are challenges in performing content analysis of encrypted video streams, including label limitations and complex calculations. In this paper, we proposed a low-dimensional embedding method based on Byte Rate Sequences (BRS), named EVS2vec (Encrypted Video Stream to Vector), to solve these problems effectively. It can represent the content of encrypted video streams with low-dimensional vectors by mapping the indefinite-length sequence into a low-dimensional Euclidean space. EVS2vec can thereby be applied for not only supervised analysis but also unsupervised analysis. Furthermore, using BRS can also save the time overhead on fine-grained network traffic parsing. In order to ensure the content-related distinguishability of the embedding result, inspired by contrastive learning, we designed a network structure based on Recurrent Neural Network (RNN) with self-attention mechanism in EVS2vec and trained it using siamese network. The experiments on a public dataset show that EVS2vec saves storage overhead while containing enough video content information. EVS2vec can achieve a high accuracy of similarity threshold, reaching 96.71%. An 8-dimensional fingerprint for each video is constructed. Moreover, classification and clustering analysis can also be performed with acceptable results. Luming Yang, Shaojing Fu, Lin Liu 0018, Yuchuan Luo |
SECON | 4 |
| 2023 | EzBoost: Fast And Secure Vertical Federated Tree Boosting Framework via EzPCabstractFederated learning (FL) has emerged as a prominent methodology for collaboratively training machine learning models among multiple participants while alleviating data privacy leakage through data localization. However, recent studies have shown that the transferred intermediate parameters still contain sensitive information that needs to be further protected. More-over, real-world institutions often possess diverse data attributes, necessitating the adoption of Vertical Federated Learning (VFL) for cooperative learning tasks. Existing researches in VFL have proposed some frameworks with privacy-preservation functionality, yet they suffer from high participant overhead or low model accuracy, etc. To address these challenges, in this paper, we propose EzBoost, a fast and secure vertical federated tree boosting framework built upon XGBoost. Specifically, we leverages the efficient Secure Multi-party Computation (MPC) framework, EzPC, to facilitate the design and implementation of EzBoost. By carefully designing our framework with two non-collusive servers for secure two-party computation, EzBoost significantly accelerates the runtime of model training and querying at most 20×, and reduces the participant overheads at most 300×. In addition, we identify a potential privacy leakage problem in recent researches and propose a more robust solution for addressing it. Through comprehensive security analysis and comparative experiments with existing approaches, we demonstrate that EzBoost achieves stronger privacy-preservation, higher accuracy and higher efficiency simultaneously. Xinwen Gao, Shaojing Fu, Lin Liu 0018, Yuchuan Luo, Luming Yang |
TrustCom | 3 |
| 2023 | SecGAN: Honest-Majority Maliciously 3PC Framework for Privacy-Preserving Image SynthesisabstractThe Generative Adversarial Network (GAN) is capable of generating high-quality images, surpassing earlier generative models by producing fake images. To effectively handle the high computational workload and the large number of parameters involved, outsourcing to cloud servers is a more suitable option. However, utilizing cloud servers for image synthesis also presents the risk of privacy breaches. Additionally, existing privacy-preserving GANs operate on a semi-honest model with two parties, which fails to resist malicious attacks. This paper introduces a novel image synthesis framework that involves three parties, enabling privacy-preserving Deep Convolutional GAN(DCGAN) on the cloud in both semi-honest and malicious scenarios. The secure data reconstruction implemented in this framework detects adversary attacks under the malicious model and aborts computation upon detection. Furthermore, we offer a range of highly efficient and accurate secure computation protocols specifically designed for DCGAN-based image synthesis. Extensive experimental results demonstrate that our secure GAN(SecGAN) can produce image quality comparable to that of the plaintext method. Lin Liu 0018, Shaojing Fu, Yuchuan Luo |
TrustCom | 2 |
| 2023 | Randomization is all you need: A privacy-preserving federated learning framework for news recommendation
Xinyi Huang 0001, Yuchuan Luo, Lin Liu 0018, Shaojing Fu |
Inf. Sci. | 3 |
| 2023 | BFCRI: A Blockchain-Based Framework for Crowdsourcing With Reputation and IncentiveabstractWith the rapid development of cloud computing and the sharing economy, crowdsourcing aroused widespread interest and adoption in providing intelligent and efficient services for humans. The majority of existing works focus on effective crowdsourcing task assignment and privacy protection, mostly relying on central servers and assuming that participants are$honest$-$and$-$curious$and proactive. However, in reality, workers may be unwilling to participate, and there may be malicious behavior among participants, thus harming the enthusiasm and interests of other participants. The central server has weaknesses such as single point of failure. To address above problems, we propose a blockchain-based framework for crowdsourcing with reputation and incentive. We first design a worker selection scheme to select credible and capable workers. We leverage reputation as a metric of workers’ credibility, which is calculated through the improved subjective logic model. Then we utilize contract theory to design incentive mechanisms to attract more workers, especially high-quality workers to participate. Experimental results show that our proposed method can detect and prevent malicious participants and resist malicious collusion when the proportion of malicious participants is no more than 1/3. And encourage more workers to actively, honestly and continuously participate in crowdsourcing. Shaojing Fu, XueLun Huang, Lin Liu 0018, Yuchuan Luo |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | Towards Strong Privacy Protection for Association Rule Mining and Query in the CloudabstractEfficiently mining frequent itemsets and association rules on the encrypted outsourced data remains a great challenge for the time-consuming ciphertext computations. Nowadays, it has been not well addressed for privacy-preserving frequent itemsets and association rule mining schemes with mining efficiency, dataset, and query confidentiality simultaneously. In this paper, we investigate the study of privacy issues on frequent itemset mining and association rule mining on outsourced data in a two-cloud model, where the data are encrypted and outsourced by multiple owners holding different public keys. We develop several secure computation protocols based on additively homomorphic cryptosystem and additive secret sharing, which enable the clouds could securely mine the frequent itemsets and association rules. Furthermore, we also design two kinds of frequent itemset and association rule query service models, i.e., service customers query the cloud-mined results, and service customers query with their own decided threshold. The proposed scheme not only supports the mining process on the data encrypted by multiple public keys without compromising the security of the datasets, query data and query results, but also offline users. In addition, the experimental results show that our query scheme is much more efficient than the state-of-the-art work. Lin Liu 0018, Jinshu Su, Ximeng Liu, Rongmao Chen, Xinyi Huang 0001, Guang Kou, Shaojing Fu |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | pCOVID: A Privacy-Preserving COVID-19 Inference Framework
Yinqiu Wang, Yuchuan Luo, Lin Liu 0018, Shaojing Fu |
ICA3PP | 3 |
| 2022 | SecRec: A Privacy-Preserving Method for the Context-Aware Recommendation SystemabstractContext-aware recommendation systems are of increasing popularity in the digital era to recommend personalized items to users. However, how to ensure user data privacy while remaining high recommendation accuracy is widely considered a challenge. In this work, we propose a privacy-preserving method for the context-aware recommendation system in the two-cloud model. In particular, we first adjust the standard additive secret sharing scheme to support secure negative integers computation, based on which we manage to design secure comparison protocol and division protocols that enjoy desirable security and efficiency. By using these new protocols, we propose a secure and efficient context-aware recommendation system that also supports offline users. Compared with the state-of-the-art, our scheme achieves stronger data privacy preservation by further protecting the intermediate data calculated during the system training. Experimental results on real-world datasets indicate that our scheme is efficient. Notable, our system could achieve more significant performance improvement by running the underlying schemes in parallel. Jinrong Chen, Lin Liu 0018, Rongmao Chen, Wei Peng 0005, Xinyi Huang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Rectifying Pseudo Labels: Iterative Feature Clustering for Graph Representation LearningabstractGraph Convolutional Networks (GCNs) are powerful representation learning methods for non-Euclidean data. Compared with the Euclidean data, labeling the non-Euclidean data is more expensive. Meanwhile, most existing GCNs only utilize few labeled data but ignore most of the unlabeled data. To address this issue, we design a novel end-to-end Iterative Feature Clustering Graph Convolutional Networks (IFC-GCN) that enhances the standard GCN with an Iterative Feature Clustering (IFC) module. The proposed IFC module constrains node features iteratively based on the predicted pseudo labels and feature clustering. Further, we design an EM-like framework for IFC-GCN training, which improves the network performance by rectifying the pseudo labels and the node features alternately. Theoretical analysis and experimental results show that our proposed IFC module can effectively modify the node features. Experimental results on public datasets demonstrate that IFC-GCN outperforms state-of-the-art methods on the semi-supervised node classification task. Guang Kou, Lin Liu 0018 |
CIKM | 6 |
| 2021 | Privacy-Preserving Swarm Learning Based on Homomorphic Encryption
Lijie Chen 0007, Shaojing Fu, Lin Liu 0018, Yuchuan Luo |
ICA3PP (3) | 3 |
| 2020 | SHOSVD: Secure Outsourcing of High-Order Singular Value Decomposition
Jinrong Chen, Lin Liu 0018, Rongmao Chen, Wei Peng 0005 |
ACISP | 2 |
| 2020 | Towards Practical Privacy-Preserving Decision Tree Training and Evaluation in the CloudabstractDue to the capacity of storing massive data and providing huge computing resources, cloud computing has been a desirable platform for doing machine learning. However, the issue of data privacy is far from being well solved and thus has been a general concern in the cloud-aided machine learning. In this work, we investigate the study of how to efficiently do decision tree training and evaluation in the cloud and meanwhile achieve privacy preservation. Unlike existing cloud server-assisted model training approaches, in our proposed solution, the whole training process is mostly done by the cloud service provider who owns the machine learning model. Since the cloud cannot directly divide the encrypted dataset according to the best attributes selected, we propose a new method for decision tree training without dataset splitting. Precisely, we design three methods for decision tree training with the different tradeoff between privacy and efficiency. In all of these methods, the outsourced data are not revealed to the cloud service provider. We also propose a privacy-preserving decision tree evaluation scheme where the cloud service provider learns nothing about the user's input and the classification result while the trained model is kept secret to the user who could only learn the classification result. Compared with previous decision tree evaluation work, our scheme achieves desirable privacy preservation against both the user and the cloud service provider, and also minimizes the user's computation and communication costs. Moreover, besides protecting the data confidentiality, our proposed scheme also supports off-line users and thus has good scalability. The real-world dataset-based experimental results demonstrate that our system is of desirable utility and efficiency. Lin Liu 0018, Rongmao Chen, Ximeng Liu, Jinshu Su, Linbo Qiao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Toward Highly Secure Yet Efficient KNN Classification Scheme on Outsourced Cloud DataabstractNowadays, outsourcing data and machine learning tasks, e.g.,$k$-nearest neighbor (KNN) classification, to clouds has become a scalable and cost-effective way for large scale data storage, management, and processing. However, data security and privacy issue have been a serious concern in outsourcing data to clouds. In this article, we propose a privacy-preserving KNN classification scheme on cloud data in a twin-cloud model based on an additively homomorphic cryptosystem and secret sharing. Compared with existing works, we redesign a set of lightweight building blocks, such as secure square Euclidean distance, secure comparison, secure sorting, secure minimum, and maximum number finding, and secure frequency calculating, which achieve the same security level but with higher efficiency. In our scheme, data owners stay offline, which is different from secure-multiparty computation-based solutions which require data owners’ stay online during computation. In addition, query users do not interact with the cloud except sending query data and receiving the query results. Our security analysis shows that the scheme protects outsourced data security and query privacy, and hides access patterns. The experiments on real-world dataset indicate that our scheme is significantly more efficient than existing schemes. Lin Liu 0018, Jinshu Su, Ximeng Liu, Rongmao Chen, Robert H. Deng, Xiaofeng Wang 0002 |
IEEE Internet Things J. | 1 |
| 2018 | Privacy-Preserving Mining of Association Rule on Outsourced Cloud Data from Multiple Parties
Lin Liu 0018, Jinshu Su, Rongmao Chen, Ximeng Liu, Xiaofeng Wang 0002, Shuhui Chen, Ho-fung Leung |
ACISP | 1 |