EDBT 2026 Demo / reviewers in the wild / expert
Wenjing Fang
dblp:20/11504
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Kona: An Efficient Privacy-Preservation Framework for KNN Classification by Communication OptimizationabstractK-nearest neighbors (KNN) classification plays a significant role in various applications due to its interpretability. The accuracy of KNN classification relies heavily on large amounts of high-quality data, which are often distributed among different parties and contain sensitive information. Dozens of privacy-preserving frameworks have been proposed for performing KNN classification with data from different parties while preserving data privacy. However, existing privacy-preserving frameworks for KNN classification demonstrate communication inefficiency in the online phase due to two main issues: (1) They suffer from huge communication size for secure Euclidean square distance computations. (2) They require numerous communication rounds to select the $k$ nearest neighbors. In this paper, we present $\texttt{Kona}$, an efficient privacy-preserving framework for KNN classification. We resolve the above communication issues by (1) designing novel Euclidean triples, which eliminate the online communication for secure Euclidean square distance computations, (2) proposing a divide-and-conquer bubble protocol, which significantly reduces communication rounds for selecting the $k$ nearest neighbors. Experimental results on eight real-world datasets demonstrate that $\texttt{Kona}$ significantly outperforms the state-of-the-art framework by $1.1\times \sim 3121.2\times$ in communication size, $16.7\times \sim 5783.2\times$ in communication rounds, and $1.1\times \sim 232.6\times$ in runtime. Guopeng Lin, Ruisheng Zhou, Weili Han, Wenjing Fang |
ICML | 6 |
| 2025 | Is MPC Secure? Leveraging Neural Network Classifiers to Detect Data Leakage Vulnerabilities in MPC ImplementationsabstractDue to the emerging privacy-protection laws and regulations (e.g. GDPR in the EU) in recent years, dozens of multi-party computation (MPC for short) protocols have been proposed and widely applied by companies and institutions. These MPC protocols enable companies and institutions to perform joint analyses and machine learning on their private data while protecting their data's privacy. However, due to the complexity of MPC protocols, their implementations of-ten contain data leakage vulnerabilities, which can critically undermine the intended privacy protection. Additionally, most existing security analyses of MPC protocols rely on theoretical proofs, neglecting to detect possible vulnerabilities in MPC im-plementations. Therefore, detecting data leakage vulnerabilities in MPC implementations is an urgent necessity. In this paper, we propose MPCGuard, a practical frame-work for detecting data leakage vulnerabilities in MPC imple-mentations. Different from traditional memory vulnerabilities, data leakage vulnerabilities in MPC implementations cannot be identified by existing sanitizers. To resolve this challenge, we first establish a leakage identifier in MPCGuard with two neural network classifiers to identify whether an MPC implementation contains data leakage vulnerabilities. To enhance identification effectiveness, the structures of neural network classifiers are designed according to the characteristics of MPC protocols. After identifying a data leakage vulnerability, we employ a delta method to assist in locating the vulnerability. To demonstrate the effectiveness of MPCGuard, we apply MPCGuard to test 29 commonly-used MPC implementations in three main-stream MPC frameworks, i.e. Crypten, TF-Encrypted, and MP-SPDZ. We discover that 12 out of 29 implementations contain data leakage vulnerabilities, some of which can lead to the reconstruction of raw data. Until the moment this paper is written, all vulnerabilities, two of which have been assigned with CVE-IDs, have been confirmed. To the best of our knowledge, these two CVE-IDs are the first CVE-IDs assigned for data leakage vulnerabilities in MPC implementations. Guopeng Lin, Xiaoning Du 0001, Lushan Song, Weili Han, Junming Ma, Wenjing Fang |
SP | 7 |
| 2024 | FedMark: Large-Capacity and Robust Watermarking in Federated LearningabstractMachine learning models are increasingly recognized as valuable intellectual property (IP), prompting the development of a range of watermarking techniques aimed at safeguarding the IP of these models. However, in the context of federated learning (FL) models involving multiple owners, such as the participants in FL model training, conventional techniques designed for single-owner models prove ineffective due to limitations in their capacity and robustness. Few work has explored how to effectively embed watermarks to FL models for multiple-owners, which is non-trivial, especially when the number of owners is large. To fill this gap, we first analyze the capacity of existing watermarking methods. Second, we propose FedMark, a general large-capacity watermarking mechanism for FL, which leverages the Bloom Filter to achieve conflict-free watermarking of a large number of participants. Moreover, we propose a secret-sharing-based verification method to improve the watermarking robustness against false positives caused by Bloom Filter. Finally, comprehensive experiments show that our design can support over 150 participants to embed watermarks while the model accuracy varies within 1 %, and is robust to non-independent identical distributed data, different participant selection rates, model modifications, permutation attacks, scaling attacks and forging attacks. Lan Zhang 0002, Chen Tang 0002, Huiqi Liu, Haikuo Yu, Xirong Zhuang, Lei Wang 0005, Wenjing Fang, Xiang-Yang Li 0001 |
ICDCS | 8 |
| 2024 | Ditto: Quantization-aware Secure Inference of Transformers upon MPCabstractDue to the rising privacy concerns on sensitive client data and trained models like Transformers, secure multi-party computation (MPC) techniques are employed to enable secure inference despite attendant overhead. Existing works attempt to reduce the overhead using more MPC-friendly non-linear function approximations. However, the integration of quantization widely used in plaintext inference into the MPC domain remains unclear. To bridge this gap, we propose the framework named Ditto to enable more efficient quantization-aware secure Transformer inference. Concretely, we first incorporate an MPC-friendly quantization into Transformer inference and employ a quantization-aware distillation procedure to maintain the model utility. Then, we propose novel MPC primitives to support the type conversions that are essential in quantization and implement the quantization-aware MPC execution of secure quantized inference. This approach significantly decreases both computation and communication overhead, leading to improvements in overall efficiency. We conduct extensive experiments on Bert and GPT2 models to evaluate the performance of Ditto. The results demonstrate that Ditto is about $3.14\sim 4.40\times$ faster than MPCFormer (ICLR 2023) and $1.44\sim 2.35\times$ faster than the state-of-the-art work PUMA with negligible utility degradation. Haoqi Wu, Wenjing Fang, Yancheng Zheng, Junming Ma |
ICML | 2 |
| 2024 | A Microfluidic Impedance Cytometer for Accurate Detection and Counting of Circulating Tumor Cells by Simultaneous Mechanical and Electrical SensingabstractMicrofluidic Impedance Cytometry (MIC) is an advanced approach for single-cell analysis, harnessing microfluidic technology and impedance-based principles, particularly applicable in cancer diagnostics based on circulating tumor cells (CTCs). However, only relying on one dimensional information from electrical sensing poses challenges when detecting smaller CTCs that exhibit comparable sizes to white blood cells. To address this issue, we propose a microfluidic impedance cytometer featuring custom-designed circuits, electrodes, and a microfluidic chip with constriction channel. This system concurrently extracts both mechanical and electrical properties from processed electrical signals, overcoming the obstacle posed by the intrinsic link between impedance signals and cell sizes. Our system was tested with A549 lung cancer cells, white blood cells, and red blood cells, demonstrating the ability to differentiate cells of similar sizes within the blood sample and accurate cell couting capabilities. This approach shows promise for early cancer detection and monitoring treatment efficacy. Xiang Ke, Rikui Xiang, Wenjing Fang, Liangzun Fu, Xiwei Huang, Jinhong Guo, Lingling Sun |
ISCAS | 5 |
| 2024 | LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
Chengkun Wei, Wenlong Meng, Zhikun Zhang 0001, Min Chen 0032, Minghu Zhao, Wenjing Fang, Lei Wang 0152, Wenzhi Chen |
NDSS | 6 |
| 2024 | Towards Practical Oblivious MapabstractOblivious map (OMAP) is an important component in encrypted databases, utilized to prevent the server inferring sensitive information about client's encrypted databases based on access patterns. Despite its widespread usage and importance, existing OMAP solutions face practical challenges, including the need for a large number of interaction rounds between the client and server, as well as substantial communication bandwidth. For example, the SOTA protocol OMIX++ in VLDB 2024 still requires O (log n ) interaction rounds and O (log 2 n ) communication bandwidth per access, where n denotes the total number of key-value pairs stored. In this work, we introduce more practical and efficient OMAP constructions. Consistent with all prior OMAPs, our constructions also adapt only the tree-based Oblivious RAM (ORAM) and oblivious data structures (ODS) to achieve OMAP for enhanced practicality. In complexity, our approach needs O (log n /log log n )+ O (log λ ) interaction rounds and O (log 2 n /log log n ) + O (log λ log n ) communication bandwidth per data access where λ is the security parameter. This new complexity results from our two main contributions. First, unlike prior works relying solely on search trees , we design a novel framework for OMAP that combines hash table with search trees. Second, we propose a more efficient tree-based ORAM named DAORAM, which is of significant independent interest. This new ORAM accelerates our constructions as it supports obliviously accessing hash tables more efficiently. We implement both our proposed constructions and prior methods to experimentally demonstrate that our constructions substantially outperform prior methods in terms of efficiency. Xinle Cao, Weiqi Feng, Jian Liu 0012, Jinjin Zhou, Wenjing Fang, Lei Wang 0251, Quanqing Xu, Chuanhui Yang, Kui Ren 0001 |
Proc. VLDB Endow. | 5 |
| 2024 | SecretFlow-SCQL: A Secure Collaborative Query pLatformabstractIn the business scenarios at Ant Group, there is a rising demand for collaborative data analysis among multiple institutions, which can promote health insurance, financial services, risk control, and others. However, the increasing concern about privacy issues has led to data silos. Secure Multi-Party Computation (MPC) provides an effective solution for collaborative data analysis, which can utilize data value while ensuring data security. Nevertheless, the performance bottlenecks of MPC and the strong demand for scalability pose great challenges to secure collaborative data analysis frameworks. In this paper, we build a secure collaborative data analysis system SCQL with a general purpose. We design more efficient MPC protocols and relational operators to meet the demand for scalability. In terms of system design, we aim to implement a system with security, usability, and efficiency. We conduct extensive experiments on SCQL to validate our optimization improvements: (1) Our optimized secure sort protocol sorts one million 64-bit data in only 4.5 minutes, 126× faster than EMP (9.4 hours). (2) The end-to-end execution time of the typical vertical scenario query is reduced by 1991× from the state-of-the-art semi-honest collaborative analysis framework Secrecy (rewritten with Additive Secret Sharing protocol), with appropriate security tradeoffs. (3) We test the system in the WAN setting with input size = 10 7 to demonstrate the scalability. We have successfully deployed SCQL to address problems in real-world business scenarios at Ant Group. Wenjing Fang, Shunde Cao, Guojin Hua, Junming Ma, Yongqiang Yu, Qunshan Huang, Xiaopeng Zan, Pu Duan |
Proc. VLDB Endow. | 1 |
| 2023 | Video-Audio Domain Generalization via Confounder DisentanglementabstractExisting video-audio understanding models are trained and evaluated in an intra-domain setting, facing performance degeneration in real-world applications where multiple domains and distribution shifts naturally exist. The key to video-audio domain generalization (VADG) lies in alleviating spurious correlations over multi-modal features. To achieve this goal, we resort to causal theory and attribute such correlation to confounders affecting both video-audio features and labels. We propose a DeVADG framework that conducts uni-modal and cross-modal deconfounding through back-door adjustment. DeVADG performs cross-modal disentanglement and obtains fine-grained confounders at both class-level and domain-level using half-sibling regression and unpaired domain transformation, which essentially identifies domain-variant factors and class-shared factors that cause spurious correlations between features and false labels. To promote VADG research, we collect a VADG-Action dataset for video-audio action recognition with over 5,000 video clips across four domains (e.g., cartoon and game) and ten action classes (e.g., cooking and riding). We conduct extensive experiments, i.e., multi-source DG, single-source DG, and qualitative analysis, validating the rationality of our causal analysis and the effectiveness of the DeVADG framework. Shengyu Zhang 0001, Xusheng Feng, Wenyan Fan, Wenjing Fang, Fuli Feng, Wei Ji 0008, Li Wang 0056, Shanshan Zhao 0001, Zhou Zhao 0001, Tat-Seng Chua, Fei Wu 0001 |
AAAI | 4 |
| 2023 | Private, Efficient, and Accurate: Protecting Models Trained by Multi-party Learning with Differential PrivacyabstractSecure multi-party computation-based machine learning, referred to as multi-party learning (MPL for short), has become an important technology to utilize data from multiple parties with privacy preservation. While MPL provides rigorous security guarantees for the computation process, the models trained by MPL are still vulnerable to attacks that solely depend on access to the models. Differential privacy could help to defend against such attacks. However, the accuracy loss brought by differential privacy and the huge communication overhead of secure multi-party computation protocols make it highly challenging to balance the 3-way trade-off between privacy, efficiency, and accuracy.In this paper, we are motivated to resolve the above issue by proposing a solution, referred to as PEA (Private, Efficient, Accurate), which consists of a secure differentially private stochastic gradient descent (DPSGD for short) protocol and two optimization methods. First, we propose a secure DPSGD protocol to enforce DPSGD, which is a popular differentially private machine learning algorithm, in secret sharing-based MPL frameworks. Second, to reduce the accuracy loss led by differential privacy noise and the huge communication overhead of MPL, we propose two optimization methods for the training process of MPL: (1) the data-independent feature extraction method, which aims to simplify the trained model structure; (2) the local data-based global model initialization method, which aims to speed up the convergence of the model training. We implement PEA in two open-source MPL frameworks: TF-Encrypted and Queqiao. The experimental results on various datasets demonstrate the efficiency and effectiveness of PEA. E.g. when ϵ = 2, we can train a differentially private classification model with an accuracy of 88% for CIFAR-10 within 7 minutes under the LAN setting. This result significantly outperforms the one from CryptGPU, one state-of-the-art MPL framework: it costs more than 16 hours to train a non-private deep neural network model on CIFAR-10 with the same accuracy. Wenqiang Ruan, Mingxin Xu, Wenjing Fang, Li Wang 0056, Lei Wang 0152, Weili Han |
SP | 3 |
| 2023 | SecretFlow-SPU: A Performant and User-Friendly Framework for Privacy-Preserving Machine Learning
Junming Ma, Yancheng Zheng, Derun Zhao, Haoqi Wu, Wenjing Fang, Chaofan Yu, Benyu Zhang, Lei Wang 0152 |
USENIX ATC | 6 |
| 2023 | ElasticDL: A Kubernetes-native Deep Learning Framework with Fault-tolerance and Elastic SchedulingabstractThe power of artificial intelligence (AI) models originates with sophisticated model architecture as well as the sheer size of the model. These large-scale AI models impose new and challenging system requirements regarding scalability, reliability, and flexibility. One of the most promising solutions in the industry is to train these large-scale models on distributed deep-learning frameworks. With the power of all distributed computations, it is desired to achieve a training process with excellent scalability, elastic scheduling (flexibility), and fault tolerance (reliability). In this paper, we demonstrate the scalability, flexibility, and reliability of our open-source Elastic Deep Learning (ElasticDL) framework. Our ElasticDL utilizes an open-source system, i.e., Kubernetes, for automating deployment, scaling, and management of containerized application features to provide fault tolerance and support elastic scheduling for DL tasks. Jun Zhou 0011, Feng Zhu 0011, Qitao Shi, Wenjing Fang, Lin Wang 0098, Yi Wang 0141 |
WSDM | 5 |
| 2021 | Large-scale Secure XGB for Vertical Federated LearningabstractPrivacy-preserving machine learning has drawn increasingly attention recently, especially with kinds of privacy regulations come into force. Under such situation, Federated Learning (FL) appears to facilitate privacy-preserving joint modeling among multiple parties. Although many federated algorithms have been extensively studied, there is still a lack of secure and practical gradient tree boosting models (e.g., XGB) in literature. In this paper, we aim to build large-scale secure XGB under vertically federated learning setting. We guarantee data privacy from three aspects. Specifically, (1) we employ secure multi-party computation techniques to avoid leaking intermediate information during training, (2) we store the output model in a distributed manner in order to minimize information release, and (3) we provide a novel algorithm for secure XGB predict with the distributed model. Furthermore, by proposing secure permutation protocols, we can improve the training efficiency and make the framework scale to large dataset. We conduct extensive experiments on both public datasets and real-world datasets, and the results demonstrate that our proposed XGB models provide not only competitive accuracy but also practical performance. Wenjing Fang, Derun Zhao, Chaochao Chen 0001, Chaofan Yu, Li Wang 0056, Lei Wang 0152, Jun Zhou 0011, Benyu Zhang |
CIKM | 1 |
| 2021 | When Homomorphic Encryption Marries Secret Sharing: Secure Large-Scale Sparse Logistic Regression and Applications in Risk ControlabstractLogistic Regression (LR) is the most widely used machine learning model in industry for its efficiency, robustness, and interpretability. Due to the problem of data isolation and the requirement of high model performance, many applications in industry call for building a secure and efficient LR model for multiple parties. Most existing work uses either Homomorphic Encryption (HE) or Secret Sharing (SS) to build secure LR. HE based methods can deal with high-dimensional sparse features, but they incur potential security risks. SS based methods have provable security, but they have efficiency issue under high-dimensional sparse features. In this paper, we first present CAESAR, which combines HE and SS to build secure large-scale sparse logistic regression model and achieves both efficiency and security. We then present the distributed implementation of CAESAR for scalability requirement. We have deployed CAESAR in a risk control task and conducted comprehensive experiments. Our experimental results show that CAESAR improves the state-of-the-art model by around 130 times. Chaochao Chen 0001, Jun Zhou 0011, Li Wang 0056, Xibin Wu, Wenjing Fang, Lei Wang 0152, Alex X. Liu, Hao Wang 0007, Cheng Hong 0001 |
KDD | 5 |
| 2020 | Practical Privacy Preserving POI RecommendationabstractPoint-of-Interest (POI) recommendation has been extensively studied and successfully applied in industry recently. However, most existing approaches build centralized models on the basis of collecting users’ data. Both private data and models are held by the recommender, which causes serious privacy concerns. In this article, we propose a novel Privacy preserving POI Recommendation (PriRec) framework. First, to protect data privacy, users’ private data (features and actions) are kept on their own side, e.g., Cellphone or Pad. Meanwhile, the public data that need to be accessed by all the users are kept by the recommender to reduce the storage costs of users’ devices. Those public data include: (1) static data only related to the status of POI, such as POI categories, and (2) dynamic data dependent on user-POI actions such as visited counts. The dynamic data could be sensitive, and we develop local differential privacy techniques to release such data to the public with privacy guarantees. Second, PriRec follows the representations of Factorization Machine (FM) that consists of a linear model and the feature interaction model. To protect the model privacy, the linear models are saved on the users’ side, and we propose a secure decentralized gradient descent protocol for users to learn it collaboratively. The feature interaction model is kept by the recommender since there is no privacy risk, and we adopt a secure aggregation strategy in a federated learning paradigm to learn it. To this end, PriRec keeps users’ private raw data and models in users’ own hands, and protects user privacy to a large extent. We apply PriRec in real-world datasets, and comprehensive experiments demonstrate that, compared with FM, PriRec achieves comparable or even better recommendation accuracy. Chaochao Chen 0001, Jun Zhou 0011, Bingzhe Wu, Wenjing Fang, Li Wang 0056, Yuan Qi 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2019 | Adapted Tree Boosting for Transfer LearningabstractSecure online transaction is an essential task for e-commerce platforms. Alipay, one of the world’s leading cashless payment platform, provides the payment service to both merchants and individual customers. The fraud detection models are built to protect the customers, but stronger demands are raised by the new scenes, which are lacking in training data and labels. The proposed model makes a difference by utilizing the data under similar old scenes and the data under a new scene is treated as the target domain to be promoted. Inspired by this real case in Alipay, we view the problem as a transfer learning problem and design a set of revise strategies to transfer the source domain models to the target domain under the framework of gradient boosting tree models. This work provides an option for the cold-start and data-sharing problems. Wenjing Fang, Chaochao Chen 0001, Li Wang 0056, Jun Zhou 0011, Kenny Q. Zhu |
IEEE BigData | 1 |
| 2018 | Unpack Local Model Interpretation for GBDT
Wenjing Fang, Jun Zhou 0011, Xiaolong Li 0005, Kenny Q. Zhu |
DASFAA (2) | 1 |