VLDB 2026 Research / reviewers in the wild / expert
Minxin Du
dblp:169/2251
· DBLP profile ↗
24ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-6620-6923ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 11 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shared Spotlight Meridian: Distributed Sparse Pseudorandom Functions for Scalable Federated Learning
Youlong Ding, Peihua Mai, Sherman S. M. Chow, Minxin Du |
SP | 5 |
| 2026 | Practical Anonymous Two-Party Gradient Boosting Decision TreeabstractStructured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutually distrustful parties. High speed and interpretability make GBDTs popular in finance and healthcare, where neural networks may fall short. Enabling secure computation for GBDTs poses unique challenges, requiring secure record alignment for comparison. Relying on private set intersection (PSI) is a de facto approach. Mistaking PSI for a safety measure actually exposes which record identifiers (IDs) are shared between the datasets. Although circuit-PSI could help, it is costly for generic uses. New ideas are needed to efficiently train in a "dark forest". Aiming to hide the IDs, we initiate the study of anonymous GBDT training on split data held by two parties. Dual circuit-PSI in our design lets the parties alternate as receiver to run pick-then-sum over local features. Via oblivious programmable pseudorandom functions, we propagate circuit-PSI outputs as shared state across runs. Avoiding universal alignment, we resolve the neglected dilemma that ID hiding incurs a cost that scales with domain size. Next, we halve the cost of ciphertext packing used to convert single-instruction multiple-data homomorphic encryption from (ring) learning with errors in prior secure GBDT (Usenix Security' 23) and related secure machine-learning computations. Comparative experiments show our protocol remains competitive with leaky approaches in efficiency. Enabling ID-hiding aggregation, our techniques can extend to other vertically partitioned analytics. Minxin Du, Sherman S. M. Chow, Huangxun Chen, Huaming Rao, Danqing Huang, Peng Chen 0021 |
SP | 3 |
| 2025 | OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsabstractLarge language models (LLMs) trained over extensive corpora risk memorizing sensitive, copyrighted, or toxic content.To address this, we propose OBLIVIATE, a robust unlearning framework that removes targeted data while preserving model utility.The framework follows a structured process: extracting target tokens, building retain sets, and fine-tuning with a tailored loss function comprising three components-masking, distillation, and world fact.Using low-rank adapters (LoRA) ensures efficiency without compromising unlearning quality.We conduct experiments on multiple datasets, including Harry Potter series, WMDP, and TOFU, using a comprehensive suite of metrics: forget quality (via a new document-level memorization score), model utility, and fluency.Results demonstrate its effectiveness in resisting membership inference attacks, minimizing the impact on retained data, and maintaining robustness across diverse scenarios. Minxin Du, Qingqing Ye 0001, Haibo Hu 0001 |
EMNLP | 2 |
| 2025 | SecEmb: Sparsity-Aware Secure Federated Learning of On-Device Recommender System with Large EmbeddingabstractFederated recommender system (FedRec) has emerged as a solution to protect user data through collaborative training techniques. A typical FedRec involves transmitting the full model and entire weight updates between edge devices and the server, causing significant burdens to edge devices with limited bandwidth and computational power. The sparsity of embedding updates provides opportunity for payload optimization, while existing sparsity-aware federated protocols generally sacrifice privacy for efficiency. A key challenge in designing a secure sparsity-aware efficient protocol is to protect the rated item indices from the server. In this paper, we propose a lossless secure recommender systems with on sparse embedding updates (SecEmb). SecEmb reduces user payload while ensuring that the server learns no information about both rated item indices and individual updates except the aggregated model. The protocol consists of two correlated modules: (1) a privacy-preserving embedding retrieval module that allows users to download relevant embeddings from the server, and (2) an update aggregation module that securely aggregates updates at the server. Empirical analysis demonstrates that SecEmb reduces both download and upload communication costs by up to 90x and decreases user-side computation time by up to 70x compared with secure FedRec protocols. Additionally, it offers non-negligible utility advantages compared with lossy message compression methods. Peihua Mai, Youlong Ding, Ziyan Lyu, Minxin Du |
ICML | 4 |
| 2024 | Machine Unlearning of Pre-trained Large Language ModelsabstractThis study investigates the concept of the 'right to be forgotten' within the context of large language models (LLMs).We explore machine unlearning as a pivotal solution, with a focus on pre-trained models-a notably under-researched area.Our research delineates a comprehensive framework for machine unlearning in pretrained LLMs, encompassing a critical analysis of seven diverse unlearning methods.Through rigorous evaluation using curated datasets from arXiv, books, and GitHub, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over 10 5 times more computationally efficient than retraining.Our results show that integrating gradient ascent with gradient descent on in-distribution data improves hyperparameter robustness.We also provide detailed guidelines for efficient hyperparameter tuning in the unlearning process.Our findings advance the discourse on ethical AI practices, offering substantive insights into the mechanics of machine unlearning for pretrained LLMs and underscoring the potential for responsible AI development.1 Eli Chien, Minxin Du, Xinyao Niu, Tianhao Wang 0001, Zezhou Cheng, Xiang Yue |
ACL (1) | 3 |
| 2024 | FastTextDodger: Decision-Based Adversarial Attack Against Black-Box NLP Models With Extremely High EfficiencyabstractRecently, achieving query-efficient adversarial example attacks targeting black-box natural language models has attracted widespread attention from researchers. This task is considered difficult due to the discrete nature of texts, limited knowledge of the target model, and strict query access limitations in real-world systems. However, existing attacks often require a large number of queries or result in low attack success rates, having not met practical requirements. To address this, we propose FastTextDodger, a simple and compact decision-based black-box textual adversarial attack that generates grammatically correct adversarial texts with high attack success rates and few queries. Experimental results show that FastTextDodger achieves an impressive 97.4% attack success rate on benchmark datasets and models, and only needs about 200 queries. Compared to state-of-the-art attacks, FastTextDodger only requires one-tenth of the number of queries in text classification and entailment tasks while maintaining comparable attack success rates and perturbed word rates. Xiaoxue Hu, Geling Liu, Baolin Zheng, Lingchen Zhao, Qian Wang 0002, Minxin Du |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Stealthy and Effective Physical Adversarial Attacks in Autonomous DrivingabstractIn autonomous assistance systems, accurate camera vision is indispensable for driving safety. Featuring this safety-critical scenario, physical adversarial examples are arguably the most threatening. However, existing physical adversarial attacks are either conspicuous or ineffective, leaving a dilemma for balancing between attack effectiveness and stealthiness. In this paper, we put forth a new type of adversarial patch attack, leveraging the behavior characteristic of high-speed shutters in autonomous driving cameras. Instead of deploying static images onto real-world objects, we embed adversarial examples into videos and cast them with projectors. By delicately deciding the frame contents and display frequencies, the adversarial frame contents are almost invisible to humans, but high-speed shutters can capture them (see our demos (https://anonymous.4open.science/r/7003)). We demonstrate the attack feasibility by fooling traffic sign detectors, altering speed limit signs to be mis-detected or undetected, and hence manipulating the driving speed of victims. We conduct extensive experiments under various conditions, confirming that our new attack is effective and robust: It can deceive state-of-the-art detector models with success rates of 99% and over 90% in untargeted and targeted manners, respectively. Our further investigations unravel the transferability of our attack to other detectors in a black-box setting. Man Zhou 0004, Wenyu Zhou, Junhui Yang, Minxin Du, Qi Li 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Cryptography-Inspired Federated Learning for Generative Adversarial Networks and Meta Learning
Yu Zheng 0021, Minxin Du, Sherman S. M. Chow, Qian Lou, Yongjun Zhao 0001, Xiuhua Wang 0009 |
ADMA (2) | 3 |
| 2023 | DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassabstractDifferentially private stochastic gradient descent (DP-SGD) adds noise to gradients in back-propagation, safeguarding training data from privacy leakage, particularly membership inference. It fails to cover (inference-time) threats like embedding inversion and sensitive attribute inference. It is also costly in storage and computation when used to fine-tune large pre-trained language models (LMs). Minxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang 0001, Huan Sun 0001 |
CCS | 1 |
| 2023 | Sanitizing Sentence Embeddings (and Labels) for Local Differential PrivacyabstractDifferentially private (DP) learning, notably DP stochastic gradient descent (DP-SGD), has limited applicability in fine-tuning gigantic pre-trained language models (LMs) for natural language processing tasks. The culprit is the perturbation of gradients (as gigantic as entire models), leading to significant efficiency and accuracy drops. Minxin Du, Xiang Yue, Sherman S. M. Chow, Huan Sun 0001 |
WWW | 1 |
| 2023 | Shielding Graph for eXact Analytics With SGXabstractGraphs nicely capture data from various domains, allowing the computations of many analytic tasks via graph queries. Graphs of real-world data are often large, albeit useful, and the involved computation can be too heavyweight for commodity computers. For secure outsourcing, we propose (SGX)$^{2}$, a forward-secure structured encryption scheme for graph data, which uses lightweight cryptographic techniques with a trusted execution environment such as SGX. To process million-scale graphs by the limited memory of SGX, we load data on-demand using Dijkstra's algorithm and Fibonacci heap. Compared with most prior graph encryption schemes, (SGX)$^{2}$supports exact shortest-distance queries instead of approximation and can be easily extended to other graph-based analytics. Minxin Du, Peipei Jiang 0002, Qian Wang 0002, Sherman S. M. Chow, Lingchen Zhao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | AdvDDoS: Zero-Query Adversarial Attacks Against Commercial Speech Recognition SystemsabstractAutomatic speech recognition (ASR) has been widely and commercially employed in health care, autonomous vehicles, and finance. Yet, recent studies have shown that universal adversarial perturbations (UAPs) pose a serious threat to white-box ASR systems, when the adversary has access to the target model. Until now, the impacts of such a threat on commercial systems are still open since their models are not publicly available. To understand the security weakness in the practical black-box setting, this paper introduces the firstzero-queryUAP attacks, called AdvDDoS, with black-box access to ASR systems: we do not need to pay any query expense to estimate UAPs. Specifically, we craft targeted UAPs under a popular feature extractor and a local ASR model by reversing the robust target-category features, in which adversarial perturbations containing robust features are believed to have better transferability. Compared with vanilla UAPs, our UAPs incorporated with target-category features lead to better attacks against commercial ASR systems. We validate the efficacy of our AdvDDoS by launching attacks against a range of commercial ASR systems,i.e., three API services (Alibaba, Tencent, and Baidu), and three personal assistants (Apple Siri, iFlytek, and Google). Extensive experimental results demonstrate the superiority of AdvDDoS. For example, AdvDDoS achieves 83.26% word error rate (WER) and 53.25% success rates of attacks (SRoA) for the universal attack against Tencent ASR API, which outperforms the vanilla UAPs by up to 61.56% on WER and 11.6% on SRoA. The success of our attack sheds light on zero-query UAP attacks against Commercial ASR systems. Yunjie Ge, Lingchen Zhao, Qian Wang 0002, Yiheng Duan, Minxin Du |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Encrypted video search: scalable, modular, and content-similarabstractVideo-based services have become popular. Clients often outsource their videos to the cloud to relieve local maintenance. However, privacy has emerged as a major concern since many videos contain sensitive information. While retrieving (unencrypted) videos has been widely studied, encrypted multimedia retrieval receives rare attention, at best in a limited form of similarity searches on images. Yu Zheng 0021, Heng Tian, Minxin Du, Chong Fu 0001 |
MMSys | 3 |
| 2022 | GraphShield: Dynamic Large Graphs for Secure Queries With Forward PrivacyabstractThe increasing amount of graph-structured data catalyzes analytics over graph databases using semantic queries. Motivated by the ubiquity of commercial cloud platforms, data owners are willing to store their graph databases remotely. However, data privacy has emerged as a widespread concern since the cloud platforms are not fully trusted. One viable solution is to encrypt sensitive data before outsourcing, which inevitably hinders data retrieval. To enable queries over encrypted data, searchable symmetric encryption (SSE) has been introduced. Yet, the most well-studied class of SSE schemes focuses on retrieving textual files given keywords, which cannot be applied to graph databases directly. This paper extends our preliminary work (FC′17) and proposes GraphShield, a structured encryption scheme for graphs. Beyond shortest distance queries, GraphShield can support other classic graph-based queries (e.g., maximum flow) and more complicated analytics (e.g., PageRank). Technically, we incorporate a suite of (efficient) cryptographic primitives and tailor some extra secure protocols for facilitating graph analytics. Our scheme also allows updates on the encrypted graph with forward privacy guaranteed. We formalize the security model and prove the adaptive security with reasonable leakage. Finally, we implement our scheme on various real-world datasets, and the experiment results demonstrate its practicality and scalability. Minxin Du, Shuangke Wu, Qian Wang 0002, Dian Chen 0004, Peipei Jiang 0002, David Mohaisen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Stargazing in the Dark: Secure Skyline Queries with SGX
Jiafan Wang 0001, Minxin Du, Sherman S. M. Chow |
DASFAA (3) | 2 |
| 2019 | ARMOR: A Secure Combinatorial Auction for Heterogeneous SpectrumabstractDynamic spectrum allocation via auction is an effective solution to spectrum shortage. Combinatorial spectrum auction enables buyers to express diversified preferences towards different combinations of channels. Despite the effort to ensure truthfulness and maximize social welfare, spectrum auction also faces potential security risks. The leakage of sensitive information such as true valuation and location of bidders may incur severe economic damage. However, there is a lack of works that can provide sufficient protection against such security risks in combinatorial spectrum auction. In this paper, we propose ARMOR, to enable combinatorial auction for heterogeneous spectrum with privacy, which can preserve bidders' privacy while guaranteeing the economic-robustness of the combinatorial auction. We leverage the cryptographic methods, including homomorphic encryption, order-preserving encryption, and garbled circuits, to shield the bid and location information of buyers from the auctioneer. We design a novel location protection algorithm, which allows the auctioneer to exploit spectrum reuse opportunities without knowing the exact locations of buyers. Furthermore, we propose a verifiable payment scheme based on digital signature to prevent the auctioneer from forging the payment. The extensive experiments confirm that ARMOR maintains the good performance of the combinatorial spectrum auction, in terms of buyer satisfactory ratio and social welfare, and achieves privacy preservation with acceptable computation and communication costs. Yanjiao Chen, Qian Wang 0002, Minxin Du, Qi Li 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2018 | InstantCryptoGram: Secure Image Retrieval ServiceabstractImage retrieval is crucial for social media sites such as Instagram to identify similar images and make recommendations for users who share similar interests. To get rid of the storage burden and computation for image retrieval, outsourcing to a remote cloud is now a trend. Yet, privacy concerns mandate the use of encryption before outsourcing the images. We need a secure way for retrieving images from a not-fully-trusted server. This paper proposes InstantCryptoGram, a secure image retrieval service. We first design a new data structure called sub-simhash, which fits for the inverted index used by many searchable symmetric encryption schemes. It leads to our modular solution that supports efficient similarity queries and updates over encrypted images. Our experiments on Amazon AWS EC2 over representative datasets show that our scheme is efficient and accurate in finding similar images while preserving privacy. Mingxue Zhang 0001, Qian Wang 0002, Sherman S. M. Chow, Minxin Du, Yanjiao Chen, Chenliang Li 0005 |
INFOCOM | 5 |
| 2018 | Searchable Encryption over Feature-Rich DataabstractStorage services allow data owners to store their huge amount of potentially sensitive data, such as audios, images, and videos, on remote cloud servers in encrypted form. To enable retrieval of encrypted files of interest, searchable symmetric encryption (SSE) schemes have been proposed. However, many schemes construct indexes based on keyword-file pairs and focus on boolean expressions of exact keyword matches. Moreover, most dynamic SSE schemes cannot achieve forward privacy and reveal unnecessary information when updating the encrypted databases. We tackle the challenge of supporting large-scale similarity search over encrypted feature-rich multimedia data, by considering the search criteria as a high-dimensional feature vector instead of a keyword. Our solutions are built on carefully-designed fuzzy Bloom filters which utilize locality sensitive hashing (LSH) to encode an index associating the file identifiers and feature vectors. Our schemes are proven to be secure against adaptively chosen query attack and forward private in the standard model. We have evaluated the performance of our scheme on real-world high-dimensional datasets, and achieved a search quality of 99 percent recall with only a few number of hash tables for LSH. This shows that our index is compact and searching is not only efficient but also accurate. Qian Wang 0002, Meiqi He, Minxin Du, Sherman S. M. Chow, Russell W. F. Lai, Qin Zou 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2018 | Privacy-Preserving Indexing and Query Processing for Secure Dynamic Cloud StorageabstractWith the increasing popularity of cloud-based data services, data owners are highly motivated to store their huge amount of potentially sensitive personal data files on remote servers in encrypted form. Clients later can query over the encrypted database to retrieve files while protecting privacy of both the queries and the database, by allowing some reasonable leakage information. To this end, the notion of searchable symmetric encryption (SSE) was proposed. Meanwhile, recent literature has shown that most dynamic SSE solutions leaking information on updated keywords are vulnerable to devastating file-injection attacks. The only way to thwart these attacks is to design forward-private schemes. In this paper, we investigate new privacy-preserving indexing and query processing protocols which meet a number of desirable properties, including the multi-keyword query processing with conjunction and disjunction logic queries, practically high privacy guarantees with adaptive chosen keyword attack (CKA2) security and forward privacy, the support of dynamic data operations, and so on. Compared with previous schemes, our solutions are highly compact, practical, and flexible. Their performance and security are carefully characterized by rigorous analysis. Experimental evaluations conducted over a large representative data set demonstrate that our solutions can achieve modest search time efficiency, and they are practical for use in large-scale encrypted database systems. Minxin Du, Qian Wang 0002, Meiqi He, Jian Weng 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Outsourced Biometric Identification With PrivacyabstractBiometric identification typically scans a large-scale database of biometric records for finding a close enough match of an individual. This paper investigates how to outsource this computationally expensive scanning while protecting the privacy of both the database and the computation. Exploiting the inherent structures of biometric data and the properties of identification operations, we first present a privacy-preserving biometric identification scheme which uses a single server. We then consider its extensions in the two-server model. It achieves a higher level of privacy than our single-server solution assuming two servers are not colluding. Apart from somewhat homomorphic encryption, our second scheme uses batched protocols for secure shuffling and minimum selection. Our experiments on both synthetic and real data sets show that our solutions outperform existing schemes while preserving privacy. Shengshan Hu, Qian Wang 0002, Sherman S. M. Chow, Minxin Du |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2018 | Privacy-Preserving Collaborative Model Learning: The Case of Word Vector TrainingabstractNowadays, machine learning is becoming a new paradigm for mining hidden knowledge in big data. The collection and manipulation of big data not only create considerable values, but also raise serious privacy concerns. To protect the huge amount of potentially sensitive data, a straightforward approach is to encrypt data with specialized cryptographic tools. However, it is challenging to utilize or operate on encrypted data, especially to perform machine learning algorithms. In this paper, we investigate the problem of training high quality word vectors over large-scale encrypted data (from distributed data owners) with the privacy-preserving collaborative neural network learning algorithms. We leverage and also design a suite of arithmetic primitives (e.g., multiplication, fixed-point representation, sigmoid function computation, etc.) on encrypted data, served as components of our construction. We theoretically analyze the security and efficiency of our proposed construction, and conduct extensive experiments on representative real-world datasets to verify its practicality and effectiveness. Qian Wang 0002, Minxin Du, Xiuying Chen, Yanjiao Chen, Pan Zhou 0001, Xiaofeng Chen 0001, Xinyi Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Learning privately: Privacy-preserving canonical correlation analysis for cross-media retrievalabstractA massive explosion of various types of data has been triggered in the “Big Data” era. In big data systems, machine learning plays an important role due to its effectiveness in discovering hidden information and valuable knowledge. Data privacy, however, becomes an unavoidable concern since big data usually involve multiple organizations, e.g., different healthcare systems and hospitals, who are not in the same trust domain and may be reluctant to share their data publicly. Applying traditional cryptographic tools is a straightforward approach to protect sensitive information, but it often renders learning algorithms useless inevitably. In this work, we, for the first time, propose a novel privacy-preserving scheme for canonical correlation analysis (CCA), which is a well-known learning technique and has been widely used in cross-media retrieval system. We first develop a library of building blocks to support various arithmetics over encrypted real numbers by leveraging additively homomorphic encryption and garbled circuits. Then we encrypt private data by randomly splitting the numerical data, formalize CCA problem and reduce it to a symmetric eigenvalue problem by designing new protocols for privacy-preserving QR decomposition. Finally, we solve all the eigenvalues and the corresponding eigenvectors by running Newton-Raphson method and inverse power method over the ciphertext domain. We carefully analyze the security and extensively evaluate the effectiveness of our design. The results show that our scheme is practically secure, incurs negligible errors compared with performing CCA in the clear and performs comparably in cross-media retrieval systems. Qian Wang 0002, Shengshan Hu, Minxin Du, Jingjun Wang, Kui Ren 0001 |
INFOCOM | 3 |
| 2016 | Catch me in the dark: Effective privacy-preserving outsourcing of feature extractions over image dataabstractAdvances in cloud computing have greatly motivated data owners to outsource their huge amount of personal multimedia data and/or computationally expensive tasks onto the semi-trusted cloud by leveraging its abundant resources for cost saving and flexibility. From the privacy perspective, however, the outsourced multimedia data and its originated applications may reveal the data owner's private information, such as the personal identity, locations or even financial profiles. This observation has recently aroused new research interest on privacy-preserving computations over outsourced multimedia data. In this paper, we propose an effective privacy-preserving computation outsourcing protocol for the prevailing scale-invariant feature transform (SIFT) over massive encrypted image data. We first show that previous solutions to this problem have either efficiency/security or practicality issues, and none can well preserve the important characteristics of the original SIFT in terms of distinctiveness and robustness. We for the first time present a new privacy-preserving outsourcing protocol for SIFT with the preservation of its key characteristics, by randomly splitting the original image data, carefully distributing the feature extraction computations to two independent cloud servers and further leveraging the garbled circuit for secure keypoints comparisons. We both carefully analyze and extensively evaluate the security and effectiveness of our design. The results show that our solution is practically secure, outperforms the state-of-the-art and performs comparably to the original SIFT in terms of various characteristics, including rotation invariance, image scale invariance, robust matching across affine distortion and change in 3D viewpoint. Qian Wang 0002, Shengshan Hu, Kui Ren 0001, Jingjun Wang, Zhibo Wang 0001, Minxin Du |
INFOCOM | 6 |
| 2015 | CloudBI: Practical Privacy-Preserving Outsourcing of Biometric Identification in the Cloud
Qian Wang 0002, Shengshan Hu, Kui Ren 0001, Meiqi He, Minxin Du, Zhibo Wang 0001 |
ESORICS (2) | 5 |