VLDB 2026 Research / reviewers in the wild / expert
Li Bai 0004
dblp:181/2902-4
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-7202-3178ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Sample-Level Evaluation and Generative Framework for Model Inversion AttacksabstractModel Inversion (MI) attacks, which reconstruct the training dataset of neural networks, pose significant privacy concerns in machine learning. Recent MI attacks have managed to reconstruct realistic label-level private data, such as the general appearance of a target person from all training images labeled on him. Beyond label-level privacy, in this paper we show sample-level privacy, the private information of a single target sample, is also important but under-explored in the MI literature due to the limitations of existing evaluation metrics. To address this gap, this study introduces a novel metric tailored for training-sample analysis, namely, the Diversity and Distance Composite Score (DDCS), which evaluates the reconstruction fidelity of each training sample by encompassing various MI attack attributes. This, in turn, enhances the precision of sample-level privacy assessments. Leveraging DDCS as a new evaluative lens, we observe that many training samples remain resilient against even the most advanced MI attack. As such, we further propose a transfer learning framework that augments the generative capabilities of MI attackers through the integration of entropy loss and natural gradient descent. Extensive experiments verify the effectiveness of our framework on improving state-of-the-art MI attacks over various metrics including DDCS, coverage and FID. Finally, we demonstrate that DDCS can also be useful for MI defense, by identifying samples susceptible to MI attacks in an unsupervised manner. Haoyang Li 0018, Li Bai 0004, Qingqing Ye 0001, Haibo Hu 0001, Yaxin Xiao, Huadi Zheng, Jianliang Xu |
AAAI | 2 |
| 2025 | Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-ExpertsabstractMachine learning models are often vulnerable to inference attacks that expose sensitive information from their training data. Shadow model technique is commonly employed in such attacks, like membership inference. However, the need for a large number of shadow models leads to high computational costs, limiting their practical applicability. Such inefficiency mainly stems from the independent training and use of these shadow models. To address this issue, we present a novel shadow pool training framework SHAPOOL, which constructs multiple shared models and trains them jointly within a single process. In particular, we leverage the Mixture-of-Experts mechanism as the shadow pool to interconnect individual models, enabling them to share some sub-networks and thereby improving efficiency. To ensure the shared models closely resemble independent models and serve as effective substitutes, we introduce three novel modules: path-choice routing, pathway regularization, and pathway alignment. These modules guarantee random data allocation for pathway learning, promote diversity among shared models, and maintain consistency with target models. We evaluate SHAPOOL in the context of various membership inference attacks and show that it significantly reduces the computational cost of shadow model construction while maintaining comparable attack performance. Li Bai 0004, Qingqing Ye 0001, Xinwei Zhang 0002, Sen Zhang 0002, Zi Liang, Jianliang Xu, Haibo Hu 0001 |
NeurIPS | 1 |
| 2025 | MER-Inspector: Assessing Model Extraction Risks from An Attack-Agnostic PerspectiveabstractInformation leakage issues in machine learning-based Web applications have attracted increasing attention. While the risk of data privacy leakage has been rigorously analyzed, the theory of model function leakage, known as Model Extraction Attacks (MEAs), has not been well studied. In this paper, we are the first to understand MEAs theoretically from an attack-agnostic perspective and to propose analytical metrics for evaluating model extraction risks. By using the Neural Tangent Kernel (NTK) theory, we formulate the linearized MEA as a regularized kernel classification problem and then derive the fidelity gap and generalization error bounds of the attack performance. Based on these theoretical analyses, we propose a new theoretical metric called Model Recovery Complexity (MRC), which measures the distance of weight changes between the victim and surrogate models to quantify risk. Additionally, we find that victim model accuracy, which shows a strong positive correlation with model extraction risk, can serve as an empirical metric. By integrating these two metrics, we propose a framework, namely Model Extraction Risk Inspector (MER-Inspector), to compare the extraction risks of models under different model architectures by utilizing relative metric values. We conduct extensive experiments on 16 model architectures and 5 datasets. The experimental results demonstrate that the proposed metrics have a high correlation with model extraction risks, and MER-Inspector can accurately compare the extraction risks of any two models with up to 89.58%. Xinwei Zhang 0002, Haibo Hu 0001, Qingqing Ye 0001, Li Bai 0004, Huadi Zheng |
WWW | 4 |
| 2025 | RMR: A Relative Membership Risk Measure for Machine Learning ModelsabstractPrivacy leakage poses a significant threat when machine learning foundation models trained on private data are released. One such threat is membership inference attacks (MIA), which determine whether a specific example was included in a model's training set. This article shifts focus from developing new MIA algorithms to measuring a model's risk under MIA. We introduce a novel metric, Relative Membership Risk (RMR), which assesses a model's MIA vulnerability from a comparative standpoint. RMR calculates the difference in prediction loss for training examples relative to a predefined reference model, enabling risk comparison across models without needing to delve into details like training strategy, architecture, or data distribution. We also explore the selection of the reference model and show that using a high-risk reference model enhances the accuracy of the RMR measure. To identify the most vulnerable reference model, we propose an efficient iterative algorithm that selects the optimal model from a set of candidates. Through extensive empirical evaluations on various datasets and network architectures, we demonstrate that RMR is an accurate and efficient tool for measuring the membership privacy risk of both individual training examples and the overall machine learning model. Li Bai 0004, Haibo Hu 0001, Qingqing Ye 0001, Jianliang Xu, Jin Li 0002, Chengfang Fang, Jie Shi 0005 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | ProVFL: Property Inference Attacks Against Vertical Federated LearningabstractRecent studies show that privacy leakages may occur in vertical federated learning (VFL), where parties hold split features of the same samples. While various attacks, including label and feature inference, focus on record-level privacy risks in VFL, few studies delve into the distribution-level privacy threat. In this paper, we explore property inference attacks (PIAs) in VFL, where an adversarial party seeks to deduce global distribution information about a target property in the victim party’s training set. Our key observation is that theLp-norm distribution of intermediate results in VFL could reflect the fraction of the target property in a training set. Inspired by this, we presentProVFL, a novel PIA framework involving distribution comparison and correlation augmentation modules. To achieve property inference, we design a distribution comparison module by creating various intermediate-result populations with different proportions, aiming to learn the relationship betweenLp-norm distributions and their fractions. Then, we theoretically analyze the factors that contribute to the attack effectiveness and develop a correlation augmentation module based on label replacement and model refinement to amplify property information leakage. Extensive experimental results demonstrate that our attacks can achieve inferences with low estimation errors as low as 1%. This poses the immediate threat of property information leakage from private training data in the VFL setting. Li Bai 0004, Xinwei Zhang 0002, Sen Zhang 0002, Qingqing Ye 0001, Haibo Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Auditing MLaaS Inference Service Quality Without Ground Truth via Mutual InformationabstractMachine Learning as a Service (MLaaS) paradigm offers an appealing solution for clients that have limited computational resources. It allows entities to train models with collected dataset and powerful cloud resources, and to deploy these models for inference. However, MLaaS currently faces significant challenges in ensuring trustworthy inference and service quality. The clients cannot verify that the inference results returned by service provider (SP) are the model’s actual inference results. Moreover, even if clients manage to ensure that the results are obtained through model inference, they are unable to determine the model’s service quality without ground truth. To address these concerns, we introduce an innovative framework to audit inference quality and integrity in MLaaS through a novel deep neural network (DNN) inspection method. In specific, our approach represents the inherent behavior of the model by collecting its intermediate layer outputs and quantifying the mutual information (MI) values derived from them. By benchmarking the model during the training process, the SP can record the characteristics of the correct model and its corresponding service quality. After receiving the auditing request, the auditor can evaluate the quality of the service by estimating its accuracy via mutual information. Moreover, it can confirm the integrity of the returned results by inspecting the intermediate layer output. In addition, we thoroughly analyze our scheme for various potential adaptive attacks. Through empirical studies, we verify the correctness, effectiveness, and robustness of our scheme for trustworthy MLaaS inference service. Qingqing Ye 0001, Haibo Hu 0001, Li Bai 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Improve Word Mover's Distance with Part-of-Speech TaggingabstractWord Mover's Distance (WMD) is a document distance metric with free parameter, intelligible interpretation and unprecedented accuracy on document classification. WMD is on the basis of word embedding and largely focuses on semantic relationships rather than syntactic relationships, which would bring some limitations on measuring document distance. To enhance the impact of syntactic information, we proposed a new method called WMD with Part-of-Speech (PWMD) that integrates part-of-speech (POS) into the original WMD model. POS is a kind of syntactic information, providing more valuable features combined with WMD in document distance metric. Two combination strategies of the POS tagging are provided in “WMD, “word level” and “document level”. The results of contrastive experiments have shown that the PWMD is able to get better document distance than WMD. Xiaojun Chen 0004, Li Bai 0004, Dakui Wang, Jinqiao Shi |
ICPR | 2 |