Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiaofan Bai

dblp:384/4279 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0004-6796-3773ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
6 papers
Security and privacy of machine learning · 100%
Artificial intelligence
3 papers
Trustworthy machine learning · 64% Language models and text generation · 36%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning › model intellectual property protection › model ownership verification
model fingerprinting
2.432025
RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language Models · EMNLP 2025
Towards Stricter Black-box Integrity Verification of Deep Neural Network Models · ACM Multimedia 2024
Intersecting-Boundary-Sensitive Fingerprinting for Tampering Detection of DNN Models · ICML 2024
Security and privacy of machine learning › model security
model integrity
1.722025
RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language Models · EMNLP 2025
SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN Models · CVPR 2025
Security and privacy of machine learning
adversarial attack
1.622025
Consensus-Robust Transfer Attacks via Parameter and Representation Perturbations · NeurIPS 2025
DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable Defenses · NDSS 2024
Security and privacy of machine learning › verifiable machine learning
model integrity verification
1.522024
Towards Stricter Black-box Integrity Verification of Deep Neural Network Models · ACM Multimedia 2024
Intersecting-Boundary-Sensitive Fingerprinting for Tampering Detection of DNN Models · ICML 2024
Security and privacy of machine learning
adversarial robustness
0.912025
Consensus-Robust Transfer Attacks via Parameter and Representation Perturbations · NeurIPS 2025
Security and privacy of machine learning › adversarial attack
transferable adversarial attack
0.912025
Consensus-Robust Transfer Attacks via Parameter and Representation Perturbations · NeurIPS 2025
Security and privacy of machine learning › adversarial attack › physical adversarial attack
adversarial patch
0.812024
DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable Defenses · NDSS 2024
Security and privacy of machine learning › adversarial robustness
certified robustness
0.812024
DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable Defenses · NDSS 2024
Natural language and speech › Language models and text generation › large language model
large language model deployment
0.312025
RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language Models · EMNLP 2025
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.212024
DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable Defenses · NDSS 2024

Methods — techniques the papers use, named apart from their topics

fingerprinting · 2.3sequential test · 1.7hypothesis testing · 1.7entropy-gradient norm · 1.7LoRA fine-tuning · 1.7KL divergence surrogate · 1.7decision boundary analysis · 1.6monte carlo sampling · 0.9gradient sensitivity analysis · 0.9consensus-robust optimization · 0.9occlusion modeling · 0.8meta-learning · 0.8distributed optimization · 0.8
YearPublicationVenuePosition
2025 SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN Models
abstract
Cloud-based AI systems offer significant benefits but also introduce vulnerabilities, making deep neural network (DNN) models susceptible to malicious tampering. This tampering may involve harmful behavior injection or resource reduction, compromising model integrity and performance. To detect model tampering, hard-label fingerprinting techniques generate sensitive samples to probe and reveal tampering. Existing fingerprinting methods are mainly based on gradient-defined sensitivity or decision boundary, with the latter showing a manifest superior detection performance. However, all existing fingerprinting methods either suffer from insufficient sensitivity or incur high computational costs.In this paper, we theoretically analyze the black-box co-optimal tampering detection sensitivity of fingerprint samples in the context of decision boundary and gradient-defined sensitivity. Based on this, we further propose Steep-Decision-Boundary Fingerprinting (SDBF), a novel lightweight approach for hard-label tampering detection that inherently and efficiently combines the strengths of existing fingerprinting techniques. SDBF places fingerprint samples near the steep decision boundary, where the outputs of samples are inherently highly sensitive to tampering. We also design a Max Boundary Coverage Strategy (MBCS), which enhances samples’ diversity over the decision boundary. Theoretical analysis and extensive experimental results show that SDBF outperforms existing SOTA hard-label fingerprinting methods in both sensitivity and efficiency.
Xiaofan Bai, Shixin Li 0001, Xiaojing Ma 0002, Bin B. Zhu, Dongmei Zhang 0001, Linchen Yu
CVPR1
2025 RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language Models
abstract
The proliferation of Machine Learning as a Service (MLaaS) has enabled widespread deployment of large language models (LLMs) via cloud APIs, but also raises critical concerns about model integrity and security.Existing black-box tamper detection methods, such as watermarking and fingerprinting, rely on the stability of model outputs-a property that does not hold for inherently stochastic LLMs.We address this challenge by formulating blackbox tamper detection for LLMs as a hypothesistesting problem.To enable efficient and sensitive fingerprinting, we derive a first-order surrogate for KL divergence-the entropy-gradient norm-to identify prompts most responsive to parameter perturbations.Building on this, we propose Regularized Entropy-Sensitive Fingerprinting (RESF), which enhances sensitivity while regularizing entropy to improve output stability and control false positives.To further distinguish tampering from benign randomness, such as temperature shifts, RESF employs a lightweight two-tier sequential test combining support-based and distributional checks with rigorous false-alarm control.Comprehensive analysis and experiments across multiple LLMs show that RESF achieves up to 98.80% detection accuracy under challenging conditions, such as minimal LoRA fine-tuning with five optimized fingerprints.RESF consistently demonstrates strong sensitivity and robustness, providing an effective and scalable solution for black-box tamper detection in cloud-deployed LLMs.
Pingyi Hu, Xiaofan Bai, Xiaojing Ma 0002, Chaoxiang He, Dongmei Zhang 0001, Bin B. Zhu
EMNLP2
2025 Consensus-Robust Transfer Attacks via Parameter and Representation Perturbations
abstract
Adversarial examples crafted on one model often exhibit poor transferability to others, hindering their effectiveness in black-box settings. This limitation arises from two key factors: (i) \emph{decision-boundary variation} across models and (ii) \emph{representation drift} in feature space. We address these challenges through a new perspective that frames transferability for \emph{untargeted attacks} as a \emph{consensus-robust optimization} problem: adversarial perturbations should remain effective across a neighborhood of plausible target models. To model this uncertainty, we introduce two complementary perturbation channels: a \emph{parameter channel}, capturing boundary shifts via weight perturbations, and a \emph{representation channel}, addressing feature drift via stochastic blending of clean and adversarial activations. We then propose \emph{CORTA} (COnsensus--Robust Transfer Attack), a lightweight attack instantiated from this robust formulation using two first-order strategies: (i) sensitivity regularization based on the squared Frobenius norm of logits’ Jacobian with respect to weights, and (ii) Monte Carlo sampling for blended feature representations. Our theoretical analysis provides a certified lower bound linking these approximations to the robust objective. Extensive experiments on CIFAR-100 and ImageNet show that CORTA significantly outperforms state-of-the-art transfer-based methods---including ensemble approaches---across CNN and Vision Transformer targets. Notably, CORTA achieves a \emph{19.1 percentage-point gain in transfer success rate over the best prior method} while using only a single surrogate model.
Shixin Li 0001, Xiaojing Ma 0002, Xiaofan Bai, Pingyi Hu, Dongmei Zhang 0001, Bin B. Zhu
NeurIPS4
2024 Intersecting-Boundary-Sensitive Fingerprinting for Tampering Detection of DNN Models
abstract
Cloud-based AI services offer numerous benefits but also introduce vulnerabilities, allowing for tampering with deployed DNN models, ranging from injecting malicious behaviors to reducing computing resources. Fingerprint samples are generated to query models to detect such tampering. In this paper, we present Intersecting-Boundary-Sensitive Fingerprinting (IBSF), a novel method for black-box integrity verification of DNN models using only top-1 labels. Recognizing that tampering with a model alters its decision boundary, IBSF crafts fingerprint samples from normal samples by maximizing the partial Shannon entropy of a selected subset of categories to position the fingerprint samples near decision boundaries where the categories in the subset intersect. These fingerprint samples are almost indistinguishable from their source samples. We theoretically establish and confirm experimentally that these fingerprint samples’ expected sensitivity to tampering increases with the cardinality of the subset. Extensive evaluation demonstrates that IBSF surpasses existing state-of-the-art fingerprinting methods, particularly with larger subset cardinality, establishing its state-of-the-art performance in black-box tampering detection using only top-1 labels. The IBSF code is available at https://github.com/CGCL-codes/IBSF.
Xiaofan Bai, Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Hai Jin 0001
ICML1
2024 Towards Stricter Black-box Integrity Verification of Deep Neural Network Models
abstract
Cloud-based machine learning services offer significant advantages but also introduce the risk of tampering with cloud-deployed deep neural network (DNN) models. Black-box integrity verification (BIV) allows model owners and end-users to determine if a cloud-deployed DNN model has been tampered with by examining only the top-1 label responses. Fingerprinting generates fingerprint samples to query the model, achieving BIV with no impact on the model's accuracy. In this paper, we present BIVBench, the first comprehensive benchmark for BIV of DNN models. BIVBench covers 16 types of model modifications, providing extensive coverage of practical modification scenarios. Our analysis reveals that existing fingerprinting methods, which are typically focused on significant tampering, lack the sensitivity needed to effectively detect subtle yet common and potentially severe modifications. To address this limitation, we propose MiSentry (Model Integrity Sentry), a novel fingerprinting method that leverages meta-learning. MiSentry strategically incorporates a few subtly modified models into the meta-learning model zoo and maximizes the divergence of output predictions between the target model and the modified models in the model zoo to generate highly sensitive, generalizable, and effective fingerprint samples. Extensive evaluations using BIVBench demonstrate that MiSentry outperforms existing state-of-the-art methods overall and significantly surpasses them in detecting subtle modifications. The BIVBench and supplementary materials are available at: https://github.com/CGCL-codes/BIVBench.
Chaoxiang He, Xiaofan Bai, Xiaojing Ma 0002, Bin B. Zhu, Pingyi Hu, Jiayun Fu, Hai Jin 0001, Dongmei Zhang 0001
ACM Multimedia2
2024 DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable Defenses
Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Yimiao Zeng, Hanqing Hu, Xiaofan Bai, Hai Jin 0001, Dongmei Zhang 0001
NDSS6