Yaxin Xiao

dblp:346/0886 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
8 papers
Security and privacy of machine learning · 72% Privacy and data protection · 28%
Artificial intelligence
4 papers
Language models and text generation · 36% Trustworthy machine learning · 36% Efficient and distributed learning · 29%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
model stealing
3.342026
Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks · AAAI 2026
Unlocking High-Fidelity Learning: Towards Neuron-Grained Model Extraction · IEEE Trans. Dependable Secur. Comput. 2025
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation · ACL (1) 2025
Security and privacy of machine learning › model intellectual property protection
model watermarking
1.012026
Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks · AAAI 2026
Natural language and speech › Language models and text generation
alignment
0.912025
Exploring Intrinsic Alignments Within Text Corpus · AAAI 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Exploring Intrinsic Alignments Within Text Corpus · AAAI 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation · ACL (1) 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Exploring Intrinsic Alignments Within Text Corpus · AAAI 2025
Security and privacy of machine learning › adversarial attack
backdoor attack
0.912025
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks? · ICML 2025
Privacy and data protection
differential privacy
0.912025
Dual Utilization of Perturbation for Stream Data Publication Under Local Differential Privacy · ICDE 2025
Privacy and data protection › differential privacy
local differential privacy
0.912025
Dual Utilization of Perturbation for Stream Data Publication Under Local Differential Privacy · ICDE 2025
Security and privacy of machine learning
machine unlearning
0.912025
Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for Privacy · ICCV 2025
Security and privacy of machine learning › privacy attack
model inversion attack
0.912025
A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks · AAAI 2025
Security and privacy of machine learning › privacy attack › model inversion attack
model inversion defense
0.912025
A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks · AAAI 2025
Security and privacy of machine learning
poisoning attack
0.912025
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks? · ICML 2025
Security and privacy of machine learning
privacy attack
0.912025
Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for Privacy · ICCV 2025
Privacy and data protection › differential privacy › privacy accounting
privacy budget allocation
0.912025
Dual Utilization of Perturbation for Stream Data Publication Under Local Differential Privacy · ICDE 2025
Privacy and data protection › differential privacy › continual release
streaming data publication
0.912025
Dual Utilization of Perturbation for Stream Data Publication Under Local Differential Privacy · ICDE 2025
Privacy and data protection › privacy-preserving machine learning
training data privacy
0.912025
A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks · AAAI 2025
Security and privacy of machine learning
training-time attack
0.912025
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks? · ICML 2025
Security and privacy of machine learning
membership inference
0.612022
MExMI: Pool-based Active Model Extraction Crossover Membership Inference · NeurIPS 2022
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.312025
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks? · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.312025
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks? · ICML 2025
Security and privacy of machine learning › model stealing
model stealing defense
0.312025
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7knowledge distillation · 1.7information theory · 1.7decision boundary analysis · 1.0class-feature watermark · 1.0adversarial attack · 1.0transfer learning · 0.9stochastic norm enlargement · 0.9sampling strategy · 0.9reinforcement learning with human feedback · 0.9neuron matching theory · 0.9neural tangent kernel · 0.9natural gradient descent · 0.9iterative perturbation parameterization · 0.9fine-tuning · 0.9entropy loss · 0.9calibration · 0.9approximate machine unlearning · 0.9
YearPublicationVenuePosition
2026 Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks
abstract
Machine learning models constitute valuable intellectual property, yet remain vulnerable to model extraction attacks (MEA), where adversaries replicate their functionality through black-box queries. Model watermarking counters MEAs by embedding forensic markers for ownership verification. Current black-box watermarks prioritize MEA survival through representation entanglement, yet inadequately explore resilience against sequential MEAs and removal attacks. Our study reveals that this risk is underestimated because existing removal methods are weakened by entanglement. To address this gap, we propose Watermark Removal attacK (WRK), which circumvents entanglement constraints by exploiting decision boundaries shaped by prevailing sample-level watermark artifacts. WRK effectively reduces watermark success rates by ≥88.79% across existing watermarking benchmarks. For robust protection, we propose Class-Feature Watermarks (CFW), which improve resilience by leveraging class-level artifacts. CFW constructs a synthetic class using out-of-domain samples, eliminating vulnerable decision boundaries between original domain samples and their artifact-modified counterparts (watermark samples). CFW concurrently optimizes both MEA transferability and post-MEA stability. Experiments across multiple domains show that CFW consistently outperforms prior methods in resilience, maintaining a watermark success rate of ≥70.15% in extracted models even under the combined MEA and WRK distortion, while preserving the utility of protected models.
Yaxin Xiao, Qingqing Ye 0001, Zi Liang, Haoyang Li 0018, Ronghua Li 0002, Huadi Zheng, Haibo Hu 0001
AAAI1
2026 Time-varying thresholds with Gaussian smoothing for one-bit DOA estimation in unequal power signals
Anqi Yan, Yaxin Xiao, Chuangrui Meng
Signal Process.3
2025 A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks
abstract
Model Inversion (MI) attacks, which reconstruct the training dataset of neural networks, pose significant privacy concerns in machine learning. Recent MI attacks have managed to reconstruct realistic label-level private data, such as the general appearance of a target person from all training images labeled on him. Beyond label-level privacy, in this paper we show sample-level privacy, the private information of a single target sample, is also important but under-explored in the MI literature due to the limitations of existing evaluation metrics. To address this gap, this study introduces a novel metric tailored for training-sample analysis, namely, the Diversity and Distance Composite Score (DDCS), which evaluates the reconstruction fidelity of each training sample by encompassing various MI attack attributes. This, in turn, enhances the precision of sample-level privacy assessments. Leveraging DDCS as a new evaluative lens, we observe that many training samples remain resilient against even the most advanced MI attack. As such, we further propose a transfer learning framework that augments the generative capabilities of MI attackers through the integration of entropy loss and natural gradient descent. Extensive experiments verify the effectiveness of our framework on improving state-of-the-art MI attacks over various metrics including DDCS, coverage and FID. Finally, we demonstrate that DDCS can also be useful for MI defense, by identifying samples susceptible to MI attacks in an unsupervised manner.
Haoyang Li 0018, Li Bai 0004, Qingqing Ye 0001, Haibo Hu 0001, Yaxin Xiao, Huadi Zheng, Jianliang Xu
AAAI5
2025 Exploring Intrinsic Alignments Within Text Corpus
abstract
Recent years have witnessed rapid advancements in the safety alignments of large language models (LLMs). Methods such as supervised instruction fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) have thus emerged as vital components in constructing LLMs. While these methods achieve robust and fine-grained alignment to human values, their practical application is still hindered by high annotation costs and incomplete human alignments. Besides, the intrinsic human values within training corpora have not been fully exploited. To address these issues, we propose ISAAC (Intrinsically Supervised Alignments by Assessing Corpus), a primary and coarse-grained safety alignment strategy for LLMs. ISAAC only relies on a prior assumption about the text corpus, and does not require preferences in RLHF or human responses selection in SFT. Specifically, it assumes a long-tail distribution of text corpus and employs a specialized sampling strategy to automatically sample high-quality responses. Theoretically, we prove that this strategy can improve the safety of LLMs under our assumptions. Empirically, our evaluations on mainstream LLMs show that ISAAC achieves a safety score comparable to current SFT solutions. Moreover, we conduct experiments on ISAAC for some RLHF-based LLMs, where we find that ISAAC can even improve the safety of these models under specific safety domains. These findings demonstrate that ISAAC can provide preliminary alignment to LLMs, thereby reducing the construction costs of existing human-feedback-based methods.
Zi Liang, Pinghui Wang, Ruofei Zhang, Haibo Hu 0001, Qingqing Ye 0001, Nuo Xu 0012, Yaxin Xiao
AAAI8
2025 "Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation
abstract
Zi Liang, Qingqing Ye, Yanyun Wang, Sen Zhang, Yaxin Xiao, RongHua Li, Jianliang Xu, Haibo Hu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zi Liang, Qingqing Ye 0001, Yanyun Wang 0003, Sen Zhang 0002, Yaxin Xiao, Ronghua Li 0002, Jianliang Xu, Haibo Hu 0001
ACL (1)5
2025 Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for Privacy
Yaxin Xiao, Qingqing Ye 0001, Huadi Zheng, Haibo Hu 0001, Zi Liang, Haoyang Li 0018, Yijie Jiao
ICCV1
2025 Dual Utilization of Perturbation for Stream Data Publication Under Local Differential Privacy
abstract
Stream data from real-time distributed systems such as IoT, tele-health, and crowdsourcing has become an important data source. However, the collection and analysis of usergenerated stream data raise privacy concerns due to the potential exposure of sensitive information. To address these concerns, local differential privacy (LDP) has emerged as a promising standard. Nevertheless, applying LDP to stream data presents significant challenges, as stream data often involves a large or even infinite number of values. Allocating a given privacy budget across these data points would introduce overwhelming LDP noise to the original stream data. Beyond existing approaches that merely use perturbed values for estimating statistics, our design leverages them for both perturbation and estimation. This dual utilization arises from a key observation: each user knows their own ground truth and perturbed values, enabling a precise computation of the deviation error caused by perturbation. By incorporating this deviation into the perturbation process of subsequent values, the previous noise can be calibrated. Following this insight, we introduce the Iterative Perturbation Parameterization (IPP) method, which utilizes current perturbed results to calibrate the subsequent perturbation process. To enhance the robustness of calibration and reduce sensitivity, two algorithms, namely Accumulated Perturbation Parameterization (APP) and Clipped Accumulated Perturbation Parameterization (CAPP) are further developed. We prove that these three algorithms satisfy$w$-event differential privacy while significantly improving utility. Experimental results demonstrate that our techniques outperform state-of-the-art LDP stream publishing solutions in terms of utility, while retaining the same privacy guarantee.
Rong Du 0001, Qingqing Ye 0001, Yaxin Xiao, Liantong Yu, Haibo Hu 0001
ICDE3
2025 Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
abstract
Low rank adaptation (LoRA) has emerged as a prominent technique for fine-tuning large language models (LLMs) thanks to its superb efficiency gains over previous methods. While extensive studies have examined the performance and structural properties of LoRA, its behavior upon training-time attacks remain underexplored, posing significant security risks. In this paper, we theoretically investigate the security implications of LoRA's low-rank structure during fine-tuning, in the context of its robustness against data poisoning and backdoor attacks. We propose an analytical framework that models LoRA’s training dynamics, employs the neural tangent kernel to simplify the analysis of the training process, and applies information theory to establish connections between LoRA's low rank structure and its vulnerability against training-time attacks. Our analysis indicates that LoRA exhibits better robustness to backdoor attacks than full fine-tuning, while becomes more vulnerable to untargeted data poisoning due to its over-simplified information geometry. Extensive experimental evaluations have corroborated our theoretical findings.
Zi Liang, Haibo Hu 0001, Qingqing Ye 0001, Yaxin Xiao, Ronghua Li 0002
ICML4
2025 Unlocking High-Fidelity Learning: Towards Neuron-Grained Model Extraction
abstract
Model extraction (ME) attacks replicate valuable black-box machine learning (ML) models via malicious query interactions. Cutting-edge attacks focus on actively designing query samples to enhance model fidelity and imprudently adhere to the standard ML training approach. This causes a deviation from the true objective of learning a model over a task. In this paper, we innovatively shift our focus from query selection to training process optimization, aiming to boost the similarity of the copy model with the victim model from neuron to model level. We leverage neuron matching theory to attain this objective and develop a general training booster framework, MEBooster, to fully exploit this theory. MEBooster comprises an initial bootstrapping phase that furnishes initial parameters and an optimal model architecture, followed by a post-processing phase that employs fine-tuning for enhanced neuron matching. Notably, MEBooster can seamlessly integrate with all existing model extraction attacks, enhancing their overall performance. Performance evaluation shows up to 58.10% fidelity gain in image classification. From a defender's perspective, we introduce a novel defensive strategy calledStochastic Norm Enlargement(SNE) to mitigate the risk of such attacks by enlarging the model parameters' norm property in training. Performance evaluation shows up to 58.81% extractability (i.e., fidelity) reduction.
Yaxin Xiao, Haibo Hu 0001, Qingqing Ye 0001, Zi Liang, Huadi Zheng
IEEE Trans. Dependable Secur. Comput.1
2024 DeepMark: A Scalable and Robust Framework for DeepFake Video Detection
abstract
With the rapid growth of DeepFake video techniques, it becomes increasingly challenging to identify them visually, posing a huge threat to our society. Unfortunately, existing detection schemes are limited to exploiting the artifacts left by DeepFake manipulations, so they struggle to keep pace with the ever-improving DeepFake models. In this work, we propose DeepMark, a scalable and robust framework for detecting DeepFakes. It imprints essential visual features of a video into DeepMark Meta (DMM) and uses it to detect DeepFake manipulations by comparing the extracted visual features with the ground truth in DMM. Therefore, DeepMark is future-proof, because a DeepFake video must aim to alter some visual feature, no matter how “natural” it looks. Furthermore, DMM also contains a signature for verifying the integrity of the above features. And an essential link to the features as well as their signature is attached with error correction codes and embedded in the video watermark. To improve the efficiency of DMM creation, we also present a threshold-based feature selection scheme and a deduced face detection scheme. Experimental results demonstrate the effectiveness and efficiency of DeepMark on DeepFake video detection under various datasets and parameter settings.
Qingqing Ye 0001, Haibo Hu 0001, Qiao Xue, Yaxin Xiao, Jin Li 0002
ACM Trans. Priv. Secur.5
2022 MExMI: Pool-based Active Model Extraction Crossover Membership Inference
abstract
With increasing popularity of Machine Learning as a Service (MLaaS), ML models trained from public and proprietary data are deployed in the cloud and deliver prediction services to users. However, as the prediction API becomes a new attack surface, growing concerns have arisen on the confidentiality of ML models. Existing literatures show their vulnerability under model extraction (ME) attacks, while their private training data is vulnerable to another type of attacks, namely, membership inference (MI). In this paper, we show that ME and MI can reinforce each other through a chained and iterative reaction, which can significantly boost ME attack accuracy and improve MI by saving the query cost. As such, we build a framework MExMI for pool-based active model extraction (PAME) to exploit MI through three modules: “MI Pre-Filter”, “MI Post-Filter”, and “semi-supervised boosting”. Experimental results show that MExMI can improve up to 11.14% from the best known PAME attack and reach 94.07% fidelity with only 16k queries. Furthermore, the precision and recall of the MI attack in MExMI are on par with state-of-the-art MI attack which needs 150k queries.
Yaxin Xiao, Qingqing Ye 0001, Haibo Hu 0001, Huadi Zheng, Chengfang Fang, Jie Shi 0005
NeurIPS1