Kanghua Mo

dblp:278/1999 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-3762-674XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TransLock: Securing LLM deployment for software applications via self-locking watermarks
Zhenxin Zhang, Kanghua Mo
Empir. Softw. Eng.3
2026 HPA: Manipulating deep reinforcement learning via adversarial interaction
Kanghua Mo, Yucheng Long, Zhengdao Li
J. Syst. Archit.1
2026 Your Non-Transferable Learning is Fragile: Practical Breach of Protected Models
abstract
Non-transferable learning (NTL) has emerged as a promising method to protect the intellectual property of deep learning models by restricting cross-domain knowledge transfer. However, the robustness of its transferability constraints against potential attacks has not been explored, especially in practical deployment scenarios. In this paper, we propose a novel black-box attack framework - distribution drift learner (DDL), which effectively bypasses NTL protection mechanisms by only accessing input-output queries of protected models. The theoretical foundation of DDL is derived from the concept of data drift, which takes advantage of the variability of the statistical distribution between the source and target domains. The core innovation of DDL is the integration of distributed perception regularization into a lightweight autoencoder architecture, enabling efficient manipulation of data distribution by optimizing dual objectives (distributed perception loss and reconstruction loss). Training for DDL involves two key steps: First, DDL reconstructs a moderate amount of target domain samples and feeds the reconstructed images into the NTL model to obtain prediction labels. The DDL parameters are then updated by optimizing distributed perception loss and reconstruction loss. Through extensive experiments against standard NTL benchmarks (Digits, CIFAR10, and STL10), we demonstrate that DDL has successfully overcome the barriers of the transferable NTL model and improved the accuracy of the target domain by 81% from 10%. Our work reveals critical vulnerabilities in the NTL framework, particularly with respect to ownership verification and applicability authorization mechanisms, providing valuable insights for developing more robust model protection strategies in real-world applications.
Anli Yan, Huali Ren, Kanghua Mo, Zhenxin Zhang, Hongyang Yan, Jin Li 0002
IEEE Trans. Inf. Forensics Secur.3
2025 Distraction is All You Need for Multimodal Large Language Model Jailbreaking
abstract
Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechanisms. To address this, we analyze the relationship between image content and task and find that the complexity of subimages, rather than their content, is key. Building on this insight, we propose the Distraction Hypothesis, followed by a novel framework called Contrasting Subimage Distraction Jailbreaking (CS-DJ), to achieve jailbreaking by disrupting MLLMs alignment through multi-level distraction strategies. CS-DJ consists of two components: structured distraction, achieved through query decomposition that induces a distributional shift by fragmenting harmful prompts into sub-queries, and visual-enhanced distraction, realized by constructing contrasting subimages to disrupt the interactions among visual elements within the model. This dual strategy disperses the model’s attention, reducing its ability to detect and mitigate harmful content. Extensive experiments across five representative scenarios and four popular closed-source MLLMs, including GPT-4o-mini, GPT-4o, GPT-4V, and Gemini-1.5-Flash, demonstrate that CS-DJ achieves average success rates of 52.40% for the attack success rate and 74.10% for the ensemble attack success rate. These results reveal the potential of distraction-based approaches to exploit and bypass MLLMs’ defenses, offering new insights for attack strategies. Our code is available at https://github.com/TeamPigeonLab/CS-DJ.Warning: This paper contains unfiltered content generated by MLLMs that may be offensive to readers
Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Kanghua Mo, Changyu Dong
CVPR7
2025 Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
abstract
Large language model (LLM) agents have demonstrated remarkable capabilities in complex reasoning and decision-making by leveraging external tools. However, this tool-centric paradigm introduces a previously underexplored attack surface, where adversaries can manipulate tool metadata---such as names, descriptions, and parameter schemas---to influence agent behavior. We identify this as a new and stealthy threat surface that allows malicious tools to be preferentially selected by LLM agents, without requiring prompt injection or access to model internals. To demonstrate and exploit this vulnerability, we propose the Attractive Metadata Attack (AMA), a black-box in-context learning framework that generates highly attractive but syntactically and semantically valid tool metadata through iterative optimization. The proposed attack integrates seamlessly into standard tool ecosystems and requires no modification to the agent’s execution framework. Extensive experiments across ten realistic, simulated tool-use scenarios and a range of popular LLM agents demonstrate consistently high attack success rates (81\%-95\%) and significant privacy leakage, with negligible impact on primary task execution. Moreover, the attack remains effective even against prompt-level defenses, auditor-based detection, and structured tool-selection protocols such as the Model Context Protocol, revealing systemic vulnerabilities in current agent architectures. These findings reveal that metadata manipulation constitutes a potent and stealthy attack surface. Notably, AMA is orthogonal to injection attacks and can be combined with them to achieve stronger attack efficacy, highlighting the need for execution-level defenses beyond prompt-level and auditor-based mechanisms. Code is available at \url{https://github.com/SEAIC-M/AMA}.
Kanghua Mo, Yucheng Long
NeurIPS1
2025 Enhancing Model Intellectual Property Protection With Robustness Fingerprint Technology
abstract
Deep neural network (DNN) models embody the intellectual property of a model owner, as the process of training the DNN model is a complex and resource-intensive task that requires significant investments in data preparation and computing resources. Numerous efforts have been made to protect the intellectual property of DNN models. However, existing methods often come with a critical limitation: they lack robustness, proving effective only in specific intellectual property threat scenarios or they either sacrifice the utility/accuracy of the model owner’s classifier because it interferes with the classifier’s training. To address these issues, we propose GMFIP, a novel generator-based model fingerprinting technology tailored for DNN intellectual property protection. GMFIP stands out for its robustness, extending its utility to various intellectual property threat scenarios rather than specific ones. Furthermore, GMFIP ensures that the utility/accuracy of the model is not affected by protection measures. Specifically, GMFIP begins with the training of the generator, which lays the groundwork for the model fingerprint. The generator generates fingerprints of the unique properties of the source model for verifying model ownership. To further improve the quality of these fingerprints, an extra selection phase dedicated to refining the fingerprints is integrated. Moreover, GMFIP is complemented by a binary classifier, which adapts the threshold setting to get optimal results. Our empirical evaluation includes an ablation study over four state-of-the-art technologies and three image benchmark datasets. Our results demonstrate that GMFIP outperforms other state-of-the-art technologies in effectively distinguishing pirated models from benign models.
Anli Yan, Huali Ren, Kanghua Mo, Zhenxin Zhang, Shaowei Wang 0003, Jin Li 0002
IEEE Trans. Inf. Forensics Secur.3
2025 Inverse correction-optimized vertical federated unlearning
Kanghua Mo, Ziyu Ding, Hongyang Yan, Jin Li 0002
J. Supercomput.3
2024 Exploring the vulnerability of self-supervised monocular depth estimation models
Ruitao Hou, Kanghua Mo, Yucheng Long, Ning Li 0050, Yuan Rao 0002
Inf. Sci.2
2023 Empirical study of privacy inference attack against deep reinforcement learning models
abstract
Most studies on privacy in machine learning have primarily focused on supervised learning, with little research on privacy concerns in reinforcement learning. However, our study has demonstrated that observation information can be extracted through trajectory analysis. In this paper, we propose a variable information inference attack targeting the observation space of policy models, which is categorised into two types: observed value inference attack and observed variable inference. Our algorithm has demonstrated a high success rate in privacy inference attacks for both types of observation information.
Huaicheng Zhou, Kanghua Mo, Yongjin Li
Connect. Sci.2
2023 Attacking Deep Reinforcement Learning With Decoupled Adversarial Policy
abstract
While Deep Reinforcement Learning (DRL) has achieved outstanding performance in extensive applications, exploiting its vulnerability with adversarial attacks is essential towards building robust DRL systems. In this work, we aim to propose a novel Decoupled Adversarial Policy (DAP) for attacking the DRL mechanism, whereas the adversarial agent can decompose the adversarial policy into two separate sub-policies: 1) the switch policy which determines if an attacker should launch the attack, and 2) the lure policy which determines the action an attacker induces the victim to take. If the adversarial agent samples an injection action from the switch policy, the attacker can query the pre-constructed database for universal perturbation in the real-time manner, misleading the victim to take the induced action sampled from the lure policy. To train the adversarial agent to learn DAP, we utilize those samples wherein both of the sub-actions from DAP are not restricted by each other or by the external constraint, but can actually affect the attacker’s behaviors. Therefore, we propose trajectory clipping and padding in data pruning, and Decoupled Proximal Policy Optimization (DPPO) in optimizing. Extensive experiments on different Atari games demonstrate the effectiveness of our proposed method. In addition, it can simultaneously implement the real-time and few-steps attack, which outperforms the existing counterparts.
Kanghua Mo, Weixuan Tang 0004, Jin Li 0002, Xu Yuan 0001
IEEE Trans. Dependable Secur. Comput.1
2022 ESM: Selfish mining under ecological footprint
Shan Ai, Guoyu Yang, Chang Chen 0003, Kanghua Mo, Wangyong Lv, Arthur Sandor Voundi Koe
Inf. Sci.4
2022 Sender anonymity: Applying ring signature in gateway-based blockchain for IoT is not enough
Arthur Sandor Voundi Koe, Shan Ai, Anli Yan, Qi Chen 0024, Kanghua Mo, Wanqing Jie, Shiwen Zhang 0004
Inf. Sci.7
2021 Querying little is enough: Model inversion attack via latent information
abstract
As machine learning (ML) technologies evolve, various online intelligent services use ML models to provide predictions. Unfortunately, attackers can obtain the private information of the model by interacting with the online service, namely model inversion attack (MIA). However, MIA requires large data sets to be transferred to an online service to obtain the predictive value of the inference model. Besides, the huge transmission may cause the administrator's active defense. To overcome this drawback, we propose a novel MIA scheme, which leverages latent information extracted by an auxiliary neural network as high-dimensional features to simplify what inversion model should learn. The core idea of our scheme is to reuse some parameters of the local pretraining model. Extensive experiments have verified the effectiveness of our method in convolutional neural networks on LFW, pubFig, MNIST data sets. Experimental results show that even with a few queries, our inversion method still work accurately and is superior to other technologies. It is worth mentioning that our method makes it more difficult for administrators to defend against the attack and elicit more investigations for privacy-preserving.
Kanghua Mo, Xiaozhang Liu, Teng Huang 0001, Anli Yan
Int. J. Intell. Syst.1