Zhixiong Zhuang

dblp:377/9411 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
3 papers
Security and privacy of machine learning · 100%
Artificial intelligence
2 papers
Generative modeling · 70% Reinforcement learning · 30%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
model stealing
2.532025
Stealix: Model Stealing via Prompt Evolution · ICML 2025
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment · AAAI 2025
Stealthy Imitation: Reward-guided Environment-free Policy Stealing · ICML 2024
Machine learning › Generative modeling
diffusion model
0.912025
Stealix: Model Stealing via Prompt Evolution · ICML 2025
Machine learning › Generative modeling
synthetic data generation
0.912025
Stealix: Model Stealing via Prompt Evolution · ICML 2025
Security and privacy of machine learning
adversarial attack
0.912025
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment · AAAI 2025
Security and privacy of machine learning › model stealing
data-free model stealing
0.912025
Stealix: Model Stealing via Prompt Evolution · ICML 2025
Medical and health informatics › medical report generation
radiology report generation
0.312025
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment · AAAI 2025

Methods — techniques the papers use, named apart from their topics

prompt evolution · 1.7genetic algorithm · 1.7data augmentation with adversarial noise · 1.7adversarial domain alignment · 1.7reward model fitting · 1.5black-box querying · 1.5pre-trained diffusion models · 0.9pre-trained diffusion model · 0.9
YearPublicationVenuePosition
2025 Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
abstract
Medical multimodal large language models (MLLMs) are becoming an instrumental part of healthcare systems, assisting medical personnel with decision making and results analysis. Models for radiology report generation are able to interpret medical imagery, thus reducing the workload of radiologists. As medical data is scarce and protected by privacy regulations, medical MLLMs represent valuable intellectual property. However, these assets are potentially vulnerable to model stealing, where attackers aim to replicate their functionality via black-box access. So far, model stealing for the medical domain has focused on image classification; however, existing attacks are not effective against MLLMs. In this paper, we introduce Adversarial Domain Alignment (ADA-Steal), the first stealing attack against medical MLLMs. ADA-Steal relies on natural images, which are public and widely available, as opposed to their medical counterparts. We show that data augmentation with adversarial noise is sufficient to overcome the data distribution gap between natural images and the domain-specific distribution of the victim MLLM. Experiments on the IU X-RAY and MIMIC-CXR radiology datasets demonstrate that Adversarial Domain Alignment enables attackers to steal the medical MLLM without any access to medical data.
Yaling Shen, Zhixiong Zhuang, Kun Yuan 0004, Maria-Irina Nicolae, Nassir Navab, Nicolas Padoy, Mario Fritz
AAAI2
2025 Stealix: Model Stealing via Prompt Evolution
abstract
Model stealing poses a significant security risk in machine learning by enabling attackers to replicate a black-box model without access to its training data, thus jeopardizing intellectual property and exposing sensitive information. Recent methods that use pre-trained diffusion models for data synthesis improve efficiency and performance but rely heavily on manually crafted prompts, limiting automation and scalability, especially for attackers with little expertise. To assess the risks posed by open-source pre-trained models, we propose a more realistic threat model that eliminates the need for prompt design skills or knowledge of class names. In this context, we introduce Stealix, the first approach to perform model stealing without predefined prompts. Stealix uses two open-source pre-trained models to infer the victim model’s data distribution, and iteratively refines prompts through a genetic algorithm, progressively improving the precision and diversity of synthetic images. Our experimental results demonstrate that Stealix significantly outperforms other methods, even those with access to class names or fine-grained prompts, while operating under the same query budget. These findings highlight the scalability of our approach and suggest that the risks posed by pre-trained generative models in model stealing may be greater than previously recognized.
Zhixiong Zhuang, Hui-Po Wang, Maria-Irina Nicolae, Mario Fritz
ICML1
2024 Stealthy Imitation: Reward-guided Environment-free Policy Stealing
abstract
Deep reinforcement learning policies, which are integral to modern control systems, represent valuable intellectual property. The development of these policies demands considerable resources, such as domain expertise, simulation fidelity, and real-world validation. These policies are potentially vulnerable to model stealing attacks, which aim to replicate their functionality using only black-box access. In this paper, we propose Stealthy Imitation, the first attack designed to steal policies without access to the environment or knowledge of the input range. This setup has not been considered by previous model stealing methods. Lacking access to the victim's input states distribution, Stealthy Imitation fits a reward model that allows to approximate it. We show that the victim policy is harder to imitate when the distribution of the attack queries matches that of the victim. We evaluate our approach across diverse, high-dimensional control tasks and consistently outperform prior data-free approaches adapted for policy stealing. Lastly, we propose a countermeasure that significantly diminishes the effectiveness of the attack.
Zhixiong Zhuang, Maria-Irina Nicolae, Mario Fritz
ICML1