Haolin Yuan

dblp:238/9094 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2024
0009-0007-8385-5221ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 PLeak: Prompt Leaking Attacks against Large Language Model Applications
abstract
Large Language Models (LLMs) enable a new ecosystem with many downstream applications, called LLM applications, with different natural language processing tasks. The functionality and performance of an LLM application highly depend on its system prompt, which instructs the backend LLM on what task to perform. Therefore, an LLM application developer often keeps a system prompt confidential to protect its intellectual property. As a result, a natural attack, called prompt leaking, is to steal the system prompt from an LLM application, which compromises the developer's intellectual property. Existing prompt leaking attacks primarily rely on manually crafted queries, and thus achieve limited effectiveness.
Bo Hui 0002, Haolin Yuan, Neil Zhenqiang Gong, Philippe Burlina, Yinzhi Cao
CCS2
2024 PFEDEDIT: Personalized Federated Learning via Automated Model Editing
Haolin Yuan, William Paul, John N. Aucott, Philippe Burlina, Yinzhi Cao
ECCV (79)1
2024 SneakyPrompt: Jailbreaking Text-to-image Generative Models
abstract
Text-to-image generative models such as Stable Diffusion and DALL•E raise many ethical concerns due to the generation of harmful images such as Not-Safe-for-Work (NSFW) ones. To address these ethical concerns, safety filters are often adopted to prevent the generation of NSFW images. In this work, we propose SneakyPrompt, the first automated attack framework, to jailbreak text-to-image generative models such that they generate NSFW images even if safety filters are adopted. Given a prompt that is blocked by a safety filter, SneakyPrompt repeatedly queries the text-to-image generative model and strategically perturbs tokens in the prompt based on the query results to bypass the safety filter. Specifically, SneakyPrompt utilizes reinforcement learning to guide the perturbation of tokens. Our evaluation shows that SneakyPrompt successfully jailbreaks DALL•E 2 with closed-box safety filters to generate NSFW images. Moreover, we also deploy several state-of-the-art, open-source safety filters on a Stable Diffusion model. Our evaluation shows that SneakyPrompt not only successfully generates NSFW images, but also outperforms existing text adversarial attacks when extended to jailbreak text-to-image generative models, in terms of both the number of queries and qualities of the generated NSFW images. SneakyPrompt is open-source and available at this repository: https://github.com/Yuchen413/text2image_safety.
Yuchen Yang 0001, Bo Hui 0002, Haolin Yuan, Neil Zhenqiang Gong, Yinzhi Cao
SP3
2023 Fortifying Federated Learning against Membership Inference Attacks via Client-level Input Perturbation
abstract
Membership inference (MI) attacks are more diverse in a Federated Learning (FL) setting, because an adversary may be either an FL client, a server, or an external attacker. Existing defenses against MI attacks rely on perturbations to either the model's output predictions or the training process. However, output perturbations are ineffective in an FL setting, because a malicious server can access the model without output perturbation while training perturbations struggle to achieve a good utility. This paper proposes a novel defense, called CIP, to fortify FL against MI attacks via a client-level input perturbation during training and inference procedures. The key insight is to shift each client's local data distribution via a personalized perturbation to get a shifted model. CIP achieves a good balance between privacy and utility. Our evaluation shows that CIP causes accuracy to drop at most 0.7% while reducing attacks to random guessing.
Yuchen Yang 0001, Haolin Yuan, Bo Hui 0002, Neil Zhenqiang Gong, Neil Fendley, Philippe Burlina, Yinzhi Cao
DSN2
2023 EdgeMixup: Embarrassingly Simple Data Alteration to Improve Lyme Disease Lesion Segmentation and Diagnosis Fairness
Haolin Yuan, John N. Aucott, Armin Hadzic, William Paul, Marcia Villegas de Flores, Philip Mathew, Philippe Burlina, Yinzhi Cao
MICCAI (4)1
2023 ImageAlly: A Human-AI Hybrid Approach to Support Blind People in Detecting and Redacting Private Image Content
Zhuohao (Jerry) Zhang, Smirity Kaushik, Jooyoung Seo, Haolin Yuan, Sauvik Das, Leah Findlater, Danna Gurari, Abigale Stangl, Yang Wang 0005
SOUPS4
2023 PrivateFL: Accurate, Differentially Private Federated Learning via Personalized Data Transformation
Yuchen Yang 0001, Bo Hui 0002, Haolin Yuan, Neil Zhenqiang Gong, Yinzhi Cao
USENIX Security Symposium3
2022 Addressing Heterogeneity in Federated Learning via Distributional Transformation
Haolin Yuan, Bo Hui 0002, Yuchen Yang 0001, Philippe Burlina, Neil Zhenqiang Gong, Yinzhi Cao
ECCV (38)1
2021 Practical Blind Membership Inference Attack via Differential Comparisons
Bo Hui 0002, Yuchen Yang 0001, Haolin Yuan, Philippe Burlina, Neil Zhenqiang Gong, Yinzhi Cao
NDSS3