Yuchen Yang 0001

dblp:06/7124-1 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CertPHash: Towards Certified Perceptual Hashing via Robust Training
Yuchen Yang 0001, Qichang Liu, Christopher Brix, Huan Zhang 0001, Yinzhi Cao
USENIX Security Symposium1
2024 Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
Yuchen Yang 0001, Kwonjoon Lee, Behzad Dariush, Yinzhi Cao, Shao-Yuan Lo
ECCV (81)1
2024 SneakyPrompt: Jailbreaking Text-to-image Generative Models
abstract
Text-to-image generative models such as Stable Diffusion and DALL•E raise many ethical concerns due to the generation of harmful images such as Not-Safe-for-Work (NSFW) ones. To address these ethical concerns, safety filters are often adopted to prevent the generation of NSFW images. In this work, we propose SneakyPrompt, the first automated attack framework, to jailbreak text-to-image generative models such that they generate NSFW images even if safety filters are adopted. Given a prompt that is blocked by a safety filter, SneakyPrompt repeatedly queries the text-to-image generative model and strategically perturbs tokens in the prompt based on the query results to bypass the safety filter. Specifically, SneakyPrompt utilizes reinforcement learning to guide the perturbation of tokens. Our evaluation shows that SneakyPrompt successfully jailbreaks DALL•E 2 with closed-box safety filters to generate NSFW images. Moreover, we also deploy several state-of-the-art, open-source safety filters on a Stable Diffusion model. Our evaluation shows that SneakyPrompt not only successfully generates NSFW images, but also outperforms existing text adversarial attacks when extended to jailbreak text-to-image generative models, in terms of both the number of queries and qualities of the generated NSFW images. SneakyPrompt is open-source and available at this repository: https://github.com/Yuchen413/text2image_safety.
Yuchen Yang 0001, Bo Hui 0002, Haolin Yuan, Neil Zhenqiang Gong, Yinzhi Cao
SP1
2023 Fortifying Federated Learning against Membership Inference Attacks via Client-level Input Perturbation
abstract
Membership inference (MI) attacks are more diverse in a Federated Learning (FL) setting, because an adversary may be either an FL client, a server, or an external attacker. Existing defenses against MI attacks rely on perturbations to either the model's output predictions or the training process. However, output perturbations are ineffective in an FL setting, because a malicious server can access the model without output perturbation while training perturbations struggle to achieve a good utility. This paper proposes a novel defense, called CIP, to fortify FL against MI attacks via a client-level input perturbation during training and inference procedures. The key insight is to shift each client's local data distribution via a personalized perturbation to get a shifted model. CIP achieves a good balance between privacy and utility. Our evaluation shows that CIP causes accuracy to drop at most 0.7% while reducing attacks to random guessing.
Yuchen Yang 0001, Haolin Yuan, Bo Hui 0002, Neil Zhenqiang Gong, Neil Fendley, Philippe Burlina, Yinzhi Cao
DSN1
2023 PrivateFL: Accurate, Differentially Private Federated Learning via Personalized Data Transformation
Yuchen Yang 0001, Bo Hui 0002, Haolin Yuan, Neil Zhenqiang Gong, Yinzhi Cao
USENIX Security Symposium1
2022 Addressing Heterogeneity in Federated Learning via Distributional Transformation
Haolin Yuan, Bo Hui 0002, Yuchen Yang 0001, Philippe Burlina, Neil Zhenqiang Gong, Yinzhi Cao
ECCV (38)3
2021 Practical Blind Membership Inference Attack via Differential Comparisons
Bo Hui 0002, Yuchen Yang 0001, Haolin Yuan, Philippe Burlina, Neil Zhenqiang Gong, Yinzhi Cao
NDSS2