Leo Hyun Park

dblp:242/7270 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-3100-2258ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Amplifying Training Data Exposure Through Fine-Tuning With Pseudo-Labeled Memberships
abstract
Large language models (LLMs) are vulnerable to training data extraction attacks due to data memorization. This paper introduces a novel attack scenario wherein an attacker adversarially fine-tunes pre-trained LLMs to amplify the exposure of the original training data. Unlike prior TDE methods that mainly rely on post-hoc querying or prompt selection to elicit memorized content from a fixed model, our strategy directly alters the model’s parameters to intensify its retention of the pre-training dataset. To achieve this, the attacker needs to collect generated texts that are closely aligned with the pre-training data. However, without knowledge of the actual dataset, quantifying the amount of pre-training data within generated texts is challenging. To address this, we propose the use of pseudo-labels for these generated texts, leveraging membership approximations indicated by machine-generated probabilities from the target LLMusing DetectGPT. We subsequently fine-tune the LLM via reinforcement learning from human feedback (RLHF) to favor generations with higher likelihoods of originating from the pre-training data, based on these membership probabilities. Our empirical findings indicate a remarkable outcome: LLMs with over 1B parameters exhibit a four to eight-fold increase in training data exposure. We discuss potential mitigations and suggest future research directions.
Myung Gyo Oh, Hong Eun Ahn, Leo Hyun Park, Taekyoung Kwon 0002
IEEE Trans. Inf. Forensics Secur.3
2025 Toward an Autonomous Purple Teaming Framework for Security and Safety in Large Language Models
abstract
Large Language Models (LLMs) have rapidly advanced in reasoning capability and accessibility, driving their deployment across diverse applications. Yet this progress has also widened the surface for safety and security vulnerabilities. Adversaries can exploit prompt diversity, dialog memory, or multimodal inputs to induce unsafe or confidential outputs, while continual fine-tuning and third-party integration render static assurance infeasible. This paper introduces our ongoing national R&D project on developing the AutoPT Framework-an Autonomous Purple Teaming architecture that extends the collaborative principles of purple teaming toward self-adaptive, continuously verifiable LLM assurance. AutoPT unifies autonomous adversarial exploration and adaptive defensive reinforcement through two co-evolving agents. The red module, AutoPT-Red, employs coverage-guided fuzzing and internal measurement metrics to autonomously uncover vulnerabilities. The blue module, AutoPT-Blue, performs self-healing adaptation by updating guardrails and detecting integrity or confidentiality violations using embedding-based feedback. Preliminary case studies on jailbreak fuzzing and backdoor-poisoning defense validate the feasibility of this closed-loop, self-adapting architecture. As part of a broader national initiative, this work lays the conceptual and technical foundation for transitioning industrial purple teaming into a fully autonomous, scalable, and measurable assurance paradigm for generative AI systems.
Leo Hyun Park, Yoonsik Kim, Eunbi Hwang, Sangsoo Han, Hyoungshick Kim, Taekyoung Kwon 0002
PRDC1
2025 Red-Teaming LLMs with Token Control Score: Efficient, Universal, and Transferable Jailbreaks
abstract
Large Language Models (LLMs) are vulnerable to jailbreak attacks, where adversaries craft malicious prompts to bypass safety mechanisms and elicit harmful, illegal, or unethical responses. To ensure robustness before deployment, red-teaming LLMs is essential. However, prior methods like GCG and AutoDAN rely on loss functions that only assess whether the output matches a fixed target, offering little insight into how well the prompt circumvents refusal behaviors. We introduce Token Control Score (TCS), a novel metric that quantifies how effectively a prompt steers the LLM toward compliance and away from refusal. TCS compares the logits of key tokens representing compliant and rejecting responses, and can be extended with gradient-based feedback into a Normalized Token Control Score (NTCS) to guide optimization. Using this metric, we propose LLM-CGF, a fuzzing-based framework that iteratively discovers effective jailbreak templates for prompt injection. LLM-CGF leverages NTCS as the fitness function to explore LLM behaviors and uncover vulnerabilities with high efficiency. Experiments show that LLM-CGF outperforms state-of-the-art methods such as LLM-Fuzzer, generating more universal, transferable, and query-efficient jailbreak prompts. These results demonstrate the utility of TCS and NTCS as new objectives for prompt injection, providing deeper insights into LLM safety and enabling more thorough red-teaming evaluations. Warning: This paper contains unfiltered content generated by LLMs that may be offensive to readers.
Leo Hyun Park, Taekyoung Kwon 0002
RAID1
2023 GradFuzz: Fuzzing deep neural networks with gradient vector coverage for adversarial examples
Leo Hyun Park, Soochang Chung, Jaeuk Kim, Taekyoung Kwon 0002
Neurocomputing1
2022 Poster: Adversarial Defense with Deep Learning Coverage on MagNet's Purification
abstract
MagNet is a defense method that adopts autoencoders to detect and purify adversarial examples. Although MagNet is robust against grey-box and black-box attacks, it is vulnerable to white-box attacks. Despite this prior knowledge, the fundamental reason for and mitigation of the vulnerability of MagNet have not been discussed. We suggest that the challenge of MagNet is the generalization of the data manifold. To explain this, in this work, we leverage deep learning coverage for the reformer of MagNet. We mutate training images through image transformation algorithms and then train the reformer using mutants with new coverage information. The selected mutants provide an interesting data manifold, that cannot be handled by the random noise of MagNet, to the reformer. In grey-box settings, our defense method classified adversarial examples for various perturbation sizes much more accurately than MagNet even with the same architecture. Based on the preliminary result of this work, we consider future work to identify whether the generalization power of deep learning coverage is effective for stronger adversaries and different architectures.
Leo Hyun Park, Jaewoo Park 0004, Soochang Chung, Jaeuk Kim, Myung Gyo Oh, Taekyoung Kwon 0002
CCS1
2022 Mixed and constrained input mutation for effective fuzzing of deep learning systems
Leo Hyun Park, Jaeuk Kim, Jaewoo Park 0004, Taekyoung Kwon 0002
Inf. Sci.1
2019 Poster: Effective Layers in Coverage Metrics for Deep Neural Networks
abstract
Deep neural networks (DNNs) gained in popularity as an effective machine learning algorithm, but their high complexity leads to the lack of model interpretability and difficulty in the verification of deep learning. Fuzzing, which is an automated software testing technique, is recently applied to DNNs as an effort to address these problems by following the trend of coverage-based fuzzing. However, new coverage metrics on DNNs may bring out the question of which layer to measure the coverage in DNNs. In this poster, we empirically evaluate the performance of existing coverage metrics. By the comparative analysis of experimental results, we compile the most effective layer for each of coverage metrics and discuss a future direction of DNN fuzzing.
Leo Hyun Park, Sangjin Oh, Jaeuk Kim, Soochang Chung, Taekyoung Kwon 0002
CCS1
2018 A Guided Approach to Behavioral Authentication
abstract
User's behavioral biometrics are promising as authentication factors in particular if accuracy is sufficiently guaranteed. They can be used to augment security in combination with other authentication factors. A gesture-based pattern lock system is a good example of such multi-factor authentication, using touch dynamics in a smartphone. However, touch dynamics can be significantly affected by a shape of gestures with regard to the performance and accuracy, and our concern is that user-chosen patterns are likely far from producing such a good shape of gestures. In this poster, we raise this problem and show our experimental study conducted in this regard. We investigate if there is a reproducible correlation between shape and accuracy and if we can derive effective attribute values for user guidance, based on the gesture-based pattern lock system. In more general, we discuss a guided approach to behavioral authentication.
Yeeun Ku, Leo Hyun Park, Sooyeon Shin, Taekyoung Kwon 0002
CCS2