EDBT 2026 Demo / reviewers in the wild / expert
Minkyoo Song
dblp:325/3084
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0004-1597-2053ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PassREfinder-FL: Privacy-preserving credential stuffing risk prediction via graph-based federated learning for representing password reuse between websites
Jaehan Kim, Minkyoo Song, Minjae Seo, Youngjin Jin, Seungwon Shin 0001, Jinwoo Kim 0006 |
Expert Syst. Appl. | 2 |
| 2025 | MOEVIL: Poisoning Experts to Compromise the Safety of Mixture-of-Experts LLMsabstractMixture-of-Experts (MoE) has emerged as a prominent architecture for scaling large language models (LLMs). In particular, leveraging readily available fine-tuned LLMs as experts provides an efficient and flexible approach to developing MoE LLMs. However, integrating highly capable but untrustworthy LLMs into an MoE system poses a significant safety risk, potentially compromising the overall safety of the MoE LLM system. To date, no study has explored how adversaries compromise an MoE LLM service by introducing a poisoned expert LLM. In this paper, we introduce MOEVIL, a novel expert poisoning attack designed to compromise the safety of MoE LLMs. We address the dissipation of harmful effects from a target expert within MoE systems by conducting harmful preference learning. Next, we strategically manipulate this expert's latent vector to deceive the gating networks. This manipulation indirectly steers routing decisions toward the poisoned expert when generating responses to harmful queries. MOEVIL demonstrates strong attack performance across diverse MoE configurations based on both Llama and Qwen LLMs, even when poisoning only a single expert. MOEVIL increases the harmfulness score from 0.58 to 79.42 in a Llama-based MoE LLM, outperforming existing harmful poisoning attacks. Furthermore, our results demonstrate that even safety alignment, when combined with an efficient MoE training strategy, fails to fully mitigate these risks. Our findings demonstrate the significant threat posed by harmful experts in MoE systems, underscoring the need for robust safety measures in MoE-based LLM development. Our implementation is available at https://github.com/jaehanwork/MoEvil. Jaehan Kim, Seung Ho Na, Minkyoo Song, Seungwon Shin 0001, Sooel Son |
ACSAC | 3 |
| 2025 | Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionabstractIn light of rising drug-related concerns and the increasing role of social media, sales and discussions of illicit drugs have become commonplace online. Social media platforms hosting user-generated content must therefore perform content moderation, which is a difficult task due to the vast amount of jargon used in drug discussions. Previous works on drug jargon detection were limited to extracting a list of terms, but these approaches have fundamental problems in practical application. First, they are trivially evaded using word substitutions. Second, they cannot distinguish whether euphemistic terms (pot, crack) are being used as drugs or as their benign meanings. We argue that drug content moderation should be done using contexts, rather than relying on a banlist. However, manually annotated datasets for training such a task are not only expensive but also prone to becoming obsolete. We present JEDIS, a framework for detecting illicit drug jargon terms by analyzing their contexts. JEDIS utilizes a novel approach that combines distant supervision and delexicalization, which allows JEDIS to be trained without human-labeled data while being robust to new terms and euphemisms. Experiments on two manually annotated datasets show JEDIS significantly outperforms state-of-the-art word-based baselines in terms of F1-score and detection coverage in drug jargon detection. We also conduct qualitative analysis that demonstrates JEDIS is robust against pitfalls faced by existing approaches. Minkyoo Song, Eugene Jang, Jaehan Kim, Seungwon Shin 0001 |
KDD (1) | 1 |
| 2025 | When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
Hanna Kim, Minkyoo Song, Seung Ho Na, Seungwon Shin 0001, Kimin Lee |
USENIX Security Symposium | 2 |
| 2025 | Refusal Is Not an Option: Unlearning Safety Alignment of Large Language Models
Minkyoo Song, Hanna Kim, Jaehan Kim, Seungwon Shin 0001, Sooel Son |
USENIX Security Symposium | 1 |
| 2024 | PassREfinder: Credential Stuffing Risk Prediction by Representing Password Reuse between Websites on a GraphabstractThe prevalence of credential stuffing has caused devastating harm to online users who tend to reuse passwords across websites. In response, researchers have made efforts to detect users who set the same passwords or malicious logins. However, existing detection methods sacrifice the usability of passwords by inhibiting password creation or website access. Moreover, the complicated mechanisms for sharing account information hinder their deployment in practice. In this work, we propose a risk prediction framework to prevent credential stuffing attacks before disrupting user behaviors rather than relying on detection. To this end, we newly define the relationship between websites in which users are highly likely to reuse passwords and represent it as an edge on a website graph using graph neural networks. We then perform a link prediction task to identify the risk of credential stuffing between websites. Our framework is applicable to a large number of arbitrary websites by utilizing public website information and linking newly observed website nodes to the graph. The evaluation on a real-world credential dataset consisting of 360 million accounts breached from 22,378 websites shows that our model successfully predicts credential stuffing risk among websites by achieving F1-scores of 0.9559 and 0.9100 in two different graph learning settings, respectively. In addition, we demonstrate the effectiveness of each design strategy and validate that the prediction results can be utilized to quantify the expected rates of password reuse as risk scores. Jaehan Kim, Minkyoo Song, Minjae Seo, Youngjin Jin, Seungwon Shin 0001 |
SP | 2 |