VLDB 2026 Research / reviewers in the wild / expert
Kiho Lee
dblp:04/5599
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdVersa: Adversarially-Robust and Practical Ad and Tracker Blocking in the Wild
Chaejin Lim, Kiho Lee, Beomjin Jin, Heewon Baek, Hyoungshick Kim |
WWW | 2 |
| 2025 | When Does Wasm Malware Detection Fail? A Systematic Analysis of Their Robustness to EvasionabstractWebAssembly (Wasm) provides a language-agnostic compilation target that delivers near-native performance for web applications, yet it also attracts adversaries who exploit Wasm to effectively steal someone else’s computer resources such as cryptojackers. While several detection tools have been proposed, their robustness against perturbations remains largely unknown.In this paper, we introduce Swamped (Systematic WebAssembly Module Perturbation Evaluation of Detectors), a framework that incorporates 22 semantics-preserving perturbation methods. Swamped generates a total of 48,840 perturbed variants from 43 cryptojacker samples and 31 additional Wasm malware binaries from real-world. We assess detection performance of six detectors: three Wasm-specific ones and three deep neural network (DNN) detectors. We find that DNN-based detectors are vulnerable to perturbations that shift the instruction distribution; profiling-based methods are disrupted by changes in instruction frequency; and semantic-aware approaches are highly sensitive to function-level dependency modifications. DNN-based detectors, which lack Wasm-specific modeling, are particularly susceptible to changes in the spatial layout of Wasm binaries. These findings highlight fundamental limitations in current Wasm malware detection approaches, relying on overly specific detection heuristics and inadequately trained or designed models. We offer suggestions to improve the robustness against perturbations. Sanghak Oh, Kiho Lee, Weihang Wang 0001, Yonghwi Kwon 0001, Sanghyun Hong 0001, Hyoungshick Kim |
ASE | 3 |
| 2025 | Open Sesame! On the Security and Memorability of Verbal PasswordsabstractDespite extensive research on text passwords, the security and memorability of verbal passwords-spoken rather than typed-remain underexplored. Verbal passwords hold significant potential for scenarios where keyboard input is impractical (e.g., smart speakers, wearables, vehicles) or users have motor impairments that make typing difficult. Through two large-scale user studies, we assessed the viability of verbal passwords. In our first study (N = 2,085), freely chosen verbal passwords were found to have a limited guessing space, with 39.76% cracked within 109guesses. However, in our second study (n = 600), applying word count and blocklist policies for verbal password creation significantly enhanced verbal password performance, achieving better memorability and security than traditional text passwords. Specifically, 65.6% of verbal password users (under the password creation policy using minimum word counts and a blocklist) successfully recalled their passwords in long-term tests, compared to 54.11% for text passwords. Additionally, verbal passwords with enforced policies exhibited a lower crack rate (6.5%) than text passwords (10.3%). These findings highlight verbal passwords as a practical and secure alternative for contexts where text passwords are infeasible, offering strong memorability with robust resistance to guessing attacks. Eunsoo Kim, Kiho Lee, Doowon Kim, Hyoungshick Kim |
SP | 2 |
| 2025 | Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection Models
Fujiao Ji, Kiho Lee, Hyungjoon Koo, Wenhao You, Euijin Choo, Hyoungshick Kim, Doowon Kim |
USENIX Security Symposium | 2 |
| 2025 | 7 Days Later: Analyzing Phishing-Site Lifespan After DetectedabstractPhishing attacks continue to be a major threat to internet users, causing data breaches, financial losses, and identity theft. This study provides an in-depth analysis of the lifespan and evolution of phishing websites, focusing on their survival strategies and evasion techniques. We analyze 286,237 unique phishing URLs over five months using a custom web crawler based on Puppeteer and Chromium. Our crawler runs on a 30-minute cycle, systematically checking the operational status of phishing websites by collecting their HTTP status codes, screenshots, HTML, and HTTP data. Temporal and survival analyses, along with statistical tests, are used to examine phishing website lifecycles, evolution, and evasion tactics. Our findings show that the average lifespan of phishing websites is 54 hours (2.25 days) with a median of 5.46 hours, indicating rapid takedown of many sites while a subset remains active longer. Interestingly, logistic-themed phishing websites (e.g., USPS) operate within a compressed timeframe (1.76 hours) compared to other brands (e.g., Facebook). We further analyze detection effectiveness using Google Safe Browsing (GSB). We find that GSB detects only 18.4% of phishing websites, taking an average of 4.5 days. Notably, 83.93% of phishing sites are already taken down before GSB detection, meaning GSB requires more prompt detection. Moreover, 16.07% of phishing sites persist beyond this point, surviving for an additional 7.2 days on average, resulting in an average total lifespan of approximately 12 days. We reveal that DNS resolution error is the main cause (67%) of phishing website takedowns. Finally, we uncover that phishing sites with extensive visual changes (more than 100 times) exhibit a median lifespan of 17 days, compared to 1.93 hours for those with minimal modifications. These results highlight the dynamic nature of phishing attacks, the challenges in detection and prevention, and the need for more rapid and comprehensive countermeasures against evolving phishing tactics. Kiho Lee, Kyungchan Lim, Hyoungshick Kim, Yonghwi Kwon 0001, Doowon Kim |
WWW | 1 |
| 2025 | What's in Phishers: A Longitudinal Study of Security Configurations in Phishing Websites and KitsabstractPhishing attacks pose a significant threat to Internet users. Understanding the security posture of phishing infrastructure is crucial for developing effective defense strategies, as it helps identify potential weaknesses that attackers might exploit. Despite extensive research, there may still be a gap in fully understanding these security weaknesses. To address this important issue, this paper presents a longitudinal study of security configurations and vulnerabilities in phishing websites and associated kits. We focus on two main areas: (1) analyzing the security configurations of phishing websites and servers, particularly HTTP headers and application-level security, and (2) examining the prevalence and types of vulnerabilities in phishing kits. We analyze data from 906,731 distinct phishing websites collected over 2.5 years, covering HTML headers, client-side resources, and phishing kits. Our findings suggest that phishing websites often employ weak security configurations, with 88.8% of the 13,344 collected phishing kits containing at least one potential vulnerability, and 12.5% containing backdoor vulnerabilities. These vulnerabilities present an opportunity for defenders to shift from passive defense to active disruption of phishing operations. Our research proposes a new approach to leverage weaknesses in phishing infrastructure, allowing defenders to take proactive actions to disable phishing sites earlier and reduce their effectiveness. Kyungchan Lim, Kiho Lee, Fujiao Ji, Yonghwi Kwon 0001, Hyoungshick Kim, Doowon Kim |
WWW | 2 |
| 2024 | Poisoned ChatGPT Finds Work for Idle Hands: Exploring Developers' Coding Practices with Insecure Suggestions from Poisoned AI ModelsabstractAI-powered coding assistant tools (e.g., ChatGPT, Copilot, and IntelliCode) have revolutionized the software engineering ecosystem. However, prior work has demonstrated that these tools are vulnerable to poisoning attacks. In a poisoning attack, an attacker intentionally injects maliciously crafted insecure code snippets into training datasets to manipulate these tools. The poisoned tools can suggest insecure code to developers, resulting in vulnerabilities in their products that attackers can exploit. However, it is still little understood whether such poisoning attacks against the tools would be practical in real-world settings and how developers address the poisoning attacks during software development. To understand the real-world impact of poisoning attacks on developers who rely on AI-powered coding assistants, we conducted two user studies: an online survey and an in-lab study. The online survey involved 238 participants, including software developers and computer science students. The survey results revealed widespread adoption of these tools among participants, primarily to enhance coding speed, eliminate repetition, and gain boilerplate code. However, the survey also found that developers may misplace trust in these tools because they overlooked the risk of poisoning attacks. The in-lab study was conducted with 30 professional developers. The developers were asked to complete three programming tasks with a representative type of AI-powered coding assistant tool (e.g., ChatGPT or IntelliCode), running on Visual Studio Code. The in-lab study results showed that developers using a poisoned ChatGPT-like tool were more prone to including insecure code than those using an IntelliCode-like tool or no tool. This demonstrates the strong influence of these tools on the security of generated code. Our study results highlight the need for education and improved coding practices to address new security issues introduced by AI-powered coding assistant tools. Sanghak Oh, Kiho Lee, Seonhye Park, Doowon Kim, Hyoungshick Kim |
SP | 2 |
| 2024 | An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
Shenao Yan, Yue Duan, Hanbin Hong, Kiho Lee, Doowon Kim, Yuan Hong 0001 |
USENIX Security Symposium | 5 |
| 2024 | AdFlush: A Real-World Deployable Machine Learning Solution for Effective Advertisement and Web Tracker PreventionabstractConventional ad blocking and tracking prevention tools often fall short in addressing web content manipulation. Machine learning approaches have been proposed to enhance detection accuracy, yet aspects of practical deployment have frequently been overlooked. This paper introduces AdFlush, a novel machine learning model for real-world browsers. To develop AdFlush, we evaluated the effectiveness of 883 features, ultimately selecting 27 key features for optimal performance. We tested AdFlush on a dataset of 10,000 real-world websites, achieving an F1 score of 0.98, thereby outperforming AdGraph (F1 score: 0.93), WebGraph (F1 score: 0.90), and WTAgraph (F1 score: 0.84). Additionally, AdFlush significantly reduces computational overhead, requiring 56% less CPU and 80% less memory than AdGraph. We also assessed AdFlush's robustness against adversarial manipulations, demonstrating superior resilience with F1 scores ranging from 0.89 to 0.98, surpassing the performance of AdGraph and WebGraph, which recorded F1 scores between 0.81 and 0.87. A six-month longitudinal study confirmed that AdFlush maintains a high F1 score above 0.97 without the need for retraining, underscoring its effectiveness. Kiho Lee, Chaejin Lim, Beomjin Jin, Hyoungshick Kim |
WWW | 1 |
| 2022 | Poster: Adversarial Perturbation Attacks on the State-of-the-Art Cryptojacking Detection System in IoT NetworksabstractThe popularity of cryptocurrency raised a new cyber security threat dubbed cryptojacking representing malicious activities for abusing victims' computing resources without their consent to mine cryptocurrency. Recently, Tekiner et al. [1] proposed an effective cryptojacking detection technique using a machine learning model with the statistical properties of the network traffic for cryptojacking in the Internet of Things (IoT) devices. In this paper, however, we demonstrate that this state-of-the-art method can effectively be evaded by maliciously manipulating the network packets for cryptojacking. Our evaluation results show that packet manipulations (packet splitting, dummy packet/payload insertion, and a proxy network) can effectively evade the model's detection -- the packet splitting technique significantly decreased the F1-score of the detection model from 0.93 to 0.30. Finally, the best combination of those packet manipulations can decrease the F1-score of the detection model to 0.21. Kiho Lee, Sanghak Oh, Hyoungshick Kim |
CCS | 1 |
| 2015 | Sigma-RF: prediction of the variability of spatial restraints in template-based modeling by random forestabstractBACKGROUND: In template-based modeling when using a single template, inter-atomic distances of an unknown protein structure are assumed to be distributed by Gaussian probability density functions, whose center peaks are located at the distances between corresponding atoms in the template structure. The width of the Gaussian distribution, the variability of a spatial restraint, is closely related to the reliability of the restraint information extracted from a template, and it should be accurately estimated for successful template-based protein structure modeling. RESULTS: To predict the variability of the spatial restraints in template-based modeling, we have devised a prediction model, Sigma-RF, by using the random forest (RF) algorithm. The benchmark results on 22 CASP9 targets show that the variability values from Sigma-RF are of higher correlations with the true distance deviation than those from Modeller. We assessed the effect of new sigma values by performing the single-domain homology modeling of 22 CASP9 targets and 24 CASP10 targets. For most of the targets tested, we could obtain more accurate 3D models from the identical alignments by using the Sigma-RF results than by using Modeller ones. CONCLUSIONS: We find that the average alignment quality of residues located between and at two aligned residues, quasi-local information, is the most contributing factor, by investigating the importance of input features used in the RF machine learning. This average alignment quality is shown to be more important than the previously identified quantity of a local information: the product of alignment qualities at two aligned residues. Kiho Lee, InSuk Joung, Keehyoung Joo, Bernard R. Brooks, Jooyoung Lee 0002 |
BMC Bioinform. | 2 |
| 2001 | The Cost Model for XML Documents in Relational Database SystemsabstractXML is emerging as the standard for representing data in the World Wide Web. If a data server can efficiently store and query the XML documents, it will allow users to manage volumes of data effectively. We describe new storage strategies and a cost model for XML documents. We suggest two storage techniques for storing XML documents in a relational database system. Our storage techniques preserve the attribute of XML documents such as the hierarchical structure of documents by considering both XML DTD and XML instances. The cost model is extended to estimate the queries on the XML documents in a relational database system. We have evaluated the performance of queries for four different storage techniques including suggested storage models and existing storage models by our cost model. JiSim Kim, Wol-Young Lee, Kiho Lee |
AICCSA | 3 |
| 2001 | Preparations for Semantics-Based XML MiningabstractXML allows users to define elements using arbitrary words and organize them in a nested structure. These features of XML offer both challenges and opportunities in information retrieval, document management, and data mining. In this paper, we propose a new methodology for preparing XML documents for quantitative determination of similarity between XML documents by taking into account XML semantics (i.e., meanings of the elements and nested structures of XML documents). Accurate quantitative determination of similarity between XML documents provides an important basis for a variety of applications of XML document mining and processing. Experiments with XML documents show that our methodology provides a 50-100% improvement in determining similarity over the traditional vector-space model that considers only term-frequency and 100% accuracy in identifying the category of each document from an on-line bookstore. Jung-Won Lee, Kiho Lee, Won Kim 0001 |
ICDM | 2 |