Kyu Hyung Lee

dblp:31/9698 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-0582-5795ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 18 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Racedb: Detecting Request Race Vulnerabilities in Database-Backed Web Applications
abstract
Request race vulnerabilities in database-backed web applications pose a significant security threat. These vulnera-bilities can lead to data inconsistencies, unexpected behavior, and even unauthorized access. Existing automated detection techniques often fall short due to the complexity of race conditions and the intricate interplay between application logic and database interactions. This paper introduces Racedb, a novel system that tackles these challenges through two key innovations. Application-aware Request Race Detection (ARD) provides a comprehensive analysis of data dependencies, considering not only the database query but also the application code. This allows RacedB to identify subtle race conditions that might be missed by existing approaches. Furthermore, Racedbemploys an automated verification technique using replay-based execution. This technique efficiently isolates true races from false positives and generates definitive exploits for verified vulnerabilities. We evaluated Racedb on a dataset of 14 real-world PHP web applications. The results demonstrate the effectiveness of Racedb compared to existing tools. Racedb achieved a superior detection rate, identifying 21 known vul-nerabilities and discovering 18 new vulnerabilities, significantly exceeding the performance of existing tools while also achieving a lower rate of false positives. Finally, we responsibly reported all newly discovered vulnerabilities to the corresponding developers, and 7 of them have been assigned CVE IDs.
Yonghwi Kwon 0001, Kyu Hyung Lee
SP3
2024 FA-SEAL: Forensically Analyzable Symmetric Encryption for Audit Logs
abstract
Audit logs serve as crucial resources for cybersecu-rity analysts conducting forensic analysis to gain insights into the constantly evolving cyberattacks. Due to the substantial volume of these logs, a growing number of companies have chosen to outsource their log storage to cloud-based solutions. To enhance the security of externally stored audit logs, which often contain sensitive data susceptible to breaches, encrypting them at all time is recommended, following guidance from the National Institute of Standards and Technology (NIST). How-ever, conducting forensic analysis on encrypted audit logs poses a non-trivial challenge, often requiring the complete decryption of the logs before analysis. Given that companies frequently outsource forensic investigations to third-party entities, the need to share fully decrypted audit logs with these parties may risk exposing sensitive data to external parties.In this paper, we introduce a new approach called FA-SEAL, to enable forensic analysis on encrypted logs without fully decrypting them, but instead selectively discloses only information related to specific incidents to third-party entities. Moreover, to handle the encryption of logs produced on a massive scale over time, FA-SEAL introduces the concept of segmentation and clustering to effectively manage the ingestion of a large volume of audit logs. This approach also allows for both backward and forward forensic analysis without neces-sitating the decryption of the entire log set. Our experiments involving a log dataset embedded with cyberattacks generated by AIT [1], along with eight simulated attack scenarios, show that FA-SEAL enables near real-time ingestion of logs. Fur-thermore, the causal graphs generated from the encrypted logs using FA-SEAL are found to be identical to those produced by the baseline system, which necessitates decrypting the entire collection of audit logs.
Basanta Chaulagain, Kyu Hyung Lee
ACSAC2
2024 RustLIVE: Reducing the Learning Barriers of Rust Through Visualization
abstract
This innovative practice full paper introduces an independent learning tool for the Rust programming language. Secure software begins with secure memory management as over 65% of software vulnerabilities in modern code bases are the result of memory safety problems. Rust is a promising memory safe language that is rapidly gaining traction as a replacement for C and C++ in systems software. However, Rust is challenging due to its novel approach to memory safety. The Rust programming paradigm is focused on the principles of ownership and borrowing. These concepts effectively achieve memory and thread safety but are enforced at compile-time, enabling run-time performance comparable to C but posing significant implementation challenges for students and even experienced programmers. We propose an innovative visualization tool, RustLIVE, that clearly depicts the most difficult concepts of the Rust language. RustLIVE is an extension for VSCode, the IDE of choice for over 60% of Rust developers. Seamlessly integrated with the Rust compiler and its borrow checker, RustLIVE requires no code annotations. Instead, it extracts necessary information directly from the compiler to provide color-coded visual timelines that illustrate the ownership of memory resources and the liveness of borrows. RustLIVE is an innovative, independent learning tool that helps learners form a correct mental model, avoid unsafe code and improve student experience.
Diane B. Stephens, Kyu Hyung Lee, Mustakimur Khandaker
FIE2
2024 Unveiling IoT Security in Reality: A Firmware-Centric Journey
Nicolas Nino, Ruibo Lu, Wei Zhou 0026, Kyu Hyung Lee, Ziming Zhao 0001, Le Guan
USENIX Security Symposium4
2023 SynthDB: Synthesizing Database via Program Analysis for Security Testing of Web Applications
Basanta Chaulagain, Yonghwi Kwon 0001, Kyu Hyung Lee
NDSS5
2022 Hiding Critical Program Components via Ambiguous Translation
abstract
Software systems may contain critical program components such as patented program logic or sensitive data. When those components are reverse-engineered by adversaries, it can cause significantly damage (e.g., financial loss or operational failures). While protecting critical program components (e.g., code or data) in software systems is of utmost importance, existing approaches, unfortunately, have two major weaknesses: (1) they can be reverse-engineered via various program analysis techniques and (2) when an adversary obtains a legitimate-looking critical program component, he or she can be sure that it is genuine.
Chijung Jung, Doowon Kim, Weihang Wang 0001, Yunhui Zheng, Kyu Hyung Lee, Yonghwi Kwon 0001
ICSE6
2022 Privacy invasion via smart-home hub in personal area networks
Omid Setayeshfar, Karthika Subramani, Xingzi Yuan, Raunak Dey, Dezhi Hong, In Kee Kim, Kyu Hyung Lee
Pervasive Mob. Comput.7
2021 A Novel AI-based Methodology for Identifying Cyber Attacks in Honey Pots
abstract
We present a novel AI-based methodology that identifies phases of a host-level cyber attack simply from system call logs. System calls emanating from cyber attacks on hosts such as honey pots are often recorded in audit logs. Our methodology first involves efficiently loading, caching, processing, and querying system events contained in audit logs in support of computer forensics. Output of queries remains at the system call level and is difficult to process. The next step is to infer a sequence of abstracted actions, which we colloquially call a storyline, from the system calls given as observations to a latent-state probabilistic model. These storylines are then accurately identified with class labels using a learned classifier. We qualitatively and quantitatively evaluate methods and models for each step of the methodology using 114 different attack phases collected by logging the attacks of a red team on a server, on some likely benign sequences containing regular user activities, and on traces from a recent DARPA project. The resulting end-to-end system, which we call Cyberian, identifies the attack phases with a high level of accuracy illustrating the benefit that this machine learning-based methodology brings to security forensics.
Muhammed AbuOdeh, Christian Adkins, Omid Setayeshfar, Prashant Doshi, Kyu Hyung Lee
AAAI5
2021 Find My Sloths: Automated Comparative Analysis of How Real Enterprise Computers Keep Up with the Software Update Races
Omid Setayeshfar, Junghwan Rhee, Kyu Hyung Lee
DIMVA4
2021 Defeating Program Analysis Techniques via Ambiguous Translation
abstract
This research explores the possibility of a new anti-analysis technique, carefully designed to attack weaknesses of the existing program analysis approaches. It encodes a program code snippet to hide, and its decoding process is implemented by a sophisticated state machine that produces multiple outputs depending on inputs. The key idea of the proposed technique is to ambiguously decode the program code, resulting in multiple decoded code snippets that are challenging to distinguish from each other. Our approach is stealthier than previous similar approaches as its execution does not exhibit different behaviors between when it decodes correctly or incorrectly. This paper also presents analyses of weaknesses of existing techniques and discusses potential improvements. We implement and evaluate the proof of concept approach, and our preliminary results show that the proposed technique imposes various new unique challenges to the program analysis technique.
Chijung Jung, Doowon Kim, Weihang Wang 0001, Yunhui Zheng, Kyu Hyung Lee, Yonghwi Kwon 0001
ASE5
2021 C^2SR: Cybercrime Scene Reconstruction for Post-mortem Forensic Analysis
Yonghwi Kwon 0001, Weihang Wang 0001, Jinho Jung 0001, Kyu Hyung Lee, Roberto Perdisci
NDSS4
2021 ChatterHub: Privacy Invasion via Smart Home Hub
abstract
Smart-home devices promise to make users’ lives more convenient. However, at the same time, such devices increase the possibility of breaching users’ privacy as they are tightly connected to the users’ daily lives and activities. To address privacy invasion through smart-home devices, we present ChatterHub. This novel approach accurately identifies smart-home devices’ activities with minimal monitoring of encrypted traffic in the home network. ChatterHub targets devices that can only connect to the Internet through a centralized smart-home hub (e.g., Samsung SmartThings) using Zigbee or Z-wave. Specifically, ChatterHub passively eavesdrops on encrypted network traffic from the hub and leverages machine learning techniques to classify events and states of smart-home devices. Using ChatterHub, an adversary can identify smart-home devices’ specific activities without prior knowledge of the target smart home (e.g., list of deployed devices, types of communication protocols). We evaluated the accuracy and efficiency of ChatterHub in three real-world smart-home environments, and the evaluation results show that an attacker can successfully disclose smart-home devices’ behaviors with over 88% F1 score. We further demonstrate that ChatterHub successfully recognizes privacy-sensitive activities, including open and close of a smart door lock and turn on and off of smart LED. Additionally, to mitigate the threats posed by ChatterHub, we introduce two approaches, packet padding and random sequence injection. These mitigation approaches can effectively prevent threats from ChatterHub with only 9.2MB of additional network traffic per day.
Omid Setayeshfar, Karthika Subramani, Xingzi Yuan, Raunak Dey, Dezhi Hong, Kyu Hyung Lee, In Kee Kim
SMARTCOMP6
2021 TRACE: Enterprise-Wide Provenance Tracking for Real-Time APT Detection
abstract
We present TRACE, a comprehensive provenance tracking system for scalable, real-time, enterprise-wide APT detection. TRACE uses static analysis to identify program unit structures and inter-unit dependences, such that the provenance of an output event includes the input events within the same unit. Provenance collected from individual hosts are integrated to facilitate construction of a distributed enterprise-wide causal graph. We describe the evolution of TRACE over a four-year period, during which our improvements to the system focused on performance, scalability, and fidelity. In this time span, the system call coverage increased (from 47 to 66) while the time and space overhead reduced by over one and two orders of magnitude, respectively. We also provide results from five adversarial engagements where an independent team of system evaluators conducted APT attacks and assessed system performance. The input from our system was used by three other teams to implement real-time APT detection logic. Retrospective analysis revealed that TRACE provided sufficient evidence to detect over 80% of the attack stages across all evaluations. By the last engagement, temporal and spatial overhead had been reduced significantly to 18% and 10%, respectively.
Hassaan Irshad, Gabriela F. Ciocarlie, Ashish Gehani, Vinod Yegneswaran, Kyu Hyung Lee, Jignesh M. Patel, Somesh Jha, Yonghwi Kwon 0001, Dongyan Xu, Xiangyu Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2020 When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising
abstract
The rapid growth of online advertising has fueled the growth of ad-blocking software, such as new ad-blocking and privacy-oriented browsers or browser extensions. In response, both ad publishers and ad networks are constantly trying to pursue new strategies to keep up their revenues. To this end, ad networks have started to leverage the Web Push technology enabled by modern web browsers. As web push notifications (WPNs) are relatively new, their role in ad delivery has not been yet studied in depth. Furthermore, it is unclear to what extent WPN ads are being abused for malvertising (i.e., to deliver malicious ads). In this paper, we aim to fill this gap. Specifically, we propose a system called PushAdMiner that is dedicated to (1) automatically registering for and collecting a large number of web-based push notifications from publisher websites, (2) finding WPN-based ads among these notifications, and (3) discovering malicious WPN-based ad campaigns. Using PushAdMiner, we collected and analyzed 21,541 WPN messages by visiting thousands of different websites. Among these, our system identified 572 WPN ad campaigns, for a total of 5,143 WPN-based ads that were pushed by a variety of ad networks. Furthermore, we found that 51% of all WPN ads we collected are malicious, and that traditional ad-blockers and malicious URL filters are remarkably ineffective against WPN-based malicious ads, leaving a significant abuse vector unchecked.
Karthika Subramani, Xingzi Yuan, Omid Setayeshfar, Phani Vadrevu, Kyu Hyung Lee, Roberto Perdisci
Internet Measurement Conference5
2019 Fuzzification: Anti-Fuzzing Techniques
Jinho Jung 0001, Hong Hu 0004, David Solodukhin, Daniel Pagan, Kyu Hyung Lee, Taesoo Kim
USENIX Security Symposium5
2018 MCI : Modeling-based Causality Inference in Audit Logging for Attack Investigation
Yonghwi Kwon 0001, Fei Wang 0001, Weihang Wang 0001, Kyu Hyung Lee, Wen-Chuan Lee, Shiqing Ma, Xiangyu Zhang 0001, Dongyan Xu, Somesh Jha, Gabriela F. Ciocarlie, Ashish Gehani, Vinod Yegneswaran
NDSS4
2018 JSgraph: Enabling Reconstruction of Web Attacks via Efficient Tracking of Live In-Browser JavaScript Executions
Bo Li 0058, Phani Vadrevu, Kyu Hyung Lee, Roberto Perdisci
NDSS3
2018 Kernel-Supported Cost-Effective Audit Logging for Causality Tracking
Shiqing Ma, Juan Zhai, Yonghwi Kwon 0001, Kyu Hyung Lee, Xiangyu Zhang 0001, Gabriela F. Ciocarlie, Ashish Gehani, Vinod Yegneswaran, Dongyan Xu, Somesh Jha
USENIX ATC4
2017 DroidForensics: Accurate Reconstruction of Android Attacks via Multi-layer Forensic Logging
abstract
The goal of cyber attack investigation is to fully reconstruct the details of an attack, so we can trace back to its origin, and recover the system from the damage caused by the attack. However, it is often difficult and requires tremendous manual efforts because attack events occurred days or even weeks before the investigation and detailed information we need is not available anymore. Consequently, forensic logging is significantly important for cyber attack investigation. In this paper, we present DroidForensics, a multi-layer forensic logging technique for Android. Our goal is to provide the user with detailed information about attack behaviors that can enable accurate post-mortem investigation of Android attacks. DroidForensics consists of three logging modules. API logger captures Android API calls that contain high-level semantics of an application. Binder logger records interactions between applications to identify causal relations between processes, and system call logger efficiently monitors low-level system events. We also provide the user interface that the user can compose SQL-like queries to inspect an attack. Our experiments show that Droid Forensics has low runtime overhead (2.9% on average) and low space overhead (105 ~ 169 MByte during 24 hours) on real Android devices. It is effective in the reconstruction of realworld Android attacks we have studied.
Xingzi Yuan, Omid Setayeshfar, Hongfei Yan, Pranav Panage, Xuetao Wei, Kyu Hyung Lee
AsiaCCS6
2017 Self Destructing Exploit Executions via Input Perturbation
Yonghwi Kwon 0001, Brendan Saltaformaggio, I Luk Kim, Kyu Hyung Lee, Xiangyu Zhang 0001, Dongyan Xu
NDSS4
2017 Enabling Reconstruction of Attacks on Users via Efficient Browsing Snapshots
Phani Vadrevu, Jienan Liu, Bo Li 0058, Babak Rahbarinia, Kyu Hyung Lee, Roberto Perdisci
NDSS5
2017 MPI: Multiple Perspective Attack Investigation with Semantic Aware Execution Partitioning
Shiqing Ma, Juan Zhai, Fei Wang 0001, Kyu Hyung Lee, Xiangyu Zhang 0001, Dongyan Xu
USENIX Security Symposium4
2016 emphaSSL: Towards Emphasis as a Mechanism to Harden Networking Security in Android Apps
abstract
The use of secure HTTP calls is a first and critical step toward securing the Android application data when the app interacts with the Internet. However, one of the major causes for the unencrypted communication is app developer's errors or ignorance. Could the paradigm of literally repetitive and ineffective emphasis shift towards emphasis as a mechanism? This paper introduces emphaSSL, a simple, practical and readily-deployable way to harden networking security in Android applications. Our emphaSSL could guide app developer's security development decisions via real-time feedback, informative warnings and suggestions. At its core of emphaSSL, we use a set of rigorous security rules, which are obtained through an in-depth SSL/TLS security analysis based on security requirements engineering techniques. We implement emphaSSL via the PMD and evaluate it against 75 open- source Android applications. Our results show that emphaSSL is effective at detecting security violations in HTTPS calls with a very low false positive rate, around 2%. Furthermore, we identified 164 substantial SSL mistakes in these testing apps, 40% of which are potentially vulnerable to man-in-the-middle attacks. In each of these instances, the vulnerabilities could be quickly resolved with the assistance of our highlighting messages in emphaSSL. Upon notifying developers of our findings in their applications, we received positive responses and interest in this approach.
Xuetao Wei, Michael Wolf, Lei Guo 0005, Kyu Hyung Lee, Ming-Chun Huang, Nan Niu
GLOBECOM4
2016 PerfGuard: binary-centric application performance monitoring in production environments
abstract
Diagnosis of performance problems is an essential part of software development and maintenance. This is in particular a challenging problem to be solved in the production environment where only program binaries are available with limited or zero knowledge of the source code. This problem is compounded by the integration with a significant number of third-party software in most large-scale applications. Existing approaches either require source code to embed manually constructed logic to identify performance problems or support a limited scope of applications with prior manual analysis. This paper proposes an automated approach to analyze application binaries and instrument the binary code transparently to inject and apply performance assertions on application transactions. Our evaluation with a set of large-scale application binaries without access to source code discovered 10 publicly known real world performance bugs automatically and shows that PerfGuard introduces very low overhead (less than 3% on Apache and MySQL server) to production systems.
Junghwan Rhee, Kyu Hyung Lee, Xiangyu Zhang 0001, Dongyan Xu
SIGSOFT FSE3
2015 Accurate, Low Cost and Instrumentation-Free Security Audit Logging for Windows
abstract
Audit logging is an important approach to cyber attack investigation. However, traditional audit logging either lacks accuracy or requires expensive and complex binary instrumentation. In this paper, we propose a Windows based audit logging technique that features accuracy and low cost. More importantly, it does not require instrumenting the applications, which is critical for commercial software with IP protection. The technique is build on Event Tracing for Windows (ETW). By analyzing ETW log and critical parts of application executables, a model can be constructed to parse ETW log to units representing independent sub-executions in a process. Causality inferred at the unit level renders much higher accuracy, allowing us to perform accurate attack investigation and highly effective log reduction.
Shiqing Ma, Kyu Hyung Lee, Junghwan Rhee, Xiangyu Zhang 0001, Dongyan Xu
ACSAC2
2014 Infrastructure-Free Logging and Replay of Concurrent Execution on Multiple Cores
Kyu Hyung Lee, Dohyeong Kim, Xiangyu Zhang 0001
ECOOP1
2014 Infrastructure-free logging and replay of concurrent execution on multiple cores
abstract
We develop a logging and replay technique for real concurrent execution on multiple cores. Our technique directly works on binaries and does not require any hardware or complex software infrastructure support. We focus on minimizing logging overhead as it only logs a subset of system calls and thread spawns. Replay is on a single core. During replay, our technique first tries to follow only the event order in the log. However, due to schedule differences, replay may fail. An exploration process is then triggered to search for a schedule that allows the replay to make progress. Exploration is performed within a window preceding the point of replay failure. During exploration, our technique first tries to reorder synchronized blocks. If that does not lead to progress, it further reorders shared variable accesses. The exploration is facilitated by a sophisticated caching mechanism. Our experiments on real world programs and real workload show that the proposed technique has very low logging overhead (2.6% on average) and fast schedule reconstruction.
Kyu Hyung Lee, Dohyeong Kim, Xiangyu Zhang 0001
PPoPP1
2013 LogGC: garbage collecting audit log
abstract
System-level audit logs capture the interactions between applications and the runtime environment. They are highly valuable for forensic analysis that aims to identify the root cause of an attack, which may occur long ago, or to determine the ramifications of an attack for recovery from it. A key challenge of audit log-based forensics in practice is the sheer size of the log files generated, which could grow at a rate of Gigabytes per day. In this paper, we propose LogGC, an audit logging system with garbage collection (GC) capability. We identify and overcome the unique challenges of garbage collection in the context of computer forensic analysis, which makes LogGC different from traditional memory GC techniques. We also develop techniques that instrument user applications at a small number of selected places to emit additional system events so that we can substantially reduce the false dependences between system events to improve GC effectiveness. Our results show that LogGC can reduce audit log size by 14 times for regular user systems and 37 times for server systems, without affecting the accuracy of forensic analysis.
Kyu Hyung Lee, Xiangyu Zhang 0001, Dongyan Xu
CCS1
2013 High Accuracy Attack Provenance via Binary-based Execution Partition
Kyu Hyung Lee, Xiangyu Zhang 0001, Dongyan Xu
NDSS1
2011 Unified debugging of distributed systems with Recon
abstract
To scale to today's complex distributed software systems, debugging and replaying techniques mostly focus on single facets of software, e.g., local concurrency, distributed messaging, or data representation. This forces developers to tediously combine different technologies such as instruction-level dynamic tracing, event log analysis, or global state reconstruction to gradually explain non-trivial defects. This paper proposes Recon, a debugging system that provides iterative and interactive homogeneous debugging services. As related systems, Recon promotes SQL-like queries for debugging distributed systems. Unlike other approaches, however, Recon allows for all system artifacts including nodes, communication channels, events, or instructions to be uniformly described by relations. Also, an application in Recon originally runs with a lightweight logger that only collects replay logs for individual nodes. Developers debug a complete program by replaying the execution with fine-grained instrumentation that is capable of exposing instruction-level information. We illustrate the effectiveness of Recon on programs as diverse as BerkeleyDB, i3/Chord, RandTree, and Pastry. Our evaluation includes executions in local clusters as well as in Amazon EC2 and exhibits an unreported bug in RandTree.
Kyu Hyung Lee, William N. Sumner, Xiangyu Zhang 0001, Patrick Eugster
DSN1
2011 Toward generating reducible replay logs
abstract
Logging and replay is important to reproducing software failures and recovering from failures. Replaying a long execution is time consuming, especially when replay is further integrated with runtime techniques that require expensive instrumentation, such as dependence detection. In this paper, we propose a technique to reduce a replay log while retaining its ability to reproduce a failure. While traditional logging records only system calls and signals, our technique leverages the compiler to selectively collect additional information on the fly. Upon a failure, the log can be reduced by analyzing itself. The collection is highly optimized. The additional runtime overhead of our technique, compared to a plain logging tool, is trivial (2.61% average) and the size of additional log is comparable to the original log. Substantial reduction can be cost-effectively achieved through a search based algorithm. The reduced log is guaranteed to reproduce the failure.
Kyu Hyung Lee, Yunhui Zheng, William N. Sumner, Xiangyu Zhang 0001
PLDI1