VLDB 2026 Research / reviewers in the wild / expert
Junghwan Rhee
dblp:27/5932 · also Junghwan John Rhee
· DBLP profile ↗
57ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-4043-9371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 31 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 14 · 8 since 2021Systems, architecture and hardware · 7 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring Attack Observability in Cloud Telemetry Logs: A Cross-Platform Analysis
Mary Grace Dhooghe, Minkyung Park, Junghwan Rhee, Yung Ryn Choe |
DSN | 3 |
| 2025 | Few-Shot Learning-Based Cyber Incident Detection with Augmented Context IntelligenceabstractIn recent years, the adoption of cloud services has been expanding at an unprecedented rate. As more and more organizations migrate or deploy their businesses to the cloud, a multitude of related cybersecurity incidents such as data breaches are on the rise. Many inherent attributes of cloud environments, for example, data sharing, remote access, dynamicity and scalability, pose significant challenges for the protection of cloud security. Even worse, cyber threats are becoming increasingly sophisticated and covert. Attack methods, such as Advanced Persistent Threats (APTs), are continually developed to bypass traditional security measures. Among the emerging technologies for robust threat detection, system provenance analysis is being considered as a promising mechanism, thus attracting widespread attention in the field of incident response. This paper proposes a new few-shot learning-based attack detection with improved data context intelligence. We collect operating system behavior data of cloud systems during realistic attacks and leverage an innovative semiotics extraction method to describe system events. Inspired by the advances in semantic analysis, which is a fruitful area focused on understanding natural languages in computational linguistics, we further convert the anomaly detection problem into a similarity comparison problem. Comprehensive experiments show that the proposed approach is able to generalize over unseen attacks and make accurate predictions, even if the incident detection models are trained with very limited samples. Fei Zuo, Junghwan Rhee, Yung Ryn Choe, Chenglong Fu 0002, Xianshan Qu |
COMPSAC | 2 |
| 2025 | IMUFuzzer: Resilience-based Discovery of Signal Injection Attacks on Robotic Aerial VehiclesabstractRobotic aerial vehicles (RAVs), particularly drones, are crucial in civil and military sectors. However, researchers have found that adversaries can inject noise into sensor measurements and cause physical impacts on the RAVs like crashes. Although identifying such signal injection attacks is essential to evaluate and improve the robustness of an RAV, it is challenging to discover them since their impact depends on the RAV’s physical states and the search space of noise signals and physical states is vast due to its dynamic nature.This paper proposes IMUFUZZER, a feedback-driven fuzzing framework, to automatically test an RAVs system and discover signal injection attacks. IMUFUZZER generates realistic noise signals for various inertial measurement unit (IMU) sensors, and monitors their impact on RAV control to detect mission failures, leveraging a high-fidelity RAV simulator. To find the physical states that attacks depend on, IMUFUZZER generates various mission paths that the RAV will fly through. We develop a novel feedback mechanism to quantify the resilience of the RAV against attacks and efficiently guide the fuzzing process to find signal injection attacks. Using IMUFUZZER, we have discovered 23 successful signal injection attacks on popular RAV control software (ArduPilot). We evaluate the correctness and effectiveness of our feedback-based sensor fuzzing and demonstrate the feasibility of the discovered attacks through physical experiments. Sudharssan Mohan, Kyeongseok Yang, Zelun Kong, Yonghwi Kwon 0001, Junghwan Rhee, Tyler Summers, Hongjun Choi, Heejo Lee |
ASE | 5 |
| 2025 | An Empirical Study on the Multi-Stage Nature of APT Attacks in Cloud ComputingabstractIn recent years, the adoption of cloud services has been expanding at an unprecedented rate. As more organizations migrate or deploy their businesses to the cloud, a multitude of related cybersecurity incidents, such as data breaches, are on the rise. Several inherent attributes of cloud environments, including data sharing, remote access, dynamic scalability, and scalability, pose significant challenges for the protection of cloud security. Even more concerning is the growing threat of Advanced Persistent Threats (APTs), which have become increasingly sophisticated and stealthy. As a more complex form of multi-stage attacks (MSAs), APTs follow a multi-step process that spreads malicious actions across different stages and blends them with legitimate operations, making intrusion detection particularly challenging. In this paper, we conduct an empirical study on the multi-stage characteristics of APT attacks specifically within cloud environments. Drawing from real-world attack scenarios, we analyze the behavioral patterns of APT attacks, evaluate the existing countermeasures, and identify ongoing challenges in defending against them. Our findings expose the limitations of conventional intrusion detection approaches and highlight the critical need for fine-grained and behavior-aware security mechanisms specifically designed for cloud environments. This study provides a deeper understanding of how APTs adapt their tactics to exploit cloud-specific vulnerabilities and offers insights for improving threat detection and response in modern cloud infrastructures. Fei Zuo, Junghwan Rhee, Shuaibing Lu, Yuqi Song |
MASS | 2 |
| 2024 | ProvIoT : Detecting Stealthy Attacks in IoT through Federated Edge-Cloud Security
Kunal Mukherjee, Josh Wiedemeier, Qi Wang 0017, Junpei Kamimura, Junghwan Rhee, James Wei, Zhichun Li, Xiao Yu 0007, Lu-An Tang, Jiaping Gui, Kangkook Jee |
ACNS (3) | 5 |
| 2024 | Differential Fuzzing for Data Distribution Service Programs with Dynamic ConfigurationabstractData Distribution Service (DDS) is a distributed network protocol widely used in cyber-physical systems. DDS provides flexible configurations defined in the formal design specification for safety and security. However, DDS programs suffer from both semantic bugs violating design specifications and software implementation bugs. To discover bugs, network protocol fuzzers have focused on testing client-server models by mutating input packets. However, they are unsuitable for fuzzing DDS programs due to a lack of consideration of the DDS-specific features, such as the DDS-specific input spaces (e.g., dynamic network topology formation and QoS and DDS security configurations) and impacts of DDS-specific semantic bugs (e.g., incorrect topology construction). Dohyun Ryu, Giyeol Kim, Seungjin Bae, Junghwan Rhee, Taegyu Kim |
ASE | 6 |
| 2024 | BinSimDB: Benchmark Dataset Construction for Fine-Grained Binary Code Similarity Analysis
Fei Zuo, Cody Tompkins, Qiang Zeng 0001, Lannan Luo, Yung Ryn Choe, Junghwan Rhee |
SecureComm (3) | 6 |
| 2024 | Automatic Configurator to Prevent Attacks for Azure Cloud SystemabstractCloud systems are integral for delivering scalable and virtualized resources globally. It also provides security updates and monitoring to keep user data safe. However, the growing complexity of these systems poses significant challenges, particularly in the realm of logging and security. It is difficult to know for users which detail is critical for further security analysis of the resources. Also, external packages used in the cloud system require updates by users to mitigate the vulnerability, but the large number of packages to manage makes them outdated versions. This paper shares the weakness of cloud logging systems we observed, which can be exploited by attackers. We propose a tool that configures alerts automatically when commands that have missing details in logs are executed and updates vulnerable versions of packages. Our tool leverages a list that includes the commands with missing details in logs and packages that need to be updated because of the known vulnerabilities. To make the list, we conduct complete enumerating for 1,279 commands in five major resources of Azure to find logs with missing details and search related communities to find vulnerable packages that require the manual update. We evaluate the proposed tool with eight attack scenarios based on real-world cases and the result shows that our tool prevents them successfully. Chijung Jung, Yung Ryn Choe, Junghwan Rhee, Yonghwi Kwon 0001 |
SERA | 3 |
| 2024 | Context Matters: Investigating Its Impact on ChatGPT's Bug Fixing PerformanceabstractIn this study, we explore the role of contextual information in enhancing ChatGPT's capabilities in bug fixing. Our focus is specifically on the “Wrong Answer” problem, where a program executes without error but fails to produce the correct output. Our approach draws inspiration from human debugging practices, which heavily rely on understanding both the intended task of the program and the specific scenarios in which it fails, such as unit test cases. We evaluate ChatGPT's performance with various types and levels of contextual data. The results reveal three key insights. First, providing the model with a mix of correct and incorrect test cases sharpens its debugging skills. Second, giving ChatGPT detailed descriptions of the problems substantially enhances its ability to identify and resolve errors. Third, merging detailed problem descriptions with various test cases leads to a synergistic outcome. This combined approach significantly elevates the efficiency of the bug-fixing process compared to employing each type of contextual information individually. Our paper presents a thorough analysis based on these findings. It offers an extensive exploration of why and how contextual information can be strategically utilized to enhance ChatGPT's debugging effectiveness. Furthermore, this investigation enriches our comprehension of the underlying mechanisms by which contextual cues amplify the model's capacity for solving problems. Xianshan Qu, Fei Zuo, Xiaopeng Li 0001, Junghwan Rhee |
SERA | 4 |
| 2023 | ProvSec: Cybersecurity System Provenance Analysis Benchmark DatasetabstractSystem provenance forensic analysis has been studied by a large body of research work. This area needs fine granularity data such as system calls along with event fields to track the dependencies of events. While prior work on security datasets has been proposed, we found a useful dataset of realistic attacks and details that can be used for provenance tracking is lacking. We created a new dataset of eleven vulnerable cases for system forensic analysis. It includes the full details of system calls including syscall parameters. Realistic attack scenarios with real software vulnerabilities and exploits are used. Also, we created two sets of benign and adversary scenarios which are manually labeled for supervised machine-learning analysis. We demonstrate the details of the dataset events and dependency analysis. Madhukar Shrestha, Jeehyun Oh, Junghwan Rhee, Yung Ryn Choe, Fei Zuo, Myung-Ah Park, Gang Qian |
SERA | 4 |
| 2023 | PowerGrader: Automating Code Assessment Based on PowerShell for Programming CoursesabstractProgramming courses in colleges often involve a myriad of coding assignments, which brings heavy grading workloads for instructors. To alleviate this problem, automatic programming evaluation tools are becoming more of a requirement than an option. However, after considering the actual requirements in our teaching practice, we have noticed that the current solutions still suffer from shortcomings and limitations. In the process of addressing the challenges, we propose and implement a brand new code assessment application based on PowerShell, which shows both extendibility and configurability. In particular, we integrate both black-box testing and the lexical analysis into the system, thus achieving a customized solution to meet specific requirements. This paper presents the architecture and design of our automatic code assessment application. Furthermore, we conduct empirical evaluations on the proposed system following the Technology Acceptance Model, and also investigate the drawbacks of manual assessment of coding assignments in terms of reliability and fairness. Finally, the evaluations demonstrate the effectiveness of our proposed auto-grader in facilitating the code assessment targeting college-level programming courses. Fei Zuo, Junghwan Rhee, Myung-Ah Park, Gang Qian |
SERA | 2 |
| 2023 | Commit Message Can Help: Security Patch Detection in Open Source Software via TransformerabstractAs open source software is widely used, the vulnerabilities contained therein are also rapidly propagated to a large number of innocent applications. Even worse, many vulnerabilities in open-source projects are secretly fixed, which leads to affected software being unaware and thus exposed to risks. For the purpose of protecting deployed software, designing an effective patch classification system becomes more of a need than an option. To this end, some researchers take advantage of the recent advancements in natural language processing to learn both commit messages and code changes. However, they often incur high false positive rates. Not only that, existing works cannot yet answer how much the textual description (such as commit messages) alone can influence the final triage. In this paper, we propose a Transformer based patch classifier, which does not use any code changes as inputs. Surprisingly, the extensive experiment shows the proposed approach can significantly outperform other state-of-the-art work with a high precision of 93.0% and low false positive rate. Therefore, our research further confirms the critical importance of well-crafted commit messages for the later software maintenance. Finally, our case study also identifies 48 silent security patches, which can benefit those affected software. Fei Zuo, Yuqi Song, Junghwan Rhee, Jicheng Fu |
SERA | 4 |
| 2022 | ShadowAuth: Backward-Compatible Automatic CAN Authentication for Legacy ECUsabstractController Area Network (CAN) is the de-facto standard in-vehicle network system. Despite its wide adoption by automobile manufacturers, the lack of security design makes it vulnerable to attacks. For instance, broadcasting packets without authentication allows the impersonation of electronic control units (ECUs). Prior mitigations, such as message authentication or intrusion detection systems, fail to address the compatibility requirement with legacy ECUs, stealthy and sporadic malicious messaging, or guaranteed attack detection. We propose a novel authentication system called ShadowAuth that overcomes the aforementioned challenges by offering backwardcompatible packet authentication to ECUs without requiring ECU firmware source code. Specifically, our authentication scheme provides transparent CAN packet authentication without modifying existing CAN packet definitions (e.g., J1939) via automatic ECU firmware instrumentation technique to locate CAN packet transmission code, and instrument authentication code based on the CAN packet behavioral transmission patterns. ShadowAuth enables vehicles to detect state-of-the-art CAN attacks, such as busoff and packet injection, responsively within 60ms without false positives. ShadowAuth provides a sound and deployable solution for real-world ECUs. Sungwoo Kim 0005, Gisu Yeo, Taegyu Kim, Junghwan Rhee, Yuseok Jeon, Antonio Bianchi, Dongyan Xu, Jing (Dave) Tian |
AsiaCCS | 4 |
| 2022 | DriveFuzz: Discovering Autonomous Driving Bugs through Driving Quality-Guided FuzzingabstractAutonomous driving has become real; semi-autonomous driving vehicles in an affordable price range are already on the streets, and major automotive vendors are actively developing full self-driving systems to deploy them in this decade. Before rolling the products out to the end-users, it is critical to test and ensure the safety of the autonomous driving systems, consisting of multiple layers intertwined in a complicated way. However, while safety-critical bugs may exist in any layer and even across layers, relatively little attention has been given to testing the entire driving system across all the layers. Prior work mainly focuses on white-box testing of individual layers and preventing attacks on each layer. Seulbae Kim, Major Liu, Junghwan Rhee, Yuseok Jeon, Yonghwi Kwon 0001 |
CCS | 3 |
| 2022 | Multi-source Inductive Knowledge Graph Transfer
Junheng Hao, Lu-An Tang, Yizhou Sun, Zhengzhang Chen, Junghwan Rhee, Zhichuan Li, Wei Wang 0010 |
ECML/PKDD (2) | 6 |
| 2021 | UTrack: Enterprise User Tracking Based on OS-Level Audit LogsabstractTracking user activities inside an enterprise network has been a fundamental building block for today's security infrastructure, as it provides accurate user profiling and helps security auditors to make informed decisions based on the derived insights from the abundant log data. Towards more accurate user tracking, we propose a novel paradigm named UTrack by leveraging rich system-level audit logs. From a holistic perspective, we bridge the semantic gap between user accounts and real users, tracking a real user's activities across different user accounts and different network hosts based on causal relationship among processes. To achieve better scalability and a more salient view, we apply a variety of data reduction and compression techniques to process the large amount of data. %and significantly reduce the data volume. We implement UTrack in a real enterprise environment consisting of 111 hosts, which generate more than 4 billion events in total during the experiment time of one month. Through our evaluation, we demonstrate that UTrack is able to accurately identify the events that are relevant to user activities. Our data reduction and compression modules largely reduce the output data size, producing a both accurate and salient overview on a user session profile. Yue Li 0002, Zhenyu Wu 0003, Haining Wang 0001, Kun Sun 0001, Zhichun Li, Kangkook Jee, Junghwan Rhee |
CODASPY | 7 |
| 2021 | Find My Sloths: Automated Comparative Analysis of How Real Enterprise Computers Keep Up with the Software Update Races
Omid Setayeshfar, Junghwan Rhee, Kyu Hyung Lee |
DIMVA | 2 |
| 2021 | SIGL: Securing Software Installations Through Deep Graph Learning
Xueyuan Han, Xiao Yu 0007, Thomas Pasquier, Ding Li 0001, Junghwan Rhee, James W. Mickens, Margo I. Seltzer |
USENIX Security Symposium | 5 |
| 2021 | PASAN: Detecting Peripheral Access Concurrency Bugs within Bare-Metal Embedded Applications
Taegyu Kim, Vireshwar Kumar, Junghwan Rhee, Jizhou Chen, Kyungtae Kim, Dongyan Xu, Jing (Dave) Tian |
USENIX Security Symposium | 3 |
| 2020 | This is Why We Can't Cache Nice Things: Lightning-Fast Threat Hunting using Suspicion-Based Hierarchical StorageabstractRecent advances in the causal analysis can accelerate incident response time, but only after a causal graph of the attack has been constructed. Unfortunately, existing causal graph generation techniques are mainly offline and may take hours or days to respond to investigator queries, creating greater opportunity for attackers to hide their attack footprint, gain persistency, and propagate to other machines. To address that limitation, we present Swift, a threat investigation system that provides high-throughput causality tracking and real-time causal graph generation capabilities. We design an in-memory graph database that enables space-efficient graph storage and online causality tracking with minimal disk operations. We propose a hierarchical storage system that keeps forensically-relevant part of the causal graph in main memory while evicting rest to disk. To identify the causal graph that is likely to be relevant during the investigation, we design an asynchronous cache eviction policy that calculates the most suspicious part of the causal graph and caches only that part in the main memory. We evaluated Swift on a real-world enterprise to demonstrate how our system scales to process typical event loads and how it responds to forensic queries when security alerts occur. Results show that Swift is scalable, modular, and answers forensic queries in real-time even when analyzing audit logs containing tens of millions of events. Wajih Ul Hassan, Ding Li 0001, Kangkook Jee, Xiao Yu 0007, Kexuan Zou, Zhengzhang Chen, Zhichun Li, Junghwan Rhee, Jiaping Gui, Adam Bates 0001 |
ACSAC | 9 |
| 2020 | Vessels: efficient and scalable deep learning prediction on trusted processorsabstractDeep learning systems on the cloud are increasingly targeted by attacks that attempt to steal sensitive data. Intel SGX has been proven effective to protect the confidentiality and integrity of such data during computation. However, state-of-the-art SGX systems still suffer from substantial performance overhead induced by the limited physical memory of SGX. This limitation significantly undermines the usability of deep learning systems due to their memory-intensive characteristics. Kyungtae Kim, Junghwan Rhee, Xiao Yu 0007, Jing (Dave) Tian, Byoungyoung Lee |
SoCC | 3 |
| 2020 | Detecting Malware Injection with Program-DNS BehaviorabstractAnalyzing the DNS traffic of Internet hosts has been a successful technique to counter cyberattacks and identify connections to malicious domains. However, recent stealthy attacks hide malicious activities within seemingly legitimate connections to popular web services made by benign programs. Traditional DNS monitoring and signature-based detection techniques are ineffective against such attacks. To tackle this challenge, we present a new program-level approach that can effectively detect such stealthy attacks. Our method builds a fine-grained Program-DNS profile for each benign program that characterizes what should be the “expected” DNS behavior. We find that malware-injected processes have DNS activities which significantly deviate from the Program-DNS profile of the benign program. We then develop six novel features based on the Program-DNS profile, and evaluate the features on a dataset of over 130 million DNS requests collected from a real-world enterprise and 8 million requests from malware-samples executed in a sandbox environment. We compare our detection results with that of previously-proposed features and demonstrate that our new features successfully detect 190 malware-injected processes which fail to be detected by previously-proposed features. Overall, our study demonstrates that fine-grained Program-DNS profiles can provide meaningful and effective features in building detectors for attack campaigns that bypass existing detection systems. Yixin Sun 0004, Kangkook Jee, Suphannee Sivakorn, Zhichun Li, Cristian Lumezanu, Lauri Korts-Pärn, Zhenyu Wu 0003, Junghwan Rhee, Mung Chiang, Prateek Mittal |
EuroS&P | 8 |
| 2020 | APTrace: A Responsive System for Agile Enterprise Level Causality AnalysisabstractWhile backtracking analysis has been successful in assisting the investigation of complex security attacks, it faces a critical dependency explosion problem. To address this problem, security analysts currently need to tune backtracking analysis manually with different case-specific heuristics. However, existing systems fail to fulfill two important system requirements to achieve effective backtracking analysis. First, there need flexible abstractions to express various types of heuristics. Second, the system needs to be responsive in providing updates so that the progress of backtracking analysis can be frequently inspected, which typically involves multiple rounds of manual tuning. In this paper, we propose a novel system, APTrace, to meet both of the above requirements. As we demonstrate in the evaluation, security analysts can effectively express heuristics to reduce more than 99.5% of irrelevant events in the backtracking analysis of real-world attack cases. To improve the responsiveness of backtracking analysis, we present a novel execution-window partitioning algorithm that significantly reduces the waiting time between two consecutive updates (especially, 57 times reduction for the top 1% waiting time). Jiaping Gui, Ding Li 0001, Zhengzhang Chen, Junghwan Rhee, Xusheng Xiao, Mu Zhang 0001, Kangkook Jee, Zhichun Li |
ICDE | 4 |
| 2020 | You Are What You Do: Hunting Stealthy Malware via Data Provenance Analysis
Qi Wang 0017, Wajih Ul Hassan, Ding Li 0001, Kangkook Jee, Xiao Yu 0007, Kexuan Zou, Junghwan Rhee, Zhengzhang Chen, Wei Cheng 0002, Carl A. Gunter |
NDSS | 7 |
| 2020 | A Generic Edge-Empowered Graph Convolutional Network via Node-Edge Mutual EnhancementabstractGraph Convolutional Networks (GCNs) have shown to be a powerful tool for analyzing graph-structured data. Most of previous GCN methods focus on learning a good node representation by aggregating the representations of neighboring nodes, whereas largely ignoring the edge information. Although few recent methods have been proposed to integrate edge attributes into GCNs to initialize edge embeddings, these methods do not work when edge attributes are (partially) unavailable. Can we develop a generic edge-empowered framework to exploit node-edge enhancement, regardless of the availability of edge attributes? In this paper, we propose a novel framework EE-GCN that achieves node-edge enhancement. In particular, the framework EE-GCN includes three key components: (i) Initialization: this step is to initialize the embeddings of both nodes and edges. Unlike node embedding initialization, we propose a line graph-based method to initialize the embedding of edges regardless of edge attributes. (ii) Feature space alignment: we propose a translation-based mapping method to align edge embedding with node embedding space, and the objective function is penalized by a translation loss when both spaces are not aligned. (iii) Node-edge mutually enhanced updating: node embedding is updated by aggregating embedding of neighboring nodes and associated edges, while edge embedding is updated by the embedding of associated nodes and itself. Through the above improvements, our framework provides a generic strategy for all of the spatial-based GCNs to allow edges to participate in embedding computation and exploit node-edge mutual enhancement. Finally, we present extensive experimental results to validate the improved performances of our method in terms of node classification, link prediction, and graph classification. Pengyang Wang, Jiaping Gui, Zhengzhang Chen, Junghwan Rhee, Yanjie Fu |
WWW | 4 |
| 2020 | CAFE: A Virtualization-Based Approach to Protecting Sensitive Cloud Application Logic ConfidentialityabstractCloud application marketplaces of modern cloud infrastructures offer a new software deployment model, integrated with the cloud environment in its configuration and policies. However, similar to traditional software distribution which has been suffering from software piracy and reverse engineering, cloud marketplaces face the same challenges that can deter the success of the evolving ecosystem of cloud software. We present a novel system named CAFE for cloud infrastructures where sensitive software logic can be executed with high secrecy protected from any piracy or reverse engineering attempts in a virtual machine even when its operating system kernel is compromised. The key mechanism is the end-to-end framework for the execution of applications, which consists of the secure encryption and distribution of confidential application binary files, and the runtime techniques to load, decrypt, and protect the program logic by isolating them from tenant virtual machines based on hypervisor-level techniques. We evaluate applications in several software categories which are commonly offered in cloud marketplaces showing that strong confidential execution can be provided with only marginal changes (around 100-220 lines of code) and minimal performance overhead. The results demonstrate the effectiveness and practicality of CAFE in cloud marketplaces. Sungjin Park 0001, Junghwan Rhee, Jong-Jin Won, Taisook Han, Dongyan Xu |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2019 | PoLPer: Process-Aware Restriction of Over-Privileged Setuid Calls in Legacy ApplicationsabstractSetuid system calls enable critical functions such as user authentications and modular privileged components. Such operations must only be executed after careful validation. However, current systems do not perform rigorous checks, allowing exploitation of privileges through memory corruption vulnerabilities in privileged programs. As a solution, understanding which setuid system calls can be invoked in what context of a process allows precise enforcement of least privileges. We propose a novel comprehensive method to systematically extract and enforce least privilege of setuid system calls to prevent misuse. Our approach learns the required process contexts of setuid system calls along multiple dimensions: process hierarchy, call stack, and parameter in a process-aware way. Every setuid system call is then restricted to the per-process context by our kernel-level context enforcer. Previous approaches without process-awareness are too coarse-grained to control setuid system calls, resulting in over-privilege. Our method reduces available privileges even for identical code depending on whether it is run by a parent or a child process. We present our prototype called PoLPer which systematically discovers only required setuid system calls and effectively prevents real-world exploits targeting vulnerabilities of the setuid family of system calls in popular desktop and server software at near zero overhead. Yuseok Jeon, Junghwan Rhee, Zhichun Li, Mathias Payer, Byoungyoung Lee, Zhenyu Wu 0003 |
CODASPY | 2 |
| 2019 | HeapTherapy+: Efficient Handling of (Almost) All Heap Vulnerabilities Using Targeted Calling-Context EncodingabstractExploitation of heap vulnerabilities has been on the rise, leading to many devastating attacks. Conventional heap patch generation is a lengthy procedure requiring intensive manual efforts. Worse, fresh patches tend to harm system dependability, hence deterring users from deploying them. We propose a heap patching system HEAPTHERAPY+ that simultaneously has the following prominent advantages: (1) generating patches without manual efforts; (2) installing patches without altering the code (so called code-less patching); (3) handling various heap vulnerability types; (4) imposing a very low overhead; and (5) no dependency on specific heap allocators. As a separate contribution, we propose targeted calling context encoding, which is a suite of algorithms for optimizing calling context encoding, an important technique with applications in many areas. The system properly combines heavyweight offline attack analysis with lightweight online defense generation, and provides a new countermeasure against heap attacks. The evaluation shows that the system is effective and efficient. Qiang Zeng 0001, Golam Kayas, Emil Mohammed, Lannan Luo, Xiaojiang Du, Junghwan Rhee |
DSN | 6 |
| 2019 | Attentional Heterogeneous Graph Neural Network: Application to Program ReidentificationabstractProgram or process is an integral part of almost every IT/OT system. Can we trust the identity/ID (e.g., executable name) of the program? To avoid detection, malware may disguise itself using the ID of a legitimate program, and a system tool (e.g., PowerShell) used by the attackers may have the fake ID of another common software, which is less sensitive. However, existing intrusion detection techniques often overlook this critical program reidentification problem (i.e., checking the program's identity). In this paper, we propose an attentional heterogeneous graph neural network model (DeepHGNN) to verify the program's identity based on its system behaviors. The key idea is to leverage the representation learning of the heterogeneous program behavior graph to guide the reidentification process. We formulate the program reidentification as a graph classification problem and develop an effective attentional heterogeneous graph embedding algorithm to solve it. Extensive experiments — using real-world enterprise monitoring data and real attacks — demonstrate the effectiveness of DeepHGNN across multiple popular metrics and the robustness to the normal dynamic changes like program version upgrades. Shen Wang 0005, Zhengzhang Chen, Ding Li 0001, Zhichun Li, Lu-An Tang, Jingchao Ni, Junghwan Rhee, Philip S. Yu |
SDM | 7 |
| 2019 | RVFuzzer: Finding Input Validation Bugs in Robotic Vehicles through Control-Guided Testing
Taegyu Kim, Junghwan Rhee, Fei Fan 0002, Zhan Tu, Gregory Walkup, Xiangyu Zhang 0001, Dongyan Xu |
USENIX Security Symposium | 3 |
| 2018 | NodeMerge: Template Based Efficient Data Reduction For Big-Data Causality AnalysisabstractToday's enterprises are exposed to sophisticated attacks, such as Advanced Persistent Threats~(APT) attacks, which usually consist of stealthy multiple steps. To counter these attacks, enterprises often rely on causality analysis on the system activity data collected from a ubiquitous system monitoring to discover the initial penetration point, and from there identify previously unknown attack steps. However, one major challenge for causality analysis is that the ubiquitous system monitoring generates a colossal amount of data and hosting such a huge amount of data is prohibitively expensive. Thus, there is a strong demand for techniques that reduce the storage of data for causality analysis and yet preserve the quality of the causality analysis. To address this problem, in this paper, we propose NodeMerge, a template based data reduction system for online system event storage. Specifically, our approach can directly work on the stream of system dependency data and achieve data reduction on the read-only file events based on their access patterns. It can either reduce the storage cost or improve the performance of causality analysis under the same budget. Only with a reasonable amount of resource for online data reduction, it nearly completely preserves the accuracy for causality analysis. The reduced form of data can be used directly with little overhead. To evaluate our approach, we conducted a set of comprehensive evaluations, which show that for different categories of workloads, our system can reduce the storage capacity of raw system dependency data by as high as 75.7 times, and the storage capacity of the state-of-the-art approach by as high as 32.6 times. Furthermore, the results also demonstrate that our approach keeps all the causality analysis information and has a reasonably small overhead in memory and hard disk. Yutao Tang, Ding Li 0001, Zhichun Li, Mu Zhang 0001, Kangkook Jee, Xusheng Xiao, Zhenyu Wu 0003, Junghwan Rhee, Fengyuan Xu, Qun Li 0001 |
CCS | 8 |
| 2018 | Towards a Timely Causality Analysis for Enterprise Security
Yushan Liu 0004, Mu Zhang 0001, Ding Li 0001, Kangkook Jee, Zhichun Li, Zhenyu Wu 0003, Junghwan Rhee, Prateek Mittal |
NDSS | 7 |
| 2017 | A Hypervisor Level Provenance System to Reconstruct Attack Story Caused by Kernel Malware
Chonghua Wang, Shiqing Ma, Xiangyu Zhang 0001, Junghwan Rhee, Xiao-chun Yun, Zhiyu Hao |
SecureComm | 4 |
| 2016 | High Fidelity Data Reduction for Big Data Security Dependency AnalysesabstractIntrusive multi-step attacks, such as Advanced Persistent Threat (APT) attacks, have plagued enterprises with significant financial losses and are the top reason for enterprises to increase their security budgets. Since these attacks are sophisticated and stealthy, they can remain undetected for years if individual steps are buried in background "noise." Thus, enterprises are seeking solutions to "connect the suspicious dots" across multiple activities. This requires ubiquitous system auditing for long periods of time, which in turn causes overwhelmingly large amount of system audit events. Given a limited system budget, how to efficiently handle ever-increasing system audit logs is a great challenge. This paper proposes a new approach that exploits the dependency among system events to reduce the number of log entries while still supporting high-quality forensic analysis. In particular, we first propose an aggregation algorithm that preserves the dependency of events during data reduction to ensure the high quality of forensic analysis. Then we propose an aggressive reduction algorithm and exploit domain knowledge for further data reduction. To validate the efficacy of our proposed approach, we conduct a comprehensive evaluation on real-world auditing systems using log traces of more than one month. Our evaluation results demonstrate that our approach can significantly reduce the size of system logs and improve the efficiency of forensic analysis without losing accuracy. Zhang Xu, Zhenyu Wu 0003, Zhichun Li, Kangkook Jee, Junghwan Rhee, Xusheng Xiao, Fengyuan Xu, Haining Wang 0001, Guofei Jiang |
CCS | 5 |
| 2016 | Detecting Stack Layout Corruptions with Robust Stack Unwinding
Yangchun Fu, Junghwan Rhee, Zhiqiang Lin 0001, Zhichun Li, Hui Zhang 0002, Guofei Jiang |
RAID | 2 |
| 2016 | PerfGuard: binary-centric application performance monitoring in production environmentsabstractDiagnosis of performance problems is an essential part of software development and maintenance. This is in particular a challenging problem to be solved in the production environment where only program binaries are available with limited or zero knowledge of the source code. This problem is compounded by the integration with a significant number of third-party software in most large-scale applications. Existing approaches either require source code to embed manually constructed logic to identify performance problems or support a limited scope of applications with prior manual analysis. This paper proposes an automated approach to analyze application binaries and instrument the binary code transparently to inject and apply performance assertions on application transactions. Our evaluation with a set of large-scale application binaries without access to source code discovered 10 publicly known real world performance bugs automatically and shows that PerfGuard introduces very low overhead (less than 3% on Apache and MySQL server) to production systems. Junghwan Rhee, Kyu Hyung Lee, Xiangyu Zhang 0001, Dongyan Xu |
SIGSOFT FSE | 2 |
| 2015 | Accurate, Low Cost and Instrumentation-Free Security Audit Logging for WindowsabstractAudit logging is an important approach to cyber attack investigation. However, traditional audit logging either lacks accuracy or requires expensive and complex binary instrumentation. In this paper, we propose a Windows based audit logging technique that features accuracy and low cost. More importantly, it does not require instrumenting the applications, which is critical for commercial software with IP protection. The technique is build on Event Tracing for Windows (ETW). By analyzing ETW log and critical parts of application executables, a model can be constructed to parse ETW log to units representing independent sub-executions in a process. Causality inferred at the unit level renders much higher accuracy, allowing us to perform accurate attack investigation and highly effective log reduction. Shiqing Ma, Kyu Hyung Lee, Junghwan Rhee, Xiangyu Zhang 0001, Dongyan Xu |
ACSAC | 4 |
| 2015 | CAFE: A Virtualization-Based Approach to Protecting Sensitive Cloud Application Logic ConfidentialityabstractCloud application marketplaces of modern cloud infrastructures offer a new software deployment model, integrated with the cloud environment in its configuration and policies. However, similar to traditional software distribution which has been suffering from software piracy and reverse engineering, cloud marketplaces face the same challenges that can deter the success of the evolving ecosystem of cloud software. We present a novel system named CAFE for cloud infrastructures where sensitive software logic can be executed with high secrecy protected from any piracy or reverse engineering attempts in a virtual machine even when its operating system kernel is compromised. The key mechanism is the end-to-end framework for the execution of applications, which consists of the secure encryption and distribution of confidential application binary files, and the runtime techniques to load, decrypt, and protect the program logic by isolating them from tenant virtual machines based on hypervisor-level techniques. We evaluate applications in several software categories which are commonly offered in cloud marketplaces showing that strong confidential execution can be provided with only marginal changes (around 100-220 lines of code) and minimal performance overhead. Sungjin Park 0001, Junghwan Rhee, Jong-Jin Won, Taisook Han, Dongyan Xu |
AsiaCCS | 3 |
| 2015 | Discover and Tame Long-running Idling Processes in Enterprise SystemsabstractReducing attack surface is an effective preventive measure to strengthen security in large systems. However, it is challenging to apply this idea in an enterprise environment where systems are complex and evolving over time. In this paper, we empirically analyze and measure a real enterprise to identify unused services that expose attack surface. Interestingly, such unused services are known to exist and summarized by security best practices, yet such solutions require significant manual effort. Jun Wang 0141, Zhiyun Qian, Zhichun Li, Zhenyu Wu 0003, Junghwan Rhee, Xia Ning, Peng Liu 0005, Guofei Jiang |
AsiaCCS | 5 |
| 2014 | DeltaPath: Precise and Scalable Calling Context Encoding
Qiang Zeng 0001, Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Peng Liu 0005 |
CGO | 2 |
| 2014 | PerfScope: Practical Online Server Performance Bug Inference in Production Cloud Computing InfrastructuresabstractPerformance bugs which manifest in a production cloud computing infrastructure are notoriously difficult to diagnose because of both the difficulty of reproducing those bugs and the lack of debugging information. In this paper, we present PerfScope, a practical online performance bug inference tool to help the developer understand how a performance bug happened during the production run. PerfScope achieves online bug inference to obviate the need for offline bug reproduction. PerfScope does not require application source code or any runtime instrumentation to the production system. PerfScope is application-agnostic, which can support both interpreted and compiled programs running inside a cloud infrastructure. Daniel Joseph Dean, Hiep Nguyen, Xiaohui Gu, Hui Zhang 0002, Junghwan Rhee, Nipun Arora, Geoff Jiang |
SoCC | 5 |
| 2014 | Software system performance debugging with kernel events feature guidanceabstractTo diagnose performance problems in production systems, many OS kernel-level monitoring and analysis tools have been proposed. Using low level kernel events provides benefits in efficiency and transparency to monitor application software. On the other hand, such approaches miss application-specific semantic information which can be effective to differentiate the trace patterns from distinct application logic. This paper introduces new trace analysis techniques based on event features to improve kernel event based performance diagnosis tools. Our prototype, AppDiff, is based on two analysis features: system resource features convert kernel events to resource usage metrics, thereby enabling the detection of various performance anomalies in a unified way; program behavior features infer the application logic behind the low level events. By using these features and conditional probability, AppDiff can detect outliers and improve the diagnosis of application performance. Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Kenji Yoshihira |
NOMS | 1 |
| 2014 | Uscope: A scalable unified tracer from kernel to user spaceabstractUnified tracing is the process of collecting trace logs across the boundary of kernel and user spaces, and has been used to understand the in-depth correspondence between low level events and application program context for diagnosing system failures and performance problems. Crossing the boundary from the kernel space to a user space to collect trace events from dual spaces imposes challenges compared to crossing the boundary in the other way from a user space to the kernel space due to multiple scheduled programs and diverse code layouts in the user space regarding the tracing target. In this paper, we propose a novel unified tracing system called Uscope to systematically trace kernel and unprecedented user code with low overhead. The key idea is to use an efficient variant of stack walking. Uscope lowers stack walking overhead by adjusting the scope of walking in two ways: (1) a highly configurable focus within the call stack, and (2) a per-application tracing that systematically tracks a dynamic set of new, exiting, or transforming processes and threads of an application software. This system is realized by using a flexible stack walking algorithm and a runtime kernel structure, Trace Map. These key features lead to low run-time overhead under 6% relative to native execution on a set of widely used benchmarks. Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Kenji Yoshihira |
NOMS | 1 |
| 2014 | CLUE: System trace analytics for cloud service performance diagnosisabstractIn this paper, we present CLUE, a system event analytics tool for black-box performance diagnosis in production Cloud Computing systems. CLUE provides an unified and extensible means of profiling service transactional behaviors, and builds structured data called event sketches. CLUE further offers a set of analytic tools for summarizing and analyzing event sketches by integrating data mining and statistical analysis. CLUE has been developed in NEC as an internal tool and applied in diagnosing a diverse set of real performance problems for multi-tiered IT applications running on multi-core servers of major platforms including Linux (Redhat, Fedora), Unix (HP-UX), and Windows (Windows Server 2008). We demonstrated the evaluation of our framework on real-world IT systems, and showed how it can enable visibility and effective diagnosis of service system performance problems. Hui Zhang 0002, Junghwan Rhee, Nipun Arora, Sahan Gamage, Guofei Jiang, Kenji Yoshihira, Dongyan Xu |
NOMS | 2 |
| 2014 | IntroPerf: transparent context-sensitive multi-layer performance inference using system stack tracesabstractPerformance bugs are frequently observed in commodity software. While profilers or source code-based tools can be used at development stage where a program is diagnosed in a well-defined environment, many performance bugs survive such a stage and affect production runs. OS kernel-level tracers are commonly used in post-development diagnosis due to their independence from programs and libraries; however, they lack detailed program-specific metrics to reason about performance problems such as function latencies and program contexts. In this paper, we propose a novel performance inference system, called IntroPerf, that generates fine-grained performance information -- like that from application profiling tools -- transparently by leveraging OS tracers that are widely available in most commodity operating systems. With system stack traces as input, IntroPerf enables transparent context-sensitive performance inference, and diagnoses application performance in a multi-layered scope ranging from user functions to the kernel. Evaluated with various performance bugs in multiple open source software projects, IntroPerf automatically ranks potential internal and external root causes of performance bugs with high accuracy without any prior knowledge about or instrumentation on the subject software. Our results show IntroPerf's effectiveness as a lightweight performance introspection tool for post-development diagnosis. Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Xiangyu Zhang 0001, Dongyan Xu |
SIGMETRICS | 2 |
| 2014 | Data-Centric OS Kernel Malware CharacterizationabstractTraditional malware detection and analysis approaches have been focusing on code-centric aspects of malicious programs, such as detection of the injection of malicious code or matching malicious code sequences. However, modern malware has been employing advanced strategies, such as reusing legitimate code or obfuscating malware code to circumvent the detection. As a new perspective to complement code-centric approaches, we propose a data-centric OS kernel malware characterization architecture that detects and characterizes malware attacks based on the properties of data objects manipulated during the attacks. This framework consists of two system components with novel features: First, a runtime kernel object mapping system which has an un-tampered view of kernel data objects resistant to manipulation by malware. This view is effective at detecting a class of malware that hides dynamic data objects. Second, this framework consists of a new kernel malware detection approach that generates malware signatures based on the data access patterns specific to malware attacks. This approach has an extended coverage that detects not only the malware with the signatures, but also the malware variants that share the attack patterns by modeling the low level data access behaviors as signatures. Our experiments against a variety of real-world kernel rootkits demonstrate the effectiveness of data-centric malware signatures. Junghwan Rhee, Ryan D. Riley, Zhiqiang Lin 0001, Xuxian Jiang, Dongyan Xu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | iProbe: A lightweight user-level dynamic instrumentation toolabstractWe introduce a new hybrid instrumentation tool for dynamic application instrumentation called iProbe, which is flexible and has low overhead. iProbe takes a novel 2-stage design, and offloads much of the dynamic instrumentation complexity to an offline compilation stage. It leverages standard compiler flags to introduce “place-holders” for hooks in the program executable. Then it utilizes an efficient user-space “HotPatching” mechanism which modifies the functions to be traced and enables execution of instrumented code in a safe and secure manner. In its evaluation on a micro-benchmark and SPEC CPU2006 benchmark applications, the iProbe prototype achieved the instrumentation overhead an order of magnitude lower than existing state-of-the-art dynamic instrumentation tools like SystemTap and DynInst. Nipun Arora, Hui Zhang 0002, Junghwan Rhee, Kenji Yoshihira, Guofei Jiang |
ASE | 3 |
| 2012 | Discovering Semantic Data of Interest from Un-mappable Memory with Confidence
Zhiqiang Lin 0001, Junghwan Rhee, Xiangyu Zhang 0001, Dongyan Xu |
NDSS | 2 |
| 2011 | Characterizing kernel malware behavior with kernel data access patternsabstractCharacterizing malware behavior using its control flow faces several challenges, such as obfuscations in static analysis and the behavior variations in dynamic analysis. This paper introduces a new approach to characterizing kernel malware's behavior by using kernel data access patterns unique to the malware. The approach neither uses malware's control flow consisting of temporal ordering of malware code execution, nor the code-specific information about the malware. Thus, the malware signature based on such data access patterns is resilient in matching malware variants.To evaluate the effectiveness of this approach, we first generated the signatures of three classic rootkits using their data access patterns, and then matched them with a group of kernel execution instances which are benign or compromised by 16 kernel rootkits. The malware signatures did not trigger any false positives in benign kernel runs; however, kernel runs compromised by 16 rootkits were detected due to the data access patterns shared with the compared signature(s). We further observed similar data access patterns in the signatures of the tested rootkits and exposed popular rootkit attack operations by ranking common data behavior across rootkits. Our experiments show that our approach is effective not only to detect the malware whose signature is available, but also to determine its variants which share kernel data access patterns. Junghwan Rhee, Zhiqiang Lin 0001, Dongyan Xu |
AsiaCCS | 1 |
| 2011 | SigGraph: Brute Force Scanning of Kernel Data Structure Instances Using Graph-based Signatures
Zhiqiang Lin 0001, Junghwan Rhee, Xiangyu Zhang 0001, Dongyan Xu, Xuxian Jiang |
NDSS | 2 |
| 2010 | Kernel Malware Analysis with Un-tampered and Temporal Views of Dynamic Kernel Memory
Junghwan Rhee, Ryan D. Riley, Dongyan Xu, Xuxian Jiang |
RAID | 1 |
| 2010 | DKSM: Subverting Virtual Machine Introspection for Fun and ProfitabstractVirtual machine (VM) introspection is a powerful technique for determining the specific aspects of guest VM execution from outside the VM. Unfortunately, existing introspection solutions share a common questionable assumption. This assumption is embodied in the expectation that original kernel data structures are respected by the untrusted guest and thus can be directly used to bridge the well-known semantic gap. In this paper, we assume the perspective of the attacker, and exploit this questionable assumption to subvert VM introspection. In particular, we present an attack called DKSM (Direct Kernel Structure Manipulation), and show that it can effectively foil existing VM introspection solutions into providing false information. By assuming this perspective, we hope to better understand the challenges and opportunities for the development of future reliable VM introspection solutions that are not vulnerable to the proposed attack. Sina Bahram, Xuxian Jiang, Zhi Wang 0004, Mike Grace, Jinku Li, Deepa Srinivasan, Junghwan Rhee, Dongyan Xu |
SRDS | 7 |
| 2009 | Defeating Dynamic Data Kernel Rootkit Attacks via VMM-Based Guest-Transparent MonitoringabstractTargeting the operating system kernel, the core of trust in a system, kernel rootkits are able to compromise the entire system, placing it under malicious control, while eluding detection efforts. Within the realm of kernel rootkits, dynamic data rootkits are particularly elusive due to the fact that they attack only data targets. Dynamic data rootkits avoid code injection and instead use existing kernel code to manipulate kernel data. Because they do not execute any new code, they are able to complete their attacks without violating kernel code integrity. We propose a prevention solution that blocks dynamic data kernel rootkit attacks by monitoring kernel memory access using virtual machine monitor (VMM) policies. Although the VMM is an external monitor, our system preemptively detects changes to monitored kernel data states and enables fine-grained inspection of memory accesses on dynamically changing kernel data. In addition, readable and writable kernel data can be protected by exposing the illegal use of existing code by dynamic data kernel rootkits. We have implemented a prototype of our system using the QEMU VMM. Our experiments show that it successfully defeats synthesized dynamic data kernel rootkits in real-time, demonstrating its effectiveness and practicality. Junghwan Rhee, Ryan D. Riley, Dongyan Xu, Xuxian Jiang |
ARES | 1 |
| 2009 | DeskBench: Flexible virtual desktop benchmarking toolkitabstractThe thin-client computing model has been recently regaining popularity in a new form known as the virtual desktop. That is where the desktop is hosted on a virtualized platform. Even though the interest in this computing paradigm is broad there are relatively few tools and methods for benchmarking virtual client infrastructures. We believe that developing such tools and approaches is crucial for the future success of virtual client deployments and also for objective evaluation of existing and new algorithms, communication protocols, and technologies. We present DeskBench, a virtual desktop benchmarking tool, that allows for fast and easy creation of benchmarks by simple recording of the user's activity. It also allows for replaying the recorded actions in a synchronized manner at maximum possible speeds without compromising the correctness of the replay. The proposed approach relies only on the basic primitives of mouse and keyboard events as well as screen region updates which are common in window manager systems. We have implemented a prototype of the system and also conducted a series of experiments measuring responsiveness of virtual machine based desktops under various load conditions and network latencies. The experiments illustrate the flexibility and accuracy of the proposed method and also give some interesting insights into the scalability of virtual machine based desktops. Junghwan Rhee, Andrzej Kochut, Kirk A. Beaty |
Integrated Network Management | 1 |
| 2006 | Autonomic Adaptation of Virtual Distributed Environments in a Multi-Domain InfrastructureabstractBy federating resources from multiple domains, a shared infrastructure provides aggregated computation resources to a large number of users. With rapid advances in virtualization technologies, we propose the concept of virtual distributed environments as a new sharing paradigm for a multi-domain shared infrastructure. Such virtual environments provide users with confined, customized platforms to execute legacy parallel/distributed applications. Furthermore, we propose to support autonomic adaptation of virtual distributed environments, driven by both dynamic availability of infrastructure resources and dynamic application resource demand. We identify new research challenges and describe our on-going work and preliminary results. 1 Dongyan Xu, Paul Ruth, Junghwan Rhee, Rick Kennell, Sebastien Goasguen |
HPDC | 3 |
| 2005 | TraceBack: first fault diagnosis by reconstruction of distributed control flowabstractFaults that occur in production systems are the most important faults to fix, but most production systems lack the debugging facilities present in development environments. TraceBack provides debugging information for production systems by providing execution history data about program problems (such as crashes, hangs, and exceptions). TraceBack supports features commonly found in production environments such as multiple threads, dynamically loaded modules, multiple source languages (e.g., Java applications running with JNI modules written in C++), and distributed execution across multiple computers. TraceBack supports first fault diagnosis-discovering what went wrong the first time a fault is encountered. The user can see how the program reached the fault state without having to re-run the computation; in effect enabling a limited form of a debugger in production code.TraceBack uses static, binary program analysis to inject low-overhead runtime instrumentation at control-flow block granularity. Post-facto reconstruction of the records written by the instrumentation code produces a source-statement trace for user diagnosis. The trace shows the dynamic instruction sequence leading up to the fault state, even when the program took exceptions or terminated abruptly (e.g., kill -9).We have implemented TraceBack on a variety of architectures and operating systems, and present examples from a variety of platforms. Performance overhead is variable, from 5% for Apache running SPECweb99, to 16%-25% for the Java SPECJbb benchmark, to 60% average for SPECint2000. We show examples of TraceBack's cross-language and cross-machine abilities, and report its use in diagnosing problems in production software. Andrew Ayers, Richard Schooler, Chris Metcalf, Anant Agarwal, Junghwan Rhee, Emmett Witchel |
PLDI | 5 |
| 2005 | Mondrix: memory isolation for linux using mondriaan memory protectionabstractThis paper presents the design and an evaluation of Mondrix, a version of the Linux kernel with Mondriaan Memory Protection (MMP). MMP is a combination of hardware and software that provides efficient fine-grained memory protection between multiple protection domains sharing a linear address space. Mondrix uses MMP to enforce isolation between kernel modules which helps detect bugs, limits their damage, and improves kernel robustness and maintainability. During development, MMP exposed two kernel bugs in common, heavily-tested code, and during fault injection experiments, it prevented three of five file system corruptions.The Mondrix implementation demonstrates how MMP can bring memory isolation to modules that already exist in a large software application. It shows the benefit of isolation for robustness and error detection and prevention, while validating previous claims that the protection abstractions MMP offers are a good fit for software. This paper describes the design of the memory supervisor, the kernel module which implements permissions policy.We present an evaluation of Mondrix using full-system simulation of large kernel-intensive workloads. Experiments with several benchmarks where MMP was used extensively indicate the additional space taken by the MMP data structures reduce the kernel's free memory by less than 10%, and the kernel's runtime increases less than 15% relative to an unmodified kernel. Emmett Witchel, Junghwan Rhee, Krste Asanovic |
SOSP | 2 |