VLDB 2026 Research / reviewers in the wild / expert
Youngjin Kwon
dblp:94/8239
· DBLP profile ↗
44ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 2 first-author · 15 since 2021Software engineering, systems software and programming languages · 16 · 4 first-author · 6 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CofferOS: Hardening OS-level Virtualization with RustabstractOS-level virtualization (e.g., Linux containers) has become a cornerstone of modern cloud systems. While it offers the illusion of isolated kernels for processes, these processes share the same underlying kernel, raising critical concerns around security, fault isolation, and the inability to customize kernels. Existing solutions address the issues by employing virtual machines that isolate kernels. However, these approaches incur significant performance overhead. Minkyu Jung, Chanshin Kwak, Junho Ahn, Sunho Park, Changjun Lee, Jongyul Kim 0001, Jeehoon Kang, Youngjin Kwon |
EuroSys | 8 |
| 2026 | BASK: Batch And SmartNIC-offloaded KSM
Chanshin Kwak, Jaehyeon Lee, Minkyu Jung, Changjun Lee, Youngjin Kwon |
EuroSys | 5 |
| 2026 | MTTM: Dynamic Fast Memory Partitioning with Bandwidth Optimization for Multi-tenant CloudabstractMemory tiering, which extends local DRAM by incorporating remote DRAM or NVM via Compute Express Link (CXL), offers a promising solution to the DRAM capacity scaling problem. However, in multi-tenant cloud environments, efficient management of fast memory requires not only optimizing data movement across tiers but also effectively distributing DRAM capacity when tenants compete. Moreover, bandwidth utilization is critically impacted by memory request rejection, which depends on the distribution of local DRAM. While previous research has extensively addressed data movement, DRAM distribution among multiple tenants remains underexplored. Changjun Lee, Sangjin Choi, Youngjin Kwon |
EuroSys | 3 |
| 2026 | PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) ArchitecturesabstractThe ability to dynamically allocate memory is fundamental in modern programming languages. However, this feature is not adequately supported in current general-purpose PIM devices. To identify key design principles that PIM must consider, we conduct a design space exploration of PIM memory allocators, examining various strategies for metadata placement and management of the allocator. Based on this exploration, we introduce PIM-malloc, a fast and scalable memory allocator for general-purpose PIM that operates on real PIM hardware, achieving a$66 \times$improvement in memory allocation performance. This design is further enhanced with a lightweight, per-PIM core hardware cache, specifically designed for dynamic memory allocation, achieving an additional 31% performance improvement. Finally, we demonstrate the applicability of PIM-malloc by developing several representative PIM workloads, demonstrating its effectiveness in enhancing programmability. Bongjoon Hyun, Youngjin Kwon, Minsoo Rhu |
HPCA | 3 |
| 2026 | Revisiting Partial Tracing for Safe, Efficient, and Concurrent Garbage Collection in Unmanaged LanguagesabstractGarbage collection (GC) remains a desirable yet elusive goal in unmanaged languages like C/C++ and Rust, where concurrent memory reclamation must be achieved without compiler or runtime support. Existing techniques face fundamental trade-offs among efficiency , safety , and ease of integration : tracing collectors like BDWGC incur costly stop-the-world pauses and unsafe conservative scanning, while reference counting schemes like CIRC are safe but introduce high overhead and require manual handling of cyclic data. We present a safe, efficient, and easy-to-integrate concurrent GC library, revisiting partial tracing (PT), a concept initially conceived by Bacon et al. 22 years ago. PT is a hybrid approach that maintains reference counts for roots, ensuring safety through precise root identification, and traces from objects with non-zero counts, offering ease of integration by handling cyclic garbage. Although PT has historically been considered inefficient due to the high cost of root mutation, we overcome this limitation in two design steps. First, Concurrent Partial Tracing (CPT) introduces phase consensus , enabling concurrent phase coordination without mutator suspension and eliminating most reference-count updates during traversal. Second, Concurrent Deferred Partial Tracing (CDPT) further reduces overhead by replacing atomic root updates with a lightweight, hazard pointer (HP)-based mechanism safeguarded by a phase barrier . We show that CDPT outperforms automatic collectors like BDWGC and CIRC while being comparable to manual schemes like RCU, through both micro-benchmarks on concurrent data structures and a macro-benchmark on Moka, a production cache library. Jongse Park, Youngjin Kwon, Jeehoon Kang |
Proc. ACM Program. Lang. | 3 |
| 2025 | Adios to Busy-Waiting for Microsecond-scale Memory DisaggregationabstractHow fast and efficiently page faults are handled determines the performance of paging-based memory disaggregation (MD) systems. Recent MD systems employ busy-waiting in page fault handling to avoid costly interrupt handling and context switching. Upon a page fault, they issue a remote fetch request and busy-wait for the completion of the request rather than yield their execution to other tasks. While these attempts succeed to cut the latency of MD systems to microseconds, they suffer from head-of-line (HOL) blocking that leads to high tail latency and causes RDMA network underutilization. Wonsup Yoon, Jisu Ok, Sue B. Moon, Youngjin Kwon |
EuroSys | 4 |
| 2025 | Scalable Address Spaces using Concurrent Interval SkiplistabstractA kernel's address space design can significantly bottleneck multi-threaded applications, as address space operations such as mmap() and munmap() are serialized by coarsegrained locks like Linux's mmap_lock. Such locks have long been known as one of the most intractable contention points in memory management. While prior works have attempted to address this issue, they either fail to sufficiently parallelize operations or are impractical for real-world kernels. Tae Woo Kim, Youngjin Kwon, Jeehoon Kang |
SOSP | 2 |
| 2025 | SAND: A New Programming Abstraction for Video-based Deep LearningabstractVideo-based deep learning (VDL) is increasingly used across diverse applications and has become highly popular, but it faces significant challenges in preprocessing highly compressed video data. Preprocessing pipelines are complex, requiring extensive engineering effort, and introduce computational bottlenecks, with latency exceeding GPU training time. Existing solutions partially mitigate these issues but remain inefficient and resource-constrained. Juncheol Ye, Seungkook Lee, Hwijoon Lim, Jihyuk Lee, Uitaek Hong, Youngjin Kwon, Dongsu Han |
SOSP | 6 |
| 2025 | SwiftSweeper: Defeating Use-After-Free Bugs Using Memory Sweeper Without Stop-the-WorldabstractUse-after-free (UAF) vulnerabilities pose severe security risks in memory-unsafe languages like C and C++. To mitigate these issues, prior work has employed memory sweeping, inspired by conservative garbage collection. However, such approaches inherit key limitations, including stop-the-world pauses, poor scalability, and high CPU usage, rendering them unsuitable for modern, latency-sensitive applications. This paper presents SwiftSweeper, a secure memory allocator designed to prevent UAF vulnerabilities in unmodified binaries. SwiftSweeper reimagines memory sweeping by eliminating stop-the-world pauses and enhancing scalability to support high-performance C and C++ workloads. It features an efficient and secure in-kernel data path, implemented using eBPF (XMP, eXpress Memory Path), and a co-designed user-level allocator and kernel. We implement SwiftSweeper on Linux and demonstrate that it delivers state-of-the-art performance, memory efficiency, and minimal latency overhead across both single-threaded and multi-threaded applications, including SPEC CPU and WebServer benchmarks. Junho Ahn, Kanghyuk Lee, Hyungon Moon, Youngjin Kwon |
SP | 5 |
| 2024 | Identifying On-/Off-CPU Bottlenecks Together with Blocked Samples
Minwoo Ahn, Jeongmin Han, Youngjin Kwon, Jinkyu Jeong |
OSDI | 3 |
| 2024 | OZZ: Identifying Kernel Out-of-Order Concurrency Bugs with In-Vivo Memory Access ReorderingabstractKernel concurrency bugs are notoriously difficult to identify, while their consequences severely threaten the reliability and security of the entire system. Especially in the kernel, developers should consider not only locks but also memory barriers to prevent out-of-order execution from breaking the correctness of concurrent execution. Incorrect use of memory barriers may cause non-intuitive concurrency bugs that manifest due to out-of-order execution, which we refer to as OoO bugs. This paper aims to identify OoO bugs in the kernel. We devise a mechanism to emulate out-of-order execution while kernel code is executed, called OEMU. Inspired by how a processor reorders memory accesses, OEMU makes the subtle and non-deterministic behavior of out-of-order execution systematically controllable. Based on OEMU, we propose Ozz , a new testing tool designed to effectively identify kernel OoO bugs. The key feature of Ozz is its ability to deterministically control both out-of-order execution and concurrent execution caused by thread interleavings, enabling comprehensive testing of their combined effects. Our evaluation shows that OEMU is effective in reproducing previously-reported kernel OoO bugs, demonstrating its strong capability of controlling out-of-order execution. Furthermore, with Ozz , we identify 11 new OoO bugs in the latest version of the Linux kernel, subsequently confirmed and patched by kernel developers. Dae R. Jeong, Yewon Choi, Byoungyoung Lee, Insik Shin, Youngjin Kwon |
SOSP | 5 |
| 2024 | BUDAlloc: Defeating Use-After-Free Bugs by Decoupling Virtual Address Management from Kernel
Junho Ahn, Jaehyeon Lee, Kanghyuk Lee, Wooseok Gwak, Minseong Hwang, Youngjin Kwon |
USENIX Security Symposium | 6 |
| 2024 | Hardware-hardened Sandbox Enclaves for Trusted Serverless ComputingabstractIn cloud-based serverless computing, an application consists of multiple functions provided by mutually distrusting parties. For secure serverless computing, the hardware-based trusted execution environment (TEE) can provide strong isolation among functions. However, not only protecting each function from the host OS and other functions, but also protecting the host system from the functions, is critical for the security of the cloud servers. Such an emerging trusted serverless computing poses new challenges: Each TEE must be isolated from the host system bi-directionally, and the system calls from it must be validated. In addition, the resource utilization of each TEE must be accountable in a mutually trusted way. However, the current TEE model cannot efficiently represent such trusted serverless applications. To overcome the lack of such hardware support, this article proposes an extended TEE model called Cloister , designed for trusted serverless computing. Cloister proposes four new key techniques. First, it extends the hardware-based memory isolation in SGX to confine a deployed function only within its TEE (enclave). Second, it proposes a trusted monitor enclave that filters and validates system calls from enclaves. Third, it provides a trusted resource accounting mechanism for enclaves that is agreeable to both service developers and cloud providers. Finally, Cloister accelerates enclave loading by redesigning its memory verification for fast function deployment. Using an emulated Intel SGX platform with the proposed extensions, this article shows that trusted serverless applications can be effectively supported with small changes in the SGX hardware. Joongun Park, Seunghyo Kang, Taehoon Kim 0001, Jongse Park, Youngjin Kwon, Jaehyuk Huh 0001 |
ACM Trans. Archit. Code Optim. | 6 |
| 2023 | Diagnosing Kernel Concurrency Failures with AITIAabstractKernel concurrency failures are notoriously difficult to identify and diagnose their fundamental reason, the root cause. Kernel concurrency bugs frequently involve challenging patterns such as multi-variable races, data races with asynchronous kernel threads, and pervasive benign races. We perform an in-depth study of real-world kernel concurrency bugs and elicit three requirements: comprehensiveness, pattern-agnostic, and conciseness. Dae R. Jeong, Minkyu Jung, Yoochan Lee, Byoungyoung Lee, Insik Shin, Youngjin Kwon |
EuroSys | 6 |
| 2023 | DiLOS: Do Not Trade Compatibility for Performance in Memory DisaggregationabstractMemory disaggregation has replaced the landscape of dat-acenters by physically separating compute and memory nodes, achieving improved utilization. As early efforts, kernel paging-based approaches offer transparent virtual memory abstraction for remote memory with paging schemes but suffer from expensive page fault handling. This paper revisits the paging-based approaches and challenges their performance in paging schemes. We posit that the overhead of the paging-based approaches is not a fundamental limitation. We propose DiLOS, a new library operating system (LibOS) specialized for paging-based memory disaggregation. We have revamped the page fault handler to get away with the swap cache and incorporated known techniques in our prefetcher, page manager, and communication module for performance optimization. Furthermore, we provide APIs to augment the LibOS with application semantics. We present two app-aware guides, app-aware prefetching and bandwidth-reducing memory allocator in DiLOS. Through extensive evaluation of microbenchmarks and applications, we demonstrate that DiLOS outperforms the state-of-the-art kernel paging-based system (Fastswap) up to 2.24× and a recent user-level system (AIFM) 1.54× on a real-world data analytic workload. Wonsup Yoon, Jisu Ok, Jinyoung Oh, Sue B. Moon, Youngjin Kwon |
EuroSys | 5 |
| 2023 | PRIMO: A Full-Stack Processing-in-DRAM Emulation Framework for Machine Learning WorkloadsabstractRecently, the size of deep learning models has significantly increased, making the excessive memory access between the AI processor and DRAM a major bottleneck of the system. The processing-in-DRAM (DRAM-PIM) concept has emerged as a promising solution, which integrates computing logic within memory, thus saving abundant access to external memory. Although many simulators have been proposed to model and analyze the benefits of DRAM-PIM, they are often too slow to run an entire application. FPGA-based emulators have been introduced to overcome this limitation. However, none of the prior works include the full software stack from the model to DRAM-PIM hardware. This paper presents a full-stack processing-in-DRAM emulation framework named PRIMO, the first emulation framework that can model and analyze DRAM-PIM for end-to-end ML inference. PRIMO enables software developers to develop and test their customized software stacks on various ML workloads without requiring a real DRAM-PIM chip. Moreover, it allows designers to explore design space and monitor memory access patterns, facilitating software and hardware co-design for efficient DRAM-PIM architectures. To achieve these goals, we develop a real-time FPGA emulator that emulates DRAM-PIM architecture and generates experimental results such as predicted cycle information and computed output at incomparably high speeds compared to the CPU-based simulation. In addition, we propose a software stack comprising a PIM compiler that enables the execution of various ML workloads, including end-to-end inference, and a PIM driver that runs the workloads with high bandwidth utilization by leveraging virtual memory scatter-gather DMA. Finally, we demonstrate that PRIMO can successfully emulate DRAM-PIM 106.64-6093.56× faster than the CPU-based simulation framework for ML workloads ranging from small microbenchmarks to end-to-end inference of ResNets. Jaehoon Heo, Yongwon Shin, Sangjin Choi, Sungwoong Yune, Hyojin Sung, Youngjin Kwon, Joo-Young Kim 0001 |
ICCAD | 7 |
| 2023 | Rearchitecting the TCP Stack for I/O-Offloaded Content Delivery
Deondre Martin Ng, Junzhi Gong, Youngjin Kwon, Minlan Yu, KyoungSoo Park |
NSDI | 4 |
| 2023 | Poster: Designing a Memory Disaggregation System for CloudabstractMemory disaggregation is a new datacenter paradigm separating compute and memory nodes. While memory disaggregation improves memory utilization and scalability, it poses challenges for cloud applications, particularly in terms of high tail latency. Existing memory disaggregation systems focus on optimizing the disaggregation stack, but it does not always guarantee excellent application performance. We review existing memory disaggregation and tail-optimized systems and explain their limitations in this context. We also propose two preliminary solutions: asynchronous page fault handling and a faulty request classifier. The emulation result shows that asynchronous page fault handling reduces tail latencies by 50% compared to synchronous handling. Wonsup Yoon, Jisu Ok, Sue B. Moon, Youngjin Kwon |
SIGCOMM | 4 |
| 2023 | SegFuzz: Segmentizing Thread Interleaving to Discover Kernel Concurrency Bugs through FuzzingabstractDiscovering kernel concurrency bugs through fuzzing is challenging. Identifying kernel concurrency bugs, as opposed to non-concurrency bugs, necessitates an analysis of possible interleavings between two or more threads. However, because the search space of thread interleaving is vast, it is impractical to investigate all conceivable thread interleavings. To explore the vast search space, most previous approaches perform random or simple heuristic searches without having coverage for thread interleaving or with an insufficient form of coverage. As a result, they either conduct wasteful searches with redundant executions or overlook concurrent bugs that their coverage cannot address.To overcome such limitations, we propose SegFuzz, a fuzzing framework for kernel concurrency bugs. When exploring the search space of thread interleavings, SegFuzz decomposes an entire thread interleaving into a set of segments, each of which represents an interleaving of the small number of instructions, and utilizes individual segments as interleaving coverage, called interleaving segment coverage. When searching for thread interleavings, SegFuzz mutates interleavings in explored interleaving segments to construct new thread interleavings that have not yet been explored. With SegFuzz, we discover new 21 concurrency bugs in Linux kernels, and demonstrate the efficiency of SegFuzz by showing that SegFuzz can identify known bugs on average 4.1 times quickly than the state-of-the-art approaches. Dae R. Jeong, Byoungyoung Lee, Insik Shin, Youngjin Kwon |
SP | 4 |
| 2023 | EnvPipe: Performance-preserving DNN Training Framework for Saving Energy
Sangjin Choi, Inhoe Koo, Jeongseob Ahn, Myeongjae Jeon, Youngjin Kwon |
USENIX ATC | 5 |
| 2023 | On-Demand Virtualization for Post-Copy OS Migration in Bare-Metal CloudabstractThe demand for bare-metal cloud services has increased rapidly because bare-metal cloud is cost-effective for various types of cloud workloads. However, as the bare-metal cloud does not utilize the abstraction of the virtualization layer, it misses the benefits of virtualization. One important benefits absent in the bare-metal cloud is the live migration of guest operating systems. Migrating an OS and applications in the OS as a single unit provides a convenient way to manage cloud services such as load balancing, fault management, and system maintenance. To enable live migration for bare-metal cloud, several approaches have been proposed but they have limitations; they require OS modifications or impose additional overheads for workloads. This paper suggests an on-demand virtualization technique for post-copy OS migration to improve manageability of the bare-metal cloud services. When live migration is requested, a lightweight virtualization layer is enabled in the host on the fly. After completion of the live migration, the virtualization layer is removed from the host. Therefore, the host returns to a bare-metal system for performance. To implement on-demand virtualization, we modify BitVisor to perform the post-copy migration on the x86 architecture. The elapsed time of on-demand virtualization is negligible. It takes only 20 ms to insert the virtualization layer and 30 ms to remove the one. The downtime of migration is reduced because of the post-copy migration. Jaeseong Im, Jongyul Kim 0001, Youngjin Kwon, Seung Ryoul Maeng |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | Personalization Trade-offs in Designing a Dialogue-based Information System for Support-Seeking of Sexual Violence SurvivorsabstractThe lack of reliable, personalized information often complicates sexual violence survivors’ support-seeking. Recently, there is an emerging approach to conversational information systems for support-seeking of sexual violence survivors, featuring personalization with wide availability and anonymity. However, a single best solution might not exist as sexual violence survivors have different needs and purposes in seeking support channels. To better envision conversational support-seeking systems for sexual violence survivors, we explore personalization trade-offs in designing such information systems. We implement a high-fidelity prototype dialogue-based information system through four design workshop sessions with three professional caregivers and interviewed with four self-identified survivors using our prototype. We then identify two forms of personalization trade-offs for conversational support-seeking systems: (1) specificity and sensitivity in understanding users and (2) relevancy and inclusiveness in providing information. To handle these trade-offs, we propose a reversed approach that starts from designing information and inclusive tailoring that considers unspecified needs, respectively. Hyeok Kim, Youjin Hwang, Youngjin Kwon, Joonhwan Lee |
CHI | 4 |
| 2022 | Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual Memory
Sangjin Choi, Taeksoo Kim, Rachata Ausavarungnirun, Myeongjae Jeon, Youngjin Kwon, Jeongseob Ahn |
USENIX ATC | 6 |
| 2022 | Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing
Seungbeom Choi, Sunho Lee 0003, Yeonjae Kim, Jongse Park, Youngjin Kwon, Jaehyuk Huh 0001 |
USENIX ATC | 5 |
| 2021 | Rethinking File Mapping for Persistent Memory
Ian Neal, Gefei Zuo, Eric Shiple, Tanvir Ahmed Khan 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci |
FAST | 5 |
| 2021 | LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismabstractIn multi-tenant systems, the CPU overhead of distributed file systems (DFSes) is increasingly a burden to application performance. CPU and memory interference cause degraded and unstable application and storage performance, in particular for operation latency. Recent client-local DFSes for persistent memory (PM) accelerate this trend. DFS offload to SmartNICs is a promising solution to these problems, but it is challenging to fit the complex demands of a DFS onto simple SmartNIC processors located across PCIe. Jongyul Kim 0001, Insu Jang, Waleed Reda, Jaeseong Im, Marco Canini, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Emmett Witchel |
SOSP | 7 |
| 2021 | Zico: Efficient GPU Memory Sharing for Concurrent DNN Training
Gangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon, Myeongjae Jeon |
USENIX ATC | 4 |
| 2020 | Perforated Page: Supporting Fragmented Memory Allocation for Large PagesabstractThe availability of large pages has dramatically improved the efficiency of address translation for applications that use large contiguous regions of memory. However, large pages can be difficult to allocate due to fragmented memory, non-movable pages, or the need to split a large page into regular pages when part of the large page is forced to have a different permission status from the rest of the page. Furthermore, they can also be expensive due to memory bloating caused by sparse accesses to application data. In this work, we enable the allocation of large 2MB pages even in the presence of fragmented physical memory via perforated pages. Perforated pages permit the OS to punch 4KB page-sized holes in the physical address range allocated to a large page and re-map them to other addresses as needed. This not only enables the system to benefit from large pages in the presence of fragmentation, but also allows for different permissions to exist within a large page, enhancing sharing flexibility. In addition, it allows unused parts of a large page to be used elsewhere, mitigating memory bloating. To minimize changes to the system, perforated pages reuse the 4KBlevel page table entries to store the hole locations and translates holes into regular 4KB pages. For performance, the proposed technique caches the translations for hole pages in the TLBs and track holes via cached bitmaps in the L2 TLB. By enabling large pages in the presence of physical memory fragmentation, perforated pages increase the applicability and resulting benefits of large pages with only minor changes to the hardware and OS. In this work, we evaluate the effectiveness of perforated pages with timing simulations under diverse and realistic fragmentation scenarios. Our results show that even with fragmented memory, perforated pages accomplish 93.2% to 99.9% of the performance achievable by ideal memory allocation, and 2.0% to 11.5% better performance over the conventional system running with fragmented memory. Chang Hyun Park 0001, Sanghoon Cha, Bokyeong Kim, Youngjin Kwon, David Black-Schaffer, Jaehyuk Huh 0001 |
ISCA | 4 |
| 2020 | Nested Enclave: Supporting Fine-grained Hierarchical Isolation with SGXabstractAlthough hardware-based trusted execution environments (TEEs) have evolved to provide strong isolation with efficient hardware supports, their current monolithic model poses challenges in representing common software structures with modules produced from potentially untrusted 3rd parties. For better mapping of such modular software designs to trusted execution environments, it is necessary to extend the current monolithic model to a hierarchical one, which provides multiple inner TEEs within a TEE. For such hierarchical compartmentalization within a TEE, this paper proposes a novel hierarchical TEE called nested enclave, which extends the enclave support from Intel SGX. Inspired by the multi-level security model, nested enclave provides multiple inner enclaves sharing the same outer enclave. Inner enclaves can access the context of the outer enclave, but they are protected from the outer enclave and non-enclave execution. Peer inner enclaves are isolated from each other while accessing the execution environment of the shared outer enclave. Both of the inner and outer enclaves are protected from vulnerable privileged software and physical attacks. Such fine-grained nested enclaves allow secure multitiered environments using software modules from untrusted 3rd parties. The security-sensitive modules run on the inner enclave with the higher security level, while the 3rd party modules on the outer enclave. It can be further extended to provide a separate inner module for each user to process privacy-sensitive data while sharing the same library with efficient hardwareprotected communication channels. This study investigates three case scenarios implemented with an emulated nested enclave support, proving the feasibility and security improvement of the nested enclave model. Joongun Park, Naegyeong Kang, Taehoon Kim 0001, Youngjin Kwon, Jaehyuk Huh 0001 |
ISCA | 4 |
| 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory SystemabstractPersistent memory programming requires failure atomicity. To achieve this in an efficient manner, recent proposals use hardware-based logging for atomic-durable updates and hardware transactional memory (HTM) for isolation. Although the unbounded HTMs are promising for both performance and programmability reasons, none of the previous studies satisfies the practical requirements. They either require unrealistic hard-ware overheads or do not allow transactions to exceed on-chip cache boundaries. Furthermore, it has never been possible to use both DRAM and NVM in HTM, though it is becoming a popular persistency model. To this end, this study proposes UHTM, unbounded hardware transactional memory for DRAM and NVM hybrid memory systems. UHTM combines the cache coherence protocol and address-signatures to detect conflicts in the entire memory space. This approach improves concurrency by significantly reducing the false-positive rates of previous studies. More importantly, UHTM allows both DRAM and NVM data to interact with each other in transactions without compromising the consistency guarantee. This is rendered possible by UHTM's hybrid version management that provides an undo-based log for DRAM and a redo-based log for NVM. The experimental results show that UHTM outperforms the state-of-the-art durable HTM, which is LLC-bounded, by 56% on average and up to 818%. Jungi Jeong, Jaewan Hong, Seung Ryoul Maeng, Changhee Jung, Youngjin Kwon |
MICRO | 5 |
| 2020 | Assise: Performance and Availability via Client-local NVM in a Distributed File System
Thomas E. Anderson, Marco Canini, Jongyul Kim 0001, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Waleed Reda, Henry Schuh, Emmett Witchel |
OSDI | 5 |
| 2020 | AGAMOTTO: How Persistent is your Persistent Memory Application?
Ian Neal, Ben Reeves, Benjamin Stoler, Andrew Quinn 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci |
OSDI | 5 |
| 2020 | Libnvmmio: Reconstructing Software IO Path with Failure-Atomic Memory-Mapped Interface
Jungsik Choi 0001, Jaewan Hong, Youngjin Kwon, Hwansoo Han |
USENIX ATC | 3 |
| 2019 | Effects of ego networks and communities on self-disclosure in an online social networkabstractUnderstanding how much users disclose personal information in Online Social Networks (OSN) has served various scenarios such as maintaining social relationships and customer segmentation. Prior studies on self-disclosure have relied on surveys or users' direct social networks. These approaches, however, cannot represent the whole population nor consider user dynamics at the community level. Young D. Kwon, Reza Hadi Mogavi, Ehsan ul Haq, Youngjin Kwon, Xiaojuan Ma, Pan Hui 0001 |
ASONAM | 4 |
| 2019 | TxFS: Leveraging File-system Crash Consistency to Provide ACID TransactionsabstractWe introduce TxFS, a transactional file system that builds upon a file system’s atomic-update mechanism such as journaling. Though prior work has explored a number of transactional file systems, TxFS has a unique set of properties: a simple API, portability across different hardware, high performance, low complexity (by building on the file-system journal), and full ACID transactions. We port SQLite, OpenLDAP, and Git to use TxFS and experimentally show that TxFS provides strong crash consistency while providing equal or better performance. Yige Hu, Zhiting Zhu, Ian Neal, Youngjin Kwon, Vijay Chidambaram, Emmett Witchel |
ACM Trans. Storage | 4 |
| 2018 | TxFS: Leveraging File-System Crash Consistency to Provide ACID Transactions
Yige Hu, Zhiting Zhu, Ian Neal, Youngjin Kwon, Vijay Chidambaram, Emmett Witchel |
USENIX ATC | 4 |
| 2018 | PerfIso: Performance Isolation for Commercial Latency-Sensitive Services
Calin Iorgulescu, Youngjin Kwon, Sameh Elnikety, Manoj Syamala, Vivek R. Narasayya, Herodotos Herodotou, Paulo Tomita, Alex Chen, Jack Zhang |
USENIX ATC | 3 |
| 2017 | From Crash Consistency to TransactionsabstractModern applications use multiple storage abstractions such as the file system, key-value stores, and embedded databases such as SQLite. Maintaining consistency of data spread across multiple abstractions is complex and error-prone. Applications are forced to copy data unnecessarily and use long sequences of system calls to update state in a consistent manner. Not only does this create implementation complexity, it also introduces potential performance problems from redundant IO and fsync() calls, which fragment disk writes into small, random IOs. In this paper, we propose that the operating system should provide transactions across multiple storage abstractions; we can build such transactions with low development cost by taking advantage of a well-tested piece of software: the file-system journal. We present the design of our cross-abstraction transactions and some preliminary results, showing such transactions can increase performance by 31% in certain cases. Yige Hu, Youngjin Kwon, Vijay Chidambaram, Emmett Witchel |
HotOS | 2 |
| 2017 | Strata: A Cross Media File SystemabstractCurrent hardware and application storage trends put immense pressure on the operating system's storage subsystem. On the hardware side, the market for storage devices has diversified to a multi-layer storage topology spanning multiple orders of magnitude in cost and performance. Above the file system, applications increasingly need to process small, random IO on vast data sets with low latency, high throughput, and simple crash consistency. File systems designed for a single storage layer cannot support all of these demands together. Youngjin Kwon, Henrique Fingler, Tyler Hunt, Simon Peter 0001, Emmett Witchel, Thomas E. Anderson |
SOSP | 1 |
| 2017 | Resource Accounting of Shared IT Resources in Multi-Tenant CloudsabstractIn today's IT platforms, the capability to accurately account overall resource usage among applications is crucial for variety of management actions (e.g., capacity planning, dynamic resource reallocation and/or load balancing). However, in the environments where small number of shared services cater to a large number of distinct entities' requests, resource accounting becomes significantly challenging. First, the overall resource consumption at the shared service is the aggregate of the resource consumption for multiple remote entities whose identities are not visible to the shared service. Second, even if such information becomes available, common monitoring tools (e.g., top, iostat) are unable to deliver accurate break-down of resource consumption since sharing occurs at sub-instance level (i.e., service instances are not exclusive). We study inherent challenges of performing resource accounting of shared resource. We compare two nonintrusive approaches having different balance between local monitoring and collective inference - (i) LR that uses easily-available tools which provide aggregate measurement and applying well-known linear regression as inference, and (ii) Rameter that puts more emphasis on gathering fine-grained per-thread information from within the hypervisor and applying light inference on the data. Evaluation shows that Rameter offers less than 1% error in accounting whereas LR's error fluctuates between 5-150%. Byung-Chul Tak, Youngjin Kwon, Bhuvan Urgaonkar |
IEEE Trans. Serv. Comput. | 2 |
| 2016 | Sego: Pervasive Trusted Metadata for Efficiently Verified Untrusted System ServicesabstractSego is a hypervisor-based system that gives strong privacy and integrity guarantees to trusted applications, even when the guest operating system is compromised or hostile. Sego verifies operating system services, like the file system, instead of replacing them. By associating trusted metadata with user data across all system devices, Sego verifies system services more efficiently than previous systems, especially services that depend on data contents. We extensively evaluate Sego's performance on real workloads and implement a kernel fault injector to validate Sego's file system-agnostic crash consistency and recovery protocol. Youngjin Kwon, Alan M. Dunn, Michael Z. Lee, Owen S. Hofmann, Yuanzhong Xu, Emmett Witchel |
ASPLOS | 1 |
| 2016 | Earp: Principled Storage, Sharing, and Protection for Mobile Apps
Yuanzhong Xu, Tyler Hunt, Youngjin Kwon, Martin Georgiev, Vitaly Shmatikov, Emmett Witchel |
NSDI | 3 |
| 2016 | Coordinated and Efficient Huge Page Management with Ingens
Youngjin Kwon, Hangchen Yu, Simon Peter 0001, Christopher J. Rossbach, Emmett Witchel |
OSDI | 1 |
| 2011 | Virtualizing performance asymmetric multi-core systemsabstractPerformance-asymmetric multi-cores consist of heterogeneous cores, which support the same ISA, but have different computing capabilities. To maximize the throughput of asymmetric multi-core systems, operating systems are responsible for scheduling threads to different types of cores. However, system virtualization poses a challenge for such asymmetric multi-cores, since virtualization hides the physical heterogeneity from guest operating systems. In this paper, we explore the design space of hypervisor schedulers for asymmetric multi-cores, which do not require asymmetry-awareness from guest operating systems. The proposed scheduler characterizes the efficiency of each virtual core, and map the virtual core to the most area-efficient physical core. In addition to the overall system throughput, we consider two important aspects of virtualizing asymmetric multi-cores: performance fairness among virtual machines and performance scalability for changing availability of fast and slow cores. Youngjin Kwon, Changdae Kim 0001, Seung Ryoul Maeng, Jaehyuk Huh 0001 |
ISCA | 1 |