VLDB 2026 Research / reviewers in the wild / expert
Miryeong Kwon
dblp:188/9926
· DBLP profile ↗
26ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-0313-1319ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 7 first-author · 16 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN PerformanceabstractGraph neural network (GNN) inference faces significant bottlenecks in preprocessing, which often dominate overall inference latency. We introduce AutoGNN, an FPGA-based accelerator designed to address these challenges by leveraging FPGA's reconfigurability and specialized components. AutoGNN adapts to diverse graph inputs, efficiently performing computationally intensive tasks such as graph conversion and sampling. By utilizing components like adder trees, AutoGNN executes reduction operations in constant time, overcoming the limitations of serialization and synchronization on GPUs. AutoGNN integrates unified processing elements (UPEs) and single-cycle reducers (SCRs) to streamline GNN preprocessing. UPEs enable scalable parallel processing for edge sorting and unique vertex selection, while SCRs efficiently handle sequential tasks such as pointer array construction and subgraph reindexing. A user-level software framework dynamically profiles graph inputs, determines optimal configurations, and reprograms AutoGNN to handle varying workloads. Implemented on a$7 n \mathrm{m}$enterprise FPGA, AutoGNN achieves up to$9.0 \times$and$2.1 \times$speedup compared to conventional and GPU-accelerated preprocessing systems, respectively, enabling high-performance GNN preprocessing across diverse datasets. Seungkwan Kang, Donghyun Gouk, Miryeong Kwon, Hyunkyu Choi, Junhyeok Jang, Sangwon Lee 0014, Huiwon Choi, Jie Zhang 0048, Wonil Choi, Mahmut T. Kandemir, Myoungsoo Jung |
HPCA | 4 |
| 2026 | A Silicon-Proven Unified Low-Latency CXL Controller and Port-Based Routing Switch for Memory-Centric Fabrics
Miryeong Kwon, Donghyun Gouk, Hongjoo Jung, Eojin Ryu, Seyeong Huh, Junseok Moon, Hyein Woo, Junhee Kim, Kyungkuk Nam, Jinwoo Baek, Hyunkyu Choi, Woojin Choi, Yongjin Cho, Myoungsoo Jung |
ISCA | 1 |
| 2024 | DockerSSD: Containerized In-Storage Processing and Hardware Acceleration for Computational SSDsabstractProcessing data in storage is an energy-efficient solution to examine massive datasets. However, a general incarnation of such well-known task-offloading model in a real system is unfortunately unsuccessful due to not only poor performance but also many practical challenges, such as limited processing capabilities and high vulnerabilities at the storage-level. We propose DockerSSD, a fully flexible in-storage processing (ISP) model that can run a variety of applications near flash without their source-level modification. Specifically, it enables lightweight OS-level virtualization in modern SSDs, which allows the storage intelligence to be well harmonized with existing computing environment and makes ISP even faster. Instead of developing a vendor-specific ISP to offload, DockerSSD can reuse existing Docker images, create containers as a self-governing execution object in storage, and process data directly where they are in real-time. To this end, we design a new communication method and virtual firmware that operate together to download Docker images and manage their container execution without a change of the existing storage interface and runtime. We further accelerate ISP and reduce the execution latency by automating container-related network and I/O handling data paths over hardware. Our evaluation shows that DockerSSD is 2.0 × faster than state-of-the-art ISP models for workloads with a high volume of system calls or file accesses. Moreover, it demonstrates a reduction in power and energy consumption by 1.6 × and 2.3 × respectively. Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Myoungsoo Jung |
HPCA | 2 |
| 2024 | Bridging Software-Hardware for CXL Memory Disaggregation in Billion-Scale Nearest Neighbor SearchabstractWe propose CXL-ANNS , a software-hardware collaborative approach to enable scalable approximate nearest neighbor search (ANNS) services. To this end, we first disaggregate DRAM from the host via compute express link (CXL) and place all essential datasets into its memory pool. While this CXL memory pool allows ANNS to handle billion-point graphs without an accuracy loss, we observe that the search performance significantly degrades because of CXL’s far-memory-like characteristics. To address this, CXL-ANNS considers the node-level relationship and caches the neighbors in local memory, which are expected to visit most frequently. For the uncached nodes, CXL-ANNS prefetches a set of nodes most likely to visit soon by understanding the graph traversing behaviors of ANNS. CXL-ANNS is also aware of the architectural structures of the CXL interconnect network and lets different hardware components collaborate with each other for the search. Furthermore, it relaxes the execution dependency of neighbor search tasks and allows ANNS to utilize all hardware in the CXL network in parallel. Our evaluation shows that CXL-ANNS exhibits 93.3% lower query latency than state-of-the-art ANNS platforms that we tested. CXL-ANNS also outperforms an oracle ANNS system that has unlimited local DRAM capacity by 68.0%, in terms of latency. Junhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Miryeong Kwon, Myoungsoo Jung |
ACM Trans. Storage | 5 |
| 2023 | Cache in Hand: Expander-Driven CXL Prefetcher for Next Generation CXL-SSDabstractIntegrating compute express link (CXL) with SSDs allows scalable access to large memory but has slower speeds than DRAMs. We present ExPAND, an expander-driven CXL prefetcher that offloads last-level cache (LLC) prefetching from host CPU to CXL-SSDs. ExPAND uses a heterogeneous prediction algorithm for prefetching and ensures data consistency with CXL.mem's back-invalidation. We examine prefetch timeliness for accurate latency estimation. ExPAND, being aware of CXL multi-tiered switching, provides end-to-end latency for each CXL-SSD and precise prefetch timeliness estimations. Our method reduces CXL-SSD reliance and enables direct host cache access for most data. ExPAND enhances graph application performance by 3.5x, surpassing CXL-SSD pools with diverse prefetching strategies. Miryeong Kwon, Sangwon Lee 0014, Myoungsoo Jung |
HotStorage | 1 |
| 2023 | GraphTensor: Comprehensive GNN-Acceleration Framework for Efficient Parallel Processing of Massive DatasetsabstractWe present GraphTensor, a comprehensive open-source framework that supports efficient parallel neural network processing on large graphs. GraphTensor offers a set of easy-to-use programming primitives that appreciate both graph and neural network execution behaviors from the beginning (graph sampling) to the end (dense data processing). Our framework runs diverse graph neural network (GNN) models in a destination-centric, feature-wise manner, which can significantly shorten training execution times in a GPU. In addition, GraphTensor rearranges multiple GNN kernels based on their system hyperparameters in a self-governing manner, thereby reducing the processing dimensionality and the latencies further. From the end-to-end execution viewpoint, GraphTensor significantly shortens the service-level GNN latency by applying pipeline parallelism for efficient graph dataset preprocessing. Our evaluation shows that GraphTensor exhibits 1.4× better training performance than emerging GNN frameworks under the execution of large-scale, real-world graph workloads. For the end-to-end services, GraphTensor reduces training latencies of an advanced version of the GNN frameworks (optimized for multi-threaded graph sampling) by 2.4×, on average. Junhyeok Jang, Miryeong Kwon, Donghyun Gouk, Hanyeoreum Bae, Myoungsoo Jung |
IPDPS | 2 |
| 2023 | CXL-ANNS: Software-Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search
Junhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Miryeong Kwon, Myoungsoo Jung |
USENIX ATC | 5 |
| 2023 | Realizing Strong Determinism Contract on Log-Structured Merge Key-Value StoresabstractWe propose Vigil-KV , a hardware and software co-designed framework that eliminates long-tail latency almost perfectly by introducing strong latency determinism. To make Get latency deterministic, Vigil-KV first enables a predictable latency mode (PLM) interface on a real datacenter-scale NVMe SSD, having knowledge about the nature of the underlying flash technologies. Vigil-KV at the system-level then hides the non-deterministic time window (associated with SSD’s internal tasks and/or write services) by internally scheduling the different device states of PLM across multiple physical functions. Vigil-KV further schedules compaction/flush operations and client requests being aware of PLM’s restrictions thereby integrating strong latency determinism into LSM KVs. We implement Vigil-KV upon a 1.92TB NVMe SSD prototype and Linux 4.19.91, but other LSM KVs can adopt its concept. We evaluate diverse Facebook and Yahoo scenarios with Vigil-KV, and the results show that Vigil-KV can reducethe tail latency of a baseline KV system by 3.19× while reducing the average latency by 34%, on average. Miryeong Kwon, Hyunkyu Choi, Jooyoung Hwang, Myoungsoo Jung |
ACM Trans. Storage | 1 |
| 2022 | Hardware/Software Co-Programmable Framework for Computational SSDs to Accelerate Deep Learning Service on Large-Scale Graphs
Miryeong Kwon, Donghyun Gouk, Sangwon Lee 0014, Myoungsoo Jung |
FAST | 1 |
| 2022 | Large-scale Graph Neural Network Services through Computational SSD and In-Storage Processing ArchitecturesabstractDemonstration Video Link: https://www.youtube.com/watch?v=b5fZBESH1TM Miryeong Kwon, Donghyun Gouk, Sangwon Lee 0014, Myoungsoo Jung |
HCS | 1 |
| 2022 | What you can't forget: exploiting parallelism for zoned namespacesabstractThis paper discusses the main benefits of ZNS and shows why ZNS can be deprived of internal parallelism when downsizing its zone writable capacity. To this end, we use two production ZNS SSDs and quantitively analyze the performance degradation caused by inter-zone interference. We then suggest a simple mechanism to detect zone-to-zone relationships generating the interference and schedule I/O requests by being aware of internal parallelism. Our evaluation results using real production ZNS devices show that our mechanism can improve the bandwidth and latency of Linux's multi-queue I/O scheduler by 1.98 × and 2.2 ×, respectively. Hanyeoreum Bae, Jiseon Kim, Miryeong Kwon, Myoungsoo Jung |
HotStorage | 3 |
| 2022 | LightPC: hardware and software co-design for energy-efficient full system persistenceabstractWe propose LightPC, a lightweight persistence-centric platform to make the system robust against power loss. LightPC consists of hardware and software subsystems, each being referred to as open-channel PMEM (OC-PMEM) and persistence-centric OS (PecOS). OC-PMEM removes physical and logical boundaries in drawing a line between volatile and nonvolatile data structures by unshackling new memory media from conventional PMEM complex. PecOS provides a single execution persistence cut to quickly convert the execution states to persistent information in cases of a power failure, which can eliminate persistent control overhead. We prototype LightPC's computing complex and OC-PMEM using our custom system board. PecOS is implemented based on Linux 4.19 and Berkeley bootloader on the hardware prototype. Our evaluation results show that OC-PMEM can make user-level performance comparable with a DRAM-only non-persistent system, while consuming 73% lower power and 69% less energy. LightPC also shortens the execution time of diverse HPC, SPEC, and In-memory DB workloads, compared to traditional persistent systems by 4.3X, on average. Sangwon Lee 0014, Miryeong Kwon, Gyuyoung Park, Myoungsoo Jung |
ISCA | 2 |
| 2022 | Direct Access, High-Performance Memory Disaggregation with DirectCXL
Donghyun Gouk, Sangwon Lee 0014, Miryeong Kwon, Myoungsoo Jung |
USENIX ATC | 3 |
| 2022 | Vigil-KV: Hardware-Software Co-Design to Integrate Strong Latency Determinism into Log-Structured Merge Key-Value Stores
Miryeong Kwon, Hyunkyu Choi, Jooyoung Hwang, Myoungsoo Jung |
USENIX ATC | 1 |
| 2021 | Empirical Guide to Use of Persistent Memory for Large-Scale In-Memory Graph AnalysisabstractWe investigate runtime environment characteristics and explore the challenges of conventional in-memory graph processing. This system-level analysis includes empirical results and observations, which are opposite to the existing expectations of graph application users. Specifically, since raw graph data are not the same as the in-memory graph data, processing a billion-scale graph exhausts all system resources and makes the target system unavailable due to out-of-memory at runtime.To address a lack of memory space problem for big-scale graph analysis, we configure real persistent memory devices (PMEMs) with different operation modes and system software frameworks. In this work, we introduce PMEM to a representative in-memory graph system, Ligra, and perform an in-depth analysis uncovering the performance behaviors of different PMEM-applied in-memory graph systems. Based on our observations, we modify Ligra to improve the graph processing performance with a solid level of data persistence. Our evaluation results reveal that Ligra, with our simple modification, exhibits 4.41× and 3.01× better performance than the original Ligra running on a virtual memory expansion and conventional persistent memory, respectively. Hanyeoreum Bae, Miryeong Kwon, Donghyun Gouk, Sungjoon Koh, Changrim Lee, Dongchul Park, Myoungsoo Jung |
ICCD | 2 |
| 2021 | Revamping Storage Class Memory With Hardware Automated Memory-Over-Storage SolutionabstractLarge persistent memories such as NVDIMM have been perceived as a disruptive memory technology, because they can maintain the state of a system even after a power failure and allow the system to recover quickly. However, overheads incurred by a heavy software-stack intervention seriously negate the benefits of such memories. First, to significantly reduce the software stack overheads, we propose HAMS, a hardware auto-mated Memory-over-Storage (MoS) solution. Specifically, HAMS aggregates the capacity of NVDIMM and ultra-low latency flash archives (ULL-Flash) into a single large memory space, which can be used as a working memory expansion or persistent memory expansion, in an OS-transparent manner. HAMS resides in the memory controller hub and manages its MoS address pool over conventional DDR and NVMe interfaces; it employs a simple hardware cache to serve all the memory requests from the host MMU after mapping the storage space of ULL-Flash to the memory space of NVDIMM. Second, to make HAMS more energy-efficient and reliable, we propose an "advanced HAMS" which removes unnecessary data transfers between NVDIMM and ULL-Flash after optimizing the datapath and hardware modules of HAMS. This approach unleashes the ULL-Flash and its NVMe controller from the storage box and directly connects the HAMS datapath to NVDIMM over the conventional DDR4 interface. Our evaluations show that HAMS and advanced HAMS can offer 97% and 119% higher system performance than a software-based NVDIMM design, while costing 41% and 45% lower energy, respectively. Jie Zhang 0048, Miryeong Kwon, Donghyun Gouk, Sungjoon Koh, Nam Sung Kim, Mahmut T. Kandemir, Myoungsoo Jung |
ISCA | 2 |
| 2020 | Scalable Parallel Flash Firmware for Many-core Architectures
Jie Zhang 0048, Miryeong Kwon, Michael M. Swift, Myoungsoo Jung |
FAST | 2 |
| 2020 | DC-Store: Eliminating Noisy Neighbor Containers using Deterministic I/O Performance and Resource Isolation
Miryeong Kwon, Donghyun Gouk, Changrim Lee, Byounggeun Kim, Jooyoung Hwang, Myoungsoo Jung |
FAST | 1 |
| 2020 | Design of a Host Interface Logic for GC-Free SSDsabstractGarbage collection (GC) and resource contention on I/O buses (channels) are among the critical bottlenecks in solid-state drives (SSDs) that cannot be easily hidden. Most existing I/O scheduling algorithms in the host interface logic (HIL) of state-of-the-art SSDs are oblivious to such low-level performance bottlenecks in SSDs. As a result, SSDs may violate quality of service (QoS) requirements by not being able to meet the deadlines of I/O requests. In this paper, we propose a novel host interface I/O scheduler that is both GC aware and QoS aware. The proposed scheduler redistributes the GC overheads across noncritical I/O requests and reduces channel resource contention. Our experiments with workloads from various application domains revealed that the proposed client-level SSD scheduler reduces the standard deviation for latency by 52.5% and the worst-case latency by 86.6%, compared to the state-of-the-art I/O schedulers used for the HIL. In addition, for I/O requests smaller than a superpage, the proposed scheduler avoids channel resource conflicts and reduces latency by 29.2% in comparison to the state-of-the-art I/O schedulers. Furthermore, we present an extension of the proposed I/O scheduler for enterprise SSDs based on the NVMe protocol. Myoungsoo Jung, Wonil Choi, Miryeong Kwon, Shekhar Srikantaiah, Joonhyuk Yoo, Mahmut T. Kandemir |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Errata to "Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters"abstractPresents corrections to author affiliation information in the above mentioned article. Sungjoon Koh, Jie Zhang 0048, Miryeong Kwon, Jungyeon Yoon, David Donofrio, Nam Sung Kim, Myoungsoo Jung |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | FlashGPU: Placing New Flash Next to GPU CoresabstractWe propose FlashGPU, a new GPU architecture that tightly blends new flash (Z-NAND) with massive GPU cores. Specifically, we replace global memory with Z-NAND that exhibits ultra-low latency. We also architect a flash core to manage request dispatches and address translations underneath L2 cache banks of GPU cores. While Z-NAND is a hundred times faster than conventional 3D-stacked flash, its latency is still longer than DRAM. To address this shortcoming, we propose a dynamic page-placement and buffer manager in Z-NAND subsystems by being aware of bulk and parallel memory access characteristics of GPU applications, thereby offering high-throughput and low-energy consumption behaviors. Jie Zhang 0048, Miryeong Kwon, Hyojong Kim, Hyesoon Kim, Myoungsoo Jung |
DAC | 2 |
| 2019 | Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array ClustersabstractLarge-scale systems with all-flash arrays have become increasingly common in many computing segments. To make such systems resilient, we can adopt erasure coding such as Reed-Solomon (RS) code as an alternative to replication because erasure coding incurs a significantly lower storage overhead than replication. To understand the impact of using erasure coding on the system performance and other system aspects such as CPU utilization and network traffic, we build a storage cluster that consists of approximately 100 processor cores with more than 50 high-performance solid-state drives (SSDs), and evaluate the cluster with a popular open-source distributed parallel file system, called Ceph. Specifically, we analyze the behaviors of a system adopting erasure coding from the following five viewpoints, and compare with those of another system using replication: (1) storage system I/O performance; (2) computing and software overheads; (3) I/O amplification; (4) network traffic among storage nodes, and (5) impact of physical data layout on performance of RS-coded SSD arrays. For all these analyses, we examine two representative RS configurations, used by Google file systems, and compare them with triple replication employed by a typical parallel file system as a default fault tolerance mechanism. Lastly, we collect 96 block-level traces from the cluster and release them to the public domain for the use of other researchers. Sungjoon Koh, Jie Zhang 0048, Miryeong Kwon, Jungyeon Yoon, David Donofrio, Nam Sung Kim, Myoungsoo Jung |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Exploring System Challenges of Ultra-Low Latency Solid State Drives
Sungjoon Koh, Changrim Lee, Miryeong Kwon, Myoungsoo Jung |
HotStorage | 3 |
| 2018 | BIBIM: A Prototype Multi-Partition Aware Heterogeneous New Memory
Gyuyoung Park, Miryeong Kwon, Pratyush Mahapatra, Michael M. Swift, Myoungsoo Jung |
HotStorage | 2 |
| 2018 | Amber*: Enabling Precise Full-System Simulation with Detailed Modeling of All SSD ResourcesabstractSSDs become a major storage component in modern memory hierarchies, and SSD research demands exploring future simulation-based studies by integrating SSD subsystems into a full-system environment. However, several challenges exist to model SSDs under a full-system simulations; SSDs are composed upon their own complete system and architecture, which employ all necessary hardware, such as CPUs, DRAM and interconnect network. Employing the hardware components, SSDs also require to have multiple device controllers, internal caches and software modules that respect a wide spectrum of storage interfaces and protocols. These SSD hardware and software are all necessary to incarnate storage subsystems under full-system environment, which can operate in parallel with the host system. In this work, we introduce a new SSD simulation framework, SimpleSSD 2.0, namely Amber, that models embedded CPU cores, DRAMs, and various flash technologies (within an SSD), and operate under the full system simulation environment by enabling a data transfer emulation. Amber also includes full firmware stack, including DRAM cache logic, flash firmware, such as FTL and HIL, and obey diverse standard protocols by revising the host DMA engines and system buses of a popular full system simulator's all functional and timing CPU models (gem5). The proposed simulator can capture the details of dynamic performance and power of embedded cores, DRAMs, firmware and flash under the executions of various OS systems and hardware platforms. Using Amber, we characterize several system-level challenges by simulating different types of full-systems, such as mobile devices and general-purpose computers, and offer comprehensive analyses by comparing passive storage and active storage architectures. Donghyun Gouk, Miryeong Kwon, Jie Zhang 0048, Sungjoon Koh, Wonil Choi, Nam Sung Kim, Mahmut T. Kandemir, Myoungsoo Jung |
MICRO | 2 |
| 2018 | FlashShare: Punching Through Server Storage Stack from Kernel to Firmware for Ultra-Low Latency SSDs
Jie Zhang 0048, Miryeong Kwon, Donghyun Gouk, Sungjoon Koh, Changlim Lee, Mohammad Alian, Myoungjun Chun, Mahmut T. Kandemir, Nam Sung Kim, Jihong Kim 0001, Myoungsoo Jung |
OSDI | 2 |