EDBT 2026 Demo / reviewers in the wild / expert
Jongsung Lee 0001
dblp:37/10922-1
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0003-4080-0611ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An LSM Tree Augmented with B+ Tree on Nonvolatile MemoryabstractModern log-structured merge (LSM) tree-based key-value stores are widely used to process update-heavy workloads effectively as the LSM tree sequentializes write requests to a storage device to maximize storage performance. However, this append-only approach leaves many outdated copies of frequently updated key-value pairs, which need to be routinely cleaned up through the operation called compaction . When the system load is modest, compaction happens in background. However, at a high system load, it can quickly become the major performance bottleneck. To address this compaction bottleneck and further improve the write throughput of LSM tree-based key-value stores, we propose LAB-DB, which augments the existing LSM tree with a pair of B + trees on byte-addressable nonvolatile memory (NVM). The auxiliary B + trees on NVM reduce both compaction frequency and compaction time, hence leading to lower compaction overhead for writes and fewer storage accesses for reads. According to our evaluation of LAB-DB on RocksDB with YCSB benchmarks, LAB-DB achieves 94% and 67% speedups on two write-intensive workloads (Workload A and F), and also a 43% geomean speedup on read-intensive YCSB Workload B, C, D, and E. This performance gain comes with a low cost of NVM whose size is just 0.6% of the entire dataset to demonstrate the scalability of LAB-DB with an ever increasing volume of future datasets. Jongsung Lee 0001, Keun Soo Lim, Jun Heo 0001, Tae Jun Ham, Jae W. Lee |
ACM Trans. Storage | 2 |
| 2023 | FlowKV: A Semantic-Aware Store for Large-Scale State Management of Stream Processing EnginesabstractWe propose FlowKV, a persistent store tailored for large-scale state management of streaming applications. Unlike existing KV stores, FlowKV leverages information from stream processing engines by taking a principled approach toward exploiting information about how and when the applications access data. FlowKV categorizes data access patterns of window operations according to how window boundaries are set and how tuples inside a window are aggregated, and deploys customized in-memory and on-disk data structures optimized for each pattern. In addition, FlowKV takes window metadata as explicit arguments of read and write methods to predict the moment when a window is read, and then loads the tuples of windows in batches from storage ahead of time. Using the NEXMark benchmark as workload, our experiments show that Apache Flink on FlowKV outperforms Flink on RocksDB or Faster with up to 4.12× throughput gain. Gyewon Lee, Jaewoo Maeng, Jinsol Park, Jangho Seo, Haeyoon Cho 0001, Youngseok Yang, Taegeon Um, Jongsung Lee 0001, Jae W. Lee, Byung-Gon Chun |
EuroSys | 8 |
| 2023 | DRAM Translation Layer: Software-Transparent DRAM Power Savings for Disaggregated MemoryabstractMemory disaggregation is a promising solution to scale memory capacity and bandwidth shared by multiple server nodes in a flexible and cost-effective manner. DRAM power consumption, which is reported to be around 40% of the total system power in the datacenter server, will become an even more serious concern in this high-capacity environment. Exploiting the low average utilization of DRAM capacity in today's datacenters, it is appealing to put unallocated/cold DRAM ranks into a power-saving mode. However, the conventional DRAM address mapping with fine-grained interleaving to maximize rank-level parallelism is incompatible with such rank-level DRAM power management techniques. Furthermore, existing DRAM power-saving techniques often require intrusive changes to the system stack, including OS, memory controller (MC), or even DRAM devices, to pose additional challenges for deployment. Thus, we propose DRAM Translation Layer (DTL) for host software/MC-transparent DRAM power management with commodity DRAM devices. Inspired by Flash Translation Layer (FTL) in modern SSDs, DTL is placed in the CXL memory controller to provide (i) flexible address mappings between host physical address and DRAM device physical address and (ii) host-transparent memory page migration. Leveraging DTL, we propose two DRAM power-saving techniques with different temporal granularities to maximize the number of DRAM ranks that can enter low-power states while provisioning sufficient DRAM bandwidth: rank-level power-down and hotness-aware self-refresh. The first technique consolidates unallocated memory pages into a subset of ranks at deallocation of a virtual machine (VM) and turns them off transparently to both OS and host MC. Our evaluation with CloudSuite benchmarks demonstrates that this technique saves DRAM power by 31.6% on average at a 1.6% performance cost. The hotness-aware self-refresh scheme further reduces DRAM energy consumption by up to 14.9% with negligible performance loss via opportunistically migrating cold pages into a rank and making it enter self-refresh mode. Wenjing Jin 0001, Wonsuk Jang, Haneul Park, Jongsung Lee 0001, Soosung Kim 0001, Jae W. Lee |
ISCA | 4 |
| 2023 | WALTZ: Leveraging Zone Append to Tighten the Tail Latency of LSM Tree on ZNS SSDabstractWe propose WALTZ, an LSM tree-based key-value store on the emerging Zoned Namespace (ZNS) SSD. The key contribution of WALTZ is to leverage the zone append command, which is a recent addition to ZNS SSD specifications, to provide tight tail latency. The long tail latency problem caused by the merging process of multiple parallel writes, called batch-group writes, is effectively addressed by the internal synchronization mechanism of ZNS SSD. To provide fast failover when the active zone becomes full for a write-ahead log (WAL) file during parallel append, WALTZ introduces a mechanism for WAL zone replacement and reservation. Finally, lazy metadata management allows a put query to be processed fast without requiring any other synchronizations to enable lock-free execution of individual append commands. For evaluation we use both mi-crobenchmarks (db_bench) with varying read/write ratios and key skewnesses, and realistic social-graph workloads (MixGraph from Facebook). Our evaluation demonstrates geomean reduction of tail latency by 2.19× and 2.45× for db_bench and MixGraph, respectively, with a maximum reduction of 3.02× and 4.73×. As a side effect of eliminating the overhead of batch-group writes, WALTZ also improves the query throughput (QPS) by up to 11.7%. Jongsung Lee 0001, Dong Uk Kim, Jae W. Lee |
Proc. VLDB Endow. | 1 |
| 2021 | FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural Networks
Jonghyun Bae, Jongsung Lee 0001, Yunho Jin, Sam Son, Shine Kim, Hakbeom Jang, Tae Jun Ham, Jae W. Lee |
FAST | 2 |
| 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingabstractTo meet surging demands for deep learning inference services, many cloud computing vendors employ high-performance specialized accelerators, called neural processing units (NPUs). One important challenge for effective use of NPUs is to achieve high resource utilization over a wide spectrum of deep neural network (DNN) models with diverse arithmetic intensities. There is often an intrinsic mismatch between the compute-to-memory bandwidth ratio of an NPU and the arithmetic intensity of the model it executes, leading to under-utilization of either compute resources or memory bandwidth. Ideally, we want to saturate both compute TOP/s and DRAM bandwidth to achieve high system throughput. Thus, we propose Layerweaver, an inference serving system with a novel multi-model time-multiplexing scheduler for NPUs. Layerweaver reduces the temporal waste of computation resources by interweaving layer execution of multiple different models with opposing characteristics: compute-intensive and memory-intensive. Layerweaver hides the memory time of a memory-intensive model by overlapping it with the relatively long computation time of a compute-intensive model, thereby minimizing the idle time of the computation units waiting for off-chip data transfers. For a two-model serving scenario of batch 1 with 16 different pairs of compute- and memory-intensive models, Layerweaver improves the temporal utilization of computation units and memory channels by 44.0% and 28.7%, respectively, to increase the system throughput by 60.1% on average, over the baseline executing one model at a time. Young H. Oh, Seonghak Kim, Yunho Jin, Sam Son, Jonghyun Bae, Jongsung Lee 0001, Yeonhong Park, Dong Uk Kim, Tae Jun Ham, Jae W. Lee |
HPCA | 6 |
| 2018 | OrcFS: Orchestrated File System for Flash StorageabstractIn this work, we develop the Orchestrated File System (OrcFS) for Flash storage. OrcFS vertically integrates the log-structured file system and the Flash-based storage device to eliminate the redundancies across the layers. A few modern file systems adopt sophisticated append-only data structures in an effort to optimize the behavior of the file system with respect to the append-only nature of the Flash memory. While the benefit of adopting an append-only data structure seems fairly promising, it makes the stack of software layers full of unnecessary redundancies, leaving substantial room for improvement. The redundancies include (i) redundant levels of indirection (address translation), (ii) duplicate efforts to reclaim the invalid blocks (i.e., segment cleaning in the file system and garbage collection in the storage device), and (iii) excessive over-provisioning (i.e., separate over-provisioning areas in each layer). OrcFS eliminates these redundancies via distributing the address translation, segment cleaning (or garbage collection), bad block management, and wear-leveling across the layers. Existing solutions suffer from high segment cleaning overhead and cause significant write amplification due to mismatch between the file system block size and the Flash page size. To optimize the I/O stack while avoiding these problems, OrcFS adopts three key technical elements. First, OrcFS uses disaggregate mapping , whereby it partitions the Flash storage into two areas, managed by a file system and Flash storage, respectively, with different granularity. In OrcFS, the metadata area and data area are maintained by 4Kbyte page granularity and 256Mbyte superblock granularity. The superblock-based storage management aligns the file system section size, which is a unit of segment cleaning, with the superblock size of the underlying Flash storage. It can fully exploit the internal parallelism of the underlying Flash storage, exploiting the sequential workload characteristics of the log-structured file system. Second, OrcFS adopts quasi-preemptive segment cleaning to prohibit the foreground I/O operation from being interfered with by segment cleaning. The latency to reclaim the free space can be prohibitive in OrcFS due to its large file system section size, 256Mbyte. OrcFS effectively addresses this issue via adopting a polling-based segment cleaning scheme. Third, the OrcFS introduces block patching to avoid unnecessary write amplification in the partial page program. OrcFS is the enhancement of the F2FS file system. We develop a prototype OrcFS based on F2FS and server class SSD with modified firmware (Samsung 843TN). OrcFS reduces the device mapping table requirement to 1/465 and 1/4 compared with the page mapping and the smallest mapping scheme known to the public, respectively. Via eliminating the redundancy in the segment cleaning and garbage collection, the OrcFS reduces 1/3 of the write volume under heavy random write workload. OrcFS achieves 56% performance gain against EXT4 in varmail workload. Jinsoo Yoo, Joontaek Oh, Seongjin Lee, Youjip Won, Jinyong Ha 0001, Jongsung Lee 0001, Junseok Shim |
ACM Trans. Storage | 6 |
| 2016 | An Empirical Evaluation of Enterprise and SATA-Based Transactional Solid-State DrivesabstractIn most file systems, performance is usually sacrificed in exchange for crash consistency, which ensures that data and metadata are restored consistently in the event of a system crash. To escape this trade-off between performance and crash consistency, recent researchers designed and implemented the transactional functionality inside Solid State Drives (SSDs). However, in order to investigate its benefit in a more realistic and standard fashion, this scheme should be re-evaluated in enterprise storage with standard interface. This paper explores the challenges and implications of a transactional SSD with extensive experiments. To evaluate the potential benefit of transactional SSD, we design and implement the transaction functionality in Samsung enterprise-class and SATA-based SSD (i.e., SM843TN) and name it TxSSD. We then modify the existing file systems (i.e., ext4 and btrfs) on topof TxSSD, making both file systems crash-consistent without redundant writes. We perform performance evaluation of two filesystems by using file I/O and OLTP benchmarks with a database. We also disclose and analyze the overhead of transactional functionality inside SSD. The experimental results show that TxSSD-aware file systems exhibit better performance compared to crash-consistent modes (i.e., data journaling mode of ext4 and cow mode of btrfs) but worse performance compared to weak consistent modes (i.e., ordered mode of ext4 and no datacow mode of btrfs). Yongseok Son, Hara Kang, Jinyong Ha 0001, Jongsung Lee 0001, Hyuck Han, Hyungsoo Jung 0001, Heon Young Yeom |
MASCOTS | 4 |
| 2013 | An empirical study of hot/cold data separation policies in solid state drives (SSDs)abstractSeparating hot data from cold data is known to allow for efficient management of NAND flash memory in Solid State Drives (SSDs). However, most of previous work has been evaluated with the trace-driven simulations under different workloads and testing conditions. The goal of this paper is to empirically study the performance, computation overhead, and memory consumption of the existing hot/cold data separation policies on a real SSD platform. After devising a general framework where a different policy can be easily plugged in, we have evaluated three hot/cold data separation policies: 2-level LRU (LRU), Multiple Bloom Filter (MBF), and Dynamic dAta Clustering (DAC). Our evaluation results show that DAC performs best, improving the performance by up to 58% in real workloads with a reasonable computation and memory overhead. Jongsung Lee 0001, Jin-Soo Kim 0001 |
SYSTOR | 1 |