Deukyeon Hwang

dblp:118/0844 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0001-3568-8322ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Focus! Fast On-disk Concurrency-control Using Sketches
abstract
Concurrency-control (CC) mechanisms are essential for ensuring consistency in large-scale key-value stores, but traditional approaches face significant challenges. Mechanisms like 2PL and OCC incur high CPU overheads. Timestamp-based mechanisms are faster but require storing timestamps for every key, resulting in substantial space overhead and numerous I/O operations in disk-based systems. We address these challenges by decomposing timestamp-based CC schemes into two components: a timestamp storage system and a CC protocol. We then show that the timestamp storage system can approximate timestamps for keys not used by ongoing transactions, substantially reducing memory requirements and I/O while maintaining correctness for various protocols (STO, MVTO, and TicToc). We introduce FPSketch, an approximate timestamp storage system, and our evaluation with SplinterDB shows that FPSketch outperforms 2PL and OCC by up to 14× on some workloads and disk-based CC systems by up to 5.9×. Remarkably, FPSketch with just 32KiB of memory yields performance comparable to an idealized in-memory implementations in our evaluation. FPSketch makes timestamp-based concurrency control mechanisms practical for disk-based key-value stores.
Deukyeon Hwang, Alexander Conway 0001, Carlos Garcia-Alvarado, Jun Yuan 0006, Naama Ben-David, Rob Johnson 0001, Adriana Szekeres
Proc. ACM Manag. Data1
2023 ScaleDB: A Scalable, Asynchronous In-Memory Database
Syed Akbar Mehdi, Deukyeon Hwang, Simon Peter 0001, Lorenzo Alvisi
OSDI2
2022 VeloxDFS: Streaming Access to Distributed Datasets to Reduce Disk Seeks
Sunghwan Ahn, Hyeongjun Park, Vicente A. B. Sanchez, Deukyeon Hwang, Wonbae Kim, Alan Sussman, Beomseok Nam
CCGRID4
2022 zIO: Accelerating IO-Intensive Applications with Transparent Zero-Copy IO
Tim Stamler, Deukyeon Hwang, Amanda Raybuck, Simon Peter 0001
OSDI2
2018 Endurable Transient Inconsistency in Byte-Addressable Persistent B+-Tree
Deukyeon Hwang, Wook-Hee Kim, Youjip Won, Beomseok Nam
FAST1
2017 EclipseMR: Distributed and Parallel Task Processing with Consistent Hashing
abstract
We present EclipseMR, a novel MapReduce framework prototype that efficiently utilizes a large distributed memory in cluster environments. EclipseMR consists of double-layered consistent hash rings - a decentralized DHT-based file system and an in-memory key-value store that employs consistent hashing. The in-memory key-value store in EclipseMR is designed not only to cache local data but also remote data as well so that globally popular data can be distributed across cluster servers and found by consistent hashing. In order to leverage large distributed memories and increase the cache hit ratio, we propose a locality-aware fair (LAF) job scheduler that works as the load balancer for the distributed in-memory caches. Based on hash keys, the LAF job scheduler predicts which servers have reusable data, and assigns tasks to the servers so that they can be reused. The LAF job scheduler makes its best efforts to strike a balance between data locality and load balance, which often conflict with each other. We evaluate EclipseMR by quantifying the performance effect of each component using several representative MapReduce applications and show EclipseMR is faster than Hadoop and Spark by a large margin for various applications.
Vicente A. B. Sanchez, Wonbae Kim, Youngmoon Eom, Kibeom Jin, Moohyeon Nam, Deukyeon Hwang, Jik-Soo Kim, Beomseok Nam
CLUSTER6
2015 EM-KDE: A locality-aware job scheduling policy with distributed semantic caches
Youngmoon Eom, Deukyeon Hwang, Jonghwan Moon, Minho Shin, Beomseok Nam
J. Parallel Distributed Comput.2
2014 Improving Multi-dimensional query processing with data migration in distributed cache infrastructure
abstract
In distributed query processing systems where caching infrastructure is distributed and scales with the number of servers, it is becoming more important to orchestrate and leverage a large number of cached objects in distributed caching systems seamlessly as the present trend is to build large scalable distributed systems by connecting small heterogeneous machines. With a large scale distributed caching system, a scheduling policy must consider both cache hit ratio and system load balance to optimize multiple queries. A scheduling policy that considers system load but not cache hit ratio often fails to reuse cached data by not assigning a query to the sever that has data objects the query needs. On the contrary, a scheduling policy that considers cache hit ratio but not system load balance may suffer from system load imbalance. To maximize the overall system throughput and to reduce query response time, a multiple query scheduling policy must balance system load and also leverage cached objects. In this paper, we present a distributed query processing framework that exhibits high cache hit ratio while achieving good system load balance. In order to seamlessly manage our distributed scalable caching system, our framework performs autonomic cached data migrations to improve cache hit ratio. Our experiments show that our proposed query scheduling policy and data migration policy significantly improve system throughput by achieving high cache hit ratio while avoiding system load imbalance.
Youngmoon Eom, Jinwoong Kim, Deukyeon Hwang, Jaewon Kwak, Minho Shin, Beomseok Nam
HiPC3
2012 High-throughput query scheduling with spatial clustering based on distributed exponential moving average
Beomseok Nam, Deukyeon Hwang, Jinwoong Kim, Minho Shin
Distributed Parallel Databases2