Duck-Ho Bae

dblp:12/7623 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-authorSystems, architecture and hardware · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 90% Cloud and datacenter computing · 5% Memory systems · 4%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 42% Query processing and optimization · 42% Transaction processing and concurrency control · 16%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › computational storage
compaction offloading
1.012026
Near-Data Compaction for LSM Tree on Rack-Scale Disaggregated Storage · IEEE Trans. Parallel Distributed Syst. 2026
Storage systems › distributed storage
disaggregated storage
1.012026
Near-Data Compaction for LSM Tree on Rack-Scale Disaggregated Storage · IEEE Trans. Parallel Distributed Syst. 2026
Storage systems › key-value storage
LSM-tree
1.012026
Near-Data Compaction for LSM Tree on Rack-Scale Disaggregated Storage · IEEE Trans. Parallel Distributed Syst. 2026
Storage systems › data reduction
data deduplication
0.712023
TiDedup: A New Distributed Deduplication Architecture for Ceph · USENIX ATC 2023
Storage systems › data reduction › data deduplication
distributed deduplication
0.712023
TiDedup: A New Distributed Deduplication Architecture for Ceph · USENIX ATC 2023
Storage systems › flash and SSD
solid-state drive
0.622018
2B-SSD: The Case for Dual, Byte- and Block-Addressable Solid-State Drives · ISCA 2018
Biscuit: A Framework for Near-Data Processing of Big Data Workloads · ISCA 2016
Storage systems › computational storage
in-storage computing
0.522016
YourSQL: A High-Performance Database System Leveraging In-Storage Computing · Proc. VLDB Endow. 2016
Biscuit: A Framework for Near-Data Processing of Big Data Workloads · ISCA 2016
Storage systems › non-volatile memory storage
byte-addressable storage
0.312018
2B-SSD: The Case for Dual, Byte- and Block-Addressable Solid-State Drives · ISCA 2018
Storage systems › i/o architecture › i/o subsystem
storage interfaces
0.312018
2B-SSD: The Case for Dual, Byte- and Block-Addressable Solid-State Drives · ISCA 2018
Information retrieval
filtering
0.212016
Biscuit: A Framework for Near-Data Processing of Big Data Workloads · ISCA 2016
Query processing and optimization › query execution › scan processing
scan acceleration
0.212016
YourSQL: A High-Performance Database System Leveraging In-Storage Computing · Proc. VLDB Endow. 2016
Storage systems
computational storage
0.212016
Biscuit: A Framework for Near-Data Processing of Big Data Workloads · ISCA 2016
Memory systems › processing-in-memory
near-data processing
0.212016
Biscuit: A Framework for Near-Data Processing of Big Data Workloads · ISCA 2016
Transaction processing and concurrency control
logging
0.112018
2B-SSD: The Case for Dual, Byte- and Block-Addressable Solid-State Drives · ISCA 2018

Methods — techniques the papers use, named apart from their topics

selective compaction admission · 1.0dual-node coordination · 1.0memory-mapped i/o · 0.7dual address space mapping · 0.7solid-state drive offloading · 0.5dataflow programming · 0.5
YearPublicationVenuePosition
2026 Near-Data Compaction for LSM Tree on Rack-Scale Disaggregated Storage
abstract
In LSM trees, background compaction tasks contend with foreground queries for CPU cycles, cache space, and also SAN bandwidth, if deployed in a disaggregated storage architecture. This study proposesNear-Data Compaction(NDC), which executes compaction on the storage node to utilize its underutilized computing resources. However, enabling NDC introduces several challenges. First, it has to support concurrent file access from both compute and storage nodes. In addition, it has to decide which compaction tasks to be executed on the storage node since the computing resources of a storage node are not unlimited. This study presentsTetherDB, an LSM tree for disaggregated storage architecture that addresses these challenges through lightweight dual-node coordination and selective NDC admission policies. Our evaluation demonstrates that TetherDB improves throughput by up to 2.1× compared to RocksDB in write-heavy workloads.
Sungho Moon, Daegyu Han, Hera Koo, Sangeun Chae, Duck-Ho Bae, Euiseong Seo, Beomseok Nam
IEEE Trans. Parallel Distributed Syst.5
2023 NVMe-Driven Lazy Cache Coherence for Immutable Data with NVMe over Fabrics
abstract
In this work, we explore opportunities to design shared storage systems that leverage the distance connectivity of NVMe over Fabrics (NVMe-oF). NVMe-oF enables the use of NVMe storage devices in a shared storage environment, where multiple servers can access the same storage device via RDMA. Leveraging the distance connectivity of NVMe-oF, we develop a shared file system called EXT4-oF by extending the EXT4 file system. EXT4-oF uses RDMA to enable a local file system to function as a shared file system without requiring remote daemon processes. EXT4-oF employs a novel NVMe-driven lazy cache coherence to maintain cache coherence of file system metadata across multiple compute nodes upon creating new files, all achieved without the need for any daemon processes. To ensure cache coherence, NVMe-driven lazy cache coherence mechanism requires compute nodes to perform a re-read of NVMe-oF to avoid false negative file open errors. Through our experiments, we demonstrate that EXT4-oF improves the performance of MinIO by minimizing network traffic between compute and storage nodes and eliminating the need for TCP/IP communication during remote reads.
Hyeongjun Jeon, Daegyu Han, Duck-Ho Bae, Youngjin Yu, Kyeungpyo Kim, Sung-Soon Park 0001, Jinkyu Jeong, Beomseok Nam
CLOUD4
2023 TiDedup: A New Distributed Deduplication Architecture for Ceph
Myoungwon Oh, Samuel Just, Youngjin Yu, Duck-Ho Bae, Sage A. Weil, Sangyeun Cho, Heon Young Yeom
USENIX ATC5
2018 2B-SSD: The Case for Dual, Byte- and Block-Addressable Solid-State Drives
abstract
Performance critical transaction and storage systems require fast persistence of write data. Typically, a non-volatile RAM (NVRAM) is employed on the datapath to the permanent storage, to temporarily and quickly store write data before the system acknowledges the write request. NVRAM is commonly implemented with battery-backed DRAM. Unfortunately, battery-backed DRAM is small and costly, and occupies a precious DIMM slot. In this paper, we make a case for dual, byte- and block-addressable solid-state drive (2B-SSD), a novel NAND flash SSD architecture designed to offer a dual view of byte addressability and traditional block addressability at the same time. Unlike a conventional storage device, 2B-SSD allows accessing the same file with two independent byte- and block-I/O paths. It controls the data transfer between its internal DRAM and NAND flash memory through an intuitive software interface, and manages the mapping of the two address spaces. 2B-SSD realizes a wholly different way and speed of accessing files on a storage device; applications can access them directly using memory-mapped I/O, and moreover write with a DRAM-like latency. To quantify the benefits of 2B-SSD, we modified logging subsystems of major database engines to store log records directly on it without buffering them in the host memory. When running popular workloads, we measured throughput gains in the range of 1.2X and 2.8X with no risk of data loss.
Duck-Ho Bae, Insoon Jo, Youra Choi, Joo Young Hwang, Sangyeun Cho, Daniel D. G. Lee
ISCA1
2016 Biscuit: A Framework for Near-Data Processing of Big Data Workloads
abstract
Data-intensive queries are common in business intelligence, data warehousing and analytics applications. Typically, processing a query involves full inspection of large in-storage data sets by CPUs. An intuitive way to speed up such queries is to reduce the volume of data transferred over the storage network to a host system. This can be achieved by filtering out extraneous data within the storage, motivating a form of near-data processing. This work presents Biscuit, a novel near-data processing framework designed for modern solid-state drives. It allows programmers to write a data-intensive application to run on the host system and the storage system in a distributed, yet seamless manner. In order to offer a high-level programming model, Biscuit builds on the concept of data flow. Data processing tasks communicate through typed and data-ordered ports. Biscuit does not distinguish tasks that run on the host system and the storage system. As the result, Biscuit has desirable traits like generality and expressiveness, while promoting code reuse and naturally exposing concurrency. We implement Biscuit on a host system that runs the Linux OS and a high-performance solid-state drive. We demonstrate the effectiveness of our approach and implementation with experimental results. When data filtering is done by hardware in the solid-state drive, the average speed-up obtained for the top five queries of TPC-H is over 15x.
Boncheol Gu, Andre S. Yoon, Duck-Ho Bae, Insoon Jo, Jonghyun Yoon, Jeong-Uk Kang, Moonsang Kwon, Chanho Yoon, Sangyeun Cho, Duckhyun Chang
ISCA3
2016 YourSQL: A High-Performance Database System Leveraging In-Storage Computing
abstract
This paper presents YourSQL , a database system that accelerates data-intensive queries with the help of additional in-storage computing capabilities. YourSQL realizes very early filtering of data by offloading data scanning of a query to user-programmable solid-state drives. We implement our system on a recent branch of MariaDB (a variant of MySQL). In order to quantify the performance gains of YourSQL, we evaluate SQL queries with varying complexities. Our result shows that YourSQL reduces the execution time of the whole TPC-H queries by 3.6×, compared to a vanilla system. Moreover, the average speed-up of the five TPC-H queries with the largest performance gains reaches over 15×. Thanks to this significant reduction of execution time, we observe sizable energy savings. Our study demonstrates that the YourSQL approach, combining the power of early filtering with end-to-end datapath optimization, can accelerate large-scale analytic queries with lower energy consumption.
Insoon Jo, Duck-Ho Bae, Andre S. Yoon, Jeong-Uk Kang, Sangyeun Cho, Daniel D. G. Lee
Proc. VLDB Endow.2
2015 Efficient Sparse Matrix Multiplication on GPU for Large Social Network Analysis
abstract
As a number of social network services appear online recently, there have been many attempts to analyze social networks for extracting valuable information. Most existing methods first represent a social network as a quite sparse adjacency matrix, and then analyze it through matrix operations such as matrix multiplication. Due to the large scale and high complexity, efficient processing multiplications is an important issue in social network analysis. In this paper, we propose a GPU-based method for efficient sparse matrix multiplication through the parallel computing paradigm. The proposed method aims at balancing the amount of workload both at fine- and coarse-grained levels for maximizing the degree of parallelism in GPU. Through extensive experiments using synthetic and real-world datasets, we show that the proposed method outperforms previous methods by up to three orders-of-magnitude.
Yong-Yeon Jo, Sang-Wook Kim, Duck-Ho Bae
CIKM3
2015 Analyzing Topological Characteristics of The Korean Blogosphere
Jiwoon Ha, Duck-Ho Bae, Minsoo Ryu, Sang-Wook Kim, Seok-Chul Baek, Byeong-Soo Jeong, Jin-Soo Cho
J. Web Eng.2
2014 On Constructing Seminal Paper Genealogy
abstract
Let us consider that someone is starting a research on a topic that is unfamiliar to them. Which seminal papers have influenced the topic the most? What is the genealogy of the seminal papers in this topic? These are the questions that they can raise, which we try to answer in this paper. First, we propose an algorithm that finds a set of seminal papers on a given topic. We also address the performance and scalability issues of this sophisticated algorithm. Next, we discuss the measures to decide how much a paper is influenced by another paper. Then, we propose an algorithm that constructs a genealogy of the seminal papers by using the influence measure and citation information. Finally, through extensive experiments with a large volume of a real-world academic literature data, we show the effectiveness and efficiency of our approach.
Duck-Ho Bae, Se-Mi Hwang, Sang-Wook Kim, Christos Faloutsos
IEEE Trans. Cybern.1
2013 Intelligent SSD: a turbo for big data mining
abstract
This paper introduces the notion of intelligent SSDs. First, we present the design considerations of intelligent SSDs, and then examine their potential benefits under various settings in data mining applications.
Duck-Ho Bae, Jin-Hyung Kim, Sang-Wook Kim, Hyunok Oh, Chanik Park
CIKM1
2012 Outlier detection using centrality and center-proximity
abstract
An outlier is an object that is considerably dissimilar with the remainder of the dataset. In this paper, we first propose the notion of centrality and center-proximity as novel outlierness measures which can be considered to represent the characteristics of all of the objects in the dataset. We then propose a graph-based outlier detection method which can solve the problems of local density, micro-cluster, and fringe objects. Finally, through extensive experiments, we show the effectiveness of the proposed method.
Duck-Ho Bae, Seo Jeong, Sang-Wook Kim, Minsoo Lee
CIKM1
2012 An efficient method for record management in flash memory environment
Duck-Ho Bae, Ji-Woong Chang, Sang-Wook Kim
J. Syst. Archit.1
2011 Constructing seminal paper genealogy
abstract
When a researcher starts with a new topic, it would be very useful if seminal papers in the topic and their relationships are provided in advance. We propose an approach to construct seminal paper genealogy and show the effectiveness and efficiency of our approach.
Duck-Ho Bae, Se-Mi Hwang, Sang-Wook Kim, Christos Faloutsos
CIKM1