Safdar Jamil

dblp:267/9736 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-9011-6431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 DEDUPKV: A Space-Efficient and High-Performance Key-Value Store via Fine-Grained Deduplication
abstract
Log-Structured Merge Tree (LSM-tree) based key-value stores excel in write-intensive environments but suffer from data duplication, consuming up to 49% of storage space in LSMtree-based key-value store deployments.Traditional solutions like compression and coarse-grained file system-level deduplication introduce overhead or have limited effectiveness.In this study, we propose DedupKV, a fine-grained deduplication framework tailored for LSM-tree, maximizing data reduction efficiency while minimizing write stalls and read overheads.DedupKV features three key innovations:(1) FLUSH-integrated inline deduplication, which removes duplicates during memory-to-storage writes; (2) WAL file-based offline deduplication, repurposing write-ahead logs to avoid double writes; and (3) elastic execution, dynamically balancing inline and offline deduplication based on memory pressure and workload intensity.Additionally, dynamic granularity management reduces deduplication metadata overhead.We implemented these four ideas in RocksDB for the first time and conducted experiments in a Linux environment.Our evaluation shows that WAL file-based offline deduplication and DedupKV outperform BlobDB by 33% and 23%, respectively, in write-heavy workloads, while reducing write amplification by 1.2×, 2×, and 1.6× for real KV datasets.
Safdar Jamil, Awais Khan 0002, Xubin He, Youngjae Kim 0001
ICS1
2025 KVACCEL: A Novel Write Accelerator for LSM-Tree-Based KV Stores with Host-SSD Collaboration
abstract
Log-Structured Merge (LSM) tree-based Key-Value Stores (KVSs) are widely adopted for their high performance in write-intensive environments, but they often face performance degradation due to write stalls during compaction. Prior solutions, such as regulating I/O traffic or using multiple compaction threads, can cause unexpected drops in throughput or increase host CPU usage, while hardware-based approaches using FPGA, GPU, and DPU aimed at reducing compaction duration introduce additional hardware costs. In this study, we propose KVACCEL, a novel hardware-software co-design framework that bypasses write stalls by leveraging a dual-interface SSD. KVACCEL allocates logical NAND flash space to support both block and key-value interfaces, using the key-value interface as a temporary write buffer during write stalls. This strategy significantly reduces write stalls, optimizes resource usage, and ensures consistency between the host and device by implementing an in-device LSM-based write buffer with an iterator-based range scan mechanism. Our extensive evaluation shows that for write-intensive workloads, KVACCEL outperforms ADOC by up to 17 % in terms of throughput and performance-to-CPU-utilization efficiency. For mixed read-write workloads, both demonstrate comparable performance.
Hyunsun Chung, Seonghoon Ahn, Junhyeok Park 0002, Safdar Jamil, Hongsu Byun, Myungcheol Lee, Jinchun Choi, Youngjae Kim 0001
IPDPS5
2024 Towards A Unified Garbage Collection Strategy in ZNS Key-Value Store File Systems Using Same-Victim GC
abstract
Zoned Namespace (ZNS) SSDs are gaining traction for eliminating in-device GC and enabling application-aware data management. BlobDB, an enhanced RocksDB with key-value separation, reduces compaction overhead but suffers from the GC over GC (GoG) problem, causing redundant data copying during BlobDB’s GC and Zone Cleaning (ZC). To address this, this paper proposes Same-Victim GC, aligning the victims and sizes of both GCs. Specifically, we introduce the BlobDB-Aware Zone Allocation (BAZA) algorithm to allocate blob files by creation order, eliminate victim file mismatch between two GCs, and $Z_{-} C u t o f f$ to minimize BlobDB’s GC size without additional overhead. Implemented in ZenFS v2.14 and RocksDB v7.4, our solution eliminates valid data copying, doubles compaction performance, and improves space utilization by $1.28 \times$.
Hamin Hwangbo, Joseph Ro, Sungjin Byeon, Safdar Jamil, Jun Young Han, Jooyoung Hwang, Youngjae Kim 0001
MASCOTS4
2024 An Analytical Model-based Capacity Planning Approach for Building CSD-based Storage Systems
abstract
The data movement in large-scale computing facilities (from compute nodes to data nodes) is categorized as one of the major contributors to high cost and energy utilization. To tackle it, in-storage processing (ISP) within storage devices, such as Solid-State Drives (SSDs), has been explored actively. The introduction of computational storage drives (CSDs) enabled ISP within the same form factor as regular SSDs and made it easy to replace SSDs within traditional compute nodes. With CSDs, host systems can offload various operations such as search, filter, and count. However, commercialized CSDs have different hardware resources and performance characteristics. Thus, it requires careful consideration of hardware, performance, and workload characteristics for building a CSD-based storage system within a compute node. Therefore, storage architects are hesitant to build a storage system based on CSDs as there are no tools to determine the benefits of CSD-based compute nodes to meet the performance requirements compared to traditional nodes based on SSDs. In this work, we proposed an analytical model-based storage capacity planner called CsdPlan for system architects to build performance-effective CSD-based compute nodes. Our model takes into account the performance characteristics of the host system, targeted workloads, and hardware and performance characteristics of CSDs to be deployed and provides optimal configuration based on the number of CSDs for a compute node. Furthermore, CsdPlan estimates and reduces the total cost of ownership (TCO) for building a CSD-based compute node. To evaluate the efficacy of CsdPlan , we selected two commercially available CSDs and four representative big data analysis workloads.
Hongsu Byun, Safdar Jamil, Jungwook Han, Sungyong Park, Myungcheol Lee, Changsoo Kim, Beongjun Choi, Youngjae Kim 0001
ACM Trans. Embed. Comput. Syst.2
2023 A Free-Space Adaptive Runtime Zone-Reset Algorithm for Enhanced ZNS Efficiency
abstract
While the state-of-the-art runtime zone-reset algorithm of the ZenFS in RocksDB is optimized for the performance of Zone Namespace (ZNS) SSDs, it does not take into account the lifetime constraint of ZNS SSDs. To address this issue, we present FAR, a Free-space Adaptive Runtime Zone-Reset algorithm for ZenFS, which dynamically adjusts the frequency of runtime zone-reset calls based on the available free-space in the ZNS SSD. We developed FAR with the ZenFS of RocksDB using a ZNS SSD prototype based on the Cosmos+ OpenSSD platform and compared it with the state-of-the-art runtime zone-reset algorithm used in ZenFS. Our extensive evaluations demonstrate that FAR improves the lifetime of ZNS SSD by 2x without compromising performance.
Sungjin Byeon, Joseph Ro, Safdar Jamil, Jeong-Uk Kang, Youngjae Kim 0001
HotStorage3
2020 A NUMA-aware NVM File System Design for Manycore Server Applications
abstract
NOVA, a state-of-the-art NVM-based file system, is known to have scalability bottlenecks when multiple I/O threads read/write data simultaneously. Recent studies have identified the cause as the coarse-grained lock adopted by NOVA to provide consistency, and proposed fine-grained range-based locks to improve the scalability of NOVA. However, these variants of NOVA only scale on Uniform Memory Access (UMA) architecture and do not scale on Non-Uniform Memory Access (NUMA) architecture. This is because NOVA has no NUMA-aware memory allocation policy and still uses non-scalable file data structures. In this paper, we propose a NUMA-aware NOVA file system which virtualizes the NVM devices located across NUMA nodes so that they can be used as a single address space. The proposed file system adopts a local-first placement policy where file data and metadata are placed preferentially on the local NVM device to reduce the remote access problem. In addition, the lock-free per-core data structures proposed in this file system allow data to be updated concurrently while mitigating the remote memory access. Extensive evaluations show that our NUMA-aware NOVA for parallel writing is scalable with respect to the increased core count and outperforms vanilla NOVA by 2.56-19.18 times.
June-Hyung Kim, Youngjae Kim 0001, Safdar Jamil, Sungyong Park
MASCOTS3