Yekang Zhan

dblp:333/1868 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0006-1057-5918ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rearchitecting Buffered I/O in the Era of High-Bandwidth SSDs
Yekang Zhan, Tianze Wang, Zheng Peng 0017, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001
FAST1
2026 eLDPC: An Elastic and Scalable LDPC-Decoder With Early Termination by Effectively Leveraging High-Level Synthesis
abstract
Emerging communication and storage embrace Low-Density Parity-Check (LDPC) codes to fully exploit their physical channels. FPGA (Field-Programmable Gate Array) is widely employed to fast prototype and accelerate the LDPC decoding with high complexity. For varying channel conditions, the FGPA decoder is desired to elastically stop iteration when meeting success condition, avoiding conservatively performing a predefined and large number of iterations. However, the dynamical-execution algorithms with adjustable parameters generally are challenging for scalable decoder structure preferred to deterministic execution logic. To overcome the problem, this paper presents an elastic and scalable HLS-based FPGA LDPC decoder architecture with early-termination to achieve high throughput and flexibility. To this end, eLDPC first provides a universal operation, fully leveraging the features of HLS to efficiently implement optimized small-scale hardware units for low-level data-update operations. Second, eLDPC presents a decoding-iteration pipeline that adds a termination-check stage to terminate the following iteration for current codeword decoding. eLDPC also presents an HLS-enhanced approach to address memory access conflicts associated with the DU pipeline. Further, eLDPC extends the number of DU decoding-iteration pipelines within a single stream to decode multiple codewords in parallel. Third, eLDPC designs elastic and independent multiple decoding streams by using FIFO queues to decouple Input, Output, and a decoding unit (DU) with variable iterations while avoiding the potential blockage of the queueing. We implement and evaluate eLDPC on a Xilinx U55C. Experiments show that eLDPC outperforms recent decoders by up to 5× with the same parameter and achieves the actual decoding throughput of up to 49.5 Gbps with high scalability and flexibility.
Qiang Cao 0001, Yifan Zhang 0012, Yekang Zhan, Jie Yao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Rethinking the Request-to-IO Transformation Process of File Systems for Full Utilization of High-Bandwidth SSDs
Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001
FAST1
2025 Repo: Proactive Swapping Exploiting Loop Patterns in Modern Applications
abstract
Modern data-intensive applications such as large language models already outrun affordable DRAM. Page swapping to fast SSDs or network-attached memory adds capacity, but existing operating system policies often struggle when an application's working set shifts, causing costly page faults and degrading performance. Meanwhile, many applications iterate over large data objects in regular loops, which is favorable for optimization. But existing eviction and prefetch policies largely miss this opportunity, because they rely on short-term recency and reactive prefetching, leading to memory thrashing and massive uncovered faults. This paper proposes Repo, a novel swap policy that identifies and exploits intrinsic loops in modern applications. Repo utilizes PEBS-based sampling and clustering for rapid and accurate delineation of loop elements, confirms loops reliably using stable load/store counts, and crucially, coordinates eviction and prefetching proactively. Experiments on real applications show that Repo reduces page faults by up to 98 % and execution time by as much as 78 % relative to state-of-the-art baselines.
Qiang Cao 0001, Yekang Zhan, Jie Yao 0001
ICCD3
2025 HeteroGNN: A Heterogeneous Stage Division Based GNN Training Framework to Maximize CPU-GPU Parallelism
abstract
Graph Neural Networks (GNNs) have become inevitable tools for extracting knowledge from massive topological structure data. However, experimental observation shows that existing GNN training frameworks exhibit low efficiency when performing memory-access-intensive data preparation stage upon CPU and computation-intensive model training stage upon GPU. This is largely due to data dependency restriction between the two stages and sequential execution of computation in training iterations. Based on these observations, this paper proposes HeteroGNN, an efficient GNN training framework, to maximize parallelism of GNN training upon heterogeneous CPU-GPU architecture. Specifically, HeteroGNN first proposes a data-dependency-aware stage division policy, which divides the two stages to six phases to offer inter-stage parallelism. Further, HeteroGNN establishes a fine-grained computation partition schema, which actively partitions typical GNN computing operations into multiple schedulable tasks suitable for CPU and GPU. Finally, HeteroGNN designs an adaptive task scheduler, which adaptively schedules the tasks upon six phases to maximize GPU efficiency. Experiment results demonstrate that HeteroGNN speeds up end-to-end training time up to 1.29× comparing to state-of-the-art GNN framework, PiPAD, and 1.30× to 2.06× comparing to DGL and PyG.
Xiangrui Yang 0001, Yekang Zhan, Qiang Cao 0001, Jie Yao 0001
ICME4
2025 AIS: An Active Idleness I/O Scheduler to Reduce Buffer-Exhausted Degradation of Solid-State Drives
abstract
Modern solid-state drives (SSDs) continue to boost storage density and I/O bandwidth at the cost of flash-access I/O latency, especially for write, hence they prevalently deploy a build-in buffer to absorb incoming writes. However, when the buffer is used up, the applications suffer from a sudden and long performance decline, i.e., buffer-exhausted degradation (BED). To holistically understand BED and recovery, we design an automated testing toolset (SSDTest) to measure six commodity NVMe SSDs and find: (1) the occurrence of the BED strictly relies on the written-data amount, (2) BED dramatically increases I/O latency of SSDs, especially write and read-after-write, (3) BED can be conditionally reduced and recovered only after a period of idle time, and (4) a read without preceding writes is largely immune to BED, but prolongs the required idle time to recover the available buffer. Furthermore, we build a black-box SSD buffer-recovery model to quantitatively characterize the idleness-recovery behaviors and design an SSD BED predictor to make BED occurrence and buffer recovery predictable. Leveraging this model, we further design an Active Idleness I/O Scheduler (AIS) with small-sized auxiliary storage to actively regulate the I/O idle-intervals to maximize the internal buffer recovery of SSD. AIS adaptively steers incoming data to the auxiliary storage to (1) strategically keep SSD idle to reduce the occurrence of BED and (2) mitigate the tail latency of SSDs caused by read-after-writes during BED. We perform extensive evaluations under a variety of workloads. The results show that AIS improves average, 99th, 99.9th, and 99.99th-percentile latencies of SSDs by up to 29.3%, 37.3%, 78.7%, and 67.2% respectively, with up to 512MB auxiliary storage.
Yekang Zhan, Xiangrui Yang 0001, Haichuan Hu, Qiang Cao 0001, Yifan Zhang 0012, Jie Yao 0001
ACM Trans. Archit. Code Optim.1
2024 RomeFS: A CXL-SSD Aware File System Exploiting Synergy of Memory-Block Dual Paths
abstract
Compute eXpress Link (CXL) based Solid-State Drives (CXL-SSDs), such as the Samsung CMM-H model, promise to offer CXL.mem memory and CXL.io block dual-mode interfaces. Nonetheless, whether and how cloud applications with diverse and varying access patterns benefit from such dual-mode CXL-SSD remains an open question for academia and industry.
Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001
SoCC1
2024 SchInFS: A File System Integrating Functions of the Block I/O Scheduler for ZNS SSDs
abstract
Emerging Zoned Namespace (ZNS) SSDs divide address space into sequentially written zones and transfer garbage collection (GC) to the host, thereby providing more stable performance, increased capacity, and extended device lifespan. However, the sequential write constraint poses some problems for file system design on ZNS devices, particularly leading to bottlenecks in multi-threaded performance for concurrent write requests. Through comprehensive experiments, we analyze the scalability issues of existing POSIX file systems on NVMe ZNS SSDs and identify the root causes: (1) current file systems generally fail to simultaneously utilize the throughput of multiple zones, and (2) their methods for concurrent writing within a single zone are inefficient. To fully exploit the concurrent performance of ZNS SSDs, we propose SchInFS, a novel multi-head logging ZNS SSD file system that integrates the functions of the block I/O scheduler. Firstly, SchInFS employs a multi-head logging design to leverage the throughput of multiple zones concurrently. Secondly, it provides an independent merge queue for each log, facilitating efficient cross-thread write blocks. Finally, SchInFS uses a Block I/O Submission Controller (BSC) to ensure the timely submission of requests and ordered writing within a single zone. Evaluation on real devices demonstrates the effectiveness of SchInFS, showcasing a substantial improvement in the concurrency performance of ZNS SSDs by up to 71.94% compared to current ZNS SSD file systems.
Jintong Zhang, Haichuan Hu, Jianxi Chen, Yekang Zhan
ICCD4
2022 HBtree: A Heterogeneous B+tree with Multi-granularity for Hybrid NVM-SSD Storage
abstract
Traditional index structures build on homogeneous storage consisting of same-granularity blocks and maintain a map between logical offset and storage-blocks. Non-Volatile Memory (NVM) and Solid-State Drive have their own optimal I/O sizes. However, these homogeneous indexes cannot uniformly and efficiently manage hybrid NVM-SSD storage space with heterogenous blocks. This paper proposes a heterogeneous B+tree, referred as to HBtree, to index multi-granularity blocks for hybrid NVM-SSD storage. The indexing node of HBtree is same to B+tree with largest-granularity blocks. However, a part of leaves in the legacy B+tree are extended to index small granularity blocks. Therefore, HBtree has the low tree-height while indexing different granularity blocks simultaneously. We implement HBtree and evaluate it on hybrid NVM-SSD storage. Compared to legacy B+tree, HBtree improves insert and search performance by up to 4.54x and 1.96x respectively and improving space utilization by up to 3.49x.
Yekang Zhan, Haichuan Hu, Qiang Cao 0001
NAS1