EDBT 2026 Demo / reviewers in the wild / expert
Yongbiao Zhu
dblp:390/4651
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0003-1738-5699ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resolving Gray Code Dilemma With Bidirectional Programming for Efficient QLC SSDsabstractQLC NAND flash is widely adopted in modern storage systems. By trading off read/write performance for storage density through a “time-for-space" approach, QLC enables ultra-high storage capacity. To mitigate performance degradation, Gray code and the two-step programming (TSP) algorithm are used. However, Gray code also has limitations: multiple Gray codes incur circuit overhead, while a single Gray code causes extra I/O latency overhead. This paradox seems unsolvable at first glance, requiring an innovative solution that maintains I/O performance without additional circuit overhead. A promising solution lies in selecting an appropriate Gray code and preventing hot data placement on slow physical pages. This paper proposes BDP, a novel Bi-Directional Programming scheme that adopts a single Gray code to fit both traditional (forward) and reverse programming directions based on TSP. The objective of BDP is to resolve the inherent contradiction between I/O performance preservation and implementation overhead. BDP optimizes the system performance through hardware/software co-design. At the hardware level, a fixed Gray code is employed to avoid additional circuit complexity. At the software level, two strategies (i.e. hotness-aware data allocation and background data migration) are proposed to further mitigate the misplacement of hot data on slow pages in QLC SSDs. The experimental results demonstrate that BDP significantly reduces the allocation of hot data to slow pages and enhances overall I/O performance compared to representative schemes. Yi Wang 0003, Shaoqi Li, Yongbiao Zhu, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Anchor First, Accelerate Next: Revolutionizing GNNs with PIM by Harnessing Stationary DataabstractSubstantial data movement caused by irregular graph topologies hinders the efficient processing of graph neural networks (GNNs). Although the emerging near-bank processing-in-memory (PIM) architecture offers a promising solution to reduce data transfer between memory and computing units, cross-bank communication remains a critical challenge, limiting the benefits of PIM architectures. Our findings indicate that only $35.6 \%$ of the data can stay stationary within PIM units on average, with the rest requiring movement due to graph dependencies. This situation worsens as the number of PIM units increases, reducing the ratio to $18.7 \%$. In this paper, we argue that to fully leverage PIM architectures, systems must maximize stationary data and minimize the movement of non-stationary data. Following this principle, we propose Anchor, a scalable PIM architecture that exploits stationary data for GNNs through a hardware-software co-design approach. To maximize stationary data, we introduce the graph partitioning algorithm Mastav, which carefully allocates vertices and edges to preserve data locality. To minimize the movement of non-stationary data, we employ a two-step strategy. First, a customized dataflow ensures that non-stationary data is accessed and distributed exactly once. Second, an optimized communication mechanism reduces redundant data transfers through critical paths. Our extensive experiments demonstrate that Anchor significantly reduces processing latency and data movement compared to representative schemes. Jiaxian Chen, Yuxuan Qi, Yongbiao Zhu, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
DAC | 3 |
| 2025 | One Gray Code Fits All: Optimizing Access Time with Bi-Directional Programming for QLC SSDsabstractGray code, a voltage-level-to-data-bit translation scheme, is widely used in QLC SSDs. However, it causes the four data bits in QLC to exhibit significantly different read and write performance with up to 8 × latency variation, severely impacting the worst-case performance of QLC SSDs. This paper presents BDP, a novel Bi-Directional Programming scheme. Based on a fixed Gray code, BDP combines both the normal (forward) and reverse programming directions to enable runtime programming direction arbitration. Experimental results show that BDP can effectively improve the read and write performance of SSD compared to representative schemes. Shaoqi Li, Tianyu Wang 0009, Yongbiao Zhu, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao |
DATE | 3 |
| 2024 | NICE: A Nonintrusive In-Storage-Computing Framework for Embedded ApplicationsabstractEmbedded machine learning applications face challenges related to massive data movement and high computational intensity, exacerbated by the limited performance of mobile devices. Computational storage devices (CSDs) pose huge potential for accelerating both data-intensive and computation-intensive embedded machine learning tasks by effectively reducing data movement and leveraging built-in accelerators. However, existing in-storage-computing (ISC) frameworks either require invasive customization of existing host driver layers or necessitate complex device firmware modifications, hindering the widespread deployment of CSDs. In addition, the lack of file semantics and the constrained internal resources within CSD implicitly compromise system performance and impact normal read/write performance. In this article, we aim to provide a nonintrusive in-storage-computing framework for embedded applications, named NICE. This framework includes an easy-to-use ISC programming interface that bypasses the kernel stack and requires no modification to the host NVMe driver, which is achieved through a novel hyper-addressing-based programming library and a file-aware page data layout within the CSD. In addition, we incorporate a lightweight kernel with coroutine-based command scheduling and several FPGA-based accelerators within the storage device firmware to enhance the performance of embedded machine learning applications while ensuring that the normal I/O performance remains unaffected. NICE is implemented on real CSD hardware integrated with ARM and FPGA. Experimental results demonstrate that our NICE framework can achieve an average latency performance improvement of$43.5\times $($9.32\times $) compared to CPU-(GPU-) based embedded machine learning solutions using the state-of-the-art NVIDIA Jetson NX platform, with$27.5\times $($4.3\times $) higher energy efficiency. NICE also has$34.2\times $less software and I/O performance overheads than state-of-the-art ISC frameworks. Tianyu Wang 0009, Yongbiao Zhu, Shaoqi Li, Jin Xue, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |