Yifan Zhang 0012

dblp:57/4707-12 · also Yi-Fan Zhang 0012 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0007-9920-4685ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 eLDPC: An Elastic and Scalable LDPC-Decoder With Early Termination by Effectively Leveraging High-Level Synthesis
abstract
Emerging communication and storage embrace Low-Density Parity-Check (LDPC) codes to fully exploit their physical channels. FPGA (Field-Programmable Gate Array) is widely employed to fast prototype and accelerate the LDPC decoding with high complexity. For varying channel conditions, the FGPA decoder is desired to elastically stop iteration when meeting success condition, avoiding conservatively performing a predefined and large number of iterations. However, the dynamical-execution algorithms with adjustable parameters generally are challenging for scalable decoder structure preferred to deterministic execution logic. To overcome the problem, this paper presents an elastic and scalable HLS-based FPGA LDPC decoder architecture with early-termination to achieve high throughput and flexibility. To this end, eLDPC first provides a universal operation, fully leveraging the features of HLS to efficiently implement optimized small-scale hardware units for low-level data-update operations. Second, eLDPC presents a decoding-iteration pipeline that adds a termination-check stage to terminate the following iteration for current codeword decoding. eLDPC also presents an HLS-enhanced approach to address memory access conflicts associated with the DU pipeline. Further, eLDPC extends the number of DU decoding-iteration pipelines within a single stream to decode multiple codewords in parallel. Third, eLDPC designs elastic and independent multiple decoding streams by using FIFO queues to decouple Input, Output, and a decoding unit (DU) with variable iterations while avoiding the potential blockage of the queueing. We implement and evaluate eLDPC on a Xilinx U55C. Experiments show that eLDPC outperforms recent decoders by up to 5× with the same parameter and achieves the actual decoding throughput of up to 49.5 Gbps with high scalability and flexibility.
Qiang Cao 0001, Yifan Zhang 0012, Yekang Zhan, Jie Yao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 AIS: An Active Idleness I/O Scheduler to Reduce Buffer-Exhausted Degradation of Solid-State Drives
abstract
Modern solid-state drives (SSDs) continue to boost storage density and I/O bandwidth at the cost of flash-access I/O latency, especially for write, hence they prevalently deploy a build-in buffer to absorb incoming writes. However, when the buffer is used up, the applications suffer from a sudden and long performance decline, i.e., buffer-exhausted degradation (BED). To holistically understand BED and recovery, we design an automated testing toolset (SSDTest) to measure six commodity NVMe SSDs and find: (1) the occurrence of the BED strictly relies on the written-data amount, (2) BED dramatically increases I/O latency of SSDs, especially write and read-after-write, (3) BED can be conditionally reduced and recovered only after a period of idle time, and (4) a read without preceding writes is largely immune to BED, but prolongs the required idle time to recover the available buffer. Furthermore, we build a black-box SSD buffer-recovery model to quantitatively characterize the idleness-recovery behaviors and design an SSD BED predictor to make BED occurrence and buffer recovery predictable. Leveraging this model, we further design an Active Idleness I/O Scheduler (AIS) with small-sized auxiliary storage to actively regulate the I/O idle-intervals to maximize the internal buffer recovery of SSD. AIS adaptively steers incoming data to the auxiliary storage to (1) strategically keep SSD idle to reduce the occurrence of BED and (2) mitigate the tail latency of SSDs caused by read-after-writes during BED. We perform extensive evaluations under a variety of workloads. The results show that AIS improves average, 99th, 99.9th, and 99.99th-percentile latencies of SSDs by up to 29.3%, 37.3%, 78.7%, and 67.2% respectively, with up to 512MB auxiliary storage.
Yekang Zhan, Xiangrui Yang 0001, Haichuan Hu, Qiang Cao 0001, Yifan Zhang 0012, Jie Yao 0001
ACM Trans. Archit. Code Optim.5
2024 HEncode: A Highly Modularized and Efficient FPGA QC-LDPC Encoder using High Level Synthesis
abstract
QC-LDPC (Quasi Cyclic Low-Density Parity-Check) codes, as a regular block-based code, have been preva-lently adopted in communication and storage fields to ensure high reliability and bandwidth of data channels. However, existing Field-Programmable Gate Array (FPGA) QC-LDPC encoders designed by RTL experts are generally dedicated to specialized LDPC codes and hardware platforms without flexibility and scalability. Recently, High-Level Synthesis (HLS) was introduced to compile a high-level encoding logic into Register Transfer Level (RTL) implementations, which are low performance and hardware efficiency due to the overlarge HLS-to-RTL design space, especially for large-scale FPGA hardware. This paper proposes a highly modularized and efficient FPGA QC-LDPC encoder, HEncoder, to fully leverage HLS to achieve high bandwidth, flexibility in both code parameters, and hardware efficiency. Firstly, HEncode presents an efficient Encode Block (EB) fully exploiting the FPGA LUT characteristic. Second, HEncode designs a low-level subword-encoding pipeline using multiple EBs and subword-parallel Encode Units (EU). Third, HEncoder designs an encode module with a pipelined data stream consecutively passing Input, EU array, and Output to balance bandwidths of accessing and encoding words. Finally, HEncode develops a design space analyzer to automatically determine the encoder parameters under constrained conditions to achieve high bandwidth. We implemented and evaluated HEncode on the Xilinx U50. The results show that compared to existing encoders, HEncode gains an increase of approximately 154.5× in the peak throughput and about 5.89 × in hardware efficiency to achieve the encoding throughput of 922.66 Gbps.
Xiangrui Yang 0001, Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Xiaodi Tan
ICCD4
2023 R-LDPC: Refining Behavior Descriptions in HLS to Implement High-throughput LDPC Decoder
abstract
High-Level Synthesis (HLS) translates high-level behavior-description to Register-Transfer Level (RTL) implemen-tation in modern Field-Programmable Gate Arrays (FPGAs), accelerating domain-specific hardware developments. Low-Density Parity-Check (LDPC), as a powerful error-correction code family, has been widely implemented in hardware for building a reliable data channel over a noisy physical channel in communication and storage applications. Leveraging HLS to fast prototype high-performance LDPC decoder is intriguing with high scalability and low hardware-dependence, but generally is sub-optimal due to the lack of accurate and precise behavior descriptions in HLS to characterize iteration- and circuit-level implementation details. This paper proposes an HLS-based QC-LDPC decoder with scalable throughput by precisely refining the LDPC behavior descriptions, R-LDPC for short. To this end, R-LDPC first adopts an HLS-based LDPC decoder microarchitecture with a module-level pipeline. Second, R-LDPC offers a multi-instance-sharing one (MSO) description to explicitly define shared parts and non-shared parts for an array of check-node updating-units (CNU), eliminating redundant function modules and addressing circuits. Third, R-LDPC designs efficient single-stage and multi-stage shifters to eliminate unnecessary bit-selection circuits. Finally, R-LDPC provides invalid-element aware loop scheduling before the compile phase to avoid some unnecessary stalls at runtime. We implement an R-LDPC decoder, compared to the original HLS-based implementation, R-LDPC reduces the hardware con-sumption up to 56%, the latency up to 67%, and the decoding throughput up to 300%. Furthermore, R-LDPC is adapted to different scales, LDPC standards, and code rates, and can achieve 9.9Gbps decoding throughput in Xilinx U50.
Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Hong Jiang 0001
DATE1
2023 HF-LDPC: HLS-friendly QC-LDPC FPGA Decoder with High Throughput and Flexibility
abstract
LDPC (Low-Density Parity-Check) codes have become a cornerstone of transforming a noise-filled physical channel into a reliable and high-performance data channel in communication and storage systems. FPGA (Field-Programmable Gate Array) based LDPC hardware, especially for decoding with high complexity, is essential to realizing the high-bandwidth channel prototypes. HLS (High-Level Synthesis) is introduced to speed up the FPGA development of LDPC hardware by automatically compiling high-level abstract behavioral descriptions into RTL-level implementations, but often sub-optimally due to lacking effective low-level descriptions. To overcome this problem, this paper proposes an HLS-friendly QC-LDPC FPGA decoder architecture, HF-LDPC, that employs HLS not only to precisely characterize high-level behaviors but also to effectively optimize low-level RTL implementation, thus achieving both high throughput and flexibility. First, HF-LDPC designs a multi-unit framework with a balanced I/O-computing dataflow to adaptively match code parameters with FPGA configurations. Second, HF-LDPC presents a novel fine-grained task-level pipeline with interleaved updating to eliminate stalls due to data interdependence within each updating task. HF-LDPC also presents several HLS-enhanced approaches. We implement and evaluate HF-LDPC on Xilinx U50, which demonstrates that HF-LDPC outperforms existing implementations by 4× to 84× with the same parameter and linearly scales to up to 116 Gbps actual decoding throughput with high hardware efficiency.
Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Hong Jiang 0001
ICCD1
2022 TLP-LDPC: Three-Level Parallel FPGA Architecture for Fast Prototyping of LDPC Decoder Using High-Level Synthesis
Yifan Zhang 0012, Qiang Cao 0001
J. Comput. Sci. Technol.1
2019 LT-TCO: A TCO Calculation Model of Data Centers for Long-Term Data Preservation
abstract
Data centers have been becoming public utilities to provide large-scale computing and storage services. The Total Cost of Ownership (TCO) models for such data centers are paramount to deeply understand their cost of investment and maintenance, the cost composition of internal components, and further cost optimization directions. Existing data center TCO models focus on either high-performance data centers or key subsystems such as IT facility, lacking of holistic analysis of the data centers designed for long-term data preservation. The long-term data centers can be built with different combinations of storage media such as HDDs, tapes, and optical discs. Meanwhile, during the long operation period, devices replacement and data migration are necessary and are not negligible in cost. In order to comprehensively and quantitatively understand the cost of long-term data preservations, we proposed LT-TCO, a TCO calculation model for data centers over time. LT-TCO simulates the construction and operation of a data center to calculate the expenditure of each year. It also introduces the cost of devices replacement and data migration during the long running period. Based on the storage media as optical discs, tapes, HDDs, and SSDs, LT-TCO evaluates the corresponding capital and operational expenditure under different developing rates. The simulation result shows that in long-term preservation, data migration cost takes more than 96% of the operational expenditure. And the TCO of optical disc data centers could be the least among four storage media.
Wenrui Yan, Jie Yao 0001, Qiang Cao 0001, Yifan Zhang 0012
NAS4