Zhi Li 0090

dblp:43/3166-90 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0000-5529-1421ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 A 0.31-V 16-Kb 9T SRAM With Enhanced Sensing Margin and Read Performance for Low-Power Applications
abstract
This brief presents a low-power 9T static random access memory (SRAM) with enhanced read sensing margin and read performance. The read decoupled port of the proposed 9T SRAM cell achieves the enhanced sensing margin by mitigating the read bitline (RBL) leakage and improves the read performance through using one-transistor read path. The multithreshold voltage devices are used in SRAM cell for improving the leakage power and performance of SRAM. Additionally, an interleaved write wordline (WWL) structure is implemented to address the write half-select issue. The measurement results of the test chip fabricated in the 22-nm FDSOI technology demonstrate that the designed 9T SRAM achieves a minimum operation voltage of 0.31 V at 1.05 MHz and can operate at 60.5 MHz when the supply voltage is 0.5 V. The minimum active energy of 18.56 fJ/access-bit is obtained at 0.33 V. Furthermore, the designed SRAM exhibits a minimum leakage power of 0.11 pW/bitcell in the retention mode.
Pengyuan Zhao, Linnan Li, Zhi Li 0090, Minglong Jia, Xiang Li 0164, Shushan Qiao
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A 230-nA Quiescent Current and Enhanced Transient Performance DC-DC Converter with Suppressed Under/Overshoot for IoT SoCs
abstract
This paper presents a low quiescent current, transient-enhanced DC-DC buck converter with an overshoot and undershoot suppression scheme, designed to support low-power SoCs. To decrease undershoot or overshoot voltage during rapid load current transitions, an overshoot and undershoot voltage detection and response scheme is proposed. The buck converter employs an adaptive constant on time control mode, and to achieve automatic transition between PWM and PFM modes, a sub-threshold zero current detector(ZCD) is introduced. To reduce quiescent current and improve efficiency under light loads, a sampling-based low-power voltage reference circuit is proposed. Implemented with a 55 nm CMOS process, the buck converter operates with a load current range of 10 μA-10 mA. It achieves a quiescent current of 230 nA and a peak efficiency of 93.2%. During a 10 mA load variation, the converter exhibits an undershoot of 36 mV and an overshoot of 34 mV.
Zhi Li 0090, Linnan Li, Shushan Qiao
ISCAS1
2025 AO-EDC: An Accuracy-Oriented Error Detection and Correction Scheme for DVFS System Based on Propagation Detection at Half-Path Points
abstract
Error detection and correction (EDC) techniques are conventionally employed in dynamic voltage and frequency scaling (DVFS) systems to eliminate timing margins preserved in ICs. However, existing error detection methods generate warning signals even when no violations occur, resulting in frequent corrections that degrade system throughput. This issue of inaccurate detection worsens with further scaling. To solve this issue, this paper presents an accuracy-oriented EDC scheme, AO-EDC. An accurate detection method, consisting of a propagation detector (PD) and an improved half-path-point insertion method, is proposed. PDs that filter out signals from other paths are inserted at the half-path points of selected paths, reducing inaccurate detection by up to 77.5%. In addition, a half-cycle clock gating circuit is proposed to reduce the redundancy of corrections. Implemented on a 22-nm process FIR circuit, post-layout simulation results demonstrate a 42.4% - 50.0% power reduction and a 106.7% - 1089.9% frequency gain at 0.9 - 0.55 V. The proposed scheme triggers fewer corrections and reduces correction timing waste, addressing the throughput degradation issue during over-scaling in DVFS systems.
Zitao Liang, Jiliang Liu, Kangning Wang 0004, Linnan Li, Zhi Li 0090, Shushan Qiao
ISCAS5
2025 A Self-Calibrated Unified Voltage-and-Frequency Regulator System Design Based on Universal Logic Line Circuit
abstract
In this brief, a unified voltage frequency regulator (UVFR) system is designed to eliminate the voltage margin induced by process, voltage, and temperature (PVT) variations. The frequency is regulated with voltage by a universal logic line oscillator (ULLO), which can protect the system from timing violations. The length of the ULLO is self-calibrated by a ULL-based time-digital converter (ULL-TDC) and an in situ half-critical path timing detector, where the ULL is designed to track the critical path delay. A fully synthesizable digital low dropout (DLDO) is designed with the ULL-TDC and a proportional differential (PD) circuit for voltage regulation. The proposed system is implemented in an ARM Cortex-M0 microcontroller in 22 nm technology. Simulation results show that the ULL can accurately track the critical path delay with a maximum variation of 3% at 0.6 V and 11.5% at 0.45 V. The UVFR system consumes 13.2–112 uW of power overhead, and eliminates the voltage margin by 22.3%–28% while reducing the power consumption by 35%–42.3%.
Jiliang Liu, Zhi Li 0090, Kangning Wang 0004, Shushan Qiao
IEEE Trans. Very Large Scale Integr. Syst.3
2025 Accelerating Unstructured Sparse DNNs via Multilevel Partial Sum Reduction and PE Array-Level Load Balancing
abstract
Unstructured pruning introduces significant sparsity in deep neural networks (DNNs), enhancing accelerator hardware efficiency. However, three critical challenges constrain performance gains: 1) complex fetching logic for nonzero (NZ) data pairs; 2) load imbalance across processing elements (PEs); and 3) PE stalls from write-back contention. This brief proposes an energy-efficient accelerator addressing these inefficiencies through three innovations. First, we propose a Cartesian-product output-row-stationary (CPORS) dataflow that inherently matches NZ data pairs by sequentially fetching compressed data. Second, a multilevel partial sum reduction (MLPR) strategy minimizes write-back traffic and converts random PE stalls into manageable load imbalance. Third, a kernel sorting and load scheduling (KSLS) mechanism resolves PE idle/stall and achieves PE array-level load balancing, attaining 76.6% average PE utilization across all sparsity levels. Implemented in 22-nmCMOS, the accelerator delivers$1.85\times $speedup and$1.4\times $energy efficiency over baseline and achieves 25.8 TOPS/W peak energy efficiency at 90% sparsity.
Chendong Xia, Zhi Li 0090, Bing Li 0017, Shushan Qiao
IEEE Trans. Very Large Scale Integr. Syst.3