Lishuo Deng

dblp:352/9385 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0009-2394-7663ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BALANCE: Bit and Layer-Aware Lightweight ECC Design Method for In-Flash Computing Based LLM Inference Accelerator
Changwei Yan, Lishuo Deng, Mingbo Hao, Weiwei Shan
ASP-DAC2
2026 A 2.2-$μ$W Keyword Spotting System With Filterbank Learning and Attention-Aware Soft-Threshold Denoising in 28-nm CMOS
abstract
Keyword spotting (KWS) is crucial for speech interactions in mobile devices, yet its widespread adoption is constrained by the trade-off between low power and high accuracy in diverse noisy environments. Here, we present an ultra-low-power and robust KWS chip to address this challenge through architecture and circuit co-optimization. First, a learnable Mel filter bank is proposed and co-designed with a group-based time-multiplexed architecture to perform adaptive spectral reshaping in a log-Mel feature extractor, reducing the downstream network computations by 25.8% while preserving feature representation capability. Second, a broadcasted residual network (BC-ResNet) is designed to support both KWS and speaker verification (SV) classification tasks, which reduces additions by 43.5% and parameters by 5% compared to a lightweight depthwise separable CNN (DSCNN). Third, an attention-aware soft-threshold denoising strategy is developed to dynamically filter noisy features, yielding up to a 5.8% increase in accuracy over the baseline without denoising at 0dB signal-to-noise ratio (SNR). Finally, a multiplier-less processing element (PE) array combined with stacked-transistor-based ultra-low-leakage custom memory is designed, reducing hardware area by 31.3% and leakage power by 20.4%. Fabricated in a 28nm CMOS process, the proposed 12-class KWS chip consumes only$2.2\mu $W at 0.42V supply. It achieves 87.6%-93.5% accuracy on GSCD across 0dB to 20dB SNR levels, demonstrating high robustness for diverse scenarios.
Zhihao Yan, Lishuo Deng, Junyi Qian, Weiwei Shan
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Efficient Hold Buffer Optimization by Supply Noise-Aware Dynamic Timing Analysis
abstract
As the CMOS process scales down, digital circuits become more susceptible to hold time violations due to increased sensitivity to supply voltage fluctuations. Since hold time violation is fatal, sufficient hold fixing buffers need to be inserted into the short paths to prevent it. However, by assuming a constant power supply level, traditional hold fixing causes imprecise and overly conservative timing analysis and hence leads to circuit overhead and degraded performance. To address this, we propose a power supply noise (PSN)-aware dynamic timing analysis for realistic hold time analysis and efficient hold buffer optimization, which integrates a machine learning-based timing model into the conventional design flow. Building on the highly effective application of the Weibull cumulative distribution function and machine learning for dynamic PSN-aware timing analysis, we propose introducing an additional parameter for PSN amplitude, which has a significant impact on delay, and narrowing the overall parameter range using real PSN waveforms extracted from the RedHawk. This approach achieves a prediction error of only 3.45% for cell delay and 5.1 % for path delay, while also reducing dataset acquisition costs. To the best of our knowledge, this work is the first to apply PSN-aware dynamic timing analysis specifically for hold optimization, mitigating the pessimism of traditional static timing analysis (STA) and effectively minimizing redundant hold fixing buffers while remaining compatible with existing design workflows. Since short paths often overlap with critical paths, reducing redundant hold buffers not only decreases area overhead but also enhances performance. Applied to a 22 nm, 64-point Fast Fourier Transform (FFT) circuit, our EDA compatible method combined with a greedy algorithm reduces hold buffers by 55%, achieving not only 6.79%circuit area reduction but also 8.1 % performance improvement due to the elimination of redundant buffers in short and critical paths.
Lishuo Deng, Changwei Yan, Zhuo Chen 0039, Weiwei Shan
DATE1
2025 Variation-Adaptive Negative Bitline and Skip Bitline Pre-charge Scheme for Low-Power SRAM
abstract
The negative bitline (NBL) write assist circuitry is widely adopted in static random access memory (SRAM) for its significant improvement in write yield and acceptable cost. Due to PVT variations and the trade-offs between yield and energy in different applications, the adaptive configuration of NBL enable timing and negative voltage level is increasingly essential, while existing NBL schemes lack a comprehensive methodology to adjust NBL options effectively. This paper presents a PVT variation-adaptive NBL (VA-NBL) write assist scheme and skip bitline pre-charge (SBP) circuitry for low-power SRAM. The VA-NBL scheme establishes a mapping relationship between NBL options and PVT conditions and employs on-chip PVT sensors to adaptively control NBL options, ensuring the lowest power consumption and satisfying target yield under PVT variations. Applied in a 64kb SRAM macro, we achieved different optimal NBL options under various PVT conditions, resulting in a maximum write energy power savings of 47%, while introducing negligible area overhead on dynamic voltage frequency scaling (DVFS) systems that already integrate PVT sensors. Additionally, the SBP circuitry for hierarchical bitline architecture avoids unnecessary pre-charge for write operations and further yields a 53% reduction in write energy.
Lishuo Deng, Changwei Yan, Xin Si, Weiwei Shan
ISCAS1
2025 A 250M-2.5GHz Two-Stage Duty-Cycle Corrector with 10%-90% Correction Range and 3-Cycle Correction Latency for Mitigating Aging Effects
abstract
Clock duty-cycle distortion caused by aging effects, which induces circuit performance degradation, has become a significant concern in advanced processes. We propose a digital two-stage duty-cycle corrector (DCC) that directly corrects the duty-cycle distortion, using a duty-cycle adjuster (DCA) for coarse duty-cycle width modulation followed by a high-accuracy half-cycle delay line (HCDL) for 50% duty-cycle correction. The two-stage structure further extends the operating frequency and duty-cycle range, while achieving low correction latency through the elimination of external complex control by employing open-loop logic as compared with state-of-the-art (SOTA) works. Implemented in a 22nm ULVT CMOS process, measurement results show that it operates at a frequency range of 250M to 2.5GHz with an acceptable input duty-cycle range of 10% to 90%. It achieves a maximum duty-cycle correction error of 1.5% at 2.5GHz and correction latency of only three clock cycles. The maximum peak-to-peak jitter of the output clock is 13.25ps.
Zhiting Li, Lishuo Deng, Changwei Yan, Zhangrui Qian, Weiwei Shan
ISCAS2
2025 A 40nm Early-Warning AVFS Design with Path Activation-Based Monitoring Point Optimization
abstract
This work presents an early-warning adaptive voltage-frequency scaling (AVFS) system, offering a practical and innovative low-power solution for commercial IoT products. First, a novel monitoring point optimization strategy based on path activation analysis is proposed, which mitigates the risk of illegal voltage scaling and reduces monitoring costs by 18.7%. Second, a dual shadow register-based monitor for memory endpoints is designed to overcome the limitation of monitoring only D-type flip-flop (DFF) endpoints. Third, a single-chip solution that integrates both frequency and voltage regulation strategies is presented to alleviate the issues of off-chip voltage regulation, such as large noise ripple and long response times. Fabricated in a 40-nm CMOS process, a Cortex M0+ CPU with AVFS works within a frequency range of 50-100MHz and a voltage range of 0.75-1.1V. Furthermore, the proposed AVFS achieves significant power savings of up to 49.3% at FF, 25°C, and 100MHz.
Kaize Zhou, Lishuo Deng, Junyi Qian, Xiaojie Yin, Kana Meng, Weiwei Shan
ISCAS4
2025 A Compound Timing Detection of Both Data Transition and Path Activation for Reliable In Situ Error Detection and Correction
abstract
Timing error detection and correction (EDAC) in resilient circuits helps eliminate excess timing margins. However, it faces misdetection risks when critical paths (CPs) remain inactive. We propose a 15-transistor timing error and path-activation detector (TEPD) capable of detecting both timing violations in late-arriving signals and CP activation, with robust operation down to 0.33 V. Regarding circuit-level error correction, we introduce an error-correcting flip-flop (ECFF), leveraging time-borrowing for zero-cycle response latency without requiring pipeline refresh. The custom-optimized ECFF adds only six transistors, increasing delay, dynamic power, and static power by 15%, 11%, and 14%, respectively, compared with a standard flip-flop, ensuring efficient error correction with minimal cost. A system-level voltage tuning strategy is further developed to handle continuous timing errors, ensuring robust adaptive voltage scaling (AVS) operation. Implemented on a neural network (NN) accelerator in the 28-nm CMOS, the system operates across a wide voltage range from 0.54 to 0.9 V. It achieves up to 52% power gain or 123% frequency gain at the near-threshold region, with negligible area and power overhead compared with the margined baseline.
Lishuo Deng, Junyi Qian, Zhengguo Shen, Jingchen Wang, Zhangrui Qian, Longning Qi, Weiwei Shan
IEEE Trans. Very Large Scale Integr. Syst.1
2023 IVATS: A Leakage Reduction Technique Based on Input Vector Analysis and Transistor Stacking in CMOS Circuits
abstract
Leakage reduction is crucial for always-on IoT applications in which static power consumption of the memory cells accounts for a large proportion of the total power. Even with high threshold voltage transistors, the leakage is still considerable. This paper proposes a novel technique based on input vector analysis and transistor stacking to analyze and suppress leakage, especially for extremely high threshold voltage (EHVT) circuits operating in the near/sub-threshold regime. At the device level, we consider the leakage ratio of each transistor terminal, which improves the universality of the method. At the circuit level, we innovatively propose the concepts of critical leakage path, leakage power components, and public leakage path to help designers locate the sources of leakage more precisely. We apply the method to a 28nm-EHVT low leakage tristate latch-like memory cell in a serial Fast Fourier Transform (FFT) circuit and find that inserting one stacking NMOS and using “01” stack to reduce the substrate leakage of PMOS can effectively suppress leakage. The average leakage power consumption of the optimized cell is reduced by 42% in the pre-layout simulation. A 26.87% and 17.52% leakage power reduction in the custom cell and the serial FFT circuit is achieved after the layout design and the synthesis.
Lishuo Deng, Weiwei Shan
ISCAS1