Changwei Yan

dblp:256/7595 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BALANCE: Bit and Layer-Aware Lightweight ECC Design Method for In-Flash Computing Based LLM Inference Accelerator
Changwei Yan, Lishuo Deng, Mingbo Hao, Weiwei Shan
ASP-DAC1
2025 Efficient Hold Buffer Optimization by Supply Noise-Aware Dynamic Timing Analysis
abstract
As the CMOS process scales down, digital circuits become more susceptible to hold time violations due to increased sensitivity to supply voltage fluctuations. Since hold time violation is fatal, sufficient hold fixing buffers need to be inserted into the short paths to prevent it. However, by assuming a constant power supply level, traditional hold fixing causes imprecise and overly conservative timing analysis and hence leads to circuit overhead and degraded performance. To address this, we propose a power supply noise (PSN)-aware dynamic timing analysis for realistic hold time analysis and efficient hold buffer optimization, which integrates a machine learning-based timing model into the conventional design flow. Building on the highly effective application of the Weibull cumulative distribution function and machine learning for dynamic PSN-aware timing analysis, we propose introducing an additional parameter for PSN amplitude, which has a significant impact on delay, and narrowing the overall parameter range using real PSN waveforms extracted from the RedHawk. This approach achieves a prediction error of only 3.45% for cell delay and 5.1 % for path delay, while also reducing dataset acquisition costs. To the best of our knowledge, this work is the first to apply PSN-aware dynamic timing analysis specifically for hold optimization, mitigating the pessimism of traditional static timing analysis (STA) and effectively minimizing redundant hold fixing buffers while remaining compatible with existing design workflows. Since short paths often overlap with critical paths, reducing redundant hold buffers not only decreases area overhead but also enhances performance. Applied to a 22 nm, 64-point Fast Fourier Transform (FFT) circuit, our EDA compatible method combined with a greedy algorithm reduces hold buffers by 55%, achieving not only 6.79%circuit area reduction but also 8.1 % performance improvement due to the elimination of redundant buffers in short and critical paths.
Lishuo Deng, Changwei Yan, Zhuo Chen 0039, Weiwei Shan
DATE2
2025 Variation-Adaptive Negative Bitline and Skip Bitline Pre-charge Scheme for Low-Power SRAM
abstract
The negative bitline (NBL) write assist circuitry is widely adopted in static random access memory (SRAM) for its significant improvement in write yield and acceptable cost. Due to PVT variations and the trade-offs between yield and energy in different applications, the adaptive configuration of NBL enable timing and negative voltage level is increasingly essential, while existing NBL schemes lack a comprehensive methodology to adjust NBL options effectively. This paper presents a PVT variation-adaptive NBL (VA-NBL) write assist scheme and skip bitline pre-charge (SBP) circuitry for low-power SRAM. The VA-NBL scheme establishes a mapping relationship between NBL options and PVT conditions and employs on-chip PVT sensors to adaptively control NBL options, ensuring the lowest power consumption and satisfying target yield under PVT variations. Applied in a 64kb SRAM macro, we achieved different optimal NBL options under various PVT conditions, resulting in a maximum write energy power savings of 47%, while introducing negligible area overhead on dynamic voltage frequency scaling (DVFS) systems that already integrate PVT sensors. Additionally, the SBP circuitry for hierarchical bitline architecture avoids unnecessary pre-charge for write operations and further yields a 53% reduction in write energy.
Lishuo Deng, Changwei Yan, Xin Si, Weiwei Shan
ISCAS2
2025 A 250M-2.5GHz Two-Stage Duty-Cycle Corrector with 10%-90% Correction Range and 3-Cycle Correction Latency for Mitigating Aging Effects
abstract
Clock duty-cycle distortion caused by aging effects, which induces circuit performance degradation, has become a significant concern in advanced processes. We propose a digital two-stage duty-cycle corrector (DCC) that directly corrects the duty-cycle distortion, using a duty-cycle adjuster (DCA) for coarse duty-cycle width modulation followed by a high-accuracy half-cycle delay line (HCDL) for 50% duty-cycle correction. The two-stage structure further extends the operating frequency and duty-cycle range, while achieving low correction latency through the elimination of external complex control by employing open-loop logic as compared with state-of-the-art (SOTA) works. Implemented in a 22nm ULVT CMOS process, measurement results show that it operates at a frequency range of 250M to 2.5GHz with an acceptable input duty-cycle range of 10% to 90%. It achieves a maximum duty-cycle correction error of 1.5% at 2.5GHz and correction latency of only three clock cycles. The maximum peak-to-peak jitter of the output clock is 13.25ps.
Zhiting Li, Lishuo Deng, Changwei Yan, Zhangrui Qian, Weiwei Shan
ISCAS4