Weiwei Shan

dblp:39/1305 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0001-5520-1326ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 2 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BALANCE: Bit and Layer-Aware Lightweight ECC Design Method for In-Flash Computing Based LLM Inference Accelerator
Changwei Yan, Lishuo Deng, Mingbo Hao, Weiwei Shan
ASP-DAC5
2026 An Accuracy-and-Efficiency-Configurable Blade-Type Approximate Multiplier With Genetic Algorithm-Based Automatic Training Framework
Kaize Zhou, Zhihao Yan, Zhangrui Qian, Zhuo Chen 0039, Jun Yin 0001, Yan Lu 0002, Weiwei Shan
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 A 2.2-$μ$W Keyword Spotting System With Filterbank Learning and Attention-Aware Soft-Threshold Denoising in 28-nm CMOS
abstract
Keyword spotting (KWS) is crucial for speech interactions in mobile devices, yet its widespread adoption is constrained by the trade-off between low power and high accuracy in diverse noisy environments. Here, we present an ultra-low-power and robust KWS chip to address this challenge through architecture and circuit co-optimization. First, a learnable Mel filter bank is proposed and co-designed with a group-based time-multiplexed architecture to perform adaptive spectral reshaping in a log-Mel feature extractor, reducing the downstream network computations by 25.8% while preserving feature representation capability. Second, a broadcasted residual network (BC-ResNet) is designed to support both KWS and speaker verification (SV) classification tasks, which reduces additions by 43.5% and parameters by 5% compared to a lightweight depthwise separable CNN (DSCNN). Third, an attention-aware soft-threshold denoising strategy is developed to dynamically filter noisy features, yielding up to a 5.8% increase in accuracy over the baseline without denoising at 0dB signal-to-noise ratio (SNR). Finally, a multiplier-less processing element (PE) array combined with stacked-transistor-based ultra-low-leakage custom memory is designed, reducing hardware area by 31.3% and leakage power by 20.4%. Fabricated in a 28nm CMOS process, the proposed 12-class KWS chip consumes only$2.2\mu $W at 0.42V supply. It achieves 87.6%-93.5% accuracy on GSCD across 0dB to 20dB SNR levels, demonstrating high robustness for diverse scenarios.
Zhihao Yan, Lishuo Deng, Junyi Qian, Weiwei Shan
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 Balancing Objective Optimization and Constraint Satisfaction for Robust Analog Circuit Optimization
abstract
Automated design of analog integrated circuits (ICs) involves balancing multiple objectives under process, voltage, and temperature (PVT) variations. An excess of constraints can ensnare algorithms in local optima, while the variations elevate the costs of simulation. To address this challenge, we propose a two-search mode multi-task evolutionary framework to balance objective optimization and constraint satisfaction under variations. Specifically, considering the inherent relationships between objective optimizations and constraint violations, our method adaptively switches between unconstrained surrogate-assisted and constrained simulation-driven search modes. Furthermore, our framework treats PVT variations as a multi-task challenge, facilitating inter-corner knowledge transfer via multi-task evolution, substantially lowering simulation costs. Our framework has been evaluated using two different sensing elements and an amplifier within a 22 nm process. Based on Monte-Carlo simulations, compared to multi-task reinforcement learning, this method attains a 60% to 80% reduction in the relative inaccuracy of sensing elements and accomplishes a 60% decrease in total runtime.
Jintao Li 0002, Haochang Zhi, Jiang Xiao 0002, Yanhan Zeng, Weiwei Shan, Yun Li 0002
ASP-DAC5
2025 Analog Circuit Transfer Method Across Technology Nodes via Transistor Behavior
abstract
In the post-Moore era, chips integrate multiple technology node chiplets, necessitating repeated implementations of the same circuit topology across nodes, highlighting the need for technology-independent circuit representation. We use a four-parameter vector---gm, ft, VDS, and ΔVGS-to represent the behavior of each transistor, called the transistor behavioral vector (TBV). The TBVs are vertically concatenated to form the transistor behavioral circuit representation (TBCR) matrix, which precisely reflects the circuit's performance and provides a technology-independent representation. Furthermore, we propose a transistor behavioral model (TBM) to convert the TBV into the corresponding sizing. Finally, we propose a method to transfer analog circuits between different technology nodes using TBM (TNT), translating the modifications in the process parameters into the corresponding adjustments in ΔVGS. The experimental results show that for a single transistor, the mapping accuracy from TBV to simulation result was reached 99%. Multiple amplifiers were transferred from 180nm to 22nm technology, compared to the conventional transfer method based on gm/id, our transfer method based on TBCR achieved a success rate of up to 5× higher, along with additional performance improvements from the scaling down.
Haochang Zhi, Jintao Li 0002, Yun Li 0002, Weiwei Shan
ASP-DAC4
2025 UVLLM: An Automated Universal RTL Verification Framework using LLMs
abstract
Verifying hardware designs in embedded systems is crucial but often labor-intensive and time-consuming. While existing solutions have improved automation, they frequently rely on unrealistic assumptions. To address these challenges, we introduce a novel framework, UVLLM, which combines Large Language Models (LLMs) with the Universal Verification Methodology (UVM) to relax these assumptions. UVLLM significantly enhances the automation of testing and repairing error-prone Register Transfer Level (RTL) codes, a critical aspect of verification development. Unlike existing methods, UVLLM ensures that all errors are triggered during verification, achieving a syntax error fix rate of 86.99% and a functional error fix rate of 71.92% on our proposed benchmark. These results demonstrate a substantial improvement in verification efficiency. Additionally, our study highlights the current limitations of LLM applications, particularly their reliance on extensive training data. We emphasize the transformative potential of LLMs in hardware design verification and suggest promising directions for future research in AI-driven hardware design methodologies. The Repo. of dataset and code: https://github.com/SEU-ACAL/reproduce-UVLLM-DAC-25/.
Junhao Ye, Xinyao Jiao, Dingrong Pan, Jie Zhou 0001, Ning Wang 0071, Weiwei Shan, Xinwei Fang, Xi Wang 0009, Nan Guan, Zhe Jiang 0004
DAC10
2025 Efficient Hold Buffer Optimization by Supply Noise-Aware Dynamic Timing Analysis
abstract
As the CMOS process scales down, digital circuits become more susceptible to hold time violations due to increased sensitivity to supply voltage fluctuations. Since hold time violation is fatal, sufficient hold fixing buffers need to be inserted into the short paths to prevent it. However, by assuming a constant power supply level, traditional hold fixing causes imprecise and overly conservative timing analysis and hence leads to circuit overhead and degraded performance. To address this, we propose a power supply noise (PSN)-aware dynamic timing analysis for realistic hold time analysis and efficient hold buffer optimization, which integrates a machine learning-based timing model into the conventional design flow. Building on the highly effective application of the Weibull cumulative distribution function and machine learning for dynamic PSN-aware timing analysis, we propose introducing an additional parameter for PSN amplitude, which has a significant impact on delay, and narrowing the overall parameter range using real PSN waveforms extracted from the RedHawk. This approach achieves a prediction error of only 3.45% for cell delay and 5.1 % for path delay, while also reducing dataset acquisition costs. To the best of our knowledge, this work is the first to apply PSN-aware dynamic timing analysis specifically for hold optimization, mitigating the pessimism of traditional static timing analysis (STA) and effectively minimizing redundant hold fixing buffers while remaining compatible with existing design workflows. Since short paths often overlap with critical paths, reducing redundant hold buffers not only decreases area overhead but also enhances performance. Applied to a 22 nm, 64-point Fast Fourier Transform (FFT) circuit, our EDA compatible method combined with a greedy algorithm reduces hold buffers by 55%, achieving not only 6.79%circuit area reduction but also 8.1 % performance improvement due to the elimination of redundant buffers in short and critical paths.
Lishuo Deng, Changwei Yan, Zhuo Chen 0039, Weiwei Shan
DATE5
2025 Variation-Adaptive Negative Bitline and Skip Bitline Pre-charge Scheme for Low-Power SRAM
abstract
The negative bitline (NBL) write assist circuitry is widely adopted in static random access memory (SRAM) for its significant improvement in write yield and acceptable cost. Due to PVT variations and the trade-offs between yield and energy in different applications, the adaptive configuration of NBL enable timing and negative voltage level is increasingly essential, while existing NBL schemes lack a comprehensive methodology to adjust NBL options effectively. This paper presents a PVT variation-adaptive NBL (VA-NBL) write assist scheme and skip bitline pre-charge (SBP) circuitry for low-power SRAM. The VA-NBL scheme establishes a mapping relationship between NBL options and PVT conditions and employs on-chip PVT sensors to adaptively control NBL options, ensuring the lowest power consumption and satisfying target yield under PVT variations. Applied in a 64kb SRAM macro, we achieved different optimal NBL options under various PVT conditions, resulting in a maximum write energy power savings of 47%, while introducing negligible area overhead on dynamic voltage frequency scaling (DVFS) systems that already integrate PVT sensors. Additionally, the SBP circuitry for hierarchical bitline architecture avoids unnecessary pre-charge for write operations and further yields a 53% reduction in write energy.
Lishuo Deng, Changwei Yan, Xin Si, Weiwei Shan
ISCAS5
2025 A 250M-2.5GHz Two-Stage Duty-Cycle Corrector with 10%-90% Correction Range and 3-Cycle Correction Latency for Mitigating Aging Effects
abstract
Clock duty-cycle distortion caused by aging effects, which induces circuit performance degradation, has become a significant concern in advanced processes. We propose a digital two-stage duty-cycle corrector (DCC) that directly corrects the duty-cycle distortion, using a duty-cycle adjuster (DCA) for coarse duty-cycle width modulation followed by a high-accuracy half-cycle delay line (HCDL) for 50% duty-cycle correction. The two-stage structure further extends the operating frequency and duty-cycle range, while achieving low correction latency through the elimination of external complex control by employing open-loop logic as compared with state-of-the-art (SOTA) works. Implemented in a 22nm ULVT CMOS process, measurement results show that it operates at a frequency range of 250M to 2.5GHz with an acceptable input duty-cycle range of 10% to 90%. It achieves a maximum duty-cycle correction error of 1.5% at 2.5GHz and correction latency of only three clock cycles. The maximum peak-to-peak jitter of the output clock is 13.25ps.
Zhiting Li, Lishuo Deng, Changwei Yan, Zhangrui Qian, Weiwei Shan
ISCAS6
2025 A 40nm Early-Warning AVFS Design with Path Activation-Based Monitoring Point Optimization
abstract
This work presents an early-warning adaptive voltage-frequency scaling (AVFS) system, offering a practical and innovative low-power solution for commercial IoT products. First, a novel monitoring point optimization strategy based on path activation analysis is proposed, which mitigates the risk of illegal voltage scaling and reduces monitoring costs by 18.7%. Second, a dual shadow register-based monitor for memory endpoints is designed to overcome the limitation of monitoring only D-type flip-flop (DFF) endpoints. Third, a single-chip solution that integrates both frequency and voltage regulation strategies is presented to alleviate the issues of off-chip voltage regulation, such as large noise ripple and long response times. Fabricated in a 40-nm CMOS process, a Cortex M0+ CPU with AVFS works within a frequency range of 50-100MHz and a voltage range of 0.75-1.1V. Furthermore, the proposed AVFS achieves significant power savings of up to 49.3% at FF, 25°C, and 100MHz.
Kaize Zhou, Lishuo Deng, Junyi Qian, Xiaojie Yin, Kana Meng, Weiwei Shan
ISCAS8
2025 An All-Digital Voltage Droop Detector Utilizing a Multi-Phase Clock assisted Nyquist Counter
abstract
The on-chip droop in processor may cause a severe voltage reduction resulting in a need for on-chip voltage droop detectors. Basically, the digital droop detector indirectly senses voltage changes by reading out the delay information affected by the voltage through a time-to-digital converter (TDC). Conventional all-digital droop detector based on the ring oscillator TDC suffers from high power and area overhead. In order to reduce power consumption and area overhead without introducing the problem of metastability, two techniques are proposed. One is the multi-phase clock sampling technique to achieve the ring stage number compression, therefor the number of sampling register is diminished. The other one is the coarse-fine sharing Nyquist counter for anti-metastability, thus reducing the number of counter. Even with the introduction of the additional all-digital multi-phase clock generator, this design can still achieve an advantage in power and area consumption. Designed at 22nm process, hspice simulations show that the power consumption is reduced by 16.4% and the area is reduced by 12.4% as compared to the SOTA.
Zhangrui Qian, Kaize Zhou, Zhiting Li, Weiwei Shan
ISCAS5
2025 A Compound Timing Detection of Both Data Transition and Path Activation for Reliable In Situ Error Detection and Correction
abstract
Timing error detection and correction (EDAC) in resilient circuits helps eliminate excess timing margins. However, it faces misdetection risks when critical paths (CPs) remain inactive. We propose a 15-transistor timing error and path-activation detector (TEPD) capable of detecting both timing violations in late-arriving signals and CP activation, with robust operation down to 0.33 V. Regarding circuit-level error correction, we introduce an error-correcting flip-flop (ECFF), leveraging time-borrowing for zero-cycle response latency without requiring pipeline refresh. The custom-optimized ECFF adds only six transistors, increasing delay, dynamic power, and static power by 15%, 11%, and 14%, respectively, compared with a standard flip-flop, ensuring efficient error correction with minimal cost. A system-level voltage tuning strategy is further developed to handle continuous timing errors, ensuring robust adaptive voltage scaling (AVS) operation. Implemented on a neural network (NN) accelerator in the 28-nm CMOS, the system operates across a wide voltage range from 0.54 to 0.9 V. It achieves up to 52% power gain or 123% frequency gain at the near-threshold region, with negligible area and power overhead compared with the margined baseline.
Lishuo Deng, Junyi Qian, Zhengguo Shen, Jingchen Wang, Zhangrui Qian, Longning Qi, Weiwei Shan
IEEE Trans. Very Large Scale Integr. Syst.9
2025 An On-Chip-Training Keyword-Spotting Chip Using Interleaved Pipeline and Computation-in-Memory Cluster in 28-nm CMOS
abstract
To improve the precision of keyword spotting (KWS) for individual users on edge devices, we propose an on-chip-training KWS (OCT-KWS) chip for private data protection while also achieving ultralow -power inference. Our main contributions are: 1) identity interchange and interleaved pipeline methods during backpropagation (BP), enabling the pipelined execution of operations that traditionally had to be performed sequentially, reducing cache requirements for loss values by 95.8%; 2) all-digital isolated-bitline (BL)-based computation-in-memory (CIM) macro, eliminating ineffective computations caused by glitches, achieving 2.03$\times$higher energy efficiency; and 3) multisize CIM cluster-based BP data flow, designing each CIM macro collaboratively to achieve all-time full utilization, reducing 47.2% of output feature map (Ofmap) access. Fabricated in 28-nm CMOS and enhanced with a refined library characterization methodology, this chip achieves both the highest training energy efficiency of 101.5 TOPS/W and the lowest inference energy of 9.9nJ/decision among current KWS chips. By retraining a three-class depthwise-separable convolutional neural network (DSCNN), detection accuracy on the private dataset increases from 80.8% to 98.9%.
Junyi Qian, Peng Cao 0002, Xin Si, Weiwei Shan
IEEE Trans. Very Large Scale Integr. Syst.8
2024 AnalogGym: An Open and Practical Testing Suite for Analog Circuit Synthesis
abstract
Recent advances in machine learning (ML) for automating analog circuit synthesis have been significant, yet challenges remain. A critical gap is the lack of a standardized evaluation framework, compounded by various process design kits (PDKs), simulation tools, and a limited variety of circuit topologies. These factors hinder direct comparisons and the validation of algorithms. To address these shortcomings, we introduced AnalogGym, an open-source testing suite designed to provide fair and comprehensive evaluations. AnalogGym includes 30 circuit topologies in five categories: sensing front ends, voltage references, low dropout regulators, amplifiers, and phase-locked loops. It supports several technology nodes for academic and commercial applications and is compatible with commercial simulators such as Cadence Spectre, Synopsys HSPICE, and the open-source simulator Ngspice. AnalogGym standardizes the assessment of ML algorithms in analog circuit synthesis and promotes reproducibility with its open datasets and detailed benchmark specifications. AnalogGym's user-friendly design allows researchers to easily adapt it for robust, transparent comparisons of state-of-the-art methods, while also exposing them to real-world industrial design challenges, enhancing the practical relevance of their work. Additionally, we have conducted a comprehensive comparison study of various analog sizing methods on AnalogGym, highlighting the capabilities and advantages of different approaches. AnalogGym is available in the GitHub repository1. The documentations are also available at2.
Jintao Li 0002, Haochang Zhi, Ruiyu Lyu, Wangzhen Li, Zhaori Bi, Keren Zhu 0001, Yanhan Zeng, Weiwei Shan, Changhao Yan, Fan Yang 0001, Yun Li 0002, Xuan Zeng 0001
ICCAD8
2024 MEIC: Re-thinking RTL Debug Automation using LLMs
abstract
The deployment of Large Language Models (LLMs) for code debugging (e.g., C and Python) is widespread, benefiting from their ability to understand and interpret intricate concepts. However, in the semiconductor industry, utilising LLMs to debug Register Transfer Level (RTL) code is still insufficient, largely due to the underrepre-sentation of RTL-specific data in training sets. This work introduces a novel framework, Make Each Iteration Count (MEIC), which contrasts with traditional one-shot LLM-based debugging methods that heavily rely on prompt engineering, model tuning, and model training. MEIC utilises LLMs in an iterative process to overcome the limitation of LLMs in RTL code debugging, which is suitable for identifying and correcting both syntax and function errors, while effectively managing the uncertainties inherent in LLM operations. To evaluate our framework, we provide an open-source dataset comprising 178 common RTL programming errors. The experimental results demonstrate that the proposed debugging framework achieves fix rate of 93% for syntax errors and 78% for function errors, with up to 48x speedup in debugging processes when compared with experienced engineers. The Repo. of dataset and code: https://github.com/SEU-ACAL/reproduce-MEIC-ICCAD.
Xinwei Fang, Weiwei Shan, Xi Wang 0009, Zhe Jiang 0004
ICCAD5
2024 Ultra-low-power one-hot transmission-gate multiplexer (OTG-MUX) scalable into large fan-in circuits in 28 nm CMOS
Yuqiang Cui, Weiwei Shan, Peng Cao 0002
Integr.2
2024 Knowledge Transfer Framework for PVT Robustness in Analog Integrated Circuits
abstract
Process, voltage, and temperature (PVT) variations in chip fabrication or operation pose a significant challenge to the robustness of analog integrated circuits. Existing design techniques for mitigating PVT variations involve analyzing offsets of DC operating points, but this approach often leads to compromises in circuit performance. To address this challenge, we developed a ‘PVT-Transfer’ framework to facilitate knowledge transfer with evolutionary design. Specifically, by cross-operating the circuit parameters under variations, design knowledge is transferred through parameter migration, thus enhancing the robustness of the resultant circuit. In addition, we leverage data-driven learning to discover potential similarities among PVT variations, thereby mitigating negative knowledge transfer. The PVT-Transfer Framework is evaluated on three integrated voltage references and compared with four state-of-the-art circuit sizing methods. Based on post-layout Monte-Carlo simulations, this framework is verified to offer superior performance to existing methods, yielding a 60% reduction in power consumption, an 80% increase in temperature resilience, and up to 70$\times$enhancement in the figure of merit. Further, it leads to a 60% reduction in the number of required circuit simulations and is suitable for parallel computation.
Jintao Li 0002, Yanhan Zeng, Haochang Zhi, Jingci Yang, Weiwei Shan, Yongfu Li 0002, Yun Li 0002
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Multi-Task Evolutionary to PVT Knowledge Transfer for Analog Integrated Circuit Optimization
abstract
Designing analog integrated circuits (ICs), particularly sensors and reference circuits, requires a significant amount of human expertise and time, largely due to the requirement of maintaining process, voltage, and temperature (PVT) consistency. So far, there has been plenty of work on tuning the circuit to meet the PVT consistency requirements by comparing the offset of the DC operating point, but this inevitably leads to circuit performance degradation. To improve, we propose a ‘PVT-Transfer’ framework that utilizes knowledge transfer among PVT corners through evolutionary multitasking. Specifically, via cross-operating the circuit parameters under different PVT corners, knowledge is transferred through parameter migration to improve the robustness of the circuit. Further, PVT-Transfer employs data-driven learning to identify potential similarities among PVT variations, thereby leading to more cost-effective optimization. This framework is evaluated on two voltage references and compared with four state-of-the-art circuit sizing methods. The post-layout Monte-Carlo simulation results verify that PVT-Transfer outperforms the existing methods. It reduces the number of simulations required by 60% compared to the GCN-RL method. Besides, PVT-transfer achieves up to 10× improvement in the figure of merit over the human design.
Jintao Li 0002, Haochang Zhi, Weiwei Shan, Yongfu Li 0002, Yanhan Zeng, Yun Li 0002
ICCAD3
2023 IVATS: A Leakage Reduction Technique Based on Input Vector Analysis and Transistor Stacking in CMOS Circuits
abstract
Leakage reduction is crucial for always-on IoT applications in which static power consumption of the memory cells accounts for a large proportion of the total power. Even with high threshold voltage transistors, the leakage is still considerable. This paper proposes a novel technique based on input vector analysis and transistor stacking to analyze and suppress leakage, especially for extremely high threshold voltage (EHVT) circuits operating in the near/sub-threshold regime. At the device level, we consider the leakage ratio of each transistor terminal, which improves the universality of the method. At the circuit level, we innovatively propose the concepts of critical leakage path, leakage power components, and public leakage path to help designers locate the sources of leakage more precisely. We apply the method to a 28nm-EHVT low leakage tristate latch-like memory cell in a serial Fast Fourier Transform (FFT) circuit and find that inserting one stacking NMOS and using “01” stack to reduce the substrate leakage of PMOS can effectively suppress leakage. The average leakage power consumption of the optimized cell is reduced by 42% in the pre-layout simulation. A 26.87% and 17.52% leakage power reduction in the custom cell and the serial FFT circuit is achieved after the layout design and the synthesis.
Lishuo Deng, Weiwei Shan
ISCAS3
2023 An efficient path delay variability model for wide-voltage-range digital circuits
Weiwei Shan, Yuqiang Cui, Wentao Dai, Xinning Liu, Peng Cao 0002, Jun Yang 0006
Sci. China Inf. Sci.1
2023 Design of high-efficiency complex multiplier for fault-tolerant computation
Zhuo Chen 0039, Boyang Cheng, Weiwei Shan
Integr.4
2023 An All-Digital, 1.92-7.32 mV/LSB, 0.5-2 GS/s Sample Rate, and 0-Latency Prediction Voltage Sensor With Dynamic PVT Calibration for Droop Detection and AVS System
abstract
The on-chip droop in processor may cause a severe voltage reduction resulting in a need for high-speed and high-resolution on-chip voltage sensors. However, traditional voltage sensors hardly achieve high resolution at GHz-level sampling rate and require multiple cycles to obtain quantized results. Therefore, we propose an on-chip all-digital voltage sensor with dynamic PVT calibration for droop detection and a unified voltage monitor and scaling (UVMS) systems for efficiency improvement. We propose a balanced ring oscillator for high resolution and Nyquist counter-based encoder for GHz-level sampling. A light-weight dynamic PVT calibration using temperature sensor is proposed to resist the PVT variations and nonlinearity. Then, an adaptive prediction mechanism is proposed for 0-latency voltage detection. Fabricated in a 28nm CMOS technology, our voltage sensor achieves a high resolution of 7.32 mV/LSB at a 2 GS/s or 1.92 mV/LSB at a 0.5 GS/s sample rate with 0-latency and a calibration error of only 1 LSB. Implemented in a BNN accelerator, proposed UVMS compresses the margin and achieves a power gain of 37% - 57% at frequencies ranging from 31 MHz to 337 MHz.
Junyi Qian, Zhuo Chen 0039, Weiwei Shan
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 An energy-efficient seizure detection processor using event-driven multi-stage CNN classification and segmented data processing with adaptive channel selection
abstract
Recently wearable EEG monitoring devices with seizure detection processor using convolutional neural network (CNN) have been proposed to detect the seizure onset of patients in real time for alert or stimulation purpose. High energy efficiency and accuracy are required for the seizure detection processor due to the tight energy constraint of wearable devices. However, the use of CNN and multi-channel processing nature of seizure detection result in significant energy consumption. In this work, an energy-efficient seizure detection processor is proposed, featuring multi-stage CNN classification, segmented data processing and adaptive channel selection to reduce the energy consumption while achieving high accuracy. The design has been fabricated and tested using a 55nm process technology. Compared with several state-of-the-art designs, the proposed design achieves the lowest energy per classification (0.32 μJ) with high sensitivity (97.78%) and low false positive rate per hour (0.5).
Jiahao Liu 0006, Zirui Zhong, Hui Qiu, Jianbiao Xiao, Jiajing Fan, Zhaomin Zhang, Sixu Li, Siqi Yang 0002, Weiwei Shan, Shuisheng Lin, Liang Chang 0002, Jun Zhou 0017
DAC11
2019 Voltage-Controlled Magnetoelectric Memory Bit-cell Design With Assisted Body-bias in FD-SOI
abstract
Voltage-controlled magnetic anisotropy (VCMA)-magnetic tunnel junction (MTJ) is incorporated into FD-SOI CMOS technology. The design space of 1 transistor-1 MTJ (1T-1M) bit-cell is explored through varied VCMA pulse duration/amplitude and scaling down transistor dimensions. The design point with 1.1 V VCMA pulse amplitude, 0.44 ns pulse duration and W/L = 400 nm/30 nm access transistor shows the ultra low write energy in VCMA-MTJ based bit-cell. It achieves a minimum 3.18 fJ/bit switching energy with 28-nm FD-SOI process. Access transistor sizing is studied, while the ultra low power implementation may lead to MTJ switching failure. Voltage assisted techniques for failure mitigation are proposed based on body-bias generator (BBG). The BBG not only provides VCMA pulse signal to control MTJ barrier, but also generates body-bias to boost the transistor performance. In the presence of forward body-bias (FBB) and increased VCMA pulse level, the proposed strategy is effective in switching failure compensation as well as writing delay improvement.
Hao Cai 0001, Menglin Han, Weiwei Shan, Jun Yang 0006, You Wang 0002, Wang Kang 0001, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI3
2018 Short-path Padding Method for Timing Error Resilient Circuits based on Transmission Gates Insertion
abstract
Resilient circuits based on timing error detection and correction can mitigate the timing margin effectively, but usually at a cost of extra area overhead. One of the major sources of area overhead is short-path padding (hold time fix), which is much severer than in traditional IC design for near-threshold operation. Therefore, we propose an insertion methodology by using transmission gates to extend short-paths, which decreases area overhead than traditional resilient methods. Because the clock-controlled transmission gate (CTG) can extend all the short paths by half a clock when working as a transparent-low latch, the short-paths problem is solved. Besides, as the transmission gates synchronize the multiple short paths, it decreases the invalid flipping of combinational logic, which reduces the glitch power. Applied on a SHA-256 algorithm circuit in a 28nm CMOS process with 0.55V supply, the proposed technique reduces the area overhead a lot compared to the conventional short-path padding techniques. For combinational circuit, its area reduces from 153.34% to 4.43%, and for sequential circuit area, it reduces from 124.33% to 19.33%.
Wentao Dai, Peiye Liu, Weiwei Shan
ACM Great Lakes Symposium on VLSI3
2018 Enabling Resilient Voltage-Controlled MeRAM Using Write Assist Techniques
abstract
Reliability concerns arise in nonvolatile magnetoelectric random access memory (MeRAM) due to continuously nanotechnology scaling down and CMOS-magnetic hybrid integration. The primary objective of this work is to investigate failure mitigation in voltage-controlled magnetic anisotropy-magnetic tunnel junction (VCMA-MTJ) based 1T-1MTJ MeRAM bit-cell, by using MTJ compact model and 28nm fully depleted silicon on insulator (FD-SOI) process design-kit. A comprehensive reliability study is performed considering process variation and aging degradations, including hot carrier injection (HCI), bias temperature instability (BTI), soft breakdown (SBD) and radiation effect. Write assist techniques are proposed to ensure failure resilient MeRAM design. Bit line (BL) boost and negative source line (SL) methods show high efficiency in writing latency improvement and failure mitigation.
Hao Cai 0001, You Wang 0002, Wang Kang 0001, Lirida A. B. Naviner, Weiwei Shan, Jun Yang 0006, Weisheng Zhao 0001
ISCAS5
2015 A Secure Reconfigurable Crypto IC With Countermeasures Against SPA, DPA, and EMA
abstract
A secure reconfigurable cryptographic co-processor supporting multiple algorithms of advanced encryption standard (AES), data encryption standard (DES), rivest cipher 6, and international data encryption algorithm is proposed using its own reconfigurable feature to resist side-channel attack (SCA). It is integrated into a system-onchip and fabricated in 0.18 μm CMOS process with 1.8 V supply voltage and 100 MHz max frequency. Several kinds of specific countermeasures are proposed to hide leakage information by utilizing idle reconfigurable processing elements to do dummy operations. Its advantages lie in its little impact on area and frequency as well as high flexibility after silicon that countermeasures can also be reconfigured. Furthermore, different protections including several kinds of global countermeasures and encryption flow related countermeasures can be stacked, thus the security level can be tuned by trading for some performance or power consumption. Experimental SCA attack results show that it resists simple power analysis and differential power analysis without revealing the subkey. For correlation-based electromagnetic analysis (EMA) of DES configuration, it increases 36× measure to disclosure when applied with partial countermeasures compared to unprotected DES. As to AES configuration with full countermeasures, it resists EMA with no sign to reveal the right subkey for up to 1.2 million electromagnetic traces.
Weiwei Shan, Xingyuan Fu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 A Side-channel Analysis Resistant Reconfigurable Cryptographic Coprocessor Supporting Multiple Block Cipher Algorithms
abstract
A side-channel analysis resistant reconfigurable cryptographic coprocessor is designed and fabricated in 0.18μm CMOS with 1.8V supply and 100MHz frequency, supporting multiple block cipher algorithms of AES, DES, RC6 and IDEA. Our countermeasure utilizes idle processing elements existed in reconfigurable array to do dummy operations to hide leakage information. This method has little impact on area and frequency, and it is flexible after silicon. It resists SPA and DPA without distinguishing the encryption region. And by correlation-based electromagnetic analysis, measurement to disclosure of DES enhances 36 times with partial countermeasures and AES discloses no subkey after more than one million electromagnetic traces with full countermeasures.
Weiwei Shan, Longxing Shi, Xingyuan Fu, Chaoxuan Tian, Jun Yang 0006, Jie Li 0057
DAC1
2009 Adaptive Fuzzy Logic Controller and Its Application in MEMS Mirror Actuation Feedback Control
Weiwei Shan, Xiqun Zhu
IDEAL1