VLDB 2026 Research / reviewers in the wild / expert
Xiaoqing Wen
dblp:38/3836
· DBLP profile ↗
191ranked-venue papers
26as first author
87since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 180 · 24 first-author · 81 since 2021Software engineering, systems software and programming languages · 14 · 5 since 2021Security and privacy · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LTTL: A Low-Overhead and Triple-Node-Upset-Tolerant Latch Design for Aerospace ApplicationsabstractAs the feature size of the CMOS technology keeps scaling down, the charge sharing caused by radiation is becoming more and more prominent, and the occurrence possibility of the triple-node upset (TNU) increases significantly. In this paper, we propose a low-overhead and TNU-tolerant latch (LTTL) that leverages three parallel storage cells and an output-level error interceptive module to achieve complete TNU tolerance while minimizing design overhead. The optimized structure eliminates redundant devices and employs a high-speed D-to-Q path, significantly reducing delay-area-power product (DAPP). Even any three nodes of the latch are flipped at the same time, the output of the latch can retain the original value. Simulation results not only confirm the TNU tolerance of the proposed latch but also demonstrate that the latch can provide a 57% reduction in delay, 20% reduction in area, and 62% reduction in DAPP on average compared to state-of-the-art TNU-tolerant latches. Zikang Ma, Zhongyu Gao, Qianhui Liu, Yi Man, Huaguo Liang, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 7 |
| 2026 | Algorithm-Aided Design and Verification for Multiple-Node-Upset-Recovery Latches
Zhiyuan Pei, Haroon Waris, M. Shahzad Younis, Aibin Yan, Jiahui Deng, Yilin Gui, Xiaoqing Wen |
ISCAS | 8 |
| 2026 | Nonvolatile Flip-Flop Designs with Soft Error Recovery Based on Magnetic Tunnel Junction and CMOS for Aerospace Applications
Zhongyu Gao, Zhiyuan Pei, Wangjin Jiang, Qijun Wang, Xiaoqing Wen |
J. Electron. Test. | 7 |
| 2026 | Automated Co-Optimization Framework of Feature Selection and Ensemble Learning for Wafer Yield PredictionabstractEfficient and automated wafer yield prediction is central to cost control and process optimization in intelligent semiconductor manufacturing. As the initial testing stage in wafer inspection, Wafer Acceptance Testing (WAT) data contain critical process-related information. However, its high dimensionality, redundancy, and nonlinear inter-dependencies pose significant challenges to conventional yield prediction models, such as high computational overhead and limited generalization capability. Moreover, existing approaches often lack an automated framework capable of jointly addressing feature redundancy and model complexity. This paper proposes a collaborative and automated prediction framework that integrates feature selection and machine learning. First, an enhanced Binary Zebra Population Optimization Algorithm (BZPOA) is introduced, which incorporates redesigned exploration and development mechanisms to automatically identify key feature subsets from high-dimensional parameters, substantially reducing data redundancy and computational dimensionality. Second, a Bayesian hyperparameter-optimized XGBoost model is constructed, utilizing the Tree-structured Parzen Estimator (TPE) to achieve deep co-optimization of model parameters and feature space, thereby overcoming the inefficiency and overfitting issues commonly associated with manual parameter tuning. Experiments on a real-world dataset demonstrate that the proposed framework achieves average Recall, Precision, and F1-scores of 0.872, 0.917, and 0.902, respectively. Compared with the full-feature baseline, the BZPOA-selected feature subset improves predictive performance by 5%–10%, attains an AUC of 0.904, and significantly reduces per-wafer prediction time. Cross-factory transfer experiments further confirm the robustness of the proposed system. Tianming Ni, Muyang Cheng, Jingchang Bian, Senling Wang, Xiaoqing Wen, Mu Nie |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | A Low-Cost Input-Split Inverter-Based Triple-Node-Upset Recoverable Latch DesignabstractAs the integration level of integrated circuits continues to increase and the feature size of nanoscale chips continues to shrink, the possibility of triple-node-upsets (TNUs) occurring in circuits increases significantly. This paper proposes a latch, namely CLTNUSL, which offers stable resilience against TNUs in radiative environments while achieving a good balance between reliability and overhead. Unlike the conventional latches consisting of multi-input C-elements, the proposed CLTNUSL latch mainly consists of 16 interlocked dual-input inverters. Due to the small number of transistors, CLTNUSL achieves low area overhead. Due to the use of high-speed paths and clock gating techniques, CLTNUSL achieves low latency and low power consumption. Simulation results demonstrate CLTNUSL’s full recovery from TNU in all scenarios. Compared with the conventional TNU-hardened latch, CLTNUSL achieves minimal latency, power consumption, area overhead, and the delay-power-area product (DPAP). CLTNUSL reduces delay by 13.22%, power by 52.74%, area by 40.90%, and DPAP by 71.44% on average, compared with state-of-the-art latches. Process-Voltage-Temperature (PVT) and Monte Carlo simulation results show that the CLTNUSL latch is less sensitive to temperature, voltage and process variations compared with conventional TNU self-recovery latches. Na Bai, Yaohua Xu, Aibin Yan, Xiaoqing Wen, Yusheng Xia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | QNU-CPN: A Low-Power Single-Event Quadruple-Node-Upset Recovery LatchabstractIntegrated circuits are increasingly sensitive to radiation-induced multi-node upset in advanced CMOS technology. This paper proposes a novel low-power quadruple-node-upset recovery latch (QNU-CPN), which is based on the feedback interconnection of twenty-two input-split C-elements with P-input and N-input (CPNs) to achieve high reliability. Post-layout simulation results for 45nm CMOS by HSPICE technology show that the proposed QNU-CPN latch exhibits a reduction in power consumption by an average of 56.45%, a reduction in power-delay product (PDP) by an average of 56.92%, a reduction in area-power-delay product (APDP) by an average of 58.59%, and a reduction in setup time by an average of 11.11%, in comparison to four other existing quadruple-node upset recovery latch (LDAVPM, QRHIL, QRHIL-LC, MURLAV). Furthermore, this paper proposes the recovery rate calculation algorithm method that can calculate the recovery rate based on the configuration of multiple fault-tolerant components. Zhengfeng Huang, Linya Qiu, Shicheng Yang, Yingchun Lu, Fan Cheng 0001, Xiaoqing Wen, Aibin Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2026 | An Accelerated Newton-Based Matrix Splitting Iteration Method for Mixed-Cell-Height Circuit LegalizationabstractThe advancement of technology nodes has intensified the focus on mixed-cell-height circuit design, posing challenges to traditional legalization techniques. In this paper, we propose a novel and efficient accelerated Newton-based matrix splitting (ANMS) iteration method to address the mixed-cell-height circuit legalization problem. Our approach reformulates this problem into a generalized absolute value equation and leverages matrix splitting and the latest estimate vector to enhance computational efficiency. We also introduce a relaxation variant within the ANMS framework, namely, the accelerated Newton-based successive overrelaxation (ANSOR) method, which is particularly effective in scenarios requiring high computational performance and precise parameter tuning. The proposed method achieves linear computational complexity. Furthermore, we perform an in-depth analysis of the sufficient convergence conditions for the ANMS method and optimize cells that have excessive displacement. Experimental results show that the proposed ANMS method achieves a speedup of 1.09× – 4.94× compared to state-of-the-art methods, while maintaining the quality of solution. This makes it highly suitable for addressing complex placement design challenges. Chencan Zhou, Yang Cao 0014, Fan Yang 0001, Xiaoqing Wen, Rong Rong, Ai-Li Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | A Parallel Feedback Obfuscation Strong PUF Against Machine-Learning Modeling Attacks and Lightweight Authentication ProtocolabstractArbiter physical unclonable function (APUF) is a hardware security primitive that generates security keys by utilizing unavoidable process variations during chip manufacturing. However, the structure based on linear additive function makes it vulnerable to machine learning (ML) attacks. This paper proposes a parallel feedback obfuscation PUF (PFO PUF) design, which uses intermediate arbitration signals of the lower-layer APUF to generate the hidden challenge of the upper-layer APUF, enhancing the overall nonlinearity of the structure. The obfuscation module makes weight judgment for intermediate arbitration signals of upper-layer and lower-layer APUFs, which obfuscates the real response of PUF. We further design a variant of PFO PUF called reconfigured challenge obfuscation PFO PUF (RPFO PUF) and propose its lightweight device authentication protocol. RPFO PUF enhances the resistance of the original PFO PUF against reverse engineering (RE) and improves its Strict Avalanche Criterion (SAC) characteristic by reordering the challenges and incorporating weak PUF responses. The proposed PFO PUF and RPFO PUF were comprehensively evaluated via Python-based simulations and FPGA measurements. In Python simulations, both designs show strong resistance to state-of-the-art ML attacks, with logistic regression (LR), support vector machine (SVM), and covariance matrix adaptation evolution strategies (CMA-ES) yielding near 50% prediction accuracies under various PUF configurations. Although deep neural network (DNN) achieves up to 69.52% prediction accuracy on the PFO PUF, it drops to ∼50% on the RPFO PUF. FPGA results further confirm this, with the (32, 11)-RPFO PUF achieving a maximum prediction accuracy of only 51.47% across all four ML attacks. Moreover, both designs incur low hardware overheads, requiring just 743 and 2145 gate equivalents (GEs), respectively. Zhengfeng Huang, Yankun Lin, Yingchun Lu, Huaguo Liang, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Dependable Secur. Comput. | 9 |
| 2026 | Mercury: Practical Cross-Chain Exchange via Trusted HardwareabstractThe proliferation of blockchain-backed cryptocurrencies has sparked the need for cross-chain exchanges of diverse digital assets. Unfortunately, current exchanges suffer from high on-chain verification costs, weak threat models of central trusted parties, or synchronous requirements, making them impractical for currency trading applications. In this paper, we present MERCURY, a practical cryptocurrency exchange that is trust-minimized and efficient without online-client requirements. MERCURY leverages Trusted Execution Environments (TEEs) to shield participants from malicious behaviors, eliminating the reliance on trusted participants and making on-chain verification efficient. Despite the simple idea, building a practical TEE-assisted cross-chain exchange is challenging due to the security and unavailability issues of TEEs. MERCURY tackles the unavailability problem of TEEs by implementing an efficient challenge-response mechanism executed on smart contracts. Furthermore, MERCURY utilizes a lightweight transaction verification mechanism and adopts multiple optimizations to reduce on-chain costs. Comparative evaluations with XClaim, ZK-bridge, and Tesseract demonstrate that MERCURY significantly reduces on-chain costs by approximately 67.87%, 45.01%, and 47.70%, respectively. Xiaoqing Wen, Quanbi Feng, Jianyu Niu, Yinqian Zhang, Chen Feng 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | A Novel Approach to Reducing Testing Costs and Minimizing Defect Escapes Using Dynamic Neighborhood Range and Shapley ValuesabstractWafer acceptance testing (WAT) is a process that is used to assess the quality and reliability of manufactured wafers. This technique for the early detection and screening of chips allows for improvements in their reliability and performance during the manufacture of semiconductor devices. The automatic test equipment (ATE) used for processing millions of wafers is susceptible to a number of issues, including the absence of data values, the presence of redundant parameters, and categorical imbalance. These issues increase the cost of data processing and impede an investigation into the relationship between WAT and feature diagnostics. In this study, we propose a method with a low test escape rate based on a multi-objective optimization algorithm to reduce the cost of testing and minimize the number of defective dice that go undetected. The proposed method retains outliers, dynamically selects the range of the neighborhood to reduce the cost of testing, and uses Shapley values to analyze a WAT dataset to determine the importance of features of the data. The multi-objective optimization algorithm ranks features by their importance and applies an adaptive method to eliminate features with a low overall correlation, thereby reducing the risk that defective dice are undetected. Tianming Ni, Wangsheng Rui, Cheng Zhuo, Yu Li 0007, Xiaoqing Wen, Mu Nie |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2026 | Functional Fault Impact Probability Prediction using Spatio-Temporal Graph Convolutional NetworkabstractLogic-level defects that escape manufacturing tests pose reliability risks in modern systems and require functional testing to identify their activation and propagation behaviors. However, effective functional testing is limited by the high cost of long-cycle fault simulations. To address this challenge, we propose a Spatio-Temporal Graph Convolutional Network framework to efficiently and accurately predict the Fault Impact Probability on the circuit's function cross-over multiple function cycles, enabling rapid quantitative assessment of functionally possible faults. Our method represents gate-level netlists as spatio-temporal graphs, capturing both structural connectivity and short-range signal-propagation dynamics. With dedicated spatial and temporal encoders, the proposed ST-GCN enables accurate prediction of multi-cycle circuit-level FIP. Experiments on ISCAS’89 benchmarks show that the approach reduces fault-simulation cost by over an order of magnitude while maintaining high accuracy (mean absolute error as low as 0.024 for 5-cycle predictions). The framework supports both testability-metric-based and simulation-based feature construction, enabling a tunable balance between efficiency and accuracy. A case study on test point selection further demonstrates that using predicted FIPs to guide observation-point placement improves the detectability of multi-cycle, hard-to-detect circuit-level faults. Overall, this work provides a scalable solution for circuit-level multi-cycle fault-impact assessment and can be readily integrated into functional test generation and other Electronic Design Automation workflows. Shaoqi Wei, Senling Wang, Hiroshi Kai, Yoshinobu Higami, Ruijun Ma 0002, Tianming Ni, Xiaoqing Wen, Hiroshi Takahashi |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2026 | High-Accuracy, Low-Utilization Multichannel 2-ps Bin Size FPGA Digital-to-Time Converter Based on Compact Multidimensional Delay ArrayabstractDigital-to-time converter (DTC) or digital delay/timing generator has been developed for quite many years and plays a crucial role in automatic test equipment (ATE) and built-in self-test (BIST) industries. This study proposes a compact multidimensional delay array DTC based on phase shift in a phase-locked loop (PLL) to further extend field-programmable gate array (FPGA) applications into the analog domain. All internal delay cells are precisely controlled by PLL from the beginning to make the output phases distributed within the reference clock period as uniformly as possible. For further resolution enhancement, a hybrid structure based on PLL and a multidimensional delay array is presented to ensure high enough accuracy and substantially reduce logic utilization through phase sorting and selection. For concept proving, the proposed four-channel DTC is implemented on an Altera Stratix IV FPGA board to achieve 20 internal control bits, 2-ps resolution with very low integral nonlinearity (INL) and differential nonlinearity (DNL) of −2.06 to 2.01 and −2.50 to 2.31 LSB, respectively. In addition, the circuit has been successfully implemented on a much cheaper Cyclone IV platform also for cost reduction to achieve the same resolution with the best fine stage INL and DNL of −3.81 to 3.56 and −3.93 to 4.41 LSB, respectively. The DTC performance has demonstrated improvements to that of prior arts by one order finer resolution and higher accuracy compared to non-Vernier prior works, while eliminating the serious dead time issues inherent in Vernier DTCs. Poki Chen, Joshua Adiel Wijaya, Xiaoqing Wen, Stefan Holst |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2026 | RMC PUF: A Highly Reliable PUF Architecture Based on Recursive Markov Chain ObfuscationabstractPhysical unclonable functions (PUFs) are critical hardware security primitives that extract entropy from intrinsic manufacturing process variations (MPVs) to generate unique cryptographic responses. They have demonstrated extensive application prospects in lightweight encryption and authentication for Internet of Things (IoT) devices. However, the classic arbiter PUF (APUF) is inherently vulnerable to machine learning (ML) modeling attacks due to the linear mathematical model of its delay circuits. To address this challenge, this article proposes a novel postprocessing obfuscation architecture named recursive Markov chain PUF (RMC PUF). The proposed architecture couples an APUF core with Markov obfuscation modules. By exploiting the intermediate stage responses of the APUF as an entropy source, the system constructs a 1-D recursive stochastic state transition mechanism. This approach recursively transforms and obfuscates the raw entropy based on Markov Chain theory, achieving a PUF system with stability and robust resistance against ML attacks. Through this efficient recursive obfuscation strategy, the design achieves precise utilization of the APUF’s intrinsic entropy. The work is validated through both numerical simulation and FPGA prototyping. Experimental results show that the prototype maintains a high reliability of 98.9%–99.9% within a voltage range of 0.8–1.2 V, and a reliability of 96.1%–99.7% within a temperature range from$- 10~{^{\circ }}$C to$80~{^{\circ }}$C. Comparative analysis indicates that the proposed RMC PUF significantly outperforms existing structures in security, limiting the prediction accuracy of four ML attack models [logistic regression (LR), artificial neural networks (ANNs), deep neural networks (DNNs), and covariance matrix adaptive evolutionary strategy (CMA-ES)] to no more than 54.72%. Zhengfeng Huang, Xuxiang Sun, Yingchun Lu, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2026 | NVLIM: MTJ and CMOS-Based Nonvolatile Latch Design With Protection Against Triple-Node-Upsets for Robust ComputingabstractSoft errors and power dissipation emerge as critical challenges in developing high-reliability and cost-sensitive embedded systems. To address these issues, the magnetic tunnel junction (MTJ) is considered a promising solution due to its nonvolatility and its compatibility with traditional CMOS manufacturing processes. In this work, we propose a novel nonvolatile (NV) latch consisting of inverters and MTJs, namely, NVLIM, which provides nonvolatility and robust partial tolerance against triple-node-upsets (TNUs) at low cost. NVLIM integrates a TNU-tolerant block based on CMOS with a backup-restore block using MTJs. Simulation results incorporating process, voltage, and temperature (PVT) variations, bias temperature instability (BTI) impact, and Monte Carlo simulations demonstrate the balanced performance in terms of nonvolatility, robust partial TNU tolerance, and comprehensive overhead of the proposed latch. Aibin Yan, Litao Wang, Zhengfeng Huang, Qingyang Zhang 0001, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | Radiation-Resistant ZnO Thin-Film Transistor Voltage ReferenceabstractThis paper presents a radiation-tolerant voltage reference circuit based on atomic layer-deposited zinc oxide thinfilm transistors (ZnO TFTs), where both active and floating gate layers utilize ZnO films. Employing a dual-exponential current source method, we accurately simulate single-event transient currents induced by high-energy particles. Measurement results demonstrate that the proposed circuit achieves $5 \mu \mathrm{~s}$ recovery time without radiation hardening, reducible to $1 \mu \mathrm{~s}$ with hardening techniques. The circuit generates a 3.1 V reference voltage operating from 6 V to 10 V supply, with $18.78 \mathrm{mV} / \mathrm{V}$ line sensitivity and $0.72 \mu \mathrm{~W}$ static power consumption. Temperature stability is achieved through a complementary temperature coefficient. From $0^{\circ} \mathrm{C}$ to $80^{\circ} \mathrm{C}$, the temperature coefficient of VREF is $79.56 \mathrm{ppm} /{ }^{\circ} \mathrm{C}$, and the PSRR reaches -36.5 dB. Na Bai, Guocheng Ge, Hanxiang Li, Aibin Yan, Xiaoqing Wen |
ATS | 6 |
| 2025 | Software-Defined Secure Island for Testing Chiplet SystemsabstractChiplet systems that stack heterogeneous dies via 2.5D/3D integration use JTAG-based test access ports (TAPs) to validate inter-die links and enable in-field diagnosis. However, these TAPs create a shared attack surface: penetrating a single die can potentially expose control of the entire stack. Current countermeasures involve embedding a complete cryptographic engine in each chiplet, which increases the area and locks the protocol at tape-out, leaving it susceptible to future unknown attacks. This paper proposes a Software-Defined Secure Island (SDSI) architecture that decouples security policies from hardwired logic while satisfying non-functional requirements. Each chiplet instantiates an ultra-lightweight Secure Island Controller (SIC) macro that performs only two XORs and two additions per handshake. Meanwhile, a centralized secure island runs heavyweight cryptography in firmware using the multi-round SASL-JTAG+ protocol. SDSI enables scalable, adaptable test access protection by decoupling security from hardware. An FPGA implementation shows that the SIC macro occupies 11% of the area of an AES- 128 core and 37% of a SHA256 core, yet it supports 256 - to 512-bit keys with negligible growth. A security analysis demonstrates immunity to replay attacks because the authentication data is refreshed with each session. All future upgrades, such as longer keys, stronger hashes, and additional rounds, are delivered via firmware, providing scalable, field-upgradable protection for heterogeneous chiplet systems. Hisashi Okamoto, Senling Wang, Hiroshi Kai, Hiroyuki Yotsuyanagi, Yoshinobu Higami, Tianming Ni, Tai Song, Hiroshi Takahashi, Xiaoqing Wen |
ATS | 9 |
| 2025 | LLM-Design Platform for Thermal-Failure-Aware 3D Chiplet Layout via Iterative Parameter AnalysisabstractAlthough advanced 3D chiplet packaging helps extend Moore’s-Law gains through vertical stacking of heterogeneous dies, it simultaneously introduces unprecedented thermal challenges. Shrinking 3D interconnects and rising current densities trap heat at through-silicon vias (TSVs) and micro-bumps, causing steep temperature gradients and thermo-mechanical stress that directly trigger early failures. Existing methods interact with finite-element analysis (FEA) tools mainly through manual, experience-based parameter tuning, inevitably making hot-spot detection time-consuming and error-prone. This paper establishes a text-interactive, thermal-failure-aware design framework for 3D chiplets that couples large language model (LLM)-accelerated thermal analysis with FEA simulation. A Python-driven Automatic Code Generation Model (ACGM) is invoked to automate geometry drawing, meshing, and boundary-condition assignment, enabling rapid analysis and prediction of hot-spot locations within densely stacked chiplets and TSVs. The proposed ACGM eliminates domain-expert dependence by abstracting cumbersome FEA parameter setting into natural-language commands, thereby lowering the entry barrier and cutting operation time. Deployment of the proposed ACGM streamlines intricate FEA operations and Python scripting into concise natural-language commands, enabling accurate layout-parameter tuning and shortened design cy-cles-establishing an AI-centric design-verification paradigm for advanced 3D chiplets. Tai Song, Senling Wang, Xiaoqing Wen |
ATS | 3 |
| 2025 | Efficient Modulated State Space Model for Mixed-Type Wafer Defect Pattern RecognitionabstractAccurate and efficient wafer defect detection is crucial in semiconductor manufacturing to maintain product quality and optimize yield. Traditional methods struggle with the complexity and diversity of modern wafer defect patterns. While deep learning approaches are effective, they are often resource-intensive, posing challenges for real-time deployment in industrial settings. To solve these problems, we propose an Efficient Modulated State Space Model (EM-SSM) for mixed-type wafer defect recognition, optimized with knowledge distillation to balance accuracy and efficiency. Our framework captures size-dependent relationships and improves defect-specific feature representation to recognize complex defects precisely. Specifically, we introduce an efficient directional modulation mechanism to refine spatial recognition of defect patterns. To further improve inference efficiency, we propose a deep-to-shallow distillation method that transfers knowledge from deeper networks to lighter networks, reducing inference time without compromising classification accuracy. Experimental results on the MixedWM38 wafer dataset with 38 defect types show that our model achieves 99.0% accuracy, outperforming traditional methods in both accuracy and efficiency. Our model offers a scalable solution for modern semiconductor defect detection. Mu Nie, Shidong Zhu, Aibin Yan, Cheng Zhuo, Xiaoqing Wen, Tianming Ni |
DATE | 5 |
| 2025 | NVSRLO: A FeFET-Based Non-Volatile and SEU-Recoverable Latch Design with Optimized OverheadabstractThis paper presents a FeFET-based non-volatile and single-event upset (SEU) recoverable latch, namely NVSRLO, which does not require any extra control signals. Simulation results show that the proposed latch provides non-volatility and SEU-recovery with optimized overhead. Compared with existing non-volatile latches, NVSRLO significantly reduces delay, power, and delay-power-area product at the cost of area. Aibin Yan, Wangjin Jiang, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
DATE | 6 |
| 2025 | Achilles: Efficient TEE-Assisted BFT Consensus via Rollback Resilient RecoveryabstractBFT consensus that uses Trusted Execution Environments (TEEs) to improve the system tolerance and performance is gaining popularity. However, existing works suffer from TEE rollback issues, resulting in a tolerance-performance tradeoff. In this paper, we propose Achilles, an efficient TEE-assisted BFT protocol that breaks the tradeoff. The key idea behind Achilles is removing the expensive rollback prevention of TEEs from the critical path of committing transactions. To this end, Achilles adopts a rollback resilient recovery mechanism, which allows nodes to assist each other in recovering their states. Besides, Achilles follows the chaining spirit in modern chained BFT protocols and leverages customized chained commit rules to achieve linear message complexity, end-to-end transaction latency of four communication steps, and fault tolerance for the minority of Byzantine nodes. Achilles is the first TEE-assisted BFT protocol in line with CFT protocols in these metrics. We implement a prototype of Achilles based on Intel SGX and evaluate it in both LAN and WAN, showcasing its outperforming performance compared to several state-of-the-art counterparts. Jianyu Niu, Xiaoqing Wen, Guanlong Wu, Shengqi Liu, Jiangshan Yu, Yinqian Zhang |
EuroSys | 2 |
| 2025 | TNURML: Triple-Node-Upset-Recovery Magnetic Latch Design with Non-Volatility for Aerospace ApplicationsabstractAs semiconductor technology advances, radiative-particle-induced soft errors and power consumption are becoming major concerns for digital circuits in aerospace applications. Radiation hardening by design and magnetic tunnel junctions (MTJs) are widely employed to address these concerns. In this paper, a novel latch, called TNURML, that can completely recover from triple-node upsets (TNUs), is proposed. The embedded MTJs provide non-volatility and are compatible with traditional CMOS processes. The TNURML employs a TNU-recovery module as well as a pair of MTJs for backup and recovery operations. Extensive simulations demonstrate the excellent TNU-recovery capability and non-volatility of the TNURML latch at the cost of slightly increased area overhead. The proposed TNURML latch reduces 43.63% of delay and 48.23% of power on average when compared to the state-of-the-art latches. Aibin Yan, Zhiyuan Pei, Cuiyun Jiang, Huaguo Liang, Xiaoqing Wen, Patrick Girard 0001 |
ISCAS | 7 |
| 2025 | A lightweight general PUF framework for resisting machine learning attacks
Tianming Ni, Zhengfeng Huang, Aibin Yan, Senling Wang, Xiaoqing Wen, Mu Nie, Jingchang Bian |
Integr. | 6 |
| 2025 | Graph-Based Multitask Transfer Learning for Fault Detection and Diagnosis of Few-Shot Analog CircuitsabstractBuilding an interpretable fault detection and diagnostic model based on few-shot circuit samples and prior information about circuit structures is of significant importance. To fill these gaps, we propose a graph-based multitask transfer learning (TL) method for fault detection and diagnosis of circuits under few-shot conditions. First, in order to model the interconnections of nodes in a circuit, the sample data is organized into a graph structure, and a semi-supervised graph-based structural feature fusion method is proposed. The proposed method can accept graph-structured data and process the data using feature fusion methods. Second, to improve the model performance under few-shot conditions, two TL mechanisms are proposed for the topological structure characteristics of analog circuits as well as circuit signal characteristics. Finally, through a parameter-shared strategy, we propose a task transfer-based fault diagnosis approach. Experimental results on three different circuits show that the proposed method has the best diagnostic accuracy compared to typical detection and diagnosis schemes. Zhongyu Gao, Aibin Yan, Zhengfeng Huang, Jie Cui 0004, Byeong-Hee Roh, Guangzhu Liu, Patrick Girard 0001, Xiaoqing Wen |
IEEE Internet Things J. | 8 |
| 2025 | TeeRollup: Efficient Rollup Design Using Heterogeneous TEEabstractRollups have emerged as a promising approach to improving blockchains’ scalability by offloading transaction execution off-chain. Existing rollup solutions either leverage complex zero-knowledge proofs or optimistically assume execution correctness unless challenged. However, these solutions suffer from high gas costs and significant withdrawal delays, hindering their adoption in decentralized applications. This paper introducesTeeRollup, an efficient rollup protocol that leverages Trusted Execution Environments (TEEs) to achieve both low gas costs and short withdrawal delays. Sequencers (i.e., system participants) execute transactions within TEEs and upload signed execution results to the blockchain with confidential keys of TEEs. Unlike most TEE-assisted blockchain designs,TeeRollupadopts a practical threat model where the integrity and availability of TEEs may be compromised. To address these issues, we first introduce a distributed system of sequencers with heterogeneous TEEs, ensuring system security even if a certain proportion of TEEs are compromised. Second, we propose a challenge mechanism to solve the redeemability issue caused by TEE unavailability. Furthermore,TeeRollupincorporates Data Availability Providers (DAPs) to reduce on-chain storage overhead and uses a laziness penalty mechanism to regulate DAP behavior. We implement a prototype ofTeeRollupin Golang, using the Ethereum test network, Sepolia. Our experimental results indicate thatTeeRollupoutperforms zero-knowledge rollups (ZK-rollups), reducing on-chain verification costs by approximately 86% and withdrawal delays to a few minutes. Xiaoqing Wen, Quanbi Feng, Hanzheng Lyu, Jianyu Niu, Yinqian Zhang, Chen Feng 0001 |
IEEE Trans. Computers | 1 |
| 2025 | Reconfigurable Radiation-Hardened SRAM Cell Design for Different Radiation EnvironmentsabstractThis article proposes a novel and effective 14-transistors (14T) reconfigurable radiation-hardened static-random access-memory cell design under the SMIC 65-nm process, featuring a unique memory reconfigurability architecture with two operation modes, namely the high reliability (HR) mode and the triple-time memory (TTM) mode for meeting different radiation environmental requirements. The proposed HR mode provides strong protection of the memory arrays in harsh radiation environments. Compared with the traditional triple modular redundancy (TMR) structure, the proposed HR mode reduces area overhead by 30%, delay by 37%, and power consumption by 16%. The TTM mode uses the enable (EN) circuit to expand the capacity threefold in less harsh radiation environments, avoiding the area wastage caused by the traditional TMR structure. By implementing the two innovative operation modes, the proposed design overcomes the limitations of the traditional TMR structure, reducing area overhead while retaining the radiation hardening capability. In addition, this article presents a mode-switching mechanism composed of a detection circuit and an EN circuit. The detection circuit can detect errors in the reconfigurable architecture. With the proposed mode-switching mechanism, two operation modes can switch in response to different radiation environments. Besides, to ensure the normal operations of the TTM mode in radiation environments, the proposed 14T cell serves as a bitcell in the memory reconfigurable architecture. Compared with typical existing designs, such as radiation-hardened based design, writability enhanced, and dual interlocked storage cell (DICE) cells, the proposed 14T cell design has better delay, critical charge, and higher hold static noise margin. Na Bai, Yaohua Xu, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | HALTRAV: Design of a High-Performance and Area-Efficient Latch With Triple-Node-Upset Recovery and Algorithm-Based VerificationsabstractWith the rapid advancement of semiconductor technologies, latches become increasingly sensitive to soft errors, especially triple node upsets (TNUs), in harsh radiation environments. In this article, we first propose a high-performance and area-efficient latch, namely, HALTRAV, featuring complete TNU-recovery. The storage portion of HALTRAV consists of 28 interlocked source-drain cross-coupled inverters (SCIs) for complete TNU-recovery with area efficiency and low delay. To mitigate the issue that node-upset-recovery verifications for existing latches highly relies on electronic design automation tools, we further propose an algorithm-based verification method that can automatically verify the node-upset-recovery of latches, which greatly simplifies the reliability-verification flow. Simulation results demonstrate the TNU-recovery of HALTRAV and also show that HALTRAV achieves 40.38%, 8.17%, and 31.89% reduction in delay, area, and delay-power–area product (DPAP) on average, respectively; however; it is at the cost of power as compared to typical latches that are TNU-recoverable. Comparison results also demonstrate the moderate sensitivity of HALTRAV to the impacts of the process, voltage, and temperature (PVT) variations. Zhenmin Li, Xiaoqing Wen, Patrick Girard 0001, Aibin Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Low-Cost Quadruple-Node-Upset Self-Recoverable Latch Based on Cross-InterlockingabstractAs the CMOS technology continues to shrink, latches are becoming increasingly susceptible to multiple-node-upset caused by charge sharing in radiation environments. In this article, a low-cost quadruple-node-upset (QNU) self-recoverable latch based on cross-interlocking (Quad-CIRC) is proposed. By utilizing four cross-interlocking self-recoverable cells (CIRCs) for interlocking, complete QNU self-recovery is achieved with reduced sensitive nodes. Meanwhile, the majority of currently available QNU self-recoverable latches are primarily composed of C-elements-based redundancy, resulting in a significant increase in area overhead. However, Quad-CIRC effectively reduces area overhead while ensuring hardened capability through cross-interlocking of CIRCs. HSPICE-based simulations in 22 nm CMOS technology demonstrate that Quad-CIRC achieves a reduction of 69.08% on average of power consumption, an increase of 16.83% on average of delay, a reduction of 63.51% on average of power-delay-product (PDP), a reduction of 51.01% on average of area, a reduction of 83.44% on average of area-PDP (APDP), and an increase of 60.49% on average of critical charge, compared to five other QNU self-recoverable latches (QRHIL, MURLAV, LDAVPM,$QR-R_{11}-C_{2}$, and low-delay QNU self-recoverable). Zhengfeng Huang, Lei Ai, Yingchun Lu, Tai Song, Xiaoqing Wen, Aibin Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | Cost Efficient Flip-Flop Designs With Multiple-Node Upset-Tolerance and Algorithm-Based VerificationsabstractThis article presents radiation-hardened flip-flop (FF) designs capable of tolerating soft errors, e.g., single-node upsets (SNUs), double-node upsets (DNUs) and multiple-node upsets (MNUs). First, a 2-input FF and a 3-input FF are proposed as the baseline FFs that not only, respectively, tolerate SNUs and DNUs but also exhibit cost efficiency in terms of delay, power, and area. Through adding two stages of c-elements, a 4-input FF and a 5-input FF are proposed as the baseline FFs as well. Utilizing the structural characteristics of these FFs, an$N-1$input FF and an N input FF are proposed as the extended FFs capable of tolerating more node upsets. Moreover, a highly efficient algorithm for verifying MNU-tolerance of these FFs is proposed. Algorithm and HSPICE-tool-based verification results both demonstrate the MNU-tolerance for the proposed FFs with more inputs. Aibin Yan, Zhengfeng Huang, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2025 | A High-Performance Low-Power Double-Node Upset Resilient Latch for Harsh Radiation EnvironmentsabstractWith the advancement of semiconductor technology, circuits have become increasingly susceptible to errors induced by radiation. Traditional approaches to enhancing the resilience of circuits against single-node upsets (SNUs) are insufficient to meet the robustness standards of modern designs. This article proposes a high-performance, low-power latch, named high-performance low-power double-node upset resilient latch (HLDRL), which is designed to exhibit exceptional resilience against double-node upsets (DNUs). Its design has six intricately interconnected C-elements (CEs) and two three-input CEs, for error interception, ensuring robust performance even in the case of DNUs. The Technology Computer Aided Design (TCAD) tool is used to validate the effectiveness of the HLDRL. Besides, comprehensive simulations are conducted utilizing the advanced SMIC 55-nm process technology. These simulation results show that our proposed HLDRL latch can autonomously recover from any DNU and thereby ensure the integrity of the system. Moreover, compared with existing DNU-resilient latches, the proposed HLDRL latch exhibits substantial improvements in terms of multiple metrics. On average, the proposed latch achieves an impressive 29.39% dynamic power saving, a remarkable 40.04% increase in speed, a notable 4.81% reduction in area, and a substantial 53.96% decrease in the power-delay–area product (PDAP). In the post-layout simulation, the proposed latch achieves a 30.39% dynamic power saving, a 36.47% increase in speed, and an impressive 52.38% decrease in PDAP. Furthermore, the proposed latch demonstrates enhanced resilience against variations in process, supply voltage, and temperature (PVT). Na Bai, Yusheng Xia, Yaohua Xu, Yi Wang 0073, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | BF PUF: A Modeling Attack-Resistant Strong PUF Based on Bent FunctionsabstractStrong physical unclonable functions (PUFs) are promising circuits for lightweight Internet of Things (IoT) authentication and security. However, existing strong PUFs exhibit very low cryptographic nonlinearity (NL), making them vulnerable to machine learning (ML) modeling and cryptanalytic attack. To address this issue, we propose the Bent function PUF (BF PUF) based on Maiorana-McFarland (M-M) constructed Bent functions, which obfuscates the responses of the strong PUF to enhance resistance against modeling attacks. The core idea is to employ the M-M construction method for Bent functions to ensure maximum cryptographic NL to resist modeling attacks. A Feistel network is configured using weak PUF responses as keys to achieve device-specific and unpredictable mappings of input challenges while meeting the requirements of the M-M Bent function construction. A Python-based model of the BF PUF was developed, and simulation results indicate that the cryptographic NL of the proposed BF PUF outperformsk-xorarbiter PUFs (APUFs) (${k} =2$, 4, 6). The proposed BF PUF was also implemented and evaluated on the FPGA hardware platform. The experimental results show that under modeling attacks using four ML algorithms—logistic regression (LR), artificial neural networks (ANNs), deep neural networks (DNNs), and covariance matrix adaptation evolution strategies (CMA-ES)—the best prediction accuracy under these four modeling attack algorithms is 52.60%. The reliability under temperature fluctuations ranging from$- 10~^{\circ }$C to$80~^{\circ }$C is between 84.20% and 99.78%. Zhengfeng Huang, Fansheng Zeng, Yanqiao Chi, Yankun Lin, Yingchun Lu, Huaguo Liang, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2025 | Highly Defect Detectable and SEU-Resilient Robust Scan-Test-Aware Latch DesignabstractSoft errors have been a severe threat to the reliability of modern integrated circuits (ICs), making hardened latch designs indispensable for masking soft errors with redundancy. However, the added redundancy also masks production defects as soft errors; this makes it hard to detect defects in hardened latches, thus significantly reducing their reliability. Our previous work proposed the scan-test-aware hardened latch (STAHL) design, the first for addressing the issue of low defect detectability of hardened latch designs. However, STAHL still suffers from two problems: 1) it is not self-resilient to soft errors and 2) a STAHL-based scan design requires one additional control signal. This article proposes a high defect detectable and single-event-upset (SEU)-resilient robust (HIDER) latch to address the issues of the low defect detectability of existing hardened latches and the STAHLs lack of SEU-resilient capability. Two scan designs [HIDER-based scan-cell-S (HIDER-SC-S) and HIDER-based scan-cell-F (HIDER-SC-F)], as well as two corresponding test procedures, are proposed to fully test HIDER latch with only one control signal. Simulation results show that the HIDER latch achieves the highest defect coverage (DC) in both single latch cell detection and scan tests among all existing hardened latch designs. In addition, the HIDER latch has much lower power and a smaller delay than STAHL. Ruijun Ma 0002, Stefan Holst, Xiaoqing Wen, Senling Wang, Jiuqi Li, Aibin Yan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A Response-Nonlinearized DEMUX-TDC PUF for Resistance Against Modeling Attacks and Secure Authentication ProtocolsabstractAs a critical hardware security primitive for the authentication within the Internet of Things (IoT), the physical unclonable function (PUF) represents an innovative security design paradigm for integrated circuits. However, the linear challenge-response mapping of the arbiter PUF (APUF) and variants render these structures more susceptible to modeling attacks due to their delayed linear structure. In this article, we propose a nonlinearized demultiplexer time-to-digital converter (DEMUX-TDC) PUF. This PUF uses a quantized delay difference technique to alter the traditional response generation mechanism, demonstrating robust resistance against modeling attacks. First, the proposed PUF employs a segmented APUF variant structure configured in both front and back segmented modes to generate a source of delay difference entropy. Additionally, the scheme incorporates a multilayer differential tapped TDC circuit to quantize the delay differences into digital codes, followed by a linear feedback shift register (LFSR) to obfuscate the final output response. We further propose a highly secure mutual authentication protocol based on reconfigurable PUF, by leveraging the characteristics of front and back segments of the PUF’s challenge. Evaluation of the proposed scheme on the implementation of Xilinx Virtex-7 and Spartan-6 field-programmable gate array (FPGA) demonstrates that the uniqueness and uniformity can reach ideal value, in the condition of prediction accuracy across six modeling attacks remaining around 50%. Tianming Ni, Mu Nie, Aibin Yan, Senling Wang, Xiaoqing Wen, Jingchang Bian |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | RHT_NoC: A Reconfigurable Hybrid Topology Architecture for Chiplet-Based Multicore SystemabstractChiplet-based system-on-chip (SoC) architectures, leveraging 2.5-D/3-D integration technologies, provide scalable solutions for a wide range of applications. Achieving high performance and cost-effectiveness in these systems relies heavily on optimizing die-to-die interconnect topologies and designs, which are essential for seamless interchiplet communication. This article introduces a reconfigurable hybrid topology (RHT) architecture designed for chiplet-based multicore systems. RHT achieves high performance and energy efficiency by dynamically reconfiguring the network topology to traffic variations, adaptively selecting transport subnets, and optimizing link bandwidth allocation, thereby minimizing congestion and maximizing packet throughput. Furthermore, RHT leverages global traffic information to dynamically combine Torus loops, maximizing opportunities for rapid packet transmission delivery while guaranteeing minimal hop counts. Moreover, RHT accelerates packet transmission via bufferless combined loops, extending the continuous sleeping periods of routers, improves power gating efficiency, and significantly reduces static power consumption. Simulation results indicate that the Mesh-DyRing achieves over a 40% reduction in network latency and more than a 20% decrease in power consumption overhead compared to the baseline design. When compared to WiNoC, an advanced hybrid wired-wireless topology design, the Mesh-DyRing-PG configuration reduces power consumption by 56.2% while maintaining equivalent average network latency. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Cost-Optimized Double-Node-Upset-Recovery Latch Designs With Aging Mitigation and Algorithm-Based Verification for Long-Term Robustness EnhancementabstractWith the continuous advancement of CMOS technologies, soft errors, such as single-node upset (SNU) and double-node upset (DNU), caused by radiation in nanoscale integrated circuits, are becoming increasingly prominent. Meanwhile, transistor aging mitigation is indispensable for long-term robustness enhancement. First, to reduce the impact of radiation on circuits, we propose a novel DNU-recovery latch with low cost, namely, DURLC, only consisting of four dual-input C-elements (CEs) and four clock-gated input-split inverters for the storage of values. Second, we propose a DNU-recovery latch with moderate cost, namely, DURMC, based on seven CEs and four inverters, for convenience to optimize the latch to alleviate aging. The proposed DNU-recovery latch with mitigated aging is called DURMA. The latch employs a high-speed path to reduce delay without sacrificing performance when mitigating aging issues. Finally, we propose an algorithm-based verification method to validate the DNU recovery of the proposed latches. The simulation results show that, compared with the state-of-the-art robust latches, the proposed latches have the advantages of DNU recovery with moderate and even low cost, and meanwhile, aging is effectively mitigated for the DURMA latch. Aibin Yan, Changli Hu, Na Bai, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | Design of Nonvolatile and Multinode-Upset Recoverable Latches Based on Magnetic Tunnel Junction and CMOSabstractSpintronic devices, such as magnetic tunnel junctions (MTJs), are promising for space applications due to their radiation hardness and nonvolatility. However, as semiconductor technology advances, CMOS peripheral circuits are becoming vulnerable to double node upset (DNU) as well as triple node upset (TNU). This brief proposes two nonvolatile and robust latch designs primarily composed of MTJs and C-elements (CEs). Both designs offer nonvolatility and self-recovery from multiple-node upsets. Simulation results demonstrate that the proposed latches provide nonvolatility and complete protection against multiple-node upsets with balanced overhead. Aibin Yan, Yongkang Xu, Na Bai, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | TUTPFL: Triple Node Upset-Tolerant and Single-Event Transient-Filtered Low-Power Latch With HSPICE and FPGA-Based VerificationsabstractIn nanoscale CMOS technology, harsh radiations in the environment can now easily cause soft errors, e.g., single-event transients (SETs) and triple node upsets (TNUs), severely affecting the reliability of space applications. In this article, TNUs tolerant and SET-pulses filtered latch (TUTPFL) with low power is proposed, which comprises from four input-stage C-elements (CEs), four inverters, and three output-stage CEs. The CEs’ delay differential enables the TUTPFL latch to effectively filter SET-pulse, while the CEs’ multilevel error-interception property enables the TUTPFL latch to tolerate any possible TNU. The results of HSPICE-based simulations and FPGA-based emulations demonstrate the TNU tolerance and SET filterability of the TUTPFL latch. Meanwhile, compared to the alternative radiation-hardened latches, the TUTPFL latch reduces power dissipation by roughly 20.43% on average. Aibin Yan, Xiumin Xu, Hanxiang Li, Na Bai, Zhengfeng Huang, Xiaoqing Wen, Patrick Girard 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | Reconfigurable Fault-Tolerant Link With Bandwidth Expansion for 2.5-D Chiplet-Based SystemsabstractThe 2.5-D chiplet-based systems offer a promising path toward higher performance and integration density, but the reliability of interchiplet links, particularly vertical links (VLs), poses a significant challenge. This article proposes reconfigurable bidirectional link (ReBL), a novel architecture providing robust fault tolerance and enhanced bandwidth for these critical interconnects. ReBL features three key innovations: 1) a robust fault-tolerant mechanism leveraging ReBLs that dynamically adapts to and mitigates permanent link failures; 2) a dynamic bandwidth expansion technique that significantly enhances the system efficiency by utilizing idle links to optimize resource allocation and throughput; and 3) a virtual channel (VC) allocation strategy that guarantees deadlock-free operations through strategic channel partitioning and assignment. Evaluations using synthetic traffic and PARSEC benchmarks demonstrate ReBL’s significant advantages under high fault rates. Compared with the state-of-the-art reliable and deadlock-free routing (ReD) approach, ReBL achieves an average reduction of 21.3% in packet latency and 4.9% in application execution time across the evaluated benchmarks. These benefits are achieved with only a 6.06% area overhead over baseline. Wu Zhou 0007, Le Luo 0002, Fulong Chen 0002, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | IDLD: Interlocked Dual-Circle Latch Design with Low Cost and Triple-Node-Upset-Recovery for Aerospace ApplicationsabstractModern powerful CMOS chips are usually highly integrated and implemented with aggressively shrunk technology nodes. In radiation environment, under charge-sharing mechanism, one particle striking can simultaneously impact multiple nodes causing double-node-upsets (DNUs) and triple-node-upsets (TNUs). In this paper, we propose an Interlocked Dual-circle Latch Design, namely IDLD, with low cost and TNU recovery for aerospace applications. IDLD consists of four transmission gates and twelve 2-input C-elements (CEs) implemented in 22nm CMOS process. Simulation results demonstrate the complete TNU recovery as well as cost-effectiveness for the proposed IDLD latch. Aibin Yan, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2024 | Nonvolatile and SEU-Recoverable Latch Based on FeFET and CMOS for Energy-Harvesting DevicesabstractNonvolatile memories are widely used in emerging energy-harvesting Internet-of-Things (IoT) applications, and nonvolatile memories constructed from FeFET devices hold great promise. This paper presents a nonvolatile and single-event-upset (SEU)-recoverable latch based on FeFET and CMOS for energyharvesting devices. The latch uses n-type FeFET devices to provide nonvolatility without any additional control signals. Moreover, since the soft error problem has become increasingly severe, radiation hardening by design gains a great attention as a promising approach to mitigate the reliability issue. The latch uses feedback interlocked loops with n-type FeFETs and C-elements, enabling it to provide nonvolatility and SEU-recovery simultaneously. Simulation results with Candence Virtuoso verifies that the proposed latch design has correct functioning with excellent performance compared to the state-of-the-art designs. Aibin Yan, Zhuoyuan Lin, Guangzhu Liu, Qingyang Zhang 0001, Zhengfeng Huang, Jie Cui 0004, Xiaoqing Wen, Patrick Girard 0001 |
ISCAS | 7 |
| 2024 | SHRCO: Design of an SRAM with High Reliability and Cost Optimization for Safety-Critical ApplicationsabstractThis paper proposes a novel radiation-hardened high-reliability SRAM cell, namely SHRCO, with 12 transistors for robust value storage as well as 6 transistors for parallel access operations. Using separated and error-interceptive feedback paths, the proposed cell has a complete self-recoverability from single-node upset (SNUs) at all single nodes and an excellent self-recoverability from double-node upsets (DNUs) at a part of node pairs. In addition, the proposed cell has superior access operation speed due to the inclusion of extra parallel access transistors. Simulation results show that the proposed cell has the largest number of node pairs that can self-recover from DNUs. Moreover, compared to the existing radiation-hardened SRAM cells, the proposed cell saves 28% of read time and 3% of write time on average. Yang Chang, Guangzhu Liu, Inam Ullah 0001, Gaoyang Shan, Xiaoqing Wen, Aibin Yan |
ITC-Asia | 5 |
| 2024 | PFO PUF: A Lightweight Parallel Feed Obfuscation PUF Resistant to Machine Learning AttacksabstractArbiter Physically Unclonable Functions (APUFs) are hardware security primitives that leverage manufacturing process variation to generate security keys. They can produce exponential challenge-response pairs (CRPs) with minimal hardware overhead. However, the symmetric nature of linear additive functions makes them vulnerable to modeling attacks rooted in machine learning. To address this issue, this paper introduces a novel design called Parallel Feed Obfuscation PUF (PFO PUF). In this approach, the intermediate decision signals from the lower APUF are used as a concealed challenge for the upper APUF, enhancing the overall nonlinearity of the dual-APUF. Additionally, obfuscation modules are employed to determine the weights of the intermediate decision signals from both the upper and lower APUFs, protecting the actual response. Experimental results demonstrate that the proposed PFO PUF effectively withstands four advanced machine learning attack algorithms, including Logistic Regression (LR), Support Vector Machine (SVM), Deep Feedforward Neural Network (DFNN), and Efficient CANDECOMP/PARAFAC Tensor Regression Network (ECPTRN). The prediction accuracy of these four algorithms is consistently below 66.30%. Compared with other enhanced structures based on APUF, PFO-PUF only uses 493 LUTs and has lower resource overhead. Zhengfeng Huang, Yankun Lin, Fansheng Zeng, Jingchang Bian, Huaguo Liang, Yingchun Lu, Xiaoqing Wen, Tianming Ni |
ITC-Asia | 9 |
| 2024 | Multiple-Error Interceptive Voter Designs for Safety-Critical ApplicationsabstractThis paper proposes multiple-error interceptive voter designs for safety-critical applications. The proposed baseline voters comprise two-stage error-filters, in which the first stage includes two parallel C-elements (CEs) and the second stage includes one CE to filter soft errors. The voters have high-speed versions, any of which is embedded with a high-speed path from its original input to its output to reduce delay. Simulation results demonstrate the soft error tolerance of the proposed voters. Moreover, compared with the triple-modular-redundancy (TMR) voter that can only tolerate single soft errors, the proposed 3-input baseline and 4-input high-speed voters can tolerate double soft errors and can reduce the area-power-delay product by 77.03% and 95.05%, respectively, due to the use of a few transistors and a high-speed path. The voters are extended to intercept N-1 soft errors, N being the number of inputs of each voter. Note that the voters are also extended to make so that they can tolerate hard/permanent errors in addition to soft errors. Xuehua Li, Chunjiong Zhang, Xiaoqing Wen, Zhengfeng Huang |
ITC-Asia | 5 |
| 2024 | CQCTL: A Cost-Optimized and Quadruple-Node-Upset Completely Tolerant Latch Design for Safety-Critical ApplicationsabstractWith the rapid development of semiconductor technologies, latches are becoming increasingly sensitive to multiple node upsets, such as triple node upsets and quadruple node upsets (QNUs). Therefore, they should be considered for safety-critical applications. To effectively tolerate QNUs, this paper proposes a QNU-tolerant latch design with moderate overhead. The latch mainly comprises two parallel storage cells, and three 2-input C-elements. When any four internal nodes are flipped at the same time, the output value of the latch will not be affected. Simulation results not only confirm the QNU tolerance of the proposed latch but also demonstrate that the latch can reduce by 40.93% delay, 40.73% area, 13.11% power, and 71.19% delay-area-power product (DAPP) on average compared to the existing QNU-tolerant latches. Qingyang Zhang 0001, Byeong-Hee Roh, Xiaoqing Wen |
ITC-Asia | 5 |
| 2024 | Test Point Selection for Multi-Cycle Logic BIST using Multivariate Temporal-Spatial GCNsabstractThis paper proposes a novel Test Point Insertion (TPI) strategy to enhance the testability for multi-cycle Built-In Self-Test (BIST) for logic circuits. The approach leverages Multivariate Temporal-Spatial Graph Convolutional Neural Networks (MTS-GCN) and Reinforcement Learning to identify optimal Test Points (TPs). The proposed TPI method treats the testability information of a logic circuit as time-series data and employs Multivariate Time-Series Graph Neural Networks (MTGNN) to capture the relationship between the circuit's structural (spatial information) attributes and the temporal variability of signal line testability across capture cycles. A subsequent Multi-Layer Perceptron (MLP) computes the metric for each signal line to pinpoint potential TPs based on the extracted temporal-spatial features. Experimental evaluation based on benchmark circuits confirms the efficacy of the proposed model, which is trained with Deep Q-Networks (DQN), in improving the fault detection for multi-cycle logic BIST. Senling Wang, Shaoqi Wei, Hisashi Okamoto, Tatusya Nishikawa, Hiroshi Kai, Yoshinobu Higami, Hiroyuki Yotsuyanagi, Ruijun Ma 0002, Tianming Ni, Hiroshi Takahashi, Xiaoqing Wen |
ITC-Asia | 11 |
| 2024 | SRBML: A Single-Event-Upset Recoverable and BTI-Mitigated Latch Design for Long-Term Reliability EnhancementabstractSoft-errors and aging are considered as two primary factors affecting the long-term reliability of aerospace integrated circuits (ICs). As one of the key components in aerospace ICs, latches play a pivotal role to ensure desirable circuit functionality. This paper presents a single-event-upset recovery latch, namely SRBML, with bias-temperature-instability (BTI)-mitigation. By optimizing its internal structure, the latch can recover from single-event-upsets (SEUs) and reduce the stress time of transistors in feedback loops to simultaneously mitigate the impact of BTI on the latch. Simulation results demonstrate that the soft error rate increase due to BTI is reduced by roughly 34% for SRBML after BTI-mitigation. In addition, the delay of SRBML is not affected, and the area and power increase are limited compared to BTI-unmitigated latches. Jehad Ali, Chunjiong Zhang, Xiaoqing Wen, Aibin Yan |
ITC-Asia | 5 |
| 2024 | ICLTR: A Input-split Inverters and C-elements based Low-Cost Latch with Triple-Node-Upset RecoveryabstractAs the semiconductor technology continues to advance, integrated circuits (ICs) are becoming increasingly sensitive to soft errors, e.g., double-node upsets (DNUs) and triplenode upsets (TNUs), induced by harsh radiation. In this paper, a low-cost latch design, namely ICLTR, using input-split inverters (ISIs) and C-elements to provide complete TNU recovery, is proposed. ICLTR consists of seven ISIs, seven 2-input C-elements and a clock-gated inverter, and all these elements are interlocked. Simulation results show the complete TNU recovery for ICLTR. The simulation results also show that ICLTR can save 59.5% of the transmission delay, 36.1% of the power consumption and 81.6% of the delay-area-power product (DAPP) on average when compared with the same type of TNU recovery latch designs. Zhenmin Li, Gaoyang Shan, Xiaoqing Wen |
ITC-Asia | 6 |
| 2024 | A new die-level flexible design-for-test architecture for 3D stacked ICsabstractA die-level design-for-test architecture for 3D stacked ICs is proposed. The main component of this architecture is a newly proposed configurable boundary cell, based on which flexible parallel test is achieved. Both of the number of parallel scan chains and their lengths can be configured during test. This test architecture features light-weight, high flexibility in parallel test configuration , modularity, and IEEE P1149.1 compatibility. In this work, both infrastructure and implementation aspects are illustrated. Experimental results demonstrate desired test acceleration. The acceleration ratio approximately reaches its limit, which equals the number of parallel scan chains, when the number of test vectors is over 300. Qingping Zhang, Wenfa Zhan, Xiaoqing Wen |
Integr. | 3 |
| 2024 | MURLAV: A Multiple-Node-Upset Recovery Latch and Algorithm-Based Verification MethodabstractIn advanced CMOS technologies, integrated circuits are sensitive to multiple-node-upsets (MNUs) induced in harsh radiation environments. The existing verification of the reliability of latches highly relies on electronic design automation (EDA) tools considering complex error-injection scenarios. In this paper, we propose a novel latch, namely MURLAV, protected against quadruple node-upsets (QNUs) induced in harsh radiation environments, as well as an algorithmic error-recovery verification method. The latch provides complete recovery from all QNUs with a formed redundant structure. The algorithm can simplify the verification process and demonstrate the QNU recovery for the proposed MURLAV latch. Simulation results demonstrate that the proposed latch can recover from any QNU and that it has lower area and delay overhead. Compared with existing latches of the same type, the proposed MURLAV latch achieves an overhead reduction of 34% in silicon area and 15% in delay on average at the cost of moderate power consumption. Aibin Yan, Zhongyu Gao, Zhengfeng Huang, Tianming Ni, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2024 | Introduction to the Special Issue on Design for Testability and Reliability of Security-aware HardwareabstractThe research on design for testability and reliability of security-aware hardware has been important in both academia and industry. With ever-growing globalization, commercial hardware design, manufacturing, transportation, and supply now involve many different countries, resulting in aggravated vulnerability from hardware design to manufacturing. Hardware with malicious purposes implanted from the third-party manufacturing process may control the operation of a circuit and tamper its functions, causing serious security issues. However, hardware includes not only devices and circuits but also systems. An important fact is that testability, reliability, and security technologies come from different design layers, but the impact evaluation is conducted at the system level. In other words, the testability, reliability, and security design of different layers can be carried out in a holistic manner to achieve optimization for the whole system. In addition, the testability, reliability, and security design technologies of each design layer can be collaboratively conducted to achieve better performance. The testability, reliability, and security tradeoff has garnered attention from academia and industry, particularly in the Post-Moore Era, due to the complexities and opportunities arising from new architectures and technologies. Tianming Ni, Xiaoqing Wen, Hussam Amrouch, Cheng Zhuo, Peilin Song |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | A Robust Newton Iteration Method for Mixed-Cell-Height Circuit Legalization Under Technology and Region ConstraintsabstractThe evolution of advanced technology nodes has prompted a shift toward mixed-cell-height circuit design, while the introduction of technology and fence region constraints further increases the complexity of placement. In this article, we innovatively transform the mixed-cell-height circuit legalization problem into a generalized absolute value equation (GAVE) and propose a novel and effective robust Newton (RN) iteration method to address the challenge of the legalization problem. 1 First, the window-based cell insertion technique is applied to obtain the initial cell row allocation and cell order, and the cells are allocated to the matching region based on the R-tree structure. Then, the legalization problem of cells within the region is transformed into a GAVE, and an RN iteration method is proposed to solve the GAVE. Finally, the maximum displacement cells and technology violation cells are optimized based on a greedy method. Experimental results confirm the efficiency and robustness of the proposed method compared with the state-of-the-art methods. Chencan Zhou, Yang Cao 0014, Lu-Xin Wang, Xiaoqing Wen |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | Nonvolatile Latch Designs With Node-Upset Tolerance and Recovery Using Magnetic Tunnel Junctions and CMOSabstractAs semiconductor technologies scale down, radiative-particle-induced soft errors and static power consumption are becoming major concerns for digital circuits. Magnetic-tunnel-junctions (MTJs) are widely used to address these concerns. MTJs are nonvolatile (NV) and compatible with traditional CMOS processes. In this article, we first propose a double-node-upset (DNU) tolerant and NV latch, i.e., M-TPDICE-V2, providing high reliability. In addition, we further propose an advanced latch, namely, M-8C, that is able to completely recover from single-node upsets (SNUs) and DNUs. M-8C uses a DNU recovery module and a backup and restore module based on a pair of MTJs. Furthermore, we propose a universal backup and restore module suitable for any latch providing nonvolatility. We simulate the proposed latches using the Synopsys HSPICE tool with a 45-nm CMOS process model. Simulation results confirm the superior capabilities of our proposed M-TPDICE-V2 and M-8C latches. M-TPDICE-V2 exhibits strong SNU and DNU tolerance and nonvolatility, while the M-8C latch provides complete DNU recovery capabilities. Aibin Yan, Litao Wang, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | Enhancing Defect Diagnosis and Localization in Wafer Map Testing Through Weakly Supervised LearningabstractDefect diagnosis and localization in wafer maps are crucial tasks in semiconductor manufacturing. Existing deep learning methods often require pixel-level annotations, making them impractical for large-scale deployment. In this paper, we propose a novel weakly supervised learning approach to achieving high-precision defect identification and effective localization with only image-level labels. By leveraging the information of defect types and locations, we introduce a weighted fusion of activation maps, called Class Activation Map (CAM), to highlight classspecific regions. We further enhance defect localization accuracy and completeness by employing optimized region growing operations to eliminate noise in defect regions. Moreover, we present an optimized inference method that provides meaningful visual explanations for defect recognition. Experimental results on real-world wafer map images demonstrate the effectiveness of our approach in accurately segmenting defect patterns with no pixel-level annotations. By training the model solely on wafer map image classification labels, our proposed model significantly improves defect recognition, facilitating efficient defect analysis in semiconductor manufacturing. The proposed weakly supervised learning approach offers a practical solution for defect diagnosis and localization, with the potential of widespread adoption in the semiconductor industry. Mu Nie, Wankou Yang, Senling Wang, Xiaoqing Wen, Tianming Ni |
ATS | 5 |
| 2023 | A High-Performance and P-Type FeFET-Based Non-Volatile LatchabstractNon-volatile memory has a significant future in the Internet of Things and computation-in-memory applications. Among them, non-volatile memories using emerging FeFET devices have garnered significant attention. This paper proposes a novel P-type FeFET -based non-volatile latch. This design takes advantage of the unique characteristics of a P-type FeFET device to achieve non-volatility with no additional control signals. The Cadence simulation tool Virtuoso verifies that our proposed design has correct functioning with excellent power, area, and delay performance compared to state-of-the-art designs. Aibin Yan, Zhengfeng Huang, Jie Cui 0004, Xiaoqing Wen |
ATS | 5 |
| 2023 | Advanced DICE Based Triple-Node-Upset Recovery Latch with Optimized Overhead for Space ApplicationsabstractWith the rapid advancement of CMOS technologies, integrated circuits are becoming more prone to soft errors, e.g., triple-node upsets (TNUs). In this paper, to effectively tolerate TNUs, an input-split C-element-based DICEs (IC-DICEs) based TNU-recovery latch is proposed. The latch employs three interlocked IC-DICEs to allow recovering from any TNU. Simulations demonstrate the TNU recovery of the latch, and also demonstrate that the proposed latch can reduce delay by 87.21%, area by 27.04%, and delay-area-power product (DAPP) by 87.44% on average, compared to the alternative latches. Aibin Yan, Xuehua Li, Zhongyu Gao, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen |
ATS | 6 |
| 2023 | High Performance and DNU-Recovery Spintronic Retention Latch for Hybrid MTJ/CMOS TechnologyabstractWith the advancement of CMOS technologies, circuits have become more vulnerable to soft errors, such as single-node-upsets (SNUs) and double-node-upsets (DNUs). To effectively provide nonvolatility as well as tolerance against DNUs caused by radiation, this paper proposes a nonvolatile and DNU resilient latch that mainly comprises two magnetic tunnel junction (MTJ), two inverters and eight C-elements. Since two MTJs are used and all internal nodes are interlocked, the latch can provide nonvolatility and recovery from all possible DNUs. Simulation results demonstrate the nonvolatility, DNU recovery and high performance of the proposed latch. Aibin Yan, Jie Cui 0004, Zhengfeng Huang, Xiaoqing Wen, Patrick Girard 0001 |
DATE | 6 |
| 2023 | BiSTAHL: A Built-In Self-Testable Soft-Error-Hardened Scan-CellabstractEnsuring the correct operation of modern VLSI circuits within safety-critical systems is essential since modern technology nodes are more susceptible to Early-Life Failures (ELFs) and radiation-induced Soft-Errors (SEs). Tackling both of these challenges leads to contradicting design requirements: Effective in-field ELF detection requires online-monitoring or periodic built-in self-testing with excellent cell-internal defect coverage. SE-hardened latch designs, however, are less testable because they are designed to mask cell-internal failures. We propose BiSTAHL, a new SE-hardened scan-cell design that is fully built-in self-testable for both production defects and ELFs. Stefan Holst, Ruijun Ma 0002, Xiaoqing Wen, Aibin Yan |
ETS | 3 |
| 2023 | Two Highly Reliable and High-Speed SRAM Cells for Safety-Critical Applications
Aibin Yan, Yang Chang, Jing Xiang, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2023 | A Low Area and Low Delay Latch Design with Complete Double-Node-Upset-Recovery for Aerospace Applications
Aibin Yan, Shaojie Wei, Jinjun Zhang, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2023 | Design of Low-Cost Approximate CMOS Full AddersabstractMany applications have an inherent tolerance for insignificant inaccuracies. Full adders are key arithmetic functions for many error-tolerant applications. Approximate full adders are considered an efficient technique to trade off energy relative to performance and accuracy. In this paper, we propose four approximate full adders with low overhead. The proposed and the existing approximate full adders are classified into two groups according to their error distances. Simulation results show that, compared with the existing approximate full adders, in the first group, the proposed ones can reduce power-area-delay product (PADP) by 61.83%, power by 54.15%, area by 44.67%, and delay by 22.78%on average; in the second group, the proposed ones can reduce PADP by 97.01%, power by 93.43%, area by 24.98%, and delay by 36.14% on average. Aibin Yan, Shaojie Wei, Jie Cui 0004, Zhengfeng Huang, Patrick Girard 0001, Xiaoqing Wen |
ISCAS | 7 |
| 2023 | Design of A Highly Reliable and Low-Power SRAM With Double-Node Upset Recovery for Safety-critical ApplicationsabstractFor high-speed operations, low power consumption and small silicon area, transistors are being scaled aggressively. Meanwhile, circuit reliability is facing greater challenges in advanced technologies. In this paper, a highly reliable and low-power SRAM with double-node-upset (DNU) recovery, namely HRLP16T, is proposed for safety-critical fields. HRLP16T can recover from single-node-upset (SNU) at all the sensitive nodes, and it has eight node pairs recoverable from DNUs. Simulation results demonstrate its advantages in terms of delay and power consumption over typical existing SRAM cell designs. Aibin Yan, Jing Xiang, Zhengfeng Huang, Tianming Ni, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 7 |
| 2023 | A Low Overhead and Double-Node-Upset Self-Recoverable LatchabstractWith the rapid advancement of semiconductor technologies, integrated circuits, especially storage elements (e.g., latches) have become increasingly vulnerable to soft errors. In order to effectively tolerate double-node-upsets (DNUs) caused by radiation and reduce the power and area of latches, this paper proposes a DNU self-recoverable latch with low overhead in terms of power and area. The proposed latch mainly comprises seven 2-input C-elements and two inverters to achieve DNU self-recovery. Simulation results show that the proposed latch can recover from all possible DNUs and that it can reduce delay by 45.7%, power by 29.1%, area by 65.9%, and area-power-delay-product by 87.4%, on average, compared to typical existing DNU self-recoverable latches. Aibin Yan, Tianming Ni, Jie Cui 0004, Zhengfeng Huang, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 7 |
| 2023 | Design of a Novel Latch with Quadruple-Node-Upset Recovery for Harsh Radiation HardnessabstractAs CMOS processes continue to shrink, nano-scale CMOS latches have become increasingly sensitive to multiple-node upset (MNU) errors caused by radiation. To tolerate MNU, a novel quadruple-node-upset (QNU) self-recoverable latch is proposed in this paper. The proposed latch is mainly constructed from six blocks of three-level C-elements (TLCEs) and six inverters. With the mutual feedback of the various TLCEs, the proposed latch can recover from any QNU. Furthermore, due to the clock gating methodology and a high-speed transmission path, the proposed latch has lower overhead in terms of power dissipation and transmission delay. Simulation results show that the proposed latch achieves high reliability with moderate overhead compared to typical existing latches. Aibin Yan, Shaojie Wei, Jie Cui 0004, Zhengfeng Huang, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 7 |
| 2023 | Trusted fingerprint localization for multimedia devices based on blockchain
Zhiting Liu, Xiaoqing Wen, Qinjun Wan |
Inf. Sci. | 3 |
| 2023 | LDAVPM: A Latch Design and Algorithm-Based Verification Protected Against Multiple-Node-Upsets in Harsh Radiation EnvironmentsabstractIn deep nano-scale and high-integration CMOS technologies, storage circuits have become increasingly sensitive to charge-sharing-induced multiple-node-upsets (MNUs) that include double, triple, and quadruple node-upsets. Currently, verifications for error recovery of existing latches highly rely on EDA tools with complex error-injection combinations. In this article, a latch design protected against MNUs in the harsh radiation as well as an algorithm-based verification process is proposed. Due to the constructed redundant feedback loops, the latch can completely recover from any MNU. Algorithm-based verification and simulations both demonstrate the MNU recovery of the proposed latch. Simulation results demonstrate the low area overhead of the proposed latch compared with the only one existing of the same type. Aibin Yan, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Design of True Random Number Generator Based on Multi-Ring Convergence Oscillator Using Short Pulse Enhanced RandomnessabstractThe entropy source structure with embedded XOR gates in a ring oscillator (RO) as a true random number generator (TRNG) can improve the speed of accumulating jitter in the oscillator. However, the XOR gate has a certain response time to the input change, and when the input changes too fast, the XOR gate will output short pulses. In this paper, we propose a TRNG design based on a multi-ring convergence oscillator (MRCO) making use of the characteristics of short pulses. We study the output of the XOR gate when facing different inputs. By modeling the time of a fibonacci ring oscillator (FIRO) as an example, we find that the loss of short pulses in an inverter chain is the reason for making the FIRO enter into periodic oscillation. This phenomenon suppresses the accumulation of jitter and occurs periodically in existing structures. Our proposed structure uses independent sub-rings to accumulate jitter, allowing the main-ring to quickly generate short pulses to provide analog randomness. The proposed TRNG design is implemented in Xilinx Virtex-6 FPGA. The experimental results show that it has the highest ratio of throughput rate to hardware resources. The generated random sequence pass both NIST SP800-22 test and NIST SP800-90B test. Tianming Ni, Qingsong Peng, Jingchang Bian, Zhengfeng Huang, Aibin Yan, Senling Wang, Xiaoqing Wen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | An AI-Driven VM Threat Prediction Model for Multi-Risks Analysis-Based Cloud CybersecurityabstractCloud virtualization technology, ingrained with physical resource sharing, prompts cybersecurity threats on users’ virtual machines (VMs) due to the presence of inevitable vulnerabilities on the offsite servers. Contrary to the existing works which concentrated on reducing resource sharing and encryption/decryption of data before transfer for improving cybersecurity which raises computational cost overhead, the proposed model operates diversely for efficiently serving the same purpose. This article proposes a novel multiple risks analysis-based VM threat prediction model (MR-TPM) to secure computational data and minimize adversary breaches by proactively estimating the VMs threats. It considers multiple cybersecurity risk factors associated with the configuration and management of VMs, along with analysis of users’ behavior. All these threat factors are quantified for the generation of respective risk score values and fed as input into a machine learning-based classifier to estimate the probability of threat for each VM. The performance of MR-TPM is evaluated using benchmark Google Cluster and OpenNebula VM threat traces. The experimental results demonstrate that the proposed model efficiently computes the cybersecurity risks and learns the VM threat patterns from historical and live data samples. The deployment of MR-TPM with existing VM allocation policies reduces cybersecurity threats up to 88.9%. Deepika Saxena, Ishu Gupta, Ashutosh Kumar Singh 0001, Xiaoqing Wen |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | RMC_NoC: A Reliable On-Chip Network Architecture With Reconfigurable Multifunctional ChannelabstractAs chip fabrication has advanced to the nano level, the increased link density has heightened the risk of failures. The potential performance drawbacks resulting from these link failures have become a critical challenge in the design of reliable network-on-chip (NoC) systems. Fault-tolerant routing algorithms have proven to be effective strategies for handling this issue by diverting packets away from failed links to prevent congestion. However, these algorithms often result in excessive packet diversion, especially in the presence of a higher failure rate, which can significantly constrain the network’s behavior. This article introduces a novel NoC design with reconfigurable multifunctional channels (RMC_NoC). This design dynamically adapts the channel functions in response to network conditions to ensure that packets from failed links follow their original paths. In addition, it presents a channel buffer bubble flow control mechanism that can resolve congestion by redistributing congested traffic within the channel buffer. The evaluation results demonstrate that our approach ensures superior network communication even in the presence of permanent link failures, with minimal area overhead and power consumption. Moreover, our system exhibits lower latency and higher throughput compared to state-of-the-art fault-tolerant methods across various link failure rates. Notably, even at a severe failure rate of 30%, RMC_NoC exhibits only a 16.3% increase in latency compared to an ideal failure-free environment (Baseline) while still maintaining system communication capabilities to a considerable extent. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | Energy-Efficient Multiple Network-on-Chip Architecture With Bandwidth ExpansionabstractAs technology feature sizes diminish to the nanometer regime, the leakage power crisis has become a major challenge in network-on-chip (NoC) design. Power gating (PG) is used to mitigate growing leakage power as an effective static power-saving technique. Applying PG in a multiple NoC (Multi-NoC) rather than a traditional NoC is a promising solution. However, limited by the channel width of the subnets, the increase in packet length will bring a severe serialization issue and performance loss. Previous Multi-NoC schemes have to wake up more subnets to minimize the performance loss, which also sacrifices their energy efficiency. In this article, we introduce an architecture, namely, BandExp, which allows subnets to expand their bandwidth by utilizing the idle physical links of other subnets. More bandwidth helps subnets mitigate the serialization issue and reduce the performance loss. Meanwhile, other subnets gain longer sleep cycles and thus save more energy. Evaluation results indicate that compared to the state-of-the-art Catnap, the proposed architecture reduces the average packet latency and execution time of different benchmarks by 19.3% and 3.2%, respectively. Also, the net static energy of the network is reduced by 23.2% on average, while the incurred area overhead is only 1.3%. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | A Radiation-Hardened Non-Volatile Magnetic Latch with High Reliability and Persistent StorageabstractWith technology scaling down, the vulnerability of circuits to radiation and the increase of static power have become severe concerns. Spintronic devices such as magnetic tunnel junction (MTJ) have been developed to cope with many concerns, among which reliability concerns [1]. Spintronic devices have attractive properties, such as non-volatility and compatibility with conventional CMOS fabrication process. Based on an advanced triple-path dual-interlocked-storage-cell (TPDICE) and MTJs, this paper proposes a radiation-hardened non-volatile magnetic latch, namely M-TPDICE, that can completely tolerate single-node upsets (SNUs) and double-node upsets (DNUs). Simulations of the proposed latch with the HSPICE tool with a 45 nm CMOS technology model have demonstrated the effectiveness of the proposed latch. Aibin Yan, Zhengfeng Huang, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
ATS | 7 |
| 2022 | SCLCRL: Shuttling C-elements based Low-Cost and Robust Latch Design Protected against Triple Node Upsets in Harsh Radiation EnvironmentsabstractAs the CMOS technology is continuously scaling down, nano-scale integrated circuits are becoming susceptible to harsh-radiation induced soft errors, such as double-node upsets (DNUs) and triple-node upsets (TNUs). This paper presents a shuttle C-elements based low-cost and robust latch (namely SCLCRL) that can recover from any TNU in harsh radiation environments. The latch comprises seven primary storage nodes and seven secondary storage nodes. Each pair of primary nodes feeds a secondary node through one C-element (CE) and each pair of secondary nodes feeds a primary node through another CE, forming redundant feedback loops to robustly retain values. Simulation results validate all key TNUs' recoverability features of the proposed latch. Simulation results also demonstrate that the proposed SCLCRL latch can approximately save 29% silicon area and 47% D-Q delay on average at the cost of moderate power, compared with the state-of-the-art TNU-recoverable reference latches of the same-type. Aibin Yan, Shiwei Huang, Zijie Zhai, Xiangyu Cheng, Jie Cui 0004, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
DATE | 8 |
| 2022 | Sextuple Cross-Coupled-DICE Based Double-Node-Upset Recoverable and Low-Delay Flip-Flop for Aerospace ApplicationsabstractThis paper proposes a novel sextuple cross-coupled dual-interlocked-storage-cell (DICE) based double-node-upset (DNU) recoverable and low-delay flip-flop (FF), namely SCDRL-FF, for aerospace applications. The SCDRL-FF mainly consists of sextuple cross-coupled DICEs controlled by clock-gating. The use of clock-gating based DICEs significantly reduces the CLK-Q transmission delay of the SCDRL-FF. Through the redundant and interlocked clock-gating based DICEs, the SCDRL-FF can provide complete DNU recoverability. Simulation results demonstrate the DNU recoverability of the SCDRL-FF and a 65% delay reduction on average compared with the state-of-the-art hardened FFs. The low delay overhead makes the proposed SCDRL-FF effectively applicable to high-performance applications and the DNU recoverability makes the proposed SCDRL-FF also suitable for aerospace applications. Aibin Yan, Shukai Song, Zijie Zhai, Jie Cui 0004, Zhengfeng Huang, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2022 | Two 0.8 V, Highly Reliable RHBD 10T and 12T SRAM Cells for Aerospace ApplicationsabstractAggressive scaling of CMOS technologies requires to pay attention to the reliability issues of circuits. This paper presents two highly reliable RHBD 10T and 12T SRAM cells, which can protect against single-node upsets (SNUs) and double-node upsets (DNUs). The 10T cell mainly consists of two cross-coupled input-split inverters and the cell can robustly keep stored values through a feedback mechanism among its internal nodes. It also has a low cost in terms of area and power consumption, since it uses only a few transistors. Based on the 10T cell, a 12T cell is proposed that uses four parallel access transistors. The 12T cell has a reduced read/write access time with the same soft error tolerance when compared to the 10T cell. Simulation results demonstrate that the proposed cells can recover from SNUs and a part of DNUs. Moreover, compared with the state-of-the-art hardened SRAM cells, the proposed 10T cell can save 28.59% write access time, 55.83% read access time, and 4.46% power dissipation at the cost of 4.04% silicon area on average. Aibin Yan, Zhihui He, Jing Xiang, Jie Cui 0004, Zhengfeng Huang, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2022 | A Highly Robust, Low Delay and DNU-Recovery Latch Design for Nanoscale CMOS TechnologyabstractWith the advancement of semiconductor technologies, nano-scale CMOS circuits have become more vulnerable to soft errors, such as single-node-upsets (SNUs) and double-node-upsets (DNUs). In order to effectively tolerate DNUs caused by radiation and reduce the delay and area consumption of latches, this paper proposes a DNU resilient latch in the nanoscale CMOS technology. The latch mainly comprises four input-split inverters and four 2-input C-elements. Since all internal nodes are interlocked, the latch can recover from all possible DNUs. Simulation results show that, compared with the state-of-the-art DNU self-recovery latch designs, the proposed latch can save 64.51% transmission delay and 56.88% delay-area-power-product (DAPP) on average, respectively. Aibin Yan, Shaojie Wei, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2022 | Effective Switching Probability Calculation to Locate Hotspots in Logic CircuitsabstractHigh power consumption in LSI testing may cause excessive IR-drop. When IR-drop becomes excessive, it causes excessive delay, resulting in test malfunction (over-testing). Excessive IR-drop does not occur in the entire area of a circuit, but in certain areas where a large number of switching activities occur (such areas are called hotspots in this work). In order to avoid test malfunction, it is important to develop a method to reduce or control IR-drop in the hotspots. Locating hotspots is a necessary technique to reduce or control IR-drop effectively and efficiently. In this work, we propose a method to locate hotspots in a logic circuit by switching probability calculation. Experimental results for IWLS2005 OpenCores circuits demonstrate the proposed method can support to locate hotspots. Taiki Utsunomiya, Ryu Hoshino, Kohei Miyase, Shyue-Kung Lu, Xiaoqing Wen, Seiji Kajihara |
ITC-Asia | 5 |
| 2022 | Cost-Optimized and Robust Latch Hardened against Quadruple Node Upsets for Nanoscale CMOSabstractWith the aggressive reduction of CMOS transistor feature sizes, the soft error rate of nano-scale integrated circuits increases exponentially. In this paper, we propose a novel cost-optimized and robust latch, namely CRLHQ, hardened against quadruple-node-upsets (QNUs) for nanoscale CMOS technologies. The latch mainly comprises a 5×5 matrix based on interlocked source-drain cross-coupled inverters to robustly store logic values. Owing to the redundant constructed feedback loops, the latch can recover from all possible QNUs. Simulation results demonstrate all key QNUs' recovery of the proposed CRLHQ latch. Simulation results also show that the proposed latch can approximately reduce the D-Q delay by 44.3%, the silicon area by 7.3% and the delay-area-power product (DAPP) by 14.2%, compared with the state-of-the-art same-type reference latches that can recover from any QNU. Aibin Yan, Shukai Song, Jixiang Zhang 0007, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
ITC-Asia | 7 |
| 2022 | A Highly Reliable and Low Power RHBD Flip-Flop Cell for Aerospace ApplicationsabstractIn space, the impact of radiative particles, such as neutrons and heavy ions, can change the node states of a flip-flop, thus resulting in loss of data. In this paper, a Highly reliable and Low power Radiation-hardened-by-design (RHBD) Flip-Flop cell, namely HLRFF, completely hardened against double-node-upsets (DNUs), is proposed for aerospace applications. The HLRFF is a master-slave structure. The master latch is mainly constructed from two 2-input C-elements (CEs) and one 2-input clock-gating based CE, while the slave latch has an additional keeper at the output stage. The verification results demonstrate that the proposed HLRFF is completely DNU-tolerant. Furthermore, compared to the state-of-the-art radiation-hardened FF cells, the proposed HLRFF can reduce power consumption by approximately 69%. However, only the proposed HLRFF is not only completely DNU-tolerant but also insensitive to high-impedance-state. Aibin Yan, Kuikui Qian, Jie Cui 0004, Ningning Cui, Zhengfeng Huang, Xiaoqing Wen, Patrick Girard 0001 |
VTS | 6 |
| 2022 | A double-node-upset completely tolerant CMOS latch design with extremely low cost for high-performance applications
Aibin Yan, Kuikui Qian, Tai Song, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen |
Integr. | 7 |
| 2022 | GoodFloorplan: Graph Convolutional Network and Reinforcement Learning-Based FloorplanningabstractElectronic design automation (EDA) comprises a series of computationally difficult optimization problems that require substantial specialized knowledge as well as a considerable amount of trial-and-error efforts. However, open challenges, including long simulation runtime and lack of generalization, continue to restrict the applications of the existing EDA tools. Recently, learning-based algorithms, especially reinforcement learning (RL), have been successfully applied to handle various combinatorial optimization problems by automatically acquiring knowledge from the past experience. In this article, we formulate the floorplanning problem, the first stage of the physical design flow, as a Markov decision process (MDP). An end-to-end learning-based floorplanning framework GoodFloorplan is proposed to explore the design space, which combines graph convolutional network (GCN) and RL. Experimental results demonstrate that compared with state-of-the-art heuristic-based floorplanners, the proposed GoodFloorplan can provide better area and wirelength. Qi Xu 0004, Hao Geng, Song Chen 0001, Bo Yuan 0006, Cheng Zhuo, Yi Kang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Fortune: A New Fault-Tolerance TSV Configuration in Router-Based Redundancy StructureabstractIn three-dimensional integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, yield is one of the key obstacles to adopt the TSV-based 3D-ICs technology in the industry. Various fault-tolerance structures using redundant TSVs to repair faulty functional TSVs have been proposed in the literature for yield and reliability enhancement. However, the TSV repair paths under delay constraint cannot always be generated due to the lack of appropriate repair algorithms. In this article, we propose an effective TSV repair strategy for the router-based TSV redundancy architecture, taking into account the delay overhead. First, we prove that the router-based fault-tolerance structure configuration (RFSC) with the delay constraint is equivalent to the length-bounded multicommodity flow (LBMCF) problem. Then, an integer linear programming (ILP) formulation with acceptable scalability is presented to solve the LBMCF problem. The experimental results demonstrate that, compared with state-of-the-art fault-tolerance designs, the proposed ILP model can provide higher yield and lower delay overhead. Qi Xu 0004, Hao Geng, Tianming Ni, Song Chen 0001, Bei Yu 0001, Yi Kang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Cellular Structure-Based Fault-Tolerance TSV Configuration in 3D-ICabstractIn 3-D integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, the yield is one of the key obstacles to adopt the TSV-based 3D-ICs technology in industry. Various fault-tolerance structures using redundant TSVs to repair faulty functional TSVs have been proposed in literature for yield and reliability enhancement. But the TSV repair paths under delay constraint cannot always be generated due to the lack of appropriate repair algorithms. In this article, we propose an effective TSV repair strategy for the cellular TSV redundancy architecture, with taking account of the delay overhead. First, we prove that the cellular structure-based fault-tolerance TSV configuration with the delay constraint (CSFTC) is equivalent to the length-bounded multicommodity flow (LBMCF) problem. Next, an integer linear programming formulation is presented to solve the LBMCF problem. Finally, to speed-up the fault-tolerance structure configuration process, an efficient Lagrangian relaxation-based heuristic method is further proposed. Experimental results demonstrate that, compared with the state-of-the-art fault-tolerance structures, the proposed method can provide high yield and low delay overhead. Qi Xu 0004, Song Chen 0001, Yi Kang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | GPU-Accelerated Timing Simulation of Systolic-Array-Based AI AcceleratorsabstractSystolic arrays are currently used in autonomous systems such as self-driving cars to accelerate the enormous amount of matrix operations necessary for DNN inference. The reliability of such accelerators are of utmost importance since any loss in DNN accuracy due to erroneous calculations can have dire consequences. We propose a novel method to measure accuracy losses caused by arbitrary timing faults in systolic arrays. Our GPU-based simulation system enables for the first time a complete and accurate timing simulation of all inference-related matrix operations on large systolic arrays. A single consumer-grade GPU can simulate a LeNet-5 at a throughput of about 13s per inference. Furthermore, our simulation approach readily scales to larger DNNs and multiple GPUs. Stefan Holst, Lim Bumun, Xiaoqing Wen |
ATS | 3 |
| 2021 | Reliability-Driven Neuromorphic Computing Systems DesignabstractIn recent years, memristive crossbar-based neuromorphic computing systems (NCS) have provided a promising solution to the acceleration of neural networks. However, stuck-at faults (SAFs) in the memristor devices significantly degrade the computing accuracy of NCS. Besides, memristors suffer from process variations, causing the deviation of actual programming resistance from its target resistance. In this paper, we propose a novel reliability-driven design framework for a memristive crossbar-based NCS in combination with general and chip-specific design optimizations. First, we design a general reliability-aware training scheme to enhance the robustness of NCS to SAFs and device variations; a dropout-inspired approach is developed to alleviate the impact of SAFs; a new weighted error function, including cross-entropy error (CEE), the l2-norm of weights, and the sum of squares of first-order derivatives of CEE with respect to weights, is proposed to obtain a smooth error curve, where the effects of variations are suppressed. Second, given the neural network model generated by the reliability-aware training scheme, we exploit chip-specific mapping and retraining to further reduce the computation accuracy loss incurred by SAFs. Experimental results clearly demonstrate that the proposed method can boost the computation accuracy of NCS and improve the NCS robustness. Qi Xu 0004, Junpeng Wang 0002, Hao Geng, Song Chen 0001, Xiaoqing Wen |
DATE | 5 |
| 2021 | A 4NU-Recoverable and HIS-Insensitive Latch Design for Highly Robust Computing in Harsh Radiation EnvironmentsabstractThis paper proposes a 4-node-upset (4NU) recoverable and high-impedance-state (HIS) insensitive latch design, namely QRHIL, for highly robust computing in harsh radiation environments. The latch mainly comprises a 5×5 looped C-element matrix to store values and provide complete 4NU recovery. Owing to the multiple-level error-interception of the 5×5 C-element matrix, the latch can recover from all possible 4NUs; thus, the latch is insensitive to HIS. Simulation results demonstrate the 4NU-recovery of the proposed latch. The results also show that the latch can approximately save 46% D-Q delay and 46% CLK-Q delay owing to the use of a high-speed D-Q path and clock-gating, compared with the state-of-the-art 3NU-recoverable latch (TNURL) that is not 4NU-recoverable. Aibin Yan, Aoran Cao, Zhengzheng Fan, Zhelong Xu, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 7 |
| 2021 | TPDICE and Sim Based 4-Node-Upset Completely Hardened Latch Design for Highly Robust Computing in Harsh RadiationabstractTechnology scaling and charge-sharing make nano- scale CMOS latches become severely vulnerable to multiple-node upsets (MNUs). This paper proposes a triple-path dual- interlocked-storage-cell (TPDICE) and soft-error interceptive module (SIM) based 4-Node-Upset (4NU) completely hardened latch, namely 4NUHL latch, that can completely tolerate soft errors, such as 4NUs. The latch mainly consists of 2 TPDICEs and a 3-level SIM which comprises six 2-input C-elements. Owing to the single-node-upset self-recoverability and multiple storage nodes of TPDICEs and the soft-error interception capability of the SIM, the latch can provide complete 4NU tolerance. Simulation results demonstrate that the proposed 4NUHL latch is completely 4NU hardened. Furthermore, we use a high-speed path, clock-gating, and a few transistors to reduce overhead of the proposed latch. We compared the proposed latch with state-of- the-art hardened latches in terms of reliability and overhead to demonstrate the advantages of the proposed latch. Aibin Yan, Chuanbo Shan, Haoran Cai, Zhanjun Wei, Zhengfeng Huang, Xiaoqing Wen |
ISCAS | 8 |
| 2021 | Parallel DICE Cells and Dual-Level CEs based 3-Node-Upset Tolerant Latch Design for Highly Robust ComputingabstractWith the rapid advancement of design and manufacturing technologies of nano-scale CMOS circuits, latches are becoming increasingly sensitive to multiple-node-upsets caused by harsh radiation effects. In this paper, a Parallel Dual-interlocked-storage-cells (DICEs) and Dual-level C-elements (CEs) based 3-node-upset (3NU)-Tolerant Latch, namely PDDCTL, design for highly robust computing, is proposed. The latch comprises five transmission gates, two DICEs and three CEs. Due to the use of two single-node-upset self-recoverable DICEs and three error-interceptive CEs, the latch can provide complete 3NU-tolerance with low cost. Simulation results not only confirm the 3NU-tolerance of the proposed latch but also demonstrate that the delay-power-area product of the PDDCTL latch is reduced by 68.82% on average compared with the state-of-the-art 3NU hardened latch designs. Aibin Yan, Zijie Zhai, Lele Wang 0011, Jixiang Zhang 0007, Ningning Cui, Tianming Ni, Xiaoqing Wen |
ITC-Asia | 7 |
| 2021 | Design of Radiation Hardened Latch and Flip-Flop with Cost-Effectiveness for Low-Orbit Aerospace Applications
Aibin Yan, Aoran Cao, Zhelong Xu, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
J. Electron. Test. | 7 |
| 2021 | A Cost-Effective TSV Repair Architecture for Clustered Faults in 3-D ICabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the through-silicon vias (TSVs) defect tend to be clustered, reducing the yield of 3-D integrated circuit significantly. To tackle this fault clustering problem, the existing TSV repair methods adopt the TSV redundancy idea, which brings a major cost to 3-D integration. In this brief, a honeycomb-TDMA TSV design is proposed to mitigate the impact of multiple clustered faults without the need of redundant TSVs (RTSVs), thereby decreasing the area overhead and enhances the yield. The yield of the honeycomb-TDMA architecture can achieve 91.38%-99.67% for different benchmark circuits from IWLS 2005, which has the highest yield. Furthermore, our design achieves total additional hardware (timing delay overhead) reduction by 83.70%-86.85% (46.01%-55.96%), 66.89%-73.25% (29.41%-38.49%), 68.02%-74.20% (41.40%-52.20%), 60.60%-68.18% (18.09%-33.18%), and 75.86%-80.52% (3.05%-20.91%), respectively, compared with router-based, ring-based, group-based, cellular-based, and honeycomb-based methods. Therefore, the proposed architecture is the best choice in terms of yields, hardware overhead, and timing delay. Tianming Ni, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | A Sextuple Cross-Coupled SRAM Cell Protected against Double-Node UpsetsabstractIn this paper, we propose a sextuple cross-coupled SRAM cell, namely SCCS18T, protected against double-node upsets. Since the proposed SCCS18T cell forms a large feedback loop for value retention and error interception, the cell can provide self-recoverability from any single-node upsets (SNUs) and partial double-node upsets (DNUs). Moreover, the proposed cell has optimized operation speed due to the use of six access transistors. Simulation results show that the SCCS18T cell can save approximately 65% read access time at the cost of 49% power dissipation and 50% silicon area on average, compared with typical hardened SRAM cells. Aibin Yan, Jun Zhou 0016, Jie Cui 0004, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
ATS | 6 |
| 2020 | HITTSFL: Design of a Cost-Effective HIS-Insensitive TNU-Tolerant and SET-Filterable Latch for Safety-Critical ApplicationsabstractThis paper proposes a cost-effective, high-impedance-state (HIS)-insensitive, triple-node-upset (TNU)-tolerant and single-event-transient (SET)-filterable latch, namely HITTSFL, to ensure high reliability with low-cost. The latch mainly comprises an output-level SET-filterable Schmitt-trigger and three inverters that make the values stored in three parallel single-node-upset (SNU)-recoverable dual-interlocked-storage-cells (DICEs) converge at a common node to tolerate any possible TNU. The latch does not use C-elements to be insensitive to the HIS. Simulation results demonstrate the TNU-tolerability and SET-filterability of the proposed HITTSFL latch. Moreover, due to the use of clock-gating technologies and fewer transistors, the proposed latch can reduce delay, power, and area by 76.65%, 6.16%, and 28.55%, respectively, compared with the state-of-the-art TNU hardened latch (TNUHL) that cannot filter SETs. Aibin Yan, Xiangfeng Feng, Jie Cui 0004, Zuobin Ying, Patrick Girard 0001, Xiaoqing Wen |
DAC | 8 |
| 2020 | Dual-Interlocked-Storage-Cell-Based Double-Node-Upset Self-Recoverable Flip-Flop Design for Safety-Critical ApplicationsabstractThis paper presents a novel dual-interlocked storage-cell (DICE)-based double-node-upset (DNU) self-recoverable, namely DURI-FF, in the nano-scale CMOS technology. The master latch of the DURI-FF cell consists of three transmission gates (TGs) and three interlocked DICEs with three common nodes. The common nodes are connected to TGs for value initialization. The slave latch of the DURI-FF cell comprises six TGs, six inverters and three interlocked DICEs. The outputs of the inverters respectively feed the internal nodes of the slave latch. The interlocked DICEs make the master latch and the slave latch DNU self-recoverable. Simulation results validate the DNU self-recoverability of the proposed DURI-FF cell. Moreover, compared with the state-of-the-art hardened flip-flop cells, the proposed DURI-FF cell achieves roughly 43% delay reduction at the cost of moderate silicon area and power dissipation. Aibin Yan, Zhelong Xu, Jie Cui 0004, Zuobin Ying, Zhengfeng Huang, Huaguo Liang, Patrick Girard 0001, Xiaoqing Wen |
ISCAS | 8 |
| 2020 | Design of a Highly Reliable SRAM Cell with Advanced Self-Recoverability from Soft ErrorsabstractIn this paper, a highly reliable SRAM cell, namely SESRS cell, is proposed. Since the cell has a special feedback mechanism among its internal nodes and has more access transistors compared to a standard SRAM cell, the SESRS cell provides the following advantages: (1) it can self-recover from single node upsets (SNUs) and double-node upsets (DNUs); (2) it can reduce power consumption by 49.78% and silicon area by 7.92%, compared with the only existing SRAM cell which can self-recover from all possible DNUs. Simulation results validate the robustness of the proposed SESRS cell. Moreover, compared with the state-of-the-art hardened SRAM cells, the proposed SESRS cell can reduce read access time by 61.93% on average. Zhengda Dou, Aibin Yan, Jun Zhou 0016, Yuanjie Hu, Tianming Ni, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 9 |
| 2020 | Logic Fault Diagnosis of Hidden Delay DefectsabstractHidden delay defects (HDDs) are small delay defects that pass all at-speed tests at nominal capture time. They are an important indicator of latent defects that lead to early-life failures and aging problems that are serious especially in autonomous and medical applications. An effective way to screen out HDDs is to use Faster-than-At-Speed Testing (FAST) to observe outputs of sensitized non-critical paths which are expected to be stable earlier than nominal capture time. To improve the reliability of current and future designs, it is important to learn about the population of HDDs using logic diagnosis. We present the very first logic fault diagnosis technique that is able to identify HDDs by analyzing fail logs produced by FAST. Even with aggressive FAST testing, HDDs generate only very few failing test response bits. To overcome this severe challenge, we propose new backtracing and response matching methods that yield high diagnostic success rates even with very limited amount of failure data. The performance and scalability of our HDD diagnosis method is validated using fault injection campaigns with large benchmark circuits. Stefan Holst, Matthias Kampmann, Alexander Sprenger, Jan Dennis Reimer, Sybille Hellebrand, Hans-Joachim Wunderlich, Xiaoqing Wen |
ITC | 7 |
| 2020 | Information Assurance Through Redundant Design: A Novel TNU Error-Resilient Latch for Harsh Radiation EnvironmentabstractIn nano-scale CMOS technologies, storage cells such as latches are becoming increasingly sensitive to triple-node-upset (TNU) errors caused by harsh radiation effects. In the context of information assurance through redundant design, this article proposes a novel low-cost and TNU on-line self-recoverable latch design which is robust against harsh radiation effects. The latch mainly consists of a series of mutually interlocked 3-input Muller C-elements (CEs) that forms a circular structure. The output of any CE in the latch respectively feeds back to one input of some specified downstream CEs, making the latch completely self-recoverable from any possible TNU, i.e., the latch is completely TNU-resilient. Simulation results demonstrate the complete TNU-resiliency of the proposed latch. In addition, due to the use of fewer transistors and a high-speed path, the proposed latch reduces the delay-power-area product by approximately 91 percent compared with the state-of-the-art TNU hardened latch (TNUHL), which cannot provide a complete TNU-resiliency. Aibin Yan, Yuanjie Hu, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Computers | 8 |
| 2020 | LCHR-TSV: Novel Low Cost and Highly Repairable Honeycomb-Based TSV Redundancy Architecture for Clustered FaultsabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the quality of through-silicon vias (TSVs) varies during the fabrication and bonding process. If one TSV exhibits a defect during its manufacturing process, the probability of multiple defects occurring in the TSVs neighboring the faulty TSV increases, i.e., the TSV defects tend to be clustered, which significantly reduces the yield of 3-D integrated circuit. To resolve the clustered TSV faults, router-based, ring-based, group-based, and cellular-based redundant TSV (RTSV) architectures were proposed. However, the repair rate is low and the hardware overhead as well as delay overhead is high. In this article, we propose a honeycomb-based RTSV architecture to utilize the area and delay more efficiently as well as to maintain high yield. The simulation results show that the proposed architecture has a 99.84% repair rate for uniform faults and an 81.42% repair rate for highly clustered faults. The proposed design achieves a 51.66% reduction of hardware overhead compared with the router-based design and a 20.69%, 46.93%, 34.17%, and 11.15% reduction of total delay compared with ring-based, router-based, group-based, and cellular-based methods, respectively. Tianming Ni, Huaguo Liang, Aibin Yan, Zhengfeng Huang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2019 | Novel Radiation Hardened Latch Design with Cost-Effectiveness for Safety-Critical Terrestrial ApplicationsabstractTo meet the requirements of both cost-effectiveness and high reliability for safety-critical terrestrial applications, this paper proposes a novel radiation hardened latch design, namely HLCRT. The HLCRT latch mainly consists of a single-node-upset self-recoverable cell, a 3-input C-element, and an inverter. If any two inputs of the C-element suffer from a double-node-upset (DNU), or if one node inside the cell together with another node outside the cell suffer from a DNU, the latch still has correct values on its output node, i.e., the latch is effectively DNU hardened. Simulation results demonstrate the DNU tolerance of the proposed latch. Moreover, due to the use of fewer transistors, clock gating technologies, and a high-speed path, the proposed latch saves about 444.80% delay, 150.50% power, 72.66% area, and 2029.63% delay-power-area product on average, compared with state-of-the-art DNU hardened latch designs. Aibin Yan, Zuobin Ying, Patrick Girard 0001, Xiaoqing Wen |
ATS | 8 |
| 2019 | Design of a Sextuple Cross-Coupled SRAM Cell with Optimized Access Operations for Highly Reliable Terrestrial ApplicationsabstractThe Aggressive technology scaling makes modern advanced SRAMs more and more sensitive to soft errors that include single-node upsets (SNUs) and double-node upsets (DNUs). This paper presents a novel Sextuple Cross-Coupled SRAM cell, namely SCCS cell, which can tolerate both SNUs and DNUs. The cell mainly consists of six cross-coupled input-split inverters, constructing a large error-interceptive feedback loop to robustly retain stored values. Since the cell has many redundant storage nodes, the cell achieves the following robustness: (1) the cell can self-recover from all possible SNU; (2) the cell can self-recover from partial DNUs; (3) the cell can avoid the occurrence of other DNUs due to node-separation. Simulation results validate the excellent robustness of the proposed cell. Moreover, compared with the state-of-the-art typical existing hardened cells, the proposed cell achieves an approximate 61% read access time as well as 12% write access time reduction at the costs of 47% power dissipation as well as 44% silicon area on average. Aibin Yan, Jun Zhou 0016, Yuanjie Hu, Zuobin Ying, Xiaoqing Wen, Patrick Girard 0001 |
ATS | 7 |
| 2019 | Single-Event Double-Upset Self-Recoverable and Single-Event Transient Pulse Filterable Latch Design for Low Power ApplicationsabstractThis paper presents a single-event double-upset (SEDU) self-recoverable and single-event transient (SET) pulse filterable latch design for low power applications in 22nm CMOS technology. The latch mainly consists of eight mutually feeding back C-elements and a Schmitt trigger. Simulation results have demonstrated both the SEDU self-recoverability and SET pulse filterability for the latch using redundant silicon area. Using clock gating technology, the latch saves about 54.85% power dissipation on average compared with the up-to-date SEDU self-recoverable latch designs which are not SET pulse filterable at all. Aibin Yan, Yuanjie Hu, Xiaoqing Wen |
DATE | 4 |
| 2019 | STAHL: A Novel Scan-Test-Aware Hardened Latch DesignabstractAs modern technology nodes become more susceptible to soft errors, many radiation hardened latch designs have been proposed. However, redundant circuitry used to tolerate soft errors in such hardened latches also reduces the test coverage of cell-internal manufacturing defects. To avoid potential test escapes that lead to soft error vulnerability and reliability issues, this paper proposes a novel Scan-Test-Aware Hardened Latch (STAHL). Simulation results show that STAHL has superior defect coverage compared to previous hardened latches while maintaining full radiation hardening in function mode. Ruijun Ma 0002, Stefan Holst, Xiaoqing Wen, Aibin Yan |
ETS | 3 |
| 2019 | A Static Method for Analyzing Hotspot Distribution on the LSIabstractPerformance degradation caused by high IR-drop in normal functional mode of LSI can be avoided by improving the power supply network in the layout design phase. However, while IR-drop increases much more in test mode than in normal functional mode, excessive IR-drop in test mode is not appropriately considered in the layout design phase. Excessive IR-drop in test mode causes over-testing, which wrongly determines a fault free LSI in normal functional mode to be faulty. In this work, we propose a method for analyzing high IR-drop areas (hotspot distribution), which is necessary to effectively and efficiently reduce excessive IR-drop. Kohei Miyase, Yudai Kawano, Shyue-Kung Lu, Xiaoqing Wen, Seiji Kajihara |
ITC-Asia | 4 |
| 2019 | A Novel Triple-Node-Upset-Tolerant CMOS Latch Design using Single-Node-Upset-Resilient CellsabstractNano-scale CMOS circuits are vulnerable to single-event triple-node-upsets (SETUs). This paper proposes the design of a novel CMOS latch to tolerate any SETU using single-node-upset-resilient cells converged at a highly reliable node. The latch makes use of three single-node-upset-resilient cells, each of which mainly consists of triple mutually feeding back 2-input C-elements. These cells have a common converged output node feeding back to the output of the latch, making the latch capable of tolerating any SETU. Simulation results not only confirm the SETU tolerance capability but also show a significant area-power-delay-product reduction of 96.81% for the proposed latch compared with the only existing SETU hardened latch. Zhiyuan Song, Aibin Yan, Jie Cui 0004, Xiaoqing Wen, Chaoping Lai, Zhengfeng Huang, Huaguo Liang |
ITC-Asia | 6 |
| 2019 | Variation-Aware Small Delay Fault Diagnosis on Compressed Test ResponsesabstractWith today's tight timing margins, increasing manufacturing variations, and new defect behaviors in FinFETs, effective yield learning requires detailed information on the population of small delay defects in fabricated chips. Small delay fault diagnosis for yield learning faces two main challenges: (1) production test responses are usually highly compressed reducing the amount of available failure data, and (2) failure signatures not only depend on the actual defect but also on omnipresent and unknown delay variations. This work presents the very first diagnosis algorithm specifically designed to diagnose timing issues on compressed test responses and under process variations. An innovative combination of variation-invariant structural analysis, GPU-accelerated time-simulation, and variation-tolerant syndrome matching for compressed test responses allows the proposed algorithm to cope with both challenges. Experiments on large benchmark circuits clearly demonstrate the scalability and superior accuracy of the new diagnosis approach. Stefan Holst, Eric Schneider, Michael A. Kochte, Xiaoqing Wen, Hans-Joachim Wunderlich |
ITC | 4 |
| 2019 | Targeted Partial-Shift For Mitigating Shift Switching Activity Hot-Spots During Scan TestabstractShifting scan chains during testing causes high switching activity in the combinational logic. Excessive shift switching activity can give rise to severe, localized IR-drop that may invalidate the test by corrupting the contents of scan flip-flops or inducing excessive shift clock skew. In this work, we propose new methods to (1) quickly analyze all shift cycles of a given scan design and a test set for potential shift switching activity hot-spots and to (2) avoid them by targeted partial shifting of the scan chains. The results on ITC'99 benchmark circuits show the computational feasibility of the analysis and demonstrate the effectiveness of targeted partial-shift for mitigating test data corruption risk with minimal impact on test time. Stefan Holst, Shiling Shi, Xiaoqing Wen |
PRDC | 3 |
| 2019 | Novel Double-Node-Upset-Tolerant Memory Cell Designs Through Radiation-Hardening-by-Design and LayoutabstractThis paper presents two novel memory cell designs that can completely tolerate double-node upsets. First, a layout dependent cell is proposed. Since the cell has many redundant storage nodes, the cell achieves the following robustness: 1) In the case of 1 being stored, the cell can self-recover from any double-node upset (DNU) as well as any single node upset (SNU); 2) in the case of 0 being stored, the cell can self-recover from any double-adjacent-node upset (DANU), partial double-separated-node upset (DSNU) as well as any SNU. Any other DSNU can be tolerated by the cell due to the use of the layout approach. Second, a layout-independent cell is proposed that can self-recover from any DNU as well as any SNU. Simulation results validate the robustness of the proposed cell designs. Furthermore, compared with typical existing radiation hardened memory cells, the proposed layout-dependent cell saves 55.60% read access time and 33.76% write access time at the costs of 4.28% power dissipation and 39.01% silicon area on average, still low compared with layout-independent cell designs. Aibin Yan, Jing Guo 0004, Xiaoqing Wen |
IEEE Trans. Reliab. | 5 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 51 |
| 2018 | Clock-Skew-Aware Scan Chain Grouping for Mitigating Shift Timing Failures in Low-Power Scan TestingabstractHigh scan shift power often leads to excessive heat as well as shift timing failures. Partial shift (shifting a subset of scan chains at a time) is a widely adopted approach for avoiding excessive heat by reducing global switching activity, we show for the first time that it may actually cause excessive IR-drop on some clock buffers and worsen shift clock skews, thus increasing the risk of shift timing failures. This paper addresses this problem with an innovative method, namely Clock-Skew-Aware Scan Chain Grouping (CSA-SCG). CSA-SCG properly groups scan chains to be shifted simultaneously so as to reduce the imbalance of switching activity around the clock paths for neighboring scan flip-flops in scan chains. Experiments on large ITC'99 benchmark circuits demonstrate the effectiveness of CSA-SCG for reducing scan shift clock skews to lower the risk of shift timing failures in partial shift. Yucong Zhang, Xiaoqing Wen, Stefan Holst, Kohei Miyase, Seiji Kajihara, Hans-Joachim Wunderlich |
ATS | 2 |
| 2018 | The impact of production defects on the soft-error tolerance of hardened latchesabstractAs modern technology nodes get more and more susceptible to soft-errors, various hardened latch cells have been proposed. The added redundancy used to tolerate transient faults in the field at the same time reduces the test coverage of cell-internal production defects. Moreover, the test escapes reduce the soft-error tolerance of the defective latches. This work introduces a new soft-error vulnerability metric called Post Test Vulnerability Factor that correctly measures the added vulnerability to transiant frults such as particle strikes caused by undiscovered production defects within hardened latches. Stefan Holst, Ruijun Ma 0002, Xiaoqing Wen |
ETS | 3 |
| 2018 | A Method to Detect Bit Flips in a Soft-Error Resilient TCAMabstractTernary content addressable memories (TCAMs) are special memories which are widely used in high-speed network applications such as routers, firewalls, and network address translators. In high-reliability network applications such as aerospace and defense systems, soft-error tolerant TCAMs are indispensable to prevent data corruption or faults caused by radiation. This paper shows a novel way in generating keys to cover the correct match. It proposes a novel soft-error tolerant TCAM for multiple-bit-flip errors using partial don't-care keys (X-keys). First, this paper observes the case of single-bit-flip errors by X-TCAM. Second, it extends the X-TCAM to the case of multiple-bit-flip errors called KX-TCAM, where K stands for the maximum number of errors k. KX-TCAM corrects up to k-bit-flip errors and enhances the tolerance of the TCAM against soft errors, where k is the maximum number of bit flips in a word of a TCAM. KX-TCAM consists of a TCAM, a preprocessed don't-care-bit index look-up memory (X look-up), and a backup error checking and correction (ECC)-SRAM. First, KX-TCAM randomly selects a search key. After that, KX-TCAM detects multiple-bit-flip errors by the generated X-keys using the X look-up. If the keys match the different locations, then a soft error is suspected and KX-TCAM refreshes the TCAM words by using the backup ECC-SRAM. Experimental results show that the soft-error tolerance capability of KX-TCAM significantly outperforms existing state-of-the-art schemes. Moreover, the hardware overhead of KX-TCAM is small due to the use of a single TCAM. KX-TCAM can be easily implemented and is useful for fault-tolerant packet classifiers. Infall Syafalni, Tsutomu Sasao, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Scan Chain Grouping for Mitigating IR-Drop-Induced Test Data CorruptionabstractLoading and unloading test patterns during scan testing causes many scan flip-flops to trigger simultaneously. This instantaneous switching activity during shift in turn may cause excessive IR-drop that can disrupt the states of some scan flip-flops and corrupt test stimuli or responses. A common design technique to even out these instantaneous power surges is to design multiple scan chains and shift only a group of the scan chains at a same time. This paper introduces a novel algorithm to optimally group scan chains so as to minimize the probability of test data corruption caused by excessive instantaneous IR-drop on scan flip-flops. The experiments show optimal results on all large ITC'99 benchmark circuits. Yucong Zhang, Stefan Holst, Xiaoqing Wen, Kohei Miyase, Seiji Kajihara |
ATS | 3 |
| 2017 | Analysis and mitigation or IR-Drop induced scan shift-errorsabstractExcessive IR-drop during scan shift can cause localized IR-drop around clock buffers and introduce dynamic clock skew. Excessive clock skew at neighboring scan flip-flops results in hold or setup timing violations corrupting test stimuli or test responses during shifting. We introduce a new method to assess the risk of such test data corruption at each scan cycle and flip-flop. The most likely cases of test data corruption are mitigated in a non-intrusive way by selective test data manipulation and masking of affected responses. Evaluation results show the computational feasibility of our method for large benchmark circuits, and demonstrate that a few targeted pattern changes provide large potential gains in shift safety and test time with negligible cost in fault coverage. Stefan Holst, Eric Schneider, Koshi Kawagoe, Michael A. Kochte, Kohei Miyase, Hans-Joachim Wunderlich, Seiji Kajihara, Xiaoqing Wen |
ITC | 8 |
| 2017 | GPU-Accelerated Simulation of Small Delay FaultsabstractDelay fault simulation is an essential task during test pattern generation and reliability assessment of electronic circuits. With the high sensitivity of current nano-scale designs toward even smallest delay deviations, the simulation of small gate delay faults has become extremely important. Since these faults have a subtle impact on the timing behavior, traditional fault simulation approaches based on abstract timing models are not sufficient. Furthermore, the detection of these faults is compromised by the ubiquitous variations in the manufacturing processes, which causes the actual fault coverage to vary from circuit instance to circuit instance, and makes the use of timing accurate methods mandatory. However, the application of timing accurate techniques quickly becomes infeasible for larger designs due to excessive computational requirements. In this paper, we present a method for fast and waveform-accurate simulation of small delay faults on graphics processing units with exceptional computational performance. By exploiting multiple dimensions of parallelism from gates, faults, waveforms, and circuit instances, the proposed approach allows for timing-accurate and exhaustive small delay fault simulation under process variation for designs with millions of gates. Eric Schneider, Michael A. Kochte, Stefan Holst, Xiaoqing Wen, Hans-Joachim Wunderlich |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 41 |
| 2017 | Low-Power Scan-Based Built-In Self-Test Based on Weighted Pseudorandom Test Pattern Generation and ReseedingabstractA new low-power (LP) scan-based built-in self-test (BIST) technique is proposed based on weighted pseudorandom test pattern generation and reseeding. A new LP scan architecture is proposed, which supports both pseudorandom testing and deterministic BIST. During the pseudorandom testing phase, an LP weighted random test pattern generation scheme is proposed by disabling a part of scan chains. During the deterministic BIST phase, the design-for-testability architecture is modified slightly while the linear-feedback shift register is kept short. In both the cases, only a small number of scan chains are activated in a single cycle. Sufficient experimental results are presented to demonstrate the performance of the proposed LP BIST approach. Xiaoqing Wen, Laung-Terng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Formal Test Point Insertion for Region-based Low-Capture-Power Compact At-Speed Scan TestabstractLaunch-Switching-Activity (LSA) is a serious problem during at-speed testing of integrated circuits, since localized LSA may lead to severe IR-drop and thus failures. The excessive LSA is conventionally mitigated by reducing the switching activity through special low-power test generation techniques, typically resulting in severe test pattern inflation and high test costs. This work introduces a novel concept of Low-Capture-Power Test Points (LCP-TPs), which are inserted to reduce switching activity in critical High-Capture-Power (HCP) regions. LCP-TPs also help in retaining high test compaction capability. An optimization- SAT based procedure is proposed to compute a small set of optimal LCP-TP locations for compact at-speed test sets with effective capture power reduction. Experimental results clearly demonstrate the advantages of LCP-TP insertion. Stephan Eggersglüß, Stefan Holst, Daniel Tille, Kohei Miyase, Xiaoqing Wen |
ATS | 5 |
| 2016 | Timing-Accurate Estimation of IR-Drop Impact on Logic- and Clock-Paths During At-Speed Scan TestabstractIR-drop induced false capture failures and test clock stretch are severe problems in at-speed scan testing. We propose a new method to efficiently and accurately identify these problems. For the first time, our approach considers the additional dynamic power caused by glitches, the spatial and temporal distribution of all toggles, and their impact on both logic paths and the clock tree without time-consuming electrical simulations. Stefan Holst, Eric Schneider, Xiaoqing Wen, Seiji Kajihara, Yuta Yamato, Hans-Joachim Wunderlich, Michael A. Kochte |
ATS | 3 |
| 2016 | A Flexible Power Control Method for Right Power Testing of Scan-Based Logic BISTabstractHigh power dissipation during scan-based logic BIST is a crucial problem that leads to over-testing. Although controlling test power of a circuit under test (CUT) to an appropriate level is strongly required, it is not easy to control test power in BIST. This paper proposes a novel power controlling method to control the toggle rate of the patterns to an arbitrary level by modifying pseudo random patterns generated by a TPG (Test Pattern Generator) of logic BIST. While many approaches have been proposed to control the toggle rate of the patterns, the proposed approach can provide higher fault coverage. Experimental results show that the proposed approach can control toggle rates to a predetermined target level and modified patterns can achieve high fault coverage without increasing test time. Takaaki Kato, Senling Wang, Yasuo Sato, Seiji Kajihara, Xiaoqing Wen |
ATS | 5 |
| 2016 | On Optimal Power-Aware Path SensitizationabstractDetailed knowledge of a circuit's timing is essential for performance optimization, timing closure, and generation of test patterns to detect small-delay defects. When an input transition is applied to the circuit's inputs, the resulting delay is not only determined by the propagation path, but also influenced by the power-supply noise. We introduce a path-sensitization procedure which precisely controls the switching activity in the circuit region surrounding the path. The procedure can maximize or minimize switching activity, or set it to a user-specified value. We study the accuracy-vs.-efficiency trade-offs for a hierarchy of timing models, from coarse zero-delay assumption to a waveform-accurate approach with sub-cycle resolution. For the first time, we present a MaxSAT formulation which guarantees maximization or minimization of switching activity, stemming from transitions and from glitches, simultaneously with path sensitization. We validate the quality of the generated test patterns using a mixed-mode IR-drop-aware timing simulator. Matthias Sauer 0002, Jie Jiang 0018, Sven Reimer, Kohei Miyase, Xiaoqing Wen, Bernd Becker 0001, Ilia Polian |
ATS | 5 |
| 2016 | SAT-based post-processing for regional capture power reduction in at-speed scan test generationabstractWith more and more sophisticated low-power design techniques being applied to modern LSI chips for aggressive functional power reduction, the risk of fault-free chips falsely failing production test grows due to excessively high test power compared with functional power. Existing low-power ATPG methods, however, suffer from severe test data inflation and often use unfocused global test power reduction. This paper proposes a novel optimization-SAT-based at-speed scan test generation method that is explicitly targeted at eliminating high-capture-power test vectors in a pre-generated compact test set. This method employs layout information in reducing capture switching activity in a focused regional manner. Experiments demonstrate that the proposed method can effectively eliminate a large number of high-capture-power test vectors with neither test data inflation nor fault coverage loss. Stephan Eggersglüß, Kohei Miyase, Xiaoqing Wen |
ETS | 3 |
| 2016 | Thermal-Aware Small-Delay Defect Testing in Integrated Circuits for Mitigating OverkillabstractAt-speed testing of deep-submicrometer or nano-scale integrated circuits (ICs) consumes excessive power and creates hotspots and temperature gradient in the chip-under-test. The problem worsens for 3-D ICs, where heat dissipation across layers is more unbalanced. These hotspots in a circuit often cause severe degradation of performance and reliability, as a rise in temperature can introduce an extra delay along paths. As a result, the delay of an otherwise fault-free path may exceed the functional clock period. Such thermal emergencies can thus lead to over-detection and undue yield loss during testing. Their effects will be more severe for small-delay defects (SDDs), which target to sensitize the long paths in a circuit. In this paper, we quantify, for the first time, the impact of thermal emergencies on SDDs and provide a solution to mitigate them. The proposed method is based on: 1) a new thermal-aware (TA) path-selection method, 2) a TA test-ordering method, and 3) an effective scan architecture and a test-application scheme. Experimental results on benchmarks demonstrate that the new method can significantly reduce the number of over-detections of SDDs. Kele Shen, Bhargab B. Bhattacharya, Xiaoqing Wen, Xijiang Lin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Test Pattern Modification for Average IR-Drop ReductionabstractThis paper presents a novel technique that modifies automatic test pattern generation test patterns to reduce time-averaged IR drop of a test pattern. We propose a fast average IR drop estimation, which is very close to the time-averaged IR drop of time-consuming transient simulation (R2= 0.99). We calculate the contribution of every node to these nodes inside IR-drop hotspot so that we can effectively modify only a few don't care bits in the test patterns to reduce IR drop. The experimental results show that our technique successively reduces time-averaged IR drop by 10% with almost no fault coverage loss and no test pattern inflation. Wei-Sheng Ding, Hung-Yi Hsieh, Cheng-Yu Han, Chien-Mo James Li, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Logic/Clock-Path-Aware At-Speed Scan Test Generation for Avoiding False Capture Failures and Reducing Clock StretchabstractIR-drop induced by launch switching activity (LSA) in capture mode during at-speed scan testing increases delay along not only logic paths (LPs) but also clock paths (Cps). Excessive extra delay along LPs compromises test yields due to false capture failures, while excessive extra delay along CPs compromises test quality due to test clock stretch. This paper is the first to mitigate the impact of LSA on both LPs and CPs with a novel LCPA (Logic/Clock Path-Aware) at-speed scan test generation scheme, featuring (1) a new metric for assessing the risk of false capture failures based on the amount of LSA around both LPs and CPs, (2) a procedure for avoiding false capture failures by reducing LSA around LPs or masking uncertain test responses, and (3) a procedure for reducing test clock stretch by reducing LSA around CPs. Experimental results demonstrate the effectiveness of the LCPA scheme in improving test yields and test quality. Koji Asada, Xiaoqing Wen, Stefan Holst, Kohei Miyase, Seiji Kajihara, Michael A. Kochte, Eric Schneider, Hans-Joachim Wunderlich |
ATS | 2 |
| 2015 | GPU-accelerated small delay fault simulation
Eric Schneider, Stefan Holst, Michael A. Kochte, Xiaoqing Wen, Hans-Joachim Wunderlich |
DATE | 4 |
| 2015 | Identification of high power consuming areas with gate type and logic level informationabstractPower-related problems in at-speed scan testing have become more and more serious, since excessive IR-drop caused by excessive power consumption results in overtesting. There are two important factors in low-power testing: one is power estimation, the other is power reduction. Several estimation methods have been proposed based on the analysis of switching activity characteristics. In order to estimate the impact of IR-drop, it is more important to consider the area containing many cells which consume excessive power than to consider the total number of switching activity in a circuit. In this paper, we propose a novel method for identifying areas where excessive IR-drop likely occurs without using test vectors. Visualized experimental results for IWLS 2005 benchmark circuits demonstrate that the proposed method can effectively identify areas containing many cells which consume higher power than others. Such areas identified can be used in low-power test generation so as to achieve effective and efficient results. Kohei Miyase, Matthias Sauer 0002, Bernd Becker 0001, Xiaoqing Wen, Seiji Kajihara |
ETS | 4 |
| 2015 | A soft-error tolerant TCAM using partial don't-care keysabstractThis paper proposes a novel soft-error tolerant TCAM using partial don't-care keys (X-keys), namely TX, which significantly enhances the tolerance of the TCAM against soft errors. Experimental results show that the soft-error tolerance of the TX outperforms existing schemes. Moreover, the overhead of the TX is very small. Infall Syafalni, Tsutomu Sasao, Xiaoqing Wen, Stefan Holst, Kohei Miyase |
ETS | 3 |
| 2014 | Data-parallel simulation for fast and accurate timing validation of CMOS circuitsabstractGate-level timing simulation of combinational CMOS circuits is the foundation of a whole array of important EDA tools such as timing analysis and power-estimation, but the demand for higher simulation accuracy drastically increases the runtime complexity of the algorithms. Data-parallel accelerators such as Graphics Processing Units (GPUs) provide vast amounts of computing performance to tackle this problem, but require careful attention to control-flow and memory access patterns. This paper proposes the novel High-Throughput Oriented Parallel Switch-level Simulator (HiTOPS), which is especially designed to take full advantage of GPUs and provides accurate timesimulation for multi-million gate designs at an unprecedented throughput. HiTOPS models timing at transistor granularity and supports all major timing-related effects found in CMOS including pattern-dependent delay, glitch filtering and transition ramps, while achieving speedups of up to two orders of magnitude compared to traditional gate-level simulators. Eric Schneider, Stefan Holst, Xiaoqing Wen, Hans-Joachim Wunderlich |
ICCAD | 3 |
| 2013 | Search Space Reduction for Low-Power Test GenerationabstractOngoing research to shrink feature sizes of LSI circuits leads to an always increasing number of logic gates in a circuit. In general, the complexity of test generation depends on the size of a circuit. Furthermore, modern test generation methods have to consider power reduction in addition to fault detection, since excessive power caused by testing may result in over testing. In this work, we propose a method to reduce the computation time of low-power test generation. The proposed method specifies gates which will cause power issues, consequently reducing the search space for X-filling technique. The reduction of search space for Xfilling also further minimizes the amount of switching activity. Experimental results for circuits of Open Cores provided by IWLS2005 benchmarks show that the proposed method achieves both a reduced computation time and at the same time increased power reduction compared to previous methods. Kohei Miyase, Matthias Sauer 0002, Bernd Becker 0001, Xiaoqing Wen, Seiji Kajihara |
Asian Test Symposium | 4 |
| 2013 | On Achieving Capture Power Safety in At-Speed Scan-Based Logic BISTabstractThe applicability of at-speed scan-based logic built-in self-test (BIST) is being severely challenged by excessive capture power that may cause erroneous test responses for good chips. Different from conventional low-power BIST, this paper is the first that has explicitly focused on achieving capture power safety with a practical scheme called capture-power-safe BIST (CPS-BIST). The basic idea is to identify all possibly erroneous test responses and use the well-known technique of mask (partial-mask or full-mask) to block them from reaching the MISR. Experiments with large benchmark and industrial circuits show that CPS-BIST can achieve capture power safety with negligible impact on both test quality and area overhead. Akihiro Tomita, Xiaoqing Wen, Yasuo Sato, Seiji Kajihara, Patrick Girard 0001, Mark Tehranipoor, Laung-Terng Wang |
Asian Test Symposium | 2 |
| 2012 | A Transition Isolation Scan Cell Design for Low Shift and Capture PowerabstractShift and capture power management has become indispensable for modern complex low-power designs. Excessive shift power increases test application time and may jeopardize the shift operation correctness, excessive capture power during at-speed scan testing may lead to yield loss. This paper proposes a scan cell design which isolates scan cells output transitions in both shift and capture modes. Experimental results on larger ISCAS'89, ITC'99, and IWLS'05 benchmark circuits show that the proposed scan cell design lowers capture power consumptions with reasonable CPU times and test set inflation. Yi-Tsung Lin, Jiun-Lang Huang, Xiaoqing Wen |
Asian Test Symposium | 3 |
| 2012 | Session Summary III: Power-Aware Testing: Present and FutureabstractSummary form only given, as follows. Power-aware testing is facing more and more challenges in terms of test power analysis as well as test power management due to the ever-growing gap between functional power and test power for low-power LSI circuits. This special session, presented by top experts from both academia and industry, will provide firsthand information on state-of-the-art solutions as well as insightful perspectives on future R&D directions about power-aware testing. It will help researchers and practitioners alike in advancing poweraware test technologies for low-power LSI circuits, the enabler of all energy-smart electronic devices that are now indispensable in our everyday life. Xiaoqing Wen, Sudhakar M. Reddy |
Asian Test Symposium | 1 |
| 2012 | Power-aware testing: The next stageabstractComplex power management circuitry in low-power designs and the excessive gap between functional power and test power have made power-aware testing (DFT and test generation) a must. Although significant progress has been made in the past decade, more is still needed in order to achieve test power safety while maximizing test quality and minimizing test cost. This paper highlights the needs for moving to the next-stage of power-aware testing, primarily characterized by a shift of focus from global test power reduction to pinpoint test power management. Xiaoqing Wen |
ETS | 1 |
| 2012 | On pinpoint capture power management in at-speed scan test generationabstractThis paper proposes a novel scheme to manage capture power in a pinpoint manner for achieving guaranteed capture power safety, improved small-delay test capability, and minimal test cost impact in at-speed scan test generation. First, switching activity around each long path sensitized by a test vector is checked to characterize it as hot (with excessively-high switching activity), warm (with normal/functional switching activity), or cold (with excessively-low switching activity). Then, X-restoration/X-filling-based rescue is conducted on the test vector to reduce switching activity around hot paths. If the rescue is insufficient to turn a hot path into a warm path, mask is then conducted on expected test response data to instruct the tester to ignore the potentially-false test response value from the hot path, thus achieving guaranteed capture power safety. Finally, X-restoration/X-filling-based warm-up is conducted on the test vector to increase switching activity around cold paths for improving their small-delay test capability. This novel approach of pinpoint capture power management has significant advantages over the conventionalapproachofglobalcapturepower management, as demonstrated by evaluation results on large ITC'99 benchmark circuits and detailed path delay analysis. Xiaoqing Wen, Y. Nishida, Kohei Miyase, Seiji Kajihara, Patrick Girard 0001, Mark Tehranipoor, Laung-Terng Wang |
ITC | 1 |
| 2012 | A novel capture-safety checking method for multi-clock designs and accuracy evaluation with delay capture circuitsabstractExcessive capture power in at-speed scan testing may cause yield loss due to timing failures. Although reducing the number of clock domains that capture test responses simultaneously is a practical and scalable solution for reducing capture power, no available capture-safety checking metric can assess its effect in an accurate-enough manner, especially when multiple clock domains capture test responses in a short period of time. This paper proposes a novel CLEAR (CLock-Edge-Arrival-Relation-based) capture-safety checking method that, for the first time, takes clock edge arrival times for different clock domains into consideration. The accuracy and usefulness of the proposed method have been clearly demonstrated by simulation-based evaluation with the largest ITC'99 benchmark circuit as well as real-chip-based evaluation with an industrial chip embedded with on-chip delay measurement circuitry. Kohei Miyase, Masao Aso, Ryou Ootsuka, Xiaoqing Wen, Hiroshi Furukawa, Yuta Yamato, Kazunari Enokimoto, Seiji Kajihara |
VTS | 4 |
| 2012 | Launch-on-Shift Test Generation for Testing Scan Designs Containing Synchronous and Asynchronous Clock DomainsabstractThis article presents a hybrid Automatic Test Pattern Generation (ATPG) technique using the staggered Launch-On-Shift (LOS) scheme followed by the one-hot launch-on-shift scheme for testing delay faults in a scan design containing asynchronous clock domains. Typically, the staggered scheme produces small test sets but needs long ATPG runtime, whereas the one-hot scheme takes short ATPG runtime but yields large test sets. The proposed hybrid technique is intended to reduce test pattern count with acceptable ATPG runtime for multimillion-gate scan designs. In case the scan design contains multiple synchronous clock domains, and each group of synchronous clock domains is treated as a clock group and tested using a launch-aligned or a capture-aligned LOS scheme. By combining these schemes together, we found the pattern counts for two large industrial designs were reduced by approximately 1.6X to 1.8X, while the ATPG runtime was increased by 40% to 50%, when compared to the one-hot clocking scheme alone. Shianling Wu, Laung-Terng Wang, Xiaoqing Wen, Wen-Ben Jone, Michael S. Hsiao, Chien-Mo James Li, Jiun-Lang Huang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2011 | Power-Aware Test Pattern Generation for At-Speed LOS TestingabstractLaunch-off-Capture (LOC) and Launch-off-Shift (LOS) are the two main test schemes for at-speed scan delay testing. In the literature, it has been shown that LOS has higher performance than LOC in terms of fault coverage and test length, but higher peak power consumption during the launch-to-capture cycle. Power reduction seems to be the key to really exploit LOS test scheme. However, it has been proven that reducing too much test power can lead to test escape due to under-test. In this context, this study proposes a smart X-filling framework able to adapt peak power consumption during the launch-to-capture cycle according to the functional power, i.e. the power consumption of the circuit in functional mode. Here, the main goal is to obtain a final test set with peak power consumption as close as possible to the functional power. Experimental results, carried out on the well-known ITC'99 benchmarks, prove the feasibility of the proposed approach. Alberto Bosio, Luigi Dilillo, Patrick Girard 0001, Aida Todri, Arnaud Virazel, Kohei Miyase, Xiaoqing Wen |
Asian Test Symposium | 7 |
| 2011 | Efficient BDD-based Fault Simulation in Presence of Unknown ValuesabstractUnknown (X) values, originating from memories, clock domain boundaries or A/D interfaces, may compromise test signatures and fault coverage. Classical logic and fault simulation algorithms are pessimistic w.r.t. the propagation of X values in the circuit. This work proposes efficient hybrid logic and stuck-at fault simulation algorithms which combine heuristics and local BDDs to increase simulation accuracy. Experimental results on benchmark and large industrial circuits show significantly increased fault coverage and low runtime. The achieved simulation precision is quantified for the first time. Michael A. Kochte, Sandip Kundu, Kohei Miyase, Xiaoqing Wen, Hans-Joachim Wunderlich |
Asian Test Symposium | 4 |
| 2011 | Effective Launch-to-Capture Power Reduction for LOS Scheme with Adjacent-Probability-Based X-FillingabstractIt has become necessary to reduce power during LSI testing. Particularly, during at-speed testing, excessive power consumed during the Launch-To-Capture (LTC) cycle causes serious issues that may lead to the overkill of defect-free logic ICs. Many successful test generation approaches to reduce IR-drop and/or power supply noise during LTC for the launch-off capture (LOC) scheme have previously been proposed, and several of X-filling techniques have proven especially effective. With X-filling in the launch-off shift (LOS) scheme, however, adjacent-fill (which was originally proposed for shift-in power reduction) is used frequently. In this work, we propose a novel X-filling technique for the LOS scheme, called Adjacent-Probability-based X-Filling (AP-fill), which can reduce more LTC power than adjacent-fill. We incorporate AP-fill into a post-ATPG test modification flow consisting of test relaxation and X-filling in order to avoid the fault coverage loss and the test vector count inflation. Experimental results for larger ITC'99 circuits show that the proposed AP-fill technique can achieve a higher power reduction ratio than 0-fill, 1-fill, and adjacent-fill. Kohei Miyase, Y. Uchinodan, Kazunari Enokimoto, Yuta Yamato, Xiaoqing Wen, Seiji Kajihara, Fangmei Wu, Luigi Dilillo, Alberto Bosio, Patrick Girard 0001, Arnaud Virazel |
Asian Test Symposium | 5 |
| 2011 | Transition-Time-Relation based capture-safety checking for at-speed scan test generationabstractExcessive capture power in at-speed scan testing may cause timing failures, resulting in test-induced yield loss. This has made capture-safety checking mandatory for test vectors. This paper presents a novel metric, called the TTR (Transition-Time-Relation-based) metric, which takes transition time relations into consideration in capture-safety checking. Capture-safety checking with the TTR metric greatly improves the accuracy of test vector sign-off and low-capture-power test generation. Kohei Miyase, Xiaoqing Wen, Masao Aso, Hiroshi Furukawa, Yuta Yamato, Seiji Kajihara |
DATE | 2 |
| 2011 | SAT-based capture-power reduction for at-speed broadcast-scan-based test compression architectures
Michael A. Kochte, Kohei Miyase, Xiaoqing Wen, Seiji Kajihara, Yuta Yamato, Kazunari Enokimoto, Hans-Joachim Wunderlich |
ISLPED | 3 |
| 2011 | Clock-gating-aware low launch WSA test pattern generation for at-speed scan testingabstractCapture power management has become a necessity to avoid at-speed scan testing yield loss, especially for modern complex and low power designs. This paper proposes a test pattern generation methodology that utilizes the available clock-gating mechanism, a popular low power design technique, to reduce the launch cycle weighted switching activity (WSA) for at-speed scan testing. Compared to previous techniques that consider clock-gating, a significant launch cycle WSA reduction is achieved without severe test pattern inflation. Yi-Tsung Lin, Jiun-Lang Huang, Xiaoqing Wen |
ITC | 3 |
| 2011 | A novel scan segmentation design method for avoiding shift timing failure in scan testingabstractHigh power consumption in scan testing can cause undue yield loss which has increasingly become a serious problem for deep-submicron VLSI circuits. Growing evidence attributes this problem to shift timing failures, which are primarily caused by excessive switching activity in the proximities of clock paths that tends to introduce severe clock skew due to IR-drop-induced delay increase. This paper is the first of its kind to address this critical issue with a novel layout-aware scheme based on scan segmentation design, called LCTI-SS (Low-Clock-Tree-Impact Scan Segmentation). An optimal combination of scan segments is identified for simultaneous clocking so that the switching activity in the proximities of clock trees is reduced while maintaining the average power reduction effect on conventional scan segmentation. Experimental results on benchmark and industrial circuits have demonstrated the advantage of the LCTI-SS scheme. Yuta Yamato, Xiaoqing Wen, Michael A. Kochte, Kohei Miyase, Seiji Kajihara, Laung-Terng Wang |
ITC | 2 |
| 2011 | Power-aware test generation with guaranteed launch safety for at-speed scan testingabstractAt-speed scan testing may suffer from severe yield loss due to the launch safety problem, where test responses are invalidated by excessive launch switching activity (LSA) caused by test stimulus launching in the at-speed test cycle. However, previous low-power test generation techniques can only reduce LSA to some extent but cannot guarantee launch safety. This paper proposes a novel & practical power-aware test generation flow, featuring guaranteed launch safety. The basic idea is to enhance ATPG with a unique two-phase (rescue & mask) scheme by targeting at the real cause of the launch safety problem, i.e., the excessive LSA in the neighboring areas (namely impact areas) around long paths sensitized by a test vector. The rescue phase is to reduce excessive LSA in impact areas in a focused manner, and the mask phase is to exclude from use in fault detection the uncertain test response at the endpoint of any long sensitized path that still has excessive LSA in its impact area even after the rescue phase is executed. This scheme is the first of its kind for achieving guaranteed launch safety with minimal impact on test quality and test costs, which is the ultimate goal of power-aware at-speed scan test generation. Xiaoqing Wen, Kazunari Enokimoto, Kohei Miyase, Yuta Yamato, Michael A. Kochte, Seiji Kajihara, Patrick Girard 0001, Mark Tehranipoor |
VTS | 1 |
| 2011 | Special session 5B: Panel How much toggle activity should we be testing with?abstractPower dissipation of an LSI circuit during scan testing, especially at-speed scan testing, can be several times higher than that during functional operations. Excessive test power causes hot spots and/or severe IR drop that may lead to chip damage, undue yield loss, or reliability degradation, especially for low-power LSI circuits. As a result, it is becoming increasingly important to reduce test power by lowering test-induced toggle activity in order to make scan test “power-safe”. However, with the stress on reducing toggle activity during scan test one might question: Have we gone too far? Should we reduce toggle activity below functional levels? Should we even plan for many test sets with different toggle activities? Can the test power problem be solved by existing DFT and ATPG solutions? What's missing in today's solutions? What's next for low-power testing? This panel provides an interactive forum to discuss these critical questions with industry experts from both semiconductor and EDA companies. It helps practitioners and researchers alike in their quest for more effective and more efficient solutions to the test power problem. Xiaoqing Wen, Mark Tehranipoor, Rohit Kapur, Anand Bhat, Amitava Majumdar 0002, LeRoy Winemberg |
VTS | 1 |
| 2011 | Using Launch-on-Capture for Testing Scan Designs Containing Synchronous and Asynchronous Clock DomainsabstractThis paper presents a hybrid automatic test pattern generation (ATPG) technique using the staggered launch-on capture (LOC) scheme followed by the one-hot LOC scheme for testing delay faults in a scan design containing asynchronous clock domains. Typically, the staggered scheme produces small test sets but needs long ATPG runtime, whereas the one-hot scheme takes short ATPG runtime but yields large test sets. The proposed hybrid technique is intended to reduce test pattern count with acceptable ATPG runtime for multi-million-gate scan designs. In case the scan design contains multiple synchronous clock domains, each group of synchronous clock domains is treated as a clock group and tested using a launch aligned or a capture aligned LOC scheme. By combining these schemes together, we found the pattern counts for two large industrial designs were reduced by approximately 1.1X to 2.1X, while the ATPG runtime was increased by 10% to 50%, when compared to the one-hot clocking scheme alone. Shianling Wu, Laung-Terng Wang, Xiaoqing Wen, Lang Tan, Yu Hu 0001, Wen-Ben Jone, Michael S. Hsiao, Chien-Mo James Li, Jiun-Lang Huang, Lizhen Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Analysis of power consumption and transition fault coverage for LOS and LOC testing schemesabstractAt-speed scan testing has become mandatory due to the extreme CMOS technology scaling. The two main at-speed scan testing schemes are namely Launch-Off-Shift (LOS) and Launch-Off-Capture (LOC). As it can be easily implemented, LOC has been widely investigated in the literature in the last few years, especially regarding test power consumption. Conversely, LOS has received much less attention. In this paper, we propose a comparison between the two testing schemes in terms of transition fault coverage and power consumption, in order to quantify the pros and cons of LOS with respect to LOC. This study shows that LOS not only exhibits higher performance in coverage but also does not require as much extra power as predicted, especially in terms of peak power. These facts may represent convincing arguments for its wider use and development. Fangmei Wu, Luigi Dilillo, Alberto Bosio, Patrick Girard 0001, Serge Pravossoudovitch, Arnaud Virazel, Junxia Ma, Wei Zhao 0010, Mark Tehranipoor, Xiaoqing Wen |
DDECS | 10 |
| 2010 | On estimation of NBTI-Induced delay degradationabstractNBTI, which is one of well-known aging phenomena, brings delay degradation in deep submicron VLSIs. In order to detect NBTI-induced delay faults, we need to estimate delay degradation and apply delay test for the circuit in the field. This paper discusses on estimation of NBTI-Induced delay degradation. We first analyze the effect of the delay degradation, and then give a procedure of path selection in which long paths after the delay degradation are selected for the delay test in the filed. Experimental results show that estimation of delay degradation significantly affects path selection, and accurate estimation is important for the test. Mitsumasa Noda, Seiji Kajihara, Yasuo Sato, Kohei Miyase, Xiaoqing Wen, Yukiya Miura |
ETS | 5 |
| 2010 | Is test power reduction through X-filling good enough?abstractThis study investigates the reasons why test power reduction through X-filling techniques works well for cycle-average power reduction but is not so efficient concerning instantaneous peak power reduction. Fangmei Wu, Luigi Dilillo, Alberto Bosio, Patrick Girard 0001, Serge Pravossoudovitch, Arnaud Virazel, Mark Tehranipoor, Kohei Miyase, Xiaoqing Wen |
ITC | 9 |
| 2010 | Using Launch-on-Capture for Testing BIST Designs Containing Synchronous and Asynchronous Clock DomainsabstractThis paper presents a new at-speed logic built-in self-test (BIST) architecture supporting two launch-on-capture schemes, namely aligned double-capture and staggered double-capture, for testing multi-frequency synchronous and asynchronous clock domains in a scan-based BIST design. The proposed architecture also includes BIST debug and diagnosis circuitry to help locate BIST failures. The aligned scheme detects and allows diagnosis of structural and delay faults among all synchronous clock domains, whereas the staggered scheme detects and allows diagnosis of structural and delay faults among all asynchronous clock domains. Both schemes solve the long-standing problem of using the conventional one-hot scheme, which requires testing each clock domain one at a time, or the simultaneous scheme, which requires adding isolation logic to normal functional paths across interacting clock domains. Physical implementation is easily achieved by the proposed solution due to the use of a slow-speed, global scan enable signal and reduced timing-critical design requirements. Application results for industrial designs demonstrate the effectiveness of the proposed architecture. Laung-Terng Wang, Xiaoqing Wen, Shianling Wu, Hiroshi Furukawa, Hao-Jan Chao, Boryau Sheu, Jianghao Guo, Wen-Ben Jone |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | CAT: A Critical-Area-Targeted Test Set Modification Scheme for Reducing Launch Switching Activity in At-Speed Scan TestingabstractReducing excessive launch switching activity (LSA) is now mandatory in at-speed scan testing for avoiding test-induced yield loss, and test set modification is preferable for this purpose. However, previous low-LSA test set modification methods may be ineffective since they are not targeted at reducing launch switching activity in the areas around long sensitized paths, which are spatially and temporally critical for test-induced yield loss. This paper proposes a novel CAT (Critical-Area-Targeted) low-LSA test modification scheme, which uses long sensitized paths to guide launch-safety checking, test relaxation, and X-filling. As a result, launch switching activity is reduced in a pinpoint manner, which is more effective for avoiding test-induced yield loss. Experimental results on industrial circuits demonstrate the advantage of the CAT scheme for reducing launch switching activity in at-speed scan testing. Kazunari Enokimoto, Xiaoqing Wen, Yuta Yamato, Kohei Miyase, H. Sone, Seiji Kajihara, Masao Aso, Hiroshi Furukawa |
Asian Test Symposium | 2 |
| 2009 | A novel post-ATPG IR-drop reduction scheme for at-speed scan testing in broadcast-scan-based test compression environmentabstractReducing IR-drop in the test cycle during at-speed scan testing has become mandatory for avoiding test-induced yield loss. An efficient approach for this purpose is post-ATPG test modification based on X-identification and X-filling since it causes no circuit/clock design change and no test vector count inflation. However, applying this approach to test compression has been considered challenging due to the limited availability of X-bits. This paper solves this serious problem by proposing a novel and practical CA (Compression-Aware) test modification scheme for reducing IR-drop in the widely-used broadcast-scan based test compression environment. This unique scheme features (1) CA circuit remodeling for minimizing the effort of applying test modification to broadcast-scan-based test compression, (2) CA X-identification for increasing X-bits for risky test vectors, and (3) CA X-filling for effectively using limited X-bits in reducing IR-drop. As a result, the CA test modification scheme can achieve significant IR-drop reduction even when a test cube only has a small number of X-bits. This advantage is clearly demonstrated by experimental results on three compression configurations created from an industrial circuit. Kohei Miyase, Yuta Yamato, Kenji Noda, Hideaki Ito, Kazumi Hatayama, Takashi Aikyo, Xiaoqing Wen, Seiji Kajihara |
ICCAD | 7 |
| 2009 | A GA-Based Method for High-Quality X-Filling to Reduce Launch Switching Activity in At-speed Scan TestingabstractPower-aware X-filling is a preferable approach to avoiding IR-drop-induced yield loss in at-speed scan testing. However, the quality of previous X-filling methods for reducing launch switching activity may be unsatisfactory, due to low effect (insufficient and global-only reduction) and/or low scalability (long CPU time). This paper addresses this quality problem with a novel, GA (Genetic Algorithm) based X-filling method, called GA-fill. Its goals are (1) to achieve both effectiveness and scalability in a more balanced manner, and (2) to make the reduction effect of launch switching activity more concentrated on critical areas that have higher impact on IR-drop-induced yield loss.Evaluation experiments are being conducted on benchmark and industrial circuits, and initial results have demonstrated the usefulness of GA-fill. Yuta Yamato, Xiaoqing Wen, Kohei Miyase, Hiroshi Furukawa, Seiji Kajihara |
PRDC | 2 |
| 2009 | Power Supply Noise Reduction for At-Speed Scan Testing in Linear-Decompression EnvironmentabstractYield loss caused by excessive power supply noise has become a serious problem in at-speed scan testing. AlthoughX-filling techniques are available to reduce the launch cycle switching activity, their performance may not be satisfactory in the linear-decompressor-based test compression environment. This paper solves this problem by proposing a novel integrated automatic test pattern generation scheme that efficiently and effectively performs compressible low-capture-powerX-filling. Related theoretical principles are established, based on which the problem size is substantially reduced. The proposed scheme is validated by benchmark circuits, as well as an industry design in the embedded deterministic test environment. Meng-Fan Wu, Jiun-Lang Huang, Xiaoqing Wen, Kohei Miyase |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | CTX: A Clock-Gating-Based Test Relaxation and X-Filling Scheme for Reducing Yield Loss Risk in At-Speed Scan TestingabstractAt-speed scan testing is susceptible to yield loss risk due to power supply noise caused by excessive launch switching activity. This paper proposes a novel two-stage scheme, namely CTX (Clock-Gating-Based Test Relaxation and X-Filling), for reducing switching activity when test stimulus is launched. Test relaxation and X-filling are conducted (1) to make as many FFs inactive as possible by disabling corresponding clock-control signals of clock-gating circuitry in Stage-1 (Clock-Disabling), and (2) to make as many remaining active FFs as possible to have equal input and output values in Stage-2 (FF-Silencing). CTX effectively reduces launch switching activity, thus yield loss risk, even with a small number of donpsilat care (X) bits as in test compression, without any impact on test data volume, fault coverage, performance, and circuit design. Hiroshi Furukawa, Xiaoqing Wen, Kohei Miyase, Yuta Yamato, Seiji Kajihara, Patrick Girard 0001, Laung-Terng Wang, Mark Tehranipoor |
ATS | 2 |
| 2008 | Practical Challenges in Logic BIST ImplementationabstractTurboBIST-Logic (TBL) is a software tool suite for incorporating logic built-in self-test (BIST) technology into digital Integrated Circuits and has been used by a variety of industrial designs globally since 2002. This abstract describes major features of TBL, and uses three industrial cases to show practical issues encountered and solved over the years. It also discusses an important new trend in going "hybrid," a flexible combination of capture-clocking schemes, with the goal to achieve an ever more optimal result over stand-alone schemes. Each of the three cases had its unique requirements for logic BIST, some needing to customize an existing solution, but all were set to achieve common BIST goals of at-speed testing, simple test interface to/from ATE, low test cost, high product reliability, and repeat testability investment reuse from IC, board, system, to in-field diagnosis. Shianling Wu, Hiroshi Furukawa, Boryau Sheu, Laung-Terng Wang, Hao-Jan Chao, Lizhen Yu, Xiaoqing Wen, Michio Murakami |
ATS | 7 |
| 2008 | Power-Aware Testing and Test Strategies for Low Power Devices
Dimitris Gizopoulos, Kaushik Roy 0001, Patrick Girard 0001, Nicola Nicolici, Xiaoqing Wen |
DATE | 5 |
| 2008 | Test Strategies for Low Power DevicesabstractUltra low-power devices are being developed for embedded applications in bio-medical electronics, wireless sensor networks, environment monitoring and protection, etc. The testing of these low-cost, low-power devices is a daunting task. Depending on the target application, there are stringent guidelines on the number of defective parts per million shipped devices. At the same time, since such devices are cost-sensitive, test cost is a major consideration. Since system-level power-management techniques are employed in these devices, test generation must be power- management-aware to avoid stressing the power distribution infrastructure in the test mode. Structural test techniques such as scan test, with or without compression, can result in excessive heat dissipation during testing and damage the package. False failures may result due to the electrical and thermal stressing of the device in the test mode of operation, leading to yield loss. This paper considers different aspects of testing low-power devices and some new techniques to address these problems. C. P. Ravikumar, Mokhtar Hirech, Xiaoqing Wen |
DATE | 3 |
| 2008 | A Capture-Safe Test Generation Scheme for At-Speed Scan TestingabstractCapture-safety, defined as the avoidance of any timing error due to unduly high launch switching activity in capture mode during at-speed scan testing, is critical for avoiding test- induced yield loss. Although point techniques are available for reducing capture IR-drop, there is a lack of complete capture-safe test generation flows. The paper addresses this problem by proposing a novel and practical capture-safe test generation scheme, featuring (1) reliable capture-safety checking and (2) effective capture-safety improvement by combining X-bit identification & X-filling with low launch- switching-activity test generation. This scheme is compatible with existing ATPG flows, and achieves capture-safety with no changes in the circuit-under-test or the clocking scheme. Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Hiroshi Furukawa, Yuta Yamato, Atsushi Takashima, Kenji Noda, Hideaki Ito, Kazumi Hatayama, Takashi Aikyo, Kewal K. Saluja |
ETS | 1 |
| 2008 | Effective IR-drop reduction in at-speed scan testing using Distribution-Controlling X-IdentificationabstractTest data modification based on test relaxation and X-filling is the preferable approach for reducing excessive IR-drop in at-speed scan testing to avoid test-induced yield loss. However, none of the existing test relaxation methods can control the distribution of identified don’t care bits (X-bits), thus adversely affecting the effectiveness of IR-drop reduction. In this paper, we propose a novel test relaxation method, called Distribution-Controlling X-Identification (DC-XID), which controls the distribution of X-bits identified from a set of fully-specified test vectors for the purpose of effectively reducing IR-drop. Experimental results on large industrial circuits demonstrate the effectiveness and practicality of the proposed method in reducing IR-drop, without any impact on fault coverage, test data volume, or test circuit size. Kohei Miyase, Kenji Noda, Hideaki Ito, Kazumi Hatayama, Takashi Aikyo, Yuta Yamato, Hiroshi Furukawa, Xiaoqing Wen, Seiji Kajihara |
ICCAD | 8 |
| 2008 | Turbo1500: Toward Core-Based Design for Test and Diagnosis Using the IEEE 1500 StandardabstractThis paper describes a core-based test and diagnosis integration and automation system, called Turbo1500, which automatically synthesizes test and diagnosis logic in accordance with the IEEE 1500 standard. Turbo1500 serves two major purposes. One is for use as a core test automation tool in a system-on-chip (SOC) environment to automatically connect multiple cores from various sources and create testbenches each targeting an individual core under the control of a chip-level test access port (TAP) controller. The other is for hierarchical (block-by-block) core test and diagnosis when chips on a printed-circuit board are embedded with 1149.1 boundary scan I/O cells and cores under test and diagnosis are surrounded with 1500-compliant wrapper cells. Application experience showed that the simplicity of the IEEE 1500 standard combined with an easy-to-use automation tool can make core-based design for test and diagnosis no longer a nightmare, especially when some cores are extremely large or complex. Laung-Terng Wang, Ravi Apte, Shianling Wu, Boryau Sheu, Kuen-Jong Lee, Xiaoqing Wen, Wen-Ben Jone, Chia-Hsien Yeh, Wei-Shin Wang, Hao-Jan Chao, Jianghao Guo, Yanlong Niu, Yi-Chih Sung, Chi-Chun Wang |
ITC | 6 |
| 2008 | Reducing Power Supply Noise in Linear-Decompressor-Based Test Data Compression Environment for At-Speed Scan TestingabstractYield loss caused by excessive power supply noise has become a serious problem in at-speed scan testing. Although X-filling techniques are available to reduce the launch cycle switching activity, their performance may not be satisfactory in the linear-decompressor-based test compression environment. This work is the first to solve this problem by proposing a novel integrated ATPG scheme that efficiently and effectively performs compressible X-filling. Related theoretical principles are established, based on which the problem size is substantially reduced. The proposed scheme is validated by large benchmark circuits as well as an industry design in the embedded deterministic test (EDT) environment. Meng-Fan Wu, Jiun-Lang Huang, Xiaoqing Wen, Kohei Miyase |
ITC | 3 |
| 2008 | Low Capture Switching Activity Test Generation for Reducing IR-Drop in At-Speed Scan Testing
Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita |
J. Electron. Test. | 1 |
| 2007 | Critical-Path-Aware X-Filling for Effective IR-Drop Reduction in At-Speed Scan TestingabstractIR-drop-induced malfunction is mostly caused by timing violations on activated critical paths during the capture cycle of at-speed scan testing. A critical-path-aware X-filling method is proposed for reducing IR-drop, especially on gates that are close to activated critical paths, thus effectively preventing test-induced yield loss. Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Yuji Ohsumi, Kewal K. Saluja |
DAC | 1 |
| 2007 | Embedded Tutorial on Low Power TestabstractExcessive power during test affects the reliability of digital integrated circuits, test throughput and manufacturing yield. Numerous low power test methods have been investigated over the past decade and new power-aware automatic test pattern generation, design-for-test and test planning techniques have emerged. This embedded tutorial introduces the topic of low power test and it overviews the basic techniques and some recent advancements in this field. Nicola Nicolici, Xiaoqing Wen |
ETS | 2 |
| 2007 | Estimation of delay test quality and its application to test generationabstractAs a method to evaluate delay test quality of test patterns, SDQM (Statistical Delay Quality Model) has been proposed for transition faults. In order to derive better test quality by SDQM, the following two things are important: for each transition fault, (1) to find out the accurate length of the longest sensitizable paths along which the fault is activated and propagated, and (2) to generate a test pattern that detects the fault through as long paths as possible. In this paper, we propose a method to calculate the length of the potentially sensitizable longest path for detection of a transition fault. In addition, we develop a procedure to extract path information that helps high quality transition ATPG. Experimental results show that the proposed method not only derives more accurate SDQL (Statistical Delay Quality Level) but also enhances the test quality of generated test patterns. Seiji Kajihara, Shohei Morishima, Masahiro Yamamoto, Xiaoqing Wen, Masayasu Fukunaga, Kazumi Hatayama, Takashi Aikyo |
ICCAD | 4 |
| 2007 | A novel scheme to reduce power supply noise for high-quality at-speed scan testingabstractHigh-quality at-speed scan testing, characterized by high small-delay-defect detecting capability, is indispensable to achieve high delay test quality for DSM circuits. However, such testing is susceptible to yield loss due to excessive power supply noise caused by high launch-induced switching activity. This paper addresses this serious problem with a novel and practical post-ATPG X-filling scheme, featuring (1) a test relaxation method, called path keeping X-identification, that finds don't-care bits from a fully-specified transition delay test set while preserving its delay test quality by keeping the longest paths originally sensitized for fault detection, and (2) an X-filling method, called justification-probability-based fill (JP-fill), that is both effective and scalable for reducing launch-induced switching activity. This scheme can be easily implemented into any ATPG flow to effectively reduce power supply noise, without any impact on delay test quality, test data volume, area overhead, and circuit timing. Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Yuta Yamato, Patrick Girard 0001, Yuji Ohsumi, Laung-Terng Wang |
ITC | 1 |
| 2006 | A dynamic test compaction procedure for high-quality path delay testingabstractWe propose a dynamic test compaction procedure to generate high-quality test patterns for path delay faults. While the proposed procedure generates a compact two-pattern test set for paths selected by a path selection criterion, the generated test set would detect not only faults on the selected paths but also faults on many unselected paths. Hence both high test quality by detecting untargeted faults and test cost reduction by reducing test patterns can be achieved. Experimental results show that the proposed procedure could generate a compact test set that detect many untargeted path delay faults certainly, compared with the static test compaction method previously proposed (Kajihara et al., 2005). Masayasu Fukunaga, Seiji Kajihara, Xiaoqing Wen, Toshiyuki Maeda, Shuji Hamada, Yasuo Sato |
ASP-DAC | 3 |
| 2006 | Test data compression based on clustered random access scanabstractWe proposed clustered random access scan (CRAS) architecture to reduce test data volume. CRAS makes use of the compatibility of the test stimuli to cluster the scan cells, and assigns every cluster a unique address. The compression ratio upper bound of CRAS is analyzed based on the random graph theory. Experimental results on ISCAS'89 benchmarks and two industry designs show that the proposed CRAS architecture can yield on average 67.3% reduction in test data volume, with reasonable area and routing overhead than scan design Yu Hu 0001, Jia Li 0022, Yinhe Han 0001, Xiaowei Li 0001, Huawei Li 0001, Laung-Terng Wang, Xiaoqing Wen |
ATS | 9 |
| 2006 | Highly-Guided X-Filling Method for Effective Low-Capture-Power Scan Test GenerationabstractX-filling is preferred for low-capture-power scan test generation, since it reduces IR-drop-induced yield loss without the need of any circuit modification. However, the effectiveness of previous X-filling methods suffers from lack of guidance in selecting targets and values for X-filling. This paper addresses this problem with a highly-guided X-filling method based on two novel concepts: (1) X-score for X-filling target selection and (2) probabilistic weighted capture transition count for Y-filling value selection. Experimental results show the superiority of the new X-filling method for capture power reduction. Xiaoqing Wen, Kohei Miyase, Yuta Yamato, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja |
ICCD | 1 |
| 2006 | A Novel and Practical Control Scheme for Inter-Clock At-Speed TestingabstractThe quality of at-speed testing is being severely challenged by the problem that an inter-clock logic block existing between two synchronous clocks is not efficiently tested or totally ignored due to complex test control. This paper addresses the problem with a novel inter-clock at-speed test control scheme, featuring a compact and robust on-chip inter-clock enable generator design. The new scheme can generate inter-clock at-speed test clocks from PLLs, and is feasible for both ATE-based scan testing and logic BIST. Successful applications to industrial circuits have proven its effectiveness in improving the quality of at-speed testing Hiroshi Furukawa, Xiaoqing Wen, Laung-Terng Wang, Boryau Sheu, Shianling Wu |
ITC | 2 |
| 2006 | A Framework of High-quality Transition Fault ATPG for Scan CircuitsabstractThis paper presents a framework of high-quality test generation for transition faults in full scan circuits. This work assumes a restricted broad-side testing as a test application method for two-pattern tests where control of primary inputs and observation of primary outputs are restricted. Because we use a modified time expansion model of a circuit-under-test during ATPG and fault simulation, conventional ATPG and fault simulation programs can work with minor change. The proposed ATPG method consists of two algorithms, which are activation-first and propagation-first, and for each fault it is decided which algorithm should be applied. Test patterns are generated such that transition faults with small delay can be detected, i.e. a path for fault excitation and propagation becomes as long as possible. In experimental results we evaluate test patterns generated by the proposed method using SDQM that calculates delay test quality, and show the effectiveness of the proposed method Seiji Kajihara, Shohei Morishima, Akane Takuma, Xiaoqing Wen, Toshiyuki Maeda, Shuji Hamada, Yasuo Sato |
ITC | 4 |
| 2006 | A New ATPG Method for Efficient Capture Power Reduction During Scan TestingabstractHigh power dissipation can occur when the response to a test vector is captured by flip-flops in scan testing, resulting in excessive JR drop, which may cause significant capture-induced yield loss in the DSM era. This paper addresses this serious problem with a novel test generation method, featuring a unique algorithm that deterministically generates test cubes not only for fault detection but also for capture power reduction. Compared with previous methods that passively conduct X-filling for unspecified bits in test cubes generated only for fault detection, the new method achieves more capture power reduction with less test set inflation. Experimental results show its effectiveness. Xiaoqing Wen, Seiji Kajihara, Kohei Miyase, Kewal K. Saluja, Laung-Terng Wang, Khader S. Abdel-Hafez, Kozo Kinoshita |
VTS | 1 |
| 2005 | Test compression for scan circuits using scan polarity adjustment and pinpoint test relaxationabstractThis paper presents a test compression method that effectively derives the capability of a run-length based encoding. The method employs two techniques: scan polarity adjustment and pinpoint test relaxation. Given a test set for a full-scan circuit, scan polarity adjustment selectively flips the values of some scan cells in test patterns. It can be realized by changing connections between two scan cells so that the inverted output of a scan cell, Q, is connected to the next scan cell. Pinpoint test relaxation flips some specified 1s in the test patterns to 0s without any fault coverage loss. Both techniques are applied by referring to a gain-penalty table to determine scan cells or bits to be flipped. Experimental results on ISCAS'89 benchmark circuits show that the proposed method could reduce test data volume by 36%. Switching activities, i.e. test power during scan testing, were also reduced. Yasumi Doi, Seiji Kajihara, Xiaoqing Wen, Lei Li 0036, Krishnendu Chakrabarty |
ASP-DAC | 3 |
| 2005 | On Improving Defect Coverage of Stuck-at Fault TestsabstractRecently design for manufacturability (DFM) has been required to achieve higher process yield. Information obtained from silicon by testing and/or fault analysis is sometimes fed back for redesign of VLSI circuits. In this paper we propose a method to maximize defect coverage of a test set initially generated for stuck-at faults in a full scan sequential circuit by using feed back information from fault analysis. If a test set for more complex faults than stuck-at faults is generated, higher defect coverage would be obtained. Such a test set, however, would have a large number of test vectors, and hence the test costs would go up. The proposed method improves defect coverage of the test set by not adding new test vectors but modifying test vectors with the information obtained from fault analysis. Therefore there are no negative impacts on test data volume and test application time. The initial fault coverage for stuck-at faults of the test set is guaranteed with modified test vectors. In this paper we focus on detecting as many as possible non-feedback AND/OR-type bridging faults. Experimental results show that the proposed method significantly decreases the number of non-feedback AND/OR-type bridging faults undetected by a test set generated for stuck-at faults. Kohei Miyase, Kenta Terashima, Seiji Kajihara, Xiaoqing Wen, Sudhakar M. Reddy |
Asian Test Symposium | 4 |
| 2005 | Path delay test compaction with process variation toleranceabstractIn this paper we propose a test compaction method for path delay faults in a logic circuit. The method generates a compact set of two-pattern tests for faults on long paths selected with a criterion. While the proposed method generates each two-pattern test for more than one fault in the target fault list as well as ordinary test compaction methods, secondary target faults are selected from the fault list such that many other faults, which may not be included in the fault list, are detected by the test pattern. Even if faults on long paths in a manufactured circuit are not included in the fault list due to a process variation or noise, the compact test set would detect the longer untargeted faults, i.e., the test set has a noise or variation tolerant nature. Experimental results show that the proposed method can generate a compact test set and it detects longer untargeted path delay faults efficiently. Seiji Kajihara, Masayasu Fukunaga, Xiaoqing Wen, Toshiyuki Maeda, Shuji Hamada, Yasuo Sato |
DAC | 3 |
| 2005 | At-Speed Logic BIST for IP CoresabstractThis paper describes a flexible logic BIST scheme that features high fault coverage achieved by fault-simulation guided test point insertion, real at-speed test capability for multi-clock designs without clock frequency manipulation, and easy physical implementation due to the use of a low-speed SE signal. Application results of this scheme to two widely used IP cores are also reported. B. Cheon, Laung-Terng Wang, Xiaoqing Wen, Po-Ching Hsu, Jin Woo Cho, Hao-Jan Chao, Shianling Wu |
DATE | 4 |
| 2005 | At-Speed Logic BIST Architecture for Multi-Clock DesignsabstractThis paper presents an at-speed logic BIST architecture for testing multi-clock, multi-frequency designs. The scheme employed allows true at-speed test quality for circuits containing multiple clocks without any clock frequency manipulation. Physical implementation is easily achieved due to the use of a low-speed scan enable (SE) signal and reduced timing-critical design requirements. Application results for two industrial designs are also reported. Laung-Terng Wang, Xiaoqing Wen, Po-Ching Hsu, Shianling Wu, Jonhson Guo |
ICCD | 2 |
| 2005 | UltraScan: using time-division demultiplexing/multiplexing (TDDM/TDM) with VirtualScan for test cost reductionabstractThis paper describes time-division demultiplexing and multiplexing of high-data-rate scan patterns applied on I/O's into low-data-rate scan patterns applied on VirtualScan compression circuitry to further reduce test application time and test pin-count without coverage loss Laung-Terng Wang, Khader S. Abdel-Hafez, Xiaoqing Wen, Boryau Sheu, Shianling Wu, Shyh-Horng Lin, Ming-Tung Chang |
ITC | 3 |
| 2005 | Low-capture-power test generation for scan-based at-speed testingabstractScan-based at-speed testing is a key technology to guarantee timing-related test quality in the deep submicron era. However, its applicability is being severely challenged since significant yield loss may occur from circuit malfunction due to excessive IR drop caused by high power dissipation when a test response is captured. This paper addresses this critical problem with a novel low-capture-power X-filling method of assigning 0's and 1's to unspecified (X) bits in a test cube obtained during ATPG. This method reduces the circuit switching activity in capture mode and can be easily incorporated into any test generation flow to achieve capture power reduction without any area, timing, or fault coverage impact. Test vectors generated with this practical method greatly improve the applicability of scan-based at-speed testing by reducing the risk of test yield loss. Xiaoqing Wen, Yoshiyuki Yamashita, Shohei Morishima, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita |
ITC | 1 |
| 2005 | Compression/Scan Co-Design for Reducing Test Data Volume, Scan-in Power Dissipation and Test Application TimeabstractTesting chips is very critical to guarantee chips are fault-free before they are integrated in a system, so as to increase the reliability of the system. Although full-scan is a widely adopted design-for-test technique for LSI design and testing, the need for reducing the test data volume, scan-in power dissipation and test application time (VPT) of the full-scan designed chip is imperative. Based on the analysis of the characteristics of the variable-to-fixed run-length coding technique and the random access scan architecture, this paper presents a novel design scheme tackling all VPT issues simultaneously. Experimental results on ISCAS'89 benchmarks have shown on average 51.2%, 99.5%, 99.3% and 85.5% reduction in test data volume, average scan-in power dissipation, peak scan-in power dissipation and test application time, respectively. Yu Hu 0001, Xiaowei Li 0001, Huawei Li 0001, Xiaoqing Wen |
PRDC | 4 |
| 2005 | On Low-Capture-Power Test Generation for Scan TestingabstractResearch on low-power scan testing has been focused on the shift mode, with little or no consideration given to the capture mode power. However, high switching activity when capturing a test response can cause excessive IR drop, resulting in significant yield loss. This paper addresses this problem with a novel low-capture-power X-filling method by assigning 0's and 1's to unspecified (X) bits in a test cube to reduce the switching activity in capture mode. This method can be easily incorporated into any test generation flow, where test cubes are obtained during ATPG or by X-bit identification. Experimental results show the effectiveness of this method in reducing capture power dissipation without any impact on area, timing, and fault coverage. Xiaoqing Wen, Yoshiyuki Yamashita, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita |
VTS | 1 |
| 2005 | Fault Diagnosis of Physical Defects Using Unknown Behavior Model
Xiaoqing Wen, Hideo Tamamoto, Kewal K. Saluja, Kozo Kinoshita |
J. Comput. Sci. Technol. | 1 |
| 2004 | On per-test fault diagnosis using the X-fault modelabstractThis work proposes a new per-test fault diagnosis method based on the X-fault model. The X-fault model represents all possible behaviors of a physical defect or defects in a gate and/or on its fanout branches by using different X symbols on the fanout branches. A novel technique is proposed for analyzing the relation between observed and simulated responses to extract diagnostic information and to score the results of diagnosis. Experimental results show the effectiveness of our method. Xiaoqing Wen, Tokiharu Miyoshi, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita |
ICCAD | 1 |
| 2004 | VirtualScan: A New Compressed Scan Technology for Test Cost ReductionabstractThis work describes the VirtualScan technology for scan test cost reduction. Scan chains in a VirtualScan circuit are split into shorter ones and the gap between external scan ports and internal scan chains are bridged with a broadcaster and a compactor. Test patterns for a VirtualScan circuit are generated directly by one-pass VirtualScan ATPG, in which multi-capture clocking and maximum test compaction are supported. In addition, VirtualScan ATPG avoids unknown-value and aliasing effects algorithmically without adding any additional circuitry. The VirtualScan technology has achieved successful tape-outs of industrial chips and has been proven to be an efficient and easy-to-implement solution for scan test cost reduction. Laung-Terng Wang, Khader S. Abdel-Hafez, Shianling Wu, Xiaoqing Wen, Hiroshi Furukawa, Fei-Sheng Hsu, Shyh-Horng Lin, Sen-Wei Tsai |
ITC | 4 |
| 2003 | Fault Diagnosis for Physical Defects of Unknown BehaviorsabstractThis paper proposes an X-fault model for fault diagnosis of physical defects with unknown behaviors by using X symbols. An efficient X-fault simulation method and an efficient X-fault diagnostic reasoning method are presented. Based on these, an X-fault diagnosis method is described to improve the failure analysis for a wide range of physical defects in complex IC circuits. Xiaoqing Wen, Hideo Tamamoto, Kewal K. Saluja, Kozo Kinoshita |
Asian Test Symposium | 1 |
| 2001 | A Flexible Logic BIST Scheme and Its Application to SoC DesignsabstractBuilt-in self-test for logic circuits or logic BIST is an effective solution for the test cost, test quality, and test reuse problems. Logic BIST implements most ATE functions on chip so that the test cost can be reduced through less test time, less tester memory requirement, or a cheaper tester. Logic BIST applies a large number of test patterns so that more defects, either modeled or un-modeled, can be detected. In addition, logic BIST makes it easy to conduct the at-speed test for detecting timing-related defects. Furthermore, a BISTed- core makes SoC testing easier. Most of logic BIST schemes are based on the STUMPS structure, which applies random patterns generated by a PRPG to a full-scan circuit in parallel and compresses the responses into a signature with a MISR. The basic BIST flow includes initialization and a shift-capture loop. Logic BIST schemes are difficult to implement due to (1) potential timing violations at the borders from a PRPG to scan chains and from scan chains to a MISR, (2) potential timing problems caused by inserting test points, especially control points, (3) potential destructive shift operations due to clock glitches, and (4) potential overtests due to false paths activated by at-speed transition generation. This paper summarizes a flexible logic BIST scheme that addresses the above problems. Xiaoqing Wen, Hsin-Po Wang 0002 |
Asian Test Symposium | 1 |
| 1998 | Design for Diagnosability of CMOS CircuitsabstractThis paper presents a new approach to improving the diagnosability of a CMOS circuit by dividing it into independent partitions and using a separate power supply for each partition. This technique makes it possible to implement multiple I/sub DDQ/ measurement points. As a result, the diagnosability of the circuit can be improved. The problem of partitioning a circuit is addressed and optimum and heuristic solutions are proposed. The effectiveness of our approach is demonstrated through experimental results. Xiaoqing Wen, Tooru Honzawa, Hideo Tamamoto, Kewal K. Saluja, Kozo Kinoshita |
Asian Test Symposium | 1 |
| 1997 | Fault Diagnosis for Static CMOS CircuitsabstractThis paper presents a new methodology for transistor leakage fault diagnosis using both I/sub DDQ/ and logic information. A method for handling intermediate faulty voltages in fault simulation is proposed. A scheme for generating diagnostic test vectors based on logic information in the presence of intermediate faulty voltages is also proposed. An example is used to demonstrate the new diagnosis methodology. Xiaoqing Wen |
Asian Test Symposium | 1 |
| 1997 | Random Pattern Testable Design with Partial Circuit DuplicationabstractThe advantage of random testing is that test application can be performed at a low cost in the BIST scheme. However, not all circuits are random pattern testable. In this paper, we present a method for improving random pattern testability of logic circuits by partial circuit duplication. The basic idea is to detect random pattern resistant faults by using the difference between the duplicated part of a circuit and the original part. Experimental results on benchmark circuits show that high fault coverage can be achieved with a very small amount of hardware overhead. Hiroshi Yokoyama, Xiaoqing Wen, Hideo Tamamoto |
Asian Test Symposium | 2 |
| 1996 | A new method towards achieving global optimality in technology mappingabstractThis paper presents a new method for covering a Boolean network by library cells. In this method, matches are classified according to their properties. Some matches are selected unconditionally into a cover and the remaining nodes are divided into independent portions. Then, a match compatibility graph (MCG) is constructed for each portion and an optimum cover is found for it using the MCG. Thus our method finds an efficient and closer to optimum cover for the complete network. Xiaoqing Wen, Kewal K. Saluja |
ICCAD | 1 |
| 1995 | Transistor leakage fault location with ZDDQ measurementabstractThis paper discusses the problem of locating transistor leakage faults with only I/sub DDQ/ measurement. A new approach of equivalence fault collapsing is proposed for reducing the number of faults that must be considered. Fault location is performed by using both random and deterministic tests in order to obtain a high diagnostic resolution with a small number of tests. The experimental results show that the diagnosed faults are confined to only a few gates in many cases and that a very high average diagnostic resolution can be achieved for a gate-array circuit. Xiaoqing Wen, Hideo Tamamoto, Kozo Kinoshita |
Asian Test Symposium | 1 |
| 1992 | Testable Designs of Sequential Circuits Under Highly Observable ConditionabstractIt is assumed that the outputs of all gates in a circuit are observable under the highly observable condition. This paper presents two kinds of testable sequential circuits: k-UCP sequential circuits and k-UCP scan circuits, under the highly observable condition. The combinational portions in both kinds of circuits are k- UCP circuits which consist of only k-input gates and inverters. All stuck-at faults in a k-UCP sequential circuit and a k-UCP scan circuit can be tested by 3(k + 1) and k + 1 test vectors respectively under the highly observable condition. Xiaoqing Wen, Kozo Kinoshita |
ITC | 1 |
| 1992 | A Testable Design of Logic Circuits under Highly Observable ConditionabstractThe concept of k-UCP circuits is proposed. In a k-UCP circuit, all stuck-at faults and stuck-open faults can be detected and located by k+1 and k(k+1)+1 tests, respectively, under the highly observable condition. A method of modifying an arbitrary combinational circuit into a k-UCP circuit is also proposed.> Xiaoqing Wen, Kozo Kinoshita |
IEEE Trans. Computers | 1 |
| 1990 | A testable design of logic circuits under highly observable conditionabstractTwo methods for modifying an arbitrary CMOS combinational circuit into a testable CMOS combinational circuit, called a k-UCP circuit, are discussed. All stuck-at faults and stuck-open faults in a k-UCP circuit can be detected by a test pattern which is a combination of no more than 2(k+1) kinds of basic sequences of fixed-length k(k +1)+1 under highly observable conditions. And all single stuck-open faults in a k-UCP circuit can be located efficiently. Experimental results show that the NAND (NOR) circuit modification is better than the general circuit modification in terms of the number of additional transistors. Experimental results also show that setting a proper backtrack limit can reduce computation time greatly. From a practical point of view, the results obtained are of value for developing CMOS ASICs (application-specific integrated circuits), where hardware overhead is allowed to some extent, but a fast diagnostic procedure is required.> Xiaoqing Wen, Kozo Kinoshita |
ITC | 1 |