EDBT 2026 Demo / reviewers in the wild / expert
Kai-Chiang Wu
dblp:28/6933
· DBLP profile ↗
52ranked-venue papers
13as first author
24since 2021 · last 2026
0009-0000-6931-6538ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 13 first-author · 15 since 2021Software engineering, systems software and programming languages · 9 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block SkippingabstractLarge language models (LLM) have achieved remarkable performance across a wide range of tasks. However, their substantial parameter sizes pose significant challenges for deployment on edge devices with limited computational and memory resources. Low-rank compression is a promising approach to address this issue, as it reduces both computational and memory costs, making LLM more suitable for resource-constrained environments. Nonetheless, naïve low-rank compression methods require a significant reduction in the retained rank to achieve meaningful memory and computation savings. For a low-rank model, the ranks need to be reduced by more than half to yield efficiency gains. Such aggressive truncation, however, typically results in substantial performance degradation. To address this trade-off, we propose SkipCat, a novel low-rank compression framework that enables the use of higher ranks while achieving the same compression rates. First, we introduce an intra-layer shared low-rank projection method, where multiple matrices that share the same input use a common projection. This reduces redundancy and improves compression efficiency. Second, we propose a block skipping technique that omits computations and memory transfers for selected sub-blocks within the low-rank decomposition. These two techniques jointly enable our compressed model to retain more effective ranks under the same compression budget. Experimental results show that, without any additional fine-tuning, our method outperforms previous low-rank compression approaches by 7% accuracy improvement on zero-shot tasks under the same compression rate. These results highlight the effectiveness of our rank-maximized compression strategy in preserving model performance under tight resource constraints. Yu-Chen Lu, Sheng-Feng Yu, Hui-Hsien Weng, Pei-Shuo Wang, Yu-Fang Hu, Liang Hung-Chun, Hung-Yueh Chiang, Kai-Chiang Wu |
AAAI | 8 |
| 2025 | FLRC: Fine-grained Low-Rank Compressor for Efficient LLM InferenceabstractAlthough large language models (LLM) have achieved remarkable performance, their enormous parameter counts hinder deployment on resource-constrained hardware.Low-rank compression can reduce both memory usage and computational demand, but applying a uniform compression ratio across all layers often leads to significant performance degradation, and previous methods perform poorly during decoding.To address these issues, we propose the Finegrained Low-Rank Compressor (FLRC), which efficiently determines an optimal rank allocation for each layer, and incorporates progressive low-rank decoding to maintain text generation quality.Comprehensive experiments on diverse benchmarks demonstrate the superiority of FLRC, achieving up to a 17% improvement in ROUGE-L on summarization tasks compared to state-of-the-art low-rank compression methods, establishing a more robust and efficient framework to improve LLM inference. Yu-Chen Lu, Chong-Yan Chen, Chi-Chih Chang, Yu-Fang Hu, Kai-Chiang Wu |
EMNLP | 5 |
| 2025 | Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI AccelerationabstractFPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic researchers have proposed in-fabric blocks that perform efficient tensor computations. However, these blocks are primarily optimized for dense computation, while most DNNs exhibit sparsity. To address this limitation, we propose incorporating structured sparsity support into FPGA architectures. We architect 2D systolic in-fabric blocks, named systolic sparse tensor (SST) slices, that support multiple degrees of sparsity to efficiently accelerate a wide variety of DNNs. SSTs support dense operation, 2:4 (50%) and 1:4 (75%) sparsity, as well as a new 1:3 (66.7%) sparsity level to further increase flexibility. When demonstrating on general matrix multiplication (GEMM) accelerators, which are the heart of most current DNN accelerators, our sparse SST-based designs attain up to 5× higher FPGA frequency and 10.9× lower area, compared to traditional FPGAs. Moreover, evaluation of the proposed SSTs on state-of-the-art sparse ViT and CNN models exhibits up to 3.52× speedup with minimal area increase of up to 13.3%, compared to dense in-fabric acceleration. Endri Taka, Ning-Chi Huang, Chi-Chih Chang, Kai-Chiang Wu, Aman Arora 0001, Diana Marculescu |
FPGA | 4 |
| 2025 | Integrating Neural Architecture Search and Rematerialization for Efficient On-Device LearningabstractDeep neural networks (DNNs) have notable performance in many fields, such as computer vision. Training a neural network on an edge device, commonly called on-device learning, has grown crucial for applications demanding real-time processing and enhanced privacy. However, existing on-device learning methods often face limitations, such as decreasing application accuracy, causing complexity in design and implementation, and increasing computational overhead, all of which hinder their effectiveness in reducing memory usage. In this paper, we address the issue by inspecting the memory usage of training a DNN, analyzing the effects of different on-device learning strategies, and introducing a framework that integrates neural architecture search (NAS) and rematerialization. The supernet of NAS can provide a population of compressed subnets/architectures to be trained without additional computational overhead, while rematerialization can mitigate memory consumption without accuracy loss. By leveraging the memory-saving effect of both supernet-based model compression and rematerialization, our proposed method can obtain suitable models that fit within the memory constraint while achieving a better trade-off between training time and model performance. In the experiments, we utilized complex datasets (CIFAR-100 and CUB-200) to fine-tune models on Raspberry Pi. The experimental results represent the effectiveness of our method in real-world on-device learning scenarios. Chih-Ling Chen, Kai-Chiang Wu, Ning-Chi Huang |
GECCO | 2 |
| 2025 | Invited Paper: 2025 ICCAD CAD Contest Problem A: Hardware Trojan Detection on Gate-Level NetlistabstractThe increasing reliance on third-party intellectual property (IP) cores in modern integrated circuit (IC) design has introduced significant security vulnerabilities, particularly the risk of Hardware Trojans (HTs). These malicious modifications can compromise system integrity, leading to data leakage, unauthorized access, or functional failures. Traditional detection methods often depend on the availability of a golden chip, which is not always feasible. This paper presents the 2025 ICCAD CAD Contest Problem A, which challenges participants to develop machine learning-based solutions for detecting HTs directly from gate-level netlists without requiring a golden reference. The problem is formulated with defined Trojan behaviors, input/output specifications, and evaluation metrics, including correctness and F1 score. The contest aims to foster innovation in HT detection by leveraging advanced data-driven techniques and scalable analysis frameworks. Chung-Han Chou, Chih-Jen Hsu, Hung-Chun Chiu, Kai-Chiang Wu, Yu-Guang Chen, Zhuo Li 0001 |
ICCAD | 4 |
| 2025 | Palu: KV-Cache Compression with Low-Rank ProjectionabstractPost-training KV-Cache compression methods typically either sample a subset of effectual tokens or quantize the data into lower numerical bit width. However, these methods cannot exploit redundancy in the hidden dimension of the KV tenors. This paper presents a hidden dimension compression approach called Palu, a KV-Cache compression framework that utilizes low-rank projection to reduce inference-time LLM memory usage. Palu decomposes the linear layers into low-rank matrices, caches compressed intermediate states, and reconstructs the full keys and values on the fly. To improve accuracy, compression rate, and efficiency, Palu further encompasses (1) a medium-grained low-rank decomposition scheme, (2) an efficient rank search algorithm, (3) low-rank-aware quantization compatibility enhancements, and (4) an optimized GPU kernel with matrix fusion. Extensive experiments with popular LLMs show that Palu compresses KV-Cache by 50% while maintaining strong accuracy and delivering up to 1.89× speedup on the RoPE-based attention module. When combined with quantization, Palu’s
inherent quantization-friendly design yields small to negligible extra accuracy degradation while saving additional memory than quantization-only methods and achieving up to 2.91× speedup for the RoPE-based attention. Moreover, it maintains comparable or even better accuracy (up to 1.19 lower perplexity) compared to quantization-only methods. These results demonstrate Palu’s superior capability to effectively address the efficiency and memory challenges of LLM inference posed by KV-Cache. Our code is publicly available at: https://github.com/shadowpa0327/Palu. Chi-Chih Chang, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, Mohamed S. Abdelfattah, Kai-Chiang Wu |
ICLR | 10 |
| 2025 | Quamba: A Post-Training Quantization Recipe for Selective State Space ModelsabstractState Space Models (SSMs) have emerged as an appealing alternative to Transformers for large language models, achieving state-of-the-art accuracy with constant memory complexity which allows for holding longer context lengths than attention-based networks. The superior computational efficiency of SSMs in long sequence modeling positions them favorably over Transformers in many scenarios. However, improving the efficiency of SSMs on request-intensive cloud-serving and resource-limited edge applications is still a formidable task. SSM quantization is a possible solution to this problem, making SSMs more suitable for wide deployment, while still maintaining their accuracy. Quantization is a common technique to reduce the model size and to utilize the low bit-width acceleration features on modern computing units, yet existing quantization techniques are poorly suited for SSMs. Most notably, SSMs have highly sensitive feature maps within the selective scan mechanism (i.e., linear recurrence) and massive outliers in the output activations which are not present in the output of token-mixing in the self-attention modules. To address this issue, we propose a static 8-bit per-tensor SSM quantization method which suppresses the maximum values of the input activations to the selective SSM for finer quantization precision and quantizes the output activations in an outlier-free space with Hadamard transform. Our 8-bit weight-activation quantized Mamba 2.8B SSM benefits from hardware acceleration and achieves a 1.72 $\times$ lower generation latency on an Nvidia Orin Nano 8G, with only a 0.9\% drop in average accuracy on zero-shot tasks. When quantizing Jamba, a 52B parameter SSM-style language model, we observe only a $1\%$ drop in accuracy, demonstrating that our SSM quantization method is both effective and scalable for large language models, which require appropriate compression techniques for deployment. The experiments demonstrate the effectiveness and practical applicability of our approach for deploying SSM-based models of all sizes on both cloud and edge platforms. Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Diana Marculescu |
ICLR | 4 |
| 2025 | Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space ModelsabstractState Space Models (SSMs) are gaining attention as an efficient alternative to Transformers due to their constant memory complexity and comparable performance. Yet, deploying large-scale SSMs on cloud-based services or resource-constrained devices faces challenges. To address this, quantizing SSMs using low bit-width data types is proposed to reduce model size and leverage hardware acceleration. Given that SSMs are sensitive to quantization errors, recent advancements focus on quantizing a specific model or bit-width to improve their efficiency while maintaining performance. However, different bit-width configurations, such as W4A8 for cloud service throughput and W4A16 for improving question-answering on personal devices, are necessary for specific scenarios.
To this end, we present Quamba2, compatible with \textbf{W8A8}, \textbf{W4A8}, and \textbf{W4A16} for both \textbf{Mamba} and \textbf{Mamba2}, addressing the rising demand for SSM deployment across various platforms. We propose an offline approach to quantize inputs of a linear recurrence in 8-bit by sorting and clustering for $x$, combined with a per-state-group quantization for $B$ and $C$. To ensure compute-invariance in the SSM output, we offline rearrange weights according to the clustering sequence. The experiments show Quamba2-8B outperforms several state-of-the-art SSMs quantization methods and delivers 1.3$\times$ and 3$\times$ speedup in the pre-filling and generation stages and 4$\times$ memory reduction with only a $1.6$% accuracy drop on average. The code and quantized models will be released at: Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Mohamed S. Abdelfattah, Diana Marculescu |
ICML | 4 |
| 2025 | Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative DecodingabstractThe immense model sizes of large language models (LLMs) challenge deployment on memory-limited consumer GPUs.
Although model compression and parameter offloading are common strategies to address memory limitations, compression can degrade quality, and offloading maintains quality but suffers from slow inference.
Speculative decoding presents a promising avenue to accelerate parameter offloading, utilizing a fast draft model to propose multiple draft tokens, which are then verified by the target LLM in parallel with a single forward pass. This method reduces the time-consuming data transfers in forward passes that involve offloaded weight transfers.
Existing methods often rely on pretrained weights of the same family, but require additional training to align with custom-trained models. Moreover, approaches that involve draft model training usually yield only modest speedups. This limitation arises from insufficient alignment with the target model, preventing higher token acceptance lengths.
To address these challenges and achieve greater speedups, we propose SubSpec, a plug-and-play method to accelerate parameter offloading that is lossless and training-free. SubSpec constructs a highly aligned draft model by generating low-bit quantized substitute layers from offloaded target LLM portions. Additionally, our method shares the remaining GPU-resident layers and the KV-Cache, further reducing memory overhead and enhance alignment.
SubSpec achieves a high average acceptance length, delivering 9.1$\times$ speedup for Qwen2.5 7B on MT-Bench (8GB VRAM limit) and an average of 12.5$\times$ speedup for Qwen2.5 32B on popular generation benchmarks (24GB VRAM limit). Pei-Shuo Wang, Jian-Jia Chen, Chun-Che Yang, Chi-Chih Chang, Ning-Chi Huang, Mohamed S. Abdelfattah, Kai-Chiang Wu |
NeurIPS | 7 |
| 2024 | Wafer-View Defect-Pattern-Prominent GDBN Method Using MetaFormer VariantabstractGood-Die-in-Bad-Neighborhood (GDBN) is a technique employed to identify chips that pass initial tests but may have defects. Previous research used neural networks and expanded observation windows but ignored the impact of isolated dice. This paper improves wafer pattern information through denoising and creates a lightweight model. It also reduces training time by annotating multiple dice simultaneously. Experiments on real-world datasets show the model effectively captures more Test Escapes, reducing Defective Parts Per Million (DPPM) and improving return merchandise authorization gains. Shu-Wen Li, Chia-Heng Yen, Shuo-Wen Chang, Ying-Hua Chu, Kai-Chiang Wu, Mango Chia-Tso Chao |
ITC | 5 |
| 2024 | Transformer and Its Variants for Identifying Good Dice in Bad NeighborhoodsabstractGood-die-in-bad-neighborhood (GDBN) is a widely adopted method utilizing the fact that manufacturing defects tend to exhibit spatial dependency and form a cluster or specific pattern of bad dice on a wafer. Existing research studies on GDBN mainly focus on learning such spatial relationships within a limited observation window through simple mechanisms such as linear regression or multilayer perceptron model. In this paper, we propose MetaFormer-GDBN, a transformer-based deep learning model with the observation window extending to the entire wafer to include broader pattern information. The enhanced neighboring information and model capacity allow our method to capture more complex patterns of bad dice. Experiments show that compared to previous work, our method can achieve up to 50 % performance improvement, reducing the DPPM (defective parts per million) with minimal yield loss. Cheng-Che Lu, Chi-Chih Chang, Chia-Heng Yen, Shuo-Wen Chang, Ying-Hua Chu, Kai-Chiang Wu, Mango Chia-Tso Chao |
VTS | 6 |
| 2024 | FLORA: Fine-grained Low-Rank Architecture Search for Vision TransformerabstractVision Transformers (ViT) have recently demonstrated success across a myriad of computer vision tasks. However, their elevated computational demands pose significant challenges for real-world deployment. While low-rank approximation stands out as a renowned method to reduce computational loads, efficiently automating the target rank selection in ViT remains a challenge. Drawing from the notable similarity and alignment between the processes of rank selection and One-Shot NAS, we introduce FLORA, an end-to-end automatic framework based on NAS. To overcome the design challenge of supernet posed by vast search space, FLORA employs a low-rank aware candidate filtering strategy. This method adeptly identifies and eliminates underperforming candidates, effectively alleviating potential undertraining and interference among subnetworks. To further enhance the quality of low-rank supernets, we design a low-rank specific training paradigm. First, we propose weight inheritance to construct supernet and enable gradient sharing among low-rank modules. Secondly, we adopt low-rank aware sampling to strategically allocate training resources, taking into account inherited information from pre-trained models. Empirical results underscore FLORA’s efficacy. With our method, a more fine-grained rank configuration can be generated automatically and yield up to 33% extra FLOPs reduction compared to a simple uniform configuration. More specific, FLORA-DeiT-B/FLORA-Swin-B can save up to 55%/42% FLOPs almost without performance degradtion. Importantly, FLORA boasts both versatility and orthogonality, offering an extra 21%-26% FLOPs reduction when integrated with leading compression techniques or compact hybrid structures. Our code is publicly available at https://github.com/shadowpa0327/FLORA. Chi-Chih Chang, Yuan-Yao Sung, Shixing Yu, Ning-Chi Huang, Diana Marculescu, Kai-Chiang Wu |
WACV | 6 |
| 2024 | Reliability Engineering in a Time of Rapidly Converging TechnologiesabstractThe convergence of technologies is happening across various aspects, such as communication, computing, medicine, and transportation. The smartphone is a perfect example of convergence, packing features, such as a camera, GPS, artificial intelligence, and Internet connectivity into one sleek device. Autonomous driving is another good example. In a time of rapidly converging technologies, reliability engineering must take into account the potential for cyber threats, the need for cyber trust, the importance of cyber security, and the criticality of cyber resilience. In this way, reliability engineers can ensure the confidentiality, integrity, and availability of computer systems and networks in the face of evolving threats and changing technologies. In this article, we introduce the challenges and current progress of reliability engineering in emerging technologies, including practices and applications of cyber trust and security, AI-empowered autonomous driving systems, modern mobile networks, blockchains and distributed ledger technologies, prognostic and health management, integrated circuit and hardware, and enterprise cybersecurity and threat hunting. Shiuh-Pyng Shieh, Jeffrey M. Voas, Phillip A. Laplante, Jason W. Rupe, Christian K. Hansen, Yu-Sung Wu, Yi-Ting Chen 0001, Chi-Yu Li 0001, Kai-Chiang Wu |
IEEE Trans. Reliab. | 9 |
| 2023 | Enhancing Good-Die-in-Bad-Neighborhood Methodology with Wafer-Level Defect Pattern InformationabstractIn semiconductor manufacturing processes, there are several causes of typical defects in silicon wafers, such as operational flaws or equipment malfunctions, which may lead to circuit failure and defective products. Therefore, testing is instrumental in improving overall yield and reliability. GDBN (good die in bad neighborhood) is a widely-used technique of rejecting potentially defective wafers in advance based on the concept that defects tend to cluster together. However, previous studies related to GDBN are limited to a local observation by using a narrow-sighted window and thus ignore the defects patterns of the wafers. In this paper, by leveraging information of wafer defect patterns and extending the observation range to the entire wafer, we strengthen the GDBN method to recognize the potentially defective dice more effectively based on the feature of the different defect patterns. The method proposed in this paper is realized by convolutional neural network technology, and it is also the first method to consider defect patterns of wafers as features to solve the problem of GDBN. Several experiments are conducted on a real-world WM-811K dataset, and the results show that our proposed method not only reduce the cost of return merchandise authorization (RMA) but the DPPM (Defective Parts Per Million) more significantly over other existing methods. Ching-Min Liu, Chia-Heng Yen, Shu-Wen Lee, Kai-Chiang Wu, Mango Chia-Tso Chao |
ITC | 4 |
| 2023 | Outlier Detection for Analog Tests Using Deep Learning TechniquesabstractWith the increasing demand for high reliability of products, how to prevent potential defective devices from shipping to customers is a serious issue about which more and more companies are concerned. Toward this end, many test methods have been developed to screen out outliers. However, basic statistical paradigm may not be enough to handle the shrinking transistor size and increasingly complex circuit design. In this paper, we propose to use the concept of Z-score derived from our proposed neural network, called single density network (SDN), to define level of abnormality. We also define new metrics called self-excluded fail rate (SE fail rate) and normalized area under curve (AUC) to be our criteria to quantify and further visualize the outcome. To filter out spatially-correlated outliers, we make use of specific information of neighboring dice and encode them into our input features for the proposed SDN. A series of experimental results on industrial data reveal the effectiveness of our methodology and the better ability to identify defective outliers than existing conventional statistical approaches for a variety of analog tests. Chin-Kuan Lin, Cheng-Che Lu, Shuo-Wen Chang, Ying-Hua Chu, Kai-Chiang Wu, Mango Chia-Tso Chao |
VTS | 5 |
| 2023 | Test Generation for Defect-Based Faults of Scan Flip-FlopsabstractWhen testing scan flip-flops (SFFs), chain test is first applied to ensure the functionality of scan chains and to detect the majority of stuck-at (SA) and transition delay (TD) faults along scan paths. However, there still exist some defects inside scan cells that cannot be effectively detected by chain test or conventional SA and TD patterns. This paper presents five cell-aware (CA) fault models to explicitly target the defects inside scan flip-flops. The proposed static shift (SS) and dynamic shift (DS) faults identify the defects detectable by chain test. For the defects escaping chain test, static single-capture (SSC) faults target the defects detectable when SFFs are in one-cycle capture mode, while static double-capture (SDC) and dynamic double-capture (DDC) faults target those detectable when SFFs are in two-cycle capture mode. The identified CA faults of SFFs are output in a format compatible with a commercial ATPG tool for pattern generation. Experimental results on large IWLS05 benchmarks demonstrate that our proposed faults cannot be fully covered by conventional SA and TD patterns and hence require dedicated test patterns to detect. Yu-Teng Nien, Chen-Hong Li, Pei-Yin Wu, Yung-Jheng Wang, Kai-Chiang Wu, Mango Chia-Tso Chao |
VTS | 5 |
| 2023 | CNN-Based Stochastic Regression for IDDQ Outlier IdentificationabstractTo reduce defect parts per million (DPPM) on IC products, IDDQ testing can be exploited for identifying the outliers which are potentially defective but not detected by sign-off functional and parametric tests. Conventional IDDQ testing paradigms depending on a simple statistical$6\sigma $rule or engineers’ experience are usually too conservative to effectively identify nontrivial outliers, especially, when spatial correlations are of great concern/influence. In article, an improved convolutional neural network (CNN)-based method can be proposed for IDDQ outlier identification. In the proposed method, the mean and the standard deviation on the IDDQ value inside a die under test (DUT) can be predicted by employing a stochastic regression model. According to the predicted mean and standard deviation, we derive an expected IDDQ interval and identify the DUT as an outlier if its actual measured IDDQ value is beyond the expected interval. From the observation of the experimental results, the improved data preprocessing and the improved CNN-based stochastic regression can be contained to enhance the prediction accuracy of the expected IDDQ intervals. In the improved method, the spatial correlations of the neighboring dice inside a window can be considered by training a CNN-based stochastic regression model with a large volume of industrial data on 28 and 65 nm products. The trained model is highly accurate prediction in the$R^{2}$(0.973) and RMSE (0.626 mA) of the expected IDDQ values on 28 nm product and the$R^{2}$(0.942) and RMSE (2.155 uA) of the expected IDDQ values on 65 nm product. Furthermore, the experimental results show that the trained model can capture the potential defective dice by identifying efficient IDDQ outliers. Chia-Heng Yen, Chun-Teng Chen, Cheng-Yen Wen, Ying-Yen Chen, Jih-Nung Lee, Shu-Yi Kao, Kai-Chiang Wu, Mango Chia-Tso Chao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Improving Cell-Aware Test for Intra-Cell Short DefectsabstractConventional fault models define their faulty behavior at the IO ports of standard cells with simple rules of fault activation and fault propagation. However, there still exist some defects inside a cell (intra-cell) that cannot be effectively detected by the test patterns of conventional fault models and hence become a source of DPPM. In order to further increase the defect coverage, many research works have been conducted to study the fault models resulting from different types of intra-cell defects, by SPICE-simulating each targeted defect with its equivalent circuit-level defect model. In this paper, we propose to improve cell-aware (CA) test methodology by concentrating on intracell bridging faults due to short defects inside standard cells. The faults extracted are based on examining the actual physical proximity of polygons in the layout of a cell, and are thus more realistic and reasonable than those (faults) determined by RC extraction. Experimental results on a set of industrial designs show that the proposed methodology can indeed improve the test quality of intra-cell bridging faults. On average, 0.36 % and 0.47% increases in fault coverage can be obtained for 1-time-frame and 2-time-frame CA tests, respectively. In addition to short defects between two metal polygons, short defects among three metal polygons are also considered in our methodology for another 9.33 % improvement in fault coverage. Dong-Zhen Lee, Ying-Yen Chen, Kai-Chiang Wu, Mango Chia-Tso Chao |
DATE | 3 |
| 2022 | Rule Generation for Classifying SLT Failed PartsabstractSystem-level test (SLT) has recently gained visibility when integrated circuits become harder and harder to be fully tested due to increasing transistor density and circuit design complexity. Albeit SLT is effective for reducing test escapes, little diagnostic information can be obtained for product improvement. In this paper, we propose an unsupervised learning (UL) method to resolve the aforementioned issue by discovering correlative, potentially systematic defects during the SLT phase. Toward this end, HDBSCAN [1] is used for clustering SLT failed devices in a low-dimensional space created by UMAP [2]. Decision trees are subsequently applied to explain the HDBSCAN results based on generating explainable quantitative rules, e.g., inequality constraints, providing domain experts additional information for advanced diagnosis. Experiments on industrial data demonstrate that the proposed methodology can effectively cluster SLT failed devices and then explain the clustering results with a promising accuracy of above 90%. Our methodology is also scalable and fast, requiring two to five orders of magnitude lower runtime than the method presented in [3]. Ho-Chieh Hsu, Cheng-Che Lu, Shih-Wei Wang, Kelly Jones, Kai-Chiang Wu, Mango Chia-Tso Chao |
VTS | 5 |
| 2022 | Methodology of Generating Timing-Slack-Based Cell-Aware TestsabstractIn order to reduce defect parts per million, cell-aware (CA) methodology was proposed to cover various types of intracell defects. In this article, we present a novel methodology for generating 2-time-frame (2tf) CA tests based on timing slack analysis. The proposed 2tf CA fault model, aware of timing slack and named TS, defines a fault: 1) on a cell instance basis and 2) based on per-instance timing criticality (according to timing slack). By comparing the derived extra delay against the timing slack of the cell instance, a delay fault can be defined, and according to its severity, the fault can be further classified into small-delay fault or gross-delay fault. In contrast to prior 2tf CA methodology that is on a cell (rather than cell instance) basis and unaware of timing criticality/slack, our methodology can identify “more realistic” faults which really need to be considered, and potentially the cost/effort for testing those 2tf CA faults can be reduced. We also propose a test quality metric, timing slack defect coverage (TSDC), to measure the effectiveness of automatic test pattern generation (ATPG) tests in terms of the ability to detect small-delay TS defects along long paths. Experimental results on a set of 22-nm industrial designs demonstrate that, due to more realistic fault identification, the number of identified small-delay faults can be reduced by 56.8%. With the slack-based ATPG for testing small-delay faults along long paths, TS can reduce the number of test patterns by 33.1% while achieving 0.49% higher TSDC, compared with the results of prior 2tf CA methodology. Yu-Teng Nien, Kai-Chiang Wu, Dong-Zhen Lee, Ying-Yen Chen, Po-Lin Chen, Mason Chern, Jih-Nung Lee, Shu-Yi Kao, Mango Chia-Tso Chao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Test Methodology for Defect-Based Bridge FaultsabstractA defect-based bridge fault represents the faulty behavior of an interconnect short defect obtained by SPICE simulating the two shorted cells with the short defect injected. In this article, we have developed a framework to automatically extract defect-based bridge faults and utilize commercial automatic test pattern generation (ATPG) to generate corresponding test patterns for a given design. A defect-based bridge fault model can not only describe the faulty behavior of a short defect precisely but also result in collapsible faults at one shorted cell pair. As a result, using a defect-based bridge fault model for ATPG can lead to a significantly smaller bridge-fault test set when compared with a conventional four-way dominance bridge fault model, where four noncollapsible faults at one shorted cell pair are considered for ATPG. In addition, some short defects can only be detected by the test set for defect-based bridge faults but not by the test set for four-way dominance bridge faults with more test patterns. The runtime required for extracting 1-time-frame (1tf) defect-based bridge faults has been proven acceptable on industrial designs and some techniques were also proposed to speed up the runtime for extracting 2tf defect-based bridge faults. All experiments in this article are conducted based on industrial designs. Shuo-Wen Chang, Yu-Teng Nien, Yu-Pang Hu, Kai-Chiang Wu, Chi Chun Wang, Fu-Sheng Huang, Yi-Lun Tang, Yung-Chen Chen, Ming-Chien Chen, Mango Chia-Tso Chao |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Timing Variability-Aware Analysis and Optimization for Variable-Latency DesignsabstractCircuit performance has been the key design constraint for over a decade. Variable-latency design (VLD) paradigm was proposed for optimizing the overall performance in terms of throughput. In addition, process variations (PVs) and aging effects manifest themselves as gate delay shifts, which in turn cause variability of circuit timing (timing variability). Required for dealing with the impact of timing variability better, detailed evaluation and analysis of circuit timing for VLD are actually not straightforward. In this article, we present a systematic methodology for analyzing a VLD circuit and identifying critical one-cycle and two-cycle paths/gates. Based on the criticality analysis, a gate sizing framework using particle swarm optimization (PSO) is proposed. Our objective is, in a less pessimistic fashion, to make constructed VLD circuits better (less vulnerable to timing variability). The experimental results show that the proposed framework can generate extra timing margins for VLD, such that process-induced error rate can be reduced to 0%–0.1% under 10% variability in gate delay. On average, an extra timing margin of 11.48% can be obtained without lengthening the clock period, and 1.52% area can be reduced simultaneously. Ning-Chi Huang, Chao-Wei Cheng, Kai-Chiang Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | An Energy-Efficient Approximate Systolic Array Based on Timing Error Prediction and PreventionabstractDeep neural networks (DNNs) have achieved out-standing accuracy on machine learning applications. However, the numbers of parameters and computational costs of DNNs have grown dramatically. To accelerate the numerous matrix multiplication operations in DNNs, a systolic array of multiply and-accumulate units (MACs) is a widely-used architecture. In this paper, both timing error prediction and approximate computing are leveraged to relax the timing constraints of MACs. Afterwards, voltage underscaling is applied to further enhance the energy efficiency of the systolic array. In the experiments, our proposed approximate systolic array can obtain 36% energy reduction with only 1% accuracy loss for CFAR-10 image classification. Ning-Chi Huang, Wei-Kai Tseng, Huan-Jan Chou, Kai-Chiang Wu |
VTS | 4 |
| 2021 | Identifying Good-Dice-in-Bad-Neighborhoods Using Artificial Neural NetworksabstractGDBN (good die in bad neighborhood) methodology has been regarded as an effective technique for reducing DPPM (defect parts per million), by identifying and rejecting suspicious dice even though they test good. Instead of examining eight immediate neighbors or exploiting simple linear regression, in this paper we propose to employ a window of larger size for broad-sighted recognition of neighborhood, and make best use of the larger window for accurate prediction of the suspicious level for any given die. The proposed methodology is realized by using an artificial neural network (NN), and is a breakthrough of NN-based work for solving the problem of GDBN. Various experiments on two sets of data clearly reveal the superiority of our NN-based methodology over other existing methods. Besides reducing DPPM, our methodology is able to achieve 1. 5X-2X better reduction in the cost for return merchandise authorization (RMA). Cheng-Hao Yang, Chia-Heng Yen, Ting-Rui Wang, Chun-Teng Chen, Mason Chern, Ying-Yen Chen, Jih-Nung Lee, Shu-Yi Kao, Kai-Chiang Wu, Mango Chia-Tso Chao |
VTS | 9 |
| 2020 | Selective Sensor Placement for Cost-Effective Online Aging Monitoring and ResilienceabstractAggressive technology scaling trends, such as thinner gate oxide without proportional downscaling of supply voltage, aggravate the aging impact and thus necessitate an aging-aware reliability verification and optimization framework during early design stages. In this paper, we propose a novel in-situ sensing strategy based on deploying transition detectors (TDs), for on-chip aging monitoring and resilience. Transformed into the set cover problem and then formulated into maximum satisfiability, the proposed problem of TD/sensor placement can be solved efficiently. Experimental results show that, by introducing at most 2.2% area overhead (for TD/sensor placement), the aging behavior of a target circuit can be effectively monitored, and the correctness of its functionality can be perfectly guaranteed with an average of 77% aging resilience achieved. In other words, with 2.2% area overhead, potential aging-induced timing errors can be detected and then eliminated, while achieving 77% recovery from aging-induced performance degradation. Hao-Chun Chang, Li-An Huang, Kai-Chiang Wu, Yu-Guang Chen |
ISPD | 3 |
| 2020 | Test Methodology for Defect-based Bridge FaultsabstractA defect-based bridge fault represents the faulty behavior of an interconnect short defect obtained by SPICE-simulating the two shorted cells with the short defect injected. In this paper, we have developed a framework to automatically extract defect-based bridge faults and utilize commercial ATPG to generate corresponding test patterns for a given design. Defect-based bridge fault model can not only describe the faulty behavior of a short defect precisely but also result in collapsible faults at one shorted cell pair. As a result, using defect-based bridge fault model for ATPG can lead to a significantly smaller bridge-fault test set when compared to conventional 4-way dominance bridge fault model, where four non-collapsible faults at one shorted cell pair are considered for ATPG. Also, some short defects can only be detected by the test set for defect-based bridge faults but not by the test set for 4-way dominance bridge faults with more test patterns. The experimental result based on industrial designs has demonstrated the effectiveness of using defect-based bridge faults for ATPG while showing an affordable runtime on extracting defect-based bridge faults. Yu-Pang Hu, Shuo-Wen Chang, Kai-Chiang Wu, Chi Chun Wang, Fu-Sheng Huang, Yi-Lun Tang, Yung-Chen Chen, Ming-Chien Chen, Mango Chia-Tso Chao |
ITC-Asia | 3 |
| 2020 | CNN-based Stochastic Regression for IDDQ Outlier IdentificationabstractIn order to reduce DPPM (defect parts per million), IDDQ testing methodology can be exploited for identifying "outliers" which are potentially defective but not detected by signoff functional and parametric tests. Conventional IDDQ testing paradigms depending on a simple statistical 6σ rule or engineers’ experience are usually too conservative to effectively identify non-trivial outliers, especially when spatial correlations are of great concern/influence. In this paper, by employing a stochastic regression model, the mean as well as the variance of the IDDQ of a die under test (DUT) can be predicted. According to the predicted mean and variance, we derive an expected IDDQ range and identify the DUT as an outlier if its actual IDDQ measurement is beyond the expected range. The proposed stochastic regression model is obtained by training a convolutional neural network (CNN) and, based on its primitive property of convolutional kernel mapping with large volume of industrial data, spatial correlations (due to spatially-correlated process variations, etc) can be considered/captured. The trained data-driven CNN is highly accurate in terms of R-square (0.958) and RMSE (0.783), and the percentage of identified outliers (0.047%) is very close to the theoretical reference (0.050%), which validates the efficacy of our proposed methodology. Chun-Teng Chen, Chia-Heng Yen, Cheng-Yen Wen, Cheng-Hao Yang, Kai-Chiang Wu, Mason Chern, Ying-Yen Chen, Chun-Yi Kuo, Jih-Nung Lee, Shu-Yi Kao, Mango Chia-Tso Chao |
VTS | 5 |
| 2020 | Making Aging Useful by Recycling Aging-induced Clock SkewabstractDevice aging, which causes significant loss on circuit performance and lifetime, has been a primary factor in reliability degradation of nanoscale designs. In this article, we propose to take advantage of aging-induced clock skews (i.e., make them useful for aging tolerance) by manipulating and recycling these time-varying skews to compensate for the performance degradation of logic networks. The goal is to assign achievable/reasonable aging-induced clock skews in a circuit, such that its effective performance degradation due to aging can be tolerated. On average, 21.21% aging tolerance can be achieved with insignificant design overhead. Moreover, we employ V th assignment on clock buffers to further tolerate the aging-induced degradation of logic networks. When V th assignment is applied on top of aforementioned aging manipulation, the average aging tolerance can be enhanced to 29.15%. Tien-Hung Tseng, Chung-Han Chou, Kai-Chiang Wu |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | Aging-aware chip health prediction adopting an innovative monitoring strategyabstractConcerns exist that the reliability of chips is worsening because of downscaling technology. Among various reliability challenges, device aging is a dominant concern because it degrades circuit performance over time. Traditionally, runtime monitoring approaches are proposed to estimate aging effects. However, such techniques tend to predict and monitor delay degradation status for circuit mitigation measures rather than the health condition of the chip. In this paper, we propose an aging-aware chip health prediction methodology that adapts to workload conditions and process, supply voltage, and temperature variations. Our prediction methodology adopts an innovative on-chip delay monitoring strategy by tracing representative aging-aware delay behavior. The delay behavior is then fed into a machine learning engine to predict the age of the tested chips. Experimental results indicate that our strategy can obtain 97.40% accuracy with 4.14% area overhead on average. To the authors' knowledge, this is the first method that accurately predicts current chip age and provides information regarding future chip health. Yun-Ting Wang, Kai-Chiang Wu, Chung-Han Chou, Shih-Chieh Chang 0001 |
ASP-DAC | 2 |
| 2019 | Sensor-Based Approximate Adder Design for Accelerating Error-Tolerant and Deep-Learning ApplicationsabstractApproximate computing is an emerging strategy which trades computational accuracy for computational cost in terms of performance, energy, and/or area. In this paper, we propose a novel sensor-based approximate adder for highperformance energy-efficient arithmetic computation, while considering the accuracy requirement of error-tolerant applications. This is the first work using in-situ sensors for approximate adder design, based on monitoring online transition activity on the carry chain and speculating on carry propagation/truncation. On top of a fully-optimized ripple-carry adder, the performance of our adder is enhanced by 2.17X. When applied in error-tolerant applications such as image processing and handwritten digit recognition, our approximate adder leads to very promising quality of results compared to the case when an accurate adder is used. Ning-Chi Huang, Szu-Ying Chen, Kai-Chiang Wu |
DATE | 3 |
| 2019 | ICE-RADAR: In-situ, Cost-Effective Razor Flip-Flop Deployment for Aging ResilienceabstractDevice aging, which causes significant loss on circuit performance and lifetime, has been a primary factor in reliability degradation of nanoscale designs. In this paper, we propose to exploit timing speculation for aging resilience, based on deploying Razor flip-flops. By formulating the problem based on Boolean satisfiability, we can determine the optimal deployment of Razor flip-flops, such that maximum degree of aging resilience can be achieved in a cost-effective manner. Experimental results show that more than 50% of aging-induced performance degradation can be recovered, while reducing the number of required Razor flip-flops by more than 3X, as compared to the case of naive Razor flip-flop deployment. Kai-Chiang Wu, Wei-Tao Huang, Chiao-Yang Huang |
IOLTS | 1 |
| 2019 | Methodology of Generating Timing-Slack-Based Cell-Aware TestsabstractIn order to reduce DPPM (defect parts per million), cell-aware (CA) methodology was proposed to cover various types of intra-cell defects. The resulting CA faults can be a 1-time-frame (1tf) or 2-time-frame (2tf) fault, and 2tf CA tests were experimentally verified to be capable of catching a significant number of defective parts not covered by other conventional tests. In this paper, we present a novel methodology for generating 2tf CA tests based on timing slack analysis. The proposed 2tf CA fault model, aware of timing slack and named TS, defines a fault (i) on a cell instance basis, and (ii) based on per-instance timing criticality (according to timing slack). More explicitly, for each cell instance with a specific defect injected, we check its output capacitive load and derive the corresponding extra delay. By comparing the extra delay against timing slack of the cell instance, a delay fault can be defined, and according to its severity, the fault can be further classified into small-delay fault or gross-delay fault. In contrast to prior 2tf CA methodology that is on a cell (rather than cell instance) basis and unaware of timing criticality/slack, our methodology can identify “more realistic” faults which really need to be considered, and potentially the cost/effort for testing those 2tf CA faults can be reduced. Experimental results on a set of 28nm industrial designs demonstrate that, due to more realistic fault identification, the numbers of identified small-delay faults and corresponding test patterns to be applied can be reduced by 35.1% and 24.1% respectively, leading to 40.7% reduction in the runtime of ATPG. Yu-Teng Nien, Kai-Chiang Wu, Dong-Zhen Lee, Ying-Yen Chen, Po-Lin Chen, Mason Chern, Jih-Nung Lee, Shu-Yi Kao, Mango Chia-Tso Chao |
ITC | 2 |
| 2019 | Layout-Based Dual-Cell-Aware TestsabstractConventional fault models define their faulty behavior at the IO ports of standard cells with simple rules of fault activation and fault propagation. However, there still exist some defects inside a cell (intra-cell) or between two cells (dual-cell) that cannot be effectively detected by the test patterns of conventional fault models and hence become a source of DPPM. In order to further increase the defect coverage, many research works have been conducted to study the fault models resulting from different types of intra-cell and dual-cell defects, by SPICE-simulating each targeted defect with its equivalent circuit-level defect model. However, it was considered computationally infeasible to simulate every possible defective scenario for a cell library and obtain a complete set of cell-level fault models. In this paper, we present a new dual-cell-aware (DCA) framework based on examining the layout of two adjacent cells (i.e., a dual cell) to identify potential defects, where time-consuming RC extraction can be avoided and the runtime for SPICE simulation can be reduced. Experimental results and silicon data on a SoC product show that the proposed DCA framework can not only save runtime significantly but also maintain the promising efficacy of DCA tests for the objective of lowering DPPM. Tse-Wei Wu, Dong-Zhen Lee, Mango Chia-Tso Chao, Kai-Chiang Wu, Shu-Yi Kao, Ying-Yen Chen, Po-Lin Chen, Mason Chern, Jih-Nung Lee |
VTS | 5 |
| 2018 | Lifetime Reliability Trojan Based on Exploring Malicious AgingabstractDue to escalating complexity of hardware design and manufacturing, integrated circuits (ICs) are designed and fabricated in multiple nations, so are software tools. It makes hardware security become more subject to various kinds of tampering in the supply chain. Hardware Trojan horses (HTHs) can be implanted to facilitate the leakage of confidential information or cause the failure of a system. Reliability Trojan is one of the main categories of HTH attacks because its behavior is progressive and thus hard to be detected, or not considered malicious. In this work, we propose to insert reliability Trojan into a circuit which can finely control the circuit lifetime as specified by attackers (or even designers), based on manipulating BTI-induced aging behavior in a statistical manner, with the consideration of process variations (PVs). Experimental results show that, given a specified lifetime target and under the influence of PVs, the circuit is highly likely to fail within a desired lifetime interval, at the cost of little area overhead. Tien-Hung Tseng, Shou-Chun Li, Kai-Chiang Wu |
ATS | 3 |
| 2018 | MAUI: Making aging useful, intentionallyabstractDevice aging, which causes significant loss on circuit performance and lifetime, has been a primary factor in reliability degradation of nanoscale designs. In this paper, we propose to take advantage of aging-induced clock skews (i.e., make them useful for aging tolerance) by manipulating these time-varying skews to compensate for the performance degradation of logic networks. The goal is to assign achievable/reasonable aging-induced clock skews in a circuit, such that its overall performance degradation due to aging can be minimized, that is, the lifespan can be maximized. On average, 25% aging tolerance can be achieved with insignificant design overhead. Kai-Chiang Wu, Tien-Hung Tseng, Shou-Chun Li |
DATE | 1 |
| 2018 | Sensor-Based Time Speculation in the Presence of Timing VariabilityabstractTime speculation has been widely used to achieve high performance in modern design as it exploits average-case timing optimization instead of worst-case timing optimization focusing on reducing longest path delay which rarely happens. Variable-latency design (VLD) style is one research category of time speculation. Since process and environmental variations are hard to predict, traditional variable-latency units (VLUs) designed at presilicon stage will suffer significant performance loss due to pessimistic assumptions for addressing variations. In this paper, we propose a novel sensor-based, transition-aware VLU (S-VLU) scheme adapting to process-voltage-temperature (PVT) variations by using in situ sensors to obtain real-time transition information in a circuit. Moreover, we also propose a sensor deployment strategy to achieve near-maximal performance gain. On average, the S-VLU achieves a 31.27% performance improvement as compared to a 19.26% improvement by using traditional HL. The area overhead of the S-VLU is 13.48%. To the best of the authors' knowledge, this is the first wok to address PVT variations in VLD style. Chung-Han Chou, Tsui-Yun Chang, Kai-Chiang Wu, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Analysis and optimization of variable-latency designs in the presence of timing variabilityabstractCircuit performance has been the key design constraint for over a decade. Variable-latency design (VLD) paradigm was proposed for optimizing the overall performance in terms of throughput. In addition, process variations and aging effects manifest themselves as gate delay shifts, and in turn cause variability of circuit timing (timing variability). Required for dealing with the impact of timing variability better, detailed evaluation and analysis of circuit timing for VLD are actually not straightforward. In this paper, we present a systematic methodology for analyzing a VLD circuit, and identifying critical 1-cycle and 2-cycle paths/gates. Based on the criticality analysis, a gate sizing framework using particle swarm optimization (PSO) is proposed. Our objective is, in a less pessimistic fashion, making constructed VLD circuits better (less vulnerable to timing variability). The proposed framework is experimentally verified to be runtime-efficient and able to provide promising results. On average, an extra timing margin of 11% can be obtained without lengthening the clock period, and only 4% area overhead is introduced. Chang-Lin Tsai, Chao-Wei Cheng, Ning-Chi Huang, Kai-Chiang Wu |
DATE | 4 |
| 2017 | Fast WAT test structure for measuring Vt variance based on latch-based comparatorsabstractAs the technology of IC manufacturing continually scales down, process variations become more and more crucial than before. To statistically characterize local process variations, the traditional array-based test structure measures threshold voltage (Vt) for a sufficiently large number of devices-under-test (DUTs). However, the array-based test structure requires long time for DUT-by-DUT measurement; furthermore, it suffers from significant IR drop or leakage current due to the large number of DUTs, which results in the loss of measurement accuracy. In this paper, we present a novel sense-amplifier-based test structure that can monitor process variations based on rapid characterization of Vt variance, with marginal error of accuracy. A test-chip containing 120 NMOS and 120 PMOS DUTs has been implemented in 28nm CMOS process technology. Various experiments reveal promising efficiency and accuracy of the proposed test structure, for characterizing Vt variance. Kao-Chi Lee, Kai-Chiang Wu, Chih-Ying Tsai, Mango Chia-Tso Chao |
VTS | 2 |
| 2014 | BTI-Aware Sleep Transistor Sizing Algorithm for Reliable Power Gating DesignsabstractPower gating is an effective way to reduce leakage power. This technique uses high Vthtransistors, called sleep transistors, to turn off the power supply. However, sleep transistors suffer from the bias temperature instability (BTI) effect, resulting in an increased Vth, and reduced reliability. This paper proposes two BTI-aware sleep transistor sizing algorithms to reduce the total width of sleep transistors based on the distributed sleep transistor network structure. The proposed algorithms reduce total width by more than 16.08%. More area can be reduced if the BTI effect on both sleep and cluster transistors is considered. Kai-Chiang Wu, Ing-Chao Lin, Yao-Te Wang, Shuen-Shiang Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | NBTI and Leakage Reduction Using ILP-Based ApproachabstractWe propose an integer linear programming-based formulation to improve the effectiveness of the transmission gate-based technique intended to reduce negative-bias temperature instability and leakage power consumption. We also propose a virtual input pin technique to improve leakage reduction and use path sensitization to reduce area overhead. Simulation results show that combining these techniques can achieve >51.18% delay improvement and 63.34% leakage power improvement with only 2.31% area overhead. Ing-Chao Lin, Kuan-Hui Li, Chia-Hao Lin, Kai-Chiang Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Power-Planning-Aware Soft Error Hardening via Selective Voltage AssignmentabstractSoft errors, which have been a significant concern in memories, are now a main factor in reliability degradation of logic circuits. This paper presents a power-planning-aware methodology using dual supply voltages for soft error hardening. Given a constraint on power overhead, our proposed framework can minimize the soft error rate (SER) of a circuit via selective voltage assignment. In the 70-nm predictive technology model, circuit SER can be reduced by 23% on top of SER-aware gate resizing. For power-planning awareness, a bi-partitioning technique based on a simplified version of the Fiduccia-Mattheyses (FM) algorithm is presented. The simplified FM-based partitioning refines the result of selective voltage assignment by decreasing the number of connections across voltage islands, while maintaining the SER reduction that has been accomplished. Kai-Chiang Wu, Diana Marculescu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | A Low-Cost, Systematic Methodology for Soft Error Robustness of Logic CircuitsabstractDue to current technology scaling trends such as shrinking feature sizes and decreasing supply voltages, circuit reliability is becoming more susceptible to radiation-induced transient faults (soft errors). Soft errors, which have been a great concern in memories, are now a main factor in reliability degradation of logic circuits as well. In this paper, we present a systematic and integrated methodology for circuit robustness to soft errors. The proposed soft error rate (SER) reduction framework, based on redundancy addition and removal (RAR), aims at eliminating those gates with large contribution to the overall SER. Several metrics and constraints are introduced to guide the RAR-based approach toward SER reduction. Furthermore, we integrate a resizing strategy into our framework, as post-RAR additive SER optimization. The strategy can identify most critical gates to be upsized and thereby, minimize area and power overheads while maintaining a high level of soft error robustness. Experimental results show that the proposed RAR-based framework can achieve up to 70% reduction in output failure probability. On average, about 23% SER reduction is obtained with less than 4% area overhead. Kai-Chiang Wu, Diana Marculescu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Mitigating lifetime underestimation: A system-level approach considering temperature variations and correlations between failure mechanismsabstractLifetime (long-term) reliability has been a main design challenge as technology scaling continues. Time-dependent dielectric breakdown (TDDB), negative bias temperature instability (NBTI), and electromigration (EM) are some of the critical failure mechanisms affecting lifetime reliability. Due to the correlation between different failure mechanisms and their significant dependence on the operating temperature, existing models assuming constant failure rate and additive impact of failure mechanisms will underestimate the lifetime of a system, usually measured by mean-time-to-failure (MTTF). In this paper, we propose a new methodology which evaluates system lifetime in MTTF and relies on Monte-Carlo simulation for verifying results. Temperature variations and the correlation between failure mechanisms are considered so as to mitigate lifetime underestimation. The proposed methodology, when applied on an Alpha 21264 processor, provides less pessimistic lifetime evaluation than the existing models based on sum of failure rate. Our experimental results also indicate that, by considering the correlation of TDDB and NBTI, the lifetime of a system is likely not dominated by TDDB or NBTI, but by EM or other failure mechanisms. Kai-Chiang Wu, Ming-Chao Lee, Diana Marculescu, Shih-Chieh Chang 0001 |
DATE | 1 |
| 2011 | Aging-aware timing analysis and optimization considering path sensitizationabstractDevice aging, which causes significant loss on circuit performance and lifetime, has been a main factor in reliability degradation of nanoscale designs. Aggressive technology scaling trends, such as thinner gate oxide without proportional down-scaling of supply voltage, necessitate an aging-aware analysis and optimization flow during early design stages. Since only a small portion of critical and near-critical paths can be sensitized and may determine the circuit delay under aging, path sensitization should also be explicitly addressed for more accurate and efficient optimization. In this paper, we first investigate the impact of path sensitization on aging-aware timing analysis and then present a novel framework for aging-aware timing optimization considering path sensitization. By extracting and manipulating critical sub-circuits accounting for the effective circuit delay, our proposed framework can reduce aging-induced performance degradation to only 1.21% or one-seventh of the original performance loss with less than 2% area overhead. Kai-Chiang Wu, Diana Marculescu |
DATE | 1 |
| 2011 | Analysis and mitigation of NBTI-induced performance degradation for power-gated circuits
Kai-Chiang Wu, Diana Marculescu, Ming-Chao Lee, Shih-Chieh Chang 0001 |
ISLPED | 1 |
| 2010 | Clock skew scheduling for soft-error-tolerant sequential circuitsabstractSoft errors have been a critical reliability concern in nanoscale integrated circuits, especially in sequential circuits where a latched error can be propagated for multiple clock cycles and affect more than one output, more than once. This paper presents an analytical methodology for enhancing the soft error tolerance of sequential circuits. By using clock skew scheduling, we propose to minimize the probability of unwanted transient pulses being latched and also prevent latched errors from propagating through sequential circuits repeatedly. The overall methodology is formulated as a piecewise linear programming problem whose optimal solution can be found by existing mixed integer linear programming solvers. Experiments reveal that 30-40% reduction in the soft error rate for a wide range of benchmarks can be achieved. Kai-Chiang Wu, Diana Marculescu |
DATE | 1 |
| 2009 | Joint logic restructuring and pin reordering against NBTI-induced performance degradationabstractNegative Bias Temperature Instability (NBTI), a PMOS aging phenomenon causing significant loss on circuit performance and lifetime, has become a critical challenge for temporal reliability concerns in nanoscale designs. Aggressive technology scaling trends, such as thinner gate oxide without proportional downscaling of supply voltage, necessitate a design optimization flow considering NBTI effects at the early stages. In this paper, we present a novel framework using joint logic restructuring and pin reordering to mitigate NBTI-induced performance degradation. Based on detecting functional symmetries and transistor stacking effects, the proposed methodology involves only wire perturbation and introduces no gate area overhead at all. Experimental results reveal that, by using this approach, on average 56% of performance loss due to NBTI can be recovered. Moreover, our methodology reduces the number of critical transistors remaining under severe NBTI and thus, transistor resizing can be applied to further mitigate NBTI effects with low area overhead. Kai-Chiang Wu, Diana Marculescu |
DATE | 1 |
| 2008 | Soft error rate reduction using redundancy addition and removalabstractDue to current technology scaling trends such as shrinking feature sizes and reducing supply voltages, circuit reliability has become more susceptible to radiation-induced transient faults (soft errors). Soft errors, which have been a great concern in memories, are now a main factor in reliability degradation of logic circuits. In this paper, we propose a novel framework based on redundancy addition and removal (RAR) for soft error rate (SER) reduction. Several metrics and constraints are introduced to guide our proposed framework towards SER reduction in an efficient manner. Experimental results show that up to 70% reduction in output failure probability can be achieved with relatively low area overhead. Kai-Chiang Wu, Diana Marculescu |
ASP-DAC | 1 |
| 2008 | Process variability-aware transient fault modeling and analysisabstractDue to reduction in device feature size and supply voltage, the sensitivity of digital systems to transient faults is increasing dramatically. As technology scales further, the increase in transistor integration capacity also leads to the increase in process and environmental variations. Despite these difficulties, it is expected that systems remain reliable while delivering the required performance. Reliability and variability are emerging as new design challenges, thus pointing to the importance of modeling and analysis of transient faults and variation sources for the purpose of guiding the design process. This work presents a symbolic approach to modeling the effect of transient faults in digital circuits in the presence of variability due to process manufacturing. The results show that using a nominal case and not including variability effects, can underestimate the SER by 5% for the 50% yield point and by 10% for the 90% yield point. Natasa Miskov-Zivanov, Kai-Chiang Wu, Diana Marculescu |
ICCAD | 2 |
| 2008 | Power-aware soft error hardening via selective voltage scalingabstractNanoscale integrated circuits are becoming increasingly sensitive to radiation-induced transient faults (soft errors) due to current technology scaling trends, such as shrinking feature sizes and reducing supply voltages. Soft errors, which have been a significant concern in memories, are now a main factor in reliability degradation of logic circuits. This paper presents a power-aware methodology using dual supply voltages for soft error hardening. Given a constraint on power overhead, our proposed framework can minimize the soft error rate (SER) of a circuit via selective voltage scaling. On average, circuit SER can be reduced by 33.45% for various sizes of transient glitches with only 11.74% energy increase. The overhead in normalized power-delay-area product per 1% SER reduction is 0.64%, 1.33X less than that of existing state-of-the-art approaches. Kai-Chiang Wu, Diana Marculescu |
ICCD | 1 |
| 2006 | Delay variation tolerance for domino circuitsabstractFactors of delay variation, such as process variation and noise effects, may cause a manufactured chip to violate the pre-specified timing constraint. In this paper, we propose a re-synthesis technique to tolerate delay variation for domino circuits. Note that the slacks of nodes along critical paths are zero; any delay addition to those zero-slack nodes worsens the final performance of a circuit. Our basic idea is to increase the slacks of nodes in the critical region by appending a redundant auxiliary subcircuit to the original circuit. The auxiliary subcircuit can cause critical paths to become false paths or imperceptible paths as stated in S. Raj et al. (2004) so as to improve the capability of delay variation tolerance. Experimental results are very encouraging. Kai-Chiang Wu, Cheng-Tao Hsieh, Shih-Chieh Chang 0001 |
ASP-DAC | 1 |
| 2004 | Re-synthesis for delay variation toleranceabstractSeveral factors such as process variation, noises, and delay defects can degrade the reliabilities of a circuit. Traditional methods add a pessimistic timing margin to resolve delay variation problems. In this paper, instead of sacrificing the performance, we propose a re-synthesis technique which adds redundant logics to protect the performance. Because nodes in the critical paths have zero slacks and are vulnerable to delay variation, we formulate the problem of tolerating delay variation to be the problem of increasing the slacks of nodes. Our re-synthesis technique can increase the slacks of all nodes or wires to be larger than a pre-determined value. Our experimental results show that additional area penalty is around 21% for 10% of delay variation tolerance. Shih-Chieh Chang 0001, Cheng-Tao Hsieh, Kai-Chiang Wu |
DAC | 3 |