VLDB 2026 Research / reviewers in the wild / expert
Cheng-Wen Wu
dblp:74/1000
· DBLP profile ↗
184ranked-venue papers
16as first author
13since 2021 · last 2025
0000-0001-8614-7908ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 180 · 14 first-author · 13 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Chiplet Interconnect Test and RepairabstractThis research work addresses various challenges in testing and repairing chiplet-based (2.5D and 3D) multi-die integrated circuits (ICs) as manufacturing shifts toward complex, ultra-dense circuits with large amounts of die-to-die interconnects by micro-bump or hybrid-bump connections and possibly through-silicon vias (TSVs). Conventional interconnect testing methods are limited to detecting hard defects, leading to ineffectiveness, and cover many unrealistic shorts, causing inefficiency. This article introduces E2I-TEST, a method which covers, for a given collection of interconnects, all hard and weak variants of only realistic short, open, and coupling defects. E2I-TEST also prevents aliasing, thus supporting fault diagnosis. While it is predicted that the number of interconnects will rise significantly in the near future, E2I-TEST provides a high-quality interconnect test for which the number of test patterns is constant and no longer dependent on the number of interconnects. To further enhance the yield and reliability of chiplet-based systems, we propose a standardized repair description language for interoperable repair logic of inter-die interconnects, allowing vendors to describe or specify repair structures for automatic verification across various repair logic. We also develop a streamlined, polynomial-time repair algorithm for UCIe designs, minimizing their impact on signal rerouting. Finally, we present micro-bump map optimizations to reduce catastrophic defects and lower the spare interconnect usage. This work provides a comprehensive testing and repair framework, enhancing defect resilience and yield in future ICs. Po-Yao Chuang, Cheng-Wen Wu, Erik Jan Marinissen |
ITC | 2 |
| 2025 | Generating Test Patterns for Chiplet Interconnects With Optimized Effectiveness and EfficiencyabstractChiplet-based (2.5-D and 3-D) multidie packages typically feature numerous die-to-die interconnects using micro-bump connections and possibly through-silicon vias or interposer wires, which are prone to manufacturing defects, such as shorts and opens, in both hard and weak (resistive) variants. Traditional I-ATPG methods only cover hard defects and scale with the logarithm of the number of interconnects. Despite being considered efficient, they cover shorts between all interconnects, including those for which shorts are unrealistic given their relative layout positions. This article proposes E2I-TEST, which covers all hard and weak variants of open defects and of only the shorts and coupling defects between physically adjacent interconnects for both 3-D and 2.5-D chips, while preventing aliasing during fault diagnosis. This article further improves E2I-TEST to prevent ground bounce, avoiding undesired voltage fluctuations during test mode. While the number of interconnects is expected to rise significantly, E2I-TEST offers a high-quality interconnect test, while maintaining a constant number of test patterns. Po-Yao Chuang, Francesco Lorenzelli, Cheng-Wen Wu, Erik Jan Marinissen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Keynote 2 - Sustainability and the Outlook of Semiconductor IndustryabstractThe semiconductor industry continues to grow rapidly with the development of AI technology and its deployment in various application domains, broadly covering homes, offices, factories, vehicles, warehouses, all kinds of public sites, green energy fields, etc. A modern semiconductor-power AI system consists of not just the AI computing servers in the cloud, but also the 5G or satellite communication network, and a lot of smart Internet-of-Things devices distributed over a large application field. The sustainability of an AI system is of paramount significance to ensure reliable and long-lasting operations without failures. In this talk, we have discussed various strategies to achieve the sustainability of a multi-die IC used in such ever-increasing mission-critical AI-related applications. The techniques include rigorous testing, online monitoring, and self-healing techniques. These techniques target not just the hard failures but also the parametric defects that cause performance hazards, thereby affecting the reliability and or lifetime of an IC product. In particular, we demonstrate with a few examples how various timing circuits such as the Phase-Locked Loop, the Delay-Locked Loop, and the Time-to-Digital Converter can be used as effective building blocks inside an IC to support built-in speed grading, online health condition checking, and online self-repair to achieve the required sustainability. Cheng-Wen Wu, Shi-Yu Huang |
ETS | 1 |
| 2023 | A Low-Bitwidth Integer-STBP Algorithm for Efficient Training and Inference of Spiking Neural NetworksabstractSpiking neural networks (SNNs) that enable energy-efficient neuromorphic hardware are receiving growing attention. Training SNNs directly with back-propagation has demonstrated accuracy comparable to deep neural networks (DNNs). However, previous direct-training algorithms require high-precision floating-point operations, which are not suitable for low-power end-point devices. The high-precision operations also require the learning algorithm to run on high-performance accelerator hardware. In this paper, we propose an improved approach that converts the high-precision floating-point operations to low-bitwidth integer operations for an existing direct-training algorithm, i.e., the Spatio-Temporal Back-Propagation (STBP) algorithm. The proposed low-bitwidth Integer-STBP algorithm requires only integer arithmetic for SNN training and inference, which greatly reduces the computational complexity. Experimental results show that the proposed STBP algorithm achieves comparable accuracy and higher energy efficiency than the original floating-point STBP algorithm. Moreover, it can be implemented on low-power end-point devices to provide learning capability during inference, which are mostly supported by fixed-point hardware. Pai-Yu Tan, Cheng-Wen Wu |
ASP-DAC | 2 |
| 2023 | Effective and Efficient Testing of Large Numbers of Inter-Die Interconnects in Chiplet-Based Multi-Die PackagesabstractChiplet-based multi-die packages implement large numbers of inter-die interconnect bundles clustered in large micro-bump islands. These micro-bumps can be subject to manufacturing defects. The most common defect types are shorts and opens. Traditional interconnect automatic test pattern generation (I-ATPG) algorithms detect, for a given collection of interconnects, all shorts between any pair of interconnects, all open interconnects, and exclude any aliasing, independent from the interconnects’ layout positions. Exploiting knowledge of their relative layout positions, we derive a new, improved I-ATPG algorithm. For a user-defined and scalable definition of realistic shorts, the new I-ATPG approach (1) increases the defect coverage significantly (in an example case, between 18% and 67%) by including realistic inter-bundle shorts between micro-bumps from adjacent bundles, and (2) reduces the overall test pattern count (and hence, the resulting test time) by 33% by providing test patterns for realistic shorts only. Po-Yao Chuang, Francesco Lorenzelli, Sreejit Chakravarty, Cheng-Wen Wu, Georges Gielen, Erik Jan Marinissen |
VTS | 4 |
| 2023 | A 40-nm 1.89-pJ/SOP Scalable Convolutional Spiking Neural Network Learning Core With On-Chip Spatiotemporal Back-PropagationabstractIn recent years, progress in spiking neural network (SNN) research has generated growing interest in specialized SNN hardware. However, most of the hardware studies are about inference-only engines, and the training process for low-power endpoint devices remains arduous due to high numerical precision requirements. In this article, we introduce a scalable convolutional SNN learning core for energy-efficient training utilizing the spatiotemporal back-propagation (STBP) algorithm. We modify the STBP algorithm with five hardware-based enhancing methods, which minimize the hardware implementation cost without losing accuracy. We propose a unified core architecture encompassing three data-flow modes for three training phases. It also offers multicore scalability for wider and deeper models. A 40-nm prototype chip has been implemented, achieving peak training and inference efficiencies of 3 and 7.7 TOPS/W, respectively, at 90% input sparsity and an energy per synaptic operation (SOP) metric of 1.89 pJ/SOP. Based on the chip, our multicore prototype system attains a competitive accuracy of 99.1% on MNIST. We also exhibit the first on-chip training results on SVHN and CIFAR10, with an average of$36\times $improvement in efficiency compared with a typical commercial GPU platform. Pai-Yu Tan, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Battery Pack Reliability and Endurance Enhancement for Electric Vehicles by Dynamic ReconfigurationabstractIn the fast-growing electric vehicle (EV) industry, key technology challenges include the improvement of battery efficiency, reliability, and endurance. In this paper, we propose a novel battery pack design methodology that supports dynamic reconfiguration of the battery pack architecture, which partitions the battery modules into the primary group and secondary group. The proposed reconfigurable design effectively improves the battery pack reliability and endurance, especially for battery packs that contain modules with uneven aging conditions. Simulation results show that, with our approach, the equivalent average aging speed among all battery modules is slowed down, and the battery pack's endurance is increased. Yu-You Chou, Cheng-Wen Wu, Ming-Der Shieh, Chao-Hsun Chen |
ATS | 2 |
| 2022 | Aging Impact of Power MOSFETs in Charger with Different Operation FrequencyabstractThe global warming and pollution issues have reached an alarming level that requires immediate actions by all governments in the world to reduce fossil fuel vehicles, and even ban them in the foreseeable future. As a result, electric vehicles (EVs) are gaining ground rapidly. Inside the EVs, especially their power train, semiconductor power devices are considered the key components. Increasing the performance, energy efficiency, and reliability of power MOSFET, therefore, is critical in the EV industry. In this work, we stress the reliability and lifetime of power MOSFET devices and propose an aging model for them. We will show the procedure to construct the efficient aging model for a power MOSFET, supported by circuit-level simulation results. The aging conditions, including high temperature and/or high voltage on the gate oxide, will result in threshold voltage shift, which in turn decreases the MOSFET's switching speed and driving capability. With the proposed aging assessment tool for power MOSFET, the device's performance under different aging conditions can be predicted. The experiment results show that aging will reduces its switching speed and driving power. Experimental results by our tool also show that the circuit efficiency, ripple voltage, switching loss, etc., are negatively affected by aging. Kuan-Hsun Duh, Cheng-Wen Wu, Ming-Der Shieh, Chao-Hsun Chen, Ming-Yan Fan |
ATS | 2 |
| 2022 | A Decision Tree-Based Screening Method for Improving Test Quality of Memory ChipsabstractThere is a growing demand for high-reliability and high-quality integrated circuit (IC) products, while their test costs should be kept as low as possible. We investigate the test process of advanced memory chips, where the high temperature operating life (HTOL) test has been used to determine their intrinsic reliability. This high temperature sampling test can run from 168 to 1,000 hours, so it is time-consuming and expensive. Recently, machine learning (ML) algorithms have been used to solve classification problems, so far as good training data can be obtained. In our case, there is already a large amount of parametric test data generated from the existing test flow. Therefore, in this work, we propose a decision tree (DT)-based screening method to predict weak (unreliable) dies that would fail the HTOL test. We show that experienced test engineers can prioritize the parametric test data for better use of the DT model. Finally, we take advantage of the high interpretability of DT to develop the multi-feature heuristics, which can be used to improve the quality of final test (FT). Keeping the overkill rate at 0%, our heuristics can screen out 25% more bad dies, i.e., we can improve the FT quality without additional cost. Ya-Chi Cheng, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao |
ITC-Asia | 3 |
| 2022 | Weak Die Screening by Feature Prioritized Random Forest for Improving Semiconductor Quality and Reliabilityabstractwith the increasing demand for safety-critical products, the quality and reliability of semiconductor components are among the top priorities. In recent years, test data analytics by machine learning (ML) algorithms are widely considered to have great potential for improving the quality and reliability of semiconductor chips. In this work, we inspect a typical test flow of advanced semiconductor products, and propose an ML-based weak die screening method for improving the quality and reliability of shipped products. We propose the feature prioritized random forest (FPRF) model, which can fit smoothly into the existing test flow. We perform experiments on an advanced SRAM product using the FPRF model. We perform feature analysis based on the test data obtained from the final test (FT). After the FPRF screening, we are able to screen out more bad dies from those that have passed the FT. For an overkill rate of 12.93 %, the bad die hit rate can be as high as 96.55%. One can explore the FPRF model for other products as well. Shian-Yu Lin, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao |
ITC-Asia | 3 |
| 2022 | Improving Test Quality of Memory Chips by a Decision Tree-Based Screening MethodabstractThere is a growing demand for high-reliability and high-quality integrated circuit (IC) products, while their test costs should be kept as low as possible. We investigate the test process of advanced memory chips, where the high temperature operating life (HTOL) test has been used to determine their intrinsic reliability. This high temperature sampling test can run from 168 to 1,000 hours, so it is time-consuming and expensive. Recently, machine learning (ML) algorithms have been used to solve classification problems, so far as good training data can be obtained. In our case, there is already a large amount of parametric test data generated from the existing test flow. Therefore, in this work, we propose a decision tree (DT)-based screening method to predict weak (unreliable) dies that would fail the HTOL test. We show that experienced test engineers can prioritize the parametric test data for better use of the DT model. Finally, we take advantage of the high interpretability of DT to develop the multi-feature heuristics, which can be used to improve the quality of final test (FT). Keeping the overkill rate at 0%, we can screen out 25% more bad dies in the 5nm SRAM case with the heuristics, and in the 4nm case, we can screen out 14% more bad dies, i.e., we can improve the FT quality without additional cost. Ya-Chi Cheng, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao |
ITC | 3 |
| 2022 | Fault Modeling and Testing of Memristor-Based Spiking Neural NetworksabstractThe resistive random-access memory (RRAM), whose core is composed of memristor cell arrays, has recently been proposed for implementing the deep neural network (DNN) and spiking neural network (SNN), which potentially can improve the performance and energy efficiency of AI computing. In this paper, we address high-quality memristor-based SNNs. We propose two functional fault models for them, i.e., the Slow Integration Fault (SIF) and Fast Integration Fault (FIF), and show the circuit-level fault simulation results, taking process variation into account. The detailed simulation and analysis results show that the SIF and FIF are feasible functional fault models for memristor-based SNNs, as the interconnect open and short defects, as well as the memristor and transistor faults, can all be covered by the proposed SIF and FIF. We have also developed a test algorithm for the SNNs, i.e., the SNN-Test algorithm. Experimental results show that the SNN-Test covers 100% of the open and short defects and the proposed functional fault models. Kuan-Wei Hou, Hsueh-Hung Cheng, Chi Tung, Cheng-Wen Wu, Juin-Ming Lu |
ITC | 4 |
| 2021 | An Improved STBP for Training High-Accuracy and Low-Spike-Count Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) that facilitate energy-efficient neuromorphic hardware are getting increasing attention. Directly training SNN with backpropagation has already shown competitive accuracy compared with Deep Neural Networks. Besides the accuracy, the number of spikes per inference has a direct impact on the processing time and energy once employed in the neuromorphic processors. However, previous direct-training algorithms do not put great emphasis on this metric. Therefore, this paper proposes four enhancing schemes for the existing direct-training algorithm, Spatio-Temporal Back-Propagation (STBP), to improve not only the accuracy but also the spike count per inference. We first modify the reset mechanism of the spiking neuron model to address the information loss issue, which enables the firing threshold to be a trainable variable. Then we propose two novel output spike decoding schemes to effectively utilize the spatio-temporal information. Finally, we reformulate the derivative approximation of the non-differentiable firing function to simplify the computation of STBP without accuracy loss. In this way, we can achieve higher accuracy and lower spike count per inference on image classification tasks. Moreover, the enhanced STBP is feasible for the on-line learning hardware implementation in the future. Pai-Yu Tan, Cheng-Wen Wu, Juin-Ming Lu |
DATE | 2 |
| 2020 | A 90nm 103.14 TOPS/W Binary-Weight Spiking Neural Network CMOS ASIC for Real-Time Object ClassificationabstractThis paper introduces a low-power 90nm CMOS binary weight spiking neural network (BW-SNN) ASIC for real-time image classification. The chip maximizes data reuse through systolic arrays that house the entire 5-layer BW-SNN, requiring a minimum off-chip bandwidth for data access. The chip achieves 97.57% accuracy for real-time bottled-drink recognition, consuming only 0.62uJ per inference. For comparison purpose, it achieves 98.73% accuracy for MNIST hand-written character recognition, consuming only 0.59uJ per inference. The bottled-drink recognition is demonstrated at 300 fps that is well enough for many other real-time applications. The peak efficiency point is 103.14TOPS/W at a voltage of 0.6V, which outperforms other designs so far as we know. By normalizing to the 28nm technology node, the proposed ASIC is about 5× more efficient and 7× lower hardware cost as compared with the state-of-the-art designs. Po-Yao Chuang, Pai-Yu Tan, Cheng-Wen Wu, Juin-Ming Lu |
DAC | 3 |
| 2020 | Tightening the Mesh Size of the Cell-Aware ATPG Net for Catching All Detectable Weakest FaultsabstractCell-aware test (CAT) explicitly targets faults caused by cell-internal short and open defects and has been shown to significantly reduce test escape rates. CAT library cell characterization is typically done for only two defect resistance values: one representing hard opens and another one representing hard shorts. In this paper, similar to fishermen tightening the mesh size of their nets to catch small fish, we perform library characterization as efficiently as possible for a set of resistances representing increasingly weaker defects, and then adjust our ATPG flow to explicitly target faults caused by the weakest still-detectable variant of each potential defect. We implemented this novel approach in an experimental ATPG tool flow script, using functions of Cadence's Modus as building blocks. To assess the effectiveness of our approach, we formulate a new dedicated test metric: the weakest fault coverage wfc. Compared to conventional CAT targeting hard defects only, experimental results show that our new approach enhances detection of weakest faults and significantly reduces wfc escapes =1-wfc, while maintaining its original (hard-defect) fault coverage fc, of course at the expense of (acceptable) increases in the required number of test patterns and associated test generation time. Min-Chun Hu 0002, Santosh Malagi, Joe Swenton, Jos Huisken, Kees Goossens, Cheng-Wen Wu, Erik Jan Marinissen |
ETS | 7 |
| 2020 | A Deep Learning-Based Screening Method for Improving the Quality and Reliability of Integrated Passive DevicesabstractIntegrated passive devices (IPDs) have been widely used in advanced packaging of semiconductor chips, to improve their power integrity and impedance matching. There is a growing demand in guaranteeing signal and power integrity for the chips used in safety-critical products, such as those used in automotive, aviation, industrial, and defense systems, where IPDs help improve quality and reliability of the chips. Therefore, IPD testing and screening itself is essential. Note that the cost of replacing failed IPDs is much higher than the cost of manufacturing them, so screening bad IPDs before mounting is also crucial. In this work, we propose a machine learning (ML) based screening methodology to identifying the IPDs that have potential reliability issues. Based on the parametric data of 360,000 IPDs collected from the wafer probing test, the proposed Semiconductor Quality Net (SQnet) is trained to predict the IPDs which have low breakdown voltage, i.e., low reliability. Keeping the overkill rate below 10%, our method can screen out 6 to 15X more bad dies than the existing industrial methods, i.e., DPAT and GDBC. Chien-Hui Chuang, Kuan-Wei Hou, Cheng-Wen Wu, Mincent Lee, Chia-Heng Tsai, Hao Chen 0053, Min-Jer Wang |
ITC-Asia | 3 |
| 2020 | A Deep Learning-Based Screening Method for Improving the Quality and Reliability of Integrated Passive DevicesabstractIntegrated passive devices (IPDs) have been widely used in advanced packaging of semiconductor chips, to improve their power integrity and impedance matching. There is a growing demand in guaranteeing signal and power integrity for the chips used in safety-critical products, such as those used in automotive, aviation, industrial, and defense systems, where IPDs help improve quality and reliability of the chips. Therefore, IPD testing and screening itself is essential. Note that the cost of replacing failed IPDs is much higher than the cost of manufacturing them, so screening bad IPDs before mounting is also crucial. In this work, we propose a machine learning (ML) based screening methodology to identifying the IPDs that have potential reliability issues. Based on the parametric data of 360,000 IPDs collected from the wafer probing test, the proposed Semiconductor Quality Net (SQnet) is trained to predict the IPDs which have low breakdown voltage, i.e., low reliability. Keeping the overkill rate below 10%, our method can screen out 6 to 15X more bad dies than the existing industrial methods, i.e., DPAT and GDBC. Chien-Hui Chuang, Kuan-Wei Hou, Cheng-Wen Wu, Mincent Lee, Chia-Heng Tsai, Hao Chen 0053, Min-Jer Wang |
ITC | 3 |
| 2019 | Asian Test Symposium - Past, Present and Future -abstractThe Asian Test Symposium (ATS) provides an international forum for engineers and researchers from all countries of the world, especially from Asia, to present and discuss various aspects of device, board and system testing with design, manufacturing and field considerations in mind. ATS has been annually held for 27 years at 23 cities (four cities, Beijing, Shanghai, Hiroshima, and Taipei hosted ATS twice). Fig. 1 shows the venues of ATS including three coming ATS. During these years, ATS has been provided the opportunity to deeply discuss test technology and enhance networking in the research and geographical regions. ATS will continuously play this role in the future. Michiko Inoue, Xiaowei Li 0001, Cheng-Wen Wu |
ITC | 3 |
| 2018 | A Built-in Self-Test Scheme for Detecting Defects in FinFET-Based SRAM CircuitabstractFinFET is a feasible solution to short-channel effects that has been encountered by planar transistors during process scaling, and is widely adopted in advanced CMOS technologies. However, the special physical structure of FinFET also brings new defect models, which are hard to detect by conventional March algorithms, thus a more effective test methodology is required. In this paper, we investigate defect candidates in a FinFET 6T-SRAM circuit, analyze their fault behaviors, and propose a built-in self-test (BIST) scheme to enhance fault coverage and reduce test time. The proposed BIST approach is able to detect all target defects with only one read cycle, which reduces the required test algorithm complexity and hence reduces test time. This BIST scheme can be used with classical March algorithms, such as the March C-, to extend fault coverage beyond static single cell or coupling faults, thus reducing the defect level. Meng-Chi Chen, Tsung-Hsuan Wu, Cheng-Wen Wu |
ATS | 3 |
| 2018 | Covering hard-to-detect defects by thermal quorum sensingabstractWith the advent of highly complex and dense modern CMOS circuits, defects caused by parametric and process variations, e.g., are more and more difficult to detect. Many hard-to-detect defects not sensitized through the critical paths can easily escape from the conventional testing methods. In order to reduce the product defect level, we introduce the notion of quorum sensing (QS) to circuit testing (sensing) for improving the quality and reliability. The proposed thermal quorum sensing (TQS) mechanism triggers a thermal chain reaction to expose the subtle variations in the circuit due to small defects, which can be observed by the common cell population behavior. A model is introduced to charactize the feature of TQS on circuit. The simualtion result verified by the ISCAS s9234 benchmark with 45nm CMOS standard cell library shows when the number of small defects injected is more than 489, the difference in total current will be higher than 2.08mA. It can discover the subtle faults compared with other state-of-the-art or traditional testing methods. Po-Yao Chuang, Cheng-Wen Wu, Harry H. Chen |
ETS | 2 |
| 2018 | RRAM-Based Neuromorphic Hardware Reliability Improvement by Self-Healing and Error CorrectionabstractNeural network (NN) has been considered as an important factor for the success of many AI applications. As the von Neumann architecture is inefficient for NN computation, researchers have been investigating new semiconductor devices and architectures for neuromorphic computing. The crossbar RRAM, which is an emerging non-volatile memory composed of memristor devices, can be used to accelerate or emulate the NN computation. However, the memristor device defects exposed during manufacturing or field use may cause performance degradation in the NN, causing reliability issues to the neuromorphic hardware. In this paper, we consider two existing fault models for the 1T1R RRAM cell, i.e., the stuck-at fault and transistor stuck-on fault. Evaluation of their influence to the NN shows that for about 10% faulty cells in the memristor array, the accuracy for the MLP model degrades about 10%, and that for the LeNet 300-100 and LeNet 5 degrades by more than 65%. Therefore, we propose a self-healing and an error correction approach to reduce the accuracy degradation, and improve the reliability (lifetime) of the neuromorphic hardware. Our simulation results show that if we limit the accuracy degradation to within 5%, then the proposed error correction approach for the MLP model will be able to tolerate up to 40% faulty cells, and even up to 60% faulty cells for LeNet 300-100 and LetNet 5 models. Also, the error correction method can extend the lifetime of the neuromorphic hardware by 5% or more. Jia-Yun Hu, Kuan-Wei Hou, Chih-Yen Lo, Yung-Fa Chou, Cheng-Wen Wu |
ITC-Asia | 5 |
| 2018 | Solutions to Multiple Probing Challenges for Test Access to Multi-Die Stacked Integrated CircuitsabstractMulti-die stacked ICs are getting increasing traction in the market, fueled by innovations in wafer processing technologies (e.g., vertical inter-die and intra-die connections), stack assembly, and advanced packaging approaches (e.g., wafer-level packaging). Given the non-perfect nature of their manufacturing processes, these stacked ICs (SICs) need all to be individually tested for manufacturing defects in an effective, yet efficient manner. This paper discusses a handful of probing challenges specific to such SICs and their solutions: probing ultra-thin wafers on a flexible tape on extra-large tape frames, probing on large arrays of dense micro-bumps, analyzing probe-to-pad alignment (PTPA) accuracy contributions from probe station and probe card on the basis of probe mark images, and efficient auto-correction of individual misalignments of singulated dies or die stacks on tape. The paper concludes with a real-life case study, in which most of the discussed challenges and solutions are combined. Erik Jan Marinissen, Ferenc Fodor, Arnita Podpod, Michele Stucchi, Yu-Rong Jian, Cheng-Wen Wu |
ITC | 6 |
| 2017 | An Enhanced Boundary Scan Architecture for Inter-Die Interconnect Leakage Measurement in 2.5D and 3D Packagesabstract3D-IC is a solution to achieve lower cost and higher performance as the transistor density doubles every 18 months following Moore's law. In recent years, Integrated Fan-Out Wafer-Level Chip-Scale Packaging (InFO WLCSP) is one of the most promising packaging technologies for 3D-IC. In InFO WLCSP, there are some inter-die interconnects. We cannot access these interconnects directly. Therefore, it causes 1-2% test coverage loss. As a result, a built-in self-test (BIST) or other design-for-test (DFT) methodology is necessary to test these interconnects. It is easy to detect open defects and short defects leading to large leakage currents with conventional test methods. However, defects leading to small leakage currents are hard to detect, so we focus on these defects in this work. In this paper, we propose a scheme to measure the small leakage current of these inter-die interconnects. The scheme uses IEEE 1149.1 boundary scan interface. Because boundary scan is a well-adopted standard, the scheme can be integrated into any electronic product easily. Beside four mandatory terminals, we add a reference current input. Users can apply current to compare it with the leakage current. For the input EBSC, two OR gates, a MUX, and a current digitizer are added and it is compared with the conventional BSC. A test chip is implemented to verify our design, which uses the one-polynine-metal (1P9M) 90nm CMOS technology. Measurement results show that the proposed scheme is able to measure the leakage current of the interconnect. In addition, with the dynamic range of 128nA, the current digitizer has 6-bit resolution. Pok Man Preston Law, Cheng-Wen Wu, Long-Yi Lin, Hao-Chiao Hong |
ATS | 2 |
| 2017 | Cell-aware test generation time reduction by using switch-level ATPGabstractIn this paper, we propose an efficient test flow for Cell-Aware Test (CAT) to drastically reduce the time for CAT-enhanced test generation at the cell level. In CAT, the detail transistor-level circuit simulation is used to find appropriate test patterns and it has been considered as very time consuming. To solve this problem, first, we exploit Switch-Level ATPG (SL-ATPG) and experimentally show that it can efficiently generate test patterns in the CAT flow. Second, based on layout-oriented defect generation method, we propose an algorithm to automatically inject those defects into the switching network used in SL-ATPG, for cells in the library. Third, note that the traditional ATPG is primarily based on the stuck-at-fault and transition-fault models, it is difficult to find small-delay faults. However, the same defects are likely to be detected by observing the short-circuit current, so we propose current-based checks for a pattern generation method which are able to detect the existence of a short-circuit path. Finally, we compare the simulation time of detailed circuit simulation and of SL-ATPG in CAT. The experiment is based on a commercial 180nm CMOS standard cell library. Moreover, it shows that SL-ATPG method can successfully reduce the simulation time by about 403X. Po-Yao Chuang, Cheng-Wen Wu, Harry H. Chen |
ITC-Asia | 2 |
| 2017 | Symbiotic system models for efficient IGT system design and testabstractIOT has seen countless potential applications that can improve our lives dramatically, but after years of efforts by numerous companies and organizations, the beautiful dreams are yet to be realized. The main obstacles are cost and energy consumption constraints of the devices and systems, which still cannot be contained. As a step forward in improving the reliability and reducing the cost and energy consumption of IOT devices and systems, we propose a high-level model for efficient design-space exploration, which is called the symbiotic system (SS) model. Based on that, we propose a quorum-sensing model for SS to improve the reliability of IOT systems that contain subsystems. Experimental result shows that, with the SS approach, the lifetime of an example IOT storage system with 50K failure-in-time/Mb nodes can be extended by 74.41 %. In addition, to double the lifetime of the system, only 5% additional repair resource is required. In short, by the SS model, it is possible for an IOT system to achieve low cost, low energy consumption, and high reliability. Cheng-Wen Wu, Bing-Yang Lin, Hsin-Wei Hung, Shu-Mei Tseng |
ITC-Asia | 1 |
| 2017 | Highly reliable and low-cost symbiotic IOT devices and systemsabstractIOT has seen countless potential applications that can improve our lives dramatically, but after years of efforts by numerous companies and organizations, the beautiful dreams are yet to be realized. The main obstacles are cost and energy consumption constraints of the devices and systems, which still cannot be contained. As a step forward in improving the reliability and reducing the cost and energy consumption of IOT devices and systems, at ITC-Asia17 we proposed a high-level model for efficient design-space exploration, which is called the symbiotic system (SS) model. Based on that, we also proposed a quorum-sensing model for SS to improve the reliability of IOT systems that contain subsystems. Preliminary experimental result shows that it is possible for an IOT system to achieve low cost, low energy consumption, and high reliability. In this extended work we discuss more evaluation results, and show that by using quorum-sensing-based peer-repair, up to 97% of the faulty devices can be repaired even when the raw yield is only 30%. It shows that by adopting the proposed approach, there is hope in eliminating production test of symbiotic IOT devices in the future. Bing-Yang Lin, Hsin-Wei Hung, Shu-Mei Tseng, Cheng-Wen Wu |
ITC | 5 |
| 2017 | A Built-Off Self-Repair Scheme for Channel-Based 3D MemoriesabstractRedundancy repair is a commonly used technique for memory yield improvement. In order to ensure high repair rate and final product yield, it is necessary to develop a repair scheme for the coming three-dimensional (3D) architecture of stacked DRAM. According to the JEDEC mobile memory technology roadmap, the interface of 3D DRAM, including the Wide I/O and High-Bandwidth Memory (HBM), is mainly classified as channel-based memories. In this paper, we propose a built-off self-test(BOSR) scheme at the controller level for channel-based 3D memory to enhance final product yield after the bonding of a memory cube to its corresponding logic die. The logic die contains the Channel controller, in which the BOSR circuit resides. Experimental results show that the repair rate is high with higher cluster failure ratio due to the flexible algorithm we choose. The area overhead is low and it decreases significantly when the memory size or channel count increases. The performance penalty is also low due to the parallel execution of address comparison and repair. Moreover, the manufacture cost is lower than conventional DRAM architecture due to allocator-based redundancies. Finally, the proposed scheme can easily be applied to other channel-based 3D memories. Hsuan-Hung Liu, Bing-Yang Lin, Cheng-Wen Wu, Wan-Ting Chiang, Mincent Lee, Hung-Chih Lin, Ching-Nen Peng, Min-Jer Wang |
IEEE Trans. Computers | 3 |
| 2016 | Efficient Cell-Aware Fault Modeling by Switch-Level Test GenerationabstractThis paper proposes methods to drastically reduce the expensive analog fault simulation currently used to create cell-aware fault models. By exploiting low-power properties of common CMOS designs, most defects in the transistor-level netlist containing parasitics can be represented by just two canonical fault classes. Via simple circuit analysis, we show that faulty behaviors are completely predictable as the defect resistance parameter value varies from zero to infinity, thus eliminating the need for circuit simulation at multiple parameter values. The two canonical fault classes can be modeled by transistor switch stuck-open and stuck-closed faults. Rather than enumerating the full combination cell input patterns to search for defect detection conditions by analog fault simulation, switch-level test generation can obtain those input conditions directly, thereby reducing significantly the role of analog simulation to that of ranking conditions in terms of detection effectiveness. Harry H. Chen, Simon Y.-H. Chen, Po-Yao Chuang, Cheng-Wen Wu |
ATS | 4 |
| 2016 | Layout-Oriented Defect Set Reduction for Fast Circuit Simulation in Cell-Aware TestabstractThe cell-aware test (CAT) methodology was previously proposed to target cell-internal faults that cannot be easily detected by gate-level stuck-at fault (SAF) patterns generated by conventional ATPG. It was shown to reduce the defect level on CMOS-based designs, with the help of detailed defect injected transistor-level circuit simulation and defect-enhanced SAF ATPG. The detailed transistor-level circuit simulation has been considered an issue in CAT, as it is very time consuming. The problem mainly lies in that all parasitic capacitors and resistors extracted from cell layout are considered as defect targets, so the defect set is large. To reduce the defect set, and therefore the circuit simulation time, we take layout into consideration when we construct the defect set for each cell, effectively removing the redundant or unnecessary defects and therefore reducing the circuit simulation time dramatically. We propose a generalized approach that can be used to build the fault models based on the cell layout, where the generated faults are closer to the realistic physical defects on the layout, so the number of faults is significantly reduced. The proposed method is verified by commercial 180nm and 350nm CMOS standard cell library, and the circuit simulation time is reduced to only 19% or even lower as compared with the original CAT methodology. Hsuan-Wei Liu, Bing-Yang Lin, Cheng-Wen Wu |
ATS | 3 |
| 2016 | Efficient probing schemes for fine-pitch pads of InFO wafer-level chip-scale packageabstractWith the increasing demand of super high scale of integration and small form factor in advanced semiconductor products, especially those that integrate DRAM and logic dies, 3D IC and Wafer-Level Chip-Scale Packaging (WLCSP) are considered promising approaches. In Integrated Fan-Out (InFO) WLCSP, a large number of fine-pitch pads, where neighboring pads cannot be probed simultaneously due to insufficient pitch, are used as the contact interfaces of inter-die interconnections. If the fine-pitch pads cannot be probed, the interconnections between the pads and boundary scan cells (BSCs) cannot be tested, which can lead to higher defect level. From industrial investigation, untested fine-pitch pads lead to 1-2% test coverage loss. To improve the overall test coverage, in this paper, we propose a pre-bond probing methodology for fine-pitch pads of InFO WLCSP. By the proposed probing schemes, open/short faults on the interconnects between the fine-pitch pads and BSCs can be all tested by the ATE. Moreover, for short faults that only occur between adjacent pads (interconnects), we propose a grouping method to determine the test patterns at each probing stage, which can minimize the test time. We also show that our method can achieve 100% test coverage of open/short faults. Yu-Chieh Huang, Bing-Yang Lin, Cheng-Wen Wu, Mincent Lee, Hao Chen 0053, Hung-Chih Lin, Ching-Nen Peng, Min-Jer Wang |
DAC | 3 |
| 2016 | A fast sweep-line-based failure pattern extractor for memory diagnosisabstractMemories have been considered as one of the major drivers of CMOS technology, due to their high density, high capacity, critical timing, sensitivity, etc. Memory diagnosis is therefore important for technology and product development. Memory failure pattern identification is traditionally an essential task for diagnosis. As memory density and capacity continue to grow, the amount of test data also keeps increasing, thus a more efficient failure pattern extraction method is required. In this paper, we propose a sweep-line-based memory failure pattern extractor to speed up the extraction process. The proposed tool can extract memory failure patterns from an industrial case of 20 state-of-the-art wafers in about 35 minutes, reducing 15% analysis time as compared with the sparse-matrix-based method that is the best so far. In our experiments of the industrial case, we are also able to identify a critical failure pattern that was not defined before, showing its more robust pattern identification capability. Sin-Yu Wei, Bing-Yang Lin, Cheng-Wen Wu |
ETS | 3 |
| 2016 | Is IoT coming to the rescue of semiconductor?abstractThe notion of Internet-of-Things (IoT) has been around for decades, and has been seriously discussed in the industry and academia in the past ten years at least. It has long been identified, or expected, as the main driving force of growth for many industries in the future. However, so far there is not so much evidence that IoT will likely give a great boost to the semiconductor industry (that we are all concerned here) in the near future. People are realizing that IoT is NOT a tangible industry, but instead just a phenomenon or notion of industry migration toward “Smart and Connected Everything.” As IC test engineers or scientists, should we or should we not bet on IoT with our future? Why? If IoT is really coming to the rescue of the struggling semiconductor industry, what will be the key factors of its success? In my talk, I will try to address these issues, from not just technological but also economic and social points of view. I will stress IC as well as system test. Cheng-Wen Wu |
ETS | 1 |
| 2016 | A Local Parallel Search Approach for Memory Failure Pattern IdentificationabstractDue to more aggressive design rules adopted by memories than logic circuits, memories have been considered as the major technology driver of advanced logic circuits, so far as CMOS process technology is concerned. Memory failure pattern identification therefore is important, and is traditionally considered a key task that can help improve the efficiency of memory diagnosis and failure analysis. Critical failure patterns (that are the yield killers), however, may change in different memory designs and process technologies. It is difficult to identify critical failure patterns from high-volume memory failure bitmaps if they are not predefined. To solve this problem, we propose a local parallel search algorithm for efficient memory failure pattern identification. In addition, the proposed system integrates the defect-spectrum-based and coordinate-distance-based methods to identify critical memory failure patterns from a large amount of memory failure bitmaps automatically, even if they are not defined in advance. In our experiment for 132,488 4-MB memory failure bitmaps, the proposed system can automatically identify six critical yet undefined failure patterns in minutes, in addition to all known patterns. In comparison, the state-of-the-art commercial tools need manual inspection of the memory failure bitmaps to identify the same failure patterns. Bing-Yang Lin, Cheng-Wen Wu, Mincent Lee, Hung-Chih Lin, Ching-Nen Peng, Min-Jer Wang |
IEEE Trans. Computers | 2 |
| 2014 | BIST-Assisted Tuning Scheme for Minimizing IO-Channel Power of TSV-Based 3D DRAMsabstractThree-dimensional dynamic random access memory (3D DRAM) using through-silicon via (TSV) has been acknowledged as one good approach for overcoming the memory wall. However, the IO-channel power of a TSV-based 3D DRAM represents a significant portion of the 3D DRAM power. In this paper, we propose a built-in self-test (BIST) -assisted tuning scheme to adjust the driving capability of programmable drivers to fit the number of stacked 3D DRAM dies such that the IO-channel power can be minimized. A BIST design supporting specific test patterns and test flow for the driver tuning is proposed as well. Simulation results show that about 6.16×10 -- 2 J energy saving can be achieved for a logic-DRAM stack with 150fF/die TSV load under 100s write operations if the proposed BIST-assisted tuning scheme is implemented in the logic die. Yun-Chao You, Chi-Chun Yang, Jin-Fu Li 0001, Chih-Yen Lo, Chao-Hsun Chen, Jenn-Shiang Lai, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
ATS | 9 |
| 2014 | Redundancy architectures for channel-based 3D DRAM yield improvementabstractThe three-dimensional integrated circuit (3D IC) is considered a promising approach that can obtain high data band-width and low power consumption for future electronic systems that require high integration level. One of the popular drivers for 3D IC is the integration of a memory stack and a logic die. Because the yield of a 3D IC is the product of respective yields of the mounted dies, the yields of the memory dies and logic die must be high enough, or the 3D IC will be too expensive to be manufactured. To obtain a high yield of 3D ICs, efficient test and repair methodologies for memories are necessary. In this paper, we target the channel-based 3D dynamic random access memory (DRAM) and propose two 3D redundancy architectures, i.e., Cubical Redundancy Architectures 1 and 2 (CRA1 and CRA2). We use Wide-IO DRAM as an example for discussion. In CRA1, spares are associated with each DRAM die as in a conventional 2D architecture. In CRA2, we use a static random access memory (SRAM) on the logic die as spares. Experimental results show that the CRA1 can achieve up to 18% higher stack yield than traditional redundancy architecture with the same area overhead. On the other hand, the CRA2 can achieve the same yield as the CRA1 with 40% less spares, but 1.3% higher area overhead. Bing-Yang Lin, Wan-Ting Chiang, Cheng-Wen Wu, Mincent Lee, Hung-Chih Lin, Ching-Nen Peng, Min-Jer Wang |
ITC | 3 |
| 2014 | DArT: A Component-Based DRAM Area, Power, and Timing Modeling ToolabstractDRAM renovation calls for a holistic architecture exploration to cope with bandwidth growth and latency reduction need. In this paper, we present DRAM area power timing (DArT), a DRAM area, power, and timing modeling tool, for array assembly and interface customization. Through proper design abstraction, our component-based modeling approach provides increased flexibility and higher accuracy, making DArT suitable for DRAM architecture exploration and performance estimation. We validate the accuracy of DArT with respect to the physical layout and circuit simulation of an industrial 68 nm commodity DRAM device as a reference. The experiment results show that the maximum deviations from the reference design, in terms of area, timing, and power, are 3.2%, 4.92%, and 1.73%, respectively. For an architectural projection by porting it to a 45 nm process, the maximum deviations are 3.4%, 3.42%, and 8.57%, respectively. The combination of modeling performance, flexibility, and accuracy of DArT allows us to easily explore new DRAM architectures in the future, including 3-D stacked DRAM. Hsiu-Chuan Shih, Pei-Wen Luo, Jen-Chieh Yeh, Shu-Yen Lin, Ding-Ming Kwai, Shih-Lien Lu, Andre Schaefer, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2014 | Low-Cost Post-Bond Testing of 3-D ICs Containing a Passive Silicon Interposer BaseabstractThrough-silicon vias (TSVs) provide high-density vertical interconnects between dies and enable the creation of 3-D ICs having higher performance and lower power consumption than traditional 2-D ICs. A practical TSV-based 3-D integration approach is to place multiple dies (or die stacks) side by side on a passive silicon interposer base, in which there are TSVs and metal wires serving as interconnects. In this paper, we propose a post-bond design-for-test architecture and a test strategy for such interposer-based 3-D ICs. Functional package pins and interconnects are reused to build multibit parallel test access mechanisms (PTAMs), which provide post-bond test access with no or low extra area costs. Four PTAM architectures are presented, and the corresponding PTAM optimization algorithms are proposed which can quickly identify the best PTAM configuration to achieve the shortest test time. We also propose an algorithm for adding dedicated test interconnects to improve test bandwidth at the expense of extra microbumps and metal wires. Experimental results show that the proposed techniques are effective in test length (and therefore test time) reduction. Moreover, cost–benefit analysis results suggest that our approaches have lower total test costs compared with a base-case one-bit JTAG-only solution. Chun-Chuan Chi, Erik Jan Marinissen, Sandeep Kumar Goel, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Application-Independent Testing of 3-D Field Programmable Gate Array Interconnect Faultsabstract3-D integration has been touted as an approach to reducing the lengths of critical paths in field-programmable gate arrays (FPGAs). A 3-D chip stacks a number of 2-D FPGA bare dies, interconnected by through-silicon vias (TSVs) and micro bumps, to attain a high packing density. However, the technology also introduces new types of defects, such as TSV void and microbump misalignment. Testing the interconnection faults becomes inevitable. In this paper, we present an automatic test pattern generator for open, short, and delay faults on 3-D FPGA interconnects by exploiting the regularity of switch matrix topology and forming repetitive paths with finite steps and with loop-back. The experimental results show that 12 test patterns (TPs) suffice to achieve 100% open fault coverage (FC). To detect all possible neighboring short faults, we need more than 40 TPs, whose number increases only slightly with the height of the 3-D FPGA. The TPs have high delay FC (96%) for 3-D FPGAs with the number of configurable logic blocks ranging from 50 × 50 × 2 to 50 × 50 × 6, demonstrating the scalability of our method. Yen-Lin Peng, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Processor and DRAM integration by TSV-based 3-D stacking for power-aware SOCsabstractWith the rapid popularization of mobile devices, the low-power and energy-efficient became far more important than the system operating frequency. This work demonstrates a processor and DRAM integration scheme by TSV-based 3-D stacking and the performance and energy efficiency is evaluated by an ESL design methodology. The integration scheme comprising Sans-Cache DRAM (SCDRAM) architecture which is designed under the power and energy considerations is explored. Experiment results show the proposed architecture can greatly reduce 80% energy while having 23.5% of system performance improvement. Shin-Shiun Chen, Chun-Kai Hsu, Hsiu-Chuan Shih, Jen-Chieh Yeh, Cheng-Wen Wu |
ASP-DAC | 5 |
| 2013 | Exploration Methodology for 3D Memory Redundancy Architectures under Redundancy ConstraintsabstractRedundancy repair is a commonly-used technique for memory yield improvement. In order to ensure high repair efficiency and final product yield, it is necessary to explore and develop the memory redundancy architecture carefully. However, due to the different failure distributions of memory arrays and various design constraints of memory architectures, it is difficult to explore the efficiency of the memory architecture thoroughly. In this paper, we propose a redundancy architecture exploration methodology to find the redundancy architecture with highest repair rate under redundancy constraints. Given a set of design constraints, failure distributions, and memory architectures, our methodology can explore at least 3(log2M* log2N* log2S) redundancy architectures systematically, where M, N, and S are the address sizes of memory row and column in a die, and the number of slices in the memory cube, respectively. In our experiments, the repair rates of 10 different 3D redundancy architectures with 3 different redundancy analysis algorithms in a given failure pattern distribution are simulated. The experimental result shows that the difference of the repair rates between the most efficient and least efficient memory redundancy architectures is up to 49.42%. Bing-Yang Lin, Mincent Lee, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 2013 | An enhanced double-TSV scheme for defect tolerance in 3D-ICabstractDie stacking based on Through-Silicon Via (TSV) is considered as an efficient way to reducing power consumption and form factor. In the current stage, the failure rate of TSV is still high, so some type of defect tolerance scheme is required. Meanwhile, the concept of double-via, which is normally used in traditional layer to layer interconnection, can be one of the feasible tolerance schemes. Double-via/TSV has a benefit compared to TSV repair: it can eliminate the fuse configuration procedure as well as the fuse layer. However, double-TSV has a problem of signal degradation and leakage caused by short defects. In this work, an enhanced scheme for double-TSV is proposed to solve the short-defect problem through signal path division and VDD isolation. Result shows that the enhanced double-TSV can tolerate both open and short defects, with reasonable area and timing overhead. Hsiu-Chuan Shih, Cheng-Wen Wu |
DATE | 2 |
| 2013 | Holistic approach to low-power system designabstractIn the past few years, we have witnessed the energy crisis and the financial tsunami that played an unwanted duo, changing the world in many aspects that affect most of us. While companies are working hard in getting out of the slump, many research organizations are rethinking how their R&D budget should be invested. We consider advanced research activities stressing ultra-low power and energy-efficient circuits and systems a top-priority direction. Among our list of R&D topics are ultra-low voltage circuits and systems, energy-harvesting circuits and systems, 3D integration based on the Through-Silicon-Via (TSV) technology, and normally-off computing technologies. To be successful in integrating the basic technologies developed, a holistic approach should be considered. Therefore, in addition to introduction and discussion of the above research topics at ITRI, I will also discuss a system-level modeling and evaluation platform, emphasizing power/energy efficiency. Cheng-Wen Wu |
ISLPED | 1 |
| 2013 | 3D-IC interconnect test, diagnosis, and repairabstractThrough-Silicon-Via (TSV)-based three-dimensional ICs (3D-ICs) have gained increasing attention due to their potential in reducing manufacturing costs and capability of integrating more functionality into a single chip. One of the most important factors that affect 3D-IC yield is the integrity of interconnects which connect different dies in a 3D-IC. This paper proposes a Design-for-Test (DIT) scheme that can 1) detect faulty interconnects in 3D-ICs, 2) pinpoint open defect locations to help yield learning, and 3) repair faulty interconnects caused by open defects to improve the 3D-IC yield. Experimental results show that the proposed scheme can achieve a diagnosis resolution of 84% for open defects. With the interconnect repair mechanism, the 3D-IC yield is improved by 10%. In addition, cost-benefit analysis reveals that the proposed technique can significantly increase the net profit, especially when the natural interconnect yield is low. Chun-Chuan Chi, Cheng-Wen Wu, Min-Jer Wang, Hung-Chih Lin |
VTS | 2 |
| 2013 | Special session 4C: Hot topic 3D-IC design and testabstractThree-dimensional (3D) integration using through silicon via (TSV) is a promising approach to coping with the challenges faced by the current 2D technology. A TSV-based 3D IC is implemented by stacking multiple dies which are vertically connected by TSVs. This may shorten the global interconnects of a 3D IC and greatly improve its performance and power consumption. High bandwidth is achieved by the increase of IO channels provided by the TSVs, which also reduce the unnecessary waste of energy during data movement. In addition, the 3D integration technology shows other advantages over 2D technology, such as high functionality, heterogeneous integration, small form factor, etc. However, there are still challenges that need to be tackled before volume production of 3D ICs using TSV becomes possible, including technology scalability, quality and reliability, yield, thermal management, equipment and infrastructure, and costs. To demonstrate the feasibility of 3D-IC technologies, many academic and industrial institutes have been working on various test vehicles, especially in recent years. An increasing attention also has been attracted around the world in the semiconductor industry by the development of related technologies. In this special session, we will discuss the evolutionary efforts toward the realization of 3-D ICs. As memory dies need to be integrated in most 3D-IC system, we will also address the challenges in the design and test of 3D memories against cross-layer process, voltage and temperature variations while suppressing thermal effect and power consumption, etc. We will share our experiences and show results from some of the test vehicles we have worked on, including processor/memory stacks, analog/logic stacks, logic/logic stacks, etc. We will show a 1,024-bit wide bus chip-to-chip interconnection using 40×40 fine-pitch TSVs to demonstrate ultra-low-power operation. For memories, we will present a 3D-RAM structure using small voltage-swing TSVs, vertical-device-stacking nonvolatile-SRAM and ReRAM, and a 3D vertical-gate NAND flash. To address the challenge of reliability and yield, we will also discuss important test techniques. Demonstration of benefits provided by 3D integration technology will also be shown. Last but not least, we will describe our development plan regarding various types of die stacking using heterogeneous process integration, especially processor/memory stacking that is widely believed to be a key technology in future generations of smart handheld devices. Jin-Fu Li 0001, Cheng-Wen Wu, Masahiro Aoyagi, Meng-Fan Chang, Ding-Ming Kwai |
VTS | 2 |
| 2013 | A hybrid ECC and redundancy technique for reducing refresh power of DRAMsabstractDynamic random access memory (DRAM) is one key component in handheld devices. It typically consumes significant portion of the energy of the device even if the device is in standby mode due to the refresh requirement. This paper proposes a hybrid error-correcting code (ECC) and redundancy (HEAR) technique to reduce the refresh power of DRAMs in standby mode. The HEAR circuit consists of a Bose-Chaudhuri-Hocquenghem (BCH) module and an error-bit repair (EBR) module to raise the error correction capability and minimize the adverse effects caused by the ECC technique such that the refresh period can be effectively prolonged and considerable refresh power reduction can be achieved. Analysis results show that the proposed HEAR scheme can achieve 40~70% of energy saving for a 2Gb DDR3 DRAM in standby mode. The area cost of parity data and ECC circuit of HEAR scheme is only about 63 % and 53 % of that of the ECC-only, respectively. Yun-Chao You, Chih-Sheng Hou, Li-Jung Chang, Jin-Fu Li 0001, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
VTS | 8 |
| 2013 | Generalization of an Enhanced ECC Methodology for Low Power PSRAMabstractError control codes (ECCs) have been widely used to maintain the reliability of memories, but ordinary ECC codes are not suitable for memories with long codewords. For portable products, power reduction in memories with DRAM-like cells can be done by reducing the refresh frequency, but the loss of data integrity should be taken care of seriously. To solve these issues, we have proposed a parallel encoding and decoding ECC scheme to reduce refresh power for an industrial pseudo-SRAM (PSRAM) with long codewords. In this paper, we briefly review the scheme and propose a systematic way to generate the parity check matrix for the new ECC scheme. We also modify the parity correction mechanism to reduce the operating power of the scheme. As for the 70 ns access time of the 256-MB PSRAM with 64-bit codewords and 16-bit I/O, experimental results show that the new ECC scheme can be integrated with the READ/WRITE operations with about 0.2 percent circuit area overhead and less than 3.5 ns encoding/decoding time. The new ECC architecture provides a flexible solution for memories with different widths of ECC codewords and I/O ports, without the error masking effect or reduction in reliability. Po-Yuan Chen, Chin-Lung Su, Chao-Hsun Chen, Cheng-Wen Wu |
IEEE Trans. Computers | 4 |
| 2013 | Low-Cost Error Tolerance Scheme for 3-D CMOS ImagersabstractThis paper presents an error tolerance scheme for 3-D CMOS imagers that are constructed by stacking a pixel array of imager sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array using microbumps$(\mu{\rm bumps})$and through silicon vias (TSVs). To deliver high-quality images in the presence of single or multiple$\mu{\rm bump}$, ADC, or TSV failures, we propose to interleave the connections from pixels to ADCs and recover the corrupted data in the ISPs. Key design parameters, such as the interleaving stride and the grouping ratio are determined by analyzing the employed error correction algorithm. Architectural simulation results demonstrate that the error tolerance scheme enhances the effective yield of an exemplar 3-D imager from 44% to 97%. Hsiu-Ming Chang 0001, Jiun-Lang Huang, Ding-Ming Kwai, Kwang-Ting Cheng, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2013 | Write Current Self-Configuration Scheme for MRAM Yield ImprovementabstractMagnetic random access memory (MRAM) is an emerging nonvolatile memory, which is widely studied for its high speed, high density, small cell size, and almost unlimited endurance. However, for deep-submicrometer process technologies, significant variation in the MRAM cells' operating condition results in write failures in cells and reduces the production yield. Memory designers have to characterize failed MRAM chips to find a suitable current level for reconfiguring their write current, which is time consuming. In this paper, we propose an efficient operating-current search method and the corresponding built-in circuit for toggle MRAM, which can rapidly find the minimal operating current. With the built-in search circuit, an MRAM chip can dynamically configure its write current through few tester channels. The resulting chip works correctly and consumes lower power. Production yield, thus, can be increased while the test cost is greatly reduced. We also present a generator of the circuit, which determines the circuit parameters according to the memory specifications and user requirements, and automatically generates the corresponding modules. Ching-Yi Chen, Sheng-Hung Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | AC-Plus Scan Methodology for Small Delay Testing and CharacterizationabstractSmall delay defects escaping traditional delay testing could cause a device to malfunction in the field and thus detecting these defects is often necessary. To address this issue, we propose three test modes in a new methodology called AC-plus scan, in which versatile test clocks can be generated on the chip by embedding an all-digital phase-locked loop (ADPLL) into the circuit under test (CUT). AC-plus scan can be executed on an in-house wireless test platform called HOY system. The first test mode of our AC-plus scan provides a more efficient way to measure the longest path delay associated with each test pattern. Experimental result shows that our method could greatly reduce the test time by 81.8%. The second test mode is designed for volume production test. It could effectively detect small delay defects and provide fast characterization on those defective chips for further processing. This mode could be used to help predict which chips are more likely to fall victim to operational failure in the field. The third test mode is to extract the waveform of each flip-flop's output in a real chip. This is made possible by taking advantage of the almost unlimited test memory our HOY test platform provides, so that we could easily store a great volume of data and reconstruct the waveform for post-silicon debugging. We have successfully fabricated a Viterbi decoder chip with such an AC-plus scan methodology inside to demonstrate its capability. Tsung-Yeh Li, Shi-Yu Huang, Hsuan-Jung Hsu, Chao-Wen Tzeng, Chih-Tsun Huang, Jing-Jia Liou, Hsi-Pin Ma, Po-Chiun Huang, Jenn-Chyou Bor, Ching-Cheng Tien, Chi-Hu Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2013 | In-Situ Method for TSV Delay Testing and Characterization Using Input Sensitivity AnalysisabstractIn this paper, we propose a method and the required architecture for characterizing the propagation delays of the through Silicon vias (TSVs) in a 3-D IC. First of all, every two TSVs are paired up to form an oscillation ring with some peripheral circuits. Their joint performance can thus be measured roughly by the oscillation period of the ring. Next, we utilize a technique called sensitivity analysis to further derive the propagation delay of each individual TSV participating in an oscillation ring-a distilling process. In this process, we perturb the strength of the two TSV drivers, and then measure their effects in terms of the change of the oscillation ring's period. By some following analysis, the propagation delay of each TSV can be revealed. On top of scheme, we also present an architecture that can activate the performance characterization process of each test unit - that consists of two TSVs - one at a time in a proper sequence. The area overhead is only 18.97 equivalent two-input NAND gate per TSV, by which one can gain the ability to profile the capacitances and the propagation delays of the TSVs on a 3-D IC. Jhih-Wei You, Shi-Yu Huang, Yu-Hsiang Lin, Meng-Hsiu Tsai, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2012 | On test and repair of 3D random access memoryabstractThe three-dimensional (3D) random access memory (RAM) using through-silicon via (TSV) has been considered as a promising approach to overcome the memory wall. However, cost and yield are two key issues for volume production of 3D RAMs, and yield enhancement increasingly requires test techniques. In this paper, we first introduce issues and existing techniques for the testing and yield enhancement of 3D RAMs. Then, a built-in self-repair (BISR) technique for 3D RAM using global redundancy is presented. According to the redundancy analysis results of each die with the BISR circuit, the die-to-die (d2d) and wafer-to-wafer (w2w) stacking problems are transferred to the bipartite maximal matching problem. Then, heuristic algorithms are also proposed to optimize the stacking yield. Cheng-Wen Wu, Shyue-Kung Lu, Jin-Fu Li 0001 |
ASP-DAC | 1 |
| 2012 | A memory yield improvement scheme combining built-in self-repair and error correction codesabstractError correction code (ECC) and built-in self-repair (BISR) schemes have been wildly used for improving the yield and reliability of memories. Many built-in redundancy-analysis (BIRA) algorithms and ECC schemes have been reported before. However, most of them focus on either BIRA algorithms or ECC schemes. In this paper, we propose an ECC-Enhanced Memory Repair (EEMR) scheme for yield improvement. Many modern memories are equipped with ECC in addition to BISR. We evaluate the back-end flow that combines both ECC and BIRA to determine whether yield can be improved by proper sequencing of the two steps. We also collect and identify important failure patterns and their distributions from over 100,000 sample memory instances, which are used to enhance the EEMR scheme that incorporates ECC. As ECC is failure pattern sensitive, careful evaluation from realistic failure bitmaps is necessary. We also verify the feasibility of implementing the proposed EEMR scheme by real test data. Experimental results from industrial 4Mb memory instances show that the proposed EEMR scheme gains over 2% instance yield on average, as compared with the traditional scheme. We also investigate the reliability of the EEMR scheme with different ECC specifications and BIRA algorithms. Tze-Hsin Wu, Po-Yuan Chen, Mincent Lee, Bin-Yen Lin, Cheng-Wen Wu, Chen-Hung Tien, Hung-Chih Lin, Hao Chen 0053, Ching-Nen Peng, Min-Jer Wang |
ITC | 5 |
| 2012 | A built-in self-test scheme for 3D RAMsabstractThree-dimensional (3D) random access memory (RAM) using through-silicon vias for inter-die interconnects has been considered as a new approach to overcome the memory wall. In this paper, we propose a built-in self-test (BIST) scheme for 3D RAMs. In the BIST scheme, a clock-domain-crossing-aware test pattern generator is proposed to cope with the clock-domain-crossing issue. An inter-die synchronization mechanism is also proposed to synchronize the BIST circuits in different dies. Furthermore, the BIST circuit provides the high-programmability feature to support the selection of RAMs in a die for testing such that it can support thermal management during the test. We design the proposed BIST scheme in a 3D IC with processor and RAM dies. Experimental results show that the area cost of the BIST circuit is very small. The area overhead of the BIST circuit for four 8192×64-bit RAMs in a die is only 0.45% using TSMC 90nm 1P9M CMOS process technology. Yun-Chao You, Che-Wei Chou, Jin-Fu Li 0001, Chih-Yen Lo, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
ITC | 7 |
| 2012 | Cost modeling and analysis for interposer-based three-dimensional ICabstractThree-dimensional (3D) integration has recently become a popular technology for integrated circuits (IC). 3D IC with the passive silicon interposer is currently the main trend in the industry, especially for processor-memory integration. Evaluating the economic efficiency of test operations in the interposer-based 3D IC thus is important. We propose a cost model for the Die-to-Wafer (D2W) and Die-to-Die (D2D) stacking, including manufacturing cost and test cost. A tool which is based on the proposed cost model is developed. We use this tool for cost analysis and for finding the most cost effective test flow. The results show that, in some applications, test flows including the iterative known-good stack (KGS) test and the pre-bond interposer test significantly reduce the cost, when the KGS test yield is lower than 98.2% and the pre-bond interposer test yield is lower than 99.38%. A Shmoo plot is depicted to show the lower bound of the yield of the final package level test, given the number of stacked dies and the final yield. For different applications, the proposed model evaluates the critical yield or cost values, which helps the designers to determine the most cost effective test flow and the system architecture. Ying-Wen Chou, Po-Yuan Chen, Mincent Lee, Cheng-Wen Wu |
VTS | 4 |
| 2012 | A Memory Failure Pattern Analyzer for memory diagnosis and repairabstractAs VLSI technology advances and memories occupy more and more area in a typical SOC, memory diagnosis has become an important issue. In this paper, we propose the Memory Failure Pattern Analyzer (MFPA), which is developed for different memories and technologies that are currently used in the industry. The MFPA can locate weak regions of the memory array, i.e., those with high failure rate. It can also be used to analyze faulty-cell/defect distributions automatically. We also propose a new defect distribution model which has 1–12 times higher accuracy than other theoretical models. Based on this model, we propose a defect-spectrum-based methodology to identify critical failure patterns from failure bitmaps. These failure patterns can further be translated to corresponding defects by our memory fault simulator (RAMSES) and physical-level failure analysis tool (FAME). In an industrial case, the MFPA fits the defect distribution with the proposed model, which has 12 times higher accuracy than the Poisson distribution. With our model, it further identifies two special failure patterns from 132,488 faulty 4-Mb macros in 1.2 minutes. Bing-Yang Lin, Mincent Lee, Cheng-Wen Wu |
VTS | 3 |
| 2011 | A self-testing and calibration method for embedded successive approximation register ADCabstractThis paper presents a self-testing and calibration method for the embedded successive approximation register (SAR) analog-to-digital converter (ADC). We first propose a low cost design-for-test (DfT) technique which tests a SAR ADC by characterizing its digital-to-analog converter (DAC) capacitor array. Utilizing DAC major carrier transition testing, the required analog measurement range is just 4 LSBs; this significantly lowers the test circuitry complexity. Then, we develop a fully-digital missing code calibration technique that utilizes the proposed testing scheme to collect the required calibration information. Simulation results are presented to validate the proposed technique. Xuan-Lun Huang, Ping-Ying Kang, Hsiu-Ming Chang 0001, Jiun-Lang Huang, Yung-Fa Chou, Yung-Pin Lee, Ding-Ming Kwai, Cheng-Wen Wu |
ASP-DAC | 8 |
| 2011 | Multi-visit TAMs to Reduce the Post-Bond Test Length of 2.5D-SICs with a Passive Silicon Interposer Baseabstract2.5D Stacked ICs (2.5D-SICs) consist of multiple active dies (or 3D towers of active dies), which are placed side-by-side on top of and interconnected through a passive silicon interposer base which contains Through-Silicon Vias (TSVs). A previously presented post-bond test and Design-for-Test(DfT) strategy for such 2.5D-SICs implements a serial Test Access Mechanism (TAM) for interposer and micro-bump testing. In addition, it tries to identify an as-wide-as-possible set of functional interposer interconnects that can be reused as parallel TAMs to the various dies. In this paper, we extend that approach with the concept of Multi-Visit TAMs, i.e., parallel TAMs which are allowed to visit the same die more than once. For minimal additional hardware costs, the Multi-Visit TAMs succeed significantly more often in identifying a valid parallel TAM and achieve significantly lower test lengths. Chun-Chuan Chi, Erik Jan Marinissen, Sandeep Kumar Goel, Cheng-Wen Wu |
Asian Test Symposium | 4 |
| 2011 | A low-cost wireless interface with no external antenna and crystal oscillator for cm-range contactless testingabstractThis work presents a low-cost wireless system design that serves as an interface to support the SoC with contactless testability feature. The communication hierarchy includes PHY, MAC, data exchange, and test wrapper functions. The wireless does not require external antennae and crystal reference, and therefore minimize the setup cost. The embedded all-digital timing generation achieves robust performance in the noisy environment. The whole wireless system occupies a small area. In a 0.18μm device-under-test, the active area of wireless front-end is 0.14mm2 and the gate count for digital processing is 112K. The maximum energy efficiency for uplink is 1.1nJ/bit and for downlink is 2.9nJ/bit when the wireless distance is set around 1cm. The prototype system includes test equipment and an SoC as the device-under-test. The SoC integrating logic, memory, and analog plug-in modules can be contactlessly tested. It is a low-cost platform controlled by a simple hand-held computer. Chin-Fu Li, Chi-Ying Lee, Chen-Hsing Wang, Shu-Lin Chang, Li-Ming Denq, Chun-Chuan Chi, Hsuan-Jung Hsu, Ming-Yi Chu, Jing-Jia Liou, Shi-Yu Huang, Po-Chiun Huang, Hsi-Pin Ma, Jenn-Chyou Bor, Cheng-Wen Wu, Ching-Cheng Tien, Chi-Hu Wang, Yung-Sheng Kuo, Chih-Tsun Huang, Tien-Yu Chang |
DAC | 14 |
| 2011 | DfT Architecture for 3D-SICs with Multiple TowersabstractThree-dimensional stacked ICs (3D-SICs) based on Through-Silicon Vias (TSVs) provide attractive benefits such as smaller form factor, higher performance, and lower power. So far, prior work on Design-for-Testability (DfT) only focused on 3D-SICs consisting of a single "tower", i.e., a 3D-SIC in which each stack level contains exactly one die. 3D stacking technology allows to place multiple dies on top of a common base die, resulting in 3D-SICs with multiple "towers". This paper presents a generic DfT architecture for 3D-SICs having any number of "towers", possibly including "sub-towers". We also present efficient test control mechanisms. Experimental results show that the proposed architecture has a negligible area cost for medium-sized and larger industrial designs, and therefore provides a cost-effective test solution for 3D-SICs. Chun-Chuan Chi, Erik Jan Marinissen, Sandeep Kumar Goel, Cheng-Wen Wu |
ETS | 4 |
| 2011 | Post-bond testing of 2.5D-SICs and 3D-SICs containing a passive silicon interposer baseabstractThrough-Silicon Vias (TSVs) enable high-density, low-latency, and low-power interconnects for system chips that consist of multiple dies. In “2.5D” Stacked ICs (2.5D-SICs), multiple dies without TSVs are stacked side-by-side on top of a passive silicon interposer base containing TSVs. In true 3D-SICs, multiple dies containing TSVs themselves are vertically stacked; one or multiple of such stacks are possibly placed on a passive silicon interposer. This paper proposes a post-bond test and design-for-test (DfT) strategy for 2.5D- and 3D-SICs containing a passive silicon interposer base. Functional interconnects in the interposer are reused as much as possible in order to keep the interposer cost low. Chun-Chuan Chi, Erik Jan Marinissen, Sandeep Kumar Goel, Cheng-Wen Wu |
ITC | 4 |
| 2011 | A built-in self-test scheme for the post-bond test of TSVs in 3D ICsabstractThree-dimensional (3D) integration using through silicon via (TSV) has been widely acknowledged as one future integrated-circuit (IC) technology. A 3D IC including multiple dies connected with TSVs offers many benefits over current 2D ICs. However, the testing of 3D ICs is much more difficult than that of 2D ICs. In this paper, we propose a cost-effective built-in self-test circuit (BIST) to test TSVs of a 3D IC. The BIST scheme, arranging the TSVs into arrays similar to memory, has the features of low test/diagnosis time and low silicon area cost. Simulation results show that the area overhead of the BIST circuit implemented with 0.18μm CMOS technology for a 16×32 TSV array in which each TSV cell size is 45μm2is 2.24%. Also, the BIST needs only 130 clock cycles to test the TSV array with stuck-at faults. In comparison with the IEEE 1500-based test approach, the BIST scheme can achieve 85.2% area cost and 93.6% test time reduction. Yu-Jen Huang, Jin-Fu Li 0001, Ji-Jan Chen, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
VTS | 6 |
| 2011 | Training-based forming process for RRAM yield improvementabstractOver the past decade, the resistive memory device known as RRAM has been studied extensively in many ways, and many of its problems have been identified, discussed, and some solved. It is time to move from material, process, and device to circuit design and yield, in order to commercialize RRAM. However, as we move from resistive device to memory circuit, new problems do appear, partly because the operating conditions of resistive devices on real RRAM circuit differ from those in an experimental environment for single devices. In this paper, an over forming problem has been identified from our analysis, and we propose a solution based on training sequence. As a result, by solving the over forming problem, RRAM yield can be improved significantly. RRAM; forming process; training sequence; yield improvement; non-volatile memory; memory testing Hsiu-Chuan Shih, Ching-Yi Chen, Cheng-Wen Wu, Chih-He Lin, Shyh-Shyuan Sheu |
VTS | 3 |
| 2011 | Special session: Hot topic design and test of 3D and emerging memoriesabstractIn this hot topic session, we are including three talks that cover design of reliable, emerging memories, with emphasis on 3D memories, DRAM and non-volatile memories. We will also introduce a new class of memory, the Storage Class Memory (SCM), and discuss its reliability issues. Cheng-Wen Wu |
VTS | 1 |
| 2011 | A Memory Built-In Self-Repair Scheme Based on Configurable SparesabstractThere is growing need for embedded memory built-in self-repair (MBISR) due to the introduction of more and more system-on-chip (SoC) and other highly integrated products, for which the chip yield is being dominated by the yield of on-chip memories, and repairing embedded memories by conventional off-chip schemes is expensive. Therefore, we propose an MBISR generator called BRAINS+, which automatically generates register transfer level MBISR circuits for SoC designers. The MBISR circuit is based on a redundancy analysis (RA) algorithm that enhances the essential spare pivoting algorithm, with a more flexible spare architecture, which can configure the same spare to a row, a column, or a rectangle to fit failure patterns more efficiently. The proposed MBISR circuit is small, and it supports at-speed test without timing-penalty during normal operation, e.g., with a typical 0.13 μm complementary metal-oxide-semiconductor technology, it can run at 333 MHz for a 512 Kb memory with four spare elements (rows and/or columns), and the MBISR area overhead is only 0.36%. With its low area overhead and zero test-time penalty, the MBISR can easily be applied to multiple memories with a distributed RA scheme. Compared with recent studies, the proposed scheme is better in not only test-time but also area overhead. Mincent Lee, Li-Ming Denq, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Yield Enhancement by Bad-Die Recycling and Stacking With Though-Silicon Viasabstract3-D integration provides a means to overcome the difficulties in design and manufacturing of system-on-chip (SOC) and memory products. Introducing a short vertical interconnect, called through-silicon via (TSV), makes it feasible to repair and recycle bad dies by stacking. We propose a method to accomplish this using a dual-TSV hardwired switch (DTHS) in which the via-hole location is programmable. With the DTHS, we activate a spare and establish inter-die routing. The spare is nothing but a good part in another bad die. To be 3-D reparable, the design is partitioned into disjoint parts. The effort for the modification is minor in view of that a typical SOC is readily composed of modules with predefined functions and supply voltages. The DTHS is used: 1) to shut off power connections of both failed and unused parts; 2) to disconnect their signal paths; and 3) to redirect them to the selected good parts in the stacked dies. Despite the speed is degraded due to the extra load incurred by the DTHS, our simulation shows that the increase in delay time can be limited below 100 ps with an over-designed buffer which occupies 0.8% of the area of a 30 μm TSV, using a 65-nm CMOS process. The performance degradation turns out to be a necessary evil, since the increased height of the die stack leads to a thermal conductivity poorer than its 2-D counterpart. The 3-D patch die helps to shorten time-to-market and turn the irreparable dies profitable. Yung-Fa Chou, Ding-Ming Kwai, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | A Built-in Self-Diagnosis and Repair Design With Fail Pattern Identification for MemoriesabstractWith the advent of deep-submicrometer VLSI technology, the capacity and performance of semiconductor memory chips is increasing drastically. This advantage also makes it harder to maintain good yield. Diagnostics and redundancy repair methodologies thus are getting more and more important for memories, including embedded ones that are popular in system chips. In this paper, we propose an efficient memory diagnosis and repair scheme based on fail-pattern identification. The proposed diagnosis scheme can distinguish among row, column, and word faults, and subsequently apply the Huffman compression method for fault syndrome compression. This approach reduces the amount of data that need to be transmitted from the chip under test to the automatic test equipment (ATE) without losing fault information. It also simplifies the analysis that has to be performed on the ATE. The proposed redundancy repair scheme is assisted by fail-pattern identification approach and a flexible redundancy structure. The area overhead for our built-in self-repair (BISR) design is reasonable. Our repair scheme uses less redundancy than other redundancy schemes under the same repair rate requirement. Experimental results show that the area overhead of the BISR design is only 4.1% for an 8 K × 64 memory and is in inverse proportion to the memory size. Chin-Lung Su, Rei-Fu Huang, Cheng-Wen Wu, Kun-Lun Luo, Wen Ching Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Is 3D integration an opportunity or just a hype?abstractThree-dimensional (3D) integration using through silicon via (TSV) is an emerging technology for integrated circuit designs. 3D integration technology provides numerous opportunities to designers looking for more cost-effective system chip solutions. In addition to stacking homogeneous memory dies, 3D integration technology supports heterogeneous integration of memories, logic, sensors, etc. It eases the interconnect performance limitation, provides higher functionality, results in small form factor, etc. On the other hand, there are challenges that should be overcome before volume production of TSV-based 3D ICs becomes possible, e.g., technological challenges, yield and test challenges, thermal and power challenges, infrastructure challenges, etc. Jin-Fu Li 0001, Cheng-Wen Wu |
ASP-DAC | 2 |
| 2010 | A Test Integration Methodology for 3D Integrated CircuitsabstractThe three-dimensional (3D) integration technology using through silicon via (TSV) provides many benefits over the 2D integration technology. Although many different manufacturing technologies for 3D integrated circuits (ICs) have been presented, some challenges should be overcome before the volume production of 3D ICs. One of the challenges is the testing of 3D ICs. This paper proposes test integration interfaces for controlling the design-for-test circuits in the dies of a 3D IC. The test integration interfaces can support the pre-bond, known-good stack, and post-bond tests. The minimum number of required test pads of the proposed test interface for pre-bond test using is only four. Furthermore, the test interface is compatible with the IEEE 1149.1 standard for the board-level testing. Simulation results show that the area overhead of the proposed test interfaces for a 3D IC with two dies in which each die implements the function of ITC'99 b19 benchmark is only about 0.15%. Che-Wei Chou, Jin-Fu Li 0001, Ji-Jan Chen, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
Asian Test Symposium | 6 |
| 2010 | Performance Characterization of TSV in 3D IC via Sensitivity AnalysisabstractIn this paper, we propose a method that can characterize the propagation delays across the Through Silicon Vias (TSVs) in a 3D IC. We adopt the concept of the oscillation test, in which two TSVs are connected with some peripheral circuit to form an oscillation ring. Upon this foundation, we propose a technique called sensitivity analysis to further derive the propagation delay of each individual TSV participating in the oscillation ring-a distilling process. In this process, we perturb the strength of the two TSV drivers, and then measure their effects in terms of the change of the oscillation ring's period. By some following analysis, the propagation delay of each TSV can be revealed. Monte-Carlo analysis of a typical TSV with 30% process variation on transistors shows that the characterization error of this method is only 2.1% with the standard deviation of 8.1%. Jhih-Wei You, Shi-Yu Huang, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
Asian Test Symposium | 5 |
| 2010 | An error tolerance scheme for 3D CMOS imagersabstractA three-dimensional (3D) CMOS imager constructed by stacking a pixel array of backside illuminated sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array using micro-bumps (μbumps) and through-silicon vias (TSVs) is promising for high throughput applications. However, due to the direct mapping from pixels to ISPs, the overall yield relies heavily on the correctness of the μbumps, ADCs and TSVs -- a single defect leads to the information loss of a tile of pixels. This paper presents an error tolerance scheme for the 3D CMOS imager that can still deliver high quality images in the presence of μbump, ADC, and/or TSV failures. The error tolerance is achieved by properly interleaving the connections from pixels to ADCs so that the corrupted data, if any, can be recovered in the ISPs. A key design parameter, the interleaving stride, is decided by analyzing the employed error correction algorithm. Architectural simulation results demonstrate that the error tolerance scheme enhances the effective yield of an exemplar 3D imager from 46% to 99%. Hsiu-Ming Chang 0001, Jiun-Lang Huang, Ding-Ming Kwai, Kwang-Ting Cheng, Cheng-Wen Wu |
DAC | 5 |
| 2010 | Fast identification of operating current for toggle MRAM by spiral searchabstractMagnetic Random Access Memory (MRAM) is a non-volatile memory which is widely studied for its high speed, high density, small cell size, and almost unlimited endurance. However, for deep-submicron process technologies, significant variation in MRAM cells' operating regions results in write failures in cells and reduces the production yield. Currently, memory designers characterize failed MRAM chips to find a suitable current level for reconfiguring their operating current, which is time-consuming. In this paper, we propose an efficient operating current search method and a built-in circuit for toggle MRAM, which can rapidly find a customized operating current for each MRAM chip. With the built-in circuit, an MRAM chip can dynamically reconfigure its operating current automatically. Production yield and product life-time thus can be increased. Sheng-Hung Wang, Ching-Yi Chen, Cheng-Wen Wu |
DAC | 3 |
| 2010 | An adaptive code rate EDAC scheme for random access memoryabstractAs the VLSI technology scaling continues and the device dimension keeps shrinking, memories are more and more sensitive to soft errors. Memory cores usually occupy a large portion of an SOC and have significant impact on the chip reliability. Therefore error detection and correction (EDAC) techniques are commonly used for protecting the system against soft errors. This paper presents a novel EDAC scheme, which provides adaptive code rate for random access memories (RAMs). Under a certain reliability restriction, the proposed design allows more error bits than a conventional EDAC design. Ching-Yi Chen, Cheng-Wen Wu |
DATE | 2 |
| 2010 | A low-cost and scalable test architecture for multi-core chipsabstractMulti-core architecture has become a mainstream in modern processor and computation-intensive chips. A widely-used multi-core architecture contains identical cores. This paper proposes a low-cost and scalable test architecture for a multi-core chip with identical cores. The test architecture provides test scalability by using a two-dimensional pipelined test access mechanism (TAM). Also, some scan cells of the cores under test are reused as the pipeline registers of the TAM such that the area cost of the proposed test architecture is low. Experimental results show that the proposed test architecture only consumes about 2.6% area for a multi-core chip with 16 Advanced Encryption Standard (AES) cores. Also, the test time for 16 AES cores is only about 1.004 times of that for a single AES core. Chun-Chuan Chi, Cheng-Wen Wu, Jin-Fu Li 0001 |
ETS | 2 |
| 2010 | On-chip testing of blind and open-sleeve TSVs for 3D IC before bondingabstractPre-bond test is preferred for a three-dimensional integrated circuit (3D IC), since it reduces stacking yield loss and thus saves cost. In this paper, we present two schemes for testing through-silicon vias (TSVs) by performing on-chip screening before wafer thinning and bonding. The first scheme is for blind TSVs, which have one end floating, using a charge-sharing technique commonly seen in DRAM. The second scheme is for open-sleeve TSVs, which have one end shorted to the substrate, using a voltage-dividing technique commonly seen in ROM. By virtue of the inherent capacitive and resistive characteristics, we detect the TSVs out of a specified range as anomalies, taking into account the effects of process variations in the detection circuitry. The statistical design by Monte Carlo simulation using TSMC 65nm low-power process shows that for blind TSVs, the best overkill ratio is below 6%. For open-sleeve TSVs, inherent limitations restrict the applicability, so more work needs to be done in the future. Our implementation enjoys little area overhead, requiring only a simple sense amplifier and a write buffer that are shared among a number of TSVs. Reducing the number of TSVs that share a test module will reduce the test time, but increase the area overhead. For blind TSVs, the parallelism also affects the overkill and escape rates. Po-Yuan Chen, Cheng-Wen Wu, Ding-Ming Kwai |
VTS | 2 |
| 2010 | Built-In Self-Repair Schemes for Flash MemoriesabstractThe advancement of deep submicrometer Integrated circuit manufacturing technology has pushed the use of embedded memory, and the strong demand of embedded nonvolatile memory for system-on-chip and system in package applications has made flash memory increasingly important as well. Nevertheless, the yield loss of memory products caused by deep submicrometer defects and manufacturing uncertainties is still a critical issue. In order to solve the yield issue, built-in self-repair (BISR) has been considered as the most cost-effective solution. However, implementing BISR on flash memories is not trivial. In this paper, we propose BISR schemes for nor flash memory and NAND flash memory, respectively. The BISR schemes perform built-in self-test, built-in redundancy analysis, and on-chip repair. For the BISR scheme of nor flash memory, a typical redundancy architecture is assumed, based on which we analyze three existing algorithms and propose a redundancy analysis (RA) algorithm. On the other hand, for NAND flash memory, an RA algorithm based on an efficient 2-D redundancy architecture is proposed, and considering the widely used page-mode operation in NAND flash memory, a method to discover currently accessed address is also proposed. A simulation tool is also developed, supporting nor flash memory and NAND flash memory. The simulation results show that our approach can effectively repair defective memories. Yu-Ying Hsiao, Chao-Hsun Chen, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | SOC Test Architecture and Method for 3-D ICsabstract3-D integration provides another way to put more devices in a smaller footprint. However, it also introduces new challenges in testing. Flexible test architecture named test access control system for 3-D integrated circuits (TACS-3D) is proposed for 3-D integrated circuits (IC) testing. Integration of heterogeneous design-for-testability methods for logic, memory, and through-silicon via (TSV) testing further reduces the usage of test pins and TSVs. To highly reuse pre-bond test circuits in post-bond test, an innovative linking mechanism shares TSVs and test pins of the 3-D IC. No matter how many layers are there in the 3-D IC, a large portion of TSVs and test pins is reserved for data application. Therefore, smaller post-bond test time is expected. A test chip composed of a network security processor platform is taken as an example. Less than 0.4% test overhead increases in area and time between 2-D and 3-D cases. Compared with the instinctively direct access, TACS-3D reveals up to 54% test time improvement under the same TSV usage. Chih-Yen Lo, Yu-Tsao Hsing, Li-Ming Denq, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Efficient BISR Techniques for Embedded Memories Considering Cluster FaultsabstractInstead of the traditional spare row/column redundancy architectures, block-based redundancy architectures are proposed in this paper. The redundant rows/columns are divided into row/column blocks. Therefore, the repair of faulty memory cells can be performed at the row/column-block level. Moreover, the redundant row/column blocks can be used to replace faulty cells anywhere in the memory array. This global characteristic is helpful for repairing cluster faults. The proposed redundancy architecture can be easily integrated with the embedded memory cores. Based on the proposed global redundancy architecture, a heuristic modified essential spare pivoting (MESP) algorithm suitable for built-in implementation is also proposed. According to experimental results, the area overhead for implementing the MESP algorithm is very low. Due to efficient usage of redundancy, the manufacturing yield, repair rate, and reliability can be improved significantly. Shyue-Kung Lu, Chun-Lin Yang, Yuang-Cheng Hsiao, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Diagnosis of MRAM Write Disturbance FaultabstractIn this paper, we propose a new test method to detect write disturbance fault (WDF) for magnetic RAM (MRAM). Furthermore, an adaptive diagnosis algorithm (ADA) is also introduced to identify and diagnose the WDF for MRAM. The proposed test method can evaluate process stability and uniformity. We also develop a built-in self-test (BIST) circuit that supports the proposed WDF diagnosis test method. A 1-Mb toggle MRAM prototype chip with the proposed BIST circuit has been designed and fabricated using a special${0.15}\hbox{-}\mu{\hbox {m}}$CMOS technology. The BIST circuit overhead is only about 0.05% with respect to the 1-Mb MRAM. The test time is reduced by about 30% as compared with the test method without using the decision write mechanism. The chip measurement results show the efficiency of our proposed method. Chin-Lung Su, Chih-Wea Tsai, Ching-Yi Chen, Wan-Yu Lo, Cheng-Wen Wu, Ji-Jan Chen, Wen Ching Wu, Chien-Chung Hung, Ming-Jer Kao |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | An Efficient Multimode Multiplier Supporting AES and Fundamental Operations of Public-Key CryptosystemsabstractThis paper presents a highly efficient multimode multiplier supporting prime field, namely, polynomial field, and matrix–vector multiplications based on an asymmetric word-based Montgomery multiplication (MM) algorithm. The proposed multimode 128$\times$32 b multiplier provides throughput rates of 441 and 511 Mb/s for 256-b operands over$GF(P)$and$GF(2^n)$at a clock rate of 100 MHz, respectively. With 21 930 additional gates for Advanced Encryption Standard (AES), the multiplier is extended to provide 1.28-, 1.06-, and 0.91-Gb/s throughput rates for 128-, 192-, and 256-b keys, respectively. The comparison result shows that the proposed integration architecture outperforms others in terms of performance and efficiency for both AES and MM that is essential in most public-key cryptosystems. Chen-Hsing Wang, Chieh-Lin Chuang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Single- and Multi-core Configurable AES Architectures for Flexible SecurityabstractAs networking technology advances, the gap between network bandwidth and network processing power widens. Information security issues add to the need for developing high-performance network processing hardware, particularly that for real-time processing of cryptographic algorithms. This paper presents a configurable architecture for Advanced Encryption Standard (AES) encryption, whose major building blocks are a group of AES processors. Each AES processor provides 219block cipher schemes with a novel on-the-fly key expansion design for the original AES algorithm and an extended AES algorithm. In this multicore architecture, the memory controller of each AES processor is designed for the maximum overlapping between data transfer and encryption, reducing interrupt handling load of the host processor. This design can be applied to high-speed systems since its independent data paths greatly reduces the input/output bandwidth problem. A test chip has been fabricated for the AES architecture, using a standard 0.25-¿m CMOS process. It has a silicon area of 6.29 mm2, containing about 200,500 logic gates, and runs at a 66-MHz clock. In electronic codebook (ECB) and cipher-block chaining (CBC) cipher modes, the throughput rates are 844.9, 704, and 603.4 Mb/s for 128-, 192-, and 256-b keys, respectively. In order to achieve 1-Gb/s throughput (including overhead) at the worst case, we design a multicore architecture containing three AES processors with 0.18-¿m CMOS process. The throughput rate of the architecture is between 1.29 and 3.75 Gb/s at 102 MHz. The architecture performs encryption and decryption of large data with 128-b key in CBC mode using on-the-fly key generation and composite field S-box, making it more cost effective (with better thousand-gate/gigabit-per-second ratio) than conventional methods. Mao-Yin Wang, Chih-Pin Su, Chia-Lung Horng, Cheng-Wen Wu, Chih-Tsun Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | A Mesh-Structured Scalable IPsec ProcessorabstractIP security (IPsec) protocols are widely used to protect sensitive data over the Internet. For equipment linked by high-bandwidth optical fibers, the throughput requirement usually results in the adoption of high-performance network security processors. In this paper, we propose a parallel mesh-structured IPsec (MIPsec) processor, which executes the IPsec protocols for Internet security gateway applications. We have developed several area-efficient cryptographic IPs embedded in MIPsec to lower silicon cost. Thanks to structural regularity, the simple deterministic programming of MIPsec guarantees high utilization of the hardware. Also, both handshake and contention issues are solved in the scheme, such that performance can be scaled up. Specifically, the 6.23-million-gate MIPsec achieves 10-Gb/s wire speed for each routing direction. The proposed MIPsec is suitable for transport mode or other crypto mix as well. Mao-Yin Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | On-Chip TSV Testing for 3D IC before Bonding Using Sense AmplificationabstractWe present a novel testing scheme for TSVs in a 3D IC by performing on-chip TSV monitoring before bonding, using a sense amplification technique that is commonly seen on a DRAM. By virtue of the inherent capacitive characteristics, we can detect the faulty TSVs with little area overhead for the circuit under test. Po-Yuan Chen, Cheng-Wen Wu, Ding-Ming Kwai |
Asian Test Symposium | 2 |
| 2009 | Test Integration for SOC Supporting Very Low-Cost TestersabstractTo reduce test cost for SOC products, it is important to reduce the cost of testers. When using low-cost testers which have a limited test bandwidth to perform testing, built-in- self-test (BIST) is necessary to reduce the data volume to be transmitted between the tester and the device-under-test (DUT). We enhance the SOC test integration tool, STEAC, so that it can support SOCs containing BISTed cores which are to be tested by low-cost testers. A test chip is implemented to verify the proposed technique. Experimental results show that the enhanced STEAC successfully works with the HOY wireless test system and other low-cost testers. Chun-Chuan Chi, Chih-Yen Lo, Te-Wen Ko, Cheng-Wen Wu |
Asian Test Symposium | 4 |
| 2009 | An Adaptive-Rate Error Correction Scheme for NAND Flash MemoryabstractECC has been widely used to enhance flash memory endurance and reliability. In this work, we propose an adaptive-rate ECC scheme with BCH codes that is implemented on the flash memory controller. With this scheme, flash memory can trade storage space for higher error correction capability to keep it usable even when there is a high noise level. Te-Hsuan Chen, Yu-Ying Hsiao, Yu-Tsao Hsing, Cheng-Wen Wu |
VTS | 4 |
| 2008 | Area and Test Cost Reduction for On-Chip Wireless Test Channels with System-Level Design TechniquesabstractWith continuing trends to embed more on-chip test circuits, increasing complexity requires more efforts on design and validation. In this paper, we use a wireless test system as an example, to demonstrate the efficiency of system-level techniques in assisting circuit specification exploration, with the goal of area and test-cost reduction. In our experiments, 30% to 50% total costs are saved compared to an initial ad-hoc setup. Chun-Kai Hsu, Li-Ming Denq, Mao-Yin Wang, Jing-Jia Liou, Chih-Tsun Huang, Cheng-Wen Wu |
ATS | 6 |
| 2008 | Test and Diagnosis Algorithm Generation and Evaluation for MRAM Write Disturbance FaultabstractWe proposed the systematic tools, RAMSES-M and TAGS-M, for test and diagnosis algorithms evaluation and development, respectively. In addition to traditional memory fault models, the tools support the MRAM specific fault model, Write Disturbance Fault (WDF) and its specific test operation, Read-previous, which is proposed in this paper, too. The concept of Weighted Fault Coverage (WFC) is introduced and adopted by RAMSES-M. Several test and diagnosis algorithms generated by the proposed tool are compared with other conventional March algorithms. The results show that the proposed algorithms have better performance for testing and diagnosis. Wan-Yu Lo, Ching-Yi Chen, Chin-Lung Su, Cheng-Wen Wu |
ATS | 4 |
| 2008 | Write Disturbance Modeling and Testing for MRAMabstractThe magnetic random access memory (MRAM) is considered one of the potential candidates that will replace current on-chip memories (RAM, EEPROM, and flash memory) in the future. The MRAM is fast and does not need a high supply voltage for read/write operations, and is compatible with the CMOS technology. It can also endure almost unlimited read/write cycles. These combined advantages of RAM and flash memory make it a potential choice for SOC. In this paper, we present the write disturbance fault (WDF) model for MRAM, i.e., a fault that affects the data stored in the MRAM cells due to excessive magnetic field during the write operation. We also construct the SPICE macro model for the magnetic tunneling junction (MTJ) device of the toggle MRAM to obtain circuit simulation results. We then present an MRAM fault simulator called RAMSES-M, based on which we derive the shortest test for the proposed WDF model. The test is shown to be better and more robust as compared with the conventional March C-test algorithm. We also present a March 17 N diagnosis algorithm for identifying WDF. A 1 Mb MRAM chip has been designed and fabricated using a CMOS-based 0.18-mum technology. The proposed WDF model is justified by chip measurement results, with the march test results reported. Finally, specific MRAM fault behavior and test issues are discussed. Chin-Lung Su, Chih-Wea Tsai, Cheng-Wen Wu, Chien-Chung Hung, Young-Shying Chen, Ding-Yeong Wang, Yuan-Jen Lee, Ming-Jer Kao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | Test Education in the Global EconomyabstractThere is an increasing demand for test and diagnosis expertise in the global semiconductor industry, in sectors ranging from foundries to test houses, to IDM companies, and from fabless design houses to EDA companies. Test education, however remains a niche, highly specialized subject area in the graduate curriculum and is seldom covered in undergraduate classes. In this panel, we evaluate the current health of academic test education and debate the role and goals of future test education, as well as those changes that need to be made to meet the global market demands. Jacob Abraham, Salvador Mir, Yinghua Min, Jeremy Wang, Cheng-Wen Wu |
ATS | 5 |
| 2007 | A Hybrid BIST Scheme for Multiple Heterogeneous Embedded MemoriesabstractIt is common that an SOC contains hundreds or even thousands of heterogeneous embedded memories. Many of these embedded memories have wide data words, leading to high routing penalty from the BIST circuits. Previous BIST schemes solve the problem using serial interface, e.g., based on the IEEE 1500 architecture and novel scan approaches, to reduce the routing area overhead. However, serial approaches do not allow at-speed test and diagnosis, and are very slow. In this paper, we propose a hybrid BIST architecture that reduces the routing penalty, while allowing at-speed test and diagnosis of the memory cores. The test time is close to that of a typical parallel BIST method. Experimental results show that the proposed BIST can effectively reduce the area overhead. Li-Ming Denq, Cheng-Wen Wu |
ATS | 2 |
| 2007 | CAMEL: An Efficient Fault Simulator with Coupling Fault Simulation Enhancement for CAMsabstractContent addressable memories (CAMs) are widely used in digital systems. A test algorithm for CAMs must be able to cover the random access memory (RAM) faults and comparison faults. However, CAM circuits are usually customized for different products, so there are no standard tests, i.e., tests should be adapted to a specific design manufactured using specific technology. This paper presents a fault simulator, called CAM Evaluation tooL (CAMEL), for the evaluation of fault coverage of CAM test algorithms. It supports five common functional outputs, i.e., Data I/O, hit, multi-hit, matchout, and priority address for various CAM specifications. Since coupling fault simulation dominates the efficiency of a memory fault simulator, a concept of observability is proposed to simplify the coupling fault behavior. By exploiting the observability, a compression technique is also proposed to speed up the fault simulation and reduce memory usage. CAMEL can support both RAM faults and comparison faults. We have demonstrated the CAMEL using widely-used March tests and CAM tests. Simulation results show that the CAMEL can evaluate the fault coverage of tests accurately and efficiently. Hsiang-Huang Wu, Jin-Fu Li 0001, Chi-Feng Wu, Cheng-Wen Wu |
ATS | 4 |
| 2007 | Diagnosis for MRAM write disturbance faultabstractTo help improve quality and yield of magnetic random access memory (MRAM), we propose an adaptive diagnosis algorithm (ADA) that can efficiently identify the write disturbance fault (WDF) for MRAM. The proposed test algorithm is a Marchbased one, i.e., it has linear time complexity and can easily be implemented with built-in self-test (BIST). However, the proposed test method can evaluate the process stability and uniformity using logical test method. We also develop a BIST circuit that supports the proposed WDF diagnosis test method. We propose the BIST scheme based on the Decision Write mechanism of the toggle MRAM to reduce total test time. A 1Mb toggle MRAM prototype chip with the proposed BIST circuit has been designed and fabricated using a special 0.15µm CMOS technology. The BIST circuit overhead is only about 0.04% with respect to the 1Mb MRAM. The test time is reduced by about 30% as compared with the test method without using the Decision Write mechanism. Chin-Lung Su, Chih-Wea Tsai, Cheng-Wen Wu, Ji-Jan Chen, Wen Ching Wu, Chien-Chung Hung, Ming-Jer Kao |
ITC | 3 |
| 2007 | SDRAM Delay Fault Modeling and Performance TestingabstractDRAM timing parameter testing has always been considered a time-consuming process. This paper presents a systematic approach to analysis and classification of the synchronous DRAM (SDRAM) delay failure modes. Four delay fault models with March expression are proposed to cover important DRAM timing parameters. By at-speed March testing of these four types of delay faults, the authors can verify the DRAM timing specifications. Yu-Tsao Hsing, Chun-Chieh Huang, Jen-Chieh Yeh, Cheng-Wen Wu |
VTS | 4 |
| 2007 | Flash Memory Testing and Built-In Self-Diagnosis With March-Like Test AlgorithmsabstractFlash memories are a type of nonvolatile memory based on floating-gate transistors. The use of commodity and embedded flash memories is growing rapidly as we enter the system-on-chip era. Conventional tests for flash memories are usually ad hoc-the test procedure is developed for a specific design. As there is a large number of possible failure modes for flash memories, long test algorithms on complicated automatic test equipment (ATE) are commonly seen. The long test time results in high test cost. We propose a systematic approach in testing flash memories, including the development of March-like test algorithms, cost-effective fault diagnosis methodology, and built-in self-test (BIST) scheme. The improved March-like test algorithms can detect disturb faults-derived from the IEEE STD 1005-and conventional faults. As the memory array architecture and/or cell structure varies, the targeted fault set may change. We have developed a flash-memory fault simulator called RAMSES-FT, with which we can easily analyze and verify the coverage of targeted faults under any given test algorithm. In addition, the RAM test algorithm generator-test algorithm generator by simulation-has been enhanced based on RAMSES-FT, so that one can easily generate tests for flash memories, whether they are bit- or word-oriented. The proposed fault diagnosis methodology helps improve the production yield. We also develop a built-in self-diagnosis (BISD) scheme-a BIST design with diagnosis support. The BISD circuit collects useful test information for off-chip diagnostic analysis. It has unique test mode control that reduces test time and diagnostic data shift-out cycles by a parallel shift-out mechanism Jen-Chieh Yeh, Kuo-Liang Cheng, Yung-Fa Chou, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | STEAC: A Platform for Automatic SOC Test IntegrationabstractThe lack of electronic design automation tools for system-on-chip (SOC) test integration increases SOC development time and cost, so SOC test integration tools are important in the success of promoting SOC. We have stressed practical SOC test integration issues, including real problems found in test scheduling, test input/output (I/O) reduction, timing of functional test, scan I/O sharing, etc. In this paper, we further consider the requirement of integrating at-speed testing of embedded cores - to detect timing-related defects, our test architecture is equipped with at-speed test capability. Test scheduling is done based on our test architecture and test access mechanism, considering I/O resource constraints. Detailed scheduling further reduces the overall test time of the system chip. All these techniques are integrated into an automatic flow to facilitate SOC test integration. The test integration platform has been applied to both academic and industrial SOC cases. The chips have been designed and fabricated. The measurement results justify the approach - simple and efficient, i.e., short test integration cost, short test time, and small hardware and pin overhead. Chih-Yen Lo, Chen-Hsing Wang, Kuo-Liang Cheng, Jing-Reng Huang, Chih-Wea Wang, Shin-Moe Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2006 | An Enhanced SRAM BISR Design with Reduced Timing PenaltyabstractRedundancy repair is an effective yield-enhancement technique for memories. There are many previously proposed repair methodologies, such as the popular repair methodology based on the concept of address remapping mechanism achieved by address comparison and address reconfiguration. However, a BISR design with a typical address remapping mechanism usually involves significant timing penalty. Therefore, we propose a new address remapping scheme with a write buffer to reduce the timing penalty. Our experiments show that with the proposed address remapping scheme and redundancy architecture, the timing penalty of our BISR scheme is the same with that of the built-in self-test (BIST) circuit-only one multiplexer delay for both the inputs and outputs Li-Ming Denq, Tzu-chiang Wang, Cheng-Wen Wu |
ATS | 3 |
| 2006 | A network security processor design based on an integrated SOC design and test platformabstractIn this paper we present a generic network security processor (NSP) design suitable for a wide range of security related protocols in wired or wireless network applications. Following the platform-based design methodology, we develop four specific platforms, i.e., architecture platform, EDA platform, design-for-testability (DFT) platform, and prototyping platform, for our NSP design. With these platforms, design of the NSP chip becomes more efficient and systematic. A prototype chip of the NSP has been implemented and fabricated with a 0.18 /spl mu/m CMOS technology. The chip area is 5 mm /spl times/ 5 mm (with 1M gates approximately), including I/O pads. The operating clock rate is 80 MHz. The best performance of the crypto-engines is 1.025 Gbps for AES, 1.652 Mbps for RSA, 125.9/157.65 Mbps for HMAC-SHA1/MD5, and 2.56 Gbps for random number generator. Comparison result shows that our NSP is efficient in terms of performance, flexibility and scalability. Chen-Hsing Wang, Chih-Yen Lo, Min-Sheng Lee, Jen-Chieh Yeh, Chih-Tsun Huang, Cheng-Wen Wu, Shi-Yu Huang |
DAC | 6 |
| 2006 | An Enhanced EDAC Methodology for Low Power PSRAMabstractAs feature size keeps shrinking, how to maintain the reliability becomes an important issue in IC production, especially for high density memory circuits. Error detection and correction (EDAC) schemes have been widely used for memory circuits for this purpose, but ordinary EDAC schemes are not suitable for memories with long codewords. The demand for low-power memory is increasing due to the growth in portable electronics markets. Power reduction in memories with DRAM-like cells can be done by reducing the refresh frequency, but the loss of data integrity should be taken care of seriously. To solve the above two issues, we propose a parallel encoding and decoding EDAC scheme, which can be used on memories with long codewords. Targeting refresh power reduction, we have implemented our scheme on an industrial pseudo SRAM (PSRAM), and have completed experiments. The major hardware penalty is the parity overhead that is 1/9, and the longest delay of our circuit is 3.6ns for the PSRAM fabricated by a 0.11mum CMOS technology. With respect to the 70ns access time of the PSRAM, the proposed EDAC scheme can be integrated with the read/write operations without increasing the latency. Experimental results show that the refresh time can be extended greatly, without sacrificing reliability Po-Yuan Chen, Chao-Hsun Chen, Jen-Chieh Yeh, Cheng-Wen Wu, Jeng-Shen Lee, Yu-Chang Lin |
ITC | 5 |
| 2006 | Testing MRAM for Write Disturbance FaultabstractThe magnetic random access memory (MRAM) is considered one of the potential candidates that will replace current on-chip memories (RAM, EEPROM, and flash memory) in the future. The MRAM is fast and does not need a high supply voltage for read/write operations. It can also endure almost unlimited read/write cycles. These combined advantages of RAM and flash memory make it a potential choice for SOC. In this paper, we present the write disturbance fault (WDF) model for MRAM, i.e., a fault that affects the data stored in the MRAM cells due to excessive magnetic field during the write operation. The proposed WDF model is justified by chip measurement results. We also construct the SPICE macro model for the magnetic tunneling junction (MTJ) device of the toggle MRAM to obtain circuit simulation results. An MRAM chip has been designed and fabricated using a CMOS-based 0.18mum technology. We also present an MRAM fault simulator called RAMSES-M, based on which we derive the shortest test for the proposed WDF model. The test is shown to be better and more robust as compared with March C. Finally, we present a March 17N diagnosis algorithm for identifying the WDF Chin-Lung Su, Chih-Wea Tsai, Cheng-Wen Wu, Chien-Chung Hung, Young-Shying Chen, Ming-Jer Kao |
ITC | 3 |
| 2006 | A Built-In Self-Repair Scheme for NOR-Type Flash MemoryabstractThe strong demand of non-volatile memory for SOC and SIP applications has made flash memory increasingly important. However, deep submicron defects and process uncertainties are causing yield loss of memory products. To solve the yield issue, built-in self-repair (BISR) is widely believed to be cost effective. It is, however, non-trivial to implement BISR on flash memories. In this paper we propose a BISR scheme for NOR-type flash memory. The BISR scheme performs built-in self-test (BIST), built-in redundancy analysis (BIRA), as well as on-chip repair. A typical redundancy architecture for NOR-type flash memory is assumed, based on which we present a redundancy analysis (RA) algorithm. Experimental result shows that the proposed BISR scheme can effectively repair most defective memories. Yu-Ying Hsiao, Chao-Hsun Chen, Cheng-Wen Wu |
VTS | 3 |
| 2006 | Session AbstractabstractIn nanometer IC manufacturing, testing plays an even more important role than ever. The concern is no longer just defect screening and test coverage, but also reliability and yield. We have three speakers representing three major IC foundries, respectively. They will talk about the challenges they have seen from past experiences, and give directions of the trend regarding test practices, including failure/defect analysis and diagnosis, yield enhancement, DFT, etc. Cheng-Wen Wu |
VTS | 1 |
| 2006 | Efficient built-in redundancy analysis for embedded memories with 2-D redundancyabstractA novel redundant mechanism is proposed for embedded memories in this paper. Redundant rows and columns are added into the memory array as in the conventional approaches. However, the redundant rows and columns are divided into row blocks and column blocks, respectively. The reconfiguration is performed at the row (column) block level instead of the conventional row (column) level. Based on the proposed redundant mechanism, we first show that the complexity of the redundancy allocation problem is NP-complete. Thereafter, an extended local repair-most (ELRM) algorithm suitable for built-in implementation is proposed. The complexity of the ELRM algorithm is O(N), where N denotes the number of memory cells. According to the simulation results, the hardware overhead for implementing this algorithm is below 0.17% for a 1024/spl times/2048-b SRAM. Due to the efficient usage of the redundant elements, the manufacturing yield, repair rate, and reliability can be improved significantly. Shyue-Kung Lu, Yu-Chen Tsai, Chih-Hsien Hsu, Kuo-Hua Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2005 | A configurable AES processor for enhanced securityabstractWe propose a configurable AES processor for extended-security communication. The proposed architecture can provide up to 219 different AES block cipher schemes within a reasonable hardware cost. Data can be encrypted not only with secret keys and initial vectors, but also by different block ciphers during the communication. A novel on-the-fly key expansion design is also proposed for 128-, 192-, and 256-bit keys. Our unified hardware can run both the original AES algorithm and the extended AES algorithm. The proposed processor design has been fabricated by a 0.25μm CMOS process, with a silicon area of 6.93mm2---about 200.5K equivalent gates. Under a 66MHz clock, the throughput rate for both the ECB and CBC operation modes are 844.8Mbps, 704Mbps, and 603.4Mbps for 128-bit, 192-bit, and 256-bit keys, respectively. Chih-Pin Su, Chia-Lung Horng, Chih-Tsun Huang, Cheng-Wen Wu |
ASP-DAC | 4 |
| 2005 | Design and test of a scalable security processorabstractThis paper presents a security processor to accelerate cryptographic processing in modern security applications. Our security processor is capable of popular cryptographic functions such as RSA, AES, hashing and random number generation, etc. With proposed Crypto-DMA controller, data gathering and scattering become flexible for security processing, using a simple descriptor-based programming model. The architecture of the security processor with its core-based platform is scalable and configurable for security variations in performance, cost and power consumption. Different number of data channels and crypto-engines can be used to meet the specifications. In addition, a DFT platform is also implemented for the design-test integration. The security processor has been fabricated with 0.18μm CMOS technology. The core area is 3.899mm x 2.296mm (525K gates approximately) and the operating clock rate is 83MHz. Chih-Pin Su, Chen-Hsing Wang, Kuo-Liang Cheng, Chih-Tsun Huang, Cheng-Wen Wu |
ASP-DAC | 5 |
| 2005 | Flash Memory Die Sort by a Sample Classification MethodabstractAs the memory cells keep scaling down and designs are getting bigger and faster, uncertainty is becoming one of the greatest challenges for the semiconductor industry. Unexpected and unpredictable behaviors of devices usually lead to poor quality and reliability. Low-cost test techniques that improves die sorting accuracy thus are critical for advanced devices. Flash memory is more prone to such problem compared with others. Large capacity, high density, and complicated cell structure makes flash memory cell behavior difficult to predict precisely. Even when we test the dies on the same wafer it can be bothering, as each of them may ask for different test condition due to geometric process variation. As a fast and easy-to-use method to solve the problem, we propose a sample classification method. It is not only effective for flash memory testing, but also for other types of circuits that face similar test problem. Experimental result shows that the method solves the flash memory die sort problem efficiently and accurately. The test time is greatly reduced-for an industrial chip, the test time is reduced from 8,817 ms to 718 ms. Moreover, the proposed approach is also suitable for design-for-testability (DFT) implementation that can easily be integrated with a commodity or embedded memory. Yu-Chun Dawn, Jen-Chieh Yeh, Cheng-Wen Wu, Chia-Ching Wang, Yung-Chen Lin, Chao-Hsun Chen |
Asian Test Symposium | 3 |
| 2005 | SOC Testing Methodology and PracticeabstractOn a commercial digital still camera (DSC) controller chip, we practice a novel SOC test integration platform, solving real problems in test scheduling, test IO reduction, timing of functional test, scan IO sharing, embedded memory built-in self-test (BIST), etc. The chip has been fabricated and tested successfully by our approach. Test results prove that short test integration cost, short test time, and small area overhead can be achieved. To support SOC testing, a memory BIST compiler and an SOC testing integration system have been developed. Cheng-Wen Wu |
DATE | 1 |
| 2005 | A BIST Scheme for FPGA Interconnect Delay FaultsabstractIn this paper, we propose a new BIST-based approach for testing FPGA interconnect delay faults. The BIST architecture utilizes the regularity of an FPGA by implementing small test circuits repetitively over FPGA's CLB arrays. Each test circuit targets a specific path and determine conformance of the path delay according to a test clock. With the target path configured as a loop back in the test circuit, test accuracy of the path delay can be increased with reduced effects from skews of the test clocks. Thus, this BIST has a higher delay fault coverage, since it is not necessary to apply guard bands for skews in test mode. Jing-Jia Liou, Yen-Lin Peng, Chih-Tsun Huang, Cheng-Wen Wu |
VTS | 5 |
| 2005 | Flash Memory Built-In Self-Diagnosis with Test Mode ControlabstractThe objective of this paper is to present a cost-effective fault diagnosis methodology for flash memory. Flash memory is enjoying a rapid market growth. The research for flash memory testing is mainly to reduce the test cost and improve the production yield. In this paper, we propose a fault diagnosis flow for flash memory. We also propose a flexible built-in self-diagnosis (BISD) design with enhanced test mode control, which reduces the test time and diagnostic data shift-out cycles by using parallel programming and erasure and employing a parallel shift-out mechanism. The area overhead of our BISD circuit is only about 0.5% for a 256Mb commodity flash memory chip. Experimental results from industrial chips show that the proposed diagnosis methodology has high accuracy in distinguishing the fault type. Jen-Chieh Yeh, Yan-Ting Lai, Yuan-Yuan Shih, Cheng-Wen Wu, Chien-Hung Ho, Yen-Tai Lin |
VTS | 4 |
| 2005 | A built-in self-repair design for RAMs with 2-D redundancyabstractThis brief presents a built-in self-repair (BISR) scheme for semiconductor memories with two-dimensional (2-D) redundancy structures, i.e., spare rows and spare columns. The BISR design is composed of a built-in self-test module and a built-in redundancy analysis (BIRA) module. The BIRA module executes the proposed RA algorithm for RAM with a 2-D redundancy structure. The BIRA module also serves as the reconfiguration unit in the normal mode. Experimental results show that a high repair rate (i.e., the ratio of the number of repaired memories to the number of defective memories) is achieved with the BISR scheme. The BISR circuit has a low area overhead-about 4.6% for an 8 K /spl times/ 64 SRAM. Jin-Fu Li 0001, Jen-Chieh Yeh, Rei-Fu Huang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2004 | SRAM delay fault modeling and test algorithm development
Rei-Fu Huang, Yan-Ting Lai, Yung-Fa Chou, Cheng-Wen Wu |
ASP-DAC | 4 |
| 2004 | An HMAC processor with integrated SHA-1 and MD5 algorithms
Mao-Yin Wang, Chih-Pin Su, Chih-Tsun Huang, Cheng-Wen Wu |
ASP-DAC | 4 |
| 2004 | A Signa-Delta Modulation Based Analog BIST System with a Wide Bandwidth Fifth-Order Analog Response Extractor for Diagnosis PurposeabstractA wide bandwidth /spl Sigma/-/spl Delta/ modulation based analog built-in self-test (BIST) system that can diagnose the prototype is presented. It consists of a low-cost design-for-testability (DJT) switched-capacitor filter as the circuit under test (CUT) and a wide bandwidth analog response extractor (ARE) to digitize the analog responses for final DSP analysis. The first stage of the DJT CUT is reconfigured to accept a repetitive /spl Sigma/-/spl Delta/ modulated bit-steam as its stimulus. This DJT technique reuses every original component and thus provides the advantages of lowering the testing cost, increasing the fault coverage as well as the accuracy, and being able to perform the at-speed tests. The ARE is a cascaded 2-1-1-I fifth-order /spl Sigma/-/spl Delta/ modulator equipped with single-bit quantizers to extend the testing bandwidth while retaining moderate tolerance of circuit imperfections. Our measurement results show that this ARE is able to provide a -95 dB spurious free dynamic range over 1 MHz bandwidth when operates at 30 MHz- A multi-tone test is performed to manifest the wide bandwidth and high accuracy of our BIST system. Based on the BIST results, a method of diagnosing the prototype to speed up the time-to-market is also proposed and demonstrated by our BIST system. Hao-Chiao Hong, Cheng-Wen Wu, Kwang-Ting Cheng |
Asian Test Symposium | 2 |
| 2004 | Fail Pattern Identification for Memory Built-In Self-RepairabstractWith the advent of deep submicron technology and system-on-chip (SOC) design methodology, we are seeing on-chip memory cores to represent a growing percentage of the chip area. The yield of an SOC is usually dominated by the memory yield, so the improvement of memory yield is crucial in SOC development. In this work, we propose a built-in self-repair (BISR) scheme for memory yield improving. The novelty of our approach is that we can identify the fail patterns so that appropriate spare elements (e.g., spare rows, columns, words, or blocks) can be allocated to repair the defective memory. Some BISR methods are discussed and compared. We select the scheme that uses fewer spare elements than others given the same repair rate. The area overhead of the BISR scheme is only 2.2% for an 8K/spl times/64 memory. Rei-Fu Huang, Chin-Lung Su, Cheng-Wen Wu, Shen-Tien Lin, Kun-Lun Luo, Yeong-Jar Chang |
Asian Test Symposium | 3 |
| 2004 | On Test and Diagnostics of Flash MemoriesabstractEmbedded flash memory has been widely used in applications that require non-volatile on-chip storage elements. However, test and diagnostics of flash memories needs further investigation so that the overall cost of the products can be reduced. This paper presents the challenges and issues for test and diagnostics of flash memories, based on our recent experiences. We also suggest improvement of the test and diagnosis flow, including design-for-testability (DFT) using built-in self-test (BIST), built-in self-repair (BISR), and failure analysis. In addition, we present a configurable flash memory tester using FPGA for low-cost testing and diagnostics. Experimental results on industrial flash chips justify the effectiveness of our test and diagnostics system. Chih-Tsun Huang, Jen-Chieh Yeh, Yuan-Yuan Shih, Rei-Fu Huang, Cheng-Wen Wu |
Asian Test Symposium | 5 |
| 2004 | MRAM Defect Analysis and Fault ModeliabstractWith the advent of system-on-chip (SOC), the demand for embedded memory cores increases rapidly. The magnetic random access memory (MRAM) is considered one of the potential candidates that replace current on-chip memories (RAM, EEPROM, and flash memory) in the future. The MRAM has a high speed and does not need high supply voltage for read/write operations, so it has the advantages of RAM and flash memory, making it a potentially good choice for SOC. The testing of MRAM, however, has not been fully investigated. In this work we classify and analyze the MRAM defects and their behavior, and propose its fault models. We have built a SPICE model of MRAM cell and performed defect injection and simulation of a real MRAM circuit. The circuit has been implemented and fabricated with a novel 0.18 m technology. The simulation results regarding the correlation between the defects and conventional fault models show that most of the defects can be covered by the stuck-at fault model. The test data based on the fabricated chips show that the stuck-at faults do cover most of the defects on the chips. However, from the experiment we also have identified two new faults, i.e., the Multi-Victims fault and Kink fault. Chin-Lung Su, Rei-Fu Huang, Cheng-Wen Wu, Chien-Chung Hung, Ming-Jer Kao, Yeong-Jar Chang, Wen Ching Wu |
ITC | 3 |
| 2004 | A Graph-Based Approach to Power-Constrained SOC Test Scheduling
Chih-Pin Su, Cheng-Wen Wu |
J. Electron. Test. | 2 |
| 2003 | A highly efficient AES cipher chipabstractWe present an efficient hardware implementation of the AES (Advanced Encryption Standard) algorithm, with key expansion capability. Instead of the widely used table-lookup implementation of S-box, the proposed basis transformation technique reduces the hardware overhead of the S-box by 64% and is easily pipelined to achieve high throughput rate. Using a typical 0.25μm CMOS technology, the throughput rate is 2.977 Gbps for 128-bit keys, 2.510 Gbps for 192-bit keys, and 2.169 Gbps for 256-bit keys with a 250MHz clock. Testability of the design is also considered. The area of the core circuit is about 1,279 x 1,271μm2. Chih-Pin Su, Tsung-Fu Lin, Chih-Tsun Huang, Cheng-Wen Wu |
ASP-DAC | 4 |
| 2003 | Design of a scalable RSA and ECC crypto-processorabstractAbstract—In this paper, we propose a scalable word-based crypto-processor that performs modular multiplication based on modified Montgomery algorithm for finite fields GF P and GF 2m. The unified crypto-processor supports scalable keys of length up to 2048 bits for RSA and 512 bits for elliptic curve cryptography (ECC). Further extension of the key length can be done easily by enlarging the memory module or using the external memory resource. With the proposed parity prediction technique, our pipelined crypto-processor achieves a 512-bit RSA encryption rate of 276 Kbps and a 160-bit ECC encryption rate of 73.3 Kbps for a 220MHz clock rate. I. Ming-Cheng Sun, Chih-Pin Su, Chih-Tsun Huang, Cheng-Wen Wu |
ASP-DAC | 4 |
| 2003 | Defect Oriented Fault Analysis for SRAMabstractFault analysis is an important step in establishing detailed fault models or subsequent diagnostics and debugging of a semiconductor memory product. We have performed defect injection in the memory cell array of an industrial SRAM circuit and analyzed the faulty behavior with respect to each defect injected. We found that although some of the defects can be mapped to existing fault models, there are many defects that result in unmodeled faults. Moreover, a defect may exhibit a different faulty behavior at a different location in the cell array. The voltage and temperature parameters can also change the faulty behavior. The simulation results show that almost all open and short defects lead to stuck-at faults, transition faults, and data retention faults. Rei-Fu Huang, Yung-Fa Chou, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 2003 | A Processor-Based Built-In Self-Repair Design for Embedded MemoriesabstractWe propose an embedded processor-based built-in self-repair (BISR) design for embedded memories. In the proposed design we reuse the embedded processor that can be found on almost every system-on-chip (SOC) product, in addition to many distinct features. By reusing the embedded processor, the controller and redundancy analysis circuit of a typical BISR design can be removed. Also, the test algorithm and redundancy analysis/allocation algorithm are easily programmable, greatly increasing the design flexibility. We also have developed a memory wrapper that allows at-speed testing of the memory cores. The area overhead of the proposed BISR scheme is low, since only the memory wrapper needs to be realized explicitly. Our experiments show that the BISR area overhead for a typical 8 K/spl times/32 SRAM is lower than 1%. Chin-Lung Su, Rei-Fu Huang, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 2003 | FAME: A Fault-Pattern Based Memory Failure Analysis Framework
Kuo-Liang Cheng, Chih-Wea Wang, Jih-Nung Lee, Yung-Fa Chou, Chih-Tsun Huang, Cheng-Wen Wu |
ICCAD | 6 |
| 2003 | A Built-In Self-Repair Scheme for Semiconductor Memories with 2-D RedundancyabstractEmbedded memories are among the most widely used cores in current system-on-chip (SOC) implementations. Memory cores usually occupy a significant portion of the chip area, and dominate the manufacturing yield of the chip. Efficient yield-enhancement techniques for embedded memories thus are important for SOC. In this paper we present a built-in self-repair (BISR) scheme for semiconductor memories with 2-D redundancy structures. The BISR design is composed of a built-in self-test (BIST) module and a built-in redundancy analysis (BIRA) module. Our BIST circuit supports three test modes: the 1) main memory testing, 2) spare memory testing, and 3) repair modes. The BIRA module executes the proposed redundancy analysis (RA) algorithm for RAM with a 2-D redundancy structure, i.e., spare rows and spare columns. The BIRA module also serves as the reconfiguration (address remapping) unit in the normal mode. Experimental results show that a high repair rate (i.e., the ratio of the number of repaired memories to the number of defective memories) is achieved with the proposed RA algorithm and BISR scheme. The BISR circuit has a low area overhead—about 4.6 % for an 8K¢64 SRAM. Jin-Fu Li 0001, Jen-Chieh Yeh, Rei-Fu Huang, Cheng-Wen Wu, Peir-Yuan Tsai, Archer Hsu, Eugene Chow |
ITC | 4 |
| 2003 | Fault Pattern Oriented Defect Diagnosis for MemoriesabstractFailure analysis (FA) and diagnosis of memory cores plays a key role in system-on-chip (SOC) product development and yield ramp-up. Conventional FA based on bitmaps and the experiences of the FA engineer is time consuming and error prone. The increasing time-to-volume pressure on semiconductor products calls for new development flow that enables the product to reach a profitable yield level as soon as possible. Demand in methodologies that allow FA automation thus increases rapidly in recent years. This paper proposes a systematic diagnosis approach based on failure patterns and functional fault models of semiconductor memories. By circuit-level simulation and analysis, we have also developed a fault pattern generator. Defect diagnosis and FA can be performed automatically by using the fault patterns, reducing the time in yield improvement. The main contribution of the paper is thus a methodology and procedure for accelerating FA and yield optimization for semiconductor memories. Chih-Wea Wang, Kuo-Liang Cheng, Jih-Nung Lee, Yung-Fa Chou, Chih-Tsun Huang, Cheng-Wen Wu, Frank Huang, Hong-Tzer Yang |
ITC | 6 |
| 2003 | Test and Diagnosis of Word-Oriented Multiport MemoriesabstractConventionally, the test of multiport memories is considered difficult because of the complex behavior of the faulty memories and the large number of inter-port faults. This paper presents an efficient approach for testing and diagnosing multiport RAMs. Our approach takes advantage of the higher access bandwidth due to the increased number of read/write ports, which also provides higher observability and controllability that effectively reduces the test time. Our key idea is that a sequence of March operations for any memory cell can be folded and executed within a single access cycle. We have also developed an efficient test algorithm for port-specific faults as well as traditional cell faults. The port-specific faults include the stuck-open, address decoder, and inter-port faults, for both bit-oriented and word-oriented RAMs. Experimental results for our folding scheme show that the test time reduction is about 28% for a commercial 8 KB embedded SRAM. An efficient diagnostic algorithm is also proposed for the port-specific faults and traditional cell faults. Chih-Wea Wang, Kuo-Liang Cheng, Chih-Tsun Huang, Cheng-Wen Wu |
VTS | 4 |
| 2003 | Testing and Diagnosis Methodologies for Embedded Content Addressable Memories
Jin-Fu Li 0001, Ruey-Shing Tzeng, Cheng-Wen Wu |
J. Electron. Test. | 3 |
| 2003 | Built-in redundancy analysis for memory yield improvementabstractWith the advance of VLSI technology, the capacity and density of memories is rapidly growing. The yield improvement and testing issues have become the most critical challenges for memory manufacturing. Conventionally, redundancies are applied so that the faulty cells can be repairable. Redundancy analysis using external memory testers is becoming inefficient as the chip density continues to grow, especially for the system chip with large embedded memories. This paper presents three redundancy analysis algorithms which can be implemented on-chip. Among them, two are based on the local-bitmap idea: the local repair-most approach is efficient for a general spare architecture, and the local optimization approach has the best repair rate. The essential spare pivoting technique is proposed to reduce the control complexity. Furthermore, a simulator has been developed for evaluating the repair efficiency of different algorithms. It is also used for determining certain important parameters in redundancy design. The redundancy analysis circuit can easily be integrated with the built-in self-test circuit. Chih-Tsun Huang, Chi-Feng Wu, Jin-Fu Li 0001, Cheng-Wen Wu |
IEEE Trans. Reliab. | 4 |
| 2003 | Cellular-array modular multiplier for fast RSA public-key cryptosystem based on modified Booth's algorithmabstractWe propose a radix-4 modular multiplication algorithm based on Montgomery's algorithm, and a fast radix-4 modular exponentiation algorithm for Rivest, Shamir, and Adleman (RSA) public-key cryptosystem. By modifying Booth's algorithm, a radix-4 cellular-array modular multiplier has been designed and simulated. The radix-4 modular multiplier can be used to implement the RSA cryptosystem. Due to reduced number of iterations and pipelining, our modular multiplier is four times faster than a direct radix-2 implementation of Montgomery's algorithm. The time to calculate a modular exponentiation is about n/sup 2/ clock cycles, where n is the word length, and the clock cycle is roughly the delay time of a full adder. The utilization of the array multiplier is 100% when we interleave consecutive exponentiations. Locality, regularity, and modularity make the proposed architecture suitable for very large scale integration implementation. High-radix modular-array multipliers are also discussed, at both the bit level and digit level. Our analysis shows that, in terms of area-time product, the radix-4 modular multiplier is the best choice. Jin-Hua Hong, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | On-chip Analog Response Extraction with 1-Bit ? - ModulatorsabstractBecause of their relative robustness to process variation, /spl Sigma/-/spl Delta/ modulation techniques are particularly suitable for VLSI implementations. In this paper, we propose to employ the 1-bit /spl Sigma/-/spl Delta/ modulation ADC (analog-to-digital converter) as the on-chip analog response extractor for analog/mixed-signal BIST (built-in self-test) applications. To validate the idea, a prototype chip with the proposed BIST circuitry has been designed and fabricated. Performance of the BIST circuitry is validated (up to 87 dB dynamic range), and measurement results of the circuit under test (CUT), a 2nd-order low-pass filter, are presented. Hao-Chiao Hong, Jiun-Lang Huang, Kwang-Ting Cheng, Cheng-Wen Wu |
Asian Test Symposium | 4 |
| 2002 | Test Scheduling and Test Access Architecture Optimization for System-on-ChipabstractWe propose an efficient test scheduling and test access architecture for system-on-chip. The test time and test control complexity are optimized under the test power and test access mechanism (TAM) resource constraints. Using our heuristic algorithms, the test scheduling can be done rapidly with small test time penalty when compared with previous works. Under an existing SoC test framework, the test access hardware can be generated from the scheduling result. Experimental results show that the proposed scheduling is hardware efficient. The system integrator can evaluate the test access architecture and perform rest scheduling systematically. Huan-Shan Hsu, Jing-Reng Huang, Kuo-Liang Cheng, Chih-Wea Wang, Chih-Tsun Huang, Cheng-Wen Wu, Youn-Long Lin |
Asian Test Symposium | 6 |
| 2002 | Test Scheduling of BISTed Memory Cores for SOCabstractThe test scheduling of memory cores can significantly affect the test time and power of system chips. We propose a test scheduling algorithm for BISTed memory cores to minimize the overall testing time under the test power constraint. The proposed algorithm combines several approaches for a near-optimal result, based on the properties of BISTed memory cores. By proper partitioning, an analytic exhaustive search finds optimal results for large memory cores, while a heuristic ordering with simulated annealing further handles a large amount of smaller memory cores. On the average, the results are within 1% difference of the optimal solution for the cases of 200 memory cores. Chih-Wea Wang, Jing-Reng Huang, Yen-Fu Lin, Kuo-Liang Cheng, Chih-Tsun Huang, Cheng-Wen Wu, Youn-Long Lin |
Asian Test Symposium | 6 |
| 2002 | A Hierarchical Test Scheme for System-On-Chip DesignsabstractSystem-on-chip (SOC) design methodology is becoming the trend in the IC industry. Integrating reusable cores from multiple sources is essential in SOC design, and different design-for-testability methodologies are usually required for testing different cores. Another issue is test integration. The purpose of this paper is to present a hierarchical test scheme for SOC with heterogeneous core test anti lest access methods. A hierarchical test manager (HTM) is proposed to generate the control signals for these cores, taking into account the IEEE P1500 Standard proposal. A standard memory BIST interface is also presented, linking the HTM and the memory BIST circuit. It can control the BIST circuit with the serial or parallel test access mechanism. The hierarchical test control scheme has low area anti pin overhead, and high flexibility. An industrial case using this scheme has been designed, showing an area overhead of only about 0.63%. Jin-Fu Li 0001, Hsin-Jung Huang, Jeng-Bin Chen, Chih-Pin Su, Cheng-Wen Wu, Chuang Cheng, Shao-I Chen, Chi-Yi Hwang, Hsiao-Ping Lin |
DATE | 5 |
| 2002 | Diagonal Test and Diagnostic Schemes for Flash MemorieabstractEmbedded flash memory plays an increasingly important role for system-on-chip (SOC), especially for battery-powered devices. Testing and diagnosis of embedded flash memory is becoming one of the key development and production issues for many SOC products. Moreover, high density, high capacity, and the integration of heterogeneous cores in an SOC results in long test time, which in turn lead to high test cost. In this paper we propose a new diagonal test algorithm for flash memory that effectively reduces the test time without sacrificing the fault coverage. Both disturb faults and conventional RAM faults are covered. A diagnostic algorithm is also presented, which can distinguish among all the disturb faults and most of the conventional RAM faults. Finally, a built-in self-diagnosis (BISD) scheme is proposed. The BISD circuit implements our algorithms and user-defined ones, and its area overhead is low, e.g., it contains only about 2,551 gates (2-3%) for a 2 Mb flash memory. The test time by our diagonal test is reduced by about 42.69% as compared with the best March-like algorithm reported so far. Sau-Kwo Chiu, Jen-Chieh Yeh, Chih-Tsun Huang, Cheng-Wen Wu |
ITC | 4 |
| 2002 | RAMSES-FT: A Fault Simulator for Flash Memory Testing and DiagnosticsabstractIn this paper we present a fault simulator for flash memory testing and diagnostics, called RAMSES-FT. The fault simulator is designed for easy inclusion of new fault models by adding their fault descriptors without modifying the simulation engine. The flash memory fault models are discussed, based on the failures defined in the IEEE 1005 Standard. Both the NOR-type and NAND-type flash memory architectures are covered. Our flash memory fault simulator uses a parallel simulation strategy to reduce the simulation time complexity from O(N/sup 3/) to O(N/sup 2/), where N is the number of cells. With the proposed scaling method for March tests, the simulation time complexity is further reduced to O(W/sup 2/), where W is the word width of the memory. The fault simulator supports March algorithms as well as single memory operations, covering most of the flash memory tests. With RAMSES-FT we have developed a diagnostic algorithm that can distinguish the target flash memory faults. Kuo-Liang Cheng, Jen-Chieh Yeh, Chih-Wea Wang, Chih-Tsun Huang, Cheng-Wen Wu |
VTS | 5 |
| 2002 | Testing and Diagnosing Embedded Content Addressable MemoriesabstractEmbedded content addressable memories (CAMs) are important components in many system chips. In this paper two efficient March-like test algorithms are proposed. In addition to typical RAM faults, they also cover CAM-specific comparison faults. The first algorithm requires 9N Read/Write operations and 2(N+W) Compare operations to cover comparison and RAM faults (but does not fully cover the intra-word coupling faults), for an N/spl times/W-bit CAM. The second algorithm uses 3N log/sub 2/ W Write and 2W log/sub 2/ W Compare operations to cover the remaining intra-word coupling faults. Compared with the previous algorithms, the proposed algorithms have higher fault coverage and lower time complexity. Moreover it can test the CAM even when its comparison result is observed only by the Hit output or the priority encoder output. Fault-location algorithms are also developed for locating the cells with comparison faults. Jin-Fu Li 0001, Ruey-Shing Tzeng, Cheng-Wen Wu |
VTS | 3 |
| 2002 | Diagnostic Data Compression Techniques for Embedded Memories with Built-In Self-Test
Jin-Fu Li 0001, Ruey-Shing Tzeng, Cheng-Wen Wu |
J. Electron. Test. | 3 |
| 2002 | A Built-in Self-Test Scheme with Diagnostics Support for Embedded SRAM
Chih-Wea Wang, Chi-Feng Wu, Jin-Fu Li 0001, Cheng-Wen Wu, Tony Teng, Kevin Chiu, Hsiao-Ping Lin |
J. Electron. Test. | 4 |
| 2002 | Neighborhood pattern-sensitive fault testing and diagnostics for random-access memoriesabstractThe authors present test algorithms for go/no-go and diagnostic test of memories, covering neighborhood pattern-sensitive faults (NPSFs). The proposed test algorithms are March based, which have linear time complexity and result in a simple built-in self-test (BIST) implementation. Although conventional March algorithms do not generate all neighborhood patterns to test the NPSFs, they can be modified by using multiple data backgrounds such that all neighborhood patterns can be generated. The proposed multibackground March algorithms have shorter test lengths than previously reported ones, and the diagnostic test algorithm guarantees 100% diagnostic resolution for NPSFs and conventional RAM faults. Based on the proposed algorithms, the authors also present a cost-effective BIST design. The BIST circuit is programmable, and it supports March algorithms, including the proposed multibackground one. Kuo-Liang Cheng, Ming-Fu Tsai, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2002 | Fault simulation and test algorithm generation for random accessmemoriesabstractThe size and density of semiconductor memories is rapidly growing, making them increasingly harder to test. New fault models and test algorithms have been continuously proposed to cover defects and failures of modern memory chips and cores. However, software tool support for automating the memory test development procedure is still insufficient. For this purpose, we have developed a fault simulator (called RAMSES) and a test algorithm generator (called TAGS) for random-access memories (RAMs). In this paper, we present the algorithms and other details of RAMSES and TAGS and the experimental results of these tools on various memory architectures and configurations. We show that efficient test algorithms can be generated automatically for bit-oriented memories, word-oriented memories, and multiport memories, with 100% coverage of the given typical RAM faults. Chi-Feng Wu, Chih-Tsun Huang, Kuo-Liang Cheng, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2002 | Efficient FFT network testing and diagnosis schemesabstractWe consider offline testing, design-for-testability, and diagnosis for fast Fourier transform (FFT) networks. A practical FFT chip can contain millions of gates, so effective testing and fault-tolerance techniques usually are required in order to guarantee high-quality products. We propose M-testability conditions for FFT butterfly, omega, and flip networks at the double-multiply-subtract-add (DMSA) module level. A novel design-for-testability technique based on the functional bijectivity property of the specified modules to detect faults other than the cell faults is presented. It guarantees 100% combinational fault coverage with negligible hardware overhead-about 0.17% for an FFT network with 16-bit operand words, independent of the network size. Our design requires fewer test vectors compared with previous ones-a factor of up to 1/(6 /spl times/ 2/sup 5n/), where n is the word length. We also propose C-diagnosability conditions and a C-diagnosable FFT network design. By property exchanging and blocking certain fault propagation paths, a faulty DMSA module can be located using a two-phase deterministic algorithm. The blocking mechanism can be implemented with no additional hardware. Compared with previous schemes, our design reduces the diagnosis complexity from O(N) to O(1). For both testing and diagnosis, the hardware overhead for our approach is only about 0.43% for 16-bit numbers regardless of the FFT network size. Jin-Fu Li 0001, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | Processor-programmable memory BIST for bus-connected embedded memoriesabstractWe present a processor-programmable built-in self-test (BIST) scheme suitable for embedded memory testing in the system-on-a-chip (SOC) environment. The proposed BIST circuit can be programmed via an on-chip microprocessor. Upon receiving the commands from the microprocessor, the BIST circuit generates pre-defined test patterns and compares the memory outputs with the expected outputs. Most popular memory test algorithms can be realized by properly programming the BIST circuit using the processor instructions. Compared with processor-based memory BIST schemes that use an assembly-language program to generate test patterns and compare the memory outputs, the test time of the proposed memory BIST scheme is greatly reduced. Ching-Hong Tsai, Cheng-Wen Wu |
ASP-DAC | 2 |
| 2001 | RSA cryptosystem design based on the Chinese remainder theoremabstractIn this paper, we present the design and implementation of a systolic RSA cryptosystem based on a modified Montgomery's algorithm and the Chinese Remainder Theorem (CRT) technique. The CRT technique improves the throughput rate up to 4 times in the best case. The processing unit of the systolic array has 100% utilization because of the proposed block interleaving technique for multiplication and square operations in the modular exponentiation algorithm. For 512-bit inputs, the number of clock cycles needed for a modular exponentiation is about 0.13M to 0.24M. The critical path delay is 6.13ns using a 0.6um CMOS technology. With a 150 MHz clock, we can achieve an encryption/decryption rate of about 328 to 578 Kb/s. Chung-Hsien Wu 0002, Jin-Hua Hong, Cheng-Wen Wu |
ASP-DAC | 3 |
| 2001 | Automatic Generation of Memory Built-in Self-Test Cores for System-on-ChipabstractMemory testing is becoming the dominant factor in testing a system-on-chip (SoC), with the rapid growth of the size and density of embedded memories. To minimize the test effort, we present an automatic generation framework of memory built-in self-test (BIST) cores for SoC designs. The BIST generation framework is a much improved one of our previous work. Test integration of heterogeneous memory architectures and clusters of memories are focused on. The automatic test grouping and scheduling optimize the overhead in test time, performance, power consumption, etc. Furthermore, with our novel BIST architecture, the BIST cores can be accessed via an on-chip bus interface (e.g., AMBA), which eases the control of testing and diagnosis in a typical SoC scenario. With a configurable and extensible architecture, the proposed framework facilitates easy memory test integration for core providers as well as system integrators. Kuo-Liang Cheng, Chia-Ming Hsueh, Jing-Reng Huang, Jen-Chieh Yeh, Chih-Tsun Huang, Cheng-Wen Wu |
Asian Test Symposium | 6 |
| 2001 | A Built-in Self-Test and Self-Diagnosis Scheme for Heterogeneous SRAM ClustersabstractTesting and diagnosis are important issues in system-on-chip (SoC) development, as more and more embedded cores are being integrated into the chips. In this paper we propose a built-in self-test (BIST) and self-diagnosis (BISD) scheme for embedded SRAMs, suitable for SoC applications. It supports manufacturing test as well as diagnosis for design verification and yield improvement. With low hardware cost, our memory BISD approach can handle various types of SRAM, including pipelined, multi-port, and multi-clock architectures. In addition, a test scheduling methodology and a BISD compiler are also implemented, which reduce the testing time as well as test development time. Chih-Wea Wang, Ruey-Shing Tzeng, Chi-Feng Wu, Chih-Tsun Huang, Cheng-Wen Wu, Shi-Yu Huang, Shyh-Horng Lin, Hsin-Po Wang 0002 |
Asian Test Symposium | 5 |
| 2001 | Simulation-Based Test Algorithm Generation and Port Scheduling for Multi-Port MemoriesabstractThe paper presents a simulation-based test algorithm generation and test scheduling methodology for multi-port memories. The purpose is to minimize the testing time while keeping the test algorithm in a simple and regular format for easy test generation, fault diagnosis, and built-in self-test (BIST) circuit implementation. Conventional functional fault models are used to generate tests covering most defects. In addition, multi-port specific defects are covered using structural fault models. Port-scheduling is introduced to take advantage of the inherent parallelism among different ports. Experimental results for commonly used multi-port memories, including dual-port, four-port, and $n$-read-1-write memories, have been obtained, showing that efficient test algorithms can be generated and scheduled to meet different test bandwidth constraints. Moreover, memories with more ports benefit more with respect to testing time. Chi-Feng Wu, Chih-Tsun Huang, Kuo-Liang Cheng, Chih-Wea Wang, Cheng-Wen Wu |
DAC | 5 |
| 2001 | Memory fault diagnosis by syndrome compressionabstractIn this paper we present a data compression technique that can be used to speed up the transmission of diagnosis data from the embedded RAM with built-in self-diagnosis (BISD) support. The proposed approach compresses the faulty-cell address and March syndrome to about 28% of the original size under the March-17N diagnostic test algorithm. The key component of the compressor is a novel syndrome-accumulation circuit, which can be realized by a content-addressable memory. Experimental results show that the area overhead is about 0.9% for a 1Mb SRAM with 164 faults. The proposed compression technique reduces the time for diagnostic test, as well as the tester storage capacity requirement. Jin-Fu Li 0001, Cheng-Wen Wu |
DATE | 2 |
| 2001 | March-based RAM diagnosis algorithms for stuck-at and coupling faultsabstractDiagnosis technique plays a key role during the rapid development of the semiconductor memories, for catching the design and manufacturing failures and improving the overall yield and quality. Investigation on efficient diagnosis algorithms is very important due to the expensive and complex fault/failure analysis process. We propose March-based RAM diagnosis algorithms which not only locate faulty cells but also identify their types. The diagnosis complexity is O(17N) and O((17+10B)N) for bit-oriented and word-oriented diagnosis algorithms, respectively, where N represents the address number and B is the data width. Using the proposed algorithms, stuck at faults, state coupling faults, idempotent coupling faults and inversion coupling faults can be distinguished. Furthermore, the coupled and coupling cells can be located in the memory array. Our word-oriented diagnosis algorithm can distinguish all of the inter-word and intra-word coupling faults, and locate the coupling cells of the intra-word inversion and idempotent coupling faults. With additional 2B-1 operations, the algorithm can further locate the intra-word state coupling faults. With improved diagnostic resolution and test time, the proposed algorithms facilitate the development and manufacturing of semiconductor memories. Jin-Fu Li 0001, Kuo-Liang Cheng, Chih-Tsun Huang, Cheng-Wen Wu |
ITC | 4 |
| 2001 | Efficient Neighborhood Pattern-Sensitive Fault Test Algorithms for Semiconductor MemoriesabstractWe present two memory test algorithms for neighborhood pattern sensitive faults (NPSFs), including static NPSF (SNPSF), passive NPSF (PNPSF) and active NPSF (ANPSF). March algorithms are widely used in memory testing because of their linear time complexity and ease in built-in self-test (BIST) implementation. Although conventional March algorithms do not generate all neighborhood patterns for testing the NPSFs, they can be modified by using multiple data backgrounds such that all neighborhood patterns can be generated. The proposed multi-background March algorithms have shorter test length than previously proposed ones (68N for detecting SNPSFs and PNPSFs and 96N for detecting all NPSFs). Also, based on the proposed algorithms, linear-time BIST circuit can be implemented with low area overhead. Kuo-Liang Cheng, Ming-Fu Tsai, Cheng-Wen Wu |
VTS | 3 |
| 2001 | Unified VLSI systolic array design for LZ data compressionabstractHardware implementation of data compression algorithms is receiving increasing attention due to exponentially expanding network traffic and digital data storage usage. In this paper, we propose several serial one-dimensional and parallel two-dimensional systolic-arrays for Lempel-Ziv data compression. A VLSI chip implementing our optimal linear array is fabricated and tested. The proposed array architecture is scalable. Also, multiple chips (linear arrays) can be connected in parallel to implement the parallel array structure and provide a proportional speedup. Shih-Arn Hwang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Radix-4 modular multiplication and exponentiation algorithms for the RSA public-key cryptosystemabstractWe propose a radix-4 modular multiplication algorithm based on Montgomery's algorithm, and a radix-4 cellulararray modular multiplier based on Booth's multiplication algorithm.The radix-4 modular multiplier can be used to implement fast RSA cryptosystem.Due to reduced number of iterations and pipelining, our modular multiplier is four times faster than the cellular-array modular multiplier based on the original Montgomery's algorithm.The time to calculate a modular exponentiation is about n 2 clock cycles, where n is the word length, and the clock cycle is roughly equal to the delay time of a full adder.The utilization of the multiplier is 100% by interleaving consecutive exponentiations.Locality, regularity, and modularity make the proposed architecture suitable for VLSI implementation. Jin-Hua Hong, Cheng-Wen Wu |
ASP-DAC | 2 |
| 2000 | A programmable built-in self-test core for embedded memoriesabstractNo abstract available. Chih-Tsun Huang, Jing-Reng Huang, Cheng-Wen Wu |
ASP-DAC | 3 |
| 2000 | An FPGA-based re-configurable functional tester for memory chipsabstractThe paper presents a prototype re-configurable tester for memory chips. The new tester consists of a memory test-circuitry compiler, a synthesis/mapping CAD tool, and an FPGA-based re-configurable hardware platform. The compiler makes user-specified parameters of memory under test (such as the address and data bus widths, the march test and the background data) as input and generates the test circuitry required to functionally, test the target memory chips. This framework not only enables the automatic synthesis/mapping of the test circuitry into the re-configurable hardware platform, it also guarantees that the hardware platform can correctly operate at the desired clock rate for the user specified parameters. The proposed solution can reduce the memory tester cost by providing hardware re-configurability to support a wide range of memory chips. We demonstrate that the prototype tester can be automatically configured to test SDRAM chips above 100 MHz. Jing-Reng Huang, Chee-Kian Ong, Kwang-Ting Cheng, Cheng-Wen Wu |
Asian Test Symposium | 4 |
| 2000 | A waveform simulator based on Boolean processabstractHigh operation frequency and strict timing behavior are characteristics of modern high performance integrated circuits, which require a digital system simulator to accurately simulate not only the logic function but also the timing behavior of a circuit. This paper presents some experimental results by SPICE to validate the analytical approach of Boolean process, and extract some data for a numerical waveform simulation. The paper also presents a waveform simulator and its results of experiments, which is very different from traditional logic simulator. Lijian Li 0001, Cheng-Wen Wu, Yinghua Min |
Asian Test Symposium | 3 |
| 2000 | A Testable/Fault Tolerant FFT Processor DesignabstractWith the advent of VLSI technology, a large collection of processing elements can be gathered to achieve high-speed computation economically. However, due to the low pin-count-to-component-count ratio, the controllability and observability of such circuits decrease significantly. As a result, the testing of such highly complex and dense circuits becomes very difficult and expensive. A testable/fault-tolerant FFT processor is proposed in this paper. We first propose a testable design scheme for FFT butterfly networks based on M-testability conditions. According to the M-testability conditions, a novel design-for-testability approach is presented and applied to the module-level systolic FFT arrays. Our M-testability conditions guarantee 100% single-module-fault testability with a minimum number of test patterns. Based on this testable design, a reconfiguration mechanism is incorporated to bypass the faulty cells and the testable/fault-tolerant structures are constructed. Special cell designs are presented to implement the design-for-testability and reconfiguration mechanisms. The reliability of the FFT system increases significantly and the hardware overhead is low-about 16% for the module-level design. Shyue-Kung Lu, Jen-Sheng Shih, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 2000 | A built-in self-test and self-diagnosis scheme for embedded SRAMabstractEmbedded memory test and diagnosis is becoming an important issue in system-on-chip (SOC) development. Direct access of the memory cores from the limited number of I/O pins is usually not feasible. Built-in self-diagnosis (BISD), which include built-in self-test (BIST), is rapidly becoming the most acceptable solution. We propose a BISD design and a fault diagnosis system for embedded SRAM. It supports manufacturing test as well as diagnosis for design verification and yield improvement. The proposed BISD circuit is on-line programmable for its March test algorithms. Test chips have been designed and implemented. Our experimental results show that the BISD hardware overhead is about 2.4% for a typical 128 Kb SRAM and only 0.65% for a 2 Mb SRAM. Chih-Wea Wang, Chi-Feng Wu, Jin-Fu Li 0001, Cheng-Wen Wu, Tony Teng, Kevin Chiu, Hsiao-Ping Lin |
Asian Test Symposium | 4 |
| 2000 | Cost and Benefit Models for Logic and Memory BISTabstractWe present cost and benefit models and analyze the economics effects of built-in self-test (BIST) for logic and memory cores. In our cost and benefit models for BIST, we take into consideration the design verification time and test development time associated with testability. Experimental results for logic BIST and memory BIST examples show that a threshold volume exists when BIST is profitable for the logic core under consideration - it is not recommended for a higher volume. However, BIST is a good choice for memory cores in general. Juin-Ming Lu, Cheng-Wen Wu |
DATE | 2 |
| 2000 | Error Catch and Analysis for Semiconductor Memories Using March TestsabstractWe present an error catch and analysis (ECA) system for semiconductor memories. The system consists of a test algorithm generator called TAGS, a fault simulator called RAMSES, and an error analyzer (ERA). We use TAGS to generate a set of test algorithms of different lengths and diagnostic resolutions for the memory under test, and use RAMSES to generate the March dictionary for each test algorithm. With the March dictionaries, ERA is able to support March algorithms for easy diagnosis of faulty RAMs. Legacy test algorithms also can be reused. When integrated with a RAM tester, our ECA system can generate RAM bitmaps that are similar to the RAM layout. The bitmaps provide detail information about the error locations and faults causing the errors. Based on the information, diagnosis of the RAM chips for yield and reliability improvement can be done more easily. Chi-Feng Wu, Chih-Tsun Huang, Chih-Wea Wang, Kuo-Liang Cheng, Cheng-Wen Wu |
ICCAD | 5 |
| 2000 | Built-in self-test and fault diagnosis for lookup table FPGAsabstractA novel built-in self-test structure for the lookup table (LUT) based field programmable gate arrays (FPGA's) is proposed in this paper. A general structure for the basic configurable logic array blocks (CLB's) is assumed. The whole chip is partitioned into disjoint one-dimensional arrays of cells. We assume that in each linear array, there is at most one faulty cell, and a faulty cell may contain multiple faulty CLB's. Our idea is to configure the cells to make each cell function bijective. In order to detect all faults defined, k+2 configurations are required. The input patterns can be easily generated with a k-bit counter and the fault coverage is 100%. The number of configurations for our BIST structures is 2k+4. Our BIST approaches also have the advantages of requiring less hardware resources for test pattern generation and output response analysis. To locate a faulty CLB, three test sessions are required. However, the maximum number of configurations is k+4 for diagnosing a faulty CLB. Shyue-Kung Lu, Jen-Sheng Shih, Cheng-Wen Wu |
ISCAS | 3 |
| 2000 | Simulation-Based Test Algorithm Generation for Random Access MemoriesabstractAlthough there are well known test algorithms that have been used by the industry for years for testing semiconductor random-access memories (RAMs), systematic evaluation of their effectiveness and efficiency has been a difficult job. In the past, it was mainly done manually by proving a certain algorithm can detect a certain type of fault. As memory technology keeps innovating, the growing complexity of the memories and number of fault types that need to be covered will require more effective and efficient test algorithms to be discovered in much shorter time. A systematic approach for developing and evaluating memory test algorithms is thus desired. We propose such an approach here: test algorithm generation by simulation (TAGS), which generates and optimizes test algorithms, given a test time budget. Experimental results show that the algorithms generated by TAGS are more efficient than the traditional test algorithms. Using TAGS, a series of test algorithms with a detailed list of faults covered by each algorithm can be generated, providing easy trade-off between test time and fault coverage. Chi-Feng Wu, Chih-Tsun Huang, Kuo-Liang Cheng, Cheng-Wen Wu |
VTS | 4 |
| 2000 | A Low-Power CAM Design for LZ Data CompressionabstractLow-power and high-performance data compressors play an increasingly important role in the portable mobile computing and wireless communication markets. Among lossless data compression algorithms for hardware implementation, LZ77 is one of the most widely used. For real-time communication, some hardware LZ compressors/decompressors have been proposed in the past. Content addressable memory (CAM) is widely considered as the most efficient architecture for pattern matching required by the LZ77 compression process. In this paper, we propose a low-power CAM-based LZ77 data compressor. By shutting down the power for unnecessary comparisons between the CAM words and the input symbol, the proposed CAM architecture consumes much lower power than the conventional ones without noticeable performance penalty. Moreover, using the proposed conditional comparison mechanism and the novel CAM cell with the NAND-type matching logic, on average we have close to two orders of improvement on power consumption, i.e., a reduction of more than 98 percent for 8-bit words. Speed is sacrificed if we use the NAND-type matching logic, but the NAND-type logic and the NOR-type logic can be combined to provide the best solution that balances power and delay. Our approach also can be applied to general-purpose CAMs which use the valid bits, so far as the proposed design techniques are adopted. Kun-Jin Lin, Cheng-Wen Wu |
IEEE Trans. Computers | 2 |
| 2000 | A fast signature computation algorithm for LFSR and MISRabstractA multiple-input signature register (MISR) computation algorithm for fast signature simulation is proposed. Based on the table look-up linear compaction algorithm and the modularity property of a single-input signature register (SISR), some new accelerating schemes-partial-input look-up tables and flying state look-up tables-are developed to boost the signature computation speed. Mathematical analysis and simulation results show that this algorithm has an order of magnitude speedup without extra memory requirement compared with the original linear compaction algorithm. Though this algorithm is derived for SISR, a simple conversion scheme exists that can convert internal-EXOR MISR to SISR. Consequently, fast MISR signature computation can be done. Bin-Hong Lin, Shao-Hui Shieh, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Testing content-addressable memories using functional fault modelsand march-like algorithmsabstractFunctional tests for content-addressable memories (CAM's) are presented in this paper. In addition to several traditional functional fault models for RAM's, we also consider the fault models based on physical defects, such as shorts between two circuit nodes and transistor stuck-on and stuck open faults. Accordingly, several functional fault models are proposed. In order to make our approach suited to various application-specific CAM's, we propose tests which require only three fundamental types of operation (i.e., write, erase, and compare), and the test results can be observed entirely from the single-bit Hit output. A complete, compact test is also proposed, which has low complexity and is suitable for modern high-density and large-capacity CAMs-it requires only 2N+3w+2 compare operations and 8N write operations to cover the functional fault models discussed, where N is the number of words and w is the word length. Kun-Jin Lin, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Hierarchical system test by an IEEE 1149.5 MTM-bus slave-module interface coreabstractAn IEEE 1149.5 module test and maintenance (MTM) bus slave module interface core is presented, which is used for direct access from the system bus to the IEEE 1149.1 chip-level or on-chip buses to facilitate hierarchical system test and diagnosis. The hierarchical test methodology also is presented, which is applicable to the system-on-chip environment, All the standard 1149.1 instructions, such as SAMPLE/PRELOAD, EXTEST, BYPASS, and even RUNBIST, can be performed within three 1149.5 read/write-data message cycles. The messages are transmitted between the MTM-bus master module (Ill-module) and the slave module (S-module). We adopt the full test access port control method to activate the 1149.1 boundary-scan paths via the 1149.5 MTM-bus. Our S-module interface circuit implements 16 CORE commands and one read/write-data command. It has been prototyped using a field-programmable gate array chip and implemented by a full-custom chip. Hierarchical test of multiple 1149.1 compatible boards has been experimentally verified. Jin-Hua Hong, Chung-Hung Tsai, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | Testing Interconnects of Dynamic Reconfigurable FPGAsabstractField Programmable Gate Arrays (FPGAs) are an increasingly popular choice for fast prototyping and for products whose time to market is relatively short. Testing FPGAs before programming them is thus becoming a major concern to the manufacturers as well as the users. In this paper we propose a universal test for the interconnects of typical dynamic reconfigurable FPGAs. The proposed test configurations and corresponding test patterns for the Xilinx XC6200 FPGAs are shown to cover all interconnect faults. In our test, the total number of test configurations is only 7, which is independent of the FPGA size. The test time for XC6216 is less than 5 ms. Chi-Feng Wu, Cheng-Wen Wu |
ASP-DAC | 2 |
| 1999 | Defect Level Prediction Using Multi-Model Fault CoverageabstractContinuous progress in VLSI technologies keeps reducing the defect density. However; a yield of 100% is considered unlikely. It is also known that VLSI circuits can never be "completely" tested. Therefore, defect level cannot be reduced to zero. The quality of a test is usually evaluated by the fault coverage. The relationship between the defect level and the fault coverage has never been proposed. In this paper; we first propose the concept of multi-model fault coverage (MFC) instead of the fault coverage based on a unique fault model. The relationship between defect level, fabrication yield, and multi-model fault coverage is then derived. We also analyze the defect level error between the predicted defect level and the physical defect level. As the number of fault models used increases, the defect level error can be reduced significantly. Our approach is very effective for product quality prediction and the concept of multi-model fault coverage is very useful for today's system-on-chip technology. Shyue-Kung Lu, Tsung-Ying Lee, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 1999 | Low-Cost Modular Totally Self-Checking Checker Design for m-out-of-n CodeabstractWe present a low-cost (hardware-efficient) and fast totally self-checking (TSC) checker for m-out-of-n code, where m/spl ges/3, 2m+1/spl les/n/spl les/4m. The checker is composed of four special adders which sum the 1s in the primary inputs added by appropriate constants, two ripple carry adders which sum the outputs of the biased-adders, and a t-variable two-rail code checker tree which compares the outputs of the two ripple carry adders, where k=[log/sub 2/(n-m)+1]. All the modules are composed of 2-input gates and inverters. Compared with previous nonmodular methods, our TSC checker has a lower hardware and time complexity. Our method reduces the hardware complexity and circuit delay of the checker from O(n/sup 2/) to O(n) and from O(n) to O(log/sub 2/n), respectively. Compared with recent modular methods, our TSC checker has about the same hardware and time complexity, but is applicable to a much broader range of n. In summary, our method is superior to existing methods for the considered range of n. In addition, our TSC checker can easily be tested (the test set size of our TSC checker is relatively small) and implemented in VLSI for its modular structure. Wen-Feng Chang, Cheng-Wen Wu |
IEEE Trans. Computers | 2 |
| 1999 | An improved Montgomery's algorithm for high-speed RSA public-key cryptosystemabstractWe revise Montgomery's algorithm such that modular multiplication can be executed two times faster. Each iteration in our algorithm requires only one addition, while that in Montgomery's requires two additions. We then propose a cellular array to implement modular exponentiation for the Rivest-Shamir-Adleman cryptosystem. It has approximately 2n cells, where n is the word length. The cell contains one full-adder and some controlling logic. The time to calculate a modular exponentiation is about 2n/sup 2/ clock cycles. The proposed architecture has a data rate of 100 kb/s for 512-b words and a 100 MHz clock. Chih-Yuang Su, Shih-Am Hwang, Po-Song Chen, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 1998 | On-Line Error Detection Schemes for a Systolic Finite-Field InverterabstractGalois-field (GF) is a number system with a finite set of elements. It is widely used in error-control coding, cryptography, etc. Among its important arithmetic operations, inversion and division have been identified as the most complicated. In this paper, concurrent error detection (CED) schemes have been presented for a systolic GF(2/sup m/) inverter that we have proposed recently. The CED circuitry tests the inverter concurrently while it is in normal operation to increase the reliability of the inverter. There is negligible performance penalty. Analysis shows that all single cell faults can be detected concurrently. The area overhead is less than 5% for 11 bits or longer words. Yu-Chun Chuang, Cheng-Wen Wu |
Asian Test Symposium | 2 |
| 1998 | Testing Embedded Memories: Is BIST the Ultimate Solution?abstractThe trend that embedded memories will play an important role in the semiconductor market over the next few years has been widely noted. The testing of embedded memories therefore is becoming an industry-wide concern. This panel will try to clarify some important issues regarding embedded memory testing, such as the use of BIST for various testing purposes, the use of IddQ for reliability screening, the possibility of built-in redundancy analysis and self-repair, etc. Cheng-Wen Wu |
Asian Test Symposium | 1 |
| 1998 | A Probabilistic Model for Path Delay FaultsabstractTesting path delay faults (PDFs) in VLSI circuits is becoming an important issue as we enter the deep submicron age. However, it is difficult in general, since the number of faults normally is very large and most faults are hard to sensitize. To make delay fault testing and test synthesis easier, we propose a probabilistic PDF model. We investigate probability density functions for wire and path delay size to model the fault effect in the circuit under test. In our approach, delay fault size is assumed to be randomly distributed. An analytical model is proposed to evaluate the PDF coverage. We show that the fault size of the undetected paths can be greatly reduced if these paths are conjoined with other detected paths. Therefore, by our approach, path selection and synthesis of PDF testable circuits can be done more accurately. Also, given a test set, fault coverage can be predicted by calculating the mean delay of the paths. Cheng-Wen Wu, Chih-Yuang Su |
Asian Test Symposium | 1 |
| 1998 | Sequential circuit fault simulation using logic emulationabstractA fast fault simulation approach based on ordinary logic emulation is proposed. The circuit configured into our system that emulates the faulty circuit's behaviour is synthesized from the good circuit and the given fault list in a novel way. Fault injection is made easy by shifting the content of a fault injection scan chain or by selecting the output of a parallel fault injection selector, with which we get rid of the time-consuming bit-stream regeneration process. Experimental results for ISCAS-89 benchmark circuits show that our serial fault emulator is about 20 times faster than HOPE. The speedup grows with the circuit size by our analysis. Two hybrid fault emulation approaches are also proposed. The first reduces the number of faults actually emulated by screening off faults not activated or with short propagation distances before emulation, and by collapsing nonstem faults into their equivalent stem faults. The second reduces the hardware requirement of the fault emulator by incorporating an ordinary fault simulator. Shih-Arn Hwang, Jin-Hua Hong, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1997 | On energy efficiency of VLSI testingabstractWe discuss the role of power and energy in computation and test efficiency. This is done by the proposal of new computation and test efficiency models that take energy into consideration, followed by the incorporation of these models with the CMOS power consumption model to establish the following observations: (1) low power and high testability need not be competing goals in the design optimization process; (2) high power dissipation during testing may not be an issue, as long as the tester limit is not reached and the chip is not over driven; (3) high-power testing due to high speed and/or high transition activity factor is better in terms of test efficiency; and (4) for a fabricated chip with a prespecified fault coverage, testing energy is roughly constant, independent of the testing power or testing time. Cheng-Wen Wu |
Asian Test Symposium | 1 |
| 1997 | Fault-Tolerant Interleaved Memory Systems with Two-Level RedundancyabstractHighly reliable interleaved memory systems for uniprocessor and multiprocessor computer architectures are presented. The memory systems are divided into groups. Each group consists of several banks and each bank has several modules. The error model is defined at the memory-module level. A module is faulty if any single or multiple faults result in loss of the entire module. Spare modules, as well as spare banks, are included in the systems to enhance reliability and availability. A faulty module is replaced by a spare module within a bank first, and, if the bank has no redundancy remaining for the faulty module, the whole bank will be replaced by a spare bank at the next higher level. The structure of the reconfigurable memory system is designed in such a way that the replacement of faulty modules (banks) by spare modules (banks) will not disturb memory references if each bank (group) has at most two spare modules (banks). If there are more than two spare modules (banks) in a bank (group), a second-level address translator is designed which can prohibit references to faulty modules by address remapping. The address translator can be implemented with a CAM or switches. Analysis results show that the system reliability can be significantly improved with little hardware overhead. Also, a typical system with one redundant row of modules has the highest cost-effectiveness during its useful lifetime period. User transparency in memory access is retained. Shyue-Kung Lu, Sy-Yen Kuo, Cheng-Wen Wu |
IEEE Trans. Computers | 3 |
| 1996 | Hierarchical Testing Using the IEEE Std 1149.5 Module Test and Maintenance Slave Interface ModuleabstractAn IEEE Std 1149.5 MTM-Bus Slave interface module is presented, which is used for direct access to 1149.1 chip-level buses and hierarchical test. All the standard 1149.1 functions, such as SAMPLE/PRELOAD, EXTEST, BYPASS, and even RUNBIST, can be performed within three 1149.5 Read/Write-Data message cycles. The messages are transmitted between the MTM-bus Master module (M-module) and the Slave module (S-module). We adopt the Full TAP Control (FTC) method to activate the 1149.1 Boundary-Scan paths via the 1149.5 MTM-Bus. A personal computer is used as the M-module. Jin-Hua Hong, Chung-Hung Tsai, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 1996 | A MISR Computation Algorithm for Fast Signature SimulationabstractA fast multiple input signature register (MISR) computation algorithm for signature simulation is proposed. Based on the linear compaction algorithm, the modularity property of a single input signature register (SISR), and the sparsity of the error-domain input, some new accelerating schemes-partial input look-up tables and reverse zero-checking policy-are developed to boost the signature computation speed. Mathematical analysis and simulation results show that this algorithm has an order of magnitude speedup without extra memory requirement compared with the linear compaction algorithm. Though originally derived for SISR, this algorithm is applicable to MISR by a simple conversion procedure or a bit-adjusting scheme with little effort. Consequently, a very fast MISR signature simulation can be achieved. Bin-Hong Lin, Shao-Hui Shieh, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 1996 | Cell delay fault testing for iterative logic arrays
Shyue-Kung Lu, Cheng-Wen Wu, Ruei-Zong Hwang |
J. Electron. Test. | 2 |
| 1995 | DC control and observation structures for analog circuitsabstractAs the complexity of electronic circuits and systems increases, so does the complexity of testing them. The level-sensitive scan-design (LSSD) structure used in a digital circuit enhances the controllability and observability of the circuit under test. For analog circuits, there also are several approaches proposed to improve their observability, based on the LSSD concept. However, none of these approaches provide control and observation capability for all test points simultaneously. In this paper, we propose two control and observation structures for analog circuits without using extra power supply. Using our approach, one is able to observe and control the DC voltage levels of all test points simultaneously, which is the basic diagnosis capability for the analog circuit under test. A calibration process is presented to ensure the accuracy of the excitation and read-out voltage levels. Yeong-Ruey Shieh, Cheng-Wen Wu |
Asian Test Symposium | 2 |
| 1995 | Cellular automata for efficient parallel logic and fault simulationabstractWe present a unilateral 2D cellular automata (CA) model and pipelining technique to parallelize logic and fault simulation. We show that given an acyclic digraph describing the Boolean function of a combinational circuit at the gate level, whose nodes are the logic gates of the circuit and whose directed edges stand for the propagating directions of signals, we can map this digraph onto a 2D CA to simulate the signal propagation of the circuit on the CA. This mapping preserves not only the electrical connectivity of the circuit but also the massive parallelism inherited from the CA. Experimental results on ISCAS-85 benchmark circuits are obtained. Compared with previous fault simulation results, the time required for simulating one test pattern on an average is shorter by three to four orders of magnitude. As to pure logic simulation, our CA performs up to 9.24 billion gate evaluations per second using a 20 MHz clock and 8-b words. Scalability and extension to sequential circuits are discussed.> Yih-Lang Li, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | C-testable design techniques for iterative logic arraysabstractA design-for-testability (DFT) approach for VLSI iterative logic arrays (ILA's) is proposed, which results in a small constant number of test patterns. Our technique applies to arrays with an arbitrary dimension, and to arrays with various connection types, e.g., hexagonal or octagonal ones. Bilateral ILA's are also discussed. The DFT technique makes general ILA's C-testable by using a truth-table augmentation approach. We propose an output-assignment algorithm for minimizing the hardware overhead. We give a CMOS systolic array multiplier as an example, and show that an overhead of no more than 5.88% is sufficient to make it C-testable, i.e., 100% single cell-fault testable with only 18 test patterns regardless of the word length of the multiplier. Our technique guarantees that the test set is easy to generate. Its corresponding built-in-self-test structures are also very simple.> Shyue-Kung Lu, Jen-Chuan Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1994 | General Modular Multiplication by Block Multiplication and Table LookupabstractThis paper deals with the problem of general modular multiplication, i.e., A/spl times/B mod M. To solve it in hardware, we suggest a simple lookup-table based approach and use the novel block multiplier which we have developed earlier on. The area (A) and time (T) complexities of this multiplier are both O(n log n) if carry-lookahead adders are used. Most previously proposed modular multipliers have put restriction on the ranges of the multiplicand A, the multiplier B, and the modulus M. Our modular multiplier is general, which releases the restriction on those operands.> Cheng-Wen Wu, Yung-Fa Chou |
ISCAS | 1 |
| 1994 | Testing Iterative Logic Arrays for Sequential Faults with a Constant Number of PatternsabstractShows that a constant number of test vectors are sufficient for fully testing a k-dimensional ILA for sequential faults if the cell function is bijective. The authors then present an efficient algorithm to obtain such a test sequence. By extending the concept of C-testability and M-testability to sequential faults, the constant-length test sequence can be obtained. A pipelined array multiplier is shown to be C-testable with only 53 test vectors for exhaustively testing the sequential faults.> Shih-Yuang Su, Cheng-Wen Wu |
IEEE Trans. Computers | 2 |
| 1991 | Designing Self-Testable Cellular ArraysabstractDesign-for-testability techniques and built-in self-test structures are presented for cellular arrays based on the M-testability condition, which results in the minimal number of tests. This technique applies to arrays with arbitrary dimensions and various connections. A systolic array multiplier is given as an example, showing an overhead of only 4% for making it M-testable. This method compares favorably with that based on pI-testability. It reduces drastically the testing costs for circuits realized as cellular arrays.> Cheng-Wen Wu, Shyue-Kung Lu |
ICCD | 1 |
| 1991 | Bit-level pipelined 2-D digital filters for real-time image processingabstractBit-level systolic arrays for real-time 2-D FIR and IIR (finite and infinite impulse response) filters are presented. Two-dimensional iteration and retiming techniques are depicted to illustrate block pipeline 2-D IIR filters, which guarantee high-throughput operation for real-time applications. The block (parallel) systolic architectures are refined down to the bit level. This increases the filter's throughput rate and decreases the filter's development and manufacturing costs. The AP figure is improved from O(N/sup 2/W/sup 3/) for the previous design to O(N/sup 2/W/sup 2/), i.e. by a factor of O(W), where W is the word length. Pipelining at the bit rate level is the major reason for this improvement. Another advantage of the proposed design is that it has simpler wire routing and control circuitry. In summary, these systolic-array realizations are more cost effective; more regular structurally; composed of bit-level cells and latches; and fully pipelined at the bit level.> Cheng-Wen Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1990 | Easily Testable Iterative Logic ArraysabstractIterative logic arrays (ILAs) are studied with respect to two testing problems. First, a variety of conditions is presented. Meeting these conditions guarantees an upper bound on the size of the test set for the ILA under consideration. Second, techniques for designing optimally testable ILAs are presented. The arrays treated are, in some cases, more general than those that have been reported by other researchers: they include multidimensional and inhomogeneous arrays. Octagonally connected arrays and bilateral arrays are also discussed. The results indicate that the characteristics of the individual cell functions (e.g. whether they are bijective) are a good guide to the test complexity of the overall array. Matrix multiplication, as an example, is shown to have several different optimally testable implementations. The results are useful for combinational and pipelined arrays and for certain systolic arrays.> Cheng-Wen Wu, Peter R. Cappello |
IEEE Trans. Computers | 1 |
| 1987 | Computer-aided design of VLSI second-order sectionsabstractA CAD tool is presented for producing very high throughput IIR filters. The architecture is a cascade of word-parallel, blocked, second-order sections. Because it is application-specific, it is a very high level CAD tool. An engineer only needs to specify 1) the word size W, 2) the block size B; and 3) each second-order section's coefficients. Using this information, the CAD tool will generate CIF files for a filter system that, operating at 10MHz, can process 5B/W million samples per second. Our purpose is to illustrate the benefits of applying both bit-level array architecture and application-specific CAD to the problem of IIR filtering. The resulting CAD system reduces the costs of very high throughput IIR filters with respect to design, fabrication, and operation. Cheng-Wen Wu, Peter R. Cappello |
ICASSP | 1 |
| 1987 | Computer-aided design of VLSI FIR filtersabstractA CAD tool is presented for producing very high-throughput FIR filters. Because the CAD tool is application-specific, it is a very high-level tool. An engineer only needs to specify 1) the filter order, N; 2) the input word size; and 3) the output word size. Using this information, the CAD tool generates CIF files for a filter system that can process 10N million samples per second. The purpose of the paper is to illustrate the benefits of applying both bit-level systolic array architecture and application-specific CAD to the problem of FIR filtering. The resulting CAD system reduces the costs of very high-throughput FIR filters with respect to design, fabrication, and operation. Peter R. Cappello, Cheng-Wen Wu |
Proc. IEEE | 2 |