EDBT 2026 Demo / reviewers in the wild / expert
Shi-Yu Huang
dblp:99/2239
· DBLP profile ↗
108ranked-venue papers
34as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 107 · 34 first-author · 22 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Error-Aware ECC Techniques for Enhancing Flash Memory Reliability
Shyue-Kung Lu, Jie-Xin Shi, Shi-Yu Huang, Kohei Miyase |
IOLTS | 3 |
| 2025 | A Speed Learning Scheme To Mitigate The Silent Data Corruption in a Multi-Core DesignabstractIn this work, we propose a speed margining scheme to mitigate silent data corruption in a multi-core system. At the hardware level, we incorporate an enhanced architecture based on dual module redundancy (Enhanced DMR), in which each computing core has two gears: standard performance gear (Standard Gear) and shadow performance gear (Shadow Gear). Such an arrangement allows us to establish a safe performance margin with an SDC-alerting capability. At the system level, we demonstrate the benefits of this method on the ZYNQ™-7010 FPGA platform. We used seven sets of representative learning programs to emulate various workload conditions. Experimental results show that our system can mitigate the SDC risk by establishing a safe performance margin. Yun-Chieh Wang, Chun-Teng Chao, Shi-Yu Huang, Chi-Kang Chen |
ITC-Asia | 3 |
| 2025 | Connectivity-Agnostic Built-In Self-Repair of Interconnects in a Chiplet ICabstractIn a chiplet IC, several dice are integrated through die-to-die interconnects. Technical challenges still exist for the repair of these die-to-die interconnects to boost the overall manufacturing yield. In this work, we propose a novel Connectivity-Agnostic Built-In Self-Repair (BISR) scheme for chiplet ICs. In our scheme, the design-for-BISR circuit inserted in each functional die except the master die is independent of the die-to-die connectivity so that a nonmaster die can be repeatedly reused in many chiplet ICs while supporting in-the-field repair of faulty interconnects to boost the manufacturing yield and in-the-field reliability. Chi Lai, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Built-In Self-Repair of Small Delay Faults Occurring to TSVs in a 3D-DRAM Using an Enhanced Pulse-Vanishing TestabstractIn a 3D-DRAM, multiple DRAM dice are stacked together and bonded vertically with through-silicon vias (TSVs). It is known that a 3D-DRAM could operate at a very high speed, and even a small delay fault could cause a failure. Even though numerous prior works have been proposed to perform built-in self-repair (BISR) for faulty TSVs in a 3D-DRAM, they cannot handle sub-100-ps small delay faults easily. In this work, we aim to fix this problem with a “progressively shrinking pulse-vanishing test (PV-Test).” Our BISR scheme streamlines the entire test-and-repair (TAR) process integrating several techniques, including small-delay-fault detection, on-the-spot diagnosis, test result broadcasting, TSV repair, and the final validation. The experimental results show that it can indeed detect and repair a small delay fault that causes a sub-100-ps extra delay on a TSV. Chen-Yu Huang, Shi-Yu Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Trojan Horse Detection for RISC-V Cores Using Cross-AuditingabstractIn security-critical applications, malicious Trojan Horses embedded in a CPU core could impose great threats on the security of an SoC. In this work, we propose a "Trojan-Horse detection framework" using a cross-auditing scheme. Our framework takes a target RISC-V core, and then pairs it up with another reference RISC-V core to conduct the functional simulation using a set of benchmark programs. The "care outputs" of both cores are compared to reveal the potential Trojan Horses in the target core. A set of well-known Trojan Horses are implanted into an open-source RISC-V core to evaluate the effectiveness of this framework. We found that we can successfully detect almost every implanted Trojan Horse as long as it has been activated and manifested by the benchmark programs. Wei-Po Huang, Shi-Yu Huang, Chi-Kang Chen, Siang-Cheng Huang |
ATS | 2 |
| 2024 | Keynote 2 - Sustainability and the Outlook of Semiconductor IndustryabstractThe semiconductor industry continues to grow rapidly with the development of AI technology and its deployment in various application domains, broadly covering homes, offices, factories, vehicles, warehouses, all kinds of public sites, green energy fields, etc. A modern semiconductor-power AI system consists of not just the AI computing servers in the cloud, but also the 5G or satellite communication network, and a lot of smart Internet-of-Things devices distributed over a large application field. The sustainability of an AI system is of paramount significance to ensure reliable and long-lasting operations without failures. In this talk, we have discussed various strategies to achieve the sustainability of a multi-die IC used in such ever-increasing mission-critical AI-related applications. The techniques include rigorous testing, online monitoring, and self-healing techniques. These techniques target not just the hard failures but also the parametric defects that cause performance hazards, thereby affecting the reliability and or lifetime of an IC product. In particular, we demonstrate with a few examples how various timing circuits such as the Phase-Locked Loop, the Delay-Locked Loop, and the Time-to-Digital Converter can be used as effective building blocks inside an IC to support built-in speed grading, online health condition checking, and online self-repair to achieve the required sustainability. Cheng-Wen Wu, Shi-Yu Huang |
ETS | 2 |
| 2024 | Small-Bridging-Fault-Aware Built-In-Self-Repair for Cycle-Based Interconnects in a Chiplet Design Using Adjusted Pulse-Vanishing TestabstractIn a chiplet design, several dies are integrated through die-to-die interconnects. Technical challenges still exist when it comes to the repair of faulty die-to-die interconnects to boost the overall manufacturing yield. In this work, we aim to perform Built-In Self-Repair (BISR) of small bridging faults occurring to cycle-based interconnects in a chiplet design. We fixed the potential "under-testing" problem in the traditional Pulse-Vanishing Test (PV-test) by adjusting the test pulse width in the test pulse signal. This adjustment involves two steps — the offline adjustment procedure and the on-chip adjustment circuit. The adjusted PV-test scheme can easily support at-speed BISR procedure suitable for arbitrary interconnects in a chiplet design. Chi Lai, Shi-Yu Huang |
ITC | 2 |
| 2024 | Instant Test and Repair for TSVs using Differential Signaling
Ching-Yi Wen, Shi-Yu Huang |
J. Electron. Test. | 2 |
| 2024 | General Fault and Soft-Error Tolerant Phase-Locked Loop by Enhanced TMR using A Synchronization-before-Voting Scheme
Shun-Hua Yang, Shi-Yu Huang |
J. Electron. Test. | 2 |
| 2024 | A Check-and-Balance Scheme in Multiphase Delay-Locked LoopabstractMultiphase delay-locked loop (MP-DLL) is a technique employed in DDR memory controllers to achieve the required fixed timing delay (tSD) for deskew functionality. Traditionally, the tunable delay line (TDL) that consists of the MP-DLL will surfer the delay period mismatch under the same control code, which is mainly caused by the process variation. In this work, a feature called check and balance (CAB) scheme is developed to fix the delay mismatch among each delay stage, thereby providing a more robust and accurate fixed tSD for double-date-rate (DDR) memory controller with a reasonable amount of area overhead. Compared to the baseline MP-DLL design, MP-DLL with CAB can lower the maximum peak-to-peak (P2P) delay mismatch between delay stages from 51.32 to 6.76 ps at 1.6 GHz, utilizing only a 90-nm CMOS process. This improvement comes with an additional area of 0.0045 mm2 and 0.55-mW power consumption at a 1.0-V supply voltage compared to the baseline MP-DLL. Shu-Yu Chang, Shi-Yu Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Low-Jitter Frequency Doubling Circuit Supporting Higher-Speed BISG and Aging Sensing in a Chiplet-Based Design EnvironmentabstractBuilt-in speed grading (BISG) is a technique that measures the maximum operating speed ($F_{\max }$) of a circuit under grading (CUG) in silicon. Recently, it has been reported as an effective aging sensor as well. In nowadays-chiplet-based design, the BISG circuit and the CUG could use different process technologies. In general, the BISG circuit needs to produce a clock signal with a frequency matching the$F_{\max }$of the CUG. In this article, we discuss how to leverage an existing flexible wide-range cell-based phased-locked loop (PLL) with a frequency doubling circuit (FDC) to support even higher$F_{\max }$for a CUG that may use a more advanced technology in another die. As demonstrated in a 90 nm process, a PLL supporting a frequency range of [40 MHz, 1.25 GHz] using a mature 90 nm CMOS process can now support up to 2 GHz. Ko-Hong Lin, Ont-Derh Lin, Shi-Yu Huang, Duo Sheng |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Low Latency Edge Classification GNN for Particle Trajectory Tracking on FPGAsabstractIn-time particle trajectory reconstruction in the Large Hadron Collider is challenging due to the high collision rate and numerous particle hits. Using GNN (Graph Neural Network) on FPGA has enabled superior accuracy with flexible trajectory classification. However, existing GNN architectures have inefficient resource usage and insufficient parallelism for edge classification. This paper introduces a resource-efficient GNN architecture on FPGAs for low latency particle tracking. The modular architecture facilitates design scalability to support large graphs. Leveraging the geometric properties of hit detectors further reduces graph complexity and resource usage. Our results on Xilinx UltraScale+ VU9P demonstrate 1625x and 1574x performance improvement over CPU and GPU respectively. Shi-Yu Huang, Yun-Chen Yang, Yu-Ru Su, Bo-Cheng Lai, Javier M. Duarte, Scott Hauck, Shih-Chieh Hsu, Jin-Xuan Hu, Mark S. Neubauer |
FPL | 1 |
| 2023 | Self-Sufficient Clock Jitter Measurement Methodology Using Dithering-Based CalibrationabstractFor a Phase-Locked Loop (PLL), the variation of its output clock cycle time known as clock jitter is an important performance metric. A complete online clock jitter measurement is often performed in two stages – quantization of the clock cycle times into digital codes, and calibration of the digital codes into absolute jitter information in pico-seconds. The latter stage could be challenging as it requires accurate training clock signals as references. In this article, we propose a dithering-based scheme to resolve this issue with ease. For a cell-based PLL using a 90nm CMOS process, the post-layout transistor-level simulation supports that this is a simple yet effective method. Yi-Hsuan Lee, Wei-Hao Chen, Shi-Yu Huang |
ITC-Asia | 3 |
| 2023 | Trustworthy Lifetime Prediction by Aging History Analysis and Multi-Level Stress TestabstractAn IC used in a safety-critical application such as automotive often requires a long lifetime of more than 10 years. Previously, stress test has been used as a means to establish the accelerated aging model for an IC product under a harsh operating condition. Then, the accelerated aging model is time-stretched to predict an IC’s normal lifetime. However, such a long-stretching prediction may not be very trustworthy. In this work, we present a more refined method to provide higher credibility in the IC lifetime prediction. We streamline in this paper a progressive lifetime prediction method with two phases – the training phase and the inference phase. During the training phase, we collect the aging histories of some training devices under various stress levels. During the inference phase, the extrapolation is performed on the “stressed lifetime” versus the “stress level” space and thereby leading to a more trustworthy prediction of the lifetime. Chen-Lin Tsai, Shi-Yu Huang |
ITC-Asia | 2 |
| 2023 | Compiler of Reed-Solomon Codec for 400-Gb/s IEEE 802.3bs StandardabstractError correction is often indispensable in a modern digital communication system that transmits data at a very high speed. Recently published IEEE Std 802.3bs requires an astounding throughput of 400 Gb/s while using the Reed–Solomon code (RS-Code) to protect the integrity of the transmitted data. An RS-Codec supporting such a high throughput demands a significant silicon area. Improper decisions on the parameters of the parallel architecture could lead to unnecessarily high costs in the implementation. We have developed a compiler to solve this problem. First, the Codec satisfying IEEE Std 802.3bs using RS(544, 514) is parameterized, in a way that the throughput can be boosted on demand by setting some “configuration.” Second, an area- and power-efficient RS-Codec design satisfying a target throughput using a specific process can be inferred by our compiler in just minutes, and thereby easy process migration is supported. Experimental results using 28 and 90-nm CMOS processes are presented to demonstrate their effectiveness. Lai Chi, Shi-Yu Huang, Ka-Yi Yeh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Clock-Latency-Aware Fault-Tolerant DLL for Multi-Die Clock SynchronizationabstractA delay-locked loop (DLL) circuit is indispensable for clock synchronization in a chip incorporating several heterogeneous dice. It has been shown previously that a fault and soft-error-tolerant DLL can be achieved by triple-module redundancy (TMR) enhanced with a timing correction scheme. However, the prior work still has a severe limitation—it does not consider the latency of the clock tree, and this limitation will make it infeasible in realistic situations. We demonstrate in this article that this limitation can be overcome by a new “clock-latency-aware” architecture, thereby making a fault and soft-error-tolerant DLL truly realistic. Yung-Chuan Su, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | A Process-Adaptive Cell-Based Cyclic Time-to-Digital Converter Using One-Way Varactor CellsabstractTime-to-digital converters (TDCs) are commonly used for on-chip instrumentation. This article proposes a low-cost Cyclic-TDC with several distinct features. It can adapt to the process variation to maintain a fine resolution across all process corners. This is made possible by a so-called one-way varactor cell (OWVC) and an autonomous calibration process. In addition, it supports easy process migration. Last but not least, we show how it can be used for clock jitter measurement with reduced output uncertainty. Simulation results using a 90-nm CMOS process and a 40-nm CMOS process are used to demonstrate their advantages. Yung-Chuan Su, Shi-Yu Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Just-Enough Stress Test for Infant-Mortality Screening Using Speed BinningabstractStress test has become increasingly more important to support reliability screening for safety-critical ICs. However, the amount of stress time that needs to be applied to an IC product is not only hard-to-decide but also too time-consuming. This paper formulates a Just-Enough Stress Test (JEST) method so that the stress time can be dramatically reduced without missing out on the weak devices with infant mortality tendency. Experimental results on examples based the on previously proposed failure time model in the literature indicate that the stress test time could be slashed significantly by an order of magnitude. Chen-Lin Tsai, Shi-Yu Huang |
ITC | 2 |
| 2022 | Process-Resilient Fault-Tolerant Delay-Locked Loop Using TMR With Dynamic Timing CorrectionabstractA delay-locked loop (DLL) circuit is often useful for the clock synchronization in a chip incorporating multiple functional dies. In this work, we present a process-resilient fault-tolerant DLL design. We will first show that a naïve triple-module redundancy (TMR) technique cannot work well unless with a simple yet powerful static timing correction scheme to nullify the adversary effect caused by the voter circuit’s delay. Furthermore, we enhance it with a dynamic scheme so that the overall performance is immune to process variation. Through the proposed schemes, the maximum phase error over 1000 clock cycles between the DLL’s input and output clock signals after locking, which is often considered as the most important performance metric, can be reduced tremendously from 130 ps using only naïve TMR to 20 ps using static timing correction, and then further down to 11 ps using the dynamic timing correction. Jun-Yu Yang, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Tiny Phase-Error Monitor for Fault and Soft-Error-Tolerant DLL to Support Graceful Degradation and Module-Level TestingabstractIn an IC used for safety-critical applications, the Fault and soft-error tolerance (or FET) is often desirable. In this work, we consider a graceful degradation scheme, as the second line of defense, for a FET delay-locked loop (DLL) we have recently developed. By doing so, a FET DLL will not operate blindly when its tolerance to faults or soft errors has been degraded. This is achieved by incorporating a novel low-cost excessive phase-error monitor. Any excessive phase error beyond a prelearned phase-error tolerance range will trigger an alarm of failure. This monitor can also be used to support an online test for deciding whether there is a faulty module in our TMR-based FET DLL at any given time. We have implemented the proposed scheme in a 90-nm CMOS process. The results show that the area of this excessive phase error monitor is as small as 60$\mu \text{m}\,\,\times $60$\mu \text{m}\,\,=$0.0036 mm2, or only 4.12% of the entire FET DLL. Jun-Yu Yang, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Rigorous Test Flow for PLL to Identify Weak DevicesabstractIn this work we propose a test method to check the robustness of a Phase-Locked Loop (PLL). The goal is to identify weak devices that could fail under hostile operating conditions. We devise a “rigorous test flow” to measure the performance of a PLL device with “mimicked” hostile operating conditions. At the end of the test, the peak-to-peak jitter is used as a “robustness indicator” to guide the screening process of weak devices. Implementation of the proposed test flow on a PLL in a 90nm CMOS process demonstrates that obscure symptoms that might have escaped a simple traditional test can now be revealed. Yi-Hsuan Lee, Shi-Yu Huang |
ITC-Asia | 2 |
| 2021 | A Duty-Cycle Monitor Supporting A Wide Frequency Range of Clock SignalabstractThis paper presents a synthesizable 50% Duty-Cycle Monitor (DCM) supporting a wide range of clock frequency from 100MHz to 1.2GHz, using in a 90nm CMOS process. We demonstrate that a wide frequency range and a high resolution can be achieved by only standard cells with a reasonable amount of area overhead. Such a design can be easily used as an off-the-shelf IP to check the duty-cycle of a clock signal that should remain at 50% at all times when used in a Double-Data-Rate (DDR) application that captures data at both the rising and the falling clock edges. In our design, a feature - called clock frequency adjustment scheme - is developed to achieve robustness in various process and environmental conditions. The benefit of using such a monitor is that an alarm of performance hazard can be raised whenever there is an excessive Duty-Cycle Error on a DDR clock signal. Chen-Lin Tsai, Wei-Hao Chen, Shi-Yu Huang |
ITC-Asia | 3 |
| 2020 | Overview of On-Chip Performance Monitors for Clock SignalsabstractFor a safety-critical IC, on-chip monitors have become more and more necessary to discover run-time performance hazards as soon as possible. The general performance metrics that have been targeted in the literature are numerous, such as, the maximum operating speed of a circuit component, the interconnect delay, the transition time or leaking current at an IO pin, etc. In this short paper, we will briefly summarize the features of some of them related to a clock signal, in particular, the performance monitors for excessive jitters, phase errors, or duty-cycle errors. Even with only mature process technology, e.g., a 90nm CMOS process, fine-resolution monitors with the time resolution of only a few picoseconds can still be easily achieved by cell-based designs. Shi-Yu Huang |
ATS | 1 |
| 2020 | Fault and Soft Error Tolerant Delay-Locked LoopabstractWe present in this paper the first fault and soft error tolerant Delay-Locked Loop (DLL) design, useful for the clock synchronization in a chip incorporating heterogeneous functional dies. In this robust DLL design, we introduce a powerful timing correction scheme to remedy the timing shortfall in a naïve Triple-Module Redundancy (TMR) architecture. Post-layout simulation results using a 90nm CMOS process is used to verify the performance of this design. In addition to the tolerance of randomly injected faults or soft errors, the Maximum Phase-Error can be improved tremendously from 117ps to just 17ps by the proposed timing correction scheme. Jun-Yu Yang, Shi-Yu Huang |
ATS | 2 |
| 2020 | Duty-Cycle Correction For A Super-Wide Frequency Range from 10MHz to 1.2GHzabstractThis paper presents a cell-based 50% Duty-Cycle Correction (DCC) design supporting a super-wide range of clock frequency from 10MHz to 1.2GHz, using a 90nm CMOS process. It can be integrated with a Delay-Locked Loop (DLL) as a convenient post-processing unit while achieving “zero phase shift” in a way that the phase locking result achieved by its precedent DLL is not affected at all. The unique features in this design include: (1) A wide-range and high-resolution Half-Period Tunable Delay Line (HP-TLD), (2) A fast-locking unit to enable our DCC to lock in to a new incoming clock frequency during frequency scaling, and (3) A wide-range and high-resolution Duty-Cycle Judge (DCJ) circuit as a feedback to guide the overall duty-cycle correction process. Post-layout simulation in a 90nm CMOS process is conducted to validate its effectiveness. Wei-Hao Chen, Shi-Yu Huang |
ICCD | 3 |
| 2020 | Rapid PLL Monitoring By A Novel min-MAX Time-to-Digital ConverterabstractFor a Phase-Locked Loop (PLL), the clock period variation is one important health condition indicator. In this paper, we present a rapid min-MAX period monitoring scheme for PLLs, using circuits made of only standard cells. The proposed scheme can monitor the clock period of a PLL 's output clock signal continuously during a designated monitoring session, while reporting the minimum and maximum clock periods in a timely manner. As a result, performance hazards can be timely exposed and an alarm can be raised earlier. The most unique contribution in this work is the design of a novel min-MAX Time-to-Digital converter (TDC). We have implemented this monitoring scheme and integrated it with a cell-based PLL in a 90nm CMOS process and post-layout simulation is conducted to verify its effectiveness. Experimental results show that the proposed scheme is able to detect some dangerous conditions when the PLL 's output clock signal exhibits abnormal clock cycle times due to some online transient fault. Wei-Hao Chen, Chu-Chun Hsu, Shi-Yu Huang |
ITC | 3 |
| 2020 | Diagnosis of Intermittent Scan Chain Faults Through a Multistage Neural Network Reasoning ProcessabstractDiagnosis of intermittent scan chain failures still remains a hard problem. In this article, we demonstrate that the use of artificial neural networks (ANNs) can lead to significantly higher accuracy. The key of this method is a multistage process incorporating ANNs with gradually refined focuses. During this process, the final fault suspect is elected through multiple rounds of ANN inference, instead of just one round. At each stage, identification of a proper Affine Group, used as the “candidate set of scan cells for the next round of ANN inference,” will influence the final diagnostic accuracy. Thus, we propose a validation-based learning procedure for Affine Group derivation to further boost the final diagnostic accuracy. The experimental results on benchmark circuits have shown that this method is, on the average, 17.46% more accurate than a state-of-the-art commercial tool for intermittent stuck-at-0 faults. Mason Chern, Shih-Wei Lee, Shi-Yu Huang, Yu Huang 0005, Gaurav Veda, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Improving scan chain diagnostic accuracy using multi-stage artificial neural networksabstractDiagnosis of intermittent scan chain failures remains a hard problem. We demonstrate that Artificial Neural Networks (ANNs) can be used to achieve significantly higher accuracy. The key is to take on domain knowledge and use a multi-stage process incorporating ANNs with gradually refined focuses. Experimental results on benchmark circuits show that this method is, on average, 20% more accurate than a state-of-the-art commercial tool for intermittent stuck-at faults, and improves the hit rate from 25.3% to 73.9% for some test-case. Mason Chern, Shih-Wei Lee, Shi-Yu Huang, Yu Huang 0005, Gaurav Veda, Kun-Han Tsai, Wu-Tung Cheng |
ASP-DAC | 3 |
| 2019 | The Ping-Pong Tunable Delay Line In A Super-Resilient Delay-Locked LoopabstractThe Tunable Delay Line (TDL) is the most important building block in a modern cell-based timing circuit such as Phase-Locked Loop (PLL) or Delay-Locked Loop (DLL). In previously proposed TDLs, one dilemma exists -- they cannot be both power efficient and environmentally adaptive at the same time. In this paper, we present an effective solution for such a dilemma - a novel "ping-pong delay line" architecture. The idea is to use two small cell-based delay lines operated in a synergistic manner in the sense that they exchange the "role of command" dynamically like in a ping-pong game, and thereby jointly reacting to severe environmental changes over a very wide range. This proposed ping-pong delay line has been incorporated in a Delay-Locked Loop (DLL) design, to demonstrate its advantages by post-layout simulation. Zheng-Hong Zhang, Shi-Yu Huang |
DAC | 3 |
| 2019 | Online Testing of Clock Delay Faults in a Clock NetworkabstractTraditionally, it has been a difficult task to characterize the quality of a Clock Delay Fault (CDF). Here, a CDF is referred to a delay fault occurring in the clock network that causes an abnormal delay when a clock signal travels through it and thereby causing large-than-expected clock skews at the clock ports of some flip-flops. In a recent work [12], a modified flush test procedure taking short pulses as the test stimuli at selected clock cycles has been proven effective in characterizing a CDF. In this work, we extend this technique to support online Built-In Self-Test (BIST). We investigate two fault detection strategies, namely, valid-range criterion, and valid-span criterion, and we compare their fault detection abilities in terms of the minimum detectable CDF. With the proposed BIST scheme, the health condition of the clock network in a device operating in the field can be inspected on a regular basis so as to take precautions before the clock network breaks down due to deteriorating fault effects. Shi-Yu Huang |
ITC-Asia | 2 |
| 2019 | Overall Strategy for Online Clock System Checking Supporting Heterogeneous IntegrationabstractThis paper discusses how to perform online integrity checking for a clock system spanning all over a heterogeneously integrated IC. A clock system in this work is assumed to consist of two major components - (1) the Delay-Locked Loop (DLL), commonly used for module-to-module clock synchronization, and (2) the clock distribution network (with clock buffers and interconnects). For the DLL part, we incorporate a “phase error monitoring scheme”, which is able to detect abnormal safety hazard, e.g., an instantaneous power glitch. For the clock distribution network, we incorporate a periodic self-test scheme featuring a “special short-pulse driven flush test procedure” to detect any worsening Clock Delay Fault (CDF). The proposed method can help identify a failure threat before it strikes havoc. Post-layout simulation results are presented to demonstrate the effectiveness of the proposed schemes. Shi-Yu Huang |
ITC | 2 |
| 2019 | International Test Conference in Asia (ITC-Asia) - Bridging ITC and Test Community in AsiaabstractPresents the title page of the proceedings record. Kuen-Jong Lee, Shi-Yu Huang, Tomoo Inoue, Yervant Zorian |
ITC | 2 |
| 2018 | Circuit and Methodology for Testing Small Delay Faults in the Clock NetworkabstractA clock network is not only difficult to design, but also challenging to test. For high-performance designs with a rigorous clock-skew requirement, small defects in a clock tree network could lead to unexpected failures in the field and thus need to be identified during the manufacturing test. In this paper, we present a novel flush test procedure to determine if a clock network has any small delay faults. This method does not require any change of the clock network, but it does require a “special test clock signal,” which can be generated on the chip by using only standard cells. Experimental results of transistor-level simulation on benchmark circuits injected with resistive open defects in the layout show that the proposed method is capable of detecting a delay fault as small as 52.8 ps. Shaofu Yang, Zhi-Yuan Wen, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Cloud-Based PVT Monitoring System for IoT DevicesabstractReliability of an IC, concerning if an IC can function reliably over its designated lifetime in the field, has become more and more important in today's safety-critical applications. It is known that reliability can be affected by PVT effects, (Process, Voltage, Temperature). These effects not only depend on the physical locations where an IC is operated, but also vary over time. In this work, we present a cloud-based PVT monitoring system for the Internet of Things (IoT) devices, by taking advantage of its inherent internet connectivity. By doing so, one can know of the PVT status of any IoT device remotely and continually at any time and any place. With the obtained information, a potential PVT-induced failure can be alarmed in advance before it actually strikes, and thereby pre-cautious actions (such as adaptive measures, online repair, or even manual replacement) can be taken in advance to avoid unnecessary system down time. Guan-Hao Lian, Shi-Yu Huang, Wei-yi Chen |
ATS | 2 |
| 2017 | DLL-Assisted Clock Synchronization Method for Multi-Die ICsabstractFor a multi-die IC, the chip-level clock synchronization problem that aims to establish a global clock signal across multiple functional dies is harder to achieve than its single-die counterpart. In this work, we investigate a process resilient solution for this problem by incorporating Delay-Locked Loops (DLLs). The basic idea is to insert a DLL (which can be generated by a DLL compiler) in each functional die so that the clock latency (from a clock source to the clock ports of a number of FFs) in different dies can be dynamically tuned and equalized. This method has a benefit that the clock network of each die can be designed independently, while the clock skew of the entire chip can still be minimized at run-time, in response to its operating environment. In a preliminary study, experimental results on a pseudo 4-die design demonstrates how the clock skew as high as 233ps initially can be reduced to 34ps after the application of the proposed method. Chia-Yuan Cheng, Shi-Yu Huang, Ding-Ming Kwai, Yung-Fa Chou |
ICCD | 2 |
| 2016 | Pre-Bond and Post-Bond Testing of TSVs and Die-to-Die InterconnectsabstractChip-to-chip interconnects can be traditionally tested for stuck-at faults and hard bridging faults via the boundary scan test standard. However, such a traditional test method may not be adequate for 3D-ICs, in which die-to-die interconnects could operate at a very high speed with an end-to-end delay of only a few hundred of picoseconds. Parametric defects (such as small delay faults, resistive open/bridging faults, leakage faults, etc.) have been identified as new threats to the quality, yield, and reliability of a 3D-IC. In response to this challenge, numerous test methods have been developed. In this chapter, we will first discuss common defect types occurring to die-to-die interconnects, and then review various interconnect test methods applicable during the pre-bond stage and the post-bond stage. Shi-Yu Huang |
ATS | 1 |
| 2016 | Die-to-Die Clock Skew Characterization and Tuning for 2.5D ICsabstractAn 2.5D IC could be composed of a number of functional dies operating in locked steps defined by a global clock signal. However, the clock network design across multiple dies may not be trivial as the die-to-die interconnects used for relaying the clock signal from a source die to the other receiver dies could be unpredictable in their delays in real silicon. In this work, we propose a characterization-and-tuning procedure for die-to-die clock skew minimization. Experimental results show that a die-to-die clock skew of 98.3ps due to the differences of process corners can be thereby reduced to only 26.1ps. Shi-Yu Huang, Chih-Chieh Zheng |
ATS | 1 |
| 2016 | Testing of small delay faults in a clock networkabstractA clock network in a 3D-IC is not only difficult to design, but also challenging to test. For high-performance designs with a rigorous clock-skew requirement, studies have shown that small defects in a clock tree network could lead to unexpected failures in the field and thus need to be identified during the manufacturing test. In this paper, we present a novel test method to determine if a clock network has any small delay faults. This method does not require any change of the clock network, and it is capable of detecting a delay fault as small as 50ps through outlier analysis, while locating the FFs affected by the fault. Shaofu Yang, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng |
ETS | 2 |
| 2016 | Online slack-time binning for IO-registered die-to-die interconnectsabstractIn today's multi-die ICs, the die-to-die interconnects are often complicated and susceptible to various kinds of manufacturing defects and stress-induced performance degradation in the field. This phenomenon has prompted a need to perform online monitoring of the signal integrity over the die-to-die interconnects for reliability critical applications. In this work, we present a slack-time binning scheme so that one can constantly quantify the margin of a timing failure threat (TFT) occurring to a registered die-to-die interconnect. The proposed scheme attaches a Slack-Time Monitor (ST-monitor) to each Flip-Flop (FF) that receives a signal transmitted through a die-to-die interconnect under monitoring. Two techniques are introduced to enhance the traditional “Timing-Violation Checker”, namely (1) a tunable guard-band technique, and (2) an offset compensation technique. With these two techniques, one can perform online slack-time binning. Experimental results using a 90nm CMOS process show that the proposed scheme has a low area overhead of only approximately 2.35 times the area of a boundary scan cell. Chih-Chieh Zheng, Shi-Yu Huang, Shyue-Kung Lu, Ting-Chi Wang, Kun-Han Tsai, Wu-Tung Cheng |
ITC | 2 |
| 2015 | Feedback-bus oscillation ring: a general architecture for delay characterization and test of interconnects
Shi-Yu Huang, Meng-Ting Tsai, Kun-Han Tsai, Wu-Tung Cheng |
DATE | 1 |
| 2015 | Monitoring the delay of long interconnects via distributed TDCabstractInterconnects are sophisticated in a multi-die IC using integration technology such as interposer or Wafer-Level Packaging (WLP), and thus they could become vulnerable to early lifetime failure or aging. Our previous work in [12] provides a way to monitor the delay of a TSV non-intrusively by a transition time binning procedure. However, it is not suitable for longer interconnects in an interposer or in the Re-Distribution Layer (RDL) of a WLP-packaged IC, as the cost could become prohibitively high due to large area overhead. To overcome this limitation, we propose a new scheme in this paper by incorporating a Distributed Time-to-Digital Converter (d-TDC). Experimental results indicate that such a scheme can support on-line delay monitoring for long interconnects, while having an area overhead equivalent to 2 boundary scan cells only for each interconnect. Meng-Ting Tsai, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng |
ITC | 2 |
| 2015 | Nonintrusive On-Line Transition-Time Binning and Timing Failure Threat Detection for Die-to-Die InterconnectsabstractDie-to-die interconnects linking multiple functional dies in a modern 3-D or 2.5-D IC by micro-bumps could experience resistance increase after certain time of field operation due to parametric defects or aging. To cope with this reliability threat, we present an “on-line transition-time binning method” that aims to continuously detect excessive transition time occurring at a target die-to-die interconnect. Our method attaches a monitor to the termination end of each target interconnect. Any transition (rising or falling) is converted into a pulse-width first, which is then further compared to a dynamically tunable threshold for a binary pass/fail judgment. By multiple runs of transition-time monitoring while sweeping the threshold incrementally, the “transition-time bin” of each target interconnect can be derived and thereby a timing failure threat can be detected by a monitor center before it actually strikes. Shi-Yu Huang, Meng-Ting Tsai, Hua-Xuan Li, Zeng-Fu Zeng, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | General Timing-Aware Built-In Self-Repair for Die-to-Die InterconnectsabstractA faulty interposer in a 2.5-D integrated circuit often results in a hefty loss as the potentially expensive known-good-dies bonded on the interposer will have to be discarded as well. To avoid such a last-minute loss during a multichip integration process, built-in self-repair (BISR) is highly valuable. Even though there have been many BISR schemes in the literature, the proposed method offers a number of distinct features. First, it can target not only catastrophic faults, but also timing faults. Second, it can be applied to general multi-pin interconnects and it can be applied to repair an interposer with multiple faulty interconnects. Third, it can perform the test-and-then-repair flow on-the-fly, and thereby eliminating the overhead of extra repair storage incurred in previous methods. Shi-Yu Huang, Meng-Ting Tsai, Zeng-Fu Zeng, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | The Design and Experiments of A SID-Based Power-Aware Simulator for Embedded Multicore SystemsabstractEmbedded multicore systems are playing increasingly important roles in the design of consumer electronics. The objective of such systems is to optimize both performance and power characteristics of mobile devices. However, currently there are no power metrics supporting popular application design platforms (such as SID) that application developers use to develop their applications. This hinders the ability of application developers to optimize power consumption. In this article we present the design and experiments of a SID-based power-aware simulation framework for embedded multicore systems. The proposed power estimation flow includes two phases: IP-level power modeling and power-aware system simulation. The first phase employs PowerMixer IP to construct the power model for the processor IP and other major IPs, while the second phase involves a power abstract interpretation method for summarizing the simulation trace, then, with a CPE module, estimating the power consumption based on the summarized trace information and the input of IP power models. In addition, a Manager component is devised to map each digital signal processor (DSP) component to a host thread and maintain the access to shared resources. The aim is to maintain the simulation performance as the number of simulated DSP components increases. A power-profiling API is also supported that developers of embedded software can use to tune the granularity of power-profiling for a specific code section of the target application. We demonstrate via case studies and experiments how application developers can use our SID-based power simulator for optimizing the power consumption of their applications. We characterize the power consumption of DSP applications with the DSPstone benchmark and discuss how compiler optimization levels with SIMD intrinsics influence the performance and power consumption. A histogram application and an augmented-reality application based on human-face-based RMS (recognition, mining, and synthesis) application are deployed as running examples on multicore systems to demonstrate how our power simulator can be used by developers in the optimization process to illustrate different views of power dissipations of applications. Cheng-Yen Lin, Chung-Wen Huang, Chi-Bang Kuan, Shi-Yu Huang, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2014 | On-Line Transition-Time Monitoring for Die-to-Die Interconnects in 3D ICsabstractThrough Silicon Vias (TSVs) are known to be susceptible to various thermal-mechanical and electro-migration effects that could lead to unexpected resistance increase after certain time of field operation. To cope with this reliability threat, we present an on-line monitoring method that aims to continuously detect excessive transition time occurring to a TSV (which often indicates performance degradation). Our method attaches a monitor to the termination end of a TSV. Any transition there is converted into a pulse-width, which is further compared to a dynamically tunable threshold. A Timing Failure Threat (TFT) is thereby detected when the pulse-width exceeds the threshold. This method can be applied to a large number of TSVs easily since the monitoring results are binary and can be stored in FFs forming a scan chain for easy access. Shi-Yu Huang, Hua-Xuan Li, Zeng-Fu Zeng, Kun-Han Tsai, Wu-Tung Cheng |
ATS | 1 |
| 2014 | On-the-fly timing-aware built-in self-repair for high-speed interposer wires in 2.5-D ICsabstractInterposer is a critical component in a 2.5-D IC as it serves as a common platform upon which multiple known good dies are bonded. Any defect in an interposer will lead to a loss of compound yield. To avoid such a last-minute yield loss, we propose a timing-aware Built-In Self-Repair method to increase the fault tolerance of high-speed interposer wires. The most unique feature of our method as compared to previous works is that ours can repair not only catastrophic faults, but also timing faults. We present an on-the-fly test-and-repair flow that can seamlessly link Pulse-Vanishing Test, a previous timing-aware interconnect test method, with a popular redundancy structure. Shi-Yu Huang, Zeng-Fu Zeng, Kun-Han Tsai, Wu-Tung Cheng |
ETS | 1 |
| 2014 | Parametric Fault Testing and Performance Characterization of Post-Bond Interposer Wires in 2.5-D ICsabstractThis paper addresses the testing and characterization of interposer wires in a 2.5-D stacked integrated circuit, which is essential for yield learning and silicon debug. The proposed method provides a number of distinctive features beyond previous works on interposer wire testing. First, we target not only catastrophic types of faults (such as stuck-at faults or hard bridging faults), but also parametric types of faults (including both resistive open faults and resistive bridging faults between interposer wires). Second, our method can also be used to characterize the propagation delay across each fault-free interposer wire. Liren Huang, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Pulse-Vanishing Test for Interposers Wires in 2.5-D ICabstractIn this paper, we present a general at-speed test method for die-to-die interconnects and demonstrate its particular application to the interposer wires in a 2.5-D IC. At the heart of this method is a pulse-vanishing test technique (called PV-test), in which a short-duration pulse signal is applied to an interposer wire under test at the driver end. If this pulse vanishes at the receiver's output, then it indicates the presence of a delay fault. This PV-test technique is effective for detecting not only resistive open faults, but also resistive bridging faults between interposer wires. This method has several other advantages. For example, the implementation is especially easy as it incorporates only logic cells and can be merged with boundary scan cells. Also, it can support on-the-spot diagnosis which is desirable in applications where subsequent built-in self-repair is needed. Shi-Yu Huang, Jeo-Yen Lee, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Parameterized All-Digital PLL Architecture and its Compiler to Support Easy Process MigrationabstractIn this paper, we propose a parameterized digitally controlled oscillator that can produce oscillating-clock signal with the tunable frequency covering an entire designated range. Moreover, we formulate the all-digital phase-locked loop optimization process as a search problem, during which we can find a good configuration that not only meets the user-defined requirement but also achieves a smaller area and lower power consumption than a typical manual design. The silicon measurement results show that this is indeed a promising new alternative for analog phase-locked loops, especially for advanced nanometer technologies. Chao-Wen Tzeng, Shi-Yu Huang, Pei-Ying Chao |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Mid-bond Interposer Wire TestabstractTesting the quality of the interposer wires in a 2.5-D stacked IC can avoid the penalty arising from bonding known-good dies to a defective interposer. However, an interposer is particularly hard to test alone due to the lack of a device layer. This paper addresses this problem by proposing amid-bond test method - in which the interposer is tested after each die is attached to the substrate. We borrow the technique of pre-bond test method for the TSV and enhance it by a length-normalization scheme to cope with interposer wires of diverse wire lengths. Simulations show that this mid-bond test method can catch a high percentage of open faults before subsequent dies are attached. Liren Huang, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng, Stephen K. Sunter |
Asian Test Symposium | 2 |
| 2013 | At-speed BIST for interposer wires supporting on-the-spot diagnosisabstractTesting the speed of post-bond interposer wires in a 2.5-D stacked IC is essential for silicon debugging, yield learning, and even for fault tolerance. In this paper, we present a novel at-speed test technique called Pulse-Vanishing test (PV-test), in which a short-duration pulse signal is applied to an interposer wire under test at the d river end. If the pulse signal can successfully propagate through the interposer wire and reach the other end, then the interposer wire is considered fault-free. Otherwise, it indicates the presence of a delay fault. This new test technique has several technical merits. For example, the Design-for-Testability (DfT) circuit for an interposer wire is similar to the boundary scan cell and can be controlled through scan chain. Also, it can be easily adapted to perform at-speed Built-In Self-Test (BIST) supporting on-the-spot diagnosis. Shi-Yu Huang, Jeo-Yen Lee, Kun-Han Tsai, Wu-Tung Cheng |
IOLTS | 1 |
| 2013 | Delay testing and characterization of post-bond interposer wires in 2.5-D ICsabstractDelay testing and characterization of interposer wires in a 2.5-D stacked IC is essential for yield learning and silicon debug. This paper addresses this problem by proposing a data analysis flow for perturbation-based oscillation test method to cope with the various wire-lengths of the interposer wires. With the proposed method, one can not only detect small delay faults but also characterize the delay across each fault-free interposer wire. Shi-Yu Huang, Liren Huang, Kun-Han Tsai, Wu-Tung Cheng |
ITC | 1 |
| 2013 | Oscillation-Based Prebond TSV TestabstractTesting the quality of prebond through-silicon vias (TSV) is a vital part of the Known-Good-Die test that is often necessary to retain a high compound yield for 3-D stacked integrated circuits. In this paper, we present a versatile prebond TSV test method applicable before wafer thinning when the deep end of the TSV is inaccessible as buried in the still-thick wafer. Technical merits include: 1) the ability to handle both the resistive open fault and the leakage fault in the same test structure; 2) a capability that allows an user to have a better measure of the severity of the fault; and 3) an all-digital and easy to implement design-for-testability circuit. Liren Huang, Shi-Yu Huang, Stephen K. Sunter, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Programmable Leakage Test and Binning for TSVs With Self-Timed Timing ControlabstractLeakage tests have been a challenge for through-silicon vias (TSVs) in a 3-D IC. Most existing methods are still inadequate in terms of the range of testable leakage currents. In this paper, we borrow the wisdom of the IO-pin leakage test while enhancing it with two features. First, we make it more suitable for a TSV, which has a much smaller capacitance than an IO pin. Second, we support a wide range of leakage test (e.g., from 0.125 μA to 16 μA), and thereby allowing for flexible test threshold setting and leakage characterization. To achieve this goal, we present two sets of techniques-1) wait-time generation by programmable delay line, and 2) wait-time propagation with a self-timed timing control scheme to overcome the timing skew problem due to signal routing. We demonstrate that the entire scheme can be done in only logic gates, making it easy to integrate into the common design flow. Shi-Yu Huang, Yu-Hsiang Lin, Liren Huang, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Parametric Delay Test of Post-Bond Through-Silicon Vias in 3-D ICs via Variable Output Thresholding AnalysisabstractA parametric delay fault could arise in a through-silicon via (TSV) of a 3-D IC due to a manufacturing defect. Identification of such a fault is essential for fault diagnosis, yield-learning, and/or reliability screening. In this paper, we present an innovative design-for-testability technique called variable output thresholding. We discovered that by dynamically switching the output of a TSV from a normal inverter to a Schmitt-Trigger inverter, the parametric delay fault on the TSV can be characterized and detected. SPICE simulation reveals that this technique remains effective even when there is significant process variation. A scalable test infrastructure indicates that the test time is modest at only 17.2 ms for 1024 TSVs and 648.8 ms for 32768 TSVs when the test clock is running at 10 MHz. Yu-Hsiang Lin, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng, Stephen K. Sunter, Yung-Fa Chou, Ding-Ming Kwai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Process-Resilient Low-Jitter All-Digital PLL via Smooth Code-JumpingabstractFor an all-digital phase-locked loop, the frequency range supported is often segmented, and this could cause significant jitter when the operating condition (such as the supply voltage and/or the temperature) changes. To address this issue, we present a scheme, called smooth code-jumping, that can stitch together the segmented frequency profile of a digitally controlled oscillator (DCO) into a continuous range, and thereby reduce the jitter significantly. This scheme incorporates a new mirror-DCO-based calibration scheme to take into account process variations. We validate this scheme by test chips in 0.18-μm CMOS technology. Measurement results show that, when operating at 1 GHz, the rms jitter is 4.3 ps (0.43%UI) and the peak-to-peak jitter is 35.6 ps (3.56%UI), respectively. Pei-Ying Chao, Chao-Wen Tzeng, Shi-Yu Huang, Chia-Chien Weng, Shan-Chien Fang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Cell-Based Process Resilient Multiphase Clock GenerationabstractMultiphase clock generation (MPCG) is a problem that aims to generate a sequence of clock signals with the same frequency and uniformly shifted phases. In this brief, we present a cell-based MPCG design with two technical merits. We use a process calibration scheme that makes the per-phase delay (defined as the timing difference between two consecutive phases of clock signals) highly accurate. We further exploit a so-called cyclic property to make the achievable per-phase delay much smaller than a buffer delay. A design with 16-phase clock signal (with the per-phase delay of only 100 ps) is used to demonstrate its effectiveness. Ruo-Ting Ding, Shi-Yu Huang, Chao-Wen Tzeng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | AC-Plus Scan Methodology for Small Delay Testing and CharacterizationabstractSmall delay defects escaping traditional delay testing could cause a device to malfunction in the field and thus detecting these defects is often necessary. To address this issue, we propose three test modes in a new methodology called AC-plus scan, in which versatile test clocks can be generated on the chip by embedding an all-digital phase-locked loop (ADPLL) into the circuit under test (CUT). AC-plus scan can be executed on an in-house wireless test platform called HOY system. The first test mode of our AC-plus scan provides a more efficient way to measure the longest path delay associated with each test pattern. Experimental result shows that our method could greatly reduce the test time by 81.8%. The second test mode is designed for volume production test. It could effectively detect small delay defects and provide fast characterization on those defective chips for further processing. This mode could be used to help predict which chips are more likely to fall victim to operational failure in the field. The third test mode is to extract the waveform of each flip-flop's output in a real chip. This is made possible by taking advantage of the almost unlimited test memory our HOY test platform provides, so that we could easily store a great volume of data and reconstruct the waveform for post-silicon debugging. We have successfully fabricated a Viterbi decoder chip with such an AC-plus scan methodology inside to demonstrate its capability. Tsung-Yeh Li, Shi-Yu Huang, Hsuan-Jung Hsu, Chao-Wen Tzeng, Chih-Tsun Huang, Jing-Jia Liou, Hsi-Pin Ma, Po-Chiun Huang, Jenn-Chyou Bor, Ching-Cheng Tien, Chi-Hu Wang, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | In-Situ Method for TSV Delay Testing and Characterization Using Input Sensitivity AnalysisabstractIn this paper, we propose a method and the required architecture for characterizing the propagation delays of the through Silicon vias (TSVs) in a 3-D IC. First of all, every two TSVs are paired up to form an oscillation ring with some peripheral circuits. Their joint performance can thus be measured roughly by the oscillation period of the ring. Next, we utilize a technique called sensitivity analysis to further derive the propagation delay of each individual TSV participating in an oscillation ring-a distilling process. In this process, we perturb the strength of the two TSV drivers, and then measure their effects in terms of the change of the oscillation ring's period. By some following analysis, the propagation delay of each TSV can be revealed. On top of scheme, we also present an architecture that can activate the performance characterization process of each test unit - that consists of two TSVs - one at a time in a proper sequence. The area overhead is only 18.97 equivalent two-input NAND gate per TSV, by which one can gain the ability to profile the capacitances and the propagation delays of the TSVs on a 3-D IC. Jhih-Wei You, Shi-Yu Huang, Yu-Hsiang Lin, Meng-Hsiu Tsai, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Programmable Leakage Test and Binning for TSVsabstractLeakage test has been a challenge for TSVs in a 3D IC. Most existing methods are still inadequate in terms of the range of leakage currents they can test. In this work, we borrow the wisdom of the IO pin leakage test while enhancing it with two features: (1) we make it more suitable for TSVs which has a much smaller capacitance than an IO pin, and (2) we support leakage binning in a wide range of currents from 1 uA to 128 uA, and thereby allowing flexible test threshold settings and leakage characterization. Since we use only logic gates in the Design-for-Testability circuit, it is also easier to be integrated into the TSV design flow than previous methods. Yu-Hsiang Lin, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng |
Asian Test Symposium | 2 |
| 2012 | Small delay testing for TSVs in 3-D ICsabstractIn this work, we present a robust small delay test scheme for through-silicon vias (TSVs) in a 3D IC. By changing the output inverter's threshold of a TSV in a testable oscillation ring structure, we can approximate the propagation delay across that TSV, and thereby detecting a small delay fault. SPICE simulation reveals that this Variable Output Thresholding (VOT) technique is still effective even when there is significant process variation in detecting a slow TSV with some resistive open defect that may escape the traditional at-speed test. Shi-Yu Huang, Yu-Hsiang Lin, Kun-Han Tsai, Wu-Tung Cheng, Stephen K. Sunter, Yung-Fa Chou, Ding-Ming Kwai |
DAC | 1 |
| 2012 | A unified method for parametric fault characterization of post-bond TSVsabstractA TSV in a 3D IC could suffer from two major types of parametric faults — a resistive open fault, or a leakage fault. Dealing with these parametric faults (which do not destroy the functionality of a TSV completely but only degrade its quality or performance) is often trickier than dealing with a stuck-at fault. Previous works have not proposed a unified test structure and method that can characterize their respective effects. Based on our previous test structure, called VOT (Variable Output Threshold) scheme for delay faults, we propose a unified in-situ characterization flow for both parametric fault types of a post-bond TSV. With this flow, one can easily derive a more insightful assessment of a parametric fault in production test, process monitoring, and/or diagnosis-driven yield learning. Yu-Hsiang Lin, Shi-Yu Huang, Kun-Han Tsai, Wu-Tung Cheng, Stephen K. Sunter |
ITC | 2 |
| 2012 | High-Performance SIFT Hardware Accelerator for Real-Time Image Feature ExtractionabstractFeature extraction is an essential part in applications that require computer vision to recognize objects in an image processed. To extract the features robustly, feature extraction algorithms are often very demanding in computation so that the performance achieved by pure software is far from real-time. Among those feature extraction algorithms, scale-invariant feature transform (SIFT) has gained a lot of popularity recently. In this paper, we propose an all-hardware SIFT accelerator-the fastest of its kind to our knowledge. It consists of two interactive hardware components, one for key point identification, and the other for feature descriptor generation. We successfully developed a segment buffer scheme that could not only feed data to the computing modules in a data-streaming manner, but also reduce about 50% memory requirement than a previous work. With a parallel architecture incorporating a three-stage pipeline, the processing time of the key point identification is only 3.4 ms for one video graphics array (VGA) image. Taking also into account the feature descriptor generation part, the overall SIFT processing time for a VGA image can be kept within 33 ms (to support real-time operation) when the number of feature points to be extracted is fewer than 890. Feng-Cheng Huang, Shi-Yu Huang, Ji-Wei Ker, Yung-Chang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | PowerDepot: integrating IP-based power modeling with ESL power analysis for multi-core SoC designsabstractIn this paper, we introduce an integrated power methodology for multi-core SoC designs. It features not only a bottom-up IP-based power modeling for all kinds of IP components ranging from hardware accelerators, processors, and memory blocks, but also a top-down system-wide ESL power estimation formulation. By linking these two methods of different levels of abstraction, one can thereby easily profile the power consumption of a multi-core SoC running a complete application while retaining high accuracy of estimation. We have realized the proposed methodology into two software tools: (1) PowerMixerIP, an IP-based power model builder that uses different strategies to build versatile power models for general IPs and processor IPs, and (2) PowerDepot, an ESL power estimation tool that can interact with the users in a simple way and then generate the needed power monitors to be embedded into the ESL design in SystemC for super-fast power estimation so as to facilitate early-stage system-wide power profiling. The application of these tools on a dual-core real-life designs executing an H. 264 shows that the average error of the ESL power estimation is less than 2%, while the speedup can be up to 2400X when comparing to gate-level simulation. Chen-Wei Hsu, Jia-Lu Liao, Shan-Chien Fang, Chia-Chien Weng, Shi-Yu Huang, Wen-Tsan Hsieh, Jen-Chieh Yeh |
DAC | 5 |
| 2011 | A low-cost wireless interface with no external antenna and crystal oscillator for cm-range contactless testingabstractThis work presents a low-cost wireless system design that serves as an interface to support the SoC with contactless testability feature. The communication hierarchy includes PHY, MAC, data exchange, and test wrapper functions. The wireless does not require external antennae and crystal reference, and therefore minimize the setup cost. The embedded all-digital timing generation achieves robust performance in the noisy environment. The whole wireless system occupies a small area. In a 0.18μm device-under-test, the active area of wireless front-end is 0.14mm2 and the gate count for digital processing is 112K. The maximum energy efficiency for uplink is 1.1nJ/bit and for downlink is 2.9nJ/bit when the wireless distance is set around 1cm. The prototype system includes test equipment and an SoC as the device-under-test. The SoC integrating logic, memory, and analog plug-in modules can be contactlessly tested. It is a low-cost platform controlled by a simple hand-held computer. Chin-Fu Li, Chi-Ying Lee, Chen-Hsing Wang, Shu-Lin Chang, Li-Ming Denq, Chun-Chuan Chi, Hsuan-Jung Hsu, Ming-Yi Chu, Jing-Jia Liou, Shi-Yu Huang, Po-Chiun Huang, Hsi-Pin Ma, Jenn-Chyou Bor, Cheng-Wen Wu, Ching-Cheng Tien, Chi-Hu Wang, Yung-Sheng Kuo, Chih-Tsun Huang, Tien-Yu Chang |
DAC | 10 |
| 2011 | Black-box leakage power modeling for cell library and SRAM compilerabstractIn this paper, we present an automatic leakage power modeling method for standard cell library as well as SRAM compiler. For this problem, there are two major challenges - (1) the high sensitivity of leakage power to the temperature (e.g., the leakage power of an inverter can be different by 19.28X when temperature rises from 25°C to 100°C in 90nm technology), and (2) the large number of models to be built (e.g., there could be 80,835 SRAM macros supported by an SRAM compiler). Our method achieves high accuracy efficiently by two formula-based prediction techniques. First of all, we incorporate a quick segmented exponential interpolation scheme to take into account the effects of the temperature. Secondly, we use a MUX-oriented linear extrapolation scheme, which is so accurate that it allows us to build the leakage power models for all SRAM macros based on linear regression using only the simulation results of 9 small-sized SRAM macros. Experimental results show that this method is not only accurate but also highly efficient. Chun-Kai Tseng, Shi-Yu Huang, Chia-Chien Weng, Shan-Chien Fang, Ji-Jan Chen |
DATE | 2 |
| 2011 | A fully cell-based design for timing measurement of memoryabstractThis work presents a scheme for measuring the timing parameters of a memory's I/O interface - including the setup/hold time and access time. For setup/hold time measurement, we incorporate a procedure that successively adjusts the timing relation between the clock signal and a controllable valid timing window to estimate the setup/hold time. For access time measurement, we propose a circuit that can capture the worst-case access time of an entire Built-In Self-Test (BIST) session as a pulse width, which is then further measured by traditional time-to-digital converter (TDC). Instead of just reporting a digital code, we also propose a calibration scheme so that we can report not just some digital codes, but also their corresponding absolute values. All the design can be constructed by standard cells. We have implemented it in TSMC 0.18nm CMOS process technology. Simulation results show that the setup time error is less than 3%, the hold time error is 7.5%, and the access time error is 4.4%, with about 6.2% area overhead when the memory size is 4096×64. Yi-Chung Chang, Shi-Yu Huang, Chao-Wen Tzeng, Jack T. Yao |
ITC | 2 |
| 2011 | A Low-Jitter ADPLL via a Suppressive Digital Filter and an Interpolation-Based Locking SchemeabstractIn this brief, we present a low-jitter and wide-range all-digital phase-locked loop (ADPLL). This ADPLL achieves low output clock jitter by a number of schemes. First, the phase is locked quickly through a predictive phase-locking scheme. Then, the jitter is further reduced by a suppressive digital loop filter. Finally, an interpolation-based locking scheme is utilized to enhance the resolution of the digitally controlled oscillator (DCO) so as to further reduce the phase error and jitter. Simulation results show that the jitter performance is very close to that of the free-running DCO. Measurement results show that the jitterPk-Pkand jitterRMSare 56 and 7.28 ps, respectively, when the output clock of the ADPLL is running at 600 MHz. Hsuan-Jung Hsu, Shi-Yu Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | PAC duo system power estimation at ESLabstractIn this work, we develop an electronic system-level (ESL) power estimation framework which uses the specified power model interface. Using the proposed power model interface we can easily integrate the various power models in ESL virtual platform. Designers can choose either the coarse-grained or fine-grained power models according to the trade-off between accuracy and computing cost. The experimental results show the proposed method can accurate estimate the system power trend immediately compared with traditional method. We also demonstrated the capability of system power and performance analysis in both hardware-view and software-view by using our approach at ESL. Meanwhile, it can be used for high level architecture exploration directly. Wen-Tsan Hsieh, Jen-Chieh Yeh, Shi-Yu Huang |
ASP-DAC | 3 |
| 2010 | Performance Characterization of TSV in 3D IC via Sensitivity AnalysisabstractIn this paper, we propose a method that can characterize the propagation delays across the Through Silicon Vias (TSVs) in a 3D IC. We adopt the concept of the oscillation test, in which two TSVs are connected with some peripheral circuit to form an oscillation ring. Upon this foundation, we propose a technique called sensitivity analysis to further derive the propagation delay of each individual TSV participating in the oscillation ring-a distilling process. In this process, we perturb the strength of the two TSV drivers, and then measure their effects in terms of the change of the oscillation ring's period. By some following analysis, the propagation delay of each TSV can be revealed. Monte-Carlo analysis of a typical TSV with 30% process variation on transistors shows that the characterization error of this method is only 2.1% with the standard deviation of 8.1%. Jhih-Wei You, Shi-Yu Huang, Ding-Ming Kwai, Yung-Fa Chou, Cheng-Wen Wu |
Asian Test Symposium | 2 |
| 2010 | Split-Masking: An Output Masking Scheme for Effective Compound Defect Diagnosis in Scan Architecture With Test CompressionabstractIn modern scan architecture, it is often desirable to compact the output response without jeopardizing the diagnostic resolution. In this paper, we propose an output masking scheme to meet such a stringent requirement. We consider a practical scenario in which an output compactor is in use. We aim to support the harshest condition called compound defect diagnosis, in which faults exist in both the scan chain and the core logic. To overcome the loss of the diagnostic resolution, we incorporate a split-masking scheme, by which one can easily separate the output responses of the faulty chains from those of the fault-free ones. The experimental results demonstrate that the proposed scheme can recover the diagnostic resolution loss induced by an output compactor almost completely without sacrificing the compaction ratio. Chao-Wen Tzeng, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | QC-Fill: An X-Fill method for quick-and-cool scan testabstractIn this paper, we present an X-Fill (QC-Fill) method for not only slashing the test time but also reducing the test power (including both capture power and shifting power). QC-Fill, built upon the existing multicasting scan architecture, can coexist with most low-capture-power (LCP) X-fill methods through a multicasting-driven X-Fill method incorporating a clique-stripping scheme. QC-Fill is independent of the ATPG patterns and does not require any area-overhead since it can directly operate on an existing scan architecture incorporating test compression. Chao-Wen Tzeng, Shi-Yu Huang |
DATE | 2 |
| 2009 | Layout-Based Defect-Driven Diagnosis for Intracell Bridging DefectsabstractThis paper presents a layout-based methodology to predict the exact physical location of a bridging defect inside a standard cell. It involves a number of techniques. First of all, most likely intracell bridging defects are identified through layout analysis and then converted into equivalent logic models. Next, we use a new defect-oriented formulation to generate test pattern for each candidate defect so as to further enhance the diagnostic resolution. Experimental results indicate that this methodology can remove 90% false defect candidates beyond gate-level diagnosis for four real designs and ISCAS'85 benchmark circuits. Chao-Wen Tzeng, Han-Chia Cheng, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | QC-Fill: Quick-and-Cool X-Filling for Multicasting-Based Scan TestabstractThis paper presents anX-fillscheme that properly utilizes the don't-care bits in test patterns to simultaneously reduce the test time as well as the test power (including both capture power and shifting power). This scheme, called Quick-and-Cool X-fill (QC-Fill), built upon the multicasting-based scan architecture, further leverages on the merits of previous low-capture-power X-fill methods through techniques likemulticasting-driven X-fillandclique stripping.QC-Fillis independent of the automatic test pattern generation patterns and does not require any extra area overhead. Experimental results demonstrate that this scheme strikes a good balance between the seemingly conflicting criteria of low power and test compression. Chao-Wen Tzeng, Shi-Yu Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Two-Gear Low-Power Scan TestabstractWe introduce in this paper a new scan test methodology that can be programmed to execute in one of two gears - that is, the low-shifting-power scan, and the low-capture-power scan, in one single scan-architecture. This two-gear method provides layered treatment to potential power-induced test failure. First, it attempts to perform scan test in gear 1 without any test time overhead over the traditional scan. Second, in case the failure is caused by excessive capture power, then the test yield loss could be recovered by switching the operation to gear 2, which is a low-capture-power mode, at the cost of extra test time. This methodology is fully compatible with existing DFT tool and requires only small area overhead. Chao-Wen Tzeng, Shi-Yu Huang |
ATS | 2 |
| 2008 | A versatile paradigm for scan chain diagnosis of complex faults using signal processing techniquesabstractScan chains are popularly used as the channels for silicon testing and debugging. However, they have also been identified as one of the culprits of silicon failure more recently. To cope with this problem, several scan chain diagnosis approaches have been proposed in the past. The existing methods, however, suffer from one common drawback—that is, they rely on fault models and matching heuristics to locate the faults. Such a paradigm may run into difficulty when the fault under diagnosis does not match the fault model exactly, for example, when there is a bridging between a flip-flop and a logic cell, or the fault is temporal and only manifests itself intermittently. In light of this, we propose in this article a more versatile model-free paradigm for locating the faulty flip-flops in a scan chain, incorporating a number of signal processing techniques, such as filtering and edge detection. These techniques performed on the test responses of the failing chip under diagnosis directly can effectively reveal the fault location(s) in a scan chain. As compared to the previous works, our approach is better capable of handling intermittent faults and bridging faults, even under nonideal conditions, for example, when the core logic is also faulty. Experimental results on several real designs indicate that this approach can indeed catch some nasty faults that previous methods could not catch. Chao-Wen Tzeng, Jheng-Syun Yang, Shi-Yu Huang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2007 | Built-In Speed Grading with a Process-Tolerant ADPLLabstractSpeed grading has becoming more and more important for nanometer technologies to support activities like process monitoring or performance diagnosis. In this work, we analyze the feasibility of providing such a capability through on-chip circuitry. This Built-in Speed Grading (BISG) methodology uses an All-Digital Phase-Locked Loop (ADPLL) as the programmable clock generator to provide various clock signals within a specific frequency range. The maximum operating speed of a circuit can thus be easily tracked down using a binary search process with multiple runs of built-in self-test. To accommodate larger process variation, we further explore a so-called binary-neighborhood-linear frequency-locking scheme for the underlying ADPLL, and thereby resulting in a higher accuracy. Experimental results show that only 2289 gates are adequate to provide this valuable infrastructure that may find numerous applications in IC testing and diagnostics. Hsuan-Jung Hsu, Chun-Chieh Tu, Shi-Yu Huang |
ATS | 3 |
| 2007 | RT-level vector selection for realistic peak power simulationabstractWe present a vector selection methodology for estimating the peak power dissipation in a CMOS logic circuit. The ultimate goal is to combine the speed of RT-level simulation with the accuracy of low-level power simulation. We rely on efficient RT-level peak power prediction heuristics to select a handful of input vector pairs from the simulation testbench. These vector pairs are highly likely to induce the worst-case peak power. After that, the low-level power simulation is performed only with these peak power candidate vector pairs to obtain the realistic peak power values. Computationally, there are three major stages in our methodology. Firstly, we analyze the structure of the circuit and calculate the peak power weight for each input pin so as to construct a so-called mountain-based model. Secondly, we perform waveform composition to compute the peak power metric for each given input vector pair. Finally, we select a few candidate vectors for low-level power simulation. By doing so, the time-consuming simulation at the low levels can be mostly avoided without losing accuracy. Experiment results show that only less than 1% of the total functional patterns defined in a testbench are needed to be selected for low-level power simulation in order to catch the worst-case peak power vector pair. Chia-Chien Weng, Ching-Shang Yang, Shi-Yu Huang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2006 | A network security processor design based on an integrated SOC design and test platformabstractIn this paper we present a generic network security processor (NSP) design suitable for a wide range of security related protocols in wired or wireless network applications. Following the platform-based design methodology, we develop four specific platforms, i.e., architecture platform, EDA platform, design-for-testability (DFT) platform, and prototyping platform, for our NSP design. With these platforms, design of the NSP chip becomes more efficient and systematic. A prototype chip of the NSP has been implemented and fabricated with a 0.18 /spl mu/m CMOS technology. The chip area is 5 mm /spl times/ 5 mm (with 1M gates approximately), including I/O pads. The operating clock rate is 80 MHz. The best performance of the crypto-engines is 1.025 Gbps for AES, 1.652 Mbps for RSA, 125.9/157.65 Mbps for HMAC-SHA1/MD5, and 2.56 Gbps for random number generator. Comparison result shows that our NSP is efficient in terms of performance, flexibility and scalability. Chen-Hsing Wang, Chih-Yen Lo, Min-Sheng Lee, Jen-Chieh Yeh, Chih-Tsun Huang, Cheng-Wen Wu, Shi-Yu Huang |
DAC | 7 |
| 2006 | Accurate Whole-Chip Diagnostic Strategy for Scan Designs with Multiple Faults
Yu-Chiun Lin, Shi-Yu Huang |
J. Electron. Test. | 2 |
| 2005 | Power estimation starategies for a low-power security processorabstractIn this paper, we present the power estimation methodologies for the development of a low-power security processor that contains significant amount of logic and memory. For the logic part, we present a highly accurate tool, called PowerMixer. This tool is a refinement of the so-called mixed-level methodology that combines the accuracy of quick SPICE and the speed of gate-level simulation. A grouping scheme is proposed so as to improve the accuracy for design blocks as large as 100K gates. For the memory part, we investigated the power consuming behavior of memories and point out the potential problems associated with the current commercial design flow. These tools, along with a previously published static peak power estimation method [4], jointly provide an evaluation platform for the power optimization and verification process of our security processor in a practical way. Yen-Fong Lee, Shi-Yu Huang, Sheng-Yu Hsu, I-Ling Chen, Cheng-Tao Shieh, Jian-Cheng Lin, Shih-Chieh Chang 0001 |
ASP-DAC | 2 |
| 2005 | Quick Scan Chain Diagnosis Using Signal ProfilingabstractIn this paper we address the scan chain diagnosis problem. We propose a new diagnosis flow based on the concept of signal profiling to accurately pinpoint the location of a faulty flip-flop in a scan chain. As compared to the conventional cause-effect or effect-cause analysis, this approach is much more computationally efficient because it does not have to simulate the behaviors of a large number of fault candidates. Also, it is general and applicable to all kinds of faults because it does not assume any specific fault model. Experimental results indicate that this approach can instantly catch a fault within a scan chain quite accurately in most cases. Jheng-Syun Yang, Shi-Yu Huang |
ICCD | 2 |
| 2003 | Chip-Level Diagnostic Strategy for Full-Scan Designs with Multiple FaultsabstractFault diagnosis of full-scan designs has been progressed significantly However, most existing techniques are aimed at a logic block with a single fault. Strategies on top of these block-level techniques are needed in order to successfully diagnose a large chip with multiple faults. In this paper, we present such a strategy. Our strategy is effective in identifying more than one fault accurately. It proceeds in two phases. In the first phase we concentrate on the identification of the so-called structurally independent faults based on a concept referred to as word-level prime candidate, while in the second phase we further trace the locations of the more elusive structural dependent faults. Experimental results show that this strategy is able to find 3 to 4 faults within 10 signal inspections for three designs randomly injected with 5 node-type or stuck-at faults. Yu-Chiun Lin, Shi-Yu Huang |
Asian Test Symposium | 2 |
| 2003 | Decomposition of Extended Finite State Machine for Low Power Design
MingHung Lee, TingTing Hwang, Shi-Yu Huang |
DATE | 3 |
| 2003 | A Symbolic Inject-and-Evaluate Paradigm for Byzantine Fault Diagnosis
Shi-Yu Huang |
J. Electron. Test. | 1 |
| 2002 | Diagnosis Of Byzantine Open-Segment FaultsabstractThis paper addresses the problem of locating the stuck-open faults in a manufactured IC with scan flip-flops. Unlike most previous methods that only aim at identifying the faulty signals, our goal is to further narrow down the faults to a few suspected segments. With such a technique, the silicon inspection time could be dramatically slashed when the fault occurs to a long-running wire with a large number of fanouts. The algorithm is based on our previous inject-and-evaluate paradigm using symbolic simulation. It is fast and accurate. For ISCAS85 benchmark circuits with only one stuck-open fault, the first-hit index is 4.5 on the average within 10 seconds of CPU time. Shi-Yu Huang |
Asian Test Symposium | 1 |
| 2002 | Speeding Up The Byzantine Fault Diagnosis Using Symbolic SimulationabstractFault diagnosis is to predict the potential fault sites in a logic IC. In this paper, we particularly address the problem of diagnosing faults that exhibit the so-called Byzantine General's phenomenon, in which a fault manifests itself as a non-logical voltage level at the fault site. Previously, explicit enumeration was suggested to deal with such a problem. However, it is often too time-consuming because the CPU time is exponentially proportional to fanout degree of the circuit under diagnosis. To speed up this process, we present an implicit enumeration technique using symbolic simulation. Experimental results show that the CPU time can be improved by several orders of magnitude for ISCAS85 benchmark circuits. Shi-Yu Huang |
VTS | 1 |
| 2001 | Towards the logic defect diagnosis for partial-scan designsabstractLogical defect diagnosis is a critical yet challenging process in VLSI manufacturing. It involves the identification of the defect spots in a logic IC that fails testing. In the last decade, algorithms for diagnosis have progressed significantly and the results are showing promise for full -scan designs. In this paper, we will first review several classical algorithms such as fault dictionary based analysis and effect cause analysis. Then, we discuss several diagnosis algorithms borrowed from the design debugging techniques. These algorithms do not require a pre-determined fault model, and thus, are more flexible and applicable to ICs in which the defects do not behave like common stuck-at or bridging faults. Finally, we will probe the possibility of extending these algorithms to designs with only partial-scan support. Shi-Yu Huang |
ASP-DAC | 1 |
| 2001 | On speeding up extended finite state machines using catalyst circuitryabstractWe propose a timing optimization technique for a complex finite state machine that consists of not only random logic but also data operators. In such a design, the timing critical path often forms a cycle and thus cannot be cut down easily by popular techniques such as pipelining or retiming. The proposed technique, based on the concept of catalyst, adds a functionally redundant block - which includes a piece of combinational logic and several other registers - to the circuits under consideration so that the timing critical paths are divided into stages. During this transformation, the circuit's functionality is not affected, while the speed is improved significantly. This technique has been successfully applied to an industrial application - a Built-In Self-Test (BIST) circuit for static random access memories (SRAMs). The synthesis result indicates a 47% clock cycle time reduction. Shi-Yu Huang |
ASP-DAC | 1 |
| 2001 | A Built-in Self-Test and Self-Diagnosis Scheme for Heterogeneous SRAM ClustersabstractTesting and diagnosis are important issues in system-on-chip (SoC) development, as more and more embedded cores are being integrated into the chips. In this paper we propose a built-in self-test (BIST) and self-diagnosis (BISD) scheme for embedded SRAMs, suitable for SoC applications. It supports manufacturing test as well as diagnosis for design verification and yield improvement. With low hardware cost, our memory BISD approach can handle various types of SRAM, including pipelined, multi-port, and multi-clock architectures. In addition, a test scheduling methodology and a BISD compiler are also implemented, which reduce the testing time as well as test development time. Chih-Wea Wang, Ruey-Shing Tzeng, Chi-Feng Wu, Chih-Tsun Huang, Cheng-Wen Wu, Shi-Yu Huang, Shyh-Horng Lin, Hsin-Po Wang 0002 |
Asian Test Symposium | 6 |
| 2001 | On Improving the Accuracy Of Multiple Defect DiagnosisabstractLogic defect diagnosis locates the defect spots in a digital IC that fail testing. It is one of the critical steps during the process of manufacturing yield improvement. Automatic defect diagnosis techniques for circuits with single defects have been improved significantly. However, the techniques for multiple defect diagnosis are still inadequate. In this paper, we propose an effective heuristic for diagnosing a full-scan design with multiple defects. Concepts called curable vectors and curable outputs are incorporated. By combining these two measures as a grading criterion, each signal's possibility of being one of the defect spots is calculated with a high accuracy. Experimental results on ISCAS85 benchmark circuits indicate that the proposed method indeed outperforms the conventional heuristics. Shi-Yu Huang |
VTS | 1 |
| 2001 | Verifying sequential equivalence using ATPG techniquesabstractIn this paper we address the problem of verifying the equivalence of two sequential circuits. State-of-the-art sequential optimization techniques such as retiming and sequential redundancy removal can handle designs with up to hundreds or even thousands of flip-flops. However, the BDD-based approaches for verifying sequential equivalence can easily run into memory explosion for such designs. In an attempt to handle larger circuits, we modify test pattern-generation techniques for verification. The suggested approach utilizes the popular efficient backward-justification technique used in most sequential ATPG programs. We present several techniques to enhance the efficiency of this approach by (1) identifying equivalent flip-flop pairs using an induction-based algorithm, and (2) generalizing the idea of exploring the structural similarity between circuits to perform verification in stages. This ATPG-based framework is suitable for verifying circuits either with or without a reset state. In order to extend this approach to verify retimed circuits, we introduce a delay-compensation-based algorithm for preprocessing the circuits. The experimental results of verifying the correctness of circuits after sequential redundancy removal and retiming with up to several hundred flip-flops are presented. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2000 | High Performance/Delay Testing
Shi-Yu Huang, Sudhakar M. Reddy |
Asian Test Symposium | 1 |
| 2000 | AQUILA: An Equivalence Checking System for Large Sequential DesignsabstractIn this paper, we present a practical method for verifying the functional equivalence of two synchronous sequential designs. This tool is based on our earlier framework that uses Automatic Test Pattern Generation (ATPG) techniques for verification. By exploring the structural similarity between the two designs under verification, the complexity can be reduced substantially. We enhance our framework by three innovative features. First, we develop a local BDD-based technique which constructs Binary Decision Diagram (BDD) in terms of some internal signals, for identifying equivalent signal pairs. Second, we incorporate a technique called partial justification to explore not only combinational similarity, but also sequential similarity. This is particularly important when the two designs have a different number of flip-flops. Third, we extend our gate-to-gate equivalence checker for RTL-to-gate verification. Two major issues are considered in this extension: (1) how to model and utilize the external don't care information for verification; and (2) how to extract a subset of unreachable states to speed up the verification process. Compared with existing approaches based on symbolic Finite State Machine (FSM) traversal techniques, our approach is less vulnerable to the memory explosion problem and, therefore, is more suitable for a lot of real-life designs. Experimental results of verifying designs with hundreds of flip-flops will be presented to demonstrate the effectiveness of this approach. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, Chung-Yang Huang, Forrest Brewer |
IEEE Trans. Computers | 1 |
| 1999 | Fault emulation: A new methodology for fault gradingabstractIn this paper, we introduce a method that uses the field programmable gate array (FPGA)-based emulation system for fault grading. The real-time simulation capability of a hardware emulator could significantly improve the performance of fault grading, which is one of the most time consuming tasks in the circuit design and test process. We employ a serial fault emulation algorithm enhanced by two speed-up techniques. First, a set of independent faults can be injected and emulated at the same time. Second, multiple dependent faults can be simultaneously injected within a single FPGA-configuration by adding extra circuitry. Because the reconfiguration time of mapping the numerous faulty circuits into the FPGA's is pure overhead and could be the bottleneck of the entire process, using extra circuitry for injecting a large number of faults can reduce the number of FPGA-reconfigurations and, thus, improving the performance significantly. In addition, we address the issue of handling potentially detected faults in this hardware emulation environment by using the dual-railed logic. The performance estimation shows that this approach could be several orders of magnitude faster than the existing software approaches for large sequential designs. Kwang-Ting Cheng, Shi-Yu Huang, Wei-Jin Dai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | ErrorTracer: design error diagnosis based on fault simulation techniquesabstractThis paper addresses the problem of locating error sources in an erroneous combinational or sequential circuit. We use a fault simulation-based technique to approximate each internal signal's correcting power. The correcting power of a particular signal is measured in terms of the signal's correctable set, namely, the maximum set of erroneous input vectors or sequences that can be corrected by resynthesizing the signal. Only the signals that can correct every given erroneous input vector or sequence are considered as a potential error source. Our algorithm offers three major advantages over existing methods. First, unlike symbolic approaches, it is applicable for large circuits. Second, it delivers more accurate results than other simulation-based approaches because it is based on a more stringent condition for identifying potential error sources. Third, it can be generalized to identify multiple errors theoretically. Experimental results on diagnosing combinational and sequential circuits with one and two random errors are presented to show the effectiveness and efficiency of this new approach. Shi-Yu Huang, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1999 | AutoFix: a hybrid tool for automatic logic rectificationabstractWe address the problem of rectifying an erroneous combinational circuit. Based on the symbolic binary decision diagram techniques, we consider the rectification process as a sequence of partial corrections. Each partial correction reduces the size of the input vector set that produces error responses. Compared with the existing approaches, this approach is more general, and thus, suitable for circuits with multiple errors and for the engineering change problem. Also, we derive the necessary and sufficient condition of general single-gate correction to improve the quality of rectification. To handle larger circuits, we develop a hybrid approach that makes use of the information of structural correspondence between specification and implementation. Experiments are performed on a suite of industrial examples as well as the entire set of ISCAS'85 benchmark circuits to demonstrate its effectiveness. Shi-Yu Huang, Kuang-Chien Chen, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | A Hybrid Power Model for RTL Power EstimationabstractWe propose a hybrid power model for estimating the power dissipation of a design at the RT-level. This new model combines the advantages of both RT-level and gate-level approaches. We investigate the relationship between steady-state transition power and overall power dissipation. We observe that, statistically, two input sequences causing similar amount of steady-state transitions will exhibit similar overall power dissipation for an RTL module. Based on this observation, we propose a method to construct a hybrid power model for RTL modules. We further propose a hierarchical power estimation method for estimating the power dissipation of data-path consisting of RTL modules. Experimental results show that, for full-chip power estimation, the estimation time of the technique based on our power models is on average 275 times faster than directly running a commercial transistor-level power simulator, and the errors are less than 6% as compared to the transistor-level power simulation results. Yi-Min Jiang, Shi-Yu Huang, Kwang-Ting Cheng, Deborah C. Wang, ChingYen Ho |
ASP-DAC | 2 |
| 1998 | Fault-Simulation Based Design Error Diagnosis for Sequential CircuitsabstractThis paper addresses the problem of locating design errors in a sequential circuit. For single-error circuits, we consider a signal ƒ as a potential error source only if the circuit can be completely rectified by re-synthesizing ƒ (i.e., changing the function of signal ƒ). In order to handle larger circuits, we do not rely on Binary Decision Diagram. Instead, we search for potential error sources by a modified sequential fault simulation process. The main contributions of this paper are two-fold: (1) we derive the necessary and sufficient condition of whether an erroneous input sequence (i.e., an input sequence producing erroneous responses) can be corrected by changing the function of a particular internal signal; and (2) we propose a modified fault simulation procedure to check this condition. Our approach does not rely on any error model, and thus, is suitable for general types of errors. Furthermore, it can be easily extended to identify multiple errors. Experimental results on ISCAS89 benchmark circuits are presented to demonstrate its capability. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, Juin-Yeu Joseph Lu |
DAC | 1 |
| 1997 | AQUILA: An equivalence verifier for large sequential circuitsabstractIn this paper, we address the problem of verifying the equivalence of two sequential circuits. A hybrid approach that combines the advantages of BDD-based and ATPG-based approaches is introduced. Furthermore, we incorporate a technique called partial justification to explore the sequential similarity between the two circuits under verification to speed up the verification process. Compared with existing approaches, our method is much less vulnerable to the memory explosion problem, and therefore can handle larger designs. The experimental results show that in a few minutes of CPU time, our tool can verify the sequential equivalence of an intensively optimized benchmark circuit with hundreds of flip-flops against its original version. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen |
ASP-DAC | 1 |
| 1997 | Error Tracer: A Fault-Simualtion-Based Approach to Design Error DiagnosisabstractThis paper addresses the problem of locating error sources in an erroneous combinational circuit. We use a fault simulation-based technique to approximate each signal's correcting power. The correcting power of a particular signal is measured in terms of the signal's correctable set, namely, the maximum set of erroneous input vectors that can be corrected by re-synthesizing the signal. Only the signals that can correct every erroneous input vector are considered as a potential error source. Our algorithm offers three major advantages over existing methods. First, unlike symbolic approaches, it is applicable for large circuits. Secondly, it delivers more accurate results than other simulation-based approaches because it is based on a more stringent condition for identifying potential error sources. Thirdly, it can be easily generalized to identify multiple errors. Experimental results on diagnosing circuits with one and two random errors are presented to show the effectiveness and efficiency of this new approach. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, David Ihsin Cheng |
ITC | 1 |
| 1997 | Incremental logic rectificationabstractWe address the problem of rectifying an incorrect combinational circuit against a given specification. Based on the symbolic BDD techniques, we consider the rectification process as a sequence of partial corrections. Each partial correction reduces the size of the input vector set producing error responses. Compared with existing approaches, this approach is more general, and able to handle circuits with multiple errors. We also formulate the necessary and sufficient condition of general single-gate correction to achieve better results for some circuits with a single error. To handle larger circuits, we develop a hybrid approach that makes use of the information of structural correspondence between specification and implementation. Experimental results on industrial examples as well as ISCAS85 benchmark circuits are presented to show the effectiveness of our approach. Shi-Yu Huang, Kuang-Chien Chen, Kwang-Ting Cheng |
VTS | 1 |
| 1996 | Error Correction Based on Verification TechniquesabstractIn this paper, we address the problem of correcting a combinational circuit that is an incorrect implementation of a given specification. Most existing error-correction approaches can only handle circuits with certain types of errors. Here, we propose a general approach that can correct a circuit with multiple errors without assuming any error model. We identify internal equivalent pairs to narrow down the possible error locations using local BDD's with dynamic support. We also employ a technique called back-substitution to correct the circuit incrementally. This approach can also be used to verify circuit equivalence. The experimental results of correcting fully SIS-optimized benchmark circuits with a number of injected errors will be presented. Shi-Yu Huang, Kuang-Chien Chen, Kwang-Ting Cheng |
DAC | 1 |
| 1996 | Compact Vector Generation for Accurate Power SimulationabstractTransistor-level power simulators have been popularly used to esi:imate the power dissipation of a CMOS circuit. These tools strike a good balance between the conventional transistor-level simulators, such as SPICE, and the logic-level power estimators with regard to accuracy and speed. However, it is still too time-consuming to run these tools for large designs. To simulate one-million functional vectors for a 5OK-gate circuit, these power simulators may take months to complete. In this paper, we propose an approach to generate a compact set of vectors that can mimic tb: transition behavior of a much larger set of functional vectors, which is given by the designer or extracted from application programs. This compact set of vectors can then replace the functional vectors for power simulation to reduce the simulation time while still retaining a high degree of accuracy. We present Shi-Yu Huang, Kuang-Chien Chen, Kwang-Ting Cheng, Tien-Chien Lee |
DAC | 1 |
| 1996 | On Verifying the Correctness of Retimed CircuitsabstractWe address the problem of verifying a retimed circuit. After retiming, some latches in a sequential circuit are repositioned to reduce the clock cycle time and thus the behavior of the combinational portion is changed. Here, we present a novel approach to check the correctness of a retimed circuit according to the definition of 3-valued equivalence. This approach is based on our verification framework using sequential ATPG techniques. We further incorporate an algorithm to pre-process the circuits and make the verification process even more efficient. We will present the experimental results of verifying the retimed circuits with hundreds of flip-flops on ISCAS89 benchmark circuits to show its capability. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen |
Great Lakes Symposium on VLSI | 1 |
| 1996 | A novel methodology for transistor-level power estimationabstractTransistor-level power simulators, which are more accurate than logic-level power estimators, are popular to estimate the power dissipation of CMOS circuits. We introduce a method which extends the Monte-Carlo approach for deriving the average power dissipation of a circuit using transistor-level power simulators. To reduce the simulation time, we propose a mixed-level extrapolation technique to speed up the convergence rate of the process, and thereby to achieve a good balance between simulation time and accuracy. Experimental results show that this is a promising method for deriving the accurate power dissipation of a circuit within a reasonable time budget. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, Mike Tien-Chien Lee |
ISLPED | 1 |
| 1996 | An ATPG-Based Framework for Verifying Sequential EquivalenceabstractIn this paper, we address the problem of verifying the equivalence of two sequential circuits. State-of-the-art sequential optimization techniques such as retiming and sequential redundancy removal can handle designs with up to hundreds or even thousands of flip-flops. The BDD-based approaches for equivalence checking can easily run into memory explosion for such designs. With an attempt to handle larger circuits, we modify the test pattern generation techniques for verification. The suggested approach utilizes the efficient backward justification technique popularly used in most sequential ATPG programs. The method explores the structural similarity between circuits under verification, and performs the verification in stages to improve the efficiency. An effective algorithm to identify equivalent hip-hops is presented. This ATPG-based framework is suitable for verifying circuits with or without a reset state. Experimental results of verifying the correctness of circuits after sequential redundancy removal are presented. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, Uwe Gläser |
ITC | 1 |
| 1995 | Fault emulation: a new approach to fault gradingabstractIn this paper, we propose a method of using an FPGA-based emulation system for fault grading. The real-time simulation capability of a hardware emulator could significantly improve the run-time of fault grading, which is one of the most resource-intensive tasks in the design process. A serial fault emulation algorithm is employed and enhanced by two speed-up techniques. First, a set of independent faults can be emulated in parallel. Second, simultaneous injection of multiple dependent faults is also possible by adding extra supporting circuitry. Because the reconfiguration time spent on mapping the numerous faulty circuits into the FPGA boards could be the bottleneck of the whole process, using extra logic for injecting a large number of faults per configuration can reduce the number of reconfigurations, and thus, significantly improve the efficiency. Some modeling issues that are unique in the fault emulation environment are also addressed. The performance estimation indicates that this approach could be several orders of magnitude faster than the existing software approaches for large designs. Kwang-Ting Cheng, Shi-Yu Huang, Wei-Jin Dai |
ICCAD | 2 |