Tong-Yu Hsieh

dblp:51/3754 · DBLP profile ↗
← Back
44ranked-venue papers
23as first author
9since 2021 · last 2026
0000-0002-7954-5569ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 43 · 22 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Late Breaking Results - Confidence-Gap-Driven Functional Test Pattern Generation for Enhancing Functional Safety of CNN Accelerators
Tong-Yu Hsieh, Ching-Hsin Hsu, Wei-Ji Chao
VTS1
2026 A Highly Cost-Effective Online Error Detection and Mitigation Scheme for CNN Hardware Accelerators Based on Approximate PE
abstract
Recent advancements in artificial intelligence have led to the widespread use of convolutional neural networks (CNNs) in various fields. To improve hardware performance, numerous hardware accelerator circuits have been developed. However, the extensive use of processing elements (PEs) in these accelerators raises potential reliability concerns. Traditional modular redundancy techniques, while effective in error mitigation, come with substantial hardware costs. In this study, we introduce a cost-effective scheme designed for efficient on-line error detection and mitigation in CNNs. This scheme is based on innovative designs of approximate PEs. Specifically, we propose two potential designs for these approximate PEs and evaluate their cost-effectiveness. These approximate PEs, used in conjunction with the original PEs, form an Approximate Dual PE Redundancy (ADPR) and Approximate Triple PE Redundancy (ATPR) structure, which is essential for verifying the quality of PE computations. Additionally, we propose a novel error mitigation technique derived from our ADPR and ATPR structure, significantly enhancing the error tolerance of PEs. Notably, our approximate PE designs require only 64% of the area compared with previous designs, while maintaining similar levels of approximation errors.
Wei-Ji Chao, Yen-Chieh Tseng, Tong-Yu Hsieh
ACM Trans. Design Autom. Electr. Syst.3
2025 Application-Aware Early-Exit Fault Classification for Video Decoder Using Miter-Based Analysis
abstract
Traditional structural testing often overlooks the actual impact of hardware faults on system-level applications, especially in error-tolerant components like video decoders. This work presents an application-aware fault classification framework that integrates a miter-based architecture with an early-exit strategy, using object detection accuracy as the evaluation target. By monitoring the accumulation of output errors during video decoding, the framework identifies application-critical faults early-typically within the first six frames of a 300-frame sequence-based on a predefined error threshold. Experimental results show that this approach achieves 100% classification precision while reducing simulation time by over 98%.
Jun-Tsung Wu, Tong-Yu Hsieh
ATS2
2025 Monitor-Like Efficiency with Detector-Level Accuracy: Frontier-Aligned Timing Monitor for AI Accelerators
abstract
Systolic-array AI accelerators operating near threshold voltage face significant timing reliability challenges due to increased PVT sensitivity. While Razor flip-flops offer accurate bit-level detection, their area overhead limits scalability. Existing timing monitors are more efficient but lack granularity and adaptability. This work presents a frontier-aligned timing monitor that enables low-overhead, bit-level visibility. By analyzing post-layout delays in a 7nm systolic array, we identify a MAC unit highly correlated with the global bit-wise delay frontier. A co-located monitor path with tunable delay buffers enables PVT-aware calibration and precise alignment. Experimental results show an average delay error of 3.1% and area overhead as low as 0.1% in large arrays. The proposed design supports scalable, energy-efficient runtime approximation and adaptive voltage/frequency scaling (AVFS), offering a practical solution for fine-grained timing management in modern AI accelerators.
Wei-Ji Chao, Tsung-Chun Chen, Chu-Cheng Chen, Tong-Yu Hsieh
ITC-Asia4
2024 On Accuracy Enhancement of No-Reference Error-Tolerability Testing for Images in Object Detection Applications Based on RGB Channel Characteristics
abstract
Video data is essential for performing computer vision applications. However, errors due to noise or soft errors can be significant, potentially rendering the system invalid. Therefore, detecting such errors is important for these applications. In a real-time video stream, there is no golden reference video data to examine video quality during the functional operation of video processing. To address this issue, the no-reference testing method provides a highly attractive solution. In this work, we propose a no-reference test method that extracts error information from the RGB channels in video data. The proposed method improves test accuracy and precision compared to the previous method designed for human vision when applied to object detection. Experimental results show that more than 90% test accuracy and precision can be achieved. The costs incurred by the proposed method are also discussed. Additionally, based on the results, we further discuss setting different criteria for dynamic and static backgrounds of the target video.
Jun-Tsung Wu, Hideyuki Ichihara, Tomoo Inoue, Tong-Yu Hsieh
ITC-Asia4
2023 Cost-Effective Error-Mitigation for High Memory Error Rate of DNN: A Case Study on YOLOv4
abstract
In a Deep Neural Network (DNN) computing platform, memory is an essential component. Unfortunately, memory errors may occur due to various factors. The memory error rate may even increase significantly when low-power technologies are used. Therefore, it is crucial to cost-effectively protect memory against errors. However, most previous memory protection work for DNN has limited capability to deal with the case of high memory error rate. In this work, we first conduct a detailed study on the inherent error-tolerability of a DNN for memory errors with various error rates. The YOLOv4 DNN model is employed as a case study, and various memory error models are considered. In particular, we also investigate the effectiveness limitation of the previous error mitigation methods. Based on these analyses, we propose a novel protection method where memory errors with high error rate can be tolerated. The experimental results show that our method can guarantee only 1% DNN accuracy degradation even when the error rate is as high as 0.1%, which is a breakthrough in the literature. Moreover, the cost overhead of our method is still comparable to previous methods.
Wei-Ji Chao, Tong-Yu Hsieh
ITC-Asia2
2023 On Development of Reliable Machine Learning Systems Based on Machine Error Tolerance of Input Images
abstract
With the rapid development of machine learning technologies, more and more practical applications arise. Representative machine learning techniques that receive much attention include object detection and image classification, which can be applied to many applications, such as self-driving cars, traffic flow calculation, and detection of product defects in factories. In this article, we investigate tolerability of errors for input images in machine learning systems and develop a generic reliability enhancement methodology. This work is based on our preliminary studies on image classification, but we put major focuses in object detection applications and the comprehensive comparisons to the prior studies. The first one error-tolerability test method to support reliability enhancement of object detection applications is then proposed based on careful error-tolerability examination of input images. The experimental results show that the test accuracy of this method can achieve 93.06%, which is the state-of-the-art. One special advantage of the proposed method is that unlike the previous error-tolerance methods in the literature, no golden reference data are required for acceptability determination by the proposed method. Hence, on-line testing can be supported. Our method is also implemented and validated in hardware. The results show that the hardware performance is up to 192 frames per second (FPS), which can thus also support real-time operations.
Tong-Yu Hsieh, Chun-Chao Cheng, Wei-Ji Chao, Pin-Xuan Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 On No-Reference Error Detection of an Image Stitching System Based on Error-Tolerance
abstract
Image stitching technology can stitch images (video frames) from multiple cameras into a single panoramic image, allowing a single image to cover a larger field of view. Image(video) stitching technology has also been widely used in real-time video applications such as online conference, VR or even Advanced Driver Assistance Systems (ADAS). Reliability of this technology is therefore of great importance. In this work, for the first time, we address the issue of error detection of an image stitching system. We show that there are inherent error-tolerability in this system, and the detection focus should be the unacceptable errors that would result in poor stitching results. In this work we also investigate the acceptability evaluation of the stitching results. In particular, a no-reference error detection technique is proposed so that the detection of unacceptable errors can be achieved in an online manner without golden data for comparison. Our experimental results show that the detection (acceptability classification) accuracy of the proposed error detection technique achieves 98.2%. In addition, the incurred performance overhead of our technique is ignorable. There is only 0.4% increase on the stitching execution time.
Tong-Yu Hsieh, Pao-Wei Tsui, Jun-Tsung Wu
ATS1
2021 Concurrent Test of Reconfigurable Scan Networks for Self-Aware Systems
abstract
Self-aware and safety-critical hardware/software systems rely on a variety of embedded instruments, sensors, monitors and design-for-test circuitry to check the system integrity. The access to these internal instruments is supported by standards commonly called iJTAG and employs so called reconfigurable scan networks (RSNs), which are more and more used at runtime, too. They collect periodically and also concurrently the information on the circuit's health state and deliver it to some dependability management unit. The integrity of RSNs is essential for the dependability of self-aware systems and can be ensured by a combination of periodic and concurrent test methods of the RSN itself. The paper at hand presents the first concurrent online test method for RSNs by adding a brief integrity test to each access operation. The presented scheme includes a hardware extension of negligible size, supports offline test, diagnosis and post-silicon validation as well, and is further referred as ROSTI: RSN Online/Offline Self-Test Infrastructure. It exploits the original RSN control signals and does not require any modification of the underlying RSN. The hardware costs are independent of the size of the RSN, and ROSTI is flexible for generating different test sequences for different types of faults. The experimental results validate these characteristics and show that ROSTI is highly scalable.
Chih-Hao Wang, Natalia Lylina, Ahmed Atteya, Tong-Yu Hsieh, Hans-Joachim Wunderlich
IOLTS4
2020 A Self-Detection and Self-Repair Methodology for Reliable Speech Recognition Considering AWGN Noises
abstract
Speech recognition is expected to be widely used in many IoT applications such as smart home and wearable devices. For successful speech recognition, the integrity of voice signals is critical. However, during transmission or storage of voice signals, there are usually noises, which may distort the voice signals and thus invalidate the speech recognition process, especially when the noises are significant. In this paper we address the issue of detecting and repairing noisy audio signals. In particular, self-detection and self-repair approaches are investigated such that on-line detection and repair is possible. In this work the AWGN noise model that is widely used in the communication area is employed to model noises. The speech recognition application is targeted such that as many correct words can be recognized as possible for the repaired audio. To the best of our knowledge, this work is the first one that addresses the audio self-detection and self-repair problem considering AWGN noises together with the speech recognition application. We propose a detection and repair methodology that can accurately detect noisy audio signals, and utilize the noise-free ones to predict the expected values of noisy audio signals for repairing. No additional golden signals are required. Compared with the related method in the literature, the proposed methodology has much higher repair effectiveness. Our experimental results show that 45.93% enhancement on the speech recognition rate can be achieved on average. As a comparison, for the previous method, the recognition rate can be enhanced by only 11.68%.
Tong-Yu Hsieh
ITC-Asia1
2020 On Enhancing Error-Tolerability of Videos via Re-Encoding with Adaptive I-Frame Insertion
abstract
Videos are expected to be widely used in IoT or AI applications. The quality of videos are thus crucial for the success of these technologies. However, noises during transmission of videos, or aging of video processing or storage related circuits may significantly degrade the quality of videos, making the video become unacceptable. In this work we investigate the issue of video quality (error-tolerability) enhancement. Different from the previous work that mainly focuses on noisy videos, this work considers erroneous videos generated by faulty circuits. Single stuck-at faults are injected to an H.264 decoding circuit for generating erroneous videos. We find that adaptive insertion of I-frames is an attractive solution. This method first examines the quality of the current videos and accordingly suggests re-encoding of future videos with a proper number of additional I-frames. An error-tolerability enhancement flow for videos is proposed, which integrates video quality grading and determination of the suggested number of I-frames to be inserted. Limitations and implementation of this flow are also discussed. Experimental results on a total of 81,412 erroneous videos show that more than 90% of unacceptable erroneous videos whose quality is within a specified range become acceptable by applying the proposed flow.
Tong-Yu Hsieh, Chen-Chia Chung, Jun-Tsung Wu
ITC-Asia1
2020 On Classification of Acceptable Images for Reliable Artificial Intelligence Systems: A Case Study on Pedestrian Detection
abstract
Images are essential data for many artificial intelligence (AI) systems such as pedestrian detection. However, image processing circuits or image storage devices may produce erroneous image data due to aging or radiation. In this paper we will show that there actually exists much tolerability in image errors. Moreover, for AI systems we find that the tolerability is even larger. This finding provides an attractive reliability enhancement solution by classifying and filtering acceptable images. This solution allows acceptable images to still go to the AI inference process, while unacceptable images are discarded, together with warning signals activated. In this work, we first evaluate and compare error tolerability of images from human and machine perspectives. Then a number of possible test methods to support machine based error-tolerance are discussed and compared in terms of their acceptability classification accuracy and computation cost. In particular, these methods should not need golden (error-free) images as the comparison basis. This greatly facilitates developing a low-cost on-line test architecture to enable a real-time reliability enhancement solution. Our experimental results show that when applying the suggested test method to pedestrian detection, 93.48% of the erroneous images can be correctly classified. The results also show that adopting machine-based error-tolerance can extend MTTF (Mean Time To Failure) of the pedestrian detection system up to additional 88.7%, while human vision based error-tolerance can extend only additional 35.1%.
Tong-Yu Hsieh, Pin-Xuan Wu, Chun-Chao Cheng
VTS1
2020 An Implication-based Test Scheme for Both Diagnosis and Concurrent Error Detection Applications
abstract
This article describes a diagnosis-aware hybrid concurrent error detection ( DAH-CED ) scheme that can facilitate both off-line and on-line test applications. By using the proposed scheme, not only the probability of detecting errors (on-line) but also the diagnosability of the target circuit (off-line) can be significantly enhanced. The proposed scheme combines the implication-based method with the parity check method. In particular, novel algorithms are developed to identify specific implications for enhancing the diagnosability for the modeled faults proactively. Furthermore, a reduction algorithm is also presented to minimize the number of the employed implications, while no loss on probability of detecting errors and diagnosability is also guaranteed. To the best of our knowledge, this issue is not addressed in the literature. To validate the proposed scheme, not only stuck-at faults but also transition faults are considered to simulate the timing-related errors. The experimental results on nine ITC’99 benchmark circuits show that the diagnosability for stuck-at (transition) faults is enhanced by 6.88% (7.78%) by applying the proposed scheme. As for the probability of detecting errors, 97.73% (97.10%) is achieved for errors caused by stuck-at (transition) faults. Moreover, only 3.11% of implications are needed.
Chih-Hao Wang, Tong-Yu Hsieh
ACM Trans. Design Autom. Electr. Syst.2
2019 A Delay-Aware Implementation Scheme for Cost-Effective Implication-Based Concurrent Error Detection
abstract
Implications have been shown to be beneficial for both concurrent error detection and diagnosis. To reduce the incurred hardware cost, one critical issue is selection of a minimum number of appropriate implications. Although the previous work developed several implication selection algorithms, the critical path delay may still be high. This is because the factors related to the critical path delay have not been well studied and considered during implication selection. In this paper, we investigate these factors and develop a new delay-aware implication selection algorithm. A buffer insertion algorithm is also developed such that the minimum number of buffers are inserted to further reduce the delay. This algorithm is integrated with the implication selection algorithm as a delay-aware implementation scheme. Experimental results on 18 ISCAS'85 and ITC'99 benchmark circuits show that 29.02% delay overhead reduction is achieved on average with only additional 0.34% implications selected.
Tong-Yu Hsieh, Kuang-Chun Lin, Hsin-Hsien Lin
ITC-Asia1
2018 On no-reference on-line error-tolerability testing for videos
abstract
In this paper we investigate how to achieve on-line error-tolerability testing on videos. In particular, a no-reference manner is considered, which means that no reference videos are needed for comparison with the videos under test. As a result, the hardware that is usually needed in conventional on-line test methods for generating reference data can be totally eliminated. This greatly reduces implementation complexity of on-line test procedures. We show that by well exploiting some simple attributes, the acceptability for 81,412 various erroneous videos can be accurately determined with more than 90% accuracy. As a comparison, the related previous work can only achieve about 80% accuracy. In addition, our attribute acquirement process requires only 33% of the computation time for the previous work.
Tong-Yu Hsieh, Shang-En Chan, Chi-Hsuan Ho
ETS1
2018 A No-Reference Error-Tolerability Test Methodology for Image Processing Applications
abstract
Error-tolerance is a notion that can extend the lifetime of a system, especially for multimedia applications. In this paper we present a no-reference error-tolerability test methodology for image processing applications. No reference images are needed in this methodology for comparison with the images under test. As a result, the hardware that is usually needed in conventional on-line test methods for generating reference data can be totally eliminated. This greatly facilitates developing on-line error-tolerability test procedures for reliability concerns. Compared with the previous error-tolerability test work in the literature, this work is the first one that can test error-tolerability of images by using a no-reference manner. In this work we develop a particular attribute and the corresponding acquiring method that can effectively quantify acceptability of errors. In particular, the proposed methodology can be adaptively re-configured according to the characteristics of the target image so as to achieve high test accuracy. We also employ 126,894 images to generally evaluate the effectiveness of the proposed test methodology. The experimental results show that up to 93.39% test accuracy is achieved on average by the proposed methodology.
Tong-Yu Hsieh, Chao-Ru Chen
ITC-Asia1
2018 Error Indication Signal Collapsing for Implication-Based Concurrent Error Detection
abstract
Implication-based concurrent error detection (CED) has been shown to have promising performance for on-line testing. However, many error indication signals may be required for this CED method, and thus incur much additional interconnection. This would result in not only complicated error checking circuits, but also a large compactor design to process the error indication signals. Both would incur high area overhead. In this paper, we present a collapsing technique that can significantly reduce the total number of required error indication signals for implications. This issue has never been addressed in the literature. We find that equivalence and dominance relationships exist between error indication signals, which are quite helpful for signal reduction. Therefore we develop an efficient algorithm to first identify these relationships, and then make good use of them to merge error indication signals without sacrificing the probability of detecting errors. We also employ 19 ISCAS'85 and ITC'99 benchmark circuits to evaluate the effectiveness of the proposed technique. The results show that 48.48% of error indication signals are reduced by our technique on average. This also leads to 39.23% and 34.52% averaged area overhead reduction to the error checking circuit and the compactor design, respectively.
Chih-Hao Wang, Chi-Hsuan Ho, Tong-Yu Hsieh
ITC-Asia3
2018 Structural Variance-Based Error-Tolerability Test Method for Image Processing Applications
abstract
Image processing circuits are expected to play an important role in Internet of Things electronic systems. For this type of circuits, errors (e.g., those due to wear out) might not be perceptible to us due to our insensitivity to small variances in images or colors. In the case that errors are acceptable (imperceptible), the systems are very likely to still be functional but with only minor performance degradation. The lifetime of the systems thus can be extended. Although there have been several attributes proposed in the literature to test acceptability of errors, these attributes do not consider human beings' sensitivity to structural variances of erroneous images. This may result in acceptability misclassification of errors. In this paper, we take this into consideration and propose a new test method to more accurately evaluate acceptability of errors. Compared with the available attributes that can also extract the structural information of images, our method has much lower computation complexity, which is advantageous for shortening test time. This also makes our method hardware efficient. It is shown that the incurred area overhead of our method is low. The proposed method is to be applied periodically in-field to examine the reliability of the target circuit. We thus consider errors that may occur during in-field use. Single/multiple stuck-at faults and noises are considered, and a total of 161 460 erroneous images are generated. The experimental results on these diverse images show that on average the test accuracy of our method is more than 99%.
Tong-Yu Hsieh, Yi-Han Peng, Kuan-Chih Cheng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 On Probability of Detection Lossless Concurrent Error Detection Based on Implications
abstract
In recent years, a new concurrent error detection method by using invariant relationships inside a circuit, called implications, has been proposed. Algorithms have also been developed to reduce the total number of required implications so as to minimize the incurred area overhead due to implication checking logic. This implication reduction process, however, would result in degradation on the probability of error detection (Pdetection) of the method. In this paper, we analyze the impact of this issue mathematically together with illustration by a real case study. Our analytical results show that just one percent degradation on Pdetectionwould result in millions more errors being undetected per second and thereby significant loss on reliability of the target circuit. To address this issue, we develop a new implication reduction algorithm that guarantees no loss on Pdetection. In our algorithm, the detectability of errors for each candidate implication is carefully evaluated. The evaluation results are then utilized to select the most efficient candidates for detecting all the detectable errors. We also analyze the computation and memory complexity of the proposed algorithm. The experimental results on 28 representative benchmark circuits from ISCAS'85, ISCAS'89, and ITC'99 show that the implication reduction rate of our method (92.59%) is close to that of the previous work (95.8%). Only a small number of additional implications need to be selected to guarantee no loss on Pdetection.
Chih-Hao Wang, Tong-Yu Hsieh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Error-Tolerability Evaluation and Test for Images in Face Detection Applications
abstract
Face detection (i.e., checking if any face can be detected in a considered image) is expected to be one critical technology for IoT (Internet of Things). However, the images to be detected are likely to be erroneous when the image capture/storage/processing circuits are aged. Fortunately, for the face detection technology only a few details of an image are required for processing so as to make the detection process efficient. Thus high tolerability for image errors exists as long as the structure of the face is not destroyed too much. Well exploiting this feature will be very helpful to extend the lifetime of the face detection based system. However, no work in the literature evaluates the error-tolerability of such application. Most of the related error-tolerance work takes advantages of human beings' insensitivity to minor vibrations in multimedia signals. In this work we will show the error tolerability of images for the machine-based face detection application is even much larger than that for human based applications. One associated critical issue is therefore how to test if an erroneous image is still acceptable for face detection. Our analysis results show that typical image quality assessment methods would result in some misclassification. This motivates us to develop an efficient test method. The proposed method captures the incurred structural variances in an erroneous image, and evaluates if this variance is significant. The implementation of the proposed method is simple, which requires only addition, counting and comparison operations. Our experimental results on 3,438 erroneous benchmark images show that 99.57% test accuracy is achievable by the developed method.
Tong-Yu Hsieh, Tai-Ang Cheng, Chao-Ru Chen
ATS1
2017 A hybrid concurrent error detection scheme for simultaneous improvement on probability of detection and diagnosability
abstract
In this work we propose a hybrid concurrent error detection (CED) scheme that combines the implication-based method with the parity check method. The parity check method is easy to implement and has high probability of detecting errors, while the implication-based method has high flexibility to be easily integrated with other CED methods for improidng the probability of detecting errors. In addition, fault diagnosis capability can even be enabled by integrating implications. We show that by combining these two methods, not only the probability of detecting errors can be significantly increased, but also the diagnosability of the target circuit can be effectively enhanced. A systematical flow is developed to add the required checking logic of implications into the most appropriate locations in the target circuit for detecting undetected errors by the parity check method. This minimizes the total number of the employed implications and thus the incurred area overhead. The experimental results show that 89.08% of area overhead is required yet the probability of detecting errors is improved from 88.32% to 97.73% by the proposed hybrid scheme. The results also show that the achievable diagnosability of our scheme enhances from 91.79% to 96.36% on average. More than 200 equivalent faults also become distinguishable in our scheme.
Chih-Hao Wang, Tong-Yu Hsieh
ITC-Asia2
2017 Cost-Effective Enhancement on Both Yield and Reliability for Cache Designs Based on Performance Degradation Tolerance
abstract
Guaranteeing functional correctness of cache memories is crucial for computer designs. In the literature, there have been several works addressing this issue. However, fault tolerability of these methods may be limited. In this paper, we present a new cache architecture that has flexible tolerability. Moreover, by using the proposed architecture, both yield and reliability of the cache can be enhanced simultaneously. In our cache, a particular type of cache blocks called tolerable block is further identified among the faulty ones. Such blocks can still be used during cache access in our architecture, while accessing to intolerable blocks will result in additional cache misses, and therefore performance degradation. The number of tolerable cache blocks is thus critical for the achievable yield and reliability enhancement, as well as the incurred cost on performance. In this work, error correcting code (ECC) methods are employed to increase the number of tolerable blocks. In particular, we propose to embed the required check bits in one of the cache ways. Analysis results show that this embedding method only incurs minor performance degradation, while the incurred area overhead due to ECC can thus be significantly reduced from 5.92% to only 0.92%. General applicability of the embedding method to ordinary ECC methods is also investigated. Experimental results show that the performance degradation can be reduced from 16% to only 1.53% by using the proposed cache. This leads to great tolerability improvement, and thus the yield and reliability are enhanced very significantly when compared with the previous work.
Tong-Yu Hsieh, Tsung-Liang Chih, Mei-Jung Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2016 A Performance Degradation Tolerable Cache Design by Exploiting Memory Hierarchies
abstract
Performance degradation tolerance (PDT) has been shown to be able to effectively improve the yield, reliability, and lifetime of an electronic product. The focus of PDT is on the particular performance degrading faults (pdef) that only incur some performance degradation of a system without inducing any computation errors. The basic idea is that as long as the defective chips containing only the pdef can provide acceptable performance for some applications, they may still be marketable. Critical issues of PDT to be addressed include the portion of the pdef in a faulty chip and their induced performance degradation. For a typical cache design, most of the possible faults are not pdef. In this brief, we propose a cache redesign method, called PDT cache, where all functional faults in the data-storage cells of a cache (major part of the cache) can be transformed into pdef. By transforming this large number of faults into pdef, a faulty cache becomes much more likely to be still marketable. The proposed design exploits the existing hardware resources and the inherent error resilience scheme to reduce the incurred hardware overhead. The logic synthesis results show that the incurred hardware overhead is only 6.29% for a 32-kB cache. We also evaluate the induced performance degradation under various fault densities using the CPU2000 and CPU2006 benchmark programs. The results show that for a 32-kB cache design, when the fault density is <;1%, only 0.31% performance degradation is incurred. In addition, the scalability of the PDT cache is also evaluated. The results show that a smaller hardware overhead is required for a larger cache, and the performance degradation is independent of the cache associativity and can even be smaller for a larger cache under a given fault density.
Tong-Yu Hsieh, Chih-Hao Wang, Tsung-Liang Chih, Ya-Hsiu Chi
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Filtering-based error-tolerability evaluation of image processing circuits
abstract
For some systems errors can be regarded as being acceptable as long as their significance is low enough. Image processing circuits are one such example due to human being's insensitivity to minor errors in colors. Significant errors usually destroy the structure of an image, and thus appear to be perceptible. This also makes larger changes to the frequency feature of the image. By examining the degree at which the frequency is varied, the acceptability of errors can be determined. In this work we propose a filtering-based test method that can quantify the frequency variance incurred by errors. According to the obtained variance value, the acceptability of an image can be determined by comparing the value with user-specified thresholds. The experimental results on a large number of erroneous benchmark images show that the proposed method can accurately differentiate unacceptable images from acceptable ones. The implementation of the proposed method is simple, and thus can facilitate implementation of a BIST (Built-In Self-Test) circuitry for efficient product grading, as well as in-field reliability determination and enhancement. This is useful when the target circuit is employed in some critical applications such as automotive or medical electronic systems.
Tong-Yu Hsieh, Yi-Han Peng
IOLTS1
2015 Performance Degradation Tolerance Analysis and Design for Effective Yield Enhancement
Tong-Yu Hsieh, Chih-Hao Wang, Chun-Wei Kuo, Shu-Yu Huang, Tsung-Liang Chih
J. Electron. Test.1
2014 Output-bit selection with X-avoidance using multiple counters for test-response compaction
abstract
Output-bit selection is a recently proposed test-response compaction approach that can effectively deal with aliasing, unknown-value, and low-diagnosis problems. This approach has been implemented using a single counter and a multiplexer without considering unknown values. Also, such an implementation may require the application of a pattern multiple times in order to observe all selected responses. In this paper, we present a multiple-counter-based architecture with a new selection algorithm that can avoid most unknown-values yet achieve high compaction ratio. The remaining small number of unknowns can then be dealt with using some simple masking logic. Experiments on IWLS'05 circuits show that even with 16% unknown responses, all unknown values can be handled with 88.92%~93.21% response-volume reduction still achieved and only a moderate increase in test-application time.
Wei-Cheng Lien, Kuen-Jong Lee, Krishnendu Chakrabarty, Tong-Yu Hsieh
ETS4
2014 Efficient Error-Tolerability Testing on Image Processing Circuits Based on Equivalent Error Rate Transformation
Tong-Yu Hsieh, Yi-Han Peng, Kuan-Hsien Li
J. Electron. Test.1
2014 Efficient LFSR Reseeding Based on Internal-Response Feedback
Wei-Cheng Lien, Kuen-Jong Lee, Tong-Yu Hsieh, Krishnendu Chakrabarty
J. Electron. Test.3
2013 An Efficient Test Methodology for Image Processing Applications Based on Error-Tolerance
abstract
Error-tolerance is a novel notion that can improve yield of VLSI circuits by identifying defective yet acceptable chips. In this paper we address two key issues related to error-tolerance, namely acceptable threshold determination and acceptability evaluation, focusing on image processing applications. We first carefully investigate the acceptability thresholds of images in terms of error rate and error significance. The investigation results show that due to human beings' various insensitivities to images with different frequencies, appropriate thresholds should be determined depending on the frequency characteristics of test images. Based on the determined thresholds we propose an efficient test methodology to help test engineers easily and quickly examine the acceptability of a circuit under test. The experimental results for a large number of erroneous benchmark images show that the proposed test methodology is as effective as the exhaustive test method. Moreover, our methodology requires much less test time and storage space. The achievable reduction ratio can be more than 99%.
Tong-Yu Hsieh, Yi-Han Peng, Chia-Chi Ku
Asian Test Symposium1
2013 A New LFSR Reseeding Scheme via Internal Response Feedback
abstract
Reseeding techniques have been adopted in BIST to enhance fault detect ability and shorten test application time for integrated circuits. In order to achieve complete fault coverage, previous reseeding methods often need large storage space to store all required seeds. In this paper, we propose a new LFSR reseeding technique that employs the internal net responses of the circuit itself as the control signals to change the states of the LFSR. A novel test architecture containing a net selection logic module and an LFSR with some inversion logic is presented that can generate all required seeds on-chip in real time without any external or internal storage requirement. Experimental results on ISCAS benchmark circuits show that the presented technique can achieve 100% stuck-at fault coverage in a short test time by using only 0.23-2.36% of internal nets for reseeding control.
Wei-Cheng Lien, Kuen-Jong Lee, Tong-Yu Hsieh, Krishnendu Chakrabarty
Asian Test Symposium3
2013 An Efficient On-Chip Test Generation Scheme Based on Programmable and Multiple Twisted-Ring Counters
abstract
Twisted-ring-counters (TRCs) have been used as built-in test pattern generators for high-performance circuits due to their small area overhead, low performance impact and simple control circuitry. However, previous work based on a single, fixed-order TRC often requires long test time to achieve high fault coverage and large storage space to store required control data and TRC seeds. In this paper, a novel programmable multiple-TRC-based on-chip test generation scheme is proposed to minimize both the required test time and test data volume. The scan path of a circuit under test is divided into multiple equal-length scan segments, each converted to a small-size TRC controlled by a programmable control logic unit. An efficient algorithm to determine the required seeds and the control vectors is developed. Experimental results on ISCAS'89, ITC'99 and IWLS'05 benchmark circuits show that, on average, the proposed scheme using only a single programmable TRC design can achieve 35.58%-98.73% reductions on the number of test application cycles with smaller storage data volume compared with previous work. When using more programmable TRC designs, 83.60%-99.59% reductions can be achieved with only slight increase on test data volume.
Wei-Cheng Lien, Kuen-Jong Lee, Tong-Yu Hsieh, Wee-Lung Ang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Counter-Based Output Selection for Test Response Compaction
abstract
Output selection is a recently proposed test response compaction method, where only a subset of output response bits is selected for observation. It can achieve zero aliasing, full X-tolerance, and high diagnosability. One critical issue for output selection is how to implement the selection hardware. In this paper, we present a counter-based output selection scheme that employs only a counter and a multiplexer, hence involving very small area overhead and simple test control. The proposed scheme is ATPG-independent and thus can easily be incorporated into a typical design flow. Two efficient output selection algorithms are presented to determine the desired output responses, one using a single counter operation for simpler test control and the other using more counter operations for achieving a better test-response reduction ratio. Experimental results show that for stuck-at faults in large ISCAS'89 and ITC'99 benchmark circuits, 48%~90% reduction ratios on test responses can be achieved with only one counter and one multiplexer employed. Even better results, i.e., 76%~95% reductions, can be obtained for transition faults. It is also shown that the diagnostic resolution of this method is almost the same as that achieved by observing all output responses.
Wei-Cheng Lien, Kuen-Jong Lee, Tong-Yu Hsieh, Krishnendu Chakrabarty, Yu-Hua Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2012 A Test-Per-Clock LFSR Reseeding Algorithm for Concurrent Reduction on Test Sequence Length and Test Data Volume
abstract
This paper proposes a new test-per-clock BIST method that attempts to minimize the test sequence length and the test data volume simultaneously. An efficient LFSR reseeding algorithm is developed by which each determined seed together with its derived patterns can detect the maximum number of so far undetected faults. During the seed determination process an adaptive X-filling process is first employed to generate a set of candidate patterns for pattern embedding. The process then derives a seed solution that can embed multiple candidate patterns at one time so as to minimize the number of seeds. To shorten the test sequence, the pattern embedding process begins with a small initial set of pseudo-random patterns and will incrementally add more patterns only when necessary. Experimental results show that compared with the previous test-per-clock techniques based on the LFSR- and twisted-ring-counter-reseeding methods, our method can reduce the test sequence length by over 60% with generally smaller numbers of storage bits. When compared with the mapping-logic-based BIST methods, our method can reduce the test sequence length by over 50% with a comparable area overhead.
Wei-Cheng Lien, Kuen-Jong Lee, Tong-Yu Hsieh
Asian Test Symposium3
2012 Accumulator-based output selection for test response compaction
abstract
Output selection is a recently proposed test response compaction method, where only a subset of output response bits is selected for observation. It can achieve zero aliasing, full X-tolerance, and high diagnosability. We propose an output selection scheme for multiple scan designs, which employs only accumulators and multiplexers, and thus involves small area overhead and simple test control. An efficient selection procedure is presented to determine a minimal test set and the corresponding output bits to select for complete fault coverage. Experimental results show that when only one accumulator and one multiplexer are employed, 100% single stuck-at fault coverage for ISCAS'89 (ITC'99) circuits can be achieved by observing only 9.84% (8.19%) of the test response bits with only 1.86% (1.18%) area overhead.
Wei-Cheng Lien, Kuen-Jong Lee, Tong-Yu Hsieh, Shih-Shiun Chien, Krishnendu Chakrabarty
ISCAS3
2012 Efficient Overdetection Elimination of Acceptable Faults for Yield Improvement
abstract
Acceptable faults in a circuit under test (CUT) refer to those faults that have no or only minor impacts on the performance of the CUT. A circuit with an acceptable fault may be marketable for some specific applications. Therefore, by carefully dealing with these faults during testing, significant yield improvement can be achieved. Previous studies have shown that the patterns generated by a conventional automatic test pattern generation procedure to detect all unacceptable faults also detect many acceptable ones, resulting in a severe loss on achievable yield improvement. In this paper, we present a novel test methodology called multiple test set detection (MTSD) to totally eliminate this overdetection problem. A basic test set generation method is first presented, which depicts a fundamental scheme to generate appropriate test sets for MTSD. We then describe an enhanced test generation method that can significantly reduce the total number of test patterns. Solid theoretical derivations are provided to validate the effectiveness of the proposed methods. Experimental results show that in general an 80%-99% reduction in the number of test patterns can be achieved compared with previous work addressing this problem.
Kuen-Jong Lee, Tong-Yu Hsieh, Melvin A. Breuer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2011 An Error-Tolerance-Based Test Methodology to Support Product Grading for Yield Enhancement
abstract
This paper presents a novel error-tolerance-based test methodology to grade defective chips according to their degree of acceptability so as to improve the effective yield of chips. We employ error rate as the attribute of error-tolerance to determine acceptability. We show that the number of test patterns that need to be applied to a circuit under test in estimating the circuit's error rate is highly dependent on how close the circuit's actual error rate is to the given grading thresholds. An iterative and adaptive error rate estimation technique is developed by which an appropriate number of test patterns can be efficiently determined and the circuit can be immediately classified into appropriate grades to fit various application requirements. Experimental results show that: 1) only a few iterations are required to classify a circuit, and 2) the total number of test patterns used is in general independent of the circuit size. Both of these observations imply that these techniques are applicable to large circuits.
Tong-Yu Hsieh, Kuen-Jong Lee, Melvin A. Breuer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 Test Response Compaction via Output Bit Selection
abstract
The conventional output compaction methods based on XOR-networks and/or linear feedback shift registers may suffer from the problems of aliasing, unknown-values, and/or poor diagnosability. In this paper, we present an alternative method called the output-bit-selection method to address the test compaction problem. By observing only a subset of output responses, this method can effectively deal with all the above-mentioned problems. Efficient algorithms that can identify near optimum subsets of output bits to cover all detectable faults in very large circuits are developed. Experimental results show that less than 10% of the output response bits of an already very compact test set are enough to achieve 100% single stuck-at fault coverage for most ISCAS benchmark circuits. Even better results are obtained for ITC 99 benchmark circuits as less than 3% of output bits are enough to cover all stuck-at faults in these circuits. The increase ratio of selected bits to cover other types of faults is shown to be quite small if these faults are taken into account during automatic test pattern generation. Furthermore, the diagnosis resolution of this method is almost the same as that achieved by observing all output response bits.
Kuen-Jong Lee, Wei-Cheng Lien, Tong-Yu Hsieh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 On-Chip SOC Test Platform Design Based on IEEE 1500 Standard
abstract
IEEE 1500 Standard defines a standard test interface for embedded cores of a system-on-a-chip (SOC) to simplify the test problems. In this paper we present a systematic method to employ this standard in a SOC test platform so as to carry out on-chip at-speed testing for embedded SOC cores without using expensive external automatic test equipment. The cores that can be handled include scan-based logic cores, BIST-based memory cores, BIST-based mixed-signal devices, and hierarchical cores. All required test control signals for these cores can be generated on-chip by a single centralized test access mechanism (TAM) controller. These control signals along with test data formatted in a single buffer are transferred to the cores via a dedicated test bus, which facilitates parallel core testing. A number of design techniques, including on-chip comparison, direct memory access, hierarchical core test architecture, and hierarchical test bus design, are also employed to enhance the efficiency of the test platform. A sample SOC equipped with the test platform has been designed. Experimental results on both FPGA prototyping and real chip implementation confirm that the test platform can efficiently execute all test procedures and effectively identify potential defect(s) in the target circuit(s).
Kuen-Jong Lee, Tong-Yu Hsieh, Chin-Yao Chang, Yu-Ting Hong, Wen-Cheng Huang
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Tolerance of performance degrading faults for effective yield improvement
abstract
To provide a new avenue for improving yield for nano-scale fabrication processes, we introduce a new notion: performance degrading faults (pdef). A fault is said to be a pdef if it cannot cause a functional error at system outputs but may result in system performance degradation. In a processor, a fault is a pdef if it causes no error in the execution of user programs but may reduce performance, e.g., decrease the number of instructions executed per cycle. By identifying faulty chips that contain pdef's that degrade performance within some limits and binning these chips based on the their resulting instruction throughput, effective yield can be improved in a radically new manner that is completely different from the current practice of performance binning on clock frequency. To illustrate the potential benefits of this notion, we analyze the faults in the branch prediction unit of a processor. Experimental results show that every stuck-at fault in this unit is a pdef. Furthermore, 97% of these faults induce almost no performance degradation.
Tong-Yu Hsieh, Melvin A. Breuer, Murali Annavaram, Sandeep Gupta 0001, Kuen-Jong Lee
ITC1
2008 An Error Rate Based Test Methodology to Support Error-Tolerance
abstract
Error-tolerance is an innovative technique to address the problem of low yields in nanometer very large scale integrated (VLSI) circuitry, which is the backbone of the system-on-a-chip (SOC) revolution. The basic principle of error-tolerance is that some chips may occasionally produce erroneous outputs, but still provide acceptable performance when used in certain systems. Using these chips in such systems results in an increase in effective yield. In this paper, a fault-oriented test methodology is presented for classifying whether or not a chip is acceptable based on error rate estimation. A sampling method is proposed to estimate error rate associated with each possible fault in the target circuit. According to this information, an approach is developed to identify a list of faults that are acceptable with respect to a specified upper bound on expected error rates of acceptable chips. Furthermore, a test pattern selection method, and an output masking technique are presented to identify tests which detect all of the unacceptable faults, and as few acceptable faults as possible, so as to maximize the effective yield. Experimental results indicate the high effectiveness of the proposed error rate estimation method, and the degree to which yield can be enhanced.
Tong-Yu Hsieh, Kuen-Jong Lee, Melvin A. Breuer
IEEE Trans. Reliab.1
2007 Test Efficiency Analysis and Improvement of SOC Test Platforms
abstract
Employing a test platform in an SOC design has been shown to be an effective method for SOC testing. However the test efficiency problem of a test platform has not been addressed. In this paper, we formally analyze the test efficiency of test platforms and seek for its optimization. We formulate the required numbers of test cycles for test platforms implemented with different test structures and/or executed with different test procedures. It is shown that up to 24X test time difference for platforms with different test structures/procedures is possible. Based on the derived formula, an appropriate test platform that can achieve best test efficiency with minimal area overhead can be determined.
Tong-Yu Hsieh, Kuen-Jong Lee, Jian-Jhih You
ATS1
2007 Reduction of detected acceptable faults for yield improvement via error-tolerance
Tong-Yu Hsieh, Kuen-Jong Lee, Melvin A. Breuer
DATE1
2006 An Error-Oriented Test Methodology to Improve Yield with Error-Tolerance
abstract
The main objective of error-tolerance is to increase the effective yield of a process by identifying defective but acceptable chips. In this paper, we propose an error-oriented test methodology to support error-tolerance in scan-based digital circuits. Error-rates of defective chips are first estimated and then compared with application-specific acceptable values of error-rates to determine the suitability of each chip. A theoretical basis to estimate error-rates of chips with a specified degree of confidence is presented. We determine an appropriate upper bound on the number of test patterns needed to satisfy a given estimation accuracy. To find out the yield improvement of the proposed test methodology, we present a method to determine the error-rate distribution of defective chips, and thus predict the fraction of defective chips that are acceptable. The proposed test methodology can support product grading, i.e., chips can be classified based on their actual error-rates such that best pricing for products used in different applications can be determined. Experimental results show that the proposed method accurately estimates error-rates of faulty chips, and the estimation results can be applied to increase the effective yield of a VLSI part as a function of various values of acceptable error-rate.
Tong-Yu Hsieh, Kuen-Jong Lee, Melvin A. Breuer
VTS1
2005 A novel test methodology based on error-rate to support error-tolerance
abstract
As the advance of VLSI technology approaches physical limitations, the yield associated with high performance system-on-chip (SOC) designs continue to decline. Conventional methodologies to address this problem, such as fault-tolerance and defect-tolerance, may become inadequate. Recently, the concept of error-tolerance has drawn much attention. Under this new concept, some defective chips (or systems) can still be labeled as acceptable, i.e., marketable, even if some outputted results are erroneous. The motivation for employing error-tolerance is to significantly increase the effective yield of some chips when used in certain applications. In this paper, we propose a novel error-rate based test methodology to support the notion of error-tolerance. Several definitions, such as various measures of yields, individual-fault and system error-rates, defect level and unacceptable defect levels are clarified or redefined. Analytically derived measures are formulated to estimate the error-rate associated with a fault, and to generate lists of faults that are acceptable with respect to a specified upper bound on the system error-rate. These results include consideration of the degree of confidence of an estimate, and provide a theoretic basis that enables the practical application of the concept of error-tolerance to both test set reduction and yield improvement. Experimental results show that the proposed test methodology can easily identify a set of acceptable faults, i.e., faults that might occur but need not cause the part to be discarded. The increase in effective yield depends on requirements imposed by end users. We show that a significant improvement in effective yield can be achieved for some applications.
Kuen-Jong Lee, Tong-Yu Hsieh, Melvin A. Breuer
ITC2