Sungho Kang 0001

dblp:95/3570 · DBLP profile ↗
← Back
117ranked-venue papers
2as first author
41since 2021 · last 2026
0000-0002-7093-2095ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 100 · 2 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Computer networks · 3Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 CLAPS: A Graph Clustering-Based Approach for Partial Scan Design
abstract
Scan is one of the representative Design for Testability (DFT) techniques designed to test sequential circuits. However, the additional hardware overhead and performance degradation caused by scan insertion can be unacceptable in specific designs. Partial scan has been applied as an alternative to the scan to balance these issues. However, previous cell selection algorithms accompany high computational complexity depending on the number of circuit components including flip-flops, and do not sufficiently consider the analysis of large-scale circuits. In this paper, a graph theory-based partial scan approach is proposed to effectively address the issues caused by scan insertion and reduce the load of structural analysis. The proposed algorithm partitions the circuit into multiple portions using graph clustering. Scan cells are selected from each subgraph to reduce sequential test generation complexity and improve testability. By partially analyzing the circuit, the proposed approach not only addresses the complexity problem of structural analysis in large-scale circuits but also can be generally applied regardless of circuit size or the number of components. The experimental results show that the proposed algorithm achieves significantly reduced processing time in seconds and reduces scan cells by approximately 11.47% with only 0.21% of test coverage loss on average compared to full scan design.
Jaeyoung Joung, Laesang Jung, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 VASE: Vector Memory Using Bit-Level Address Segmentation for High-Speed Memory Testing
abstract
To achieve high test coverage for scaled-down high-speed memory, the hardware complexity of the algorithmic pattern generator (ALPG) in automatic test equipment (ATE) has increased due to the demands of high-speed operation. However, the potential for further speed-up is constrained by the challenges associated with pipeline insertion and signal integrity preservation. Unlike ALPGs, test patterns in vector memory (VM) are generated by simply fetching pre-stored patterns without performing complex operations. Although its fetching logic is advantageous in high-speed operation, the speed of VM-based pattern generation is limited by the VM load speed. Furthermore, the limited VM capacity restricts the storage of extensive test patterns. To address the limitations of both approaches, a vector memory using bit-level address segmentation (VASE) that enables high-speed memory testing is proposed. VASE improves pattern generation speed by cyclically reusing test patterns while reducing the required VM capacity. Although VASE may impose some limitations on test algorithm coverage and incur overhead in logic area and power consumption, it still supports a wide range of commonly used test algorithms. Considering the speed improvements achieved, these trade-offs are acceptable for ATE applications.
Sooryeong Lee, Hayoung Lee, Sungho Kang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 TOPS: Topology-Based Partial Scan for Stuck-At and Delay Testing
abstract
Partial scan design reduces the shortcomings of full-scan design while retaining its advantages by selectively converting flip-flops in a circuit. In many cases, specific flip-flops cannot be converted to scan cells due to design constraints. In this paper, a topology-based partial scan method is proposed to consider both stuck-at and delay testing. The need for delay testing is growing, yet conventional partial scan methods are ineffective for delay fault models. The proposed method should consider not only stuck-at fault models but also delay fault models. The proposed method introduces three processes. First, controllable flip-flops are selected for the shift operation without reducing circuit controllability. Second, observable flip-flops are selected for the capture operation without reducing circuit observability. Third, a logic topology-based algorithm is applied to ensure that both stuck-at and delay testing are considered in the selection process. As a result, the chosen flip-flops preserve controllability and observability close to that of a full-scan design, allowing the partial scan to remain effective for delay testing. Experimental results show that the proposed method achieves higher test coverage than the previous methods and enables efficient scan-based testing for both stuck-at and transition-delay fault models.
Jaeyoung Joung, Laesang Jung, Sungho Kang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 STAR-PIM: Self-Test and Repair Structure for Processing-in-Memory With Adder Tree-Based MAC
abstract
Processing-in-memory (PIM) architectures alleviate memory bottlenecks and improve latency and energy efficiency for AI and ML workloads by accelerating general matrix-vector multiplication (GEMV) operations in DNNs. However, permanent faults in arithmetic units (AUs) within processing units (PUs) critically impact yield and inference accuracy. Although the hybrid built-in self-test (HBIST) method has been proposed, it has limited capabilities in diagnosing and repairing faulty AUs within PUs. In this study, a novel Self-Test And Repair structure for PIM (STAR-PIM) is proposed to enable both fault diagnosis and repair by incorporating a bypass mechanism. A scan-path-like approach enables the testing and precise localization of faulty AUs, while faulty adders are bypassed using a redundant adder structure integrated within the memory die. Furthermore, faulty multipliers are masked using the weight-swapping logic. Experimental results demonstrate that STAR-PIM achieves high AU-level test coverage, ranging from 98.89% to 100% with reasonable area overhead. Recovery experiments show that STAR-PIM maintains low relative errors under fault rates up to 1% for GPT-2 and preserves inference accuracy under fault rates up to 3% for MNIST-MLP. Power measurements on GDDR6-AiM indicate an average overhead of 7.39% with only a 0.08% latency increase. Consequently, STAR-PIM significantly enhances the yield and reliability of PIM while reducing test costs, making it a highly practical solution.
Seung Ho Shin, Younwoo Yoo, Youngki Moon, Nuri Son, Dahoon Kim, Sungho Kang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 A Dual-Mode Online BISR Architecture for Interconnect and Memory Repair in Chiplet-Based Systems
abstract
Chiplet-based systems provide a scalable platform for heterogeneous integration, but permanent faults in interconnects and memory cells challenge their reliability. While standards like Universal Chiplet Interconnect Express (UCIe) employ single-domain system Error Correction Code (ECC) simplify architecture, they often limit fault isolation and recovery. This paper proposes a dual-mode Online Built-In Self-Repair (OBISR) architecture for unified detection and repair of both interconnect and memory faults during post-bond testing and in-field operation. Unlike conventional methods that merely expand ECC capacity, the OBISR targets the root causes of persistent errors by repurposing existing Content-Addressable Memory (CAM) and pathfinding logic. A hierarchical CAM structure classifies faults using spatial and recurrence patterns of ECC metadata to distinguish permanent defects from transients. Memory faults are masked via logical redirection, while interconnect faults are rerouted, both operating outside the datapath to ensure zero performance degradation. Evaluations demonstrate that the OBISR improves system-level resilience by 235.49% and reduces hardware area by 16.69% compared to conventional schemes. Its integrated dual-domain and dual-mode capabilities offer a robust, scalable solution for next-generation chiplet reliability.
Donghyun Han, Sooryeong Lee, Seungtae Kim, Youngkwang Lee, Sungho Kang 0001
IEEE Trans. Reliab.7
2026 Test Cycle Reduction for TSV Test Using Streaming Scan Network on 3-D IC
Donghyun Han, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2026 A Low Area Built-In Self-Repair Using Hybrid Fault Address Memory for HBM
abstract
The massive computational requirements of large language model (LLMs) have increased the need for high-bandwidth memory (HBM), which involves high-volume data transfers. The high cell capacity of HBM results in extended test and repair times, leading to increased manufacturing costs. To reduce test time, a built- in self-repair (BISR) circuit, integrated into the HBM base die to detect and repair faults, tests multiple banks in parallel. Conventional BISR approaches adopt content-addressable memory (CAM) for fault classification to reduce repair time. However, dedicated CAM on each bank leads to substantial area overhead associated with its comparison logic. To address these issues, a novel BISR architecture that decouples fault classification and storage is proposed in this article. By introducing a linked CAM design with low area and sharing it across banks for fault classification, while small-area first-in first-out (FIFO) memories allocated to each bank store the classified fault information, the proposed architecture substantially reduces overall area overhead. Furthermore, the proposed architecture reorders the repair solution search sequence toward the most promising candidates by swapping fault entries during test idle periods, thereby significantly reducing repair time. Experimental results demonstrate that the proposed BISR architecture achieves low area overhead and fast repair time for high-density HBM.
Seung Ho Shin, Youngki Moon, Eugene Jeong, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2026 An Area-Efficient Hybrid Architecture for High-Speed ALPG
abstract
Conventional architectures for the implementation of an algorithmic pattern generator (ALPG) are subject to practical limitations. The shared-resource ALPG utilizes common arithmetic operations across multiple input/output (I/O) pins, but its operating frequency is limited by the delay of arithmetic logic units (ALUs). In contrast, the per-pin ALPG achieves high-speed operations by generating one bit of the test vector for each I/O pin independently. However, the hardware area increases significantly since an individual pin pattern generator (PPG) is required for each I/O pin. To address these limitations, an area-efficient hybrid architecture for high-speed ALPG is proposed in this article. In the proposed hybrid ALPG, the lower bits of each test vector, where bit-level transitions occur frequently, are generated by a high-speed lower-bit pattern generator (LBPG), while the upper bits, where transitions are infrequent, are generated by a low-speed upper-bit pattern generator (UBPG). The maximum test rate of the proposed hybrid ALPG is determined by the high-speed LBPG, as it generates the lower bits that must be updated every test cycle. Consequently, the proposed hybrid architecture achieves high-speed ALPG while reducing the total hardware area, as the UBPG is designed with a simple structure, and the per-pin hardware resources in the LBPG are incorporated for only a limited number of I/O pins.
Duyeon Won, Gyeonggyu Park, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2026 FADE: Fault-Aware Adaptive On-Die ECC for Improving Robustness
abstract
The increasing density of dynamic random access memory (DRAM) renders permanent faults and soft errors more prevalent, which critically reduces yield and reliability. Although error correction code (ECC) can mitigate this issue, existing ECCs are not optimized for fault correction. As a result, fault tolerance remains insufficient, and the error correction capability in the presence of faults is degraded. Therefore, to improve DRAM robustness by efficiently addressing both permanent faults and soft errors, this brief proposes a fault-aware adaptive on-die ECC (FADE) in which two ECC engines independently operate in either fault mode (FM) or error mode (EM) according to the number of faulty symbols (FSs). In FM, a fault polynomial is reconstructed by reusing the fault addresses that the built-in self-repair (BISR) stores in content-addressable memory (CAM). To calculate the corresponding fault magnitudes, a modified decoding equation is employed. As a result, the number of correctable FSs in FM doubles compared to the conventional ECC. Moreover, with the proposed symbol-based fault isolation, both fault tolerance and error correction capability in the presence of faults are drastically enhanced. Additionally, the experimental results show that the proposed design can be implemented with a reasonable overhead in terms of delay and area.
Youngki Moon, Nayeun Kim, Yeonho Choi, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2026 Memory-Optimized Block Compression for High-Speed Memory Testing
Gyeonggyu Park, Duyeon Won, Youngki Moon, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2025 PASS: Pattern-Sequence-Authentication-Based Secure Scan Against Reverse Engineering Attacks
abstract
Scan-based testing is a widely used design for testability method to ensure the ease of testing. In this method, the enhanced observability and controllability provided by the inserted scan chains significantly improve the ability to analyze circuit data. However, since the enhanced testability can be exploited by malicious users as a backdoor for attacks, countermeasures need to be implemented to prevent scan chain access by unauthorized users. Although much research on secure scan designs has been conducted, most proposed methods are vulnerable to architecture exposure by reverse engineering. Moreover, even the latest proposed methods are affected by issues related to untrustworthy test engineers. This study proposes a pattern-sequence-authentication-based secure scan that not only defends against reverse engineering-based attacks but also prevents test engineers from launching attacks using additional patterns other than the given pattern. The proposed method effectively addresses the issue of secret key leakage through valid test patterns by untrustworthy test engineers, which is a limitation of the existing methods. The experimental results show that the proposed method effectively defends against existing attack techniques and ensures high security performance.
Seokjun Jang, Youngki Moon, Duyeon Won, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 An Efficient Test Scheduling Method Based on Dynamic Pairing
abstract
Test scheduling is a process that manages tests within a System-on-Chip (SoC) to minimize test time by allocating test resources and adjusting priorities. Efficient test scheduling offers cost saving opportunities by reducing test time without compromising test coverage. Since test scheduling is a NP-hard problem, conventional methods adopt optimization or heuristic algorithms that leverage metrics of each test. However, due to the interdependence of tests based on how limited resources are allocated, finding the optimal solution to minimize test time is challenging. In this article, a dynamic pairing algorithm is proposed to consider the mutual influence of tests on each other in the test scheduling process. The proposed method identifies the available test resources for a specific pair and minimizes the test time of the paired modules under identified constraints. Additionally, the proposed algorithm employs a heuristic-based approach to test scheduling to reduce the CPU time needed for scheduling. The proposed heuristic algorithm sequentially schedules test target modules and determines the optimal test schedule through dynamic pairing. Experiments have been carried out on diverse benchmarks and under various constraint conditions. The results indicate that the proposed method achieves shorter test times on average in comparison to conventional methods.
Heetae Kim, Hyojoon Yun, Doohyun Yoon, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 A Robust Test Architecture for Low-Power AI Accelerators
abstract
With the rapid advancement of artificial intelligence (AI), there has been extensive research on AI accelerators to meet the demand for data-intensive analytics. Recently, low-power AI accelerators have been also developed to support battery-operated edge devices and minimize power consumption. However, traditional test architectures are insufficient for effectively testing such low-power AI accelerators. To address this issue, a robust test architecture for low-power AI accelerators has been proposed in this article. The proposed test architecture employs a simple clock-gating technique in systolic array-based low-power AI accelerators and conducts testing through their functional paths. Accordingly, it can achieve 100% test coverage for both stuck-at and transition-delay faults with a minimal number of test patterns. Additionally, the proposed test architecture requires negligible area overhead since only one AND gate is implemented for the entire systolic array in low-power AI accelerators.
Hayoung Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 A New Pipelined Output Data Reducer of BOST for Improved Parallelism
abstract
To reduce the cost of memory production, built-off self-test (BOST) enables low-speed automatic test equipment to test the high-speed memory. To maximize the cost reduction benefit of BOST, it is crucial to test as many memories as possible using as few test output pins as possible. For this purpose, a new pipelined output data reducer called PODR is proposed for the output data reduction, and channel sharing between memories tested in parallel is introduced. The proposed structure is adopted to reduce hardware complexity while facilitating test output channel sharing between concurrently tested memories. Additionally, further output data reduction can be achieved by integrating the output data code into the pipelined structure. Output data reduction is also attainable by transmitting fault cell addresses using relative distance from the previously detected fault cells rather than the absolute addresses. Reducing the total code length can be achieved by the adoption of relative addressing, but this requires additional code transmission as its overhead. To mitigate this overhead, a revised approach to relative addressing is introduced. Consequently, as the number of memories tested in parallel increases, the amount of output data of PODR decreases and the number of normalized test output pins usages is reduced in half compared to the previous works.
Sooryeong Lee, Hayoung Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 PETRA: Powerful Early Termination-Based Redundancy Analysis
abstract
In dynamic random-access memory (DRAM), memory redundancy analysis (RA) is a crucial process for enhancing memory yield and reducing production costs. It finds a memory repair solution by efficiently allocating the limited number of spare cell lines that replace a faulty cell. However, it is challenging to quickly find a memory repair solution because RA is an NP-complete problem. To address this issue more effectively, we present a powerful early termination-based high-speed RA method. This method rapidly assesses memory repairability, terminating the RA process early in cases where repair is impossible, or a solution can be easily found. Additionally, by dividing faulty cells into several groups, the proposed RA method finds fast and approximate albeit nonoptimal solution sets for each group. This facilitates the rapid acquisition of a memory repair solution without the need to search for all the optimal solution sets. These features enable RA to be promptly executed while ensuring the repair solution for any repairable memory. Experimental results demonstrate that the proposed RA method can find a repair solution faster than the existing RA methods.
Youngkwang Lee, Hyojun Yun, Younwoo Yoo, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 An Efficient Low-Power BIST for Automotive SoC With Periodic Pattern Type Selection
abstract
In the realm of automotive System-on-Chip (SoC), scan-based logic built-in self-test (LBIST) is commonly utilized for in-system testing, primarily for its cost-effectiveness. Nevertheless, this approach encounters challenges, particularly in attaining high-test coverage within constrained test times. The challenge intensifies when implementing low-power patterns, as it further complicates the achievement of adequate test coverage. To overcome these hurdles, this article introduces a novel testing methodology that enhances test coverage using low-toggled patterns. This method consists of two primary phases. Initially, it involves grouping and pairing scan cells (SCs), ensuring adjacent placement of paired cells. The following phase involves the generation of low-toggled patterns, tailored to the specific arrangement of SCs. To optimally detect as many previously undetected faults as possible, this method applies the low-power pattern selectively to particular scan groups. Furthermore, the scan group subjected to low-power patterns alternates after a certain number of patterns. Experimental results demonstrate the superiority of this proposed method in both fault detection and power reduction, in comparison to earlier methods.
Hyemin Kim, Jaeyoung Joung, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 A Novel CNN-Based Redundancy Analysis Using Parallel Solution Decision
abstract
The increase in memory cell density and capacity has resulted in more faulty cells, necessitating the use of redundant memory row and column lines for repairs. However, existing redundancy analysis (RA) algorithms face a critical issue that RA time increases exponentially with the number of faulty cells. Furthermore, RA solutions for multiple memory chips cannot be derived simultaneously. In this study, a novel RA method is proposed using a convolutional neural network (CNN). The proposed RA algorithm also includes preprocessing to improve training accuracy. The solution locations on the fault map are predicted using multi-label classification. Moreover, parallel solution decision methods ensure that even if the CNN does not find the correct RA solution, an accurate final solution can still be derived, and PyCUDA is used to process multiple memories in parallel. From the experimental results, the normalized repair rate of the proposed RA is 100%. The RA time of the proposed RA is not affected by the number of faults but rather by the CNN execution time. Moreover, RA solutions for multiple memories can be quickly derived simultaneously by utilizing GPU parallel processing. In conclusion, a high yield and low test cost can be achieved.
Seung Ho Shin, Minho Cheong, Hayoung Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 A Built-In Self-Repair With Maximum Fault Collection and Fast Analysis Method for HBM
abstract
High bandwidth memory (HBM) represents a significant advancement in memory technology, requiring quick and accurate data processing. Built-in self-repair (BISR) is crucial for ensuring high-capacity and reliable memories, as it automatically detects and repairs faults within memory systems, preventing data loss and enhancing overall memory reliability. The proposed BISR aims to enhance the repair rate and reliability by using a content-addressable memory structure that operates effectively in both offline and online modes. Furthermore, a new redundancy analysis algorithm reduces both analysis time and area overhead by converting fault information into a matrix format and focusing on fault-free areas for each repair solution. Experimental results demonstrate that the proposed BISR improves repair rates and derives a final repair solution immediately after the test sequences are completed. Moreover, hardware comparisons have shown that the proposed approach reduces the area overhead as memory size increases. Consequently, the proposed BISR enhances the overall performance of BISR and the reliability of HBM.
Joonsik Yoon, Hayoung Lee, Youngki Moon, Seung Ho Shin, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Multistage Enhanced Diagnosis With Fault Candidate Reduction
abstract
Logic diagnosis is essential for improving reliability and yield. In conventional diagnosis methods, although various methods are proposed to enhance the accuracy and resolution of logic diagnosis, there are still diagnosis results where the reported locations of defects are incorrect. Particularly in logic circuits, which contain a large number of gates, multiple faults can occur, not just single faults. Since the number of possible cases for multiple faults is significantly greater compared to single faults, the diagnosis of multiple faults is complicated. To address this problem, a new diagnosis method that uses a multistage process with fault candidate reduction is proposed. In the proposed method, machine learning is used with fault candidate reduction, and post-processing is performed after the use of machine learning. This proposed method allows for the analysis of multiple faults using only the test responses for single faults, demonstrating that this method can maintain sufficient accuracy and resolution for unexpected faults.
Hyojoon Yun, Hyeonchan Lim, Hayoung Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 TSV Built-In Self-Repair Architecture for Lifespan Reliability Enhancement of HBM
abstract
High-bandwidth memory (HBM) is one of the 3-D stacked memory standards that demonstrate high performance, including high bandwidth, large capacity, and low power consumption. However, despite these advantages, issues related to reliability and yield have imposed limitations on mass production. Various methodologies to enhance the reliability of HBM have been proposed, such as built-in self-repair (BISR) architectures and error correction code algorithms. Nevertheless, ensuring the reliability of through-silicon vias (TSV) remains a challenging problem. Existing built-in architectures aiming to enhance TSV reliability often incur significant hardware overhead, limiting practical applications. In this article, an innovative TSV BISR architecture that can detect and repair permanent TSV faults in real time at the user stage is proposed. The proposed architecture significantly enhances the reliability of HBM while implementing it with minimal hardware overhead. Furthermore, it effectively identifies both temporary errors and permanent TSV faults, enabling efficient TSV repairs. Through fast and accurate TSV fault repair, the proposed architecture substantially improves the reliability of HBM.
Donghyun Han, Duyeon Won, Sungho Kang 0001
IEEE Trans. Reliab.4
2025 SPOT: Fast and Optimal Built-In Redundancy Analysis Using Smart Potential Case Collection
abstract
With advancements in manufacturing and design technology, memory integration density has improved. However, as integration density increases, the cost of testing and repairing memory has also risen, posing a significant challenge in memory production. To address this challenge, built-in self-repair (BISR) has been proposed. Traditional built-in redundancy analysis (BIRAs) performs limited analysis of faults during the fault collection process, resulting in a significant delay in generating a repair solution after the test sequence is completed. This inefficiency arises from the time required to repair the memory posttest. This article proposes a new fast and optimal BIRA using smart potential case collection. The proposed BIRA conducts a detailed analysis of detected faults during the test process. Using this novel fault collection results, a potential case is generated. This is a repair case that can repair the memory with a high probability and is generated immediately after the test sequence ends. If the memory cannot be repaired by the potential case, an exhaustive search is conducted for the faults requiring further analysis to generate an optimal repair solution. Compared to previous studies, the proposed BIRA demonstrates extremely low analysis time with an optimal repair rate.
Donghyun Han, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2025 An Efficient Test Architecture Using Hybrid Built-In Self-Test for Processing-in-Memory
abstract
With the rapid advances in artificial intelligence (AI), the demand for data-intensive analytics has surged. Consequently, extensive research on AI acceleration has been conducted to enhance AI performance. Processing-in-memory (PiM) has emerged as a promising AI acceleration architecture, offering an unprecedented high-bandwidth connection between compute and memory. However, integrating many components in PiM can lead to yield degradation. To address this issue, we propose an efficient test architecture that utilizes a hybrid built-in self-test (BIST) for PiM. This architecture utilizes the structural and operational characteristics of PiM to facilitate testing. It can execute testing through the existing functional paths without requiring any additional hardware implementation in PiM. Furthermore, it achieves a 100% test coverage with the small number of test patterns. In addition, the functionality of self-test can be realized for PiM through reconfiguration of the existing hardware, resulting in a very small area overhead.
Hayoung Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2025 A Cost-Effective Per-Pin ALPG for High-Speed Memory Testing
abstract
An algorithmic pattern generator (ALPG) has been developed within automatic test equipment (ATE) due to the extensive number of test patterns required for testing the memories. Since shared-resource ALPG generates the test pattern using the same arithmetic instruction and timing across multiple input/output (I/O) pins, the maximum operating frequency is limited by the delay of the arithmetic operation. On the other hand, per-pin ALPG can achieve high-speed operations by generating one bit of the test pattern for each I/O pin. However, the hardware cost is significantly increased due to the need for individual instruction and pattern generator (PG) for each I/O pin. To address these limitations, a cost-effective per-pin ALPG for high-speed memory testing is proposed. The proposed per-pin ALPG can achieve high-speed operations, and the hardware resources for storing and decoding the instructions are shared among multiple I/O pins to reduce the hardware cost. The experimental results indicate that the proposed ALPG can achieve a higher speed than the conventional per-pin ALPG with a reasonable hardware cost comparable to the conventional shared-resource ALPG.
Hayoung Lee, Sooryeong Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A Novel Prediction-Based Two-Tiered ECC for Mitigating SWD Errors in HBM
abstract
Errors emerge as a major issue in the reliability of dynamic random access memory (DRAM). To enhance reliability, a two-tiered error correction code (ECC) architecture that comprises on-die ECC (OD-ECC) and system ECC (S-ECC) is adopted as a part of the standard for state-of-the-art high-bandwidth memory (HBM). However, conventional ECCs are insufficient to mitigate malfunctions of subwordline drivers (SWDs), a primary cause of errors. Moreover, the efficient co-design of two-tiered ECCs has not been sufficiently studied. To address these issues without increasing the size of check bits, this article proposes a two-tiered ECC architecture comprising an OD-ECC based on prediction and an S-ECC with data deinterleaving. The proposed OD-ECC predicts the SWD errors by leveraging the detection capabilities of two interleaved Reed-Solomon (RS) engines. In addition, the proposed S-ECC not only preserves strong error detection capability but also masks the misprediction effect of OD-ECC, where data deinterleaving renders additional errors caused by misprediction of OD-ECC to be bounded in the detectable range of the employed cyclic redundancy check (CRC). The experimental results demonstrate that the proposed two-tiered ECC can significantly enhance the error correction capability for SWD errors while maintaining the correction capability for other types of errors.
Youngki Moon, Seung Ho Shin, Seokjun Jang, Duyeon Won, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Effective Parallel Redundancy Analysis Using GPU for Memory Repair
abstract
The rapid increment of the memory density leads to an increment of fault occurrence in memory cells. To improve the memory yield, effective memory test and repair methodologies for automatic test equipment (ATE) have been studied. Multiple memory chips are tested simultaneously by the ATE to improve throughput and reduce costs. In general, redundancy analysis (RA) is used for memory repair. However, since conventional RA methods store fault information in the respective failure bitmaps and operate sequentially, those have limitations due to the high area and analysis time. To address these problems, a novel graphic processing unit (GPU)-based RA method has been proposed which significantly enhances the efficiency of searching for repair solutions for multiple memories. The proposed RA method strategically focuses on the pivot line to efficiently utilize parallel processing and reduce the solution search space. Moreover, the proposed method does not require the extensive use of failure bitmaps since all process is conducted on the GPU. The process involves real-time fault collection, analysis, spare allocation, and solution decision process dynamically during the memory test. Experimental results demonstrate that the performance of the proposed RA method achieves an optimal repair rate and high analysis speed for multiple memories.
Seung Ho Shin, Hayoung Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2024 A New Fail Address Memory Architecture for Cost-Effective ATE
abstract
Memory test and repair has been generally applied to improve memory yield. However, due to the high cost of automatic test equipment (ATE) equipment, which has been employed for memory test and repair, there is a significant focus on reducing the ATE expense. One of the major problems, which contribute to the increase in ATE cost, is fail address memory. The size of fail address memory, where memory fault information is stored during the memory test, has continuously grown in line with the memory capacity increase. To address the problem, a new fail address memory architecture for cost-effective ATE is proposed in this article. In the proposed architecture, memory fault information is compressed and unrequired memory fault information is eliminated. In addition, a new structure of fail address memory is used to efficiently store memory fault information. Accordingly, the size of fail address memory is highly reduced in the proposed architecture. Furthermore, since some information, which can be used during the memory repair, can be collected during the memory test, the redundancy analysis time required to find memory repair solutions is also reduced in the proposed architecture. The advantages were verified experimentally.
Hayoung Lee, Sooryeong Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 GRAP: Efficient GPU-Based Redundancy Analysis Using Parallel Evaluation for Cross Faults
abstract
Various memory repair methodologies based on redundancy analysis (RA) have been developed to improve the memory yield. However, conventional RAs often encounter difficulties in finding repair solutions for cases involving a large number of faults and redundancies. To address this problem, an efficient graphics processing unit (GPU)-based RA is proposed using Parallel evaluation for cross faults (GRAP). GRAP involves a preprocessing stage during memory testing, leveraging the parallel processing capacities of the GPU. Preprocessing facilitates rapid solution search by analyzing the fault information. After the test, the solution search is performed. The GPU threads are used to implement all possible cases of redundancy allocation, focusing on cross faults. The remaining faults are categorized by allocating the corresponding redundancies using an efficient method. Given that the solution search process efficiently exploits the multiple threads, GRAP can rapidly find a solution even in cases with a large number of faults and redundancies. Experiments are performed using the compute unified device architecture (CUDA) library for GPU parallel processing, and the performance of the GRAP is compared with those of conventional RA methodologies. The results demonstrated that the proposed RA method can achieve an optimal repair rate with a high analysis speed by leveraging efficient parallel computing.
Seung Ho Shin, Hayoung Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 RA-Aware Fail Data Collection Architecture for Cost Reduction
abstract
As fault occurrence probability has increased with corresponding increases in memory density and capacity, memory test and repair have been widely used. However, the total cost for these has increased dramatically due to the increased cost of automatic test equipment (ATE) and the time required for redundancy analysis (RA). The increase in ATE cost has been caused by the increased size of fail address memory, where memory fault information is stored during memory test. The RA time has also increased because the difficulty encountered during fault analysis has increased in proportion to the increase in the number of memory faults. To address these problems, an RA-aware fail data collection architecture is proposed. This includes a new fail address memory structure that can significantly reduce the fail address memory size. Moreover, the architecture can integrate memory fault information using simple calculations without any data losses. In addition, unnecessary memory fault information can be eliminated easily with efficient data encoding to reduce the data size. Furthermore, since some information required for fault analysis in memory repair can be collected during memory test, the RA time needed to find memory repair solutions is also reduced without any degradation in the repair rate. Consequently, the total cost for memory test and repair can be considerably reduced by reducing the cost of ATE and the RA time. Experimental results reveal that the fail address memory size and RA time can be reduced by an average of 63% and 41%, respectively, with the proposed architecture.
Hayoung Lee, Sooryeong Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2024 An Area-Efficient Systolic Array Redundancy Architecture for Reliable AI Accelerator
abstract
The increasing demand for data-intensive analytics, driven by the rapid advances in artificial intelligence (AI), has led to the proposal of various AI accelerators. However, as AI-based solutions are being applied to applications that require high accuracy and reliability, ensuring the dependability of these solutions has become a critical issue. In this brief, we present an area-efficient systolic array redundancy architecture for reliable AI accelerator. In the proposed architecture, computations assigned to faulty multiply-accumulate (MAC) units are bypassed using dedicated routes. Subsequently, the same computations are executed in shiftable redundant MACs or selectable redundant MACs. This ensures the correct completion of calculations all without performance reduction. Moreover, the reassignment of computations can be efficiently managed through a simple scheduling algorithm. As a result, the proposed architecture achieves a high repair rate through the redundant MACs and effective computation reassignment. Despite these capabilities, the proposed architecture incurs only a small area overhead.
Hayoung Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2023 Scan Chain Architecture With Data Duplication for Multiple Scan Cell Fault Diagnosis
abstract
Scan chain diagnosis is an important step in solving yield problems in early manufacturing processes. The higher the diagnostic resolution, the better the yield in the initial process. In the conventional scan chain architecture, if failed scan chain data are observed during scan mode, all the shifted data through a stuck-at faulty cell are contaminated. As a result, only a single fault can be diagnosed, and there are difficulties in diagnosing multiple stuck-at fault locations. The number of multiple faults has increased significantly, increasing the cost of physical fault analysis. In addition, it is difficult to diagnose faults with a high resolution because there are many candidate faults early in the process. In this article, a new hardware architecture with data duplication is proposed to diagnose fault locations by deliberate voltage collision even if multiple faults occur. The resources required for the diagnosis include a minimum of one good scan chain, a diagnosis circuit, and a diagnostic line. Experimental results show that the proposed method has lower hardware and routing overhead and fewer additional pin counts than the existing method. The diagnostic speed is also faster, and the larger the number of scan chains (the larger the circuit), the higher the diagnostic performance, which has advantages over conventional methods.
Seokjun Jang, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Novel Error-Tolerant Voltage-Divider-Based Through-Silicon-Via Test Architecture
abstract
A voltage-divider-based through-silicon-via (TSV) test architecture tests the TSV by using the voltage value differently divided according to TSV defects. This architecture is widely used for TSV testing owing to its small hardware overhead and high test speed. However, the existing voltage-divider-based TSV test architectures are vulnerable to process–voltage–temperature (PVT) variations and noise. In addition, they cannot effectively detect pinhole defects. This study proposes a novel error-tolerant voltage-divider-based TSV test architecture to address these problems. The proposed architecture reduces the test errors by appropriately adjusting the on-resistance value of each MOSFET and adding a compensator circuit. In addition, it effectively detects the pinhole defects by modifying the voltage divider structure and changing the MOSFET control method. Experimental results reveal that the proposed architecture promptly tests various TSV defects and significantly reduces the test errors.
Youngkwang Lee, Donghyun Han, Sooryeong Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 STRAIT: Self-Test and Self-Recovery for AI Accelerator
abstract
As the demand for data-intensive analytics has increased with the rapid advance in artificial intelligence (AI), various AI accelerators have been proposed. However, as AI-based solutions have been adapted to applications requiring accuracy and reliability, the reliability of them has become a critical issue. For this reason, self-test and self-recovery for AI accelerator (STRAIT) is proposed in this article. It facilitates self-test, self-diagnosis, and self-recovery by utilizing the structural and operational characteristics of systolic array in AI accelerator. The proposed self-test is progressed using scan chains composed of functional paths and can achieve a 100% test coverage (for both stuck-at and transition-delay faults) with a small number of test patterns and reduced test power. The proposed self-diagnosis is progressed with the proposed self-test in real time and allows accurate fault localization with fault type analysis. The proposed self-recovery is progressed using efficient pruning for faulty processing elements with weight allocation, and the reliability of AI accelerators can drastically increase with negligible performance degradation. However, STRAIT can be implemented with a small area overhead.
Hayoung Lee, Jihye Kim 0002, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 TRUST: Through-Silicon via Repair Using Switch Matrix Topology
abstract
To address the demand for memory scaling capabilities, 3-D integrated circuits (3D-ICs) based on short and dense through-silicon vias (TSVs) have been introduced. However, the defects of TSVs considerably influence the yield and reliability of 3D-ICs. For this reason, TSV repair using switch matrix (SM) topology (TRUST) is proposed in this article. TRUST adopts an SM, which has a high routing flexibility, to realize TSV connections. Consequently, a 100% repair rate can be achieved for the 3D-ICs that have faulty TSVs smaller than or equal to redundant TSVs. Furthermore, TRUST utilizes content-addressable memories in built-in self-repair to identify TSV repair paths via a simple TSV repair path search algorithm. For this reason, TRUST can be applied to repair manufacturing and aging defects of TSVs. Nevertheless, TRUST can be applied with reasonable area and delay overheads, such as 58.3% area reduction and 55.1% delay reduction compared to the only conventional TSV repair architecture that can achieve the optimal repair rate. In addition, the area ratios in high bandwidth memory (HBM) and HBM2 are only 5.3% and much smaller than 0.1%, respectively. The advantages are experimentally verified.
Hayoung Lee, Seung Ho Shin, Younwoo Yoo, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 TSV Built-In Self-Repair Architecture for Improving the Yield and Reliability of HBM
abstract
High-bandwidth memory (HBM) is the latest 3-D-stacked dynamic random access memory (DRAM) standard adopted in Joint Electron Device Engineering Council (JEDEC). It has many advantages, such as high bandwidth, large capacity, and low power consumption, but mass production is challenging due to its low yield and reliability. One of the reasons is that through-silicon-vias (TSVs) are prone to defects. Therefore, HBM requires a TSV built-in self-repair (TBISR) architecture that can repair the TSV at a high repair rate even after chip shipment; however, implementing it through existing TSV repair architectures is difficult. They have a large area overhead or a low repair rate. In addition, they lack consideration for bidirectional TSV repair. To address these issues, this article proposes a novel TBISR architecture that can repair bidirectional TSVs and has a small area overhead and a high repair rate. Experimental results show that the proposed architecture, capable of bidirectional TSV repair, has a high repair rate, despite the small size compared to other architectures.
Youngkwang Lee, Donghyun Han, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2022 SPAR: A New Test-Point Insertion Using Shared Points for Area Overhead Reduction
abstract
Test-point insertion (TPI) is an effective technique for improving the random pattern testability of digital circuits. However, it introduces area and performance overhead. Because the test-point area takes a significant portion of the test logic area, many techniques have been studied to reduce the area impact, such as sharing a control point (CP) driver with multiple CPs or replacing a dedicated CP driver with an existing flip-flop. This article proposes shared point insertion for area overhead reduction (SPAR) to simultaneously reduce the area impact of CPs and observation points (OPs). SPAR inserts a shared point instead of inserting a pair of CP and OP individually. Consequently, the pair of CP and OP is provided requirements, such as a control signal or propagation path from each other through the shared point—accordingly, the shared point simultaneously functions as the CP and OP. Furthermore, a signal that drives a shared point can be chosen to ensure the fair propagation of faults. The proposed flow searches for appropriate CP–OP pairs to insert shared points while avoiding potential issues from the newly created path. Experimental results on benchmarks demonstrate that SPAR can significantly reduce area overhead caused by test points while achieving almost identical or even slightly improved test coverage.
Gyungbin Kim, Minho Cheong, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Multibank Optimized Redundancy Analysis Using Efficient Fault Collection
abstract
With technological advancements, the density and capacity of memory are rapidly increasing. As the number of memory cells increases, the difficulty of fault analysis and the number of faults also increase. Hence, the yield and test cost of memory have become essential issues in memory manufacturing. Many manufacturers have used redundancy analysis (RA) to improve the memory yield and decrease the test cost. However, most conventional RA methods require a lengthy analysis time to find a repair solution, and it is difficult to obtain an optimal repair rate with conventional RA algorithms. Although several algorithms using various spare structures to achieve performance improvement have been proposed, those improvements have not been ground breaking. In this article, a new multibank optimized RA (MORA) algorithm is proposed. It achieves a very high repair rate and a drastic reduction in the analysis time compared with conventional RA algorithms using various spare structures. During testing, the proposed algorithm stores the faulty cell information efficiently. Therefore, the analysis time can be shortened through the presolution process of the repair analysis using the proposed fault storage spaces. Additionally, the proposed spare structures are used to increase the repair rate. The experimental results reveal that the proposed algorithm can achieve a very high repair rate at a faster speed than conventional RA algorithms.
Hogyeong Kim, Hayoung Lee, Donghyun Han, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Reduced-Pin-Count BOST for Test-Cost Reduction
abstract
Built-off self-test (BOST) is a widely used technique to reduce the test cost. It makes it possible to test high-speed dynamic random-access memory (DRAM) without using a costly high-performance automatic test equipment (ATE). However, the currently used BOSTs require many ATE connection pins, which degrade the cost reduction effect. In this article, we propose a novel reduced-pin-count BOST to reduce the test cost. The proposed BOST uses bidirectional pins to employ the pins as efficiently as possible. Thus, even if the same amount of data is transferred, fewer pins are required than the previous BOSTs. In addition, it reduces the amount of output data by sending only the information necessary for a DRAM repair process. This is possible because the DRAM repair process requires only the location information of some faulty cells. Therefore, the proposed BOST can send output data with fewer pins compared with the previous BOSTs. Experimental results indicate that the proposed BOST can test high-speed DRAMs using a few ATE connection pins.
Youngkwang Lee, Sungyoul Seo, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 A Hybrid Test Scheme for Automotive IC in Multisite Testing
abstract
This article proposes a hybrid test scheme based on interleaved test applications to increase the efficiency of multisite testing of integrated circuits (ICs) by solving problems related to test coverage, which is an important feature in automotive IC, and test costs, including the number of test pins and the volume of test data. To solve the problems, a hybrid test scheme that combines the merits of broadcast-based scan compression and built-in self-test (BIST) is proposed. The proposed scheme not only generates deterministic test patterns, similar to scan compression, but also requires fewer test pins, similar to BIST, thereby increasing the efficiency of multisite testing. With appropriately designed linear-feedback shift registers (LFSRs) and seeds periodically inserted through the reduced test pins, the test application process combines deterministic testing and random testing. The random testing stage can compensate the seed-insertion time for the subsequent deterministic testing stage and reduce the number of required deterministic patterns by detecting easy-to-detect faults. For further reduction of the volume of test data and hardware overhead, some compatible scan chains (SCs) are grouped and receive test data from an LFSR. The experimental results show that the proposed method increases the efficiency of multisite testing with a reduced pin count.
Hyeonchan Lim, Hyojoon Yun, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Herringbone-Based TSV Architecture for Clustered Fault Repair and Aging Recovery
abstract
Three-dimensional integrated circuits (3-D ICs) utilizing through-silicon via (TSV) technology have many advantages over 2-D ICs, including high bandwidth, high density, and low power consumption. However, TSV, which is a key feature of 3-D ICs, has not only problems due to defects in the manufacturing process but also potential problems due to aging. Various solutions have been proposed to address each of these issues, but no one solution has been proposed considering both. In practice, to improve the overall reliability of the TSV, the two problems should be solved together, not separately. In this article, a new TSV architecture is proposed to cope with both issues. The proposed TSV architecture uses redundant TSVs (RTSVs) to repair faulty TSVs due to manufacturing defects and uses unused RTSVs in this way to solve the aging-related problems. Experimental results show that the proposed architecture achieves similar repair rate with less than 1% difference in less than six clustered faults using smaller hardware overhead, and also shows that unused RTSVs are available with a 98.5% high probability, resulting in a 1.5 times improvement in lifetime.
Minho Cheong, Donghyun Han, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 ECMO: ECC Architecture Reusing Content-Addressable Memories for Obtaining High Reliability in DRAM
abstract
Advances in the density and capacity of dynamic random access memories (DRAMs) have resulted in emerging reliability issues. The error correction code (ECC) is widely used as a promising technique to improve the reliability of high-density memories. For this reason, many studies on ECC have been conducted to address the increased cell failure rates. However, conventional ECCs have shown limited achievements owing to area, latency, and power overheads. This study proposes ECC architecture reusing content-addressable memories (CAMs) for obtaining high reliability in DRAM, which can be called ECMO. The proposed architecture reuses CAMs in built-in self-repair, which can be used to repair memory hard faults during manufacturing as data storage to replace error data words. This achieves high reliability along with an additional 9155 h DRAM lifetime. Nevertheless, it can be implemented with a 3.04% area overhead due to the reuse of CAMs. Moreover, only 0.21 ns is added to the critical path. Furthermore, the power overhead is 0.1% compared to the total power consumption of DDR3 and DDR4.
Hayoung Lee, Younwoo Yoo, Seung Ho Shin, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2021 Enhanced Postbond Test Architecture for Bridge Defects Between the TSVs
abstract
The use of through silicon vias (TSVs) is essential for vertically connecting the individual circuit layers of a 3-D IC. However, because the pitch between TSVs is being decreased, it becomes more vulnerable to bridge defects between the TSVs. In the previous architectures, the parallel topology between the power supply and the bridge defects reduces the resolution of the test results and makes them very sensitive to process variations. In this article, an enhanced postbond test architecture that improves the testability of bridge defects screening with a floating control circuit of supply voltage driver is proposed. The proposed structure allows the selective formation of a series circuit that includes the supply voltage driver and bridge defects between the TSVs; the structure can improve the output voltage resolution of the bridge defects by 27.2 times over that of the previous architectures. According to the experimental results, while exhibiting a proper test time, hardware overhead, and peak current consumption for mass production, the system guarantees testability for the detection of bridge defects. Moreover, the Monte Carlo simulation results show that the proposed architecture has a higher test reliability under process variations than the current test architectures.
Jungil Mok, Hyeonchan Lim, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2020 W-ERA: One-Time Memory Repair with Wafer-Level Early Repair Analysis for Cost Reduction
abstract
Since the probability of fault occurrence on memory has increased with the advance of memory density and capacity, memory repairs in wafer-level and package-level are widely used with redundancy analysis (RA) to improve memory yield. However, as the costs for memory repair also have increased in proportion to the memory density and capacity, the repair costs have occupied a significant portion of the total costs. To address the problem, one-time memory repair with wafer-level early repair analysis (W-ERA) for cost reduction is proposed in this paper. The proposed W-ERA facilitates that all unrepairable memories are classified rapidly without searching memory repair solutions in wafer-level and repairable memory faults occurred in wafer-level are repaired in package-level with additional faults occurred in package-level simultaneously. It means, as the costs of memory repair can be highly reduced since memory repair is skipped in wafer-level, the total costs also can be highly reduced. In addition, memory redundancies can be efficiently used for memory repair in package-level and it results a high repair rate achievement.
Hayoung Lee, Donghyun Han, Hogyeong Kim, Sungho Kang 0001
ITC-Asia4
2020 Fail Memory Configuration Set for RA Estimation
abstract
Since the redundancy analysis (RA) has been introduced for memory yield, many RA researches have been conducted. However, objective comparisons of them are difficult by the absence of real memory models with realistic fault distributions. This paper presents a fail memory configuration set for RA estimation, called as ITC'2020 RA Benchmarks. It enables objective estimations of RAs with respect to effectiveness and efficiency. The fail memory configuration set includes memory models which have various redundancy structures and a fault generation algorithm with fault distribution which can be criteria for objective comparisons of RA. Simulations for estimations and comparisons of RA researches including BIRA are progressed utilizing the fail memory configuration set.
Hayoung Lee, Keewon Cho, Sungho Kang 0001, Wooheon Kang, Seungtaek Lee, Woosik Jeong
ITC3
2020 A 3-D Rotation-Based Through-Silicon via Redundancy Architecture for Clustering Faults
abstract
Three-dimensional integrated circuits (3-D ICs), which feature many benefits, such as high bandwidth and a high degree of integration, have recently received considerable attention from the semiconductor industry. However, these chips feature through-silicon vias (TSVs), which vertically connect multiple dies, and these TSVs may fail, resulting in a decreased yield. Unfortunately, previously proposed methods to repair TSVs cannot handle certain failure patterns. For example, existing techniques cannot repair clustered TSV faults, which commonly occur in practice. Furthermore, the number of signal TSVs typically determines the number of redundant TSVs, which may result in wasteful and redundant TSVs. In this paper, a new TSV repair scheme is proposed that replaces defective TSVs with redundant TSVs by utilizing the architecture of a cube, which can replace any face with any of the other faces. Both signal TSVs and redundant TSVs are placed in the face of cube, so any faulted TSVs can be replaced with redundant TSVs. The experimental results indicate that the new method guarantees 100% coverage with any number of signal TSVs and redundant TSVs.
Minho Cheong, Ingeol Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Robust Secure Shield Architecture for Detection and Protection Against Invasive Attacks
abstract
Invasive attacks, such as microprobing or focused-ion-beam (FIB) circuit editing are serious threats to security-related semiconductors. To ensure that there is security against invasive attacks, an effective countermeasure is to use a protective layer as a secure shield. Previous secure shield methods can be classified into one of two categories; detection circuits based on the delay difference or block ciphers. For the former, timing asymmetries caused by the capacitance of the probe are detected. The main drawback of this method is that it is highly vulnerable to chip editing by FIB equipment. FIB circuit editing can easily cripple the detection circuits of the secure shield. In contrast, the cryptographically secure shield based on the block cipher can provide strong protection against FIB circuit editing. However, it is prone to microprobing attacks because of its inability to detect the capacitance load of the probe. In this article, we propose a robust secure shield architecture against invasive attacks, including both probe attempts and the FIB circuit editing. The proposed method is based on the detection circuits with low hardware overhead and fast-analysis time and includes protection circuits to prevent information from being leaked.
Hyeonchan Lim, Youngkwang Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 GPU-Based Redundancy Analysis Using Concurrent Evaluation
abstract
Redundancy analysis (RA) is essential for improving memory yield. The recent increase in memory size has made RA more complicated. This article presents graphics processing unit (GPU)-based RA using concurrent evaluation (GRACE), which is an efficient RA technique. In GRACE, to perform dynamic RA, memory faults found during the test are directly analyzed instead of being stored in the fault bitmap in the automatic test equipment (ATE). Therefore, RA is performed simultaneously with the memory test, and the RA latency is eliminated after the test time. Using the GPU, all possible repair cases are examined in parallel; thus, a high memory repair rate is achieved in a short period of time. Also, GRACE can be applied to practical environments where the structure of memory redundancy is complicated. Experimental results indicate that GRACE is faster than other ATE-based RA methods since it completes the RA almost simultaneously at the end of the test. Additionally, the repair rate of GRACE is always higher than those of the other RA methods.
Hayoung Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2019 An Efficient BIRA Utilizing Characteristics of Spare Pivot Faults
abstract
The current growth of micro-semiconductor technologies requires that an effective solution be found to address the yield and reliability issues associated with embedded memories. A common solution is built-in redundancy analysis (BIRA), which is utilized to guarantee reasonable memory yields. The most common form of BIRA is a module that stores and analyzes fault addresses with a 2-D spare architecture. When the performance of BIRA is evaluated, numerous different parameters are considered, such as repair rate, area overhead, and analysis speed. Because there is a tradeoff between these criteria, many BIRA approaches have been studied so that an ideal BIRA can be found. A novel BIRA approach that focuses on a 100% repair rate and a minimal area overhead is proposed in this paper. In the fault collection phase, the proposed BIRA stores only the essential part of fault addresses in content addressable memories (CAMs), with the rest of the fault addresses being stored in spare memories. After the fault collection phase, a redundancy analysis procedure is performed with the minimum amount of fault information that is stored in the proposed CAM structure. By doing so, the proposed BIRA algorithm can repair all repairable faulty memories while maintaining a minimal area overhead. Our experimental results confirm that the proposed approach exhibits outstanding performance for area overhead, especially when compared to other BIRA approaches that have 100% repair rates.
Keewon Cho, Sungyoul Seo, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 TSV Repair Architecture for Clustered Faults
abstract
The poor quality of the die stacking process for 3-D integrated circuits can result in the failure of the process in the through-silicon-vias (TSVs) in dense regions. Previous works use the same number of redundant TSVs and architectures that do not consider the TSV density. A repair architecture and an appropriate number of redundant TSVs, which are chosen considering the TSV density, are required for an improved repair rate. This paper proposes a method that demonstrates such an architecture and calculates the required number of TSVs. The method has a high repair rate for clustered faults, and wire-length problems are solved using the shift-based repair method.
Jaewon Jang, Minho Cheong, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Highly Reliable Redundant TSV Architecture for Clustered Faults
abstract
A three-dimensional (3-D) integration technology involving the use of a through-silicon-via (TSV) offers advantages such as low power consumption, small form factor, and large bandwidth. However, owing to the incompleteness of the 3-D manufacturing process, TSVs may exhibit some inherent defects; hence, switching and shifting repairing methods have been proposed. These methods repair the TSV faults by rerouting the signals of the faulty TSVs to the other signal TSVs or redundant TSVs using a simple repair algorithm. However, if one TSV exhibits a defect during its manufacturing process, the probability of multiple defects occurring in the TSVs neighboring the FTSV increases, i.e., the TSV defects tend to be clustered. Therefore, recently proposed repair solutions, such as ring/router-based repair architectures, have focused on clustered TSV faults. However, the implementation of these existing repair solutions for clustered faults involves an extremely high hardware overhead. This study proposes a TSV redundancy architecture to repair clustered TSV faults with a high repair rate and low hardware overhead. The proposed architecture divides the TSVs into several groups and connects the TSVs of each group using a 2:1 multiplexer chain. Simulation results show that the proposed architecture exhibits a repair rate of 98.52% for uniformly distributed faults and 69.86% for highly clustered faults. These repair rates are higher than those of other TSV redundancy architectures and the difference in the repair rate becomes even greater if the faults are more clustered. Moreover, the approach yields a 58.55% reduced area as compared to that of the router-based redundancy architecture, which also targets the repair of clustered faults.
Ingeol Lee, Minho Cheong, Sungho Kang 0001
IEEE Trans. Reliab.3
2019 Test-Friendly Data-Selectable Self-Gating (DSSG)
abstract
Low-power design is a key consideration in modern design. XOR self-gating (data-driven self-gating) is used for power reduction in clock networks, which is one of the main factors of dynamic power consumption. When applying XOR self-gating, dynamic power consumption is reduced, but the number of required test patterns on the testing side is inflated. In critical cases, more than three times the regular number of scan test patterns may be required for industrial designs, such as GPUs. In this brief, we propose a novel self-gating structure. Data-selectable self-gating (DSSG) is designed to use functional data and scan data selectively to eliminate the unnecessary clock toggling of flip-flops. With this structure, the self-gating function can be used in the scan test mode, as well as the function mode. When the self-gating logic is used during scan shift operations, the stuck-at faults in the self-gating logic can be tested with short test sequences; therefore, the rise in test costs can be mitigated. It is possible to test the stuck-at faults in self-gating logic using only four scan test patterns. The experimental results show that the average of the stuck-at test pattern increase ratio has been dropped from more than 90% to less than 8%. The low-power performance of the proposed method in the mission mode is the same as that of the conventional self-gating structures. When the DSSG method is used, the dynamic power of the shift operation which may increase excessively during the scan test can be reduced.
Jihye Kim 0002, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2019 Dynamic Built-In Redundancy Analysis for Memory Repair
abstract
As advances in memory density and capacity result in an increase in the probability of fault occurrence, many studies on built-in redundancy analysis (BIRA) have been conducted to address this problem. However, conventional BIRAs cannot directly find a final repair solution as soon as test sequences of the built-in self-test (BIST) are over, because they require starting the fault analyses after finishing the test sequences to achieve an optimal repair rate. For this reason, additional analysis time is inevitable, which affects total test costs. In this paper, a dynamic BIRA is proposed for memory repair. It can find a final repair solution directly as soon as test sequences if the BIST are over and achieve an optimal repair rate. The proposed BIRA can restore faults in fault-storing content-addressable memories whenever the spaces in them can be reduced via dynamic fault analysis. Furthermore, the proposed BIRA can be implemented with a reasonable hardware size. This is demonstrated via experiments.
Hayoung Lee, Donghyun Han, Seungtaek Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2018 3D Memory Formed of Unrepairable Memory Dice and Spare Layer
abstract
With the development of memory manufacturing technology, the density of memory die has been increased and more data can be stored in a small area than before. However, due to the complexity of the manufacturing process, faults in memory have increased. And it leads to poor yield and quality of memory. To improve yield and quality of the memory, the importance of memory test and repair is growing to maintain memory productivity. This paper presents solutions for test and repair in pre-bond. In the pre-bond, proposed method makes a new 3D stacked memory by using unrepairable memory dice which cannot be repaired with existing spare memories. Discard the bank with the largest number of faults in the unrepairable memory die and repair the remaining banks. The memory dice and a spare layer which made of the known good die or unrepairable memory die are stacked to create a 3D memory. A bank of the spare layer is mapped to discarded bank of unrepairable memory die to operate as one normal working memory die. The proposed method can lead to high yields of 3D stacked memory.
Donghyun Han, Hayoung Lee, Seungtaek Lee, Minho Moon, Sungho Kang 0001
TENCON5
2018 Neural Network Reliability Enhancement Approach Using Dropout Underutilization in GPU
abstract
Recently, the researches on DNN (deep neural network) using GPUs has been actively conducted. The reason for using GPUs in DNN is that it reduces the learning time by using many computational cores. However, GPUs have no implements to support the reliable computing operations. It can exacerbate the reliability of the deep neural network. To ensure the reliability of the deep neural network, the proposed approach is to utilize the dropout technique used in MLP (Multi-Layer Perceptron) learning. In case of the dropout method, some threads in GPUs do not participate in the calculation and it causes GPUs underutilization. The proposed method uses the GPUs underutilization to support a reliable deep neural network. The proposed approach is available through using the idle neurons with adjacent calculating neurons in the dropout process. The experiment results show that the proposed approach is able to support the reliability issues in GPUs while executing deep neural network algorithms.
Dongsu Lee, Hyunyul Lim, Sungho Kang 0001
TENCON4
2018 Test Resource Reused Debug Scheme to Reduce the Post-Silicon Debug Cost
abstract
In this paper, a design for debug (DFD) method that reuses test resources is proposed to reduce the debug cost in post-silicon validation. With the proposed method, the trace buffer is shared for embedded cores to capture the signatures of each core concurrently by reusing a test access mechanism. In this case, the depth of the trace buffer allocated to the core is reconfigurable and variable according to debug scheme. The experimental results indicate that the proposed DFD significantly reduces the debug time when the trace buffer is shared by cores in various debug cases.
Inhyuk Choi, Hyunggoy Oh, Sungho Kang 0001
IEEE Trans. Computers4
2018 Fault Group Pattern Matching With Efficient Early Termination for High-Speed Redundancy Analysis
abstract
Advances in memory density and capacity have had the consequence of increasing the probability of memory faults. For this reason, redundancy analysis (RA) and repair are used as effective solutions to improve memory yield. However, as the growth of the number of memory cells increases, it causes increase of the number of faulty cells and results in increase of difficulty of fault analysis. Although various RA methodologies have been proposed, most of them require a long analysis time or fast analysis speed without achieving a 100% normalized repair rate. Furthermore, research on conventional RA methodologies has not included effective early termination methods. Therefore, in this paper, fault group pattern matching (FGPM) is proposed for high speed RA with an effective early termination method. It can achieve very fast analysis with a 100% normalized repair rate. Additionally, it can finish the analysis rapidly by the proposed early termination method when a memory cannot be repaired. Experimental results demonstrate that the FGPM is highly effective in reducing analysis time with the achievement of a 100% normalized repair rate. In addition, the effectiveness of the proposed early termination is shown.
Hayoung Lee, Keewon Cho, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Thermal Aware Test Scheduling for NTV Circuit
abstract
Although the near threshold voltage (NTV) design has achieved energy efficiency, certain challenges remain regarding its application. In this paper, we describe the analysis of thermally induced reliability concern in test process. In an NTV environment, the thermal dependency of a circuit delay is changed, and a difference in thermal constraints from that in a nominal voltage design exists. In addition, we propose a new test scheduling method for NTV circuits that alleviates the thermal constraints in system-on-chip test processes. Our simulation results show that the test time could be reduced while minimizing the reliability loss.
Jaeil Lim, Hyunggoy Oh, Heetae Kim, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Fast Built-In Redundancy Analysis Based on Sequential Spare Line Allocation
abstract
Built-in redundancy analysis (BIRA) is widely used for memory yield improvement. However, increases in fault occurrence probability inevitably lead to the use of various spare lines to achieve a high repair rate. Generally, it is difficult to apply conventional BIRAs for memories with various spare lines because they focus on a simple spare structure. Therefore, this study examines a BIRA that focuses on a various spare lines structure. The proposed BIRA achieves a high repair rate through the use of various spare lines. Although long analysis time is typically required due to the use of various spare lines, the proposed BIRA solves the problem through sequential spare line allocation. Additionally, it achieves hardware overhead reduction through a simple analyzer. These advantages of the proposed BIRA are demonstrated experimentally.
Hayoung Lee, Keewon Cho, Sungho Kang 0001
IEEE Trans. Reliab.4
2018 An Area-Efficient BIRA With 1-D Spare Segments
abstract
The growing capacity and density of embedded memories increases the probability of defects and affects the yield. To improve the yield, built-in redundancy analysis (BIRA) has been developed to replace faulty cells with healthy redundant cells. BIRA requires a high repair rate and a feasible hardware size for implementation. Although many BIRAs have been proposed, most of them still demonstrate a low repair rate or a large required hardware size. The proposed BIRA employs an intuitive algorithm with a small-area analyzer that uses 1-D spare segments in the 2-D spare structure. Because most faults in the memory are single faults, spare segments can be used to efficiently allocate redundancies. In terms of the yield, 1-D spare segments are effective when used with an intuitive algorithm that can be implemented with a small hardware overhead. Experimental results show that the proposed BIRA has a higher repair rate and relatively low hardware overhead than state-of-the-art BIRAs and has the advantages of 1-D spare segments.
Hayoung Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2017 DRAM-Based Error Detection Method to Reduce the Post-Silicon Debug Time for Multiple Identical Cores
abstract
In the post-silicon debug of multicore designs, the debug time has increased significantly because the number of cores undergoing debug has increased; however the resources available to debug the design are limited. This paper proposes a new DRAM-based error detection method to overcome this challenge. The proposed method requires only three debug sessions even if multiple cores are present. The first debug session is used to detect the error intervals of each core using golden signatures. The second session is used to detect the error clock cycles in each core using a golden data stream. Instead of storing all of the golden data, the golden data stream is generated by selecting error-free debug data for each interval which are guaranteed by the first session. Finally, the error data in all cores are only captured during the third session. The experimental results on various debug cases show significant reductions in total debug time and the amount of DRAM usage compared to previous methods.
Hyunggoy Oh, Inhyuk Choi, Sungho Kang 0001
IEEE Trans. Computers3
2017 An On-Chip Error Detection Method to Reduce the Post-Silicon Debug Time
abstract
Debug time has become a major issue in post silicon debug because of the increasingly complicated nature of circuit design. However, reducing debug time is a major challenge because of the limited size of the trace buffer used to observe internal signals in the circuit. This study proposes an on-chip error detection method to overcome this challenge. The on-chip process detects the error-suspect window using the pre-calculated golden data stored in the trace buffer. This allows the selective compaction and capture of the debug data in the trace buffer during the error-containing interval. As a result, reducing the number of debug sessions significantly reduces the total debug time. The experimental results on various debug cases show significant reductions in total debug time compared to previous work.
Hyunggoy Oh, Taewoo Han 0001, Inhyuk Choi, Sungho Kang 0001
IEEE Trans. Computers4
2017 Grouping-Based TSV Test Architecture for Resistive Open and Bridge Defects in 3-D-ICs
abstract
After the 3-D stacking, 3-D-ICs based on through-silicon-vias (TSVs) must be inspected for any TSV defects such as resistive open or bridge defects. In some research studies, several effective testing techniques have been developed such as parallel or serial test architectures, which measure the voltage across a single TSV with a comparator. However, in the current test architectures, hardware overhead and test time are proportional to the number of TSVs. In this paper, we propose a new unified test architecture for screening of TSV defects in 3-D-ICs. Depending on the number of assembled TSVs, the proposed grouping-based test architecture can effectively reduce the cumulative test time and hardware overhead without compromising the test quality.
Hyeonchan Lim, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 FRESH: A New Test Result Extraction Scheme for Fast TSV Tests
abstract
Three-dimensional integrated circuits (3-D ICs) are considered to meet the performance needs of future ICs. The core components of 3-D ICs are through-silicon vias (TSVs), which should pass appropriate prebond and post-bond tests in 3-D IC fabrication processes. The test inputs must be injected into the TSVs, and the test results must be extracted. This paper proposes a new test result extraction scheme [fast result extraction by selective shift-out (FRESH)] for prebond and post-bond TSV testing. With additional hardware, the proposed scheme remarkably reduces the TSV test time. FRESH avoids unnecessary test result extraction when the number of faulty TSVs in the TSV set is 0 or exceeds the number of TSV redundancies in the set. These early fault analyses are executed in the checkers of TSV groups. The experimental results show that the proposed scheme can reduce the result extraction time in practical environments.
Jaeseok Park, Hyunyul Lim, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 R2-TSV: A Repairable and Reliable TSV Set Structure Reutilizing Redundancies
abstract
Recently, three-dimensional integrated circuit (3-D IC) design has attracted much attention, and the reliability of these systems has become increasingly important. In this paper, a new repairable and reliable through-silicon via (TSV) set structure is proposed. This proposed TSV set structure can be applied to the previous TSV repair structures which require TSV redundancies, and detects a defect or error reutilizing residual TSV redundancies for high reliability of 3-D ICs. Both online test and soft error detection/analysis are supported by the proposed approach. Furthermore, a redundancy-sharing structure is introduced to guarantee a balanced detection rate among TSV sets and a reasonable full detection rate. The experimental results prove that the new approach guarantees high redundancy utilization efficiency and reliability of TSV. Also, they show that defect or error detection is achieved by the proposed TSV set structure.
Jaeseok Park, Minho Cheong, Sungho Kang 0001
IEEE Trans. Reliab.3
2017 Chain-Based Approach for Fast Through-Silicon-Via Coupling Delay Estimation
abstract
A chain-based coupling delay estimation method for through-silicon-vias (TSVs) in 3-D integrated circuits is proposed. Existing works target the worst case scenarios and this leads to inaccurate TSV coupling delay estimations, as the worst case may not occur during normal operation. The proposed method calculates the TSV coupling delay using simulation-based switching data. In addition, our TSV chain method allows us to capture the effects of nonneighboring TSVs accurately. Our simulations show that the error introduced by our method without using HSPICE is less than 10 ps even in TSV-crowded regions.
Jaewon Jang, Minho Cheong, Jin-Ho Ahn, Sung Kyu Lim, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Hardware-Efficient Built-In Redundancy Analysis for Memory With Various Spares
abstract
Memory capacity continues to increase, and many semiconductor manufacturing companies are trying to stack memory dice for larger memory capacities. Therefore, built-in redundancy analysis (BIRA) is of utmost importance because the probability of fault occurrence increases with a larger memory capacity. A traditional spare structure that consists of simple rows and columns is somewhat inadequate for multiple memory blocks BIRA because the hardware overhead and spare allocation efficiency are degraded. The proposed BIRA uses various types of spares and can achieve a higher yield than a simple row and column spare structure. Herein, we propose a BIRA that can achieve an optimal repair rate using various spare types. The proposed analyzer can exhaustively search not only row and column spare types but also global and local spare types. In addition, this paper proposes a fault-storing content-addressable memory (CAM) structure. The proposed CAM is small and collects faults efficiently. The experimental results show a high repair rate with a small hardware overhead and a short analysis time.
Woosung Lee, Keewon Cho, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Parallelized Network-on-Chip-Reused Test Access Mechanism for Multiple Identical Cores
abstract
This paper proposes a new network-on-chip (NoC)-reused test access mechanism (TAM) for testing multiple identical cores. It can test multiple cores concurrently and identify faulty cores to derate the chip by excluding the core. In order to minimize the test time, the TAM utilizes the majority value of test response data. All of the cores can thereby be tested in parallel and test costs (in both test pins and test time) are exactly the same as those for a single core. The hardware overhead is minimized by reusing the NoC infrastructures and transfer-counters are designed as a majority analyzer. The experimental results in this paper show that the proposed TAM can test multiple cores in the same time as a single core and with negligible hardware overhead.
Taewoo Han 0001, Inhyuk Choi, Hyunggoy Oh, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 A New 3-D Fuse Architecture to Improve Yield of 3-D Memories
abstract
A new 3-D fuse architecture is proposed to improve the yield of 3-D memories. Because the 2-D memories are stacked to form a 3-D memory, the repair status of the prebond is kept as the good status. However, if faults occur in the postbond on the same cells which were repaired in the prebond, they must be identified and repaired because they cannot be repaired by the previous methods. There is no research on the repair the same faulty cells which occur in the prebond and postbond yet. Therefore, the new 3-D fuse architecture is proposed to repair the faulty cells which occur in the prebond and postbond. The redundancies which repair the faulty cells in the prebond are invalidated. The faulty cells are repaired by other redundancies in the postbond by the proposed 3-D fuse architecture. Thus, the proposed technique can improve the yield of 3-D memories. The experimental results show that the proposed technique can achieve higher yields of 3-D memories because only the proposed technique can repair the same faulty cells occurring in the prebond and postbond, and verify the good repair status of the 3-D memories.
Wooheon Kang, Changwook Lee, Hyunyul Lim, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 Tri-State Coding Using Reconfiguration of Twisted Ring Counter for Test Data Compression
abstract
As technology processes scale up and design complexities grow, system-on-chip integration continues to rise rapidly. According to these trends, increasing test data volume is one of the biggest challenges in the testing industry. In this paper, we present a new test data compression method based on reusing a stored set with tri-state coding (TSC). For improving the compression efficiency, a twisted ring counter is used to reconfigure twist function. It is useful to reuse previously used data for making next data by using the function of feedback of the ring counter. Moreover, the TSC is used to increase the range information transmission without additional input ports. Experimental results show that this compression method improves a compression ratio and a test time on both International Symposium on Circuits and Systems'89 and large International Test Conference'99 benchmark circuits in most cases compared to the results of the previous work without a heavy burden on the hardware.
Sungyoul Seo, Yong Lee 0002, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Optimized Built-In Self-Repair for Multiple Memories
abstract
A new built-in self-repair (BISR) scheme is proposed for multiple embedded memories to find optimum point of the performance of BISR for multiple embedded memories. All memories are concurrently tested by the small dedicated built-in self-test to figure out the faulty memories, the number of faults, and irreparability. After all memories are tested, only faulty memories are serially tested and repaired by the shared built-in redundancy analysis according to the sizes of memories in descending order. Thus, the fast test and repair are performed with low area overhead. To accomplish an optimal repair rate and a fast analysis speed, an exhaustive search for all combinations of spare rows and columns is proposed based on the optimized fault collection. Experimental results show that the proposed BISR has the optimal repair rate because of the exhaustive search. The performance of the proposed BISR is located in the optimum point between the test and repair time, and the area overhead. For example, the proposed BISR requires 49.6% of the area and 1.3 times of the test and repair time in comparison with parallel BISR scheme for four memories (one 128 K, two 256 K, and one 512 K memories). Furthermore, the more there are memories, the more superior performance in terms of the test and repair time, and the area overhead is shown.
Wooheon Kang, Changwook Lee, Hyunyul Lim, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Scan Chain Reordering-Aware X-Filling and Stitching for Scan Shift Power Reduction
abstract
As a scan-based testing enables higher test coverage and faster test time than alternative ways, it is widely used by most system-on-chip (SoC) designers. However, since the number of logic gates is over one hundred million gates, a number of scan cells lead to excessive power consumption and it produces a low shifting frequency during the scan shifting mode. In this paper, we present a new scan shift power reduction method based on a scan chain reordering (SR)-aware X-filling and a stitching method. There is no need to require an additional logic for reducing the scan shift power, just a little routing overhead. Experimental results show that this method improves scan shift power consumption on benchmark circuits in most cases compared to the results of the previous works.
Sungyoul Seo, Yong Lee 0002, Hyeonchan Lim, Joohwan Lee, Hongbom Yoo, Yojoung Kim, Sungho Kang 0001
ATS7
2015 3-D Stacked DRAM Refresh Management With Guaranteed Data Reliability
abstract
The 3-D integrated dynamic random-access memory (DRAM) structure with a processor is being widely studied due to advantages, such as a large band-width and data communication power reduction. In these structures, the massive heat generation of the processor results in a high operating temperature and a high refresh rate of the DRAM. Thus, in the 3-D DRAM over processor architecture, temperature-aware refresh management is necessary. However, temperature determination is difficult, because in the 3-D DRAM, the temperature changes dynamically and temperature variation in a DRAM die is complicated. In this paper, a thermal guard-band set-up method for 3-D stacked DRAM is proposed. It considers the latency of the temperature data and the position difference between the temperature sensor and the DRAM cell. With this method, the data reliability of the on-chip temperature sensor-dependent adaptive refresh control is guaranteed. In addition, an efficient temperature sensor built-in and refresh control method is analyzed. The expected refresh power reduction is examined through a simulation.
Jaeil Lim, Hyunyul Lim, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2015 A 3 Dimensional Built-In Self-Repair Scheme for Yield Improvement of 3 Dimensional Memories
abstract
A 3-dimensional Built-In Self-Repair (3D BISR) scheme is proposed for 3-dimensional (3D) memories. The proposed 3D BISR scheme consists of two phases: a parallel test-repair phase, and a serial test-repair phase. After all memory dice are simultaneously tested, only the faulty memory dice are serially tested and repaired using one Built-In Redundancy Analysis (BIRA) module. Thus, it is a faster test-repair with low area overhead. The proposed BIRA algorithm with a post-share redundancy scheme performs exhaustive searches for all combinations of spare rows and columns. Experimental results show that the proposed 3D BISR is up to two times faster than the 3D serial test-serial repair BISR when seven 2048 × 2048 bit memory dice are stacked. The proposed 3D BISR requires 44.55% of the area in comparison to a 3D parallel test-parallel repair BISR for four stacked memory dice (one 128 K RAM, two 256 K RAMs, and 512 K RAM). The yield of 3D memories is the highest due to the exhaustive search BIRA algorithm with the post-share redundancy scheme as shown in various experimental results.
Wooheon Kang, Changwook Lee, Hyunyul Lim, Sungho Kang 0001
IEEE Trans. Reliab.4
2015 A Novel Massively Parallel Testing Method Using Multi-Root for High Reliability
abstract
Wafer testing (wafer sort) is used in the semiconductor industry to reduce test costs. High parallelism is important to reduce test application time. However, increasing parallelism is becoming more difficult because the elements that drive test costs are increases in the pin count of the system on chip (SOC), and required automated test equipment (ATE) channels. While the need for parallelism has been growing, a reliability problem in which fault distribution causes good devices under test (DUT) to be improperly tested is becoming a concern. To achieve high parallelism and reliability, we propose a novel, massively parallel testing method using multi-root. In addition, we develop a test setup for setting the root-DUT location and the path of all DUTs. Using the proposed wafer testing method, test data can be transferred between the ATE and dies using multi-roots. The experimental results using ITC 02 SOC benchmarks show that the proposed method can reduce test costs up to 90%, and achieve nearly 94.84% test reliability without affecting yield.
Haksong Kim, Yong Lee 0002, Sungho Kang 0001
IEEE Trans. Reliab.3
2015 Majority-Based Test Access Mechanism for Parallel Testing of Multiple Identical Cores
abstract
The increased use of multicore chips diminishes per-core complexity and also demands parallel design and test technologies. An especially important evolution of the multicore chip has been the use of multiple identical cores, providing a homogenous system with various merits. This paper introduces a novel test access mechanism (TAM) for parallel testing of multiple identical cores and identifying faulty cores to derate the chip by excluding it. Instead of typical test response data from the cores, the test output data used in this paper are the majority values, that is, the typical test responses from the cores. All the cores can thereby be tested in parallel and test costs (in both test pins and test time) are exactly the same as for a single core. The proposed TAM can be implemented with on-chip comparators and majority analyzers. The experimental results in this paper show that the proposed TAM can test multiple cores with minimal test pins and test time and with hardware overhead of <;0.1%.
Taewoo Han 0001, Inhyuk Choi, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2014 A Scalable and Parallel Test Access Strategy for NoC-Based Multicore System
abstract
This paper proposes a new parallel test access strategy for multiple identical cores in a network-on-chip (NoC). The proposed test strategy takes advantage of the regular design of NoC to reduce both test area overhead and test time. The proposed NoC reused test access mechanism (TAM) adopted a pipelining structure and a deterministic test data routing algorithm in order to reuse the full bandwidth of links in the NoC. Also, the architecture has complete scalability according to the number of cores and applications for 3D environment are also represented. Experimental results show that the proposed TAM can test multiple cores with the same test time as a single core and negligible hardware overhead.
Taewoo Han 0001, Inhyuk Choi, Hyunggoy Oh, Sungho Kang 0001
ATS4
2014 A New Fuse Architecture and a New Post-Share Redundancy Scheme for Yield Enhancement in 3-D-Stacked Memories
abstract
3-D-stacked memory using through-silicon-vias (TSVs) has emerged as a good alternative for overcoming the limitation of 2-D memory technology. Among many issues with 3-D-stacked memory, yield is one of the major challenges for mass production. This paper proposes a new fuse architecture and redundancy scheme to improve the yield of 3-D-stacked memories. The new fuse architecture is developed based on the fact that the unused redundancies in prebond repair cause the inefficiency. Therefore, the new fuse architecture provides a way to share redundancies in prebond and postbond repairs. There are two kinds of operation modes. One is an enable mode for collecting the used redundancy information. The other is a mask mode for obtaining faulty redundancy information using a short test algorithm. Using the new fuse architecture, a new redundancy scheme called the post-share scheme is developed to achieve optimal yield. The post-share scheme allocates the fixed number of spare rows and columns for each repair just like other schemes. However, only allocated redundancies are used in prebond repair, while both the redundancies allocated for postbond repair and unused redundancies in prebond repair can be used for postbond repair. Experimental results show that the post-share redundancy scheme significantly increases the final yield of 3-D-stacked memories and the increase of area overhead is small.
Changwook Lee, Wooheon Kang, Donkoo Cho, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2014 A BIRA for Memories With an Optimal Repair Rate Using Spare Memories for Area Reduction
abstract
The test cost and yield improvement of embedded memories have become very important as memory capacity and density have grown. For embedded memories, built-in redundancy analysis (BIRA) is widely used to improve yield by replacing faulty cells with a 2-D redundancy architecture. However, the most important factor in BIRA is the reduction of hardware overhead while keeping optimal repair rate. Most BIRA approaches require extra hardware overhead in order to store and analyze faults in the memory. These approaches do not utilize spare memories during the redundancy analysis (RA) procedure. However, the proposed BIRA minimizes area overhead by utilizing a part of the spare memory as an address mapping table (AMT). Since storing the faulty memory addresses take most of the extra hardware overhead, the reduced logical addresses produced by the AMT are used to reduce the extra hardware overhead. In addition, the reduced addresses are stored in content-addressable memories (CAMs) and used in the RA procedure. The proposed BIRA can achieve an optimal repair rate by using an exhaustive search RA algorithm. The proposed RA algorithm compares the repair solution candidates with all the fault addresses stored in the proposed CAMs to guarantee an exhaustive search. The experimental results show that the proposed BIRA requires a smaller area overhead than that of the previous state-of-the-art BIRA with an optimal repair rate.
Wooheon Kang, Hyungjun Cho, Joohwan Lee, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Interleaving Test Algorithm for Subthreshold Leakage-Current Defects in DRAM Considering the Equal Bit Line Stress
abstract
Since the minimum feature size of dynamic RAM has been down-scaled, several studies have been carried out to determine ways to protect cell data from leakage current in many areas. In the field of testing, more appropriate test algorithms are required to detect weak cells with leakage-current sources. In this paper, we propose an interleaving test algorithm that takes into account the equal bit-line stress regardless of the cell location. The proposed test algorithm allows screening of weak cells that cannot hold cell data due to the subthreshold leakage current. During the stress period, the algorithm can also detect other leakage currents. This paper presents the maximum stress differences according to the cell location, and determines the influence of the refresh operation on the maximum stress time. Therefore, this paper suggests a correlation between the refresh and read time to give maximum stress time.
Hyoyoung Shin, Youngkyu Park, Gihwa Lee, Jungsik Park, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2014 A Delay Test Architecture for TSV With Resistive Open Defects in 3-D Stacked Memories
abstract
The limits of technology scaling for smaller chip size, higher performance, and lower power consumption are being reached. For this reason, the memory semiconductor industry is searching for new technology. 3-D stacked memory using through-silicon via (TSV) has been considered as a promising solution for overcoming this challenge. However, to guarantee quality and yield for mass production of 3-D stacked memories, effective test techniques for TSV are required. In this paper, a new test architecture for testing TSVs in 3-D stacked memories is proposed. By comparing voltage changes generated due to resistive open defects with a reference voltage applied externally, the test circuit estimates delay across the TSV. This allows the possibility of a delay test with low-frequency test equipment. Experimental results demonstrate that the proposed test architecture can be effective in the testing of TSV with resistive open defects, and have lower area overhead and lower peak current consumption.
Hyungsu Sung, Keewon Cho, Kunsang Yoon, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2013 A Die Selection and Matching Method with Two Stages for Yield Enhancement of 3-D Memories
abstract
Three-dimensional (3-D) memories using through-silicon-vias (TSVs) as vertical buses across memory layers has regarded as one of 3-D integrated circuits (ICs) technology. The memory dies to stack together in a 3-D memory are selected by a die selection method. In order to improve yield of 3-D memories, redundancy sharing between inter-die using TSVs is an effective strategy. With the redundancy sharing strategy, the bad memory dies can become good 3-D memories through matching the good memory dies. To support die selection and matching efficiently, a novel redundancy analysis (RA) algorithm, which considers various repair solutions, is proposed. Because the repair solutions can be various, the proposed die selection and matching is performed with two stages; general die selection and matching method in the first stage and re-matched remained memory dies, after the first stage, applying other repair solutions in the second stage. Thus, the proposed die selection and matching algorithm using the proposed RA algorithm can improve yield of 3-D memories. The experimental results show that the proposed die selection and matching method can achieve higher yield of 3-D memories than that of the previous state-of-the-art the die selection and matching methods.
Wooheon Kang, Changwook Lee, Keewon Cho, Sungho Kang 0001
Asian Test Symposium4
2012 An efficient IP address lookup algorithm based on a small balanced tree using entry reduction
Hyuntae Park, Hyejeong Hong, Sungho Kang 0001
Comput. Networks3
2011 Communication-aware task scheduling and voltage selection for total energy minimization in a multiprocessor system using Ant Colony Optimization
Sungho Kang 0001
Inf. Sci.2
2011 A Lossless Color Image Compression Architecture Using a Parallel Golomb-Rice Hardware CODEC
abstract
In this paper, a high performance lossless color image compression and decompression architecture to reduce both memory requirement and bandwidth is proposed. The proposed architecture consists of differential-differential pulse coded modulation (DDPCM) and Golomb-Rice coding. The original image frame is organized as m by n sub-window arrays, to which DDPCM is applied to produce one seed and m × n - 1 pieces of differential data. Then the differential data are encoded using the Golomb-Rice algorithm to produce losslessly compressed data. According to the experimental results on benchmark images, the proposed architecture can guarantee high enough compression rate and throughput to perform real-time lossless CODEC operations with a reasonable hardware area.
Hong-Sik Kim, Joohong Lee, Sungho Kang 0001, Woo-Chan Park
IEEE Trans. Circuits Syst. Video Technol.4
2011 A Memory-Efficient Bit-Split Parallel String Matching Using Pattern Dividing for Intrusion Detection Systems
abstract
For the low-cost hardware-based intrusion detection systems, this paper proposes a memory-efficient parallel string matching scheme. In order to reduce the number of state transitions, the finite state machine tiles in a string matcher adopt bit-level input symbols. Long target patterns are divided into subpatterns with a fixed length; deterministic finite automata are built with the subpatterns. Using the pattern dividing, the variety of target pattern lengths can be mitigated, so that memory usage in homogeneous string matchers can be efficient. In order to identify each original long pattern being divided, a two-stage sequential matching scheme is proposed for the successive matches with subpatterns. Experimental results show that total memory requirements decrease on average by 47.8 percent and 62.8 percent for Snort and ClamAV rule sets, in comparison with several existing bit-split string matching methods.
Hong-Sik Kim, Sungho Kang 0001
IEEE Trans. Parallel Distributed Syst.3
2010 An Advanced BIRA for Memories With an Optimal Repair Rate and Fast Analysis Speed by Using a Branch Analyzer
abstract
As memory capacity and density grow, a corresponding increase in the number of defects decreases the yield and quality of embedded memories for systems-on-chip as well as commodity memories. For embedded memories, built-in redundancy analysis (BIRA) is widely used to solve quality and yield issues by replacing faulty cells with healthy redundant cells. Many BIRA approaches require extra hardware overhead in order to achieve optimal repair rates, or they suffer a loss of repair rate in minimizing the hardware overhead. An innovative BIRA approach is proposed to achieve optimal repair rates, lower area overhead, and increase analysis speed. The proposed BIRA minimizes area overhead by eliminating some storage coverage for only must-repair faulty information. The proposed BIRA analyzes redundancies quickly and efficiently by evaluating all nodes of a branch in parallel with a new analyzer which is simple and easy-to-implement. Experimental results show that the proposed BIRA allows for a much faster analysis speed than that of the state-of-the-art BIRA, as well as the optimal repair rate, and relatively small area overhead.
Woosik Jeong, Joohwan Lee, Taewoo Han 0001, Kaangchil Lee, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2010 EOF: Efficient Built-In Redundancy Analysis Methodology With Optimal Repair Rate
abstract
Faulty cell repair with redundancy can improve memory yield. In particular, built-in redundancy analysis (BIRA) is widely used to enhance the yield of embedded memories. We propose an efficient BIRA algorithm to achieve the optimal repair rate with a very short analysis time and low hardware cost. The proposed algorithm can significantly reduce the number of backtracks in the exhaustive search algorithm: it uses early termination based on the number of orthogonal faulty cells and fault classification in fault collection. Experimental results show that the proposed BIRA methodology can achieve optimal repair rate with low hardware overhead and short analysis time, as compared to previous BIRA methods.
Myung-Hoon Yang, Hyungjun Cho, Wooheon Kang, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 A High-Level Signal Integrity Fault Model and Test Methodology for Long On-Chip Interconnections
abstract
In this paper, considering the interconnection topology information, an abstract model and a new test pattern generation method of signal integrity problems on interconnects are proposed. In addition, previous SPICE-based pattern generation methods are too complex and time consuming to generate test patterns for signal integrity faults. To more accurately detect signal integrity defects on practical on-chip interconnection lines and avoid time consuming for interconnection analysis, in this paper, we propose a new high-level signal integrity fault model to estimate noise effects based on process variation and interconnect signal transition. Experimental results show that the proposed signal integrity fault model is more exact for long interconnects than previous approaches. In addition, the proposed method is much faster than the SPICE-based pattern generation method.
Sunghoon Chun, YongJoon Kim, Sungho Kang 0001
VTS4
2008 XPDF-ATPG: An Efficient Test Pattern Generation for Crosstalk-Induced Faults
abstract
In this paper, we propose a new test generation method for delay faults considering crosstalk-induced delay effects, based on a conventional delay ATPG technique in order to reduce the complexity of previous ATPG algorithm for crosstalk delay faults and to consider multiple aggressor crosstalk faults to maximize the noise of the victim line. Since the proposed ATPG for crosstalk-induced delay faults uses the physical and timing information, the proposed ATPG can reduce the search space of the backward implication of the aggressor's constraints and it is helpful for reducing the time cost of the ATPG than previous works. In addition, since the proposed technique targets on the critical path for the original delay test as the victim lines, it can improve test effectiveness of delay testing. Experimental results demonstrate the effectiveness of the proposed method.
Sunghoon Chun, YongJoon Kim, Myung-Hoon Yang, Sungho Kang 0001
ATS5
2008 An Effective Hybrid Test Data Compression Method Using Scan Chain Compaction and Dictionary-Based Scheme
abstract
In this paper, we propose a new test data compression method for reducing test data volume and test application time. The proposed method consists of two steps: scan chain compaction and dictionary-based compression scheme. The scan chain compaction provides a minimum scan chain depth by using compaction of the compatible scan cells in the scan chain. The compacted scan chain is partitioned to the multiple internal scan chains for using the fixed-length index dictionary-based compression scheme that provides the high compression ratio and the fast testing time. The proposed compression method delivers compressed patterns from the ATE to the chip and drives a large number of multiple internal scan chains using only a single ATE input and output. Experimental results for the ISCAS-89 test benches show that the test data volume and testing time for the proposed method are less than previous compression schemes.
Sunghoon Chun, YongJoon Kim, Myung-Hoon Yang, Sungho Kang 0001
ATS5
2008 An Efficient Scan Chain Diagnosis Method Using a New Symbolic Simulation
abstract
Locating the scan chain faults is very important for dedicated IC manufacturers to guide the failure analysis process for yield improvement. In this paper, we propose a new symbolic simulation based scan chain diagnosis method to solve the scan chain diagnosis resolution problem as well as the multiple faults problem. The proposed method uses a new symbolic simulation with the faulty probabilities of a set of candidate faulty scan cells in a bounded range and to analyze the effects caused by faulty scan cells in good scan chains. In addition, we use the faulty information in good scan chains that are not contaminated by the faults while unloading scan out responses. In addition, a new score matching method is proposed to effectively handle multiple faults and to improve the diagnostic resolution by ranking the candidate scan cells in the candidate list. Experimental results demonstrate the effectiveness of the proposed method.
Sunghoon Chun, YongJoon Kim, Sungho Kang 0001
VTS4
2008 A New Scan Architecture for Both Low Power Testing and Test Volume Compression Under SOC Test Environment
Hong-Sik Kim, Sungho Kang 0001, Michael S. Hsiao
J. Electron. Test.2
2008 An Effective Power Reduction Methodology for Deterministic BIST Using Auxiliary LFSR
Myung-Hoon Yang, YongJoon Kim, Sunghoon Chun, Sungho Kang 0001
J. Electron. Test.4
2008 Total Energy Minimization of Real-Time Tasks in an On-Chip Multiprocessor Using Dynamic Voltage Scaling Efficiency Metric
abstract
This paper proposes an algorithm that provides both dynamic voltage scaling and power shutdown to minimize the total energy consumption of an application executed on an on-chip multiprocessor. The proposed algorithm provides an extended schedule and stretch method, where task computations are iteratively stretched within the slack of a time-constrained dependent task set. In addition, the break-even threshold interval for amortizing the shutdown overhead is considered. By evaluating each set of stretched task computations, an energy-efficient set is obtained. The proposed dynamic voltage scaling efficiency metric is the ratio of the reduced energy to the increased cycle time when the supply voltage is scaled, which can be used to determine the task computation cycle to be stretched. Experimental results show that the proposed algorithm outperforms the traditional schedule and stretch method in the various evaluations of target real applications.
Hyejeong Hong, Hong-Sik Kim, Jin-Ho Ahn, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2007 High-MDSI: A High-level Signal Integrity Fault Test Pattern Generation Method for Interconnects
abstract
Unacceptable loss of signal integrity may cause permanent or intermittent harm to the functionality and performance of SoCs. In this paper, considering the interconnection topology information, an abstract model and a new test pattern generation method of signal integrity problems on interconnects are proposed. In addition, previous SPICE- based pattern generation methods are too complex and time consuming to generate test patterns for signal integrity faults. To overcome this problem, we also develop a new high-level test pattern generation method by using the abstract signal integrity fault model. Experimental results show that the proposed signal integrity fault model is more exact for long interconnects than previous approaches. In addition, the proposed method is much faster than the SPICE-based pattern generation method.
Sunghoon Chun, YongJoon Kim, Sungho Kang 0001
ATS3
2007 MDSI: Signal Integrity Interconnect Fault Modeling and Testing for SoCs
Sunghoon Chun, YongJoon Kim, Sungho Kang 0001
J. Electron. Test.3
2006 An Effective Test Pattern Generation for Testing Signal Integrity
abstract
As more cores are integrated in a single chip with sophisticated process like nano technology, testing signal integrity between the cores needs much effort due to complicate coupling effects. In this paper, we propose a novel test pattern generation method for testing signal integrity. Using this method, short and effective test patterns are generated with low hardware overhead. It can be used for self-test scheme and experimental results show the effectiveness of the proposed scheme.
YongJoon Kim, Myung-Hoon Yang, Youngkyu Park, Daeyeal Lee, Sungho Kang 0001
ATS5
2006 SoC Test Scheduling Algorithm Using ACO-Based Rectangle Packing
Jin-Ho Ahn, Sungho Kang 0001
ICIC (2)2
2006 System on a Chip Implementation of Social Insect Behavior for Adaptive Network Routing
Jin-Ho Ahn, Byung In Moon, Sungho Kang 0001
ICIC (2)4
2006 An Efficient Dictionary Organization for Maximum Diagnosis
Sunghoon Chun, Hong-Sik Kim, Sungho Kang 0001
J. Electron. Test.4
2006 Increasing encoding efficiency of LFSR reseeding-based test compression
abstract
A new methodology to increase the encoding efficiency of test compression based on linear feedback shift registers (LFSRs) is proposed. The proposed method combines LFSR reseeding and bit fixing. Deterministic test patterns tend to have a biased probability of the logic value 1 or 0 at each primary input. If such biased inputs are fixed to the logic value 1 or 0 with some combinational logic, then the amount of data to be encoded by the LFSR will be considerably reduced. Additionally, in order to reduce the encoded data volume much further, a variable degree of the LFSR polynomial is employed. In the variable-degree LFSR scheme, a test cube with less specified bits is encoded with an LFSR polynomial of lower degree, while a test cube with more specified bits is encoded with an LFSR polynomial of higher degree. Experimental results for the larger ISCAS 89 benchmark circuits show that the proposed scheme can increase the encoding efficiency with little hardware overhead compared to previous schemes.
Hong-Sik Kim, Sungho Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 MICRO: a new hybrid test data compression/decompression scheme
abstract
To overcome the limitation of the automatic test equipment (ATE), test data compression/decompression schemes become a more important issue of testing for a system-on-chip (SoC). In order to alleviate the limitation of previous works, a new hybrid test data compression/decompression scheme for an SoC is developed. The new scheme is based on analyzing the factors that influence test parameters: compression ratio and hardware overhead. To improve compression ratio, the proposed scheme, called the Modified Input reduction and CompRessing One block (MICRO), uses the modified input reduction, the one block compression, a novel mapping, and reordering algorithms. Unlike previous approaches using the cyclic scan register architecture, the proposed scheme is to compress original test data and to decompress the compressed test data without the cyclic scan register architecture. Therefore, the proposed scheme leads to high-compression ratio with low-hardware overhead. Experimental results on ISCAS '89 and ITC '99 benchmark circuits prove the efficiency of the new method.
Sunghoon Chun, YongJoon Kim, Jung-Been Im, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2005 A New Low Power Test Pattern Generator using a Transition Monitoring Window based on BIST Architecture
abstract
This paper presents a new low power BIST TPG scheme. It uses a transition monitoring window (TMW) that is comprised of a transition monitoring window block and a MUX. When random test patterns are generated by an LFSR, transitions of those patterns satisfy pseudo-random Gaussian distribution. The proposed technique represses transitions of patterns using the k-value which is a standard that is obtained from the distribution of TMW to observe over transitive patterns causing high power dissipation in a scan chain. Experimental results show that the proposed BIST TPG schemes can reduce scan transition by about 60% without performance loss in ISCAS’89 benchmark circuits that have large number scan inputs.
Youbean Kim, Myung-Hoon Yang, Yong Lee 0002, Sungho Kang 0001
Asian Test Symposium4
2004 RAIN (RAndom Insertion) Scheduling Algorithm for SoC Test
abstract
This paper presents a new SoC (system-on-chip) test scheduling algorithm. Reducing the test application time is an important issue for a core-based SoC test. In this paper, each core is represented by a rectangle, the height of which is equal to the TAM width and the width of which is equal to the test time. A 'one-element-exchange' algorithm is used for optimizing the test time of each core and the 'RAIN' scheduling algorithm is used for optimizing the test time of SoC. The RAIN scheduling algorithm uses a sequence pair data structure to represent the placement of rectangles, and obtains the optimized results by inserting into a random position on the sequence pair. The results of the experiments conducted using ITC '02 SoC benchmarks show that the proposed algorithm gives the shortest test application time compared with earlier researches for most of the cases.
Jung-Been Im, Sunghoon Chun, Jin-Ho Ahn, Sungho Kang 0001
Asian Test Symposium5
2004 An In-Order SMT Architecture with Static Resource Partitioning for Consumer Applications
Byung In Moon, Hongil Yoon, Ilgu Yun, Sungho Kang 0001
PDCAT4
2004 A new maximal diagnosis algorithm for interconnect test
abstract
Interconnect test for highly integrated environments becomes more important in terms of its test time and a complete diagnosis, as the complexity of the circuit increases. Since the board-level interconnect test is based on boundary scan technology, it takes a long test time to apply test vectors serially through a long scan chain. Complete diagnosis is another important issue. Since the board-level test is performed for repair, noticing the faulty position is an essential element of any interconnect test. Generally, the interconnect test algorithms that need a short test time cannot perform the complete diagnosis and the algorithms that perform the complete diagnosis need a lengthy test time. To overcome this problem, a new interconnect test algorithm is developed. The new algorithm can provide the complete diagnosis of all faults with a shorter test time compared to the previous algorithms.
YongJoon Kim, Hyun-Don Kim, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2003 A New Maximal Diagnosis Algorithm for Bus-structured Systems
YongJoon Kim, DongSub Song, YongSeung Shin, Sunghoon Chun, Sungho Kang 0001
ITC5
2003 Test-decompression mechanism using a variable-length multiple-polynomial LFSR
abstract
A new test-decompression methodology using a variable-rank multiple-polynomial linear feedback shift register (MP-LFSR) is proposed. In the proposed reseeding scheme, a test cube with a large number of specified bits is encoded with a high-rank polynomial, while a test cube with a small number of specified bits is encoded with a low-rank polynomial. Therefore, according to the number of specified bits in each test cube, the size of the encoded data can be optimally reduced. A variable-rank MP-LFSR can be implemented with a slight modification of a conventional MP-LFSR. The experimental results on the largest ISCAS'89 benchmark circuits show that the proposed methodology can provide much better encoding efficiency than the previous methods with adequate hardware overhead.
Hong-Sik Kim, YongJoon Kim, Sungho Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2002 DPSC SRAM Transparent Test Algorithm
abstract
We present a new transparent SRAM test algorithm, which uses dynamic power supply current. The proposed test scheme employs the dynamic power supply current instead of making signatures, so that it does not need the additional steps and additional hardware to generate signatures. This paper describes how to convert a traditional March algorithm to a transparent one. The transformed algorithm is much simpler and the test time can be reduced very much. In addition, it can detect some additional faults that the original algorithm cannot detect.
Hong-Sik Kim, Sungho Kang 0001
Asian Test Symposium2
2001 Control-theoretic max-min flow control with minimum rate guarantee
abstract
We present a novel control-theoretic explicit rate (ER) allocation algorithm for the max-min flow control of elastic traffic services with minimum rate guarantee in the context of the ATM ABR service. The proposed ER algorithm is simple in that the number of operations required to compute it at a switch is minimized, scalable in that per-VC (virtual circuit) operations including per-VC queueing, per-VC accounting and per-VC state management are virtually removed, and stable in that by employing it the user transmission rates and the network queues are asymptotically stabilized at a unique equilibrium point at which max-min fairness with minimum rate guarantee and target queue lengths are achieved respectively. To improve the speed of convergence we normalize the controller gains of the algorithm by the estimate of the number of locally-bottlenecked VCs. The estimation scheme is also computationally simple and scalable since it does not require per-VC accounting either. We analyze the theoretical performance of the proposed algorithm and verify its agreement with the practical performance through simulations in the case of multiple bottleneck nodes. We believe that the proposed algorithm will serve as an encouraging solution to the max-min flow control not only in the context of ATM ABR service but also in general elastic traffic services.
Song Chong, Sangho Lee 0003, Sungho Kang 0001
GLOBECOM3
2001 A Heuristic for Multiple Weight Set Generation
abstract
The number of weighted random patterns depends on the number of deterministic test patterns with a low sampling probability. The weight set that is extracted from the deterministic pattern set with high sampling probability reduces the number of test patterns. In this paper we present a new deterministic pattern selection algorithm which generates high performance weight sets by removing deterministic patterns with low sampling frequencies. Simulation results using ISCAS 85 benchmark circuits prove the effectiveness of the new weight set generation algorithm.
Hong-Sik Kim, Jin-kyue Lee, Sungho Kang 0001
ICCD3
2001 A new multiple weight set calculation algorithm
abstract
The number of weighted random patterns depends on the sampling probability of the corresponding deterministic test pattern. Therefore if the weight set is extracted from the deterministic pattern set with high sampling probabilities, the test length can be shortened. In this paper we present a new multiple weight set generation algorithm that generates high performance weight sets by removing deterministic patterns with low sampling probabilities. In addition, the weight set that makes the variance of sampling probabilities for deterministic test patterns small, reduces the number of the deterministic test patterns with low sampling probability. Henceforth we present a new weight set calculation algorithm that uses the optimal candidate list and reduces the variance of the sampling probability. The results on ISCAS 85 and ISCAS 89 benchmark circuits prove the effectiveness of the new weight set calculation algorithm.
Hong-Sik Kim, Jin-kyue Lee, Sungho Kang 0001
ITC3
2001 A simple, scalable, and stable explicit rate allocation algorithm for MAX-MIN flow control with minimum rate guarantee
abstract
We present a novel control-theoretic explicit rate (ER) allocation algorithm for the max-min flow control of elastic traffic services with minimum rate guarantee in the setting of the ATM available bit rate (ABR) service. The proposed ER algorithm is simple in that the number of operations required to compute it at a switch is minimized, scalable in that per-virtual-circuit (VC) operations including per-VC queueing, per-VC accounting, and per-VC state management are virtually removed, and stable in that by employing it, the user transmission rates and the network queues are asymptotically stabilized at a unique equilibrium point at which max-min fairness with minimum rate guarantee and target queue lengths are achieved, respectively. To improve the speed of convergence, we normalize the controller gains of the algorithm by the estimate of the number of locally bottlenecked VCs. The estimation scheme is also computationally simple and scalable since it does not require per-VC accounting either. We analyze the theoretical performance of the proposed algorithm and verify its agreement with the practical performance through simulations in the case of multiple bottleneck nodes. We believe that the proposed algorithm will serve as an encouraging solution to the max-min flow control of elastic traffic services, the deployment of which has been debated long due to their lack of theoretical foundation and implementation complexity.
Song Chong, Sangho Lee 0003, Sungho Kang 0001
IEEE/ACM Trans. Netw.3
1999 At-Speed Boundary-Scan Interconnect Testing in a Board with Multiple System Clocks
abstract
As an at-speed solution to board-level interconnect testing, an enhanced boundary-scan architecture utilizing a combination of slightly modified boundary-scan cells and a user-defined register is proposed. Test methods based on the new architecture can accomplish cost-effective at-speed testing and propagation delay measurements on board-level interconnects. Particularly when the board under test has multiple domains of interconnects controlled by different clock speeds, our at-speed solution is much more efficient than other previous works.
Jongchul Shin, Sungho Kang 0001
DATE3
1999 An Efficient Interconnect Test Using BIST Module in a Boundary-Scan Environment
abstract
In this paper, an efficient built-in self-test (BIST) method for applying tests is developed without collisions of the test data in three-state nets in a system. A new interconnect test algorithm in multiple boundary scan chains and a BIST module based on the new BIST method are presented. The new algorithm can be easily applied to any net configurations with high flexibility.
Jongchul Shin, Sungho Kang 0001
ICCD3
1995 Massively Parallel Array Processor for Logic, Fault, and Design Error Simulation
abstract
Digital logic, fault, and error simulation of large VLSI circuits is one of the most compute-intensive tasks in digital systems analysis. This paper describes a massively parallel special purpose array processor, or hardware accelerator, for digital logic, fault, and error simulation. Hardware simulation is a viable approach for simulation of large systems, since simulation time increases rapidly as a function of the size and complexity of the systems to be simulated. In order to reduce the cost and to achieve high performance, a massively parallel array processor and new algorithms have been introduced. By executing an efficient and direct model of the design on the PE array, the architecture can provide high performance, similar to prototyping. Simulation results show that the hardware accelerator is orders of magnitude faster than software simulation.>
Youngmin Hur, Stephen A. Szygenda, E. Scott Fehr, Granville E. Ott, Sungho Kang 0001
HPCA5
1994 The simulation automation system (SAS); concepts, implementation, and results
abstract
The simulation automation system (SAS) was developed to provide an efficient simulation environment, by automating the entire simulation process. This system can be classified, by its salient unique features, into: automatic model generator (AMG), automatic simulator developer (ASD), and design error simulation and test system (DEST). The system can automatically generate multivalued simulation models and automatically develop various simulators, using domain specific automatic programming techniques. The automatic model generation feature can be used when a new model library is built or when an existing library is upgraded. The automatic simulator development feature allows a user who may not be knowledgeable about simulators, to easily develop unique simulators, which can be used for special purposes or special designs. SAS can also verify designs using the Design Error Simulation and Test System. It provides a confidence measure of the verification, as well as simulation results. When users are not satisfied with the confidence level achieved after simulation, they can automatically generate additional simulation patterns for design errors in order to achieve a higher confidence level. Using this approach, design verification time and cost can be considerably reduced, and an actual measure of verification is provided. Consequently, the design cycle can be considerably reduced. This is especially significant for large, complex systems.>
Sungho Kang 0001, Stephen A. Szygenda
IEEE Trans. Very Large Scale Integr. Syst.1
1992 Modeling and Simulation of Design Errors
abstract
When using simulation for design verification, only a subset of possible simulation patterns is used, since exhaustive simulation is usually not practical. When this subset is used, the designer wants to know how efficiently the design has been verified. The authors provide a measure of simulation pattern effectiveness, based on the concepts of design error modeling. This measure gives insight into the actual level of design validation that has been achieved. These concepts provide a basis for discussing a design validation metric.>
Sungho Kang 0001, Stephen A. Szygenda
ICCD1