VLDB 2026 Research / reviewers in the wild / expert
Seung Ho Shin
dblp:308/0210
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0002-1153-0034ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR-PIM: Self-Test and Repair Structure for Processing-in-Memory With Adder Tree-Based MACabstractProcessing-in-memory (PIM) architectures alleviate memory bottlenecks and improve latency and energy efficiency for AI and ML workloads by accelerating general matrix-vector multiplication (GEMV) operations in DNNs. However, permanent faults in arithmetic units (AUs) within processing units (PUs) critically impact yield and inference accuracy. Although the hybrid built-in self-test (HBIST) method has been proposed, it has limited capabilities in diagnosing and repairing faulty AUs within PUs. In this study, a novel Self-Test And Repair structure for PIM (STAR-PIM) is proposed to enable both fault diagnosis and repair by incorporating a bypass mechanism. A scan-path-like approach enables the testing and precise localization of faulty AUs, while faulty adders are bypassed using a redundant adder structure integrated within the memory die. Furthermore, faulty multipliers are masked using the weight-swapping logic. Experimental results demonstrate that STAR-PIM achieves high AU-level test coverage, ranging from 98.89% to 100% with reasonable area overhead. Recovery experiments show that STAR-PIM maintains low relative errors under fault rates up to 1% for GPT-2 and preserves inference accuracy under fault rates up to 3% for MNIST-MLP. Power measurements on GDDR6-AiM indicate an average overhead of 7.39% with only a 0.08% latency increase. Consequently, STAR-PIM significantly enhances the yield and reliability of PIM while reducing test costs, making it a highly practical solution. Seung Ho Shin, Younwoo Yoo, Youngki Moon, Nuri Son, Dahoon Kim, Sungho Kang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2026 | A Low Area Built-In Self-Repair Using Hybrid Fault Address Memory for HBMabstractThe massive computational requirements of large language model (LLMs) have increased the need for high-bandwidth memory (HBM), which involves high-volume data transfers. The high cell capacity of HBM results in extended test and repair times, leading to increased manufacturing costs. To reduce test time, a built- in self-repair (BISR) circuit, integrated into the HBM base die to detect and repair faults, tests multiple banks in parallel. Conventional BISR approaches adopt content-addressable memory (CAM) for fault classification to reduce repair time. However, dedicated CAM on each bank leads to substantial area overhead associated with its comparison logic. To address these issues, a novel BISR architecture that decouples fault classification and storage is proposed in this article. By introducing a linked CAM design with low area and sharing it across banks for fault classification, while small-area first-in first-out (FIFO) memories allocated to each bank store the classified fault information, the proposed architecture substantially reduces overall area overhead. Furthermore, the proposed architecture reorders the repair solution search sequence toward the most promising candidates by swapping fault entries during test idle periods, thereby significantly reducing repair time. Experimental results demonstrate that the proposed BISR architecture achieves low area overhead and fast repair time for high-density HBM. Seung Ho Shin, Youngki Moon, Eugene Jeong, Sungho Kang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | A Novel CNN-Based Redundancy Analysis Using Parallel Solution DecisionabstractThe increase in memory cell density and capacity has resulted in more faulty cells, necessitating the use of redundant memory row and column lines for repairs. However, existing redundancy analysis (RA) algorithms face a critical issue that RA time increases exponentially with the number of faulty cells. Furthermore, RA solutions for multiple memory chips cannot be derived simultaneously. In this study, a novel RA method is proposed using a convolutional neural network (CNN). The proposed RA algorithm also includes preprocessing to improve training accuracy. The solution locations on the fault map are predicted using multi-label classification. Moreover, parallel solution decision methods ensure that even if the CNN does not find the correct RA solution, an accurate final solution can still be derived, and PyCUDA is used to process multiple memories in parallel. From the experimental results, the normalized repair rate of the proposed RA is 100%. The RA time of the proposed RA is not affected by the number of faults but rather by the CNN execution time. Moreover, RA solutions for multiple memories can be quickly derived simultaneously by utilizing GPU parallel processing. In conclusion, a high yield and low test cost can be achieved. Seung Ho Shin, Minho Cheong, Hayoung Lee, Sungho Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | A Built-In Self-Repair With Maximum Fault Collection and Fast Analysis Method for HBMabstractHigh bandwidth memory (HBM) represents a significant advancement in memory technology, requiring quick and accurate data processing. Built-in self-repair (BISR) is crucial for ensuring high-capacity and reliable memories, as it automatically detects and repairs faults within memory systems, preventing data loss and enhancing overall memory reliability. The proposed BISR aims to enhance the repair rate and reliability by using a content-addressable memory structure that operates effectively in both offline and online modes. Furthermore, a new redundancy analysis algorithm reduces both analysis time and area overhead by converting fault information into a matrix format and focusing on fault-free areas for each repair solution. Experimental results demonstrate that the proposed BISR improves repair rates and derives a final repair solution immediately after the test sequences are completed. Moreover, hardware comparisons have shown that the proposed approach reduces the area overhead as memory size increases. Consequently, the proposed BISR enhances the overall performance of BISR and the reliability of HBM. Joonsik Yoon, Hayoung Lee, Youngki Moon, Seung Ho Shin, Sungho Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | A Novel Prediction-Based Two-Tiered ECC for Mitigating SWD Errors in HBMabstractErrors emerge as a major issue in the reliability of dynamic random access memory (DRAM). To enhance reliability, a two-tiered error correction code (ECC) architecture that comprises on-die ECC (OD-ECC) and system ECC (S-ECC) is adopted as a part of the standard for state-of-the-art high-bandwidth memory (HBM). However, conventional ECCs are insufficient to mitigate malfunctions of subwordline drivers (SWDs), a primary cause of errors. Moreover, the efficient co-design of two-tiered ECCs has not been sufficiently studied. To address these issues without increasing the size of check bits, this article proposes a two-tiered ECC architecture comprising an OD-ECC based on prediction and an S-ECC with data deinterleaving. The proposed OD-ECC predicts the SWD errors by leveraging the detection capabilities of two interleaved Reed-Solomon (RS) engines. In addition, the proposed S-ECC not only preserves strong error detection capability but also masks the misprediction effect of OD-ECC, where data deinterleaving renders additional errors caused by misprediction of OD-ECC to be bounded in the detectable range of the employed cyclic redundancy check (CRC). The experimental results demonstrate that the proposed two-tiered ECC can significantly enhance the error correction capability for SWD errors while maintaining the correction capability for other types of errors. Youngki Moon, Seung Ho Shin, Seokjun Jang, Duyeon Won, Sungho Kang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Effective Parallel Redundancy Analysis Using GPU for Memory RepairabstractThe rapid increment of the memory density leads to an increment of fault occurrence in memory cells. To improve the memory yield, effective memory test and repair methodologies for automatic test equipment (ATE) have been studied. Multiple memory chips are tested simultaneously by the ATE to improve throughput and reduce costs. In general, redundancy analysis (RA) is used for memory repair. However, since conventional RA methods store fault information in the respective failure bitmaps and operate sequentially, those have limitations due to the high area and analysis time. To address these problems, a novel graphic processing unit (GPU)-based RA method has been proposed which significantly enhances the efficiency of searching for repair solutions for multiple memories. The proposed RA method strategically focuses on the pivot line to efficiently utilize parallel processing and reduce the solution search space. Moreover, the proposed method does not require the extensive use of failure bitmaps since all process is conducted on the GPU. The process involves real-time fault collection, analysis, spare allocation, and solution decision process dynamically during the memory test. Experimental results demonstrate that the performance of the proposed RA method achieves an optimal repair rate and high analysis speed for multiple memories. Seung Ho Shin, Hayoung Lee, Sungho Kang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | GRAP: Efficient GPU-Based Redundancy Analysis Using Parallel Evaluation for Cross FaultsabstractVarious memory repair methodologies based on redundancy analysis (RA) have been developed to improve the memory yield. However, conventional RAs often encounter difficulties in finding repair solutions for cases involving a large number of faults and redundancies. To address this problem, an efficient graphics processing unit (GPU)-based RA is proposed using Parallel evaluation for cross faults (GRAP). GRAP involves a preprocessing stage during memory testing, leveraging the parallel processing capacities of the GPU. Preprocessing facilitates rapid solution search by analyzing the fault information. After the test, the solution search is performed. The GPU threads are used to implement all possible cases of redundancy allocation, focusing on cross faults. The remaining faults are categorized by allocating the corresponding redundancies using an efficient method. Given that the solution search process efficiently exploits the multiple threads, GRAP can rapidly find a solution even in cases with a large number of faults and redundancies. Experiments are performed using the compute unified device architecture (CUDA) library for GPU parallel processing, and the performance of the GRAP is compared with those of conventional RA methodologies. The results demonstrated that the proposed RA method can achieve an optimal repair rate with a high analysis speed by leveraging efficient parallel computing. Seung Ho Shin, Hayoung Lee, Sungho Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | TRUST: Through-Silicon via Repair Using Switch Matrix TopologyabstractTo address the demand for memory scaling capabilities, 3-D integrated circuits (3D-ICs) based on short and dense through-silicon vias (TSVs) have been introduced. However, the defects of TSVs considerably influence the yield and reliability of 3D-ICs. For this reason, TSV repair using switch matrix (SM) topology (TRUST) is proposed in this article. TRUST adopts an SM, which has a high routing flexibility, to realize TSV connections. Consequently, a 100% repair rate can be achieved for the 3D-ICs that have faulty TSVs smaller than or equal to redundant TSVs. Furthermore, TRUST utilizes content-addressable memories in built-in self-repair to identify TSV repair paths via a simple TSV repair path search algorithm. For this reason, TRUST can be applied to repair manufacturing and aging defects of TSVs. Nevertheless, TRUST can be applied with reasonable area and delay overheads, such as 58.3% area reduction and 55.1% delay reduction compared to the only conventional TSV repair architecture that can achieve the optimal repair rate. In addition, the area ratios in high bandwidth memory (HBM) and HBM2 are only 5.3% and much smaller than 0.1%, respectively. The advantages are experimentally verified. Hayoung Lee, Seung Ho Shin, Younwoo Yoo, Sungho Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | ECMO: ECC Architecture Reusing Content-Addressable Memories for Obtaining High Reliability in DRAMabstractAdvances in the density and capacity of dynamic random access memories (DRAMs) have resulted in emerging reliability issues. The error correction code (ECC) is widely used as a promising technique to improve the reliability of high-density memories. For this reason, many studies on ECC have been conducted to address the increased cell failure rates. However, conventional ECCs have shown limited achievements owing to area, latency, and power overheads. This study proposes ECC architecture reusing content-addressable memories (CAMs) for obtaining high reliability in DRAM, which can be called ECMO. The proposed architecture reuses CAMs in built-in self-repair, which can be used to repair memory hard faults during manufacturing as data storage to replace error data words. This achieves high reliability along with an additional 9155 h DRAM lifetime. Nevertheless, it can be implemented with a 3.04% area overhead due to the reuse of CAMs. Moreover, only 0.21 ns is added to the critical path. Furthermore, the power overhead is 0.1% compared to the total power consumption of DDR3 and DDR4. Hayoung Lee, Younwoo Yoo, Seung Ho Shin, Sungho Kang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |