VLDB 2026 Research / reviewers in the wild / expert
Joon-Sung Yang
dblp:18/5332
· DBLP profile ↗
65ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0002-1502-5353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 10 first-author · 22 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Computer networks · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable GNN-Driven Test Point Insertion on Uncontrollable I/OsabstractTest coverage degradation from uncontrollable I/Os is a critical challenge in modern SoC design. In area-sensitive applications, such as the peripheral circuits of memory devices, standard DFT solutions like wrapper chains are prohibitively expensive due to their high area overhead. This necessitates a surgical Test Point Insertion (TPI) strategy that maximizes testability while adhering to strict cost constraints. To address this challenge, we propose a novel TPI framework using an explainable Graph Neural Network (GNN). Our GNN accurately predicts test coverage in circuits with masked I/Os, and an integrated saliency map (XAI) technique then identifies the most critical I/Os for TPI. Compared to a leading commercial tool, our framework achieves the target test coverage with 7.53% fewer TPs and improves coverage by 4.34% with the same TP budget on average. The scalability on large circuits (>100k gates) and technology independence confirm its practical applicability for minimizing die cost in constrained, real-world designs. Sung-Hyuk Cho, Tae-Min Park 0002, Jeongyeol Lee, Jae-Youn Hong, Andreas Gerstlauer, Joon-Sung Yang |
DATE | 6 |
| 2025 | DBC: Drift-aware Binary Code for Drift-tolerant Deep Neural NetworksabstractDeep neural networks (DNNs) have demonstrated outstanding performance across a wide range of applications. However, their substantial number of weights necessitates scalable and efficient storage solutions. Emerging non-volatile memory technologies, such as phase change memory (PCM) with multi-level cell (MLC) operation, are promising candidates due to their high scalability and non-volatility compared to conventional charge-based storage devices. Despite these advantages, MLC PCM suffers from reliability issues, particularly conductance drift, where the conductance of a PCM cell changes over time. This drift can lead to significant accuracy degradation in DNNs, as their weights are stored in PCM cells. In this paper, we propose Drift-aware Binary Code (DBC), a novel binary code designed to improve the tolerance of DNNs to conductance drift. DBC maps smaller decimal values to less error-prone MLC PCM cell levels and ensures that values shift to smaller magnitudes when conductance drift occurs. This approach helps maintain the accuracy of the DNN over an extended period compared to conventional binary code, as dominant DNN weights are stored at levels less prone to errors and DNNs exhibit better tolerance when weight values decrease rather than increase due to drift. Additionally, DBC requires no additional hardware overhead for auxiliary bits and can be combined with other fault-tolerant approaches, such as error correction code (ECC). Experimental results based on the real PCM device developed by IBM Research demonstrate that DBC improves the drift tolerance of DNNs by up to $55.18 \times$ compared to conventional binary code. Insu Choi, Jaeyong Chung, Joon-Sung Yang |
DAC | 3 |
| 2025 | Bit-slice Architecture for DNN Acceleration with Slice-level Sparsity Enhancement and ExploitationabstractDeep Neural Networks (DNNs) demand significant computational resources, prompting the emergence of bit-slice architectures designed to efficiently accelerate DNNs by leveraging high bit-precision reconfigurability and fine-grained sparsity through slice-level computation. However, fully utilizing slice-level sparsity remains challenging, leading conventional bit-slice architectures to exploit either input or weight sparsity at a coarse-grained level. In this paper, we introduce a Bit-slice Architecture for DNN Acceleration (BADA) that simultaneously leverages both input and weight sparsity at coarse- and fine-grained levels. BADA features a novel architecture that skips computations for bit-slice chunks containing zero values and also skips any set of operands where either the input or weight bit-slice is zero. The design comprises a front-end unit responsible for generating bit-slices and selectively gathering only the non-zero slices, and a back-end unit equipped with a signed multiply-and-accumulate (MAC) unit to process these collected non-zero bit-slices. Additionally, we present two algorithmic optimizations to further enhance the efficiency and performance of BADA. First, we propose a novel bit-slice representation that supports 8-bit data without incurring additional hardware overhead, whereas conventional bit-slice representations are limited to 6-bit or 7-bit data under similar constraints. Second, we introduce a method to narrow the weight distribution during the training process, thereby increasing the proportion of zero-valued higher-order bit-slices and further enhancing slice-level weight sparsity. Experimental results demonstrate that BADA achieves a 2.67× increase in throughput, a 1.52× improvement in area efficiency, and a 2.15× enhancement in energy efficiency compared to LUTein, the state-of-the-art bit-slice architecture. Insu Choi, Young-Seo Yoon, Joon-Sung Yang |
HPCA | 3 |
| 2025 | Self-Error Detection and Correction Techniques for Reliable and Efficient Selector-Only MemoryabstractStorage class memory (SCM) has been proposed as a solution to bridge the performance gap between main memory and storage in data-intensive applications. To overcome the limitations of prior technologies like 3D cross-point memory (3DXP), Selector-Only Memory (SOM) has recently emerged as a promising candidate for next-generation SCM. However, SOM devices face various reliability challenges that limit their ability to function as ideal memory devices. Among various non-ideal characteristics, threshold voltage (Vth) drift, which increases Vthover time, has been identified as a critical issue that raises the error rate of SOM. Moreover, device variation caused by immature manufacturing processes results in Vthdiscrepancies across devices, reducing the read window margin (RWM) and thus contributing to an increased error rate. In response to these issues, we propose Securer, consisting of self-error detection and correction techniques that inherently detect and correct errors for SOM, without relying on external methods such as error correction codes (ECC). The first technique effectively identifies all errors through a dual-polarity read operation, leveraging the unique features of SOM. For the second technique, an analytical drift estimation model is introduced to estimate the magnitude of Vthdrift experienced by erroneous cells. Using this information, an adaptive error-aware read operation, which tracks the drifted Vth, is employed to correct the detected errors. Experimental results across real workloads from various applications demonstrate that Securer achieves error rates below 10−13without significant overhead, and confirm error-free data integrity in conjunction with single-error correction code. Hyunjun Lee, Joon-Sung Yang |
ICCAD | 2 |
| 2025 | Accelerating Retrieval Augmented Language Model via PIM and PNM Integration
Je-Woo Jang, Junyong Oh, Youngbae Kong, Jae-Youn Hong, Sung-Hyuk Cho, Jeongyeol Lee, Hoeseok Yang, Joon-Sung Yang |
MICRO | 8 |
| 2025 | Reducing Errors and Powers in LPDDR for DNN Inference: A Compression and IECC-Based Approach
Jae-Youn Hong, Je-Woo Jang, Sung-Hyuk Cho, Youngbae Kong, Sungkyu Kim, Youngjung Kang, Jaehyung Ko, Jaeyong Chung, Joon-Sung Yang |
J. Syst. Archit. | 9 |
| 2025 | AGD: Analytic Gradient Descent for Discrete Optimization in EDA and its Use to Gate SizingabstractIn electronic design automation (EDA), simulation models are often non-differentiable, and many design choices are discrete. As a result, greedy optimization methods based on numerical gradients are widely used, although they often lead to suboptimal solutions. In contrast, analytical methods may provide better solutions but require significant research effort. Reinforcement learning (RL) has been employed to address this problem; however, RL also suffers from notorious sample inefficiency, which is exaggerated in EDA because data sampling in EDA is very expensive due to slow simulations. This article proposes an alternative to RL for EDA, namely analytic gradient descent (AGD). Our method starts with a differentiable performance model, which can be either a learned surrogate or a static model. It then applies transformations similar to Shannon decomposition for each design variable in the performance model. Finally, one design option for each variable is selected using a one-hot variable, which is trained via a straight-through estimator (STE) through gradient descent. We demonstrate AGD on the well-known gate sizing problem using both a learned surrogate and a static model across 20 industrial benchmark circuits. Our experimental results show that the proposed method can outperform a several-decade-old commercial tool in the gate sizing task for 19 out of the 20 circuits. Phuoc Pham, Tae-Min Park 0001, Sung-Hyuk Cho, Tayyeb Mahmood, Joon-Sung Yang, Jaeyong Chung |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | ViT-slice: End-to-end Vision Transformer Accelerator with Bit-slice AlgorithmabstractVision Transformers have demonstrated remarkable performance in various vision tasks. However, general-purpose processors, such as CPUs and GPUs, face challenges in efficiently handling the inference of Vision Transformers. To address the issue, prior works have focused on accelerating only attention due to its high computational cost in NLP Transformers. In contrast, Vision Transformers demonstrate a higher computational cost due to linear modules such as linear transformation, linear projection and Feed-Forward Network (FFN), compared to attention. In this paper, we present ViT-slice, an algorithm-architecture co-design that enhances end-to-end performance and energy efficiency by optimizing not only attention but also linear modules. At the algorithm level, we propose bit-slice compression that avoids storing the redundant most significant bits (MSBs). Additionally, we present bit-slice dot product with early skip to efficiently compute the dot product using bit-sliced data. To enable early skip during the dot product computation, we leverage a trainable threshold. On the hardware level, we introduce a specialized bit-slice dot product unit (BSDPU) to efficiently process the bit-slice dot product with early skip algorithm. Additionally, we present a bit-slice encoder and decoder for on-chip bit-slice compression. ViT-slice achieves 244×, 35.3×, 16.8×, 10.4×, 5.0× end-to-end speedup over Xeon CPU, EdgeGPU, TITAN Xp GPU, Sanger accelerator and ViTCoD accelerator, respectively. Dongjin Shin, Insu Choi, Joon-Sung Yang |
DAC | 3 |
| 2024 | LOCo: LPDDR Optimization with Compression and IECC scheme for DNN InferenceabstractWith the increasing demand for on-device Artificial Intelligence (AI), various compression schemes have been proposed for DNN models to be efficiently executed on mobile devices. Although various studies have proposed compression schemes, their compatibility with mobile device memory (i.e., LPDDR) is not considered; LPDDR is prone to error as it operates at low voltage due to its strict power constraint. Currently, DRAM vendors adopt an ECC engine with SEC(136,128) code inside each LPDDR bank (i.e., IECC) for reliable operation. While IECC enhances reliability, it lowers performance due to Read-Modify-Write (RMW) and parity storage costs. Thus, for both power-efficient and reliable DNN operation in mobile environments, this paper introduces LOCo, a DNN weight compression scheme with 3-staged protection. LOCo reduces the IECC engine operation granularity from SEC(136,128) to SECDED(72,64), thus eliminating power-intensive internal reads but enhancing reliability. To mitigate the storage overhead, this paper proposes a compression scheme with protection based on the characteristics of DNN weights. Our evaluations show that LOCo provides a power reduction of 16.94%, latency reduction of 16.81%, and energy reduction of 30.81% when compared to conventional LPDDR with SEC engine, while demonstrating robustness under 17959X more bit error rate when compared to existing compression scheme. Jae-Youn Hong, Sungkyu Kim, Je-Woo Jang, Joon-Sung Yang |
ISLPED | 4 |
| 2024 | HYDRA: A Hybrid Resistance Drift Resilient Architecture for Phase Change Memory-Based Neural Network AcceleratorsabstractIn-memory Computing (IMC) using Phase Change Memory (PCM) has proven to be effective for efficient processing of Deep Neural Networks (DNNs). However, with the use of multi-level cell PCM (MLC-PCM) in NVMs-based accelerators, errors due to resistance drift in MLC-PCM can severely degrade the DNNs accuracy. In this paper, an analysis of the impact of resistance drift errors on accuracy of MLC-PCM based DNN accelerator shows that the drift errors alone can significantly impact the accuracy. This paper proposes Hydra, which is a hybrid resistance drift resilient architecture for MLC-PCM based DNN accelerators which use IMC for efficient computations. Hydra utilizes Tri-level cell PCM, which has a negligible resistance drift error rate, to store the critical bits of DNNs parameters and MLC-PCM (4-level cell), which has a higher error rate (but offers more storage density), for the non-critical bits. Experimental results on various DNN architectures, configurations and datasets show that, with the presence of resistance drift errors in PCM, Hydra can maintain the baseline accuracy of DNNs for up to 1 year (resistance drift is time-dependent), whereas conventional drift tolerance techniques lead to a significant accuracy drop in just a few seconds. Thai-Hoang Nguyen, Muhammad Imran 0010, Jaehyuk Choi 0001, Joon-Sung Yang |
IEEE Trans. Computers | 4 |
| 2023 | RQ-DNN: Reliable Quantization for Fault-tolerant Deep Neural NetworksabstractDeep Neural Networks (DNNs) are deployed in many real-time and safety-critical applications such as autonomous vehicles and medical diagnosis. In such applications, quantization is used to compress the model for storage and computation reduction. However, recent research has shown that faults in memory can cause a significant drop in DNN accuracy and conventional quantization methods focus only on model compression. This paper proposes a novel method that performs model quantization while remarkably improving the fault-tolerance of the model. It can be incorporated with other hardware approaches such as Error Correcting Code to further improve fault-tolerance. The proposed method reduces possible error patterns that negatively impact classification accuracy by modifying weight distributions and applying a novel masking-based clipping function. Experimental results show that the proposed method enhances the fault-tolerance of the quantized DNN, which can tolerate 1803× higher bit error rates than the conventional method. Insu Choi, Jae-Youn Hong, JaeHwa Jeon, Joon-Sung Yang |
DAC | 4 |
| 2023 | PIE-DRAM: Postponing IECC to Enhance DRAM performance with access tableabstractThis paper proposes a novel memory architecture, PIE-DRAM, to mitigate the performance overhead caused by IECC. Unlike conventional IECC architectures, the proposed method separates IECC from data path to enable independent IECC operations from memory read or write access. Based on recent memory access histories, memory controller selectively determines the usage of IECC to alleviate IECC overhead, thereby the proposed architecture enhances the memory performance. Experimental results show that, from memory intensive workloads, 6% IPC (Instructions Per Cycle) improvement is achieved solely by applying the proposed DRAM architecture utilizing the locality and modified RMW. JaeHwa Jeon, Jae-Youn Hong, Insu Choi, Joon-Sung Yang |
DAC | 5 |
| 2023 | VECOM: Variation-Resilient Encoding and Offset Compensation Schemes for Reliable ReRAM-Based DNN AcceleratorabstractResistive Random-Access Memory (ReRAM)-based Processing In-Memory (PIM) Accelerator has emerged as a promising computing architecture for memory-intensive applications, such as Deep Neural Networks (DNNs). However, due to its immaturity, ReRAM devices often suffer from various reliability issues, which hinder the practicality of the PIM architecture and lead to a severe degradation in DNN accuracy. Among various reliability issues, device variation and offset current from High Resistance State (HRS) cell have been considered as major problems in a ReRAM-based PIM architecture. Due to these problems, the throughput of the ReRAM-based PIM is reduced as fewer wordlines are activated. In this paper, we propose VECOM, a novel approach that includes a variation-resilient encoding technique and an offset compensation scheme for a robust ReRAM-based PIM architecture. The first technique (i.e., VECOM encoding) is built based on the analysis of the weight pattern distribution of DNN models, along with the insight into the ReRAM's variation property. The second technique, VECOM offset compensation, tolerates offset current in PIM by mapping the conductance of each Multi-level Cell (MLC) level added with a specific offset conductance. Experimental results in various DNN models and datasets show that the proposed techniques can increase the throughput of the PIM architecture by up to 9.1 times while saving 50% of energy consumption without any software overhead. Additionally, VECOM is also found to endure low R-ratio ReRAM cell (up to 7) with a negligible accuracy drop. Je-Woo Jang, Thai-Hoang Nguyen, Joon-Sung Yang |
ICCAD | 3 |
| 2023 | A convertible neural processor supporting adaptive quantization for real-time neural networks
Hongju Kal, Hyoseong Choi, Ipoom Jeong, Joon-Sung Yang, Won Woo Ro |
J. Syst. Archit. | 4 |
| 2023 | CRAFT: Criticality-Aware Fault-Tolerance Enhancement Techniques for Emerging Memories-Based Deep Neural NetworksabstractDeep neural networks (DNNs) have emerged as the most effective programming paradigm for computer vision and natural language processing applications. With the rapid development of DNNs, efficient hardware architectures for deploying DNN-based applications on edge devices have been extensively studied. Emerging nonvolatile memories (NVMs), with their better scalability, nonvolatility, and good read performance, are found to be promising candidates for deploying DNNs. However, despite the promise, emerging NVMs often suffer from reliability issues, such as stuck-at faults, which decrease the chip yield/memory lifetime and severely impact the accuracy of DNNs. A stuck-at cell can be read but not reprogrammed, thus, stuck-at faults in NVMs may or may not result in errors depending on the data to be stored. By reducing the number of errors caused by stuck-at faults, the reliability of a DNN-based system can be enhanced. This article proposes CRAFT, i.e., criticality-aware fault-tolerance enhancement techniques to enhance the reliability of NVM-based DNNs in the presence of stuck-at faults. A data block remapping technique is used to reduce the impact of stuck-at faults on DNNs accuracy. Additionally, by performing bit-level criticality analysis on various DNNs, the critical-bit positions in network parameters that can significantly impact the accuracy are identified. Based on this analysis, we propose an encoding method which effectively swaps the critical bit positions with that of noncritical bits when more errors (due to stuck-at faults) are present in the critical bits. Experiments of CRAFT architecture with various DNN models indicate that the robustness of a DNN against stuck-at faults can be enhanced by up to$10^{5}$times on the CIFAR-10 dataset and up to 29 times on ImageNet dataset with only a minimal amount of storage overhead, i.e., 1.17%. Being orthogonal, CRAFT can be integrated with existing fault-tolerance schemes to further enhance the robustness of DNNs against stuck-at faults in NVMs. Thai-Hoang Nguyen, Muhammad Imran 0010, Jaehyuk Choi 0001, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Bipolar vector classifier for fault-tolerant deep neural networksabstractDeep Neural Networks (DNNs) surpass the human-level performance on specific tasks. The outperforming capability accelerate an adoption of DNNs to safety-critical applications such as autonomous vehicles and medical diagnosis. Millions of parameters in DNN requires a high memory capacity. A process technology scaling allows increasing memory density, however, the memory reliability confronts significant reliability issues causing errors in the memory. This can make stored weights in memory erroneous. Studies show that the erroneous weights can cause a significant accuracy loss. This motivates research on fault-tolerant DNN architectures. Despite of these efforts, DNNs are still vulnerable to errors, especially error in DNN classifier. In the worst case, because a classifier in convolutional neural network (CNN) is the last stage determining an input class, a single error in the classifier can cause a significant accuracy drop. To enhance the fault tolerance in CNN, this paper proposes a novel bipolar vector classifier which can be easily integrated with any CNN structures and can be incorporated with other fault tolerance approaches. Experimental results show that the proposed method stably maintains an accuracy with a high bit error rate up to 10−3 in the classifier. Suyong Lee, Insu Choi, Joon-Sung Yang |
DAC | 3 |
| 2022 | Value-aware Parity Insertion ECC for Fault-tolerant Deep Neural NetworkabstractDeep neural networks (DNNs) are deployed on hardware devices and are widely used in various fields to perform inference from inputs. Unfortunately, hardware devices can become unreliable by incidents such as unintended process, voltage and temperature variations, and this can introduce the occurrence of erroneous weights. Prior study reports that the erroneous weights can cause a significant accuracy degradation. In safety-critical applications such as autonomous driving, it can bring catastrophic results. Retraining or fine-tuning can be used to adjust corrupted weights to prevent the accuracy degradation. However, training-based approaches would incur a significant computational overhead due to a massive size of training datasets and intensive training operations. Thus, this paper proposes a value-aware parity insertion error correction code (ECC) to recover erroneous weights with a reduced parity storage overhead and no additional training processes. Previous ECC-based reliability improvement methods, Weight Nulling and In-place Zero-space ECC, are compared with the proposed method. Experimental results demonstrate that DNNs with the value-aware parity insertion ECC can perform inference without the accuracy degradation, on average, in 122.5× and 15.1× higher bit error rate conditions over Weight Nulling and In-place Zero-space ECC, respectively. Seo-Seok Lee, Joon-Sung Yang |
DATE | 2 |
| 2022 | DynaPAT: A Dynamic Pattern-Aware Encoding Technique for Robust MLC PCM-Based Deep Neural NetworksabstractAs the effectiveness of Deep Neural Networks (DNNs) is rising over time, so is the need for highly scalable and efficient hardware architectures to capitalize this effectiveness in many practical applications. Emerging non-volatile Phase Change Memory (PCM) technology has been found to be a promising candidate for future memory systems due to its better scalability, non-volatility and low leakage/dynamic power consumption, compared to conventional charged-based memories. Additionally, with its cell's wide resistance span, PCM also has the Flash-like Multi-Level Cell (MLC) capability, which has enhanced storage density, providing an opportunity for the deployment of data-intensive applications such as DNNs on resource-constrained edge devices. However, the practical deployment of MLC PCM is hampered by certain reliability challenges, among which, the resistance drift is considered to be a critical concern. In a DNN application, the presence of resistance drift in MLC PCM can cause a severe impact to DNN's accuracy if no drift-error-tolerance technique is utilized. This paper proposes DynaPAT, a low-cost and effective pattern-aware encoding technique to enhance the drift-error-tolerance of MLC PCM-based Deep Neural Networks. DynaPAT has been constructed on the insight into DNN's vulnerability against different data pattern switching. Based on this insight, DynaPAT efficiently maps the most-frequent data pattern in DNN's parameters to the least-drift-prone level of the MLC PCM, thus significantly enhancing the robustness of the system against drift errors. Various experiments on different DNN models and configurations demonstrate the effectiveness of DynaPAT. The experimental results indicate that DynaPAT can achieve up to 500× enhancement in the drift-errors-tolerance capability over the baseline MLC PCM based DNN while requiring only a negligible hardware overhead (below 1% storage overhead). Being orthogonal, DynaPAT can be integrated with existing drift-tolerance schemes for even higher gains in reliability. Thai-Hoang Nguyen, Muhammad Imran 0010, Joon-Sung Yang |
ICCAD | 3 |
| 2022 | CEnT: An Efficient Architecture to Eliminate Intra-Array Write Disturbance in PCMabstractPhase Change Memory (PCM), with its better scaling potential compared to DRAM, is seen as a promising candidate to replace or complement DRAM. The heat generated from a RESET programming pulse to a PCM cell can disturb the neighboring cells which are not being programmed. Write disturbance (WD) poses a critical reliability challenge in high-density PCM memory with scaling below 20nm process technology node. Increasing the intra-cell space can eliminate the WD, however, it reduces the storage density which counteracts the benefits of scalability in PCM. At architectural level, a verify and correct (VnC) technique can be used to address this problem. However, this leads to an increased number of write operations, thus degrading performance, energy efficiency and memory lifetime. Due to its dependence on the type of programming operation and the state of the neighboring cell, WD is a data-dependent problem. Exploiting this property, encoding techniques have been proposed to reduce the frequency of WD-vulnerable data patterns. These techniques, however, do not eliminate the WD in an array and ultimately rely on the VnC method to ensure reliable memory operation. This article introduces a novel architecture, based on encoding and multi-level programming characteristics of PCM, to eliminate the intra-array WD in PCM. By eliminating WD and hence the need for a VnC operation, the proposed architecture improves performance, energy efficiency and memory lifetime. Our evaluation of the proposed architecture shows an average reduction of 57 percent in the number of writes (to service one write request) over the existing state-of-the-art intra-array WD-mitigation technique. Depending on the PCM write bandwidth, the proposed architecture can reduce the write service time by up to 27 percent, on average, compared to the existing best-performing technique. This leads to an average improvement of 15 percent in IPC. Additionally, by eliminating the overhead of a verify operation, the write energy efficiency is also improved by 8 percent over the previous art. Finally, with an average reduction of 26 percent in bit flips, the proposed method also improves the memory lifetime. The proposed method is also proven to be effective when considering WD both within the word-lines and across the bit-lines. Muhammad Imran 0010, Nur A. Touba, Joon-Sung Yang |
IEEE Trans. Computers | 4 |
| 2022 | ADAPT: A Write Disturbance-Aware Programming Technique for Scaled Phase Change MemoryabstractPhase change memory (PCM) is an emerging, resistance-based, nonvolatile memory. With a promising scaling potential, PCM can replace the existing charge-based memory technologies. A highly scaled PCM is prone to write disturbance (WD) because of a high-current RESET programming pulse. Exploiting the data-dependent nature of WD, encoding techniques have been proposed to reduce the frequency of WD-vulnerable data patterns. These techniques work along with a verify and correct (VnC) method to ensure memory reliability. However, the effectiveness of these techniques varies depending on the data patterns. Unlike the conventional methods, this article introduces a WD-aware programming technique to mitigate WD in PCM. The proposed method encodes the data based on the number of WD-vulnerable cells and the bit flips. By reducing the number of WD-vulnerable cells as well as the bit flips, the proposed method is more effective than the existing encoding techniques, in mitigating WD as well as improving the memory lifetime. Evaluation using various realistic workloads shows that the proposed method can reduce the average word-line WD errors by 62%, compared to the existing state of the art. With a reduced number of WD errors, the frequency of a VnC operation is also reduced. This leads to a reduction of 44% in the number of extra writes and 20% in the average write time. With reduction in the number of writes and the write time, instructions-per-cycle is improved by 9% and the write energy by 11% over the existing art. By reducing the number of bit flips compared to the previous state of the art, the proposed method improves the PCM memory lifetime by 13% to 33%, considering the asymmetry of SET and RESET operations in impacting the cell endurance. Muhammad Imran 0010, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Low-Cost and Effective Fault-Tolerance Enhancement Techniques for Emerging Memories-Based Deep Neural NetworksabstractDeep Neural Networks (DNNs) have been found to outperform conventional programming approaches in several applications such as computer vision and natural language processing. Efficient hardware architectures for deploying DNNs on edge devices have been actively studied. Emerging memory technologies with their better scalability, non-volatility, and good read performance are ideal candidates for DNNs which are trained once and deployed over many devices. Emerging memories have also been used in DNNs accelerators for efficient computations of dot-product. However, due to immature manufacturing and limited cell endurance, emerging resistive memories often result in reliability issues like stuck-at faults, which reduce the chip yield and pose a challenge to the accuracy of DNNs. Depending on the state, stuck-at faults may or may not cause error. Fault-tolerance of DNNs can be enhanced by reducing the impact of errors resulting from the stuck-at faults. In this work, we introduce simple and light-weight Intra-block Address remapping and weight encoding techniques to improve the fault-tolerance for DNNs. The proposed schemes effectively work at the network deployment time while preserving the network organization and the original values of the parameters. Experimental results on state-of-the-art DNN models indicate that, with a small storage overhead of just 0.98%, the proposed techniques achieve up to 300× stuck-at faults tolerance capability on Cifar10 dataset and 125× on Imagenet datatset, compared to the baseline DNNs without any fault-tolerance method. By integrating with the existing schemes, the proposed schemes can further enhance the fault resilience of DNNs. Thai-Hoang Nguyen, Muhammad Imran 0010, Jaehyuk Choi 0001, Joon-Sung Yang |
DAC | 4 |
| 2021 | Reliability Enhanced Heterogeneous Phase Change Memory Architecture for Performance and Energy EfficiencyabstractNext-generation memories have been actively researched to replace the existing memories like DRAM and flash in deep sub-micron process technology. Unlike the conventional charge-based memories, next-generation memories utilize the resistive properties of different materials to store and read a data. Among the next-generation memories, Phase Change Memory (PCM) is seen as a good choice for future memory systems, given its good read performance, process compatibility and scaling potential. To enhance the storage density, multi-level cell (MLC) operation is seemed promising which can store more than one bit in each PCM cell. However, MLC operation significantly degrades the reliability of PCM, thus requiring a strong Error Correction Code (ECC) to guarantee correct memory operation. The use of heavyweight ECC comes at cost of significant degradations in storage density, performance and energy efficiency. In this article, we propose a heterogeneous PCM architecture which uses both multi-level cell and single-level cell (SLC) together for a single word line. With highly-reliable SLC cells, the overall array reliability is enhanced. To improve the reliability further, a dynamic self-encoding/decoding scheme is performed before the data is written to the PCM cells. The dynamic scheme automatically determines the locations of MLC and SLC cells and sets the corresponding resistance levels to be programmed. Since the proposed encoding/decoding scheme does not require any additional stages or storages for encoding and decoding, the overhead is negligible. The improved reliability allows to use lighter ECC scheme which in turn helps to improve performance and energy efficiency of the MLC PCM. The experimental results show that the reliability is improved by approximately$10^6$times compared to the conventional 4LC and more than$10^3$times compared to the existing encoding methods. The performance improvement is 21.5 percent over the conventional 4LC and is more than 4.1 percent higher than the prior encoding techniques. The proposed method is 30.3 percent more energy efficient than the conventional 4LC and this is similar or higher than other energy efficiency improvement methods. Muhammad Imran 0010, Joon-Sung Yang |
IEEE Trans. Computers | 3 |
| 2020 | PAIR: Pin-aligned In-DRAM ECC architecture using expandability of Reed-Solomon codeabstractThe computation speed of computer systems is getting faster and the memory has been enhanced in performance and density through process scaling. However, due to the process scaling, DRAMs are recently suffering from numerous inherent faults. DRAM vendors suggest In-DRAM Error Correcting Code (IECC) to cope with the unreliable operation. However, the conventional IECC schemes have concerns about miscorrection and performance degradation. This paper proposes a pin-aligned In-DRAM ECC architecture using the expandability of a Reed-Solomon code (PAIR), that aligns ECC codewords with DQ pin lines (data passage of DRAM). PAIR is specialized in managing widely distributed inherent faults without the performance degradation, and its correction capability is sufficient to correct burst errors as well. The experimental results analyzed with the latest DRAM model show that the proposed architecture achieves up to 106times higher reliability than XED with 14% performance improvement, and 10 times higher reliability than DUO with a similar performance, on average. Sangmok Jeong, SeungYup Kang, Joon-Sung Yang |
DAC | 3 |
| 2020 | Factored Radix-8 Systolic Array for Tensor ProcessingabstractSystolic arrays are re-gaining the attention as the heart to accelerate machine learning workloads. This paper shows that a large design space exists at the logic level despite the simple structure of systolic arrays and proposes a novel systolic array based on factoring and radix-8 multipliers. The factored systolic array (FSA) extracts out the booth encoding and the hard-multiple generation which is common across all processing elements, reducing the delay and the area of the whole systolic array. This factoring is done at the cost of an increased number of registers, however, the reduced pipeline register requirement in radix-8 offsets this effect. The proposed factored 16-bit multiplier achieves up to 15%, 13%, and 23% better delay, area, and power, respectively, compared with the radix-4 multipliers even if the register overhead is included. The proposed FSA architecture improves delay, area, and power up to 11%, 20% and 31%, respectively, for different bitwidths when compared with the conventional radix-4 systolic array. Inayat Ullah, Kashif Inayat, Joon-Sung Yang, Jaeyong Chung |
DAC | 3 |
| 2020 | Effective Write Disturbance Mitigation Encoding Scheme for High-density PCMabstractWrite Disturbance (WD) is a crucial reliability concern in a high-density PCM with below 20nm scaling. WD occurs because of the inter-cell heat transfer during a RESET operation. Being dependent on the type of programming pulse and the state of the vulnerable cell, WD is significantly impacted by the data patterns. Existing encoding techniques to mitigate WD reduce the percentage of a single WD-vulnerable pattern in the data. However, it is observed that reducing the frequency of a single bit pattern may not be effective to mitigate WD for certain data patterns. This work proposes a significantly more effective encoding method which minimizes the number of vulnerable cells instead of a single bit pattern. The proposed method mitigates WD both within a word-line and across the bit-lines. In addition to WD-mitigation, the proposed method encodes the data to minimize the bit flips, thus improving the memory lifetime compared to the conventional WD-mitigation techniques. Our evaluation using SPEC CPU2006 benchmarks shows that the proposed method can reduce the aggregate (word- line+bit-line) WD errors by 42% compared to the existing state- of-the-art (SD-PCM). Compared to the state-of-the-art SD-PCM method, the proposed method improves the average write time, instructions-per-cycle (IPC) and write energy by 12%, 12% and 9%, respectively, by reducing the frequency of verify and correct operations to address WD errors. With reduction in bit flips, memory lifetime is also improved by 18% to 37% compared to SD-PCM, given an asymmetric cost of the bit flips. By integrating with the orthogonal techniques of SD-PCM, the proposed method can further enhance the performance and energy efficiency. Muhammad Imran 0010, Joon-Sung Yang |
DATE | 3 |
| 2020 | Reliable and Lightweight PUF-based Key Generation using Various Index Voting ArchitectureabstractPhysical Unclonable Functions (PUFs) can be utilized for secret key generation in security applications. Since the inherent randomness of PUF can degrade its reliability, most of the existing PUF architectures have designed post-processing logic to enhance the reliability such as an error correction function for guaranteeing reliability. However, the structures incur high cost in terms of implementation area and power consumption. This paper introduces a Various Index Voting Architecture (VIVA) that can enhance the reliability with a low overhead compared to the conventional schemes. The proposed architecture is based on an index-based scheme with simple computation logic units and iterative operations to generate multiple indices for the accuracy of key generation. Our evaluation results show that the proposed architecture reduces the hardware implementation overhead by 2 to more than 5 times, without losing a key generation failure probability compared to conventional approaches. Jeong-Hyeon Kim, Ho-Jun Jo, Kyung-Kuk Jo, Sung-Hee Cho, Jaeyong Chung, Joon-Sung Yang |
DATE | 6 |
| 2020 | Pattern-Aware Encoding for MLC PCM Storage Density, Energy Efficiency, and Performance EnhancementabstractWith the scaling limitations and increasing leakage power of the existing charge-based memories, next-generation memory technologies to overcome the issues are in development. Among the various emerging memories, phase change memory (PCM) is considered as a promising candidate due to its scalability potential and negligible leakage power. For enhanced storage density, the multilevel cell (MLC) operation has been proposed for PCM. This, however, comes at cost of poor reliability, write energy increase and performance degradation. Unlike DRAM, the MLC PCM has a much higher soft error rate due to the resistance drift phenomenon. Error correction code (ECC) schemes can be utilized to improve the MLC PCM reliability, however, this would lead to a lower storage density and an increase in write energy and latency. The iterative programming required for the MLC PCM also degrades its energy efficiency and performance. This paper introduces a simple yet effective encoding scheme to mitigate the problems of the MLC PCM. By using a simple XOR-based encoding, the proposed architecture minimizes the most drift-prone state in the data. The method divides the original data into several encoding blocks and analyzes initial pattern frequencies for each 2-bit pattern. Based on the initial pattern frequencies, the inputs for the XOR encoding are selected that result in minimal frequency of the drift-prone state. This considerably enhances the MLC PCM reliability, leading to a high storage density with a reduced ECC overhead. The energy efficiency and performance are also improved due to reduction in iterative current pulses and ECC overhead. The simulation results show a reduction of about$10^{5}$X in soft error rate. The improvements in energy efficiency and performance over the conventional 4-level cell (4LC) PCM are 11.5% and 31.9%, respectively. Muhammad Imran 0010, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Virtual-Tile-Based Flip-Flop Alignment Methodology for Clock Network Power OptimizationabstractClock network plays the most significant role in power consumption in IC design. Since a clock network normally has a high switching ratio, power optimization of the clock network is one of the best solutions to minimize dynamic power and total power in modern IC designs. The clock network is synthesized based on an initial flip-flop placement. The number of clock buffers and their sizes are decided by the initial placement. Moreover, clock wires, which are the major sources of clock power consumption, are also constructed based on the flip-flop placement. As a result, the flip-flop placement determines the quality of the clock network. In this article, we propose a new clock network optimization method to reduce the dynamic power consumption of clock network. The method first creates virtual tiles over the entire design area and selects the most effective columns to align flip-flops in lines. Once the effective columns are determined, flip-flops are relocated based on the virtual tiles in the columns considering the minimum moving distance. By aligning flip-flops, it is possible to significantly reduce both wire capacitance and wire length. Since it does not change the clock structure, unlike the conventional clock network optimization techniques which use multibit flip-flop or register bank, there is no degradation in timing or other constraints. Experimental results show that the proposed method reduces the wire capacitance, wire length, and via count up to 23.2%, 10.2%, and 16.4%, respectively, in five industrial intellectual property (IP) designs. The reduction in clock network power is 14.1% on average. Muhammad Imran 0010, David Z. Pan, Joon-Sung Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | ER-TCAM: A Soft-Error-Resilient SRAM-Based Ternary Content-Addressable Memory for FPGAsabstractStatic random access memory (SRAM)-based ternary content-addressable memory (TCAM) on field-programmable gate arrays (FPGAs) is used for packet classification in software-defined networking (SDN) and OpenFlow applications. SRAMs implementing TCAM contents constitute the major part of a TCAM design on FPGAs, which are vulnerable to soft errors. The protection of SRAM-based TCAMs against soft errors is challenging without compromising critical path delay and maintaining a high search performance. This brief presents a lowcost and low-response-time technique for the protection of SRAM-based TCAMs. This technique uses simple, single-bit parity for fault detection which has a minimal critical path overhead. This technique exploits the binary-encoded TCAM table maintained in SRAM-based TCAMs for update purposes to implement a low-response-time error-correction mechanism at low cost. The error-correction process is carried out in the background, allowing lookup operations to be performed simultaneously, thus maintaining a high search performance. The proposed technique provides protection against soft errors with a response time of 293 ns, whereas maintaining a search rate of 222 million searches per second on a 1024 × 40 size TCAM on Artix-7 FPGA. Inayat Ullah, Joon-Sung Yang, Jaeyong Chung |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | DRIS-3: Deep Neural Network Reliability Improvement Scheme in 3D Die-Stacked Memory based on Fault AnalysisabstractVarious studies have been carried out to improve the operational efficiency of the Deep Neural Networks (DNNs). However, the importance of the reliability in DNNs has generally been overlooked. As the underlying semiconductor technology decreases in reliability, the probability that some components of computing devices fail also increases, preventing high accuracy in DNN operations. To achieve high accuracy, ensuring operational reliability, even if faults occur, is necessary. Jae-San Kim, Joon-Sung Yang |
DAC | 2 |
| 2019 | MRLoc: Mitigating Row-hammering based on memory LocalityabstractWith the increasing integration of semiconductor design, many problems have emerged. Row-hammering is one of these problems. The row-hammering effect is a critical issue for reliable memory operation because it can cause some unexpected errors. Hence, it is necessary to address this problem. Mainly, there are two different methods to deal with the row-hammering problem. One is a counter based method, and the other is a probabilistic method. This paper proposes the improved version of the latter method and compares it with other probabilistic methods, PARA and PRoHIT. According to the evaluation results, comparing the proposed method with conventional ones, the proposed one has increased row-hammering reduction per refresh 1.82 and 7.78 times against PARA and PRoHIT in average, respectively. Jung Min You, Joon-Sung Yang |
DAC | 2 |
| 2019 | Flipcy: Efficient Pattern Redistribution for Enhancing MLC PCM Reliability and Storage DensityabstractPhase change memory (PCM) is a scalable, non-volatile emerging memory. The storage density of the PCM can be enhanced by using the multi-level cell (MLC) operation. However, the MLC PCM suffers from low reliability due to resistance drift. The rate of resistance drift is proportional to the initial resistance of the cell with intermediate storage levels being particularly vulnerable. Using heavy Error Correction Codes (ECC) results in poor effective storage density (data bits per cell) of the MLC PCM. The state-of-the-art Tri-level cell technique improves reliability by using only three out of four storage levels, thus eliminating the ECC overhead. However, its storage density is much less than the ideal MLC PCM. Moreover, its storage density is fixed even if practically the MLC PCM reliability is improved. This paper introduces a more flexible pattern redistribution technique, Flipcy, to improve the MLC PCM reliability and effective storage density. The proposed method proportions the data-patterns according to the rate of resistance drift for different storage levels. A simple flip or a complement operation is used to reduce the percentage of the most error-prone pattern. The simulation results show up to 107X reduction in the error rate and 31% improvement in performance compared to the conventional MLC PCM. With a reduced overhead of the auxiliary bits and ECC parity bits, the proposed method can achieve about 25% improvement in effective storage density over the Tri-level cell approach for a similar level of reliability while incurring about 11% degradation in performance compared to the Tri-level cell approach. This performance degradation can be reduced when the proposed method is accompanied with orthogonal techniques to improve MLC PCM reliability and efficient scrubbing methods. Muhammad Imran 0010, Jung Min You, Joon-Sung Yang |
ICCAD | 4 |
| 2019 | Simplifying Deep Neural Networks for FPGA-Like Neuromorphic SystemsabstractDeep learning using deep neural networks is taking machine intelligence to the next level in computer vision, speech recognition, natural language processing, etc. Brain-like hardware platforms for the brain-inspired computational models are being studied, but the maximum size of neural networks they can evaluate is often limited by the number of neurons and synapses equipped with the hardware. This paper presents two techniques, factorization and pruning, that not only compress the models but also maintain the form of the models for the execution on neuromorphic architectures. We also propose a novel method to combine the two techniques. The proposed method shows significant improvements in reducing the number of model parameters over standalone use of each method while maintaining the performance. Our experimental results show that the proposed method can achieve 30× reduction rate within 1% budget of accuracy for the largest layer of AlexNet. Jaeyong Chung, Taehwan Shin, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Weight Partitioning for Dynamic Fixed-Point Neuromorphic Computing SystemsabstractNeuromorphic computing systems consist of neurons and synapses with limited programmability, and neural networks are modified to be mapped for such a system. In order to map a perceptron with a large number of connections into a hardware neuron with a fixed, small number of synapses, it is decomposed into a tree of perceptrons, which substantially affects the neuron usage and predictive performance. In this paper, we propose two decomposition algorithms that take advantage of plastic connections and the dynamic scaling capability of neurons. One algorithm based on sorting considers the neuron usage first, and the other algorithm based on packing considers the predictive performance first. Our experimental results on two popular deep convolutional neural networks showed that the sorting-based algorithm substantially improved the accuracy at no cost of neurons compared to a previous work, and the packing-based algorithm improved it even further at a small cost of neurons. Yongshin Kang, Joon-Sung Yang, Jaeyong Chung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | System level performance analysis and optimization for the adaptive clocking based multi-core processorabstractA supply voltage droop, temperature variation and aging effects can generate timing failures during operation. Various adaptive clocking methods have been introduced to resolve the problems. They use a tunable clock to avoid the timing failures rather than using wide design guard bands. However, the system performance analysis becomes complicated in a multi-core system with the adaptive clocking method. In this paper, a queueing theory based system level performance model is proposed to estimate an average response time and power by a closed form equation. Furthermore, for multi-core system with the adaptive clocking, an optimal job scheduling method using the inequality of arithmetic and geometric means is proposed. The proposed optimal job scheduling method relieves a system performance degradation arising from the adaptive clocking. The proposed performance model can analyze the system level performance within ~3% error compared with a JMT system simulation tool. Experimental results also show that the proposed job scheduling method can obtain a significant performance enhancement than the conventional round-robin method. Byung-Su Kim, Joon-Sung Yang |
ASP-DAC | 2 |
| 2018 | Test cost reduction for X-value elimination by scan slice correlation analysisabstractX-values in test output responses corrupt an output response compaction and can cause a fault coverage loss. X-Masking and X-Canceling MISR methods have been suggested to eliminate X-values, however, there are control data volume and test time overhead issues. These issues become significant as the complexity and the density of the circuits increase. This paper proposes a method to eliminate X's by applying a scan slice granularity X-value correlation analysis. The proposed method exploits scan slice correlation analysis, determines unique control data for the scan slice groups sharing the same control data, and applies them for each scan slice. Hence, the volume of control data can be significantly reduced. The simulation results demonstrate that the proposed method achieves greater control data and test time reduction compared to the conventional methods, without loss of fault coverage. Hyunsu Chae, Joon-Sung Yang |
DAC | 2 |
| 2018 | Optimized I/O determinism for emerging NVM-based NVMe SSD in an enterprise systemabstractNon-volatile memory express (NVMe) over peripheral component interconnect express (PCIe) has been adopted in the storage system to provide low latency and high throughput. NVMe allows a host system to reduce latency because it offers a high parallel operation and optimized command processing flow. In addition, an introduction of emerging non-volatile memory (NVM) significantly reduces the solid state drive (SSD) latency. The latency reduction in the host system and SSD makes a relative ratio of PCIe fabric latency to total I/O latency considerably grow. Therefore, this paper proposes a novel I/O optimization method using the PCIe feature, virtual channel. Unlike conventional approaches with the same priority data path, based on SSD's internal latency, an emerging NVM-based NVMe SSD with the proposed architecture selects a prioritized virtual channel to provide deterministic I/O latency. Experimental results show that the proposed method with phase-change memory (PCM) SSD improves I/O determinism by processing 45 ∼ 74% more commands within the predictable I/O latency than a conventional PCM SSD. Seonbong Kim, Joon-Sung Yang |
DAC | 2 |
| 2018 | Bayesian theory based switching probability calculation method of critical timing path for on-chip timing slack monitoringabstractAccurate in-situ monitoring is urgently required for an adaptive performance control system and post silicon validation. For accurate in-situ monitoring, a direct probing method is presented in which monitors directly measure a path delay from real critical timing paths. However, we may not be able to predict when the timing slack monitors would activate since the activation depends on a design structure and input patterns. If a timing slack monitor is rarely activated by timing critical paths, the observability from this monitor would be low and the monitor possibly can be discarded. For this reason, we propose a novel timing slack monitoring methodology based on switching probability of timing critical paths. Switching probability and correlation on critical timing paths are formulated, and the proposed method finds a list of critical path endpoints for the timing slack monitor insertion under given power and area constraints. Experimental results with ISCAS'89 circuits show that, compared to the method which places monitors for all worst critical paths, 16.67 ~ 97.2% of timing slack monitors are removed and 32.56 ~ 96.88% of dynamic power reduction from the monitors is achieved by the proposed method. Byung-Su Kim, Joon-Sung Yang |
DATE | 2 |
| 2018 | Heterogeneous PCM array architecture for reliability, performance and lifetime enhancementabstractConventional DRAM and flash memory are reaching their scaling limits thus motivating research in various emerging memory technologies as a potential replacement. Among these, phase change memory (PCM) has received considerable attention owing to its high scalability and multi-level cell (MLC) operation for high storage density. However, due to the resistance drift over time, the soft error rate in MLC PCM is high. Additionally, the iterative programming in MLC negatively impacts performance and cell endurance. The conventional methods to overcome the drift problem incur large overheads, impact memory lifetime and are inadequate in terms of acceptable soft error rate (SER). In this paper, we propose a new PCM memory architecture with heterogeneous PCM arrays to increase reliability, performance and lifetime. The basic storage unit in the proposed architecture consists of two single-level cells (SLCs) and one four-level cell (4LC). Using the reduced number of 4LCs compared to conventional homogeneous 4LC PCM arrays, the drift-induced error rate is considerably reduced. By alternating each cell operation between SLC and 4LC over time, the overall lifetime can also be significantly enhanced. The proposed architecture achieves up to 105times lower soft error rate with considerably less ECC overhead. With simple ECC scheme, about 22% performance improvement is achieved and additionally, the overall lifetime is also enhanced by about 57%. Muhammad Imran 0010, Jung Min You, Joon-Sung Yang |
DATE | 4 |
| 2018 | READ: Reliability Enhancement in 3D-Memory Exploiting Asymmetric SER Distributionabstract3D-memory is one of promising applications in 3D-IC technology. With a 3D integration technology, the effective density of memories can increase and the interconnect distance from processor to memory can be shortened. Due to its stacked structure, the upper dies behave as shields blocking outer particles from reaching lower dies, and it makes error rate of the top layer largest among all layers. From a heat perspective, the lower dies would suffer from reliability problems since the lower dies are placed on top of logic die. The heat dissipation can more influence lower dies than upper dies. This creates unequal a reliability distribution for each layer in 3D-memories. A novel ECC organization scheme for 3D-memory to secure reliable operations under soft error rate (SER) profiles is introduced in this paper. The proposed scheme does not require additional redundant arrays. Instead, it utilizes unused spare columns of relatively reliable layer memories to store additional check-bits of less reliable layer memories. It forms a heterogeneous ECC organization across different layers which enhances ECC capabilities in less reliable layers. In addition, redundancy sharing scheme for yield enhancement can be implemented with the proposed scheme. Experimental results show that a memory with the proposed method can tolerate more than three times of a bit-error rate compared to the conventional memory. Hyunseung Han, Jaeyong Chung, Joon-Sung Yang |
IEEE Trans. Computers | 3 |
| 2018 | Mitigating Observability Loss of Toggle-Based X-Masking via Scan Chain PartitioningabstractThe Toggle-based X-masking method requires a single toggle at a given cycle, there is a chance that non-Xvalues are also masked. Hence, the non-Xvalue over-masking problem may cause a fault coverage degradation. In this paper, a scan chain partitioning scheme is described to alleviate non-Xbit over-masking problem arising from Toggle-based X-Masking method. The scan chain partitioning method finds a scan chain combination that gives the least toggling conflicts. The experimental results show that the amount of over-masked bits is significantly reduced, and it is further reduced when the proposed method is incorporated with X-canceling method. However, as the number of scan chain partitions increases, the control data for decoder increases. To reduce a control data overhead, this paper exploits a Huffman coding based data compression. Assuming two partitions, the size of control bits is even smaller than the conventional X-toggling method that uses only one decoder. In addition, selection rules of X-bits delivered to X-Canceling MISR are also proposed. With the selection rules, a significant test time increase can be prevented. Sae-Eun Kim, Jaeyong Chung, Joon-Sung Yang |
IEEE Trans. Computers | 3 |
| 2018 | Clock Network Optimization With Multibit Flip-Flop Generation Considering Multicorner Multimode Timing ConstraintabstractClock network should be optimized to reduce clock power dissipation. The power efficient clock network can be constructed by multibit flip-flop generation and gated clock tree aware flip-flop clumping to pull flip-flops close to the same integrated clock gating cell. It is capable of providing an attractive solution to reduce clock power. This paper considers multicorner and multimode timing constraints for the two combined approach. This proposed method is applied to five industrial digital intellectual property blocks of state-of-the-art mobile system-on-a-chip fabricated in 14-nm CMOS process. Experimental results show that MBFF generation algorithm achieves 22% clock power reduction. Applying a gated clock tree aware flip-flop clumping on top of the MBFF generation further reduces the power to around 32%. Taehee Lee 0003, David Z. Pan, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Improving NVMe SSD I/O determinism with PCIe virtual channel: work-in-progressabstractNVMe SSD over PCIe is attractive since it provides high throughput and low latency. However, complex internal SSD operations may cause a non-deterministic I/O latency which is one of the most important factors in a storage system. While conventional approaches to enhance I/O latency prediction are based on host systems, this paper proposes a novel SSD-based deterministic latency enhancement scheme. The proposed method exploits the fact that multiple virtual channels can be utilized. For each virtual channel, the proposed method assigns a different priority for data transmission. NVMe SSD analyses its internal latency and dynamically chooses the virtual channels to compensate the latency. The experimental results show that, using a PCIe switch model, the proposed method can save 41.6% of the latency for each transaction layer packet. Seonbong Kim, Joon-Sung Yang |
CASES | 2 |
| 2017 | MVP ECC : Manufacturing process variation aware unequal protection ECC for memory reliabilityabstractWith a development of process technology, a memory density has been increased. However, a smaller feature size makes the memory susceptible to soft errors. For reliability enhancement, ECC with single bit error correction and double bit error detection is widely used. As multiple bit cell upset become dominant, there is a need for stronger ECC. ECC such as RS or BCH code requires significantly large overhead and longer latency. To overcome the problem, this paper introduces an unequal protection ECC assigning stronger level of protection to weak memory cells and normal level to normal cells. Information from manufacturing characterization test is utilized to identify weak memory cells with low design margins. Instead of equally treating all memory cells, the proposed ECC focuses more on the weak cells since they are more susceptible to soft errors. Compared to conventional ECCs, experimental results show that the proposed ECC considerably enhances memory reliability with the same code length. Seung-Yeob Lee, Joon-Sung Yang |
DATE | 2 |
| 2017 | PUFSec: Device fingerprint-based security architecture for Internet of ThingsabstractA low-end embedded platform for Internet of Things (IoT) often suffers from a critical trade-off dilemma between security enhancement and computation overhead. We propose PUFSec, a new device fingerprint-based security architecture for IoT devices. By leveraging intrinsic hardware characteristics, we aim to design a computationally lightweight security software system architecture so that complex cryptography computation can dramatically be prohibited. We exploit the innovative idea of Public Physical Unclonable Functions (PPUFs) that fundamentally protects attackers from recovering the secret key from public gate delay information. We implement its hardware logic in a real-world FPGA board. On top of the PPUF fingerprint hardware, we present an adaptive security control mechanism consisting of adaptive key generation and key exchange protocol, which adjusts security strength depending on system load dynamics. We demonstrate that our PPUF FPGA implementation embeds distinctive variability enough to distinguish between two different PPUFs with high fidelity. We validate our PUFSec architecture by implementing necessary algorithms and protocols in a real-world IoT platform, and performing empirical evaluation in terms of computation and memory usages, proving its practical feasibility. So-Yeon Park, Sunil Lim, Dahee Jeong, Jungjin Lee, Joon-Sung Yang, HyungJune Lee |
INFOCOM | 5 |
| 2017 | Non-linear library characterization method for FinFET logic cells by L1-minimizationabstractState-of-the-art process technology offers ultra-low power devices operating at ultra-low voltages. However, they show a considerable level of non-linear characteristics. Hence, the accuracy of cell delay and variation modeling for logic cells is expected to be very low with a linear interpolation. In this paper, we propose a compressive sensing based high non-linear cell delay and variation modeling. This paper introduces accuracy optimization methods to fit the delay and variation modeling by pre-processing. Pre-processing is a hybrid approach combining a linear interpolation and compressive sensing for accurate restoration with using less samples. With FinFET cell delay and variation modeling, the experimental results show that the proposed method can obtain a similar or better accuracy with a half of measurement samples than a conventional linear interpolation based modeling. Byung-Su Kim, Hyo-Sig Won, Tae Hee Han, Joon-Sung Yang |
ISCAS | 4 |
| 2017 | Exploiting Unused Spare Columns and Replaced Columns to Enhance Memory ECCabstractDue to the emergence of extremely high density memory along with the growing number of embedded memories, memory yield is an important issue. Memory self-repair using redundancies to increase the yield of memories is widely used. Because high density memories are vulnerable to soft errors, memory error correction code (ECC) plays an important role in memory design. In this paper, methods to exploit spare columns including replaced defective columns are proposed to improve memory ECC. To utilize replaced defective columns, the defect information needs to be stored. Two approaches to store defect information are proposed-one is to use a spare column and the other is to use a content-addressable-memory. Experimental results show that the proposed method can significantly enhance the ECC performance. Hyunseung Han, Nur A. Touba, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Enhancing Test Compression With Dependency Analysis for Multiple Expansion RatiosabstractScan test data compression is widely used in industry to reduce test data volume (TDV) and test application time (TAT). This paper shows how multiple scan chain expansion ratios can help to obtain high test data compression in system-on-chips. Scan chains are partitioned with a higher expansion ratio than normal in scan compression mode and then are gradually concatenated based on a cost function to detect any faults that could not be detected at the higher expansion ratios. It improves the overall test compression ratio since it potentially allows faults to be detected at the highest expansion ratio. This paper introduces a new cost function to choose scan chain concatenation candidates for concatenation for multiple expansion ratios. To avoid TDV and TAT increase by scan concatenation, the proposed method takes a logic structure and scan chain length into consideration. Experiment results show the proposed method reduces TAT and TDV by 53%-64% compared with a traditional scan compression method. Taehee Lee 0003, Nur A. Touba, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Reducing control bit overhead for X-masking/X-canceling hybrid architecture via pattern partitioningabstractAn X-masking scheme prevents unknown (X) values from shifting into an output response compactor, whereas an X-canceling MISR methodology allows X's to enter the compactor, but then cancels them out through selective XORing. However, both approaches require significantly high volume of the control bits to remove X values to generate X-free output signatures. This paper proposes a method to reduce the control bit overhead by combining X-masking and X-canceling methodologies and exploiting the fact that unknown values tend to have high correlation in the scan cells. In this paper, correlation is considered across whole patterns in order to enhance reuse of control bits. The proposed hybrid method of X-canceling and X-masking reduces test time without losing fault coverage. The experimental results show that the proposed method significantly reduces control bits and test time compared to a conventional X-canceling MISR methodology. Jin-Hyun Kang, Nur A. Touba, Joon-Sung Yang |
DAC | 3 |
| 2016 | AFSEM: Advanced frequent subcircuit extraction method by graph mining approach for optimized cell library developmentsabstractThe optimization of cells and cell combinations used in design is critical to enhance the performance. If frequently used cell combinations are known in advance, a new cell development can be significantly optimized using the cell combinations for chip design. However, extracting frequent cell combinations is an NP hard problem. We propose a new framework, referring as AFSEM, to extract frequent cell combinations for design optimization. To solve this problem, we use a frequent subgraph mining method which is a process of discovering subgraphs. We present an advanced graph modeling and optimized frequent subgraph mining platform for a practical use. The experimental results with various designs demonstrate that the proposed method can discover various types of subcircuits for design optimization with various runtime optimization methods. Byung-Su Kim, Hyo-Sig Won, Tae Hee Han, Joon-Sung Yang |
ISCAS | 4 |
| 2016 | Multi-bit flip-flop generation considering multi-corner multi-mode timing constraintabstractClock power is a significant portion of chip power in System-on-chip (SoC). Applying Multi-bit flip-flop (MBFF) is capable of providing attractive solution to reduce clock power. To our best knowledge, this is the first work in the literature that considers multi-corner and multi-mode (MCMM) timing constraint for the MBFF generation. This proposed method is applied to five industrial digital intellectual property (IP) blocks of state-of-the-art System-on-chip (SoC). Experimental results show that our proposed MBFF generation algorithm achieves 22% clock power reduction. Taehee Lee 0003, JongWon Yi, Joon-Sung Yang |
ISCAS | 3 |
| 2016 | Enhancing Superset X-Canceling Method With Relaxed Constraints on Fault ObservationabstractAn X-tolerant multiple-input signature register (MISR) compaction methodology that compacts output streams containing unknown (X) values, called X-canceling, is an alternative to masking X values (i.e., X-masking). A number of control bits that is linear in the number of X's to be canceled are required to perform the X-canceling operation for existing X-canceling approaches. This paper proposes a new X-canceling method significantly reducing the number of control bits for X-canceling. We exploit the fact that 1) unknown values tend to be highly correlated in the scan cells (i.e., X's tend to be generated in certain portions of design) and 2) fault effects can typically be observed in a multiplicity of scan cells. Instead of custom generating the control bits to cancel out only the X's in one MISR signature, the proposed approach finds a general superset solution which can cancel out the X's for many MISR signatures without losing fault coverage. This allows the same control bits to be reused many times thereby significantly improving the amount of compression that can be obtained. Architectures for implementing superset X-canceling are described along with experimental results. Joon-Sung Yang, Jinsuk Chung, Nur A. Touba |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Robust via-programmable ROM design based on 45nm process considering process variation and enhancement Vmin and yieldabstractThis paper presents a Via programmable read only memory (Via-ROM) for Vminand macro-yield enhancement through robust ROM designs based on 45nm process. The main stability issues in ROM are 1) lower on-cell (NMOS) current, 2) higher keeper (PMOS) current, and 3) higher bit-line (BL) parasitic value. To improve the Vminand macro-yield, the robust ROM design schemes are implemented as follows. 1) ROM bit cell size optimization without increasing a bit cell area, 2) BL loading reduction to use a rom code pattern optimization, 3) selective full BL pre-charge and keeper control to use an external pin named as KCS (Keeper Control Signal) and 4) wide pulse width generator using an asynchronous 3-bit ripple binary counter. These schemes to improve 0 read margin were confirmed by both the simulation and the measurement. Experimental results show that macro-yield improved from 0% to 100% at 1.1V (Voperation) and -40°C. Byung-Jun Jang, Chan-Ho Lee, Sung-Hun Sim, Kyu-Won Choi, Do-Hun Byun, Yeon-Ho Jung, Ki-Man Park, Dong-Yeon Heo, Gyu-Hong Kim, Joon-Sung Yang |
ISCAS | 10 |
| 2015 | Devil in a box: Installing backdoors in electronic door locksabstractElectronic door locks must be carefully designed to allow valid users to open (or close) a door and prevent unauthorized people from opening (or closing) the door. However, lock manufacturers have often ignored the fact that door locks can be modified by attackers in the real world. In this paper, we demonstrate that the most popular electronic door locks can easily be compromised by inserting a malicious hardware backdoor to perform unauthorized operations on the door locks. Attackers can replay a valid DC voltage pulse to open (or close) the door in an unauthorized manner or capture the user's personal identification number (PIN) used for the door lock. Seongyeol Oh, Joon-Sung Yang, Andrea Bianchi, Hyoungshick Kim |
PST | 2 |
| 2014 | 3-D Probe: Low-Cost Variation Modeling Using Intertest-Item CorrelationsabstractProcess variation models for variation tolerant designs are developed through expensive silicon characterization. This paper presents a low-cost variation characterization method that takes advantage of correlations between test items. The proposed method is based on compressed sensing (CS), a new innovative theory in signal processing and information theory, and we formulate the problem of accounting for the correlations in the form of standard CS problems, allowing us to leverage advances in CS theory. We consider wafer-level measurement results for multiple test items a 3-D signal and propose the sparsifying transform that combines the 2-D discrete cosine transform and the Karhunen-Loéve transform. Our experimental results show that the proposed method reduces the number of samples required for the same accuracy up to 2X compared to virtual probe when two test items are used. Jaeyong Chung, Joon-Sung Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Utilizing ATE Vector Repeat With Linear Decompressor for Test Vector CompressionabstractPrevious approaches for utilizing automatic test equipment (ATE) vector repeat are based on identifying runs of repeated scan data and directly generating that data using ATE vector repeat. Each run requires a separate vector repeat instruction, so the amount of compression is limited by the amount of ATE instruction memory available and the length of the runs (which typically will be much shorter than the length of a scan vector). In this paper, a new and more efficient approach is proposed for utilizing ATE vector repeat. The scan vector sequence is partitioned and decomposed into a common sequence which is the same for an entire cluster of test cubes and a unique sequence that is different for each test cube. The common sequence can be generated very efficiently using ATE vector repeat. Experimental results demonstrate that the proposed approach can achieve much greater compression while using many fewer vector repeat instructions compared with previous methods. Joon-Sung Yang, Jinkyu Lee 0005, Nur A. Touba |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Enhanced algorithm of combining trace and scan signals in post-silicon validationabstractAs the complexity of integrated circuit design increases and production schedules become shorter, the dependency on post-silicon validation for capturing design errors that escape from pre-silicon verification also increases. A major challenge in post-silicon validation is the limited observability of internal states caused by the limited storage capacity available for post-silicon validation. Recent research has shown that observability can be enhanced if trace and scan signals are combined together, compared with the debugging scenario where only trace signals are monitored. This paper proposes an enhanced and systematic algorithm for the efficient combination of trace and scan signals to maximize the observability of internal circuit states. Experimental results on benchmark circuits show that the proposed technique provides a higher number of restored states compared to the existing techniques. Kihyuk Han, Joon-Sung Yang, Jacob A. Abraham |
VTS | 2 |
| 2013 | Improved Trace Buffer Observation via Selective Data Capture Using 2-D Compaction for Post-Silicon DebugabstractThis paper presents a novel technique for extending the capacity of trace buffers when capturing debug data during post-silicon debug. It exploits the fact that is it not necessary to capture error-free data in the trace buffer since that information can be obtained from simulation. A selective data capture method is proposed in this paper that only captures debug data during clock cycles in which errors are present. The proposed debug method requires only three debug sessions. The first session estimates a rough error rate, the second session identifies a set of suspect clock cycles where errors may be present, and the third session captures the suspect clock cycles in the trace buffer. The suspect clock cycles are determined through a 2-D compaction technique using multiple-input signature register signatures and cycling register signatures. Intersecting both signatures generates a small number of suspect clock cycles for which the trace buffer needs to capture. The effective observation window of the trace buffer can be expanded significantly, by up to orders of magnitude. Experimental results indicate very significant increases in the effective observation window for a trace buffer can be obtained. Joon-Sung Yang, Nur A. Touba |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Test Point Insertion with Control Points Driven by Existing Functional Flip-FlopsabstractThis paper presents a novel test point insertion method for pseudorandom built-in self-test (BIST) to reduce the area overhead. The proposed method replaces dedicated flip-flops for driving control points by existing functional flip-flops. For each control point, candidate functional flip-flops are identified by using logic cone analysis that investigates the path inversion parity, logical distance, and reconvergence from each control point. Four types of new control point structures are introduced based on the logic cone analysis results to avoid degrading the testability. Experimental results indicate that the proposed method significantly reduces test point area overhead by replacing the dedicated flip-flops and achieves essentially the same fault coverage as conventional test point implementations using dedicated flip-flops driving the control points. Joon-Sung Yang, Nur A. Touba, Benoit Nadeau-Dostie |
IEEE Trans. Computers | 1 |
| 2012 | Efficient Trace Signal Selection for Silicon Debug by Error Transmission AnalysisabstractIn this paper, a technique is presented for selecting signals to observe during silicon debug. Internal signals are used to analyze, understand, and debug circuit misbehavior. An automated procedure to select which signals to observe is proposed to facilitate early detection of circuit malfunction and to enhance the utilization of hardware resources for storage. Signals that are most often sensitized to possible errors are observed in sequential circuits. Given a functional input vector set, an error transmission matrix is generated by analyzing which flip-flops are sensitized to other flip-flops. Relatively independent flip-flops are identified and a set of signals that maximally cover the possible error sites with given constraints are identified through integer linear programming. Experimental results show that the proposed approach can rapidly and precisely identify the nonconforming chip behavior and thereby can speed up the post-silicon debug process. Joon-Sung Yang, Nur A. Touba |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | X-Canceling MISR Architectures for Output Response Compaction With Unknown ValuesabstractIn this paper, anX-tolerant multiple-input signature register (MISR) compaction methodology that compacts output responses containing unknownXvalues is described. Each bit of the MISR signature is expressed as a linear combination in terms ofXs by symbolic simulation. Linearly dependent combinations of the signature bits are identified with Gaussian elimination and XORed to removeXvalues and yield deterministic values. TwoX-canceling MISR architectures are proposed and analyzed with industrial designs. This paper also shows the correlation between the estimated result based on idealized modeling and the actual data for real circuits for error coverage, hardware overhead, and other metrics. Experimental results indicate that high error coverage can be achieved withX-canceling MISR configurations and it highly correlates with actual results. Joon-Sung Yang, Nur A. Touba |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Test point insertion using functional flip-flops to drive control pointsabstractThis paper presents a novel method for reducing the area overhead introduced by test point insertion. Test point locations are calculated as usual using a commercial tool. However, the proposed method uses functional flip-flops to drive control test points instead of test-dedicated flip-flops. Logic cone analysis that considers the distance and path inversion parity from candidate functional flip-flops to each control point is used to select an appropriate functional flip-flop to drive the control point which avoids adding additional timing constraints. Reconvergence is also checked to avoid degrading the testability. Experimental results indicate that the proposed method significantly reduces test point area overhead and achieves essentially the same fault coverage as the implementations using dedicated flip-flops driving the control points. Joon-Sung Yang, Benoit Nadeau-Dostie, Nur A. Touba |
ITC | 1 |
| 2009 | An industrial case study for X-canceling MISRabstractAn X-tolerant multiple-input signature register (MISR) compaction methodology that compacts output streams containing unknown (X) values was described in [Touba 07]. Unlike conventional approaches, it does not use X-masking logic at the input of the MISR. Instead it uses symbolic simulation to express each bit of the MISR signature as a linear equation in terms of the X's. Linearly dependent combinations of the signature bits are identified with Gaussian elimination and XORed together to cancel out all X values and yield deterministic values. This new X-canceling approach was applied to some industrial designs under the constraints imposed by an industrial test environment. Practical issues for implementing X-canceling are discussed, and a new architecture for implementing X-canceling based on using a shadow register with multiple selective XORs is presented. Experimental results are shown for industrial designs comparing the performance of X-canceling with X-compact. Joon-Sung Yang, Nur A. Touba, Shih-Yu Yang |
ITC | 1 |
| 2009 | Automated Selection of Signals to Observe for Efficient Silicon DebugabstractInternal signals of a circuit are observed to analyze, understand, and debug nonconforming chip behavior. The number of signals that can be observed is limited by bandwidth and storage requirements. This paper presents an automated procedure to select which signals to observe to facilitate early detection of circuit malfunction to help find the root cause of a bug. This paper exploits the nature of error propagation in sequential circuits by observing signals which are most often sensitized to possible errors. Given a functional input vector set, an error transmission matrix is generated by analyzing which flip-flops are sensitized to other flip-flops. Signal observability is enhanced by merging data from relatively independent flip-flops. The final set of signals to observe is determined through integer linear programming (ILP) which provides a set of locations that maximally cover the possible error sites within given constraints. Experimental results indicate that the cycle in which a bug first appears can be more rapidly and precisely found with the proposed approach thereby speeding up the post-silicon debug process. Joon-Sung Yang, Nur A. Touba |
VTS | 1 |
| 2008 | Expanding Trace Buffer Observation Window for In-System Silicon Debug through Selective CaptureabstractTrace buffers are commonly used to capture data during in-system silicon debug. This paper exploits the fact that it is not necessary to capture error-free data in the trace buffer since that information is obtainable from simulation. The trace buffer need only capture data during clock cycles in which errors are present. A three pass methodology is proposed. During the first pass, the rough error rate is measured, in the second pass, a set of suspect clock cycles where errors may be present is determined, and then in the third pass, the trace buffer captures only during the suspect clock cycles. In this manner, the effective observation window of the trace buffer can be expanded significantly, by up to orders of magnitude. This greatly increases the effectiveness of a given size trace buffer and can rapidly speed up the debug process. The suspect clock cycles are determined through a two dimensional (2-D) compaction technique using a combination of multiple-input signature register (MISR) signatures and cycling register signatures. By intersecting the signatures, the proposed 2-D compaction technique generates a small set of remaining suspect clock cycles for which the trace buffer needs to capture data. Experimental results indicate very significant increases in the effective observation window for a trace buffer can be obtained. Joon-Sung Yang, Nur A. Touba |
VTS | 1 |