VLDB 2026 Research / reviewers in the wild / expert
Ing-Chao Lin
dblp:53/2026
· DBLP profile ↗
43ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0003-1994-7512ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 10 first-author · 19 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | G-PathGen: An Efficient GPU-Parallel k-Critical Path Generation AlgorithmabstractCritical path generation (CPG) plays a key role in many circuit timing analysis (CTA) applications. As the design complexity continues to increase, CPG runtime has become a major bottleneck in many timing-driven applications. To mitigate this runtime challenge, several CPU-based algorithms have been introduced by both the CTA and parallel computing communities, but they remain slow for large CPG problems. While GPU-accelerated solutions exist, they are often inexact and incur significant overhead from iterative CPU–GPU data transfers, limiting their practical use in CTA applications. To overcome this challenge, we propose G-PathGen, an exact GPU-parallel CPG algorithm targeting CTA applications. G-PathGen introduces efficient kernel algorithms for generating critical paths in parallel and dynamically adjusts the generated path count to maximize GPU utilization while minimizing redundant work. Compared to a state-of-the-art GPU solution, G-PathGen is 1.6 × –243.8 × faster when generating one million critical paths on industrial circuit graphs. Che Chang, Yi-Hua Chung, Cheng-Hsiang Chiu, Wan-Luan Lee, Boyang Zhang 0007, Ulf Schlichtmann, Ing-Chao Lin, Xiangyao Yu, Tsung-Wei Huang |
ICS | 7 |
| 2026 | TT-RRAM: Joint Improvement of Sparsity and Repetition in Tensor-Train Inference With Fused Processing on RRAM AcceleratorsabstractDeep neural networks compressed using Tensor-Train decomposition , known as TT-DNNs, significantly reduce model size to decrease storage requirements but suffer from more frequent data movement during inference. To address this issue, RRAM-based DNN accelerators, by leveraging computing-in-memory, can mitigate the data movement and efficiently execute the VMM operations required during inference, making them an ideal candidate for accelerating the inference for TT-DNNs. However, RRAM-based accelerators face three challenges in TT-format DNN inference: (1) weights distribution on the RRAM crossbar with low sparsity and repetition; (2) low input sparsity; and (3) high data dependencies and storage overhead during inference. To address these challenges, we propose three corresponding methods: (1) Base-Offset Splitting (BOS) to improve the weights distribution on the crossbar, thereby enhancing performance potential; (2) Inverse Input Reusing (IIR) to enhance the sparsity of inputs during inference; and (3) Fused Scheduling (FS) to reorganize the computation order and leveraging the parallel processing capabilities of the RRAM-based accelerator to improve overall inference efficiency while also reducing data dependencies and storage overhead. Experimental results demonstrate that our proposed methods achieve a 3.44× to 5.53× performance improvement and 55.6% to 67.1% energy savings compared to the state-of-the-art RRAM-based accelerator, while also reducing storage overhead by 66.7% to 99.1%. Fang-Yi Gu, Pin-Hong Liu, Ing-Chao Lin, Bo Yuan 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | SegTransformer: Enhancing Softmax Performance Through Segmentation with a ReRAM-Based PIM AcceleratorabstractReRAM-based Processor-In-Memory (PIM) architectures have demonstrated their ability to accelerate the matrix multiplication performed by the Transformer. However, these approaches often shift the performance bottleneck from the attention mechanism to the softmax computation. Moreover, data sharding, commonly employed for acceleration, hinters the Transformer's ability to find global softmax value. This paper proposes SegTransformer, a ReRAM-based PIM accelerator that improves softmax performance with high accuracy through segmentation techniques. Experimental results indicate that the SegTransformer significantly outperforms state-of-the-art Transformer accelerators. Yu Chen Wang, Ing-Chao Lin, Yuan-Hao Chang 0001 |
DATE | 2 |
| 2025 | Optimizing Reliability and Energy Efficiency in Heterogeneous Multicore Systems: A Novel Task Deployment StrategyabstractThe rapid advancement of CMOS technology has empowered modern integrated circuits (ICs) to handle increasingly complex tasks, which are crucial across various applications. However, achieving a balance between performance and energy consumption remains a challenge, particularly for edge devices such as cell phones and IoT sensors. Heterogeneous multicore systems, combining high-performance and energy-efficient cores, offer a promising solution for optimizing power while meeting performance requirements. Nevertheless, the continuous evolution of technology exacerbates reliability concerns in multicore systems, notably transient errors and aging effects. Previous studies have proposed various strategies, including task replication and Dynamic Voltage and Frequency Scaling (DVFS), to mitigate reliability degradation. However, the lack of holistic consideration for these mitigation techniques may inadvertently shorten the system's lifespan due to potential conflicts between them. Hence, this paper thoroughly analyzes the interplay between strategies for mitigating transient errors and aging effects. Subsequently, we introduce a novel reliability-aware task deployment framework aimed at significantly extending the system lifespan while adhering to a specified reliability target and minimizing overall energy consumption. Experimental findings demonstrate the efficacy of our framework, showcasing a 4.98x improvement in system lifespan and a 50% reduction in energy consumption compared to prior studies. Yu-Guang Chen, Yin-Rong Zhuo, Zheng-Wei Chen, Ing-Chao Lin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | A Lifetime Extension Framework for Communication-Intensive Systems Based on Wavelength-Routed Optical Networks-on-Chip
Zhidan Zheng, Liaoyuan Cheng, Jeng-De Chang, Tsun-Ming Tseng, Ing-Chao Lin, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | WROXIM: A Network-Level Simulation Platform for Wavelength-Routed Optical Networks-on-ChipabstractTo meet the increasing demand for high-speed communication in many-cores systems, optical networks-on-chip (ONoCs) have gained attention for their ability to deliver low-latency and high-bandwidth data transmission. As a specific type of ONoCs, wavelength-routed ONoCs (WRONoCs) offer exclusive benefits such as collision-free and arbitration-free communication between cores. To guide the design and optimization of WRONoCs, many simulators have been developed to model WRONoC behavior at the device and circuit levels, offering detailed insights into photonic components. These tools have significantly advanced photonic design. As WRONoCs move closer to practical applications, network-level simulation becomes increasingly important for evaluating end-to-end communication behavior. Despite the need, existing simulators rarely support WRONoC modeling at the network level, leaving a critical gap in current toolchains. To address that, we propose the first network-level WRONoC simulation platform, WROXIM, adapted from an open-source NoC simulator, Noxim. It supports cycle-accurate modeling of optical communication behaviors and retains full compatibility with traditional electrical NoC simulations. To model realistic communication behaviors, WROXIM incorporates both optical and electrical components, including serializers, deserializers, waveguides, buffers, and processing elements (PEs). This modeling approach allows the simulator to capture end-to-end data movement from a PE through the optical interconnect to another PE. Given an application with multiple tasks, WROXIM supports task-driven workloads and provides key performance metrics such as latency, throughput, and energy consumption, offering a platform for design exploration and verification for WRONoCs. Jeng-De Chang, Zhidan Zheng, Liaoyuan Cheng, Liu-Xuan-Wei Zhang, Tsun-Ming Tseng, Ing-Chao Lin, Ulf Schlichtmann |
ICCAD | 6 |
| 2025 | Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model CompressionabstractLarge Language Models (LLMs) have achieved remarkable breakthroughs. However, the huge number of parameters in LLMs require significant amount of memory storage in inference, which prevents their practical deployment in many applications. To reduce memory storage of LLMs, singular value decomposition (SVD) provides a promising solution to approximate weight matrices for compressing LLMs. In this paper, we take a step further to explore parameter sharing across different layers with SVD to achieve more effective compression for LLMs. Specifically, weight matrices in different layers are decomposed and represented with a linear combination of a set of shared basis vectors and unique coefficients. The types of weight matrices and the layer selection for basis sharing are examined when compressing LLMs to maintain the performance. Comprehensive experiments demonstrate that Basis-Sharing outperforms state-of-the-art SVD-based compression approaches, especially at large compression ratios. Jingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li 0005, Grace Li Zhang |
ICLR | 3 |
| 2025 | Efficient Model Switching in RRAM-Based DNN AcceleratorsabstractResistive random access memory (RRAM) has emerged as a promising technology for deep neural network (DNN) accelerators, but programming every weight in a DNN onto RRAM cells for inference can be both time-consuming and energy-intensive, especially when switching between different DNN models. This article introduces a hardware-aware multimodel merging (HA3M) framework designed to minimize the need for reprogramming by maximizing weight reuse, while taking into account the hardware constraints of the accelerator. The framework includes three key approaches: 1) crossbar (XB)-aware model mapping (XAMM); 2) block-based layer matching (BLM); and 3) multimodel retraining (MMR). XAMM reduces the XB usage of the preprogrammed model on RRAM XBs while preserving the model’s structure. BLM reuses preprogrammed weights in a block-based manner, ensuring the inference process remains unchanged. MMR then equalizes the block-based matched weights across multiple models. Experimental results show that the proposed framework significantly reduces programming cycles in multi-DNN switching scenarios while maintaining or even enhancing accuracy, and eliminating the need for reprogramming. Fang-Yi Gu, Ing-Chao Lin, Bing Li 0005, Ulf Schlichtmann, Grace Li Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Class-Aware Pruning for Efficient Neural NetworksabstractDeep neural networks (DNNs) have demonstrated remarkable success in various fields. However, the large number of floating-point operations (FLOPs) in DNNs poses challenges for their deployment in resource-constrained applications, e.g., edge devices. To address the problem, pruning has been introduced to reduce the computational cost in executing DNNs. Previous pruning strategies are based on weight values, gradient values and activation outputs. Different from previous pruning solutions, in this paper, we propose a class-aware pruning technique to compress DNNs, which provides a novel perspective to reduce the computational cost of DNNs. In each iteration, the neural network training is modified to facilitate the class-aware pruning. Afterwards, the importance of filters with respect to the number of classes is evaluated. The filters that are only important for a few number of classes are removed. The neural network is then retrained to compensate for the incurred accuracy loss. The pruning iterations end until no filter can be removed anymore, indicating that the remaining filters are very important for many classes. This pruning technique outperforms previous pruning solutions in terms of accuracy, pruning ratio and the reduction of FLOPs. Experimental results confirm that this class-aware pruning technique can significantly reduce the number of weights and FLOPs, while maintaining a high inference accuracy. Our code is available at https://github.com/HWAI-TUDa/Class-Aware-Pruning Mengnan Jiang, Jingcun Wang, Amro Eldebiky, Xunzhao Yin, Cheng Zhuo, Ing-Chao Lin, Grace Li Zhang |
DATE | 6 |
| 2024 | GNN-Based INC and IVC Co-Optimization for Aging MitigationabstractAs semiconductor processes advance, circuit aging becomes prominent. One of the most severe aging effects is Negative Bias Temperature Instability (NBTI), which increases the threshold voltage and the propagation delay of PMOS transistors. To mitigate NBTI, aging mitigation methods such as Internal Node Control (INC) and Input Vector Control (IVC) have been proposed. INC applies designed logic gates, while IVC uses appropriate input patterns during circuit idle. However, INC leads to extra area overhead and power consumption, and the circuit structure limits the controllability of IVC. Although various approaches have proposed aging tolerance methods with INC or IVC, only a few of them consider co-optimization. In this paper, we introduce a GNN-based INC and IVC co-optimization framework to minimize aging-induced delay. The key concept of our framework is using GNN to identify serious-aged gates in a circuit, and then using INC and IVC to mitigate the aging effect under an area overhead constraint. The experimental results indicate that our method reduces aging-induced delay and area by 2.16 times and 29.5%, respectively, compared to previous work. Yu-Guang Chen, Hsiu-Yi Yang, Ing-Chao Lin |
ETS | 3 |
| 2024 | BasisN: Reprogramming-Free RRAM-Based In-Memory-Computing by Basis Combination for Deep Neural NetworksabstractDeep neural networks (DNNs) have made breakthroughs in various fields including image recognition and language processing. DNNs execute hundreds of millions of multiply-and-accumulate (MAC) operations. To efficiently accelerate such computations, analog in-memory-computing platforms have emerged leveraging emerging devices such as resistive RAM (RRAM). However, such accelerators face the hurdle of being required to have sufficient on-chip crossbars to hold all the weights of a DNN. Otherwise, RRAM cells in the crossbars need to be reprogramed to process further layers, which causes huge time/energy overhead due to the extremely slow writing and verification of the RRAM cells. As a result, it is still not possible to deploy such accelerators to process large-scale DNNs in industry. To address this problem, we propose the BasisN framework to accelerate DNNs on any number of available crossbars without reprogramming. BasisN introduces a novel representation of the kernels in DNN layers as combinations of global basis vectors shared between all layers with quantized coefficients. These basis vectors are written to crossbars only once and used for the computations of all layers with marginal hardware modification. BasisN also provides a novel training approach to enhance computation parallelization with the global basis vectors and optimize the coefficients to construct the kernels. Experimental results demonstrate that cycles per inference and energy-delay product were reduced to below 1% compared with applying reprogramming on crossbars in processing large-scale DNNs such as DenseNet and ResNet on ImageNet and CIFAR100 datasets, while the training and hardware costs are negligible. Amro Eldebiky, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ing-Chao Lin, Ulf Schlichtmann, Bing Li 0005 |
ICCAD | 5 |
| 2024 | Aging-Aware Energy-Efficient Task Deployment of Heterogeneous Multicore SystemsabstractHeterogeneous multicore systems, which consist of high-performance and power-efficient cores, are emerging to satisfy the various demands on performance and power consumption. On the other hand, as CMOS technology continues to shrink in size, the aging effect, which can cause performance degradation or timing failures, has become a non-negligible threat to lifetime reliability. To overcome the challenges under the aging effect, various approaches have been proposed in previous studies. Most previous studies, however, did not consider the different characteristics of big and little cores. In addition, most of them do not consider critical tasks with the strict timing requirements present in real-time applications, resulting in early system failure. Therefore, considering different characteristics of cores and the presence of critical tasks, we propose an aging-aware task deployment framework for real-time systems. In this framework, for high-performance big cores, we propose a novel asymmetric aging-aware strategy. This strategy finds an energy-efficient task-to-core assignment to reserve some healthy cores at the early system life stage. The reserved cores are kept idle with the lowest voltage and can execute critical tasks at the late system life stage, extending the system lifetime. Meanwhile, the non-reserved cores use lower voltages to execute tasks, reducing the aging effect. For energy-efficient little cores, we adopt the symmetric aging-aware strategy to balance out the aging effect of each little core. With a balanced aging effect, the utilization of little cores is improved. In addition, we propose voltage/frequency boosting and task migration techniques to increase the number of cores that can meet the task timing constraints. Compared to the state of the art, the proposed framework can achieve 1.10x lifetime improvement and 5% energy reduction. Yu-Guang Chen, Chieh-Shih Wang, Ing-Chao Lin, Zheng-Wei Chen, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | AttentionRC: A Novel Approach to Improve Locality Sensitive Hashing Attention on Dual-Addressing MemoryabstractAttention is a crucial component of the Transformer architecture and a key factor in its success. However, it suffers from quadratic growth in time and space complexity as input sequence length increases. One popular approach to address this issue is the Reformer model, which uses locality-sensitive hashing (LSH) attention to reduce computational complexity. LSH attention hashes similar tokens in the input sequence to the same bucket and attends tokens only within the same bucket. Meanwhile, a new emerging nonvolatile memory (NVM) architecture, row column NVM (RC-NVM), has been proposed to support row- and column-oriented addressing (i.e., dual addressing). In this work, we present AttentionRC, which takes advantage of RC-NVM to further improve the efficiency of LSH attention. We first propose an LSH-friendly data mapping strategy that improves memory write and read cycles by 60.9% and 4.9%, respectively. Then, we propose a sort-free RC-aware bucket access and a swap strategy that utilizes dual-addressing to reduce 38% of the data access cycles in attention. Finally, by taking advantage of dual-addressing, we propose transpose-free attention to eliminate the transpose operations that were previously required by the attention, resulting in a 51% reduction in the matrix multiplication time. Chun-Lin Chu, Yun-Chih Chen, Wei Cheng 0006, Ing-Chao Lin, Yuan-Hao Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | A Hardware Friendly Variation-Tolerant Framework for RRAM-Based Neuromorphic ComputingabstractEmerging resistive random access memory (RRAM) attracts considerable interest in computing-in-memory by its high efficiency in multiply-accumulate operation, which is the key computation in the neural network (NN). However, due to the imperfect fabrication, RRAM cells suffer from the variations, which make the values in RRAM cells deviate from the target values so that the accuracy of the RRAM-based NN accelerator degrades significantly. Moreover, in a practical hardware design of RRAM-based NN accelerators, if the number of wordlines and bitlines in a crossbar array activated at the same time increases, ADCs with a high resolution are required and the power consumption of ADC increases. This paper proposes a novel methodology to mitigate the impact of variations in RRAM-based neural network accelerators. The methodology includes a unary-based non-uniform quantization method and a variation-aware operation unit (OU) based framework. The unary-based non-uniform quantization method equalizes the significance of weights stored in each RRAM cell to reduce the impact of variations. The variation-aware OU-based framework activates only RRAM cells in the same OU at the same time, which reduces the power consumption of ADCs. Additionally, the framework introduces three methods, including OU skipping, OU recombination, and OU compensation, to further mitigate the impact of variations. The experiments show that the proposed approach outperforms the state-of-the-art among four NN models on two datasets with 2-bit cell resolution. Fang-Yi Gu, Cheng-Han Yang, Ing-Chao Lin, Da-Wei Chang, Darsen D. Lu, Ulf Schlichtmann |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | A Novel and Efficient Block-Based Programming for ReRAM-Based Neuromorphic ComputingabstractReRAM-based accelerators have emerged as promising accelerators for deep neural networks (DNNs). How-ever, programming every ReRAM cell to its corresponding conductance before inference can be time-consuming and energy-intensive using existing one-by-one/row-by-row programming mechanisms. Although a two-phase multi-row programming scheme has been proposed to enhance programming efficiency, there are situations where multiple rows cannot be programmed together and only row-by-row programming can be employed. Therefore, this paper proposes a new block-based programming architecture for ReRAM crossbars that enables precise control of wordline and bitline transistors. In addition, a block-based programming framework, including the approximation phase and the fine-tuning phase, along with a multi-line programming algorithm and a programming-aware model retraining are proposed to reduce programming cycles and energy consumption. Experimental results demonstrate that our proposed method can reduce programming cycles and energy consumption by 46%-49 % and 63 % -64 %, respectively, compared to the state of the art. Additionally, the area and power overhead are negligible. Wei-Lun Chen, Fang-Yi Gu, Ing-Chao Lin, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
ICCAD | 3 |
| 2022 | WRAP: Weight RemApping and Processing in RRAM-based Neural Network Accelerators Considering Thermal EffectabstractResistive random-access memory (RRAM) has shown great potential for computing in memory (CIM) to support the requirements of high memory bandwidth and low power in neuromorphic computing systems. However, the accuracy of RRAM-based neural network (NN) accelerators can degrade significantly due to the intrinsic statistical variations of the resistance of RRAM cells, as well as the negative effects of high temperatures. In this paper, we propose a subarray-based thermal-aware weight remapping and processing framework (WRAP) to map the weights of a neural network model into RRAM subarrays. Instead of dealing with each weight individually, this framework maps weights into subarrays and performs subarray-based algorithms to reduce computational complexity while maintaining accuracy under thermal impact. Experimental results demonstrate that using our framework, inference accuracy losses of four DNN models are less than 2% compared to the ideal results and 1% with compensation applied even when the surrounding temperature is around 360K. Po-Yuan Chen, Fang-Yi Gu, Yu-Hong Huang, Ing-Chao Lin |
DATE | 4 |
| 2022 | GraphRC: Accelerating Graph Processing on Dual-Addressing Memory with Vertex MergingabstractArchitectural innovation in graph accelerators attracts research attention due to foreseeable inflation in data sizes and the irregular memory access pattern of graph algorithms. Conventional graph accelerators ignore the potential of Non-Volatile Memory (NVM) crossbar as a dual-addressing memory and treat it as a traditional single-addressing memory with higher density and better energy efficiency. In this work, we present GraphRC, a graph accelerator that leverages the power of dual-addressing memory by mapping in-edge/out-edge requests to column/row-oriented memory accesses. Although the capability of dual-addressing memory greatly improves the performance of graph processing, some memory accesses still suffer from low-utilization issues. Therefore, we propose a vertex merging (VM) method that improves cache block utilization rate by merging memory requests from consecutive vertices. VM reduces the execution time of all 6 graph algorithms on all 4 datasets by 24.24% on average. We then identify the data dependency inherent in a graph limits the usage of VM, and its effectiveness is bounded by the percentage of mergeable vertices. To overcome this limitation, we propose an aggressive vertex merging (AVM) method that outperforms VM by ignoring the data dependency inherent in a graph. AVM significantly reduces the execution time of ranking-based algorithms on all 4 datasets while preserving the correct ranking of the top 20 vertices. Wei Cheng 0006, Chun-Feng Wu, Yuan-Hao Chang 0001, Ing-Chao Lin |
ICCAD | 4 |
| 2022 | An Efficient Implementation of Convolutional Neural Network With CLIP-Q Quantization on FPGAabstractConvolutional neural networks (CNNs) have achieved tremendous success in the computer vision domain recently. The pursue for better model accuracy drives the model size and the storage requirements of CNNs as well as the computational complexity. Therefore, Compression Learning by InParallel Pruning-Quantization (CLIP-Q) was proposed to reduce a vast amount of weight storage requirements by using a few quantized segments to represent all weights in a CNN layer. Among various quantization strategies, CLIP-Q is suitable for hardware accelerators because it reduces model size significantly while maintaining the full-precision model accuracy. However, the current CLIP-Q approach did not consider the hardware characteristics and it is not straightforward when mapped to a CNN hardware accelerator. In this work, we propose a software-hardware codesign platform that includes a modified version of CLIP-Q algorithm and a hardware accelerator, which consists of$5\times 5$reconfigurable convolutional arrays with input and output channel parallelization. Additionally, the proposed CNN accelerator maintains the same accuracy of a full-precision CNN in Cifar-10 and Cifar-100 datasets. Wei Cheng 0006, Ing-Chao Lin, Yun-Yang Shih |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | An efficient NBTI-aware wake-up strategy: Concept, design, and manipulation
Yu-Guang Chen, Ing-Chao Lin, Kun-Wei Chiu, Cheng-Hsuan Liu |
Integr. | 2 |
| 2021 | Rescuing RRAM-Based Computing From Static and Dynamic FaultsabstractEmerging resistive random access memory (RRAM) has shown the great potential of in-memory processing capability, and thus attracts considerable research interests in accelerating memory-intensive applications, such as neural networks (NNs). However, the accuracy of RRAM-based NN computing can degrade significantly, due to the intrinsic statistical variations of the resistance of RRAM cells. In this article, we propose SIGHT, a synergistic algorithm-architecture fault-tolerant framework, to holistically address this issue. Specifically, we consider three major types of faults for RRAM computing: 1) nonlinear resistance distribution; 2) static variation; and 3) dynamic variation. From the algorithm level, we propose a resistance-aware quantization to compel the NN parameters to follow the exact nonlinear resistance distribution as RRAM, and introduce an input regulation technique to compensate for RRAM variations. We also propose a selective weight refreshing scheme to address the dynamic variation issue that occurs at runtime. From the architecture level, we propose ageneralandlow-costarchitecture accordingly for supporting our fault-tolerant scheme. Our evaluation demonstrates almost no accuracy loss for our three fault-tolerant algorithms, and the proposed SIGHT architecture incurs performance overhead as little as 7.14%. Jilan Lin, Cheng-Da Wen, Xing Hu 0001, Tianqi Tang 0001, Ing-Chao Lin, Yu Wang 0002, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Overview of 2020 CAD Contest at ICCADabstractThe "CAD Contest at ICCAD" is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2020 CAD Contest hits a record high of 186 teams from all over the world, which represents more than 50% growth compared to last year. The contest keeps enhancing impact and boosting EDA research. Ing-Chao Lin, Ulf Schlichtmann, Tsung-Wei Huang, Mark Po-Hung Lin |
ICCAD | 1 |
| 2020 | Global Clean Page First Replacement and Index-Aware Multistream Prefetcher in Hybrid Memory ArchitectureabstractAs cloud computing and big data applications become more popular, the demand for large capacity memory and data preservation in memory increases. Therefore, nonvolatile memory (NVM) with high capacity is being actively developed. A hybrid memory that comprises both NVM and DRAM and provides both high access speed and nonvolatility has become a major trend. However, compared to DRAM, NVM in the hybrid memory typically suffers from a shorter lifetime and higher latency. To improve the lifetime and address the latency issues associated with hybrid memory, we propose a global clean page first replacement (GCPF) to reduce the write operations to NVM. We also propose an index-aware multistream prefetcher (IAMSP) that considers the indexes of prefetch candidates individually so as to prefetch pages from NVM more accurately. Benchmarks with a large memory footprint are used to evaluate the proposed schemes. The experimental results show that GCPF enhances lifetime by 56.8% as compared to LRU, on average. When applying prefetching schemes on GCPF, the lifetime is insignificantly degraded. In addition, IAMSP reduces DRAM misses by 42.0% as compared to LRU, while a modern prefetcher that can change the prefetch degree dynamically only reduces DRAM misses by 38.0%, on average. When applying both GCPF and IAMSP, the average access latency can be reduced by 28.8% as compared to LRU. Ing-Chao Lin, Da-Wei Chang, Wei-Jun Chen, Jian-Ting Ke |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | ROAD: Improving Reliability of Multi-core System via Asymmetric AgingabstractNegative-Bias Temperature Instability (NBTI), which may lead to performance degradation or even timing failure, has become one of the most drastic challenges in modern multi-core systems. To tolerate NBTI and extend the lifetime of the system, previous researchers proposed maintaining all cores in the multi-core system under similar aging conditions (symmetric aging) through various task assignment algorithms and/or dynamic voltage frequency scaling. Although the concept of symmetric aging provides efficient approaches to tolerating NBTI, it may reduce the lifetime of a multi-core system. If a critical task (i.e., a task with tight timing constraints) arrives when the system has already operated for years, it is possible that none of the equivalently aged cores will be able to complete the critical task within its timing constraints. This unavoidable timing failure then will shorten the lifetime of the system. In contrast, if a few cores are kept robust, these cores can be used to execute the critical task even if all the other cores are aged (asymmetric aging), which avoids timing failure and extends the system lifetime. Based on the above observation, this paper proposes a novel reliability improvement framework that consists of task graph Retiming, task Ordering, task Assignment under asymmetric aging, and Dynamic voltage selection (ROAD) for multi-core systems. With our framework, asymmetric aging can extend the system lifetime through successfully executing critical tasks at the later life stages of the system. The experimental results show that our approach can significantly increase the system lifetime with no or insignificant energy overhead. Yu-Guang Chen, Ing-Chao Lin, Jian-Ting Ke |
ICCAD | 2 |
| 2019 | Overview of 2019 CAD Contest at ICCADabstractThe “CAD Contest at ICCAD” is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has published many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The contest keeps enhancing impact and boosting EDA research. Ulf Schlichtmann, Sabya Das, Ing-Chao Lin, Mark Po-Hung Lin |
ICCAD | 3 |
| 2019 | TAP: Reducing the Energy of Asymmetric Hybrid Last-Level Cache via Thrashing Aware Placement and MigrationabstractEmerging non-volatile memories (NVMs) have favorable properties, such as low leakage and high density, and have attracted a lot of attention in recent years. Among them, spin-transfer torque magnetoresistive random access memory (STT-MRAM) with SRAM-comparable read speed is a good candidate to build large last-level caches (LLCs). However, STT-MRAM suffers from long write latency and high write energy. To mitigate the impact of asymmetric read/write energy and latency, hybrid cache designs have been proposed to combine the merits of STT-MRAM and SRAM. In such a hybrid SRAM/STT-MRAM LLC, intelligent block placement and migration policies are needed to improve the energy efficiency. Prior studies map write-intensive blocks to SRAM and keep read-intensive blocks in STT-MRAM for reducing the energy consumption of hybrid LLCs. The write-intensive/read-intensive blocks are usually captured by sampling the address (PC) of memory access instructions or adding simple access counters in each cache line. Nevertheless, these prior approaches cannot fully capture the energy-harmful access behavior in STT-MRAM, especially the writes caused by repetitive data transfer between the LLC and upper-level caches. In this paper, we find that conflict misses in L2 often generate thrashing blocks which move back and forth between L2 and LLC. If dirty thrashing blocks that incur extensive writes are placed in STT-MRAM, energy consumption would excessively increase, especially when running memory-bound workloads. Thus, we propose a thrashing aware placement and migration policy (TAP) to tackle the challenge. TAP places dirty thrashing blocks into SRAM and migrates clean thrashing blocks from SRAM to STT-MRAM. Evaluation results show that TAP can provide significant energy savings with minimal performance loss. Jing-Yuan Luo, Hsiang-Yun Cheng, Ing-Chao Lin, Da-Wei Chang |
IEEE Trans. Computers | 3 |
| 2019 | OCMAS: Online Page Clustering for Multibank Scratchpad MemoryabstractScratchpad memory (SPM), a software-controlled on-chip memory, is being increasingly used in embedded systems to reduce on-chip memory energy consumption. To further reduce energy consumption, multibank SPM architecture is proposed. In multibank SPM, each bank can be accessed independently, and unused banks can enter the low power mode, thus reducing leakage energy. However, if both frequently and infrequently used data exist in the same bank, the bank will not be able to enter the low power mode, resulting in less energy reduction. To address this issue, we propose online page clustering for multibank SPM (OCMAS) to reduce the leakage energy in multibank SPM. OCMAS groups SPM pages with similar access frequencies into the same bank, allowing banks containing infrequently used data to stay in low power mode longer. We also propose a method to dynamically adjust the thresholds for determining cold pages (pages that contain infrequently used data), so banks that contain cold pages can enter the low power mode with a shorter idle timeout. Compared to conventional timeout-based, periodic drowsy, and bank-based methods, OCMAS can reduce the energy delay product by up to 37.67% (18.14% on average), 39.53% (22.38% on average), and 132.32% (23.34% on average) in 32 KB 4-bank SPM, by up to 25.33% (15.71% on average), 29.87% (15.92% on average), and 72.25% (22.64% on average) in 32 KB 8-bank SPM, by up to 28.99% (13.45% on average), 30.71% (14.74% on average), and 96.67% (19.94% on average) in 16 KB 4-bank SPM, and by up to 30.18% (10.05% on average), 32.13% (11.62% on average), and 65.56% (16.2% on average) in 16 KB 8-bank SPM. The area overhead is approximately 0.72%, which is insignificant. Da-Wei Chang, Ing-Chao Lin, Yi-Chiao Lin, Wen-Zhi Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Infection-Based Dead Page Prediction in Hybrid Memory ArchitectureabstractWith the widespread use of cloud computing and the Internet, applications that require a large memory footprint, such as in-memory databases, have gained in popularity. These applications depend on a high capacity, reliable memory architecture. To achieve these two goals, hybrid memory that uses both DRAM and nonvolatile memory (NVM) provides benefits that include large capacity and nonvolatility. However, NVM is usually accompanied by high write latency and endurance problems. It is important to reduce NVM writes and improve the latency and lifetime of hybrid memory. One way to reduce NVM writes is to reduce dead pages that occupy DRAM and have not been accessed for a long time. When dead pages are removed from the memory, more space can be reserved for frequently accessed data, reducing DRAM misses and NVM writes. Currently, there are several dead block prediction techniques that can identify dead blocks and reduce miss rates at the cache level. However, they are not effective at the memory level because CPU memory accesses exhibit less locality when accesses are filtered by caches. To propose an application that is suitable at the memory level and to achieve a reduction in NVM writes, this paper proposes a simple but effective dead page predictor, called the infection-based dead page predictor (IDP), for the memory level. IDP uses the access counts of evicted pages to determine if nearby pages are also dead pages (i.e., other pages are infected by evicted pages). The simulation results show that compared to related work, the proposed predictor significantly reduces DRAM misses and enhances lifetime. Ing-Chao Lin, Da-Wei Chang, Chen-Tai Kao, Sheng-Xuan Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | An efficient NBTI-aware wake-up strategy for power-gated designsabstractThe wake-up process of a power-gated design may induce an excessive surge current and threaten the signal integrity. A proper wake-up sequence should be carefully designed to avoid surge current violations. On the other hand, PMOS sleep transistors may suffer from the negative-bias temperature instability (NBTI) effect which results in decreased driving current. Conventional wake-up sequence decision approaches do not consider the NBTI effect, which may result in a longer or unacceptable wake-up time after circuit aging. Therefore, in this paper, we propose a novel NBTI-aware wake-up strategy to reduce the average wake-up time within a circuit lifetime. Our strategy first finds a set of proper wake-up sequences for different aging scenarios (i.e. after a certain period of aging), and then dynamically reconfigures the wake-up sequences at runtime. The experimental results show that compared to a traditional fixed wake-up sequence approach, our strategy can reduce average wake-up time by as much as 45.04% with only 3.7% extra area overhead for the reconfiguration structure. Kun-Wei Chiu, Yu-Guang Chen, Ing-Chao Lin |
DATE | 3 |
| 2018 | Mitigating BTI-Induced Degradation in STT-MRAM Sensing SchemesabstractSpin-transfer torque magnetic RAM (STT-MRAM), which uses a magnetic tunnel junction to store binary data, is a promising memory technology. With many benefits, such as low leakage power, high density, high endurance, and nonvolatility, it has been explored as an SRAM replacement for cache design or a DRAM replacement for main memory. Meanwhile, along with the continuous shrinking of CMOS process technology, the bias temperature instability (BTI) effect has become a major reliability issue. Prior work has investigated the influence of the BTI effect on the SRAM sense amplifier, but no investigation has been done for the STT-MRAM sense amplifier. Therefore, this paper investigates the BTI effect on STT-MRAM sense amplifiers. We propose a majority-based technique and an alternative sensing technique to reduce circuit degradation. To further improve sensing delay, we propose using forward body bias (FBB) on an access transistor with a positive voltage. Extensive simulation results are done to show the effectiveness of the proposed techniques. The sensing delay for reading zeros and ones can be reduced by 10.61% and 4.35%, respectively, on average, with the majority-based technique. The sensing delay for reading zeros and ones can be reduced by 4.42% and 1.83%, respectively, on average, using the alternative sensing technique. The sensing delay for reading zeros and ones can be reduced by 15.37% and 6.25%, respectively, on average, by using both techniques simultaneously. When using the majority-based and alternative sensing techniques with the FBB technique, the sensing delay for reading zeros and ones can be improved by 29.93% and 57.67%, respectively, on average. We also analyze the BTI-induced degradation of a high-performance sense amplifier and a low power sense amplifier with the proposed techniques. The simulation results show that our proposed technique and simulation flow can be easily extended to other sense amplifiers. Ing-Chao Lin, Yun Kae Law, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | ROHOM: Requirement-Aware Online Hybrid On-Chip Memory Management for Multicore SystemsabstractMany studies have shown that energy consumption of on-chip memory is a critical issue for multicore embedded systems. In order to reduce energy consumption, scratchpad memory (SPM), a software controlled on-chip memory, has been increasingly used. In a typical multicore embedded system that uses SPM in the on-chip memory, each core has local SPM and can access data in both local SPM and SPMs of other cores (i.e., remote SPM). Since the latency and energy of accessing remote SPMs is higher than accessing local SPM, how data are allocated in local and remote SPMs influences the performance and energy consumption of the system. This paper proposes a requirement-aware online hybrid on-chip memory management (ROHOM) method. This method contains an SPM supervisor that determines the SPM allocation of each task according to the dynamic access behavior and SPM requirements of the task. In addition, two policies: 1) free remote SPM space first and 2) get local SPM space first, are proposed in ROHOM to reduce the access frequency of remote SPMs. The experimental results show ROHOM can reduce energy delay product up to 69% (42% on average) in an 8-core system and up to 69% (50% on average) in a 16-core system when compared to a contention aware SPM allocation method. The hardware area overhead is insignificant (about 1.8%). Da-Wei Chang, Ing-Chao Lin, Lin-Chun Yong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | High-Endurance Hybrid Cache Design in CMP Architecture With Cache Partitioning and Access-Aware PoliciesabstractIn recent years, nonvolatile memory (NVM) technologies, such as spin-transfer torque random-access memory (RAM) (STT-RAM) and phase change RAM, have drawn a lot of attention due to their low leakage and high density. However, both of these NVMs suffer from high write latency and limited endurance problems. To mitigate the write pressure on NVM, many static RAM (SRAM)/NVM hybrid cache designs have been proposed with write management policies. Unfortunately, existing hybrid cache designs do not consider the unbalanced workload of each core in (chip multiprocessor) architecture, resulting in unbalanced wear out of hybrid caches. This paper considers the unbalanced write distribution of a hybrid cache for CMP architecture as well as a novel hybrid cache design that includes SRAM cache, STT-RAM cache, and STT-RAM/SRAM hybrid cache banks. Based on the proposed hybrid cache design, two access-aware policies are proposed to mitigate unbalanced wearout of the STT-RAM region, and a wearout-aware dynamic cache partitioning scheme is proposed to dynamically partition the hybrid cache, improving the unbalanced write pressure among different cache partitions. The experimental results show that our proposed scheme and policies can achieve an average of 89 times improvement in cache lifetime and are able to reduce energy consumption by 58% compared with a SRAM cache. Ing-Chao Lin, Jeng-Nian Chiou |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Aging-Aware Reliable Multiplier Design With Adaptive Hold LogicabstractDigital multipliers are among the most critical arithmetic functional units. The overall performance of these systems depends on the throughput of the multiplier. Meanwhile, the negative bias temperature instability effect occurs when a pMOS transistor is under negative bias (Vgs= -Vdd), increasing the threshold voltage of the pMOS transistor, and reducing multiplier speed. A similar phenomenon, positive bias temperature instability, occurs when an nMOS transistor is under positive bias. Both effects degrade transistor speed, and in the long term, the system may fail due to timing violations. Therefore, it is important to design reliable high-performance multipliers. In this paper, we propose an aging-aware multiplier design with a novel adaptive hold logic (AHL) circuit. The multiplier is able to provide higher throughput through the variable latency and can adjust the AHL circuit to mitigate performance degradation that is due to the aging effect. Moreover, the proposed architecture can be applied to a columnor row-bypassing multiplier. The experimental results show that our proposed architecture with 16 × 16 and 32 × 32 column-bypassing multipliers can attain up to 62.88% and 76.28% performance improvement, respectively, compared with 16×16 and 32×32 fixed-latency column-bypassing multipliers. Furthermore, our proposed architecture with 16 × 16 and 32 × 32 row-bypassing multipliers can achieve up to 80.17% and 69.40% performance improvement as compared with 16×16 and 32 × 32 fixed-latency row-bypassing multipliers. Ing-Chao Lin, Yu-Hung Cho, Yi-Ming Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | High-Performance Low-Power Carry Speculative Addition With Variable LatencyabstractAdders are one of the most critical arithmetic circuits in a system and their throughput affects the overall performance of the system. Traditional n-bit adders provide accurate results, but the lower bound of their critical path delay is Ω(log n). To achieve a critical path delay lower than Ω(log n), many approximate adders have been proposed. These approximate adders decrease the critical path delay and improve the speed by sacrificing computation accuracy or predicting the computation results. This paper proposes a high-performance low-power carry speculative adder (CSPA). This adder separates the carry generator and sum generator. Only one sum generator is used in a block adder to reduce the critical path delay and area overhead. In addition, to generate 100% accurate results, error detection and recovery circuits are added to the proposed CSPA to construct a variable-latency carry speculative adder (VLCSPA). Instead of recalculating all results, the error detection and recovery circuits find and correct the block adder that generates incorrect partial sum bits, reducing power consumption. The experimental results show that the proposed CSPA achieves a 26.59% delay reduction, a 14.06% area reduction, and a 19.03% power consumption reduction compared to the corresponding values for an existing speculative carry-select adder. The experimental results also show the proposed CSPA can be used to improve image denoising results as well. Ing-Chao Lin, Yi-Ming Yang, Cheng-Chian Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | NBTI tolerance and leakage reduction using gate sizingabstractLeakage power is a major design constraint in deep submicron technology and below. Meanwhile, transistor degradation due to Negative Bias Temperature Instability (NBTI) has emerged as one of the main reliability concerns in nanoscale technology. Gate sizing is a widely used technique to reduce circuit leakage, and this approach has recently attracted much attention with regard to improving circuits to tolerate NBTI. However, these studies only consider timing and area constraints, and many other important issues, such as slew and max-load, are missing. In this article, we present an efficient gate sizing framework that can reduce leakage and improve circuit reliability under timing constraints. Our algorithms consider slack, slew and max-load constraints. The benchmarks are those from ISPD 2012, which feature industrial design properties, including discrete cell sizes, nonconvex cell timing models, slew dependencies and constraints, as well as large design sizes. The experimental results obtained from ISPD 2012 benchmark circuits demonstrate that our approach can meet all the constraints and tolerated NBTI degradation with a power savings of 6.54% as compared with the traditional method. Ing-Chao Lin, Shun-Ming Syu, Tsung-Yi Ho |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2014 | CASA: Contention-Aware Scratchpad Memory Allocation for Online Hybrid On-Chip Memory ManagementabstractScratchpad memory (SPM) has been increasingly used in embedded systems due to its higher efficiency in terms of energy and area compared to that of ordinary cache. A hybrid on-chip memory architecture that combines SPM with a mini-cache has been proposed. One key issue for hybrid on-chip memory architectures is to reduce the number of off-chip memory accesses and energy consumption. Existing methods achieve this by moving the most frequently accessed data into SPM. However, these methods may be ineffective because the main source of off-chip memory accesses may not be the most frequently accessed data. Instead, most off-chip memory accesses are caused by cache misses, so reducing the latter will reduce the former. Cache misses are mainly caused by data contending for cache lines. Therefore, this paper proposes a contention-aware SPM allocation method for hybrid on-chip management. The number of cache misses for a page is used as a metric to determine whether a page should be moved to SPM. When the number of misses for a page exceeds a threshold, the page is moved to SPM, reducing cache contention. Experimental results show that the proposed method can reduce the energy delay product by 35% to 53% compared to a cache-only on-chip memory architecture and 19% to 31% compared to an existing hybrid on-chip memory architecture. Da-Wei Chang, Ing-Chao Lin, Yu-Shiang Chien, Ching-Lun Lin, Alvin Wen-Yu Su, Chung-Ping Young |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Capture-Power-Safe Test Pattern Determination for At-Speed Scan-Based TestingabstractDuring an at-speed scan-based test, excessive capture power may cause significant current demand, resulting in the IR-drop problem and unnecessary yield loss. Many methods address this problem by reducing the switching activities of power-risky patterns. These methods may not be efficient when the number of power-risky patterns is large or when some of the patterns require extremely high power. In this paper, we propose discarding all power-risky patterns and starting with power-safe patterns only. Our test generation procedure includes two processes, namely, test pattern refinement and low-power test pattern regeneration. The first process is used to refine the power-safe patterns to detect faults originally detected only by power-risky patterns. If some faults are still undetected after this process, the second process is applied to generate new power-safe patterns to detect these faults. The patterns obtained using the proposed procedure are guaranteed to be power-safe for the given power constraints. To the best of our knowledge, this is the first method that refines only the power-safe patterns to address the capture power problem. Experimental results on ISCAS'89 and ITC'99 benchmark circuits show that an average of 75% of faults originally detected only by power-risky patterns can be detected by refining power-safe patterns and that most of the remaining faults can be detected by the low-power test generation process. Furthermore, the required test data volume can be reduced by 12.76% on average with little or no fault coverage loss. Yi-Hua Li, Wei-Cheng Lien, Ing-Chao Lin, Kuen-Jong Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | BTI-Aware Sleep Transistor Sizing Algorithm for Reliable Power Gating DesignsabstractPower gating is an effective way to reduce leakage power. This technique uses high Vthtransistors, called sleep transistors, to turn off the power supply. However, sleep transistors suffer from the bias temperature instability (BTI) effect, resulting in an increased Vth, and reduced reliability. This paper proposes two BTI-aware sleep transistor sizing algorithms to reduce the total width of sleep transistors based on the distributed sleep transistor network structure. The proposed algorithms reduce total width by more than 16.08%. More area can be reduced if the BTI effect on both sleep and cluster transistors is considered. Kai-Chiang Wu, Ing-Chao Lin, Yao-Te Wang, Shuen-Shiang Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | NBTI and Leakage Reduction Using ILP-Based ApproachabstractWe propose an integer linear programming-based formulation to improve the effectiveness of the transmission gate-based technique intended to reduce negative-bias temperature instability and leakage power consumption. We also propose a virtual input pin technique to improve leakage reduction and use path sensitization to reduce area overhead. Simulation results show that combining these techniques can achieve >51.18% delay improvement and 63.34% leakage power improvement with only 2.31% area overhead. Ing-Chao Lin, Kuan-Hui Li, Chia-Hao Lin, Kai-Chiang Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | High-endurance hybrid cache design in CMP architecture with cache partitioning and access-aware policyabstractIn recent years, NVM (non-volatile memory) technologies, such as STT-RAM (spin transfer torque RAM) and PRAM (phase change RAM), have drawn a lot of attention due to their low leakage and high density. However, both NVMs suffer from high write latency and limited endurance problems. To overcome these problems, the SRAM/NVM hybrid cache architecture has been proposed, and the write pressure on NVM can be mitigated with appropriate write management policy. Moreover, many wear leveling techniques have been proposed to extend the lifetime of NVM in the hybrid cache. In this paper, we proposed a hybrid cache design that includes SRAM cache, STT-RAM cache, and STT-RAM/SRAM hybrid cache banks for CMP (chip multi-processors) architecture. We also propose a partition-level wear leveling scheme and access-aware policies to mitigate unbalanced wear-out of STT-RAM lines within a partition and among different cache partitions. Experimental results show that, our proposed scheme and policies can achieve an average of 89 times improvement in cache lifetime and are able to save 58% power consumption compared to SRAM cache. Shun-Ming Syu, Yu-Hui Shao, Ing-Chao Lin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | High accuracy approximate multiplier with error correctionabstractApproximate computing has gained significant attention due to the popularity of multimedia applications. In this paper, we propose a novel inaccurate 4:2 counter that can effectively reduce the partial product stages of the Wallace Multiplier. Compared to the normal Wallace multiplier, our proposed multiplier can reduce 10.74% of power consumption and 9.8% of delay on average, with an error rate from 0.2% to 13.76% The accuracy of amplitude is higher than 99% In addition, we further enhance the design with error-correction units to provide accurate results. The experimental results show that the extra power consumption of correct units is lower than 6% on average. Compared to the normal Wallace multiplier, the average latency of our proposed multiplier with EDC is 6% faster when the bit-width is 32, and the power consumption is still 10% lower than that of the Wallace multiplier. Chia-Hao Lin, Ing-Chao Lin |
ICCD | 2 |
| 2013 | Leakage and Aging Optimization Using Transmission Gate-Based TechniqueabstractNegative bias temperature instability (NBTI), which can degrade the switching speed of PMOS transistors, has become a major reliability challenge. Reducing leakage consumption is one of the major design goals. The gate replacement (GR) technique is an effective way to reduce both the NBTI effect and leakage. This technique, however, has less flexibility because the replaced gate can only produce one output value and careful algorithms are needed to decide the output value of the replaced gate. In this paper, we propose a novel transmission gate-based technique to minimize NBTI-induced degradation and leakage. This technique, which can offer logic 1 for NBTI mitigation and logic 0 for leakage reduction, provides higher flexibility, as compared to the GR technique. Simulation results show that our proposed technique has up to 20× and 2.16×, on average, improvement on NBTI-induced degradation with comparable leakage power reduction. With a 19.19% area penalty, combining our technique and the GR can reduce 17.92% of the total leakage power and 32.36% of NBTI-induced circuit degradation. Ing-Chao Lin, Chin-Hung Lin, Kuan-Hui Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Analyzing throughput of power and thermal-constraint multicore processor under NBTI effectabstractNBTI (Negative Bias Temperature Instability) which can degrade the switching speed of PMOS transistors has become a major reliability challenge. In this paper, we investigate the throughput impact of NBTI on power and thermal-constraint multicore processors and show up to 30 % degradation when both process variation and NBIT are considered. Then we evaluate the effectiveness of core rotation, adaptive voltage scaling and adaptive body biasing to improve the throughput of power and thermal constrained multicore processors. Our experimental results demonstrate 11.1 % improvement in VDD is sufficient to guarantee throughput after 10-yr NBTI influence when processor variation is not considered. In contract, ABB technique is not able to recover throughput loss caused by NBTI. Shi-Qun Zheng, Ing-Chao Lin, Yen-Han Lee |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | TG-based technique for NBTI degradation and leakage optimization
Chin-Hung Lin, Ing-Chao Lin, Kuan-Hui Li |
ISLPED | 2 |