Soontae Kim

dblp:k/SoontaeKim · DBLP profile ↗
← Back
74ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0001-5106-8409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 66 · 7 first-author · 11 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorComputer networks · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 C2C: A Framework for Critical Token Classification in Transformer-Based Inference Systems
abstract
Because embedding vectors in a Transformer-based model represent crucial information about input texts, attacks or errors affecting them can cause severe accuracy degradation. We observe critical tokens for the first time, that determine the overall accuracy but their embedding vectors take only a small portion of the embedding table. Therefore, we propose a framework called C2C that classifies the critical tokens to facilitate their protection in a Transformer-based inference system with a small overhead. Using BERT with GLUE datasets, critical embedding vectors take only 13.8% of the embedding table. Compromising critical embedding vectors can reduce accuracy by up to 44.8% even if other parameters are not corrupted.
Myeongjae Jang, Jesung Kim, Haejin Nam, Sihyun Kim, Soontae Kim
DATE5
2024 Zero and Narrow-Width Value-Aware Compression for Quantized Convolutional Neural Networks
abstract
Convolutional neural networks are normally used in systems with dedicated neural processing units for CNN-related computations. For high performance and low hardware overheads, CNN datatype quantization is applied. As an additional optimization, to further reduce DRAM accesses, compression algorithms have been used for CNN data. However, conventional zero value-aware compression algorithms suffer from a reduction in compression ratio with the latest quantized CNNs, owing to the small number of zero values. Moreover, the appropriate zero run-length code width can be changed dynamically based on the CNNs, layers, and quantization datatypes. As another compressible data value for increasing the compression ratio, the latest quantized CNNs have many narrow-width values. Because low-precision quantization reduces the data bit width, CNN data are gathered into a few discrete values and incur a biased data distribution. These discrete values become narrow-width values, and constitute a large proportion of the biased distribution. In this article, we propose an efficient compression algorithm for quantized CNNs, ENCORE, which utilizes variable zero run-length encoding and compresses narrow-width values. With the latest quantized CNNs, ENCORE shows higher compression ratios, 93.55% and 50.85% in Mobilenet v1 and Tiny YOLO v3, respectively, than conventional zero value-aware CNN data compression algorithms.
Myeongjae Jang, Jinkwon Kim, Haejin Nam, Soontae Kim
IEEE Trans. Computers4
2024 Highly VM-Scalable SSD in Cloud Storage Systems
abstract
Solid-state drives (SSDs) are widely used in cloud storage. As the capacity of an SSD has been increasing, it has become common for many virtual machines (VMs) to share a single SSD to maximize resource utilization. However, this sharing can degrade the efficiency of internal operations, such as garbage collection, resulting in increased latencies. Existing literature in this field has mostly focused on interdevice isolation considering the storage device as a black-box entity or presumed an SSD to be shared by up to only eight VMs. In this study, we first analyze a realistic SSD usage environment in cloud systems and identify that block-level data isolation (BDI) should be guaranteed to efficiently scale up the number of VMs in an SSD with minimum latency increases. However, previous schemes cannot work efficiently with BDI when the SSD is shared by dozens of VMs. Based on this analysis, we propose an SSD internal resource management scheme in a cloud environment, called highly VM-scalable SSD (VMS). VMS dynamically partitions physical resources and allocates them to VMs, while the VMs share global buffer blocks to lower latency during abrupt fluctuations of write I/O intensities. Our experimental results show up to 29% of latency reduction. VMS exhibits reduced latencies even in the experiment with 64 VMs, where existing schemes do not function normally.
Wonyoung Lee 0001, Mincheol Kang, Soontae Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 HARP: Hardware-Based Pseudo-Tiling for Sparse Matrix Multiplication Accelerator
abstract
General sparse matrix-matrix multiplication (SpGEMM) is a memory-bound workload, due to the compression format used. To minimize data movements for input matrices, outer product accelerators have been proposed. Since these accelerators access input matrices only once and then generate numerous partial products, managing the generated partial products is the key optimization factor. To reduce the number of partial products handled, the state-of-the-art accelerator uses software to tile an input matrix. However, the software-based tiling has three limitations. First, a user manually executes the tiling software and manages the tiles. Second, generating a compression format for each tile incurs memory-intensive operations. Third, an accelerator that uses the compression format cannot skip ineffectual accesses for input matrices.
Jinkwon Kim, Myeongjae Jang, Haejin Nam, Soontae Kim
MICRO4
2023 PR-SSD: Maximizing Partial Read Potential by Exploiting Compression and Channel-Level Parallelism
abstract
Recent NAND flash memories provide a partial read operation that can read a page partially and has lower latency than a normal read operation. In order to maximize the benefit of the partial read operation, compression techniques can be applied to improve performance by generating additional partial page requests by compressing pages into smaller ones. Unfortunately, existing compression support SSDs suffer from a huge decompression latency that eventually cancels the benefit of the partial read operation. In this paper, we propose Partial Read-aware SSD (PR-SSD) for fully exploiting partial read operations. In order to mitigate the decompression latency, we propose a new compression algorithm, called Dominant Pattern Compression (DPC), which has extremely low decompression latency. Because uncompressed page requests cannot exploit the partial read operation, we propose split Flash Translation Layer (FTL) that can split the requests into smaller ones and allocate them to different channels for exploiting channel-level parallelism in SSD. Experimental results reveal that PR-SSD can reduce the read response time by 18% on average and also the number of writes and write response time by 29% and 24% on average, respectively
Mincheol Kang, Wonyoung Lee 0001, Jinkwon Kim, Soontae Kim
IEEE Trans. Computers4
2022 ENCORE Compression: Exploiting Narrow-width Values for Quantized Deep Neural Networks
abstract
Deep Neural Networks (DNNs) become a practical machine learning algorithm running on various Neural Processing Units (NPUs). For higher performance and lower hardware overheads, DNN datatype reduction through quantization is proposed. Moreover, to solve the memory bottleneck caused by large data size in DNNs, several zero value-aware compression algorithms are used. However, these compression algorithms do not compress modern quantized DNNs well because of decreased zero values. We find that the latest quantized DNNs have data redundancy due to frequent narrow-width values. Because low-precision quantization reduces DNN datatypes to a simple datatype with less bits, scattered DNN data are gathered to a small number of discrete values and incur a biased data distribution. Narrow-width values occupy a large proportion of the biased distribution. Moreover, an appropriate zero run-length bits can be dynamically changed according to DNN sparsity. Based on this observation, we propose a compression algorithm that exploits narrow-width values and variable zero run-length for quantized DNNs. In experiments with three quantized DNNs, our proposed scheme yields an average compression ratio of 2.99.
Myeongjae Jang, Jinkwon Kim, Jesung Kim, Soontae Kim
DATE4
2022 Salvaging Runtime Bad Blocks by Skipping Bad Pages for Improving SSD Performance
abstract
Recent research has revealed that runtime bad blocks are found in the early lifespan of solid state drives. The reduction in overprovisioning space due to runtime bad blocks may well have a negative impact on performance as it weakens the chances of selecting a better victim block during garbage collection. Moreover, previous studies focused on reusing worn-out bad blocks exceeding a program/erase cycle threshold, leaving the problem of runtime bad blocks unaddressed. Based on this observation, we present a salvation scheme for runtime bad blocks. This paper reveals that these blocks can be identified when a page write fails at runtime. Furthermore, we introduce a method to salvage functioning pages from runtime bad blocks. Consequently, the loss in the overprovisioning space can be minimized even after the occurrence of runtime bad blocks. Experimental results show a 26.3% reduction in latency and a 25.6% increase in throughput compared to the baseline at a conservative bad block ratio of 0.45%. Additionally, our results confirm that almost no overhead was observed.
Junoh Moon, Mincheol Kang, Wonyoung Lee 0001, Soontae Kim
DATE4
2022 Exploiting Inter-block Entropy to Enhance the Compressibility of Blocks with Diverse Data
abstract
As higher memory bandwidth is required for data-intensive environments, memory compression can be a simple but effective solution to increase memory bandwidth. However, previous intra-block compression techniques do not provide sufficient bandwidth improvement owing to the incompressibility of blocks with diverse data while previous inter-block compression techniques suffer from huge additional memory access overheads or low compression coverages. To overcome the limitations of the previous intra-and inter-block compression techniques, we leverage both the naturally observed low-entropy among blocks and the artificially generated low-entropy resulting from our optimization techniques. Based on these two low-entropies, we propose an Entropy-based Pattern Compression (EPC), which generates an inter-block pattern from the same low-entropy region in numerous blocks and then compresses these blocks by using the selected pattern. Our evaluations show that EPC achieves up to 13% (3% on average) higher speedup and 13% (4% on average) DRAM energy consumption reduction with 160x (20x on average) fewer patterns(groups) compared to the state-of-the-art inter-block compression technique.
Jinkwon Kim, Mincheol Kang, Jeongkyu Hong, Soontae Kim
HPCA4
2021 ECC-United Cache: Maximizing Efficiency of Error Detection/Correction Codes in Associative Cache Memories
abstract
Error Detection/Correction Codes (EDCs/ECCs) are the most conventional approaches to protect on-chip caches against radiation-induced soft errors. The overhead of EDCs/ECCs is a major concern and is of decisive importance when a higher protection capability is required to tolerate multiple adjacent bit errors (burst errors). This article proposes the ECC-United Cache (EUC) architecture to improve the efficiency of EDCs/ECCs in set-associative L1 caches. EUC architecture extends the data protection granularity from a single word to multiple words by exploiting the parallel cache lines access, which is inherently available in the cache. As compared with the conventional architecture, EUC can be configured to provide: 1) the same protection capability with a significantly lower overhead, 2) a significantly higher protection capability with the same number of check bits, or 3) a trade-off between the former two features. Simulation results show that, when configured to minimize the overhead, EUC reduces the number of check bits by 69 and 75 percent in data-cache and instruction-cache, respectively. When configured to maximize the protection capability, EUC provides fourfold higher burst error detection/correction capability. Moreover, EUC is orthogonal to previous protection schemes and they can be redesigned based on the EUC architecture to further improve their efficiency.
Hamed Farbeh, Leila Delshadtehrani, Hyeonggyu Kim, Soontae Kim
IEEE Trans. Computers4
2021 CID: Co-Architecting Instruction Cache and Decompression System for Embedded Systems
abstract
Code compression is widely used to reduce the footprint of code memory in cost-sensitive embedded systems. However, despite the small code size, the decompressor and the address translator required to support the code compression incur energy and area overheads. To reduce such overheads while still supporting code compression, we co-architect the instruction cache and decompression system (CID). In CID, each component is placed at the optimal location and the instruction cache is redesigned to recognize the compression state and retain the original address, through the cache division and address space decompression process. As a result of the cache division, the energy consumption and area overheads of the CID instruction cache are reduced. Since the decompressor overhead depends on the code compression technique, we propose a new code compression technique called entropy-based pattern code compression, which reduces overheads of the decompressor. Our experimental results show that the total energy consumption of the instruction cache and decompression system is reduced by up to 29.7 percent and their area is reduced by up to 15.4 percent compared to the post-cache architecture with almost no performance degradation, while achieving an 18.8 percent improvement in the compression ratio compared to the state-of-the-art code compression technique.
Jinkwon Kim, Seokin Hong, Jeongkyu Hong, Soontae Kim
IEEE Trans. Computers4
2021 Update Frequency-Directed Subpage Management for Mitigating Garbage Collection and DRAM Overheads
abstract
The increased flash page sizes cause a large number of subpage requests due to the difference in host and flash I/O units. The subpage requests may degrade the space utilization, response time, and lifetime of NAND flash memories. Addressing the subpage issue, a few studies in the past have proposed merging subpages to generate full pages. Although these subpage schemes may improve performance and lifetime, they incur immense DRAM space for storing the sector information of merged pages. Moreover, they do not consider the update frequencies of the subpages for merging and thus an update request to a subpage causes partial page invalidation, which leads to garbage collection (GC) overhead. To address these issues, our proposed scheme considers the update frequencies of the subpages for merging in order to avoid partial page invalidation, which in turn improves the GC efficiency. Further, the proposed scheme uses a sector information table (SIT) in flash pages to store the fine-grained sector information of merged subpages. In the light of experiment results, our scheme, on average, reduces DRAM footprint, flash writes, and block erasures by 29%, 18%, and 13%, respectively.
Imran Fareed, Mincheol Kang, Wonyoung Lee 0001, Soontae Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Leveraging intra-page update diversity for mitigating write amplification in SSDs
abstract
A solid state drive (SSD) receives requests in multiple of sectors from the host system, which are then mapped to logical pages, the basic I/O units of the flash memory. As the SSD receives requests in sector units, the sectors in a logical page tend to exhibit diverse update frequencies. Therefore, frequent updates to some sectors of a page cause other sectors of the same page to be unnecessarily read and written to other free pages, thereby increasing write amplification and harming the flash memory lifetime. To eliminate unnecessary sector movement and to reduce write amplification, we propose a sector-level classification (SLC) technique. SLC considers the diversity in the update frequencies of sectors and merges sectors with similar update frequencies to generate full, homogeneous pages. Thus, multiple update operations can be converged to a single flash page, thereby reducing write amplification and increasing flash memory lifetime. SLC handles the merged sectors using the proposed shared-page mapping table (SMT), whereas pages whose sectors remain unmerged are handled by a conventional page mapping table. Despite the SMT overhead, SLC does not require excessive resources to accommodate SMT. The capability of SLC is evaluated by a series of experiments, which provides highly encouraging results. It is demonstrated that SLC reduces flash writes, flash reads, block erasures, and flash writes execution time by 42%, 23%, 45%, and 37%, respectively.
Imran Fareed, Mincheol Kang, Wonyoung Lee 0001, Soontae Kim
ICS4
2020 Freezing: Eliminating Unnecessary Drawing Computation for Low Power
abstract
Numerous functionalities are provided by smartphones and prolonging their battery lifetime is an undeniably critical support issue. In this paper, we propose a low-power scheme called unnecessary region freezing (URF) that achieves significant total power consumption reduction in smartphones by reducing unnecessary drawing computation that originates from a user's uninteresting display region. We implement URF on a real smartphone (Nexus 6). Then, we evaluate our scheme on top of an existing low-power scheme, partial display darkening (PDD). The results indicate that the URF scheme reduces total power consumption by average 5.1% through 12.5% depending on the size of the unnecessary region, compared with the PDD scheme.
Bohun Seo, Hyeonggyu Kim, Soontae Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 SALE: Smartly Allocating Low-Cost Many-Bit ECC for Mitigating Read and Write Errors in STT-RAM Caches
abstract
Spin-transfer torque RAM (STT-RAM) is a future technology for ON-chip caches. However, it suffers from high read and write error rates. Concurrently dealing with these errors is quite challenging and incurs large performance overhead. This article proposes a smartly allocating low-cost many-bit ECC (SALE) scheme, which makes use of the low-cost many-bit error correction coding (ECC) to overcome this performance overhead. The low-cost many-bit ECC can fix many errors with low logic complexity and latency overheads. However, it requires a large number of parity bits. Therefore, SALE smartly uses low-cost many-bit ECC for only a certain type of cache lines and manages the corresponding large number of parity bits in the data array. SALE also introduces an ECC-free partition to reduce the ECC storage requirement for the STT-RAM caches. The cache lines belonging to an ECC-free partition do not have dedicated storage space for the ECC parity bits, thereby reducing the ECC storage requirement for the STT-RAM caches. Our experimental results demonstrate that SALE achieves performance close to that of an error-free cache by improving performance by 13% (16%) over the baseline scheme in single-core (quad-core) systems while requiring 50% less storage space for the ECC parity bits.
Muhammad Avais Qureshi, Jungwoo Park, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2019 MH Cache: A Mult Stephen Jarvisi-retention STT-RAM-based Low-power Last-level Cache for Mobile Hardware Rendering Systems
abstract
Mobile devices have become the most important devices in our life. However, they are limited in battery capacity. Therefore, low-power computing is crucial for their long lifetime. A spin-transfer torque RAM (STT-RAM) has become emerging memory technology because of its low leakage power consumption. We herein propose MH cache, a multi-retention STT-RAM-based cache management scheme for last-level caches (LLC) to reduce their power consumption for mobile hardware rendering systems. We analyzed the memory access patterns of processes and observed how rendering methods affect process behaviors. We propose a cache management scheme that measures write-intensity of each process dynamically and exploits it to manage a power-efficient multi-retention STT-RAM-based cache. Our proposed scheme uses variable threshold for a process’ write-intensity to determine cache line placement. We explain how to deal with the following issue to implement our proposed scheme. Our experimental results show that our techniques significantly reduce the LLC power consumption by 32% and 32.2% in single- and quad-core systems, respectively, compared to a full STT-RAM LLC.
Jungwoo Park, Myoungjun Lee, Soontae Kim, Minho Ju, Jeongkyu Hong
ACM Trans. Archit. Code Optim.3
2019 Time-sensitivity-aware shared cache architecture for multi-core embedded systems
Myoungjun Lee, Soontae Kim
J. Supercomput.2
2019 Interpage-Based Endurance-Enhancing Lower State Encoding for MLC and TLC Flash Memory Storages
abstract
During the past decade, the endurance of NAND flash memory has severely deteriorated. The maximum number of program and erase cycles has fallen significantly with emerging of multilevel cell (MLC) and triple-level cell (TLC) technology, and scaling down of the cell size. Wear leveling is a general solution used to alleviate this issue; it enables cells to wear down evenly but it cannot actually mitigate the wearing of the cells. Accordingly, techniques are required to minimize the actual cell degradation. This paper proposesendurance-enhancing lower state encoding. The key insight leveraged by the proposed technique is the data pattern-related characteristic of MLC and TLC NAND flash memories, in which the lower the state of the cells, the lower the occurrence of wear out. Thus, our proposed scheme encodes input data to make the cell state as low as possible in consideration of interpage relation. As a result, the wear out of the memory cells can be minimized and their lifetime is improved by 62.7% in a file type and 43.0% in MySQL. Experimental results indicate that our scheme shows better lifetime improvement than other schemes in most cases.
Wonyoung Lee 0001, Mincheol Kang, Seokin Hong, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2019 A Restore-Free Mode for MLC STT-RAM Caches
abstract
Spin-transfer torque RAM (STT-RAM) caches are foreseen to replace traditional static RAM caches because of their nonvolatile nature and high density. Multilevel cell (MLC) STT-RAMs further enhance the storage density of single-level cell STT-RAMs. However, the two-step read/write process in MLC STT-RAMs adversely affects performance, energy consumption, and lifetime. Moreover, technology scaling makes the read operations disturb the stored data in MLC STT-RAMs, giving rise to an issue called read disturbance (RD). Restore operations, which are required to cope with RD, further add to the problems of using MLCs. In this brief, we propose a Restore-free mode for frequently reused cache lines in MLC STT-RAMs that leverages single-step read/write operations to the MLC without the need for restore operations. Our proposed scheme in single-core (quad-core) systems, achieves a 27.4% (23%) dynamic energy reduction, a 3.7% (7%) increase in performance, and an 81% (62.5%) lifetime improvement.
Muhammad Avais Qureshi, Hyeonggyu Kim, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2018 EAR: ECC-aided refresh reduction through 2-D zero compression
abstract
Continuous DRAM scaling and integration is making refresh operations not scalable because more rows have to be refreshed in the same refresh interval. Thus, refresh operations are expected to consume more power in future DRAMs. To alleviate this problem, we propose a compression-based refresh-reducing DRAM architecture. To exploit prevalent zero and small memory values, we devise a novel 2-D ZERO compression scheme to increase compression coverage significantly with simple hardware support. 2-D ZERO compression can achieve 77% compression coverage compared to 50% of the conventional zero-value compression. The freed space of memory blocks obtained by 2-D zero compression is exploited to store ECC bits, which can correct bit-errors that occur when the refresh interval is lengthened to reduce refresh operations. Experimental results show that more than 98% of refresh operations can be removed and overall DRAM power consumption is reduced by 9%.
Jeongkyu Hong, Hyeonggyu Kim, Soontae Kim
PACT3
2018 Subpage-Aware Solid State Drive for Improving Lifetime and Performance
abstract
The manufacturers of NAND flash-based solid-state drives (SSDs) are increasing capacity and throughput by enlarging their page size, which is the minimum I/O unit in the NAND flash chips. Because the host and NAND flash chips have different I/O granularity units, the number of subpage requests increases. However, these subpage requests, especially writes, can cause internal fragmentation and endurance problems. Furthermore, subpage write requests inevitably involve read-modify-write (RMW) operations that increase the write response time because of the out-place-update feature in the NAND flash chips. In this paper, we propose a subpage-aware SSD to increase the lifetime and performance by reducing the number of NAND writes and eliminating unnecessary RMW operations. Our scheme attempts to merge subpage write requests to full page write requests in the write buffer to reduce the number of NAND writes and adds size information to the mapping table to detect unnecessary RMW operations. Our proposed scheme reduces the number of NAND writes by up to 30 and 19 percent on average and the write response time by up to 22 and 13 percent on average.
Mincheol Kang, Wonyoung Lee 0001, Soontae Kim
IEEE Trans. Computers3
2018 OnNetwork+: Network Delay-Aware Management for Mobile Systems
abstract
Network errors such as packet losses consume large amounts of energy. We analyzed the reason for this through measurements using the latest smartphones and full-system simulation. We found that on packet losses the smartphones maintain high frequencies for CPU without doing useful work. To address this problem, we propose a method for reducing the energy consumption by lowering the performance level by exploiting a dynamic voltage and frequency scaling mechanism when long network delays are expected. According to our experiments, our method reduces the total energy consumption of web browsing on two different smartphones by up to 10.0% and 11.5%, respectively.
Hyeonggyu Kim, Minho Ju, Soontae Kim
ACM Trans. Embed. Comput. Syst.3
2017 Partial Row Activation for Low-Power DRAM System
abstract
Owing to increasing demand of faster and larger DRAM system, the DRAM system accounts for a large portion of the total power consumption of computing systems. As memory traffic and DRAM bandwidth grow, the row activation and I/O power consumptions are becoming major contributors to total DRAM power consumption. Thus, reducing row activation and I/O power consumptions has big potential for improving the power and energy efficiency of the computing systems. To this end, we propose a partial row activation scheme for memory writes, in which DRAM is rearchitected to mitigate row overfetching problem of modern DRAMs and to reduce row activation power consumption. In addition, accompanying I/O power consumption in memory writes is also reduced by transferring only a part of cache line data that must be written to partially opened rows. In our proposed scheme, partial rows ranging from a one-eighth row to a full row can be activated to minimize row activation granularity for memory writes and the full bandwidth of the conventional DRAM can be maintained for memory reads. Our partial row activation scheme is shown to reduce total DRAM power consumption by up to 32% and 23% on average, which outperforms previously proposed schemes in DRAM power saving with almost no performance loss.
Yebin Lee, Hyeonggyu Kim, Seokin Hong, Soontae Kim
HPCA4
2017 Smart ECC Allocation Cache Utilizing Cache Data Space
abstract
Conventional error correcting codes (ECC) for caches are applied to all cache lines and stored in dedicated SRAM storage, which incurs both space and energy overheads. In contrast, we propose a Smart ECC Allocation (SEA) cache that utilizes cache data space for low-cost error protection of last-level caches. SEA cache avoids the requirement for dedicated storage for ECC check bits, which are stored in cache lines as data. To effectively utilize cache space, we group several cache sets and manage them according to program behavior. SEA cache eliminates the considerable space overheads of conventional ECC schemes without noticeable reliability and performance degradation.
Jeongkyu Hong, Soontae Kim
IEEE Trans. Computers2
2017 TLB Index-Based Tagging for Reducing Data Cache and TLB Energy Consumption
abstract
Conventional cache tag matching identifies the requested data based on a memory address. However, this address-based tag matching is inefficient because it requires unnecessarily many tag bits. Previous studies show that translation look-aside buffer (TLB) index-based tagging (TLBIT) can be adopted in instruction caches because there are not many different tags at a given moment due to spatial locality, and those tags can be captured by TLBs. For the TLBIT scheme, extra TLB indices are added to each TLB entry and conventional cache tags are replaced with TLB indices to identify the requested data in the cache. TLBIT reduces the number of required tag bits in tag arrays; therefore, the cache energy consumption and area are decreased. In this paper, we show that naively adopting TLBIT for data caches is inefficient, in terms of performance and energy consumption, because of cache line searches and invalidations on TLB misses. To achieve the true potential of TLBIT, we propose four novel techniques: search zone, c-LRU, TLB buffer and demand address fetching. The search zone reduces unnecessary cache line searching and c-LRU reduces the cache line invalidations. The TLB buffer prevents immediate cache line invalidations on TLB misses. Furthermore, we present demand address fetching to reduce energy consumption in the TLB. From our experiments, we observed that the proposed techniques reduce the overall dynamic energy consumption of the data cache by 14.3 percent on average. The overall tag array area and leackage power of the data cache are also reduced by 54 and 45 percent, respectively. The TLB energy consumption is reduced by 22.7 percent. The performance impact is small, less than 0.4 percent on average. We also demonstrate that TLBIT can be applied to large caches, and set-associative TLBs.
Jesung Kim, Jongmin Lee 0002, Soontae Kim
IEEE Trans. Computers3
2017 Write-Amount-Aware Management Policies for STT-RAM Caches
abstract
Spin-transfer torque random access memory (STT-RAM) technology has emerged as one of the most promising memory technologies owing to its nonvolatility, high density, and low-leakage power characteristics. However, STT-RAM has certain drawbacks such as high write energy consumption and limits to the number of write cycles. To enable the adoption of STT-RAM in the implementation of cache memories, new cache hierarchy management policies are required to overcome such drawbacks. In this brief, we evaluated several cache hierarchy management policies in the context of static random access memory L1 caches and an STT-RAM L2 cache. We found that a nonexclusive policy is superior to noninclusive and exclusive policies in terms of energy consumption and endurance. We also propose a sub-block-based management policy because the write energy consumption and endurance are proportional and inversely proportional to the amount of written data, respectively. A combination of the proposed policy with a nonexclusive policy reduces the L2 cache energy consumption by 33.3% (31.5%) and improves the lifetime by 56.3% (56.8%) in a single-core (quad-core) system.
Hyeonggyu Kim, Soontae Kim, Jooheung Lee
IEEE Trans. Very Large Scale Integr. Syst.2
2017 A Way-Filtering-Based Dynamic Logical-Associative Cache Architecture for Low-Energy Consumption
abstract
Last-level caches (LLCs) help improve performance but suffer from energy overhead because of their large sizes. An effective solution to this problem is to selectively power down several cache ways, which, however, reduces cache associativity and performance and thus limits its effectiveness in reducing energy consumption. To overcome this limitation, we propose a new cache architecture that can logically increase cache associativity of way-powered-down LLCs. Our proposed scheme is designed to be dynamic in activating an appropriate number of cache ways in order to eliminate the need for static profiling to determine an energy-optimized cache configuration. The experimental results show that our proposed dynamic scheme reduces the energy consumption of LLCs by 34% and 40% on single- and dual-core systems, respectively, compared with the best performing conventional static cache configuration. The overall system energy consumption including CPU, L2 cache, and DRAM is reduced by 9.2% on quad-core systems.
Jungwoo Park, Jongmin Lee 0002, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Network delay-aware energy management for mobile systems
Minho Ju, Hyeonggyu Kim, Soontae Kim
DATE3
2016 MofySim: A mobile full-system simulation framework for energy consumption and performance analysis
abstract
The analysis of energy consumption and performance is essential to design and optimize mobile systems because of their limited battery capacity. Full-system simulation provides detailed performance metrics for an entire system. Thus it has been widely used for designing and optimizing microarchitectures and mobile systems. The gem5 simulator provides full-system simulation based on the ARM architecture and Android for mobile systems. However, gem5 for mobile systems does not support wireless network interfaces and can not configure various networking environments such as network errors and network types. Furthermore, gem5 provides only performance statistics without power consumption data. This paper presents a mobile full-system simulation framework based on an enhanced gem5 that includes a simulated mobile system, a simulated server system, and a simulated Ethernet, which enables us to configure various networking environments, in addition to power models for the main components of mobile systems: CPU/caches, DRAM, network interfaces, and display. Using mobile applications and SPEC CPU2006 benchmarks, we show that the proposed mobile full-system simulator achieves performance accuracy within 26.8% error rate for various network packet loss rates, and power modeling accuracy within 12.8% error rate, compared with Nexus 5. This mobile full-system simulator considering the real networking environments provides the energy consumption and performance analysis of not only hardware components, but also application processes and threads at the same time. We also discovered energy-inefficient tasks and the inefficiency of the DVFS ondemand governor on network delays using the proposed mobile full-system simulation framework.
Minho Ju, Hyeonggyu Kim, Soontae Kim
ISPASS3
2016 Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches
abstract
Spin-Transfer Torque RAM (STT-RAM) is a promising alternative to SRAM for implementing on-chip L2 and L3 caches. One of the most critical challenges in STT-RAM is reliability due to limited write endurance, which results in insufficient lifetime, as well as various types of errors. Previous studies have focused on either presenting various cache architectures/management techniques to improve the lifetime of STT-RAM caches or utilizing different Error Correcting Codes (ECCs) to protect against the permanent and transient errors. However, there is no quantitative analysis in the literature to determine the impact of ECCs on the lifetime of the STT-RAM caches. This paper formulates this impact and demonstrates that ECCs shorten the lifetime of STT-RAM cache lines by more than 50 percent due to ECCs high write activity. Then, we propose the Floating-ECC architecture for increasing the lifetime of the STT-RAM caches. The main idea is to evenly distribute the ECC write activity over all bits of cache lines by periodically relocating the ECC bits inside the cache lines. The simulation results for the most conventional ECC scheme, i.e., interleaved Single Error Correction-Double Error Detection (SEC-DED), show that Floating-ECC increases the lifetime of L2 and L3 caches by more than 318 percent and 254 percent, respectively.
Hamed Farbeh, Hyeonggyu Kim, Seyed Ghassem Miremadi, Soontae Kim
IEEE Trans. Computers4
2016 Designing a Resilient L1 Cache Architecture to Process Variation-Induced Access-Time Failures
abstract
Continuous scaling of process technology increases variations in transistors. The process variations cause large fluctuations in the access times of static random-access memory (SRAM) cells. Caches made of those SRAM cells cannot be accessed within the target clock cycle time, which reduces the yield of processors. Many schemes have been proposed to combat these access time failures in caches. However, these schemes are limited in their coverage and do not scale well at high failure rates. We propose a new level one (L1) cache architecture employing multi-cycle cell access (MCCA) and subarray-level parallel access (SLPA). MCCA eliminates all access-time failures in L1 caches. SLPA minimizes the performance impact of cache bandwidth loss due to MCCA. For further performance improvement, architectural techniques are proposed. Our experimental results show that our proposed L1 cache architecture incurs a performance hit of less than 1.2 percent compared to the conventional cache architecture with no access time failure. Our proposed architecture is not sensitive to access time failure rates and has a low overhead compared with previously proposed competitive schemes.
Seokin Hong, Soontae Kim
IEEE Trans. Computers2
2016 RAMS: DRAM Rank-Aware Memory Scheduling for Energy Saving
abstract
DRAMs are one of the main players of the computer system energy consumption. Thus, reducing DRAM energy consumption has a big potential to save the entire system energy consumption. Because the standby power consumption of DRAM is significant, modern DRAMs provide low-power modes for reducing idle energy consumption. However, the use of low-power modes can degrade the performance because state transitions to/from low-power states involve a time penalty. To effectively utilize low-power modes, we propose DRAM rank-aware memory scheduling schemes. One scheme utilizes a prioritized cache block replacement method considering the power states of DRAM ranks to select victim blocks for the late level cache. Through this scheme, DRAM traffic and the number of state transitions of DRAM ranks can be reduced. The other scheme utilizes the memory controller by controlling write traffic to DRAM with the awareness of the DRAM rank states. DRAM rank idle times and state transitions can be reduced by this scheme. Our proposed schemes are shown to reduce DRAM energy consumption by 15.2 percent on average.
Yebin Lee, Soontae Kim
IEEE Trans. Computers2
2016 Write Buffer-Oriented Energy Reduction in the L1 Data Cache for Embedded Systems
abstract
In resource-constrained embedded systems, on-chip cache memories play an important role in both performance and energy consumption. In contrast to read operations, scant regard has been paid to optimizing write operations even though the energy consumed by write operations in the data cache constitutes a large portion of the total energy consumption. Consequently, this paper proposes a write buffer-oriented (WO) cache architecture that reduces energy consumption in the L1 data cache. Observing that write operations are very likely to be merged in the write buffer because of their high localities, we construct the proposed WO cache architecture to utilize two schemes. First, the write operations update the write buffer but not the L1 data cache, which is updated later by the write buffer after the write operations are merged. Write merging significantly reduces write accesses to the data cache and, consequently, energy consumption. Second, we further reduce energy consumption in the write buffer by filtering out unnecessary read accesses to the write buffer using a read hit predictor. In this paper, we show that the proposed WO cache architecture is applicable to the conventional embedded processors that support both write-through and write-back policies. Further, the experimental results verify that the proposed cache architecture reduces energy consumption in data caches up to 14%.
Jongmin Lee 0002, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Flexible ECC Management for Low-Cost Transient Error Protection of Last-Level Caches
abstract
The conventional error correcting code (ECC) schemes for caches are based on a fixed mapping between cache data words and ECC check bits, and fixed ECC word granularity. This leads to inefficient usage of the ECC check bits. We propose to manage the check bits flexibly for low-cost error protection of last-level caches. The proposed ECC schemes work at the word level, whereas the conventional ECC schemes work at the cache line or set level. The proposed schemes protect only dirty words with ECC check bits using a flexible mapping. Moreover, the proposed schemes utilize variable ECC word granularities. Dirty (modified) words that are unlikely to be modified further before being evicted are collectively protected with a larger ECC word granularity. The proposed schemes reduce DRAM and data bus energy overheads by 28% and 45%, respectively, with the same area overhead as previously proposed competitive schemes. Our schemes show more energy reduction results for multicore systems without noticeable performance degradation.
Jeongkyu Hong, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2016 CLAP: Clustered Look-Ahead Prefetching for Energy-Efficient DRAM System
abstract
DRAM is one of the main sources of energy consumption in computer systems. Thus, reducing the energy consumption of DRAM can prolong the lifetime of battery-operated embedded/mobile systems. To this end, we propose a DRAM energy-aware prefetching scheme to increase row buffer hits and idle periods of DRAM by clustering its accesses. Although prefetching schemes have traditionally been used to improve the system performance, utilizing them for the energy conservation of DRAM has yet to be investigated. For such energy conservation, our scheme accurately predicts and clusters potential future DRAM accesses. Clustered DRAM accesses exploit a popular first-ready first-come first-serve memory request scheduling and a power-down mode of DRAM more effectively; the probability of row buffer hits and idle periods is significantly increased by our clustering scheme. As a result, large amounts of row activation and idle energy consumption, which are major energy consumption factors in modern DRAM, can be saved. Our prefetching-based memory traffic-clustering scheme was shown to reduce the power and energy consumption of DRAM and improve its performance by an average of 0.2%, 28.9%, and 15.7%, respectively, for memory-intensive programs.
Yebin Lee, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Filter Data Cache: An Energy-Efficient Small L0 Data Cache Architecture Driven byMiss Cost Reduction
abstract
On-chip cache memories play an important role in resource-constrained embedded systems by filtering out most off-chip memory accesses. Because cache latency and energy consumption are generally proportional to cache sizes, a small cache at the top level of the memory hierarchy is desirable. Previous work has presented a novel cache architecture called a filter cache to reduce hit time and energy consumption of the L1 instruction cache. However, consideration to the data cache requires a different approach and has not been researched much. In this paper, we propose a filter data cache architecture to effectively adopt the filter cache to the data cache hierarchy. We observed that cache misses occur considerably and they are likely to be continuous when the filter cache is used for the data cache. Those misses cost performance and energy consumption by increasing cache latency and uploading unnecessary data. The proposed filter data cache architecture reduces miss costs using three schemes: early cache hit predictor (ECHP), locality-based allocation (LA), and No Tag Matching Write (NTW). Experimental results show that the proposed filter data cache reduces energy consumption of the data caches by 21% compared with the filter cache, and the energy consumption of the ALU by 27.2 percentage on average. The overheads in terms of area and leakage power are small and the proposed filter data cache architecture does not hurt performance.
Jongmin Lee 0002, Soontae Kim
IEEE Trans. Computers2
2015 A Low-Cost Mechanism Exploiting Narrow-Width Values for Tolerating Hard Faults in ALU
abstract
Digital circuits are expected to increasingly suffer from more hard faults due to technology scaling. Especially, a single hard fault in ALU (Arithmetic Logic Unit) might lead to a total failure in processors or significantly reduce their performance. To address these increasingly important problems, we propose a novel cost-efficient fault-tolerant mechanism for the ALU, called LIZARD. LIZARD employs two half-word ALUs, instead of a single full-word ALU, to perform computations with concurrent fault detection. When a fault is detected, the two ALUs are partitioned into four quarter-word ALUs. After diagnosing and isolating a faulty quarter-word ALU, LIZARD continues its operation using the remaining ones, which can detect and isolate another fault. Even though LIZARD uses narrow ALUs for computations, it adds negligible performance overhead through exploiting predictability of the results in the arithmetic computations. We also present the architectural modifications when employing LIZARD for scalar as well as superscalar processors. Through comparative evaluation, we demonstrate that LIZARD outperforms other competitive fault-tolerant mechanisms in terms of area, energy consumption, performance and reliability.
Seokin Hong, Soontae Kim
IEEE Trans. Computers2
2015 Ensuring Cache Reliability and Energy Scaling at Near-Threshold Voltage With Macho
abstract
Nanoscale process variations in conventional SRAM cells are known to limit voltage scaling in microprocessor caches. Recently, a number of novel cache architectures have been proposed which substitute faulty words of one cache line with healthy words of others, to tolerate these failures at low voltages. These schemes rely on the fault maps to identify faulty words, inevitably increasing the chip area. Besides, the relationship between word sizes and the cache failure rates is not well studied in these works. In this paper, we analyze the word substitution schemes by employing Fault Tree Model and Collision Graph Model. A novel cache architecture (Macho) is then proposed based on this model. Macho is dynamically reconfigurable and is locally optimized (tailored to local fault density) using two algorithms: 1) a graph coloring algorithm for moderate fault densities and 2) a bipartite matching algorithm to support high fault densities. An adaptive matching algorithm enables on-demand reconfiguration of Macho to concentrate available resources on cache working sets. As a result, voltage scaling down to 400 mV is possible, tolerating bit failure rates reaching 1 percent (one failure in every 100 cells). This near-threshold voltage (NTV) operation achieves 44 percent energy reduction in our simulated system (CPU+DRAM models) with a 1 MB L2 cache.
Tayyeb Mahmood, Seokin Hong, Soontae Kim
IEEE Trans. Computers3
2015 Exploiting Same Tag Bits to Improve the Reliability of the Cache Memories
abstract
With the trend of increasing transient error rate, it is becoming important to prevent transient errors and provide a correction mechanism for hardware circuits, especially for SRAM cache memories. Caches are the largest structures in current microprocessors and, hence, are most vulnerable to the transient errors. Tag bits in cache memories are also exposed to transient errors but a few efforts have been made to reduce their vulnerability. In this paper, we propose to exploit prevalent same tag bits to improve error protection capability of the tag bits in the caches. When data are fetched from the main memory, it is checked if adjacent cache lines have the same tag bits as those of the data fetched. This same tag bit information is stored in the caches as extra bits to be used later. When an error is detected in the tag bits, the same tag bit information is used to recover from the error in the tag bits. The proposed scheme has small area, energy, and performance overheads with error protection coverage of 97.9% on average. Even with large working sets and various cache sizes, our scheme shows protection coverage of higher than 95% on average.
Jeongkyu Hong, Jesung Kim, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Low-Cost Control Flow Protection via Available Redundancies in the Microprocessor Pipeline
abstract
Miniaturization of very large scale integration circuits, higher frequencies and reduction of supply voltages make embedded systems more susceptible to soft errors (or transient errors). Soft errors affect the processor's pipeline and hence its data and control flows. Specifically, errors in control flows can change program's execution sequence, which might be catastrophic for safety-critical applications. Several state-of-the-art techniques are available for control flow error checking (CFEC). Software-based techniques suffer from increased code size overhead and can have a negative impact on energy consumption. On the other hand, hardware-based schemes incur high hardware and area costs. In this paper, a low-cost CFEC scheme is proposed that exploits available redundancies in the processor's pipeline; a branch target buffer stores the target addresses of taken branches, a short backward branch detector stores short loop branch targets and an arithmetic logic unit generates branch target addresses using the low-order branch displacement bits of branch instructions. The proposed CFEC scheme uses these redundancies to detect and recover from control flow errors in the pipeline with low energy overhead of 0.9% and performance overhead of 0.8%, while its error coverage ranges from 86% to 99%.
Mohammad Abdur Rouf, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Ternary cache: Three-valued MLC STT-RAM caches
abstract
Spin-transfer torque random access memory (STT-RAM) has become a promising non-volatile memory technology for cache memories. Recently, 2-bit multi-level cell (MLC) STT-RAM has been proposed to enhance data density, but it suffers from low reliability of its read and write operations. In this paper, we propose a novel cache design called Ternary cache. In Ternary cache, a memory cell can store three values (i.e., 0,1,2) while MLC STT-RAM can store four values. In this way, Ternary cache achieves much higher read stability than MLC STT-RAM-based caches. To enhance writability, a write operation is performed with high current and terminated as soon as the data is written. Evaluation results show that Ternary cache achieves the data density benefit of MLC STT-RAM and the reliability benefit of SLC STT-RAM.
Seokin Hong, Jongmin Lee 0002, Soontae Kim
ICCD3
2013 AVICA: an access-time variation insensitive L1 cache architecture
abstract
Ever scaling process technology increases variations in transistors. The process variations cause large fluctuations in the access times of SRAM cells. Caches made of those SRAM cells cannot be accessed within the target clock cycle time, which reduces yield of processors. To combat these access time failures in caches, many schemes have been proposed, which are, however, limited in their coverage and do not scale well at high failure rates. We propose a new L1 cache architecture (AVICA) employing asymmetric pipelining and pseudo multi-banking. Asymmetric pipelining eliminates all access time failures in L1 caches. Pseudo multi-banking minimizes the performance impact of asymmetric pipelining. For further performance improvement, architectural techniques are proposed. Our experimental results show that our proposed L1 cache architecture incurs less than 1% performance hit compared to the conventional cache architecture with no access time failure. Our proposed architecture is not sensitive to access time failure rates and has low overheads compared to the previously proposed competitive schemes.
Seokin Hong, Soontae Kim
DATE2
2013 Skinflint DRAM system: Minimizing DRAM chip writes for low power
abstract
DRAMs are one of the main players of computer system energy consumption due to their large capacities and frequent accesses. Consequently, many schemes have been proposed to reduce DRAM power/energy consumption. Some of them propose new DRAM system and chip organizations, which are effective in reducing power consumption but intrusive. In contrast, we minimize DRAM write accesses at chip level with minimal modification of the conventional DRAM system organization and small addition to caches. When all data going to the same DRAM chips are not modified, the chips are not accessed. Consequently, chips are accessed selectively in our scheme while all chips are accessed simultaneously in the conventional DRAM system. Our chip-based selective DRAM write scheme is shown to reduce DRAM power and energy consumptions by 17% and 14%, respectively, on average. The overheads of our scheme are small in terms of performance, area, and energy consumption.
Yebin Lee, Soontae Kim, Seokin Hong, Jongmin Lee 0002
HPCA2
2013 Macho: A failure model-oriented adaptive cache architecture to enable near-threshold voltage scaling
abstract
Recent interest in CMOS voltage scaling has produced a class of cache architectures which tolerate parametric SRAM failures at low voltage by substituting faulty words of one cache line with healthy words of another line. These caches rely on the fault maps (which grow reciprocally with smaller word sizes) for fault identification. Therefore, the benefits of cache voltage scaling must be rigorously investigated against the cost of their fault map overheads, especially in large caches. This paper reviews the word substitution caches and develops their parametric failure model. Our developed model leads to a non-intrusive and reconfigurable cache (Macho) which can be locally optimized (based on local fault density) by two graph-based algorithms. Specifically, our adaptive matching algorithm increases effective cache capacity by dynamically concentrating healthy cache blocks into active cache sets. Macho enables voltage scaling down to 400mV by tolerating high SRAM-failure rates (≥ 1%) and achieves better energy reduction (44%) than other substitution caches with similar area overheads.
Tayyeb Mahmood, Soontae Kim, Seokin Hong
HPCA2
2013 Performance-controllable shared cache architecture for multi-core soft real-time systems
abstract
Multi-core processors with shared L2 caches can improve performance and integrate several functions of real-time systems on a single chip. However, tasks running on different cores increase interferences in the shared L2 cache, resulting in more deadline misses and, consequently, worse quality of real-time tasks. This is mainly because of the blind sharing of the L2 cache by multiple tasks running on different cores.We propose a novel performance-controllable shared L2 cache architecture that can alleviate these problems. First, our proposed L2 cache architecture is made to be aware of instructions/data belonging to real-time tasks by adding a real-time indication bit to each L2 cache block. Second, it can control the performance of real-time tasks and non-real-time tasks. Our experimental results show that our proposed L2 cache architecture reduces more deadline misses of real-time tasks than the conventional L2 cache architecture and partitioning schemes.
Myoungjun Lee, Soontae Kim
ICCD2
2012 Low-cost control flow error protection by exploiting available redundancies in the pipeline
abstract
Due to device miniaturization and reducing supply voltage, embedded systems are becoming more susceptible to transient faults. Specifically, faults in control flow can change the execution sequence, which might be catastrophic for safety critical applications. Many techniques are devised using software, hardware or software-hardware co-design for control flow error checking. Software techniques suffer from a significant amount of code size overhead, and hence, negative impact on performance and energy consumption. On the other hand, hardware-based techniques have a significant amount of hardware and area cost. In this research we exploit the available redundancies in the pipeline. The branch target buffer stores target addresses of taken branches, and ALU generates target addresses using the low-order branch displacement bits of branch instructions. To exploit these redundancies in the pipeline, we propose a control flow error checking (CFEC) scheme. It can detect control flow errors and recover from them with negligible energy and performance overhead.
Mohammad Abdur Rouf, Soontae Kim
ASP-DAC2
2012 ECC string: Flexible ECC management for low-cost error protection of L2 caches
abstract
Conventional error correcting codes (ECC) scheme for caches is based on fixed mapping between cache words and ECC check bits, and fixed ECC word granularity, which leads to inefficient usage of ECC check bits. In contrast, we propose to use the ECC check bits flexibly for low-cost error protections of L2 caches. Our ECC scheme works at word level while the conventional ECC scheme works at cache line or set level; Our scheme protects only dirty words. In addition, our scheme utilizes variable ECC word granularities; Dirty words that are unlikely to be modified further are protected together with larger ECC word granularity. Our scheme reduces DRAM and data bus energy overheads by 28% and 45% on average, respectively, with the same area overhead as the previously proposed competitive scheme.
Jeongkyu Hong, Soontae Kim
ICCD2
2012 DRAM power-aware rank scheduling
abstract
Modern DRAMs provide multiple low-power states to save their energy consumption during idle times. The use of low-power states, however, can cause performance degradation because state transitions from low-power states to an active state incur time penalty. To effectively utilize the low-power states, we propose DRAM power-aware rank scheduling schemes applied to the last-level cache and the memory controller. Our scheme utilizing the last-level cache reduces write requests to DRAM and the state transitions by replacing cache blocks based on their dirty states and DRAM rank power states. Our scheme utilizing the memory controller decreases the state transitions with rank power state-aware batch writes. With the second scheme, the states transitions are reduced by 21.2%, on average. Consequently DRAM energy consumption is reduced by 11.2%, on average, with no performance loss.
Sukki Kim, Soontae Kim, Yebin Lee
ISLPED2
2012 Adopting TLB index-based tagging to data caches for tag energy reduction
abstract
Conventional cache tag matching is based on addresses to identify requested data. However, this address-based tagging scheme is not efficient because unnecessarily many tag bits are used. Previous studies show that TLB index-based tagging (TLBIT) can be used in caches because there are not many different tags at a moment due to spatial locality, and those tags are conventionally captured by TLBs.
Jongmin Lee 0002, Soontae Kim
ISLPED2
2012 Traffic management strategy for delay-tolerant networks
Kwangcheol Shin, Kyungjun Kim, Soontae Kim
J. Netw. Comput. Appl.3
2012 Resuscitating privacy-preserving mobile payment with customer in complete control
Divyan M. Konidala, Made Harta Dwijaksara, Kwangjo Kim, Dongman Lee, Byoungcheon Lee, Daeyoung Kim 0001, Soontae Kim
Pers. Ubiquitous Comput.7
2012 Predictive routing for mobile sinks in wireless sensor networks: a milestone-based approach
Kwangcheol Shin, Soontae Kim
J. Supercomput.2
2011 ADSR: Angle-Based Multi-hop Routing Strategy for Mobile Wireless Sensor Networks
abstract
Though Dynamic Source Routing (DSR) is a popular on-demand algorithm designed to restrict the bandwidth consumed by control packets, the flooding of Route Request (RREQ) packets to find paths to a destination makes it difficult to apply DSR to wireless sensor networks (WSNs), which have tiny sensor nodes of limited memory and power. In this study, an angle-based multi-hop routing algorithm for mobile WSNs, which is termed ADSR, is proposed. ADSR is a variation of DSR utilizing the location information of a neighboring mobile node and a stationary sink to make the decision of dropping or rebroadcasting the RREQ packet received. ADSR focuses on reducing the number of control packets and, as a result, ADSR improves data deliveries and End-to-End delay. The experimental results using Qualnet modeler software prove that ADSR outperforms DSR in terms of overhead, data delivery, and End-to-End delay regardless of the moving speed or number of sensor nodes.
Kwangcheol Shin, Kyungjun Kim, Soontae Kim
APSCC3
2011 Realizing near-true voltage scaling in variation-sensitive l1 caches via fault buffers
abstract
Voltage scaling can be applied to cache memories to reduce their energy consumptions. However, reduced supply voltage to the cache memories increases defective SRAM cells due to process variations, which will decrease their yields and performance nullifying the benefits of voltage scaling. To mitigate this problem, we propose a fault buffer-based scheme for L1 caches. Faults are identified and isolated at the granularity of individual words in the L1 caches. Actively used faulty cache words are allocated in the fault buffers dynamically. The fault buffers are organized as multiple banks for low cost implementation and can be reconfigured dynamically to reflect varying performance demands of programs. This dynamic scheme is shown to be more energy- and area-efficient than, and to be performing comparably to the previously proposed static schemes.
Tayyeb Mahmood, Soontae Kim
CASES2
2011 Dynamic scheduling algorithm and its schedulability analysis for certifiable dual-criticality systems
abstract
Real-time embedded systems are becoming more complex to include multiple functionalities. Sharing a computing platform is a natural and effective solution to reducing the cost of those systems. However, the sharing can cause serious problems in mixed-criticality systems where applications have different levels of criticality. Certifying the mixed-criticality systems requires efficient scheduling algorithms and schedulability tests different from the ones used in single criticality systems.
Taeju Park, Soontae Kim
EMSOFT2
2011 DRAM energy reduction by prefetching-based memory traffic clustering
abstract
DRAMs consume a large portion of total system energy consumption. Thus, reducing DRAM energy consumption is able to prolong the lifetime of battery-operated embedded/portable systems. To this end, we propose DRAM energy-aware data prefetching scheme to lengthen DRAM idle periods by clustering DRAM accesses. Low-power modes of DRAMs can better exploit longer idle times. We performed experiments with a cycle-accurate simulator with built-in DRAM power model. The experimental results show that our proposed DRAM-aware prefetching is effective in reducing DRAM energy consumption. Up to 77% and average 59% of DRAM energy consumption is saved.
Yebin Lee, Soontae Kim
ACM Great Lakes Symposium on VLSI2
2011 TLB index-based tagging for cache energy reduction
Jongmin Lee 0002, Seokin Hong, Soontae Kim
ISLPED3
2011 Residue cache: a low-energy low-area L2 cache architecture via compression and partial hits
abstract
L2 cache memories are being adopted in the embedded systems for high performance, which, however, increases energy consumption due to their large sizes. We propose a low-energy low-area L2 cache architecture, which performs as well as the conventional L2 cache architecture with 53% less area and around 40% less energy consumption. This architecture consists of an L2 cache and a small cache called residue cache. L2 and residue cache lines are half sized of the conventional L2 cache lines. Well compressed conventional L2 cache lines are stored only in the L2 cache while other poorly compressed lines are stored in both the L2 and residue caches. Although many conventional L2 cache lines are not fully captured by the residue cache, most accesses to them do not incur misses because not all their words are needed immediately, which are termed as partial hits in this paper. The residue cache architecture consumes much lower energy and area than conventional L2 cache architectures, and can be combined synergistically with other schemes such as the line distillation and ZCA. The residue cache architecture is also shown to perform well on a 4-way superscalar processor typically used in high performance systems.
Soontae Kim, Jongmin Lee 0002, Jesung Kim, Seokin Hong
MICRO1
2011 Enhanced buffer management policy that utilises message properties for delay-tolerant networks
abstract
A delay-tolerant network is a network designed so that temporary or intermittent communication problems and limitations have the least possible adverse impact. Two major issues should be considered to achieve data delivery in such challenging networking environments: a routing strategy for the network and a buffer management policy for each node in the network. The routing strategy determines which messages should be forwarded when nodes meet and the buffer management policy determines which message is purged when the buffer overflows in a node. This study proposes an enhanced buffer management policy that utilises message properties. For maximisation of the message deliveries and minimisation of the average delay, two utility functions are proposed on the basis of message properties, particularly the number of replicas, the age and the remaining time-to-live. The experimental results on two types of well-known real-world mobility trace data and synthetic data show that the proposed buffer management policy yields better results over the history-based drop and traditional policies, such as the shortest lifetime first, the most forwarded first in terms of the number of message deliveries and average delay.
Kwangcheol Shin, Soontae Kim
IET Commun.2
2010 SimTag: Exploiting tag bits similarity to improve the reliability of the data caches
abstract
Though tag bits in the data caches are vulnerable to transient errors, few effort has been made to reduce their vulnerability. In this paper, we propose to exploit prevalent same tag bits to improve error protection capability of the tag bits in the data caches. When data are fetched from the main memory, it is checked if adjacent cache lines have the same tag bits as those of the data fetched. This similarity information is stored in the data caches as extra bits to be used later. When an error is detected in the tag bits, the similarity information is used to recover from the error in the tag bits. The proposed scheme has small area, energy, and performance overheads with error protection coverage of 97.9% on average. In contrast, the previously proposed In-Cache Replication scheme is shown to incur large performance and energy overheads.
Jesung Kim, Soontae Kim, Yebin Lee
DATE2
2010 Write buffer-oriented energy reduction in the L1 data cache of two-level caches for the embedded system
abstract
In resource-constrained embedded systems, on-chip cache memories play an important role in both performance and energy consumption points of view. In contrast with read operations, little effort has been made to write operations though write energy consumption in the data cache constitutes a large portion of total energy consumption. To this end, this paper proposes write buffer-oriented energy reduction schemes in the L1 data cache of two-level cache architecture. First, the write operations update only the write buffer but not the L1 data cache which is updated later by the write buffer after the write operations are merged. Write merging significantly reduces write accesses to the data cache and, consequently, energy consumption. Second, many write operations from the write buffer to the L1 data cache do not require tag matching by recording way numbers of the L1 data cache in the write buffer. The two optimizations in combination are shown to reduce write energy consumption by 77% in the L1 data cache and total L1 data cache energy consumption by 27%.
Soontae Kim, Jongmin Lee 0002
ACM Great Lakes Symposium on VLSI1
2010 Lizard: Energy-efficient hard fault detection, diagnosis and isolation in the ALU
abstract
Digital circuits are expected to increasingly suffer from more hard faults due to technology scaling. Especially, a single hard fault in the ALU might lead to a total failure in the embedded systems. In addition, energy efficiency is critical in these systems. To address these increasingly important problems in the ALU, we propose a novel energy-efficient fault-tolerant ALU design called Lizard. Lizard utilizes two 16-bit ALUs to perform 32-bit computations with fault detection and diagnosis. By exploiting predictable operations, fault detection is performed in a single cycle. The 16-bit ALUs can be partitioned into two 8-bit ALUs. When a fault occurs in one of the four 8-bit ALUs, Lizard diagnoses and isolates a faulty 8-bit ALU for itself. After the faulty 8-bit ALU is isolated, Lizard continues its operation using the remaining three 8-bit ALUs, which can detect and isolate another fault. In this way, Lizard can survive faults on at most two sub-ALUs increasing its lifetime and fault tolerance. We conducted comparative evaluations with an unprotected ALU, triple modular redundancy ALU, and quadruple time redundancy ALU in terms of area, energy consumption, performance, and reliability. It is demonstrated that Lizard outperforms other ALU designs in most cases, especially in energy efficiency.
Seokin Hong, Soontae Kim
ICCD2
2010 Modeling and Evaluation of Control Flow Vulnerability in the Embedded System
abstract
Faults in control flow-changing instructions are critical for correct execution because the faults could change the behavior of programs very differently from what they are expected to show. The conventional techniques to deal with control flow vulnerability typically add extra instructions to detect control flow-related faults, which increase both static and dynamic instructions, consequently, execution time and energy consumption. In contrast, we make our own control flow vulnerability model to evaluate the effects of different compiler optimizations. We find that different programs show very different degrees of control flow vulnerabilities and some compiler optimizations have high correlation to control flow vulnerability. The results observed in this work can be used to generate more resilient code against control flow-related faults.
Mohammad Abdur Rouf, Soontae Kim
MASCOTS2
2009 An energy-delay efficient 2-level data cache architecture for embedded system
abstract
We propose a 2-level data cache architecture with a low energy-delay product tailored for the embedded systems. The L1 data cache is small and direct-mapped, and employs a write-through policy. In contrast, the L2 data cache is set-associative and adopts a write-back policy. Consequently, the L1 data cache is accessed fast and is able to provide high cache bandwidth while the L2 data cache is effective in reducing global miss rate. To reduce the penalty of high miss rates caused by the small L1 data cache, we propose an ECP (Early Cache hit Predictor) scheme. The ECP predicts if the L1 cache has the requested data using both partial address generation and L1 cache hit prediction. If so, the L2 data cache is directly accessed. To reduce high energy cost of accessing the L2 data cache due to heavy write-through traffic between the two cache levels, we propose a one-way write scheme. From our simulation-based experiments, the proposed 2-level data cache architecture shows average 3.6% and 50% improvements in overall system performance and energy consumption of the data cache and address generation, respectively.
Jongmin Lee 0002, Soontae Kim
ISLPED2
2009 Reducing Area Overhead for Error-Protecting Large L2/L3 Caches
abstract
Due to increasing concern about various errors, current processors adopt error protection mechanisms for their on-chip components. Especially, protecting caches in current processors incurs as much as 12.5 percent area overhead due to error-correcting codes (ECCs). Considering large L2/L3 caches employed in current high-performance processors, the area overhead is very high, consuming a large number of on-chip transistors. As an attempt to reduce that overhead, this paper proposes an area-efficient error protection architecture for large L2/L3 caches. First, it selectively applies ECC to only dirty cache lines, and other clean cache lines are protected by using simple parity check codes. Second, the dirty cache lines are periodically cleaned by exploiting the generational behavior of cache lines in order not to increase traffic to the off-chip main memory. Experimental results show that the cleaning technique effectively reduces the average number of dirty cache lines per cycle. The ECCs of the reduced dirty cache lines can be confined in a small ECC array or ECC cache. Our proposed error protection architecture has been shown to reduce the area overhead of a 1-Mbyte L2 cache for error protection by 59 percent with less than 1 percent performance degradation, on the average, using SPEC2000 benchmarks running on a typical four-issue superscalar processor. A dirty cache line cleaning scheme is also beneficial for reducing the vulnerability of tag arrays to soft errors.
Soontae Kim
IEEE Trans. Computers1
2009 A Framework for Correction of Multi-Bit Soft Errors in L2 Caches Based on Redundancy
abstract
With the continuous decrease in the minimum feature size and increase in the chip density due to technology scaling, on-chip L2 caches are becoming increasingly susceptible to multi-bit soft errors. The increase in multi-bit errors could lead to higher risk of data corruption and potentially result in the crashing of application programs. Traditionally, the L2 caches have been protected from soft errors using techniques such as: 1) error detection/correction codes; 2) physical interleaving of cache bit lines to convert multi-bit errors into single-bit errors; and 3) cache scrubbing. While the first two methods incur large area overheads for multi-bit errors, identifying the time interval for scrubbing could be tricky. In this paper, we investigate in detail the multi-bit soft error rates in large L2 caches and propose a framework of solutions for their correction based on the amount of redundancy present in the memory hierarchy. We investigate several new techniques for reducing multi-bit errors in large L2 caches, in which, the multi-bit errors are detected using simple error detection codes and corrected using the data redundancy in the memory hierarchy. We also propose several techniques to control/mine the redundancy in the memory hierarchy to further improve the reliability of the L2 cache. The proposed techniques were implemented in the Simplescalar framework and validated using the SPEC 2000 integer and floating point benchmarks for L2 cache vulnerability, global cache miss-rate, average cycle count and main memory write back rate, considering the area and power overheads. Experimental results indicate that the vulnerability of L2 caches can be decreased by 40% on the average for integer benchmarks and 32% on the average for floating point benchmarks, with an average multi-bit error coverage of about 96%, with significantly less area and power overheads and with virtually no performance penalty. The proposed techniques are applicable to both single and multi-core processor-based systems.
Koustav Bhattacharya, N. Ranganathan, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2007 Improving the reliability of on-chip L2 cache using redundancy
abstract
The reliability of large on-chip L2 cache poses a significant challenge due to technology scaling trends. As the minimum feature size continues to decrease, the L2 caches become more vulnerable to multi-bit soft errors. Traditionally, L2 caches have been protected from multi-bit soft errors using techniques like using error detection/correction codes or employing physical interleaving of cache bit lines to convert multi-bit errors into single-bit errors. These methods, however, incur large overheads in area and power. In this work, we investigate several new techniques for reducing multi-bit errors in large L2 caches, in which the multi-bit errors are detected using simple error detection codes and corrected using the data redundancy in the memory hierarchy. Further, we develop a reliability aware replacement policy that dynamically trades performance for reliability whenever the soft-error budget is exceeded. In order to further improve reliability, we propose the duplication of the data values in cache lines by exploiting their small data widths. The proposed techniques were implemented in the Simplescalar framework and validated using the SPEC 2000 integer and floating point benchmarks. The proposed techniques improve the reliability of L2 caches by 40% and 32% on the average, for integer and floating point applications respectively, with little impact on performance and area.
Koustav Bhattacharya, Soontae Kim, N. Ranganathan
ICCD2
2007 Reducing ALU and Register File Energy by Dynamic Zero Detection
abstract
Register files and ALU are dominant energy consumers in the datapath of typical pipelined processors. Reducing energy consumption of these components has a big impact on processor energy budget. Consequently, several techniques have been proposed at circuit and architectural levels. In contrast, we propose a dynamic zero detection technique for reducing energy in the register files and ALU in this paper. We find that zero is the most frequent value used in the register files and ALU. When a register value is determined to be zero before reading a register file, an access to the register file is prevented to save energy and zero is directly provided to the datapath. When one of operand register values of an add instruction is zero, it is prevented from being executed on ALU since its result is just the value of the other operand register. Since adds/subs constitute most of arithmetic and logic instructions, this optimization saves large ALU energy. Our dynamic zero detection technique is demonstrated to save 9.0% and 20.2% of energy in the register files and the adder of ALU for SPEC2000 floating-point and integer benchmarks, respectively.
Soontae Kim
IPCCC1
2006 Area-efficient error protection for caches
abstract
Due to increasing concern about various errors, current processors adopt error protection mechanisms. Especially, protecting L2/L3 caches incur as much as 12.5% area overhead due to error correcting codes. Considering large L2/L3 caches of current processors, the area overhead is very high. This paper proposes an area-efficient error protection scheme for L2/L3 caches. First, it selectively applies ECC (error correcting code) to only dirty cache lines and other clean cache lines are protected using simple parity check codes. Second, the dirty cache lines are periodically cleaned by exploiting the generational behavior of cache lines. Experimental results show that the cleaning technique effectively reduces the number of dirty cache lines per cycle. The ECCs of this reduced number of dirty cache lines can be maintained in a small storage. Our proposed scheme is shown to reduce the area overhead of a 1MB L2 cache for error protection by 59% for SPEC2000 benchmarks running on a typical four-issue superscalar processor
Soontae Kim
DATE1
2004 Scheduling Reusable Instructions for Power Reduction
abstract
In this paper, we propose a new issue queue design that is capable of scheduling reusable instructions. Once the issue queue is reusing instructions, no instruction cache access is needed since the instructions are supplied by the issue queue itself. Furthermore, dynamic branch prediction and instruction decoding can also be avoided permitting the gating of the front-end stages of the pipeline (the stages before register renaming). Results using array-intensive codes show that up to 82% of the total execution cycles, the pipeline front-end can be gated, providing a power reduction of 72% in the instruction cache, 33% in the branch predictor, and 21% in the issue queue, respectively, at a small performance cost. Our analysis of compiler optimizations indicates that the power savings can be further improved by using optimized code.
Jie S. Hu, Narayanan Vijaykrishnan, Soontae Kim, Mahmut T. Kandemir, Mary Jane Irwin
DATE3
2004 Data Organization and Retrieval on Parallel Air Channels: Performance and Energy Issues
J. Juran, Ali R. Hurson, Narayanan Vijaykrishnan, Soontae Kim
Wirel. Networks4
2003 Masking the Energy Behavior of DES Encryption
Hendra Saputra, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Richard R. Brooks, Soontae Kim, Wei Zhang 0002
DATE6
2003 On load latency in low-power caches
abstract
Many of the recently proposed techniques to reduce power consumption in caches introduce an additional level of non-determinism in cache access latency. Due to this additional latency, instructions speculatively issued and dependent on a non-deterministic load must be re-executed. Our exper-iments show that there is a large performance degradation and associated energy wastage due to these effects of in-struction re-execution. To address this problem, we propose an early cache set resolution scheme. It is based on the observation that the displacement values used for address generation are generally small. Our experimental evaluation shows that this technique is quite effective in mitigating this problem.
Soontae Kim, Narayanan Vijaykrishnan, Mary Jane Irwin, Lizy Kurian John
ISLPED1
2003 Partitioned instruction cache architecture for energy efficiency
abstract
The demand for high-performance architectures and powerful battery-operated mobile devices has accentuated the need for low-power systems. In many media and embedded applications, the memory system can consume more than 50% of the overall system energy, making it a ripe candidate for optimization. To address this increasingly important problem, this article studies energy-efficient cache architectures in the memory hierarchy that can have a significant impact on the overall system energy consumption.Existing cache optimization approaches have looked at partitioning the caches at the circuit level and enabling/disabling these cache partitions (subbanks) at the architectural level for both performance and energy. In contrast, this article focuses on partitioning the cache resources architecturally for energy and energy-delay optimizations. Specifically, we investigate ways of splitting the cache into several smaller units, each of which is a cache by itself (called a subcache ). Subcache architectures not only reduce the per-access energy costs, but can potentially improve the locality behavior as well.The proposed subcache architecture employs a page-based placement strategy, a dynamic page remapping policy, and a subcache prediction policy in order to improve the memory system energy behavior, especially on-chip cache energy. Using applications from the SPECjvm98 and SPEC CPU2000 benchmarks, the proposed subcache architecture is shown to be very effective in improving both the energy and energy-delay metrics. It is more beneficial in larger caches as well.
Soontae Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin
ACM Trans. Embed. Comput. Syst.1
2001 Power-aware partitioned cache architectures
abstract
This paper focuses on partitioning the cache resources architecturally for energy and energy-delay optimizations. Specifically, we investigate ways of splitting the cache into several smaller units, each of which is a cache by itself (called subcache). Subcache architectures not only reduce the peraccess energy costs but can potentially improve the locality behavior as well. We present a unified framework for designing, implementing and evaluating different subcache architectures. Different techniques for data placement, subcache prediction, and selective probing are proposed and evaluated using a diverse set of applications. The results show that intelligent subcache mechanisms proposed in this paper are effective.
Soontae Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin, E. Geethanjali
ISLPED1