Eui-Young Chung

dblp:36/3050 · DBLP profile ↗
← Back
49ranked-venue papers
6as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 43 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 DH-PIM: Maximizing Computing Unit Utilization in Digital PIM by Dual Half Mode Extension
abstract
Transformer-based Large Language Models (LLMs) rely on both General Matrix-Matrix Multiplication (GEMM) and General Matrix-Vector Multiplication (GEMV) for inference. While existing Processing-in-Memory (PIM) architectures, like HBM-PIM and Newton, efficiently handle GEMV operations, they suffer from low computing unit utilization during GEMM due to frequent input preparation and activation cycles. To address this, we propose Dual Half-mode PIM (DH-PIM), an extension of the Half Division Mode from previous PipePIM. DH-PIM splits sense amplifiers, and computing units into halves, allowing them to execute different operations independently. Dual half mode, which allows both halves of the computing unit to process data broadcast from either half of the sense amplifier, significantly enhances utilization via overlapped and pipelined execution. Our simulations show DH-PIM improves GEMM performance by up to 2.06x in GEMM kernels, with about 1.46x speedup in LLM tasks, compared to HBM-PIM. These performance improvements are achieved with minimal architectural modifications, resulting in low area and energy overhead.
Byoung Jin Kim, Taeyang Jeong, SeongJun Yun, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 PipePIM: Maximizing Computing Unit Utilization in ML-Oriented Digital PIM by Pipelining and Dual Buffering
abstract
A digital processing-in-memory (PIM) that integrates computing units (CUs) with DRAM banks emerges as a promising technique for accelerating matrix–vector multiplication (MV). However, activating and precharging all banks incur significant overheads in a digital PIM based on conventional DRAM, which is limited to activating only a single subarray in a bank. Moreover, a digital PIM utilizes a vector buffer to store and reuse the input vector. This necessitates repeated buffer writes, incurring substantial overhead for large MV. Consequently, these overheads reduce CU utilization in a digital PIM, degrading the performance. To overcome these issues, we propose PipePIM, which maximizes CU utilization in a digital PIM by pipelining and dual buffering. PipePIM consists of two primary schemes: 1) subarray-level pipelining (SAPI) and 2) dual vector buffer. They exploit and extend the features of a multitude of activated subarrays (MASA) introduced by subarray-level parallelism (SALP). SAPI enables a digital PIM to perform activation, precharging, and computation on different subarrays in a pipelined manner. Through SAPI, these operations are overlapped, and activation and precharging overheads are hidden. A dual vector buffer employs two vector buffers and manages them as ping-pong buffering, one for computation and another for buffer write simultaneously. To facilitate it, PipePIM proposes a half-division mode (HDM) enabling independent access to two activated subarrays with marginal area increase. We demonstrate the improvements by PipePIM on the state-of-the-art digital PIMs, Newton and HBM-PIM. Our simulation results indicate that the average speedups of Newton and HBM-PIM on MV are$2.16\times $and$1.74\times $, respectively.
Taeyang Jeong, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Operand-Oriented Virtual Memory Support for Near-Memory Processing
abstract
Virtual memory support is one of the major challenges of near-memory processing (NMP). Many previous works focused on this issue, but there are practical limitations that conventional CPU hardware or memory allocation schemes should be modified. Another technique uses a specialized page table for NMP to avoid such limitations. However, the previous work proposed NMP-specific page table that has static page table walk latency regardless of data size. This causes unnecessarily long address translation time for relatively small data. In this paper, we propose an operand-oriented technique for virtual memory support. Our scheme does not pre-determine the size of shared space; rather, it allocates shared space depending on the size of operands data for NMP. Then, we significantly reduce page table walk latency by using our flexible page table, which adapts the page table hierarchy to the size of shared spaces. To prove our concept, we implement our scheme in a full-system simulator and an FPGA-based verification platform. We then compared it with CPU's page table and the previous NMP-specific page table. The experimental results show that our technique outperforms page table walk latency by 69.3 percent and 43.8 percent compared to the CPU's page table and the comparison, respectively.
Duheon Choi, Taeyang Jeong, Joonhyeok Yeom, Eui-Young Chung
IEEE Trans. Computers4
2023 Unified Wear-Leveling Technique for NVM-Based Buffer of SSD
abstract
Since the gradual emergence of DRAM cost issues from the perspective of solid-state drive (SSD) system design, studies to replace DRAM buffer with a nonvolatile memory (NVM) have been conducted. However, the inferior durability of NVM requires additional wear-leveling algorithms for the buffer. Related works on the wear-leveling of the SSD internal buffer are significant only for specific applications or restricted partitions separated by user data and metadata. This article proposes a novel wear-leveling technique that covers the entire buffer without being constrained by specific applications. We introduce metadata characteristic-aware allocation (MCAA), which utilizes an appropriate hot data detection scheme depending on the type of update pattern on metadata. We also propose unified wear-leveling (UWL) that hierarchically integrates the wear-leveling techniques of each partition, including MCAA. The proposed method improves the maximum write count by 39.3% and the standard deviation of the wear-out by 25.9% on average compared to the conventional technique.
Young Min Park, Joonhyeok Yeom, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Page Type-Aware Data Migration Technique for Read Disturb Management of NAND Flash Memory
abstract
Read disturb is a phenomenon in which the cells of all unselected wordlines (WLs) within a block are soft programed while reading a selected WL, which causes errors. To solve the problem of read disturb, read reclaim (RR) operations, which migrate deteriorated block data to a new free block, are executed when the read count of the deteriorated block exceeds a preset threshold. RR requires additional operations, such as error correction, reprogramming, block erasure, and occupying I/O. So as the number of RR operations increases, the performance and reliability of the memory degrade. We observe that read disturb affects pages within a block differently according to page types. With this perspective, we propose Page Type-aware Data Migration technique (PTDM) technique, which migrates pages within blocks according to the vulnerability to read disturb. Shifting RR from the conventional block level to the page level improves the efficiency of the memory by reducing data migration. Also, we propose distributing read-hot data into several blocks during migration to reduce read disturb impact. Our experimental results show that the proposed method can reduce the amount of RR data migration and the overall erasure counts by 39.53% and 21.69%, respectively.
Sangwoo Han, Minjung Cho, Gi Lee, Eui-Young Chung
IEEE Trans. Very Large Scale Integr. Syst.4
2022 STT-MRAM-Based Multicontext FPGA for Multithreading Computing Environment
abstract
The demand for high-performance computing and rapidly increasing power consumption has increased the necessity for application-specific accelerators. In the datacenter and mobile system, more applications are increasingly relying on accelerators. Field-programmable gate arrays (FPGAs) emerge as a good candidate because they have high programmability and power efficiency. As the number of applications requiring acceleration increases, there is huge demand for FPGAs that support multiple contexts. Previous FPGA designs that support multicontext have various shortcomings such as volatility, poor power efficiency, large performance, area, and reconfiguration overhead. In this article, we propose a spin-transfer torque magnetic RAM (STT-MRAM)-based nonvolatile multicontext FPGA (NVMC-FPGA) that overcomes these shortcomings. We introduce the NVMC-FPGA architecture and operation modes that take advantage of nonvolatility and support multicontext. We also develop the multicontext-aware FPGA computer aided design flow to make the most of the NVMC-FPGA. Compared to the conventional SRAM-based FPGA, when eight identical circuits are mapped, the NVMC-FPGA improves the performance by 15.3% on average and reduces the power consumption by 11.2%–80.7%, depending on the number of simultaneously activated circuits. Moreover, when eight different circuits are mapped, the NVMC-FPGA improves the performance by 58.5% on average and reduces the power consumption by 6.2%–63.3%, depending on the number of simultaneously activated circuits.
Jeongbin Kim 0001, Yongwoon Song, Kyungseon Cho, Hyuk-Jun Lee, Hongil Yoon, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache
abstract
Recently, a hybrid cache consisting of SRAM and STT-RAM has attracted much attention as a future memory by complementing each other with different memory characteristics. Prior works focused on developing data allocation and migration techniques considering write-intensity to reduce write energy at STT-RAM. However, these works often neglect the impact of operation-specific reusability of a cache line. In this paper, we propose an energy-efficient per-operation reusability-based allocation and migration policy (ORAM) with a unified LRU replacement policy. First, to select an adequate memory type for allocation, we propose a cost function based on per-operation reusability - gain from an allocated cache line and loss from an evicted cache line for different memory types - which exploits the temporal locality. Besides, we present a migration policy, victim and target cache line selection scheme, to resolve memory type inconsistency between replacement policy and the allocation policy, with further energy reduction. Experiment results show an average energy reduction in the LLC and the main memory by 12.3 and 21.2 percent, and the improvement of latency and execution time by 21.2 and 8.8 percent, respectively, compared with a baseline hybrid cache management. In addition, the Energy-Delay Product (EDP) is improved by 36.9 percent over the baseline.
Minsik Oh, Kwangsu Kim, Duheon Choi, Hyuk-Jun Lee, Eui-Young Chung
IEEE Trans. Computers5
2020 Multitoken-Based Power Management for NAND Flash Storage Devices
abstract
NAND flash-based storage devices (NFSDs) have been widely employed in various systems, including cloud servers as well as mobile devices. The core component of NFSDs is NAND flash memory (NFM) which has several advantages over the conventional hard disk drives (HDDs). An NFSD typically adopts a bunch of NFMs which are operated in parallel for maximizing the I/O throughput. However, optimizing for performance may not be desirable from the power budget (PB) perspective. In other words, concurrent operations of NFMs often drain inordinate current, which leads to the violation of the PB allocated for a storage device. In this article, we propose a novel power management scheme which maximizes concurrent operations of NFMs under the given power constraint. The proposed method quantizes the given power constraint of an NFSD. A quantum also called token is the basic unit of power management. The proposed power management scheme allocates tokens to NFMs and only the NFMs having enough tokens can perform their operations. We call this method multitoken-based power management (MTPM). The critical issue of MTPM is a deadlock which is resolved with the key allocation scheme. Furthermore, we enhance MTPM to improve performance. The extended method called keyless MTPM (KMTPM) improves the overall performance by relaxing the key acquisition requirement and allowing subatomic operations. In the experimental results, we confirm that the proposed methods always meet the given power constraint. The proposed KMTPM improves throughput by 22.85% compared to state of the art technique. In addition, KMTPM only incurs 3.8% of performance overhead and 0.015% of area overhead.
Tae-Hee You, Sangwoo Han, Young Min Park, Hyuk-Jun Lee, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 A Locality-Aware Compression Scheme for Highly Reliable Embedded Systems
abstract
Dynamic random access memory (DRAM) reliability has become one of the critical issues in embedded systems, as DRAM process technology advances with the increase in bit error probability. Unfortunately, redundant error-correction code (ECC) chips cannot be applied to embedded systems since cores and DRAMs are tightly coupled without a dual in-line memory module (DIMM) slot to account for the form factor, cost, and limited pin count. Therefore, ECC parities are typically placed in the same physical array where the user and system data reside. This coexistence eventually deteriorates data locality, which could be the critical factor in DRAM performance degradation. To address this issue, we propose an ECC scheme called locality-aware compression (LoComp) which integrates a compression algorithm, DRAM data layout, and memory controller especially optimized for embedded systems. We focus on the locality of the dataset and its corresponding metadata, as well as spatial data locality in the design of DRAM data layout, which reduces the number of row activations. The major feature in a compression algorithm is adjusting the misalignment of data streams caused by the data packing in many embedded systems. Moreover, we specialize the memory controller to reduce DRAM access for ECC parities and compression flags. The core technologies for the memory controller are the adoption of a set of small caches for metadata and the support of partial write operation without changing the DRAM interface. LoComp+, an enhanced version of LoComp, further reduces DRAM access for metadata by placing the metadata close to the corresponding data. In the experiment, previous works increase the DRAM access time from 68% to over twice the value compared to ECC DIMM. Whereas, LoComp and LoComp+ show reduced performance degradation by 33% and 48%, respectively. In other words, LoComp and LoComp+ substantially improved performance from between 13% and 33% compared to previous embedded ECC schemes.
JuHyung Hong, Jeongbin Kim 0001, Sangwoo Han, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 A Novel NAND Flash Memory Architecture for Maximally Exploiting Plane-Level Parallelism
abstract
Solid-state drive (SSD) has become one of the most dominant storage devices and is rapidly replacing conventional storage devices. The core component of SSD is NAND flash memory (NFM), where the actual data are stored. Cost pressure is the most critical factor limiting the further deployment of SSDs and past researches have focused on developing cost-effective high-density NFM. Although the cost-driven technology development increases per-chip capacity, it reduces channel-/way-level parallelisms for the given device capacity, resulting in the performance degradation. Such observation directs us to focus on a novel NFM architecture exploiting plane-level parallelism. The distinct features of this architecture are: 1) enabling a decoupled word-line (WL) selection for the mated planes and 2) segmenting each plane into subplanes for further maximizing the plane-level parallelism. The experimental results show that decoupled WL selection improves the throughput by up to 21.3% with a marginal overhead of less than 1%, compared to the conventional NFM architecture. In addition, adopting the plane segmentation improves the throughput by up to 43.9% with an additional overhead of 14%. Considering the tradeoff between performance and overhead, the proposed NFM architecture is a cost-efficient method to secure high performance under decreasing channel-/way-level parallelisms in high-density NFM.
Myeongjin Kim, Wontaeck Jung, Hyuk-Jun Lee, Eui-Young Chung
IEEE Trans. Very Large Scale Integr. Syst.4
2018 EXTREME: Exploiting Page Table for Reducing Refresh Power of 3D-Stacked DRAM Memory
abstract
For future exascale computing systems, ultra-high-density memories would be required that consume low power to process massive data. Of the various memory devices, 3D-stacked DRAMs using TSVs are a perfect solution for this purposes. In addition to providing high capacity, these provide functional flexibility to the computing system by attaching a logic die in each 3D-stacked DRAM chip. However, the high capacity 3D-stacked DRAMs suffer from a significant loss of refresh power, which is solely required to maintain data. Although various schemes have been proposed to mitigate this issue, they cannot be adopted by commercial products due to compatibility issues. To tackle this issue, we propose EXTREME, which effectively reduces the refresh power of 3D-Stacked DRAMs. In order to retain the compatibility with OS, a simple page table manager is implemented at the logic die of 3D-stacked DRAM devices, which pins the page table to a specific memory space. The experiment results demonstrate that this reduces the refresh power at idle time by up to 98 percent with 16 KB of SRAM (static RAM) and 64 KB of DRAM register overhead for a 2 GB 3D-stacked DRAM device.
Ho Hyun Shin, Young Min Park, Duheon Choi, Byoung Jin Kim, Dae-Hyung Cho, Eui-Young Chung
IEEE Trans. Computers6
2018 Replacement Policy Adaptable Miss Curve Estimation for Efficient Cache Partitioning
abstract
Cache replacement policies and cache partitioning are well-known cache management techniques which aim to eliminate inter- and intra-application contention caused by co-running applications, respectively. Since replacement policies can change applications' behavior on a shared last-level cache, they have a massive impact on cache partitioning. Furthermore, cache partitioning determines the capacity allocated to each application affecting incorporated replacement policy. However, their interoperability has not been thoroughly explored. Since existing cache partitioning methods are tailored to specific replacement policies to reduce overheads for characterization of applications' behavior, they may lead to suboptimal partitioning results when incorporated with the up-to-date replacement policies. In cache partitioning, miss curve estimation is a key component to relax this restriction which can reflect the dependency between a replacement policy and cache partitioning on partitioning decision. To tackle this issue, we propose a replacement policy adaptable miss curve estimation (RME) which estimates dynamic workload patterns according to any arbitrary replacement policy and to given applications with low overhead. In addition, RME considers asymmetry of miss latency by miss type, thus the impact of miss curve on cache partitioning can be reflected more accurately. The experimental results support the efficiency of RME and show that RME-based cache partitioning cooperated with high-performance replacement policies can minimize both inter- and intra-application interference successfully.
Byunghoon Lee, Kwangsu Kim, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 ICS: Interrupt-Based Channel Sneaking for Maximally Exploiting Die-Level Parallelism of NAND Flash-Based Storage Devices
JuHyung Hong, Sangwoo Han, Young Min Park, Eui-Young Chung
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Timing window wiper: A new scheme for reducing refresh power of DRAM
abstract
DRAM refresh power, which is consumed solely to preserve data, is rapidly increasing as capacity increases. A study predicts that the power will account for up to 35% of total DRAM power in the high capacity DRAM device. While various schemes were proposed to reduce the refresh power, those could not be adopted to commercial products due to cost overhead and design complexity issues. In this paper, we propose a simple refresh power saving scheme called TWW. We implement it with a much smaller amount of register size than previous works without modification on OS and DRAM controller. It eliminates unnecessary refresh operation of pre-activated rows in a specific timing window with optimum register size. We can save the refresh power up to 16% with only 6.2 KB registers in a DRAM device. This paper explains the implementation of the proposed scheme and shows the power saving effectiveness with gem5 full system simulator.
Ho Hyun Shin, Hyeokjun Seo, Byunghoon Lee, Jeongbin Kim 0001, Eui-Young Chung
ASP-DAC5
2017 Scalable Bandwidth Shaping Scheme via Adaptively Managed Parallel Heaps in Manycore-Based Network Processors
abstract
Scalability of network processor-based routers heavily depends on limitations imposed by memory accesses and associated power consumption. Bandwidth shaping of a flow is a key function, which requires a token bucket per output queue and abuses memory bandwidth. As the number of output queues increases, managing token buckets becomes prohibitively expensive and limits scalability. In this work, we propose a scalable software-based token bucket management scheme that can reduce memory accesses and power consumption significantly. To satisfy real-time and low-cost constraints, we propose novel parallel heap data structures running on a manycore-based network processor. By using cache locking, the performance of heap processing is enhanced significantly and is more predictable. In addition, we quantitatively analyze the performance and memory footprint of the proposed software scheme using stochastic modeling and the Lyapunov central limit theorem. Finally, the proposed scheme provides an adaptive method to limit the size of heaps in the case of oversubscribed queues, which can successfully isolate the queues showing unideal behavior. The proposed scheme reduces memory accesses by up to three orders of magnitude for one million queues sharing a 100Gbps interface of the router while maintaining stability under stressful scenarios.
Jongbum Lim, Jinku Kim, Woo-Cheol Cho, Eui-Young Chung, Hyuk-Jun Lee
ACM Trans. Design Autom. Electr. Syst.5
2016 NAND Flash Memory With Multiple Page Sizes for High-Performance Storage Devices
abstract
In recent years, the demand for NAND flash-based storage devices has rapidly increased because of the popularization of various portable devices. NAND flash memory (NFM) offers many advantages, such as nonvolatility, high performance, the small form factor, and low-power consumption, while achieving high chip integration with a specialized architecture for bulk data access. A unit of NFM's read and program operations, the page, has continuously grown. Although increasing page size reduces costs, it adversely affects performance because of the resultant side effects, such as fragmentation and wasted space, caused by the incongruity of data and page sizes. To address this issue, we propose a multiple-page-size NFM architecture and its management. Our method dramatically improves write performance through adopting multiple page sizes without requiring additional area overhead or manufacturing processes. Based on the experimental results, the proposed NFM improves write latency and NFM lifetime by up to 65% and 62%, respectively, compared with the single-page-size NFM.
Sang-Hoon Park, Hyeokjun Seo, Ki-Whan Song, Sungroh Yoon, Eui-Young Chung
IEEE Trans. Very Large Scale Integr. Syst.6
2014 Exploiting Implementation Diversity and Partial Connection of Routers in Application-Specific Network-on-Chip Topology Synthesis
abstract
This paper proposes a novel application-specific Network-on-Chip (NoC) topology synthesis method, in which the partial connection and the implementation diversity of routers are exploited. NoC has emerged as a promising solution to future system-on-chip (SoC), and many researchers have focused on the automatic synthesis of NoC topology. In our observation, those NoC topology synthesis methods resemble the logic synthesis in the following sense: both the NoC topology synthesis and the logic synthesis determine the connections among the components where the components are the routers in the former and the logic cells in the latter. However, an outstanding difference is that the existing NoC topology synthesis methods consider only a single implementation for each size of router, whereas modern logic synthesis tools utilize multiple implementations of a cell to produce better netlist by the feature called technology mapping. To tackle this drawback, we propose a novel NoC topology synthesis methodology where the implementation diversity of routers is exploited to produce optimal topologies in terms of area and/or power consumption. Two different approaches, the post-process approach and the in-process approach, are proposed for exploiting the implementation diversity to provide flexibility between synthesis time and design quality. Also, the proposed method for characterizing and modeling routers makes it feasible to consider the implementation diversity even when the partial connection of routers is considered during the synthesis. Compared to the method in which the implementation diversity is exploited but the partial connection is not, the experimental results demonstrate that the proposed method can reduce the power consumption by up to 67.8% and 40.0% on average. On the other hand, compared to the method in which the partial connection is exploited but the implementation diversity is not, the power consumption is reduced by up to 12.0% and 3.4% on average.
Minje Jun, Won Woo Ro, Eui-Young Chung
IEEE Trans. Computers3
2014 $C\!\!-\!\!Lock$ : Energy Efficient Synchronization for Embedded Multicore Systems
abstract
Data synchronization among multiple cores has been one of the critical issues which must be resolved in order to optimize the parallelism of multicore architectures. Data synchronization schemes can be classified as lock-based methods (“pessimistic”) and lock-free methods (“optimistic”). However, none of these methods consider the nature of embedded systems which have demanding and sometimes conflicting requirements not only for high performance, but also for low power consumption. As an answer to these problems, we propose$C\!\!- \!\! Lock$, an energy- and performance-efficient data synchronization method for multicore embedded systems.$C\!\!- \!\! Lock$achieves balanced energy- and performance-efficiency by combining the advantages of lock-based methods and transactional memory (TM) approaches; in$C\!\!- \!\! Lock$, the core is blocked only when true conflicts exist (advantage of TM), while avoiding roll-back operations which can cause huge overhead with regard to both performance and energy (this is an advantage of locks). Also, in order to save more energy,$C\!\!- \!\! Lock$disables the clocks of the cores which are blocked for the access to the shared data until the shared data become available. We compared our$C\!\!- \!\! Lock$approach against traditional locks and transactional memory systems and found that$C\!\!- \!\! Lock$can reduce the energy-delay product by up to 1.94 times and 13.78 times compared to the baseline and TM, respectively.
Sang Hyong Lee, Minje Jun, Byunghoon Lee, Won Woo Ro, Eui-Young Chung, Jean-Luc Gaudiot
IEEE Trans. Computers6
2014 An Adaptive Idle-Time Exploiting Method for Low Latency NAND Flash-Based Storage Devices
abstract
The market share of NAND flash-based storage devices (NFSDs) has rapidly grown in recent years since many characteristics, such as non-volatility, low latency, and high reliability, meet the requirements for various types of storage devices. However, the unique characteristic of NAND flash memories (NFMs), erase-before-write, causes problems for NFSDs from a performance perspective. Specifically, performance degradation is incurred by extra operations that serve to hide the bad characteristics of NFMs. In order to resolve this problem, many attractive methods have been proposed. Various algorithms for flash translation layers (FTLs) are representative methods that provide space redundancy to NFSDs for better performance. However, the amount of space redundancy is limited by the capacity of NFMs and thus, space redundancy is still insufficient for improving the performance of NFSDs. Consequently, a new type of redundancy, termed temporal redundancy, has recently been introduced for NFSDs. More precisely, the idleness of NFSDs is exploited so as to precede extra operations for NFSDs while minimizing the overhead of extra operations. In this paper, we propose an adaptive time-out method based on the Hidden-Markov Model (HMM) to efficiently utilize idle periods. In addition, we also suggest a simple scheduling scheme for extra operations that can be customized for general FTLs. The experimental results demonstrate that the proposed method yields performance improvements in terms of average write latency and peak latency, 74% and 76% better than the existing method, respectively, and approaching within average 9% and 5% of the optimal case, respectively.
Sang-Hoon Park, Donggun Kim 0005, Kwanhu Bang, Hyuk-Jun Lee, Sungjoo Yoo, Eui-Young Chung
IEEE Trans. Computers6
2012 Asymmetric DRAM synthesis for heterogeneous chip multiprocessors in 3D-stacked architecture
abstract
Various computational requirements of real-world applications have leveraged moving to heterogeneous chip multiprocessors (CMPs) from homogeneous ones. In the meantime, three-dimensional integration of DRAMs and processors using Through Silicon Vias (TSVs) has emerged as the most viable solution for breaking the memory wall in CMP environment by bringing much higher memory bandwidth compared to current PCB level processor-DRAM integration. However, most researches on 3D-stacked DRAM have focused on increasing the memory bandwidth to improve the overall throughput of a system, even though the memory access requirements of real-world applications are various just as the computational requirements. To tackle this problem, we propose an asymmetric 3D-stacked DRAM architecture where the DRAM die is divided into multiple segments and the segments are optimized for different memory requirements. Also, since the optimal architecture of the DRAM can be different for different heterogeneous CMPs, we propose an automatic synthesis method for the asymmetric 3D-stacked DRAM architecture. The experimental results show that the area-power-product is reduced by 65.1% on average compared to the conventional architectures for the four realistic benchmarks and many of their derivatives.
Minje Jun, Myoung-Jin Kim, Eui-Young Chung
ICCAD3
2012 Partial Connection-Aware Topology Synthesis for On-Chip Cascaded Crossbar Network
abstract
The crossbar (also called bus matrix) solution is known as one of the most effective communication architectures for modern high-performance embedded systems. To make it even more effective, several topology synthesis methods have been proposed. They mostly generate a crossbar network in a cascaded fashion under the assumption that each crossbar switch is fully connected (i.e., each input has a connection to every output). This assumption often limits optimizing the area efficiency and/or performance of the network due to the unnecessary connections inside the crossbar switches. Some existing methods marginally improve their synthesis results by eliminating the unnecessary connections after the synthesis step. Such postprocessing approaches make sense since considering partially connected crossbar switches earlier in the synthesis flow can greatly increase the optimal topology search space, thereby increasing the runtime. However, the result from these postprocessing techniques is typically far inferior to that from the exhaustive search. In this work, we tackle such limitations of previous methods by introducing a heuristic method based on iterative switch merging. To the best of authors' knowledge, none of previous methods consider the partial connection of crossbar switches in the middle of the topology synthesis. Our experimental results prove the effectiveness of the proposed method by showing up to 30.35 percent of area saving against those methods that consider the partial connection only in a postprocess. The results also show the superiority of the proposed method against the existing topology synthesis methods, showing up to 49.09 percent area saving and synthesis time reduction by several orders of magnitude.
Minje Jun, Deumji Woo, Eui-Young Chung
IEEE Trans. Computers3
2012 Application-Support Particle Filter for Dynamic Voltage Scaling of Multimedia Applications
abstract
Dynamic Voltage and Frequency Scaling (DVFS) is an effective low-power technique for real-time workloads. Its effectiveness critically depends on the accurate prediction of the task execution time. Many DVFS approaches have been proposed, but they are insufficient for highly nonstationary workloads. Several recent DVFS techniques adopted adaptive filters to improve accuracy. However, their improvement was rather limited, since they mainly focused on applying a filter framework to the target application without tuning it. We address this issue by proposing Particle Filter (PF)-based video decoders (MPEG2 and H.264) which exploit application-specific characteristics. More specifically, our PF-based video decoders utilize the size of each frame for the prediction of its decoding time. Compared to previous work, the PF is more suitable for our purpose, since it achieves higher prediction accuracy, even for highly nonstationary workloads such as H.264 clips. Our results show that the energy saved by the proposed approach is comparable to that of the ideal policy called oracle-DVFS, while the existing methods we tested were far inferior to oracle-DVFS in terms of H.264 video decoding. Additionally, when our method was used, only 0.40 and 6.88 percent of the frames missed their deadlines with negligible computational overhead for MPEG and H.264, respectively.
Jae-Beom Lee, Myoung-Jin Kim, Sungroh Yoon, Eui-Young Chung
IEEE Trans. Computers4
2010 Exploiting multiple switch libraries in topology synthesis of on-chip interconnection network
abstract
On-chip interconnection network is a crucial design component in high-performance System-on-Chips (SoCs). Many of previous works have focused on the automation of its topology design, since the topology largely determines its overall performance. For this purpose, they mostly require a switch library which includes all possible switch configurations (e.g. the number of in/output ports and data width) with their implementation costs such as delay, area, and power. More precisely, they characterize the switches by synthesizing them with a common design objective (e.g. minimizing area) and common design constraints for a given gate-level design library. The implementation costs are used in evaluating the topologies throughout the topology synthesis. The major drawback of single switch library approach is that it forces the topology synthesis methods to search the best topology with the assumption that all the switches comprising a topology will be implemented (synthesized) with a common design objective and common design constraints. Such assumption prevents them from exploring diverse combinations of the switches for a topology from the implementation perspective. To tackle this issue, we propose a topology synthesis method with multiple switch libraries, where the switch libraries are prepared with different design objectives and design constraints. The experimental results show that the power consumption and the area of optimal topologies can be saved by up to 67.1% and 27.2%, respectively, by the proposed method with negligible synthesis time overhead.
Minje Jun, Sungroh Yoon, Eui-Young Chung
DATE3
2010 Design of On-Chip Crossbar Network Topology Using Chained Edge Partitioning
abstract
This paper proposes an efficient topology synthesis method for on-chip interconnection network based on crossbar switches. The efficiency of topology synthesis methods is often measured by two metrics—the quality of the synthesized topology and synthesis time. These two metrics are critically determined by the definition of the topology design space and the exploration method. Furthermore, an efficient representing method for the design space is required to tightly link the design space and the exploration method. Even though topology synthesis methods have actively been researched, most of the previous methods were not deep in thought for these factors. Unlike the previous methods, we propose a topology synthesis method with a careful consideration of these factors. Our method efficiently defines the design space by a technique called chained edge partitioning, in conjunction with a representing method for the points in the space, called enhanced restricted growth function. We also provide an exploration method which well incorporates with the aforementioned search space. To prove the effectiveness of our method, we compared our method with previous methods. The experimental results show that our method outperforms the compared methods by up to 49.8% and 104.6× in the quality of the synthesized topology and the synthesis time, respectively.
Minje Jun, Eui-Young Chung
Comput. J.2
2010 A fast and simple system performance emulator for enhanced solid state disks: a case study of long read operations
abstract
In this paper, we propose a fast and simple system emulator, called a system performance emulator (SPE), to evaluate long read operations. The SPE estimates how much system-wide performance is enhanced by using a faster solid state disk (SSD). By suspending a CPU for a certain time during direct memory access (DMA) transfer and subtracting this suspended time from the total DMA time, the SPE estimates the improvement in system performance expected from an enhanced SSD prior to its manufacture. We also examine the relation between storage performance and system performance using the SPE.
Do Yeun Kim, Chanik Park, Eui-Young Chung, Sung Woo Chung
J. Zhejiang Univ. Sci. C3
2010 Architecture Exploration of High-Performance PCs with a Solid-State Disk
abstract
As the cost per bit of NAND flash memory devices rapidly decreases, NAND-flash-based Solid-State Disks (SSDs) are replacing Hard Disk Drives (HDDs) used in a wide spectrum of consumer computing devices. Although typical SSDs can deliver higher performances than HDDs can, the full capabilities of SSDs are currently not exploited in most systems. This is because an SSD is interfaced with its host system using the architectures and interface protocols designed for HDDs, due to compatibility issues. Given the pace at which the stand-alone performance of SSDs improves, the performance loss of SSDs due to the legacy interface and system architecture will soon become intolerable. To address this issue, we propose several architectural choices to fully exploit the performance of SSDs used in consumer PC architectures. More specifically, we explore its interface scheme, and data transfer concurrency with the change of the conventional PC architecture if necessary. We evaluated the performance of the architectural choices by prototyping them with SystemC. The experimental results guide us how to trade off the performance enhancement and the change of the PC architecture. The performance improvement was maximized by 2.67 times when the PC architecture is changed to support a dual-port SSD connected to the North Bridge via the Double-Data Rate (DDR) interface in real trace environments.
Dong Kim, Kwanhu Bang, Seung-Hwan Ha, Sungroh Yoon, Eui-Young Chung
IEEE Trans. Computers5
2010 On the Thermal Attack in Instruction Caches
abstract
The instruction cache has been recognized as one of the least hot units in microprocessors, which leaves the instruction cache largely ignored in on-chip thermal management. Consequently, thermal sensors are not allocated near the instruction cache. However, malicious codes can exploit the deficiency in this empirical design and heat up fine-grain localized hotspots in the instruction cache, which might lead to physical damages. In this paper, we show how instruction caches can be thermally attacked by malicious codes and how simple techniques can be utilized to protect instruction caches from the thermal attack.
Joonho Kong, Johnsy K. John, Eui-Young Chung, Sung Woo Chung, Jie S. Hu
IEEE Trans. Dependable Secur. Comput.3
2010 Energy-Optimal Dynamic Thermal Management: Computation and Cooling Power Co-Optimization
abstract
Conventional dynamic thermal management (DTM) assumes that the thermal resistance of a heat-sink is a given constant determined at design time. However, the thermal resistance of a common forced-convection heat sink is inversely proportional to the flow rate of the air or coolant at the expense of the cooling power consumption. The die temperature of the silicon devices strongly affects its leakage power consumption and reliability, and it can be changed by adjusting the thermal resistance of the cooling devices. Different from conventional DTM which aims to avoid the thermal emergency, our proposed DTM regards the thermal resistance of a forced-convection heat sink as a control variable, and minimize the total power consumption both for computation and cooling. We control the cooling power consumption together with the microprocessor clock frequency and supply voltage, and track the energy-optimal die temperature. Consequently, we reduce a significant amount of the temperature-dependent leakage power consumption of the microprocessor while spending a bit higher cooling power than conventional DTM, and eventually consume less total power. Experimental results show the proposed DTM saves up to 8.2% of the total energy compared with a baseline DTM approach. Our proposed DTM also enhances the Failures in Time (FIT) up to 80% in terms of the electromigration lifetime reliability.
Donghwa Shin, Sung Woo Chung, Eui-Young Chung, Naehyuck Chang
IEEE Trans. Ind. Informatics3
2009 Energy-optimal dynamic thermal management for green computing
abstract
Existing thermal management systems for microprocessors assume that the thermal resistance of the heat-sink is constant and that the objective of the cooling system is simply to avoid thermal emergencies. But in fact the thermal resistance of the usual forced-convection heat-sink is inversely proportional to the fan speed, and a more rational objective is to minimize the total power consumption of both processor and cooling system. Our new method of dynamic thermal management uses both the fan speed and the voltage/frequency of the microprocessor as control variables. Experiments show that tracking the energy-optimal steady-state temperature can saves up to 17.6% of the overall energy, when compared with a conventional approach that merely avoids over-heating.
Donghwa Shin, Naehyuck Chang, Jinhang Choi, Sung Woo Chung, Eui-Young Chung
ICCAD6
2009 A robust peak detection method for RNA structure inference by high-throughput contact mapping
abstract
MOTIVATION: For high-throughput prediction of the helical arrangements of large RNA molecules, an innovative method termed multiplexed hydroxyl radical (*OH) cleavage analysis (MOHCA) has been proposed. A key step in this promising technique is to detect peaks accurately from noisy radioactivity profiles. Since manual peak finding is laborious and prone to error, an automated peak detection method to improve the accuracy and throughput of MOHCA is required. Existing methods were not applicable to MOHCA due to their high false positive rates. RESULTS: We developed a two-step computational method that can detect peaks from MOHCA profiles in a robust manner. The first step exploits an ensemble of linear and non-linear signal processing techniques to find true peak candidates. In the second step, a binary classifier trained with the characteristics of true and false peaks is used to eliminate false peaks out of the peak candidates. We tested the proposed approach with 2002 MOHCA cleavage profiles and obtained the median recall, precision and F-measure values of 0.917, 0.750 and 0.830, respectively. Compared with the alternatives considered, the proposed method was able to handle false peaks substantially better, thus resulting in 51.0-71.8% higher median values of precision and F-measure. AVAILABILITY: The software and supplementary data are available at http://dna.korea.ac.kr/pub/mohca.
Jinkyu Kim 0001, Seunghak Yu, Byonghyo Shim, Hanjoo Kim, Hyeyoung Min, Eui-Young Chung, Rhiju Das, Sungroh Yoon
Bioinform.6
2009 Run-Time Adaptive Workload Estimation for Dynamic Voltage Scaling
abstract
Dynamic voltage scaling (DVS) is a popular energy-saving technique for real-time tasks. The effectiveness of DVS critically depends on the accuracy of workload estimation, since DVS exploits the slack or the difference between the deadline and execution time. Many existing DVS techniques are profile based and simply utilize the worst-case or average execution time without estimation. Several recent approaches recognize the importance of workload estimation and adopt statistical estimation techniques. However, these approaches still require extensive profiling to extract reliable workload statistics and furthermore cannot effectively handle time-varying workloads. Feedback-control-based adaptive algorithms have been proposed to handle such nonstationary workloads, but their results are often too sensitive to parameter selection. To overcome these limitations of existing approaches, we propose a novel workload estimation technique for DVS. This technique is based on the Kalman filter and can estimate the processing time of workloads in a robust and accurate manner by adaptively calibrating estimation error by feedback. We tested the proposed method with workloads of various characteristics extracted from eight MPEG video clips. To thoroughly evaluate the performance of our approach, we used both a cycle-accurate simulator and an XScale-based test board. Our simulation result demonstrates that the proposed technique outperforms the compared alternatives with respect to the ability to meet given timing and Quality of Service constraints. Furthermore, we found that the accuracy of our approach is almost comparable to the oracle accuracy achievable only by offline analysis. Experimental results indicate that using our approach can reduce energy consumption by 57.5% on average, only with negligible deadline miss ratio (DMR) around 6.1%. Moreover, the average of computational overheads for the proposed technique is just 0.3%, which is the minimum value compared to other methods. More importantly, the DMR of our method is bounded by 11.7% in the worst case, while those of other methods are twice or more than ours.
Sung-Yong Bang, Kwanhu Bang, Sungroh Yoon, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 Topology Synthesis of Cascaded Crossbar Switches
abstract
Performance requirements of on-chip network increase as system-on-chips (SoCs) are becoming more and more complex. For high-performance applications, crossbar switch-based networks are replacing the traditional shared buses as the backbone networks in SoCs. In this paper, we tackle the topology design of on-chip networks with crossbar switches in a cascaded fashion. We also resolve the unacceptable complexity of our previous method based on mixed integer linear programming by a heuristic method. Experimental results show that the proposed method overcomes the frequency limitation of the single crossbar-based design, particularly when the wire delay effect is considered. The proposed heuristic method also achieves more area reduction (up to 69.5%) over the existing methods, and finds as good solutions as the exact method while the synthesis time is saved by orders of magnitude.
Minje Jun, Sungjoo Yoo, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 High-Speed Post-Layout Logic Simulation Using Quasi-Static Clock Event Evaluation
abstract
The post-layout gate-level simulation constitutes a critical design step for timing closure. The major drawback of traditional post-layout gate-level simulation is its long analysis time, which becomes exacerbated as design complexity increases. An alternative method isstatictiminganalysis(STA), which can drastically reduce analysis time. However, STA sacrifices accuracy for speed and often produces unrealistic results such as false paths and overly pessimistic estimates. In this paper, we propose a hybrid analysis method that can significantly reduce analysis time, while preserving accuracy, with respect to the traditional gate-level simulation. Our key idea is that a large speedup would be possible by removing those events that are repetitious and unnecessary for simulation. In particular, we focus on reducing the number of clock-related events, which account for a major portion of all the events handled by a simulator. We tested the proposed method extensively with various benchmark circuits as well as industrial designs. Our experimental results exhibit that the proposed approach accelerates the total simulation speed by two times on average, yet maintaining the accuracy acquired by the traditional gate-level simulation.
Myeongjin Kim, Eui-Young Chung, Sungroh Yoon
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 Mixed integer linear programming-based optimal topology synthesis of cascaded crossbar switches
abstract
We present a topology synthesis method for high performance System-on-Chip (SoC) design. Our method provides an optimal topology of on-chip communication network for the given bandwidth, latency, frequency and/or area constraints. The optimal topology consists of multiple crossbar switches and some of them can be connected in a cascaded fashion for higher clock frequency and/or area efficiency. Compared to previous works, the major contribution of our work is the exactness of the solution from two aspects. First, the solving method of our work is exact by employing the mixed integer linear programming (MILP) method. Second, we generalize the crossbar switch representation in MILP in order that the optimal topology can include any arbitrary sizes of crossbar switches together. The experimental results show that the topologies optimized for the clock frequency (area) give up to 37.3% (12.7%) improvements compared to the conventional single large crossbar switch networks for two industrial strength SoC designs.
Minje Jun, Sungjoo Yoo, Eui-Young Chung
ASP-DAC3
2008 Energy and Performance Optimization of Demand Paging With OneNAND Flash
abstract
New fusion memory devices consisting of multiple heterogeneous memory components in a single die or package offer efficient ways to optimize embedded systems in terms of energy, performance, and cost. Samsung Electronics recently announced the OneNAND fusion memory, in which a NAND flash array is integrated with dual SRAM buffers to provide a nor-type I/O interface. OneNAND has the low cost and large capacity of a NAND flash but also permits eXecution-in-Place (XIP) like a nor flash. The deployment of such devices requires careful system-level resource management because of their impact on energy consumption and performance, and existing memory optimization techniques, such as the demand paging used with NAND flash, may no longer be appropriate for systems with a fusion memory. We introduce a new online demand paging scheme that fully exploits the XIP capability of OneNAND flash by classifying pages as load preferred (residing in the on-chip SRAM) and XIP preferred (accessed directly from the OneNAND flash and discarded after use). This achieves, on average, a 26% reduction in energy consumption and a 19% increase in performance, compared with conventional NAND flash demand paging.
Yongsoo Joo, Yongseok Choi, Jaehyun Park 0005, Chanik Park, Sung Woo Chung, Eui-Young Chung, Naehyuck Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2008 Scalable QoS-Aware Memory Controller for High-Bandwidth Packet Memory
abstract
This paper proposes a high-performance scalable quality-of-service (QoS)-aware memory controller for the packet memory where packet data are stored in network routers. A major challenge in the packet memory controller design is to make the design scalable. As the input and output bandwidth requirement and the number of output queues for routers increase, the memory system becomes a bottleneck that limits the performance and scalability. Existing schemes require an input and output buffer that store packet data temporarily before they are written into or read from the memory. With the buffer size proportional to the number of output queues, the buffer becomes a limiting factor for scalability. Our scheme consists of a hashing logic and a reorder buffer whose size is not proportional to the number of output queues and is scalable with the increasing number of output queues. Another major challenge in the packet memory controller design is supporting QoS. As an increasing number of Internet packets become latency sensitive, it is critical that the memory controller is capable of providing different QoS to packets belonging to different classes. To the best of our knowledge, no published work on the packet memory controller supports QoS. In this paper, we show our scheme reduces the SRAM buffer size of the existing schemes by an order of magnitude whereas guaranteeing a packet loss probability as low as 10-20. Our QoS-aware scheduler shows that it meets the latency requirements assigned to multiple service classes under dynamically changing input loads for multiple classes using a feedback control loop.
Hyuk-Jun Lee, Eui-Young Chung
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Slack-based Bus Arbitration Scheme for Soft Real-time Constrained Embedded Systems
abstract
We present a bus arbitration scheme for soft real-time constrained embedded systems. Some masters in such systems are required to complete their work for given timing constraints, resulting in the satisfaction of system-level timing constraints. The computation time of each master is predictable, but it is not easy to predict its data transfer time since the communication architecture is mostly shared by several masters. Previous works solved this issue by minimizing the latencies of several latency-critical masters, but the side effect of these methods is that it can increase the latencies of other masters, hence they may violate the given timing constraints. Unlike previous works, our method uses the concept of "slack" in order to make the latency as close as its given constraint, resulting in the reduction of the side effect. The proposed arbitration scheme consists of bandwidth-conscious arbiter and scheduler. The arbiter can be any existing bandwidth-conscious arbiter and the scheduler implements the latency-awareness proposed in this paper. The scheduler is involved in the arbitration only when it observes a request whose slack is not sufficient for the given timing constraint. The experimental results show that our method outperforms the conventional round-robin arbiter by more than 100% in the best case in terms of the longest violated cycles.
Minje Jun, Kwanhu Bang, Hyuk-Jun Lee, Naehyuck Chang, Eui-Young Chung
ASP-DAC5
2007 Clustering protein environments for function prediction: finding PROSITE motifs in 3D
abstract
BACKGROUND: Structural genomics initiatives are producing increasing numbers of three-dimensional (3D) structures for which there is little functional information. Structure-based annotation of molecular function is therefore becoming critical. We previously presented FEATURE, a method for describing microenvironments around functional sites in proteins. However, FEATURE uses supervised machine learning and so is limited to building models for sites of known importance and location. We hypothesized that there are a large number of sites in proteins that are associated with function that have not yet been recognized. Toward that end, we have developed a method for clustering protein microenvironments in order to evaluate the potential for discovering novel sites that have not been previously identified. RESULTS: We have prototyped a computational method for rapid clustering of millions of microenvironments in order to discover residues whose surrounding environments are similar and which may therefore share a functional or structural role. We clustered nearly 2,000,000 environments from 9,600 protein chains and defined 4,550 clusters. As a preliminary validation, we asked whether known 3D environments associated with PROSITE motifs were "rediscovered". We found examples of clusters highly enriched for residues that share PROSITE sequence motifs. CONCLUSION: Our results demonstrate that we can cluster protein environments successfully using a simplified representation and K-means clustering algorithm. The rediscovery of known 3D motifs allows us to calibrate the size and intercluster distances that characterize useful clusters. This information will then allow us to find new clusters with similar characteristics that represent novel structural or functional sites.
Sungroh Yoon, Jessica C. Ebert, Eui-Young Chung, Giovanni De Micheli, Russ B. Altman
BMC Bioinform.3
2006 PowerViP: Soc power estimation framework at transaction level
abstract
In this work, we propose a SoC power estimation framework built on our system-level simulation environment. Our framework provides designers with the system-level power profile in a cycle-accurate manner. We target the framework to run fast and accurately, which is enabled by adopting different modeling techniques depending on the power characteristics of various IP blocks. The framework can be applied to any target SoC design
Ikhwan Lee, Sungjoo Yoo, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
ASP-DAC5
2006 Worst case execution time analysis for synthesized hardware
abstract
We propose a hardware performance estimation flow for fast design space exploration, based on worst-case execution time analysis algorithms for software analysis. Test cases on some real-world applications show that our flow provides a fight upper bound of the execution time, and many useful hints to the designer.
Jun-hee Yoo, Xingguang Feng, Kiyoung Choi, Eui-Young Chung, Kyu-Myung Choi
ASP-DAC4
2005 Fast and Accurate Transaction Level Modeling of an Extended AMBA2.0 Bus Architecture
abstract
A transaction level modeling (TLM) approach is used to meet the simulation speed as well as cycle accuracy for large scale SoC performance analysis. We implemented the transaction-level model of a proprietary bus called AHB+ which supports an extended AMBA2.0 protocol. The AHB+ transaction-level model is shown to be 353 times faster than the pin-accurate RTL model, while maintaining 97% accuracy on average. We also present the TLM development procedure of a bus architecture.
Young-Taek Kim, Youngduk Kim, Chulho Shin, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
DATE5
2004 Fast Exploration of Parameterized Bus Architecture for Communication-Centric SoC Design
abstract
For successful SoC design, efficient and scalable communication architecture is crucial. Some bus interconnects now provide configurable structures to meet this requirement of an SoC design. Furthermore, bus IP vendors provide software tools that automatically generate RTL codes of a bus once its designer configures it. Configurability, however, imposes more challenges upon designers because complexity involved in optimization increases exponentially as the number of parameters grows. In this paper, we present a novel approach with which effort requirement can be dramatically reduced. An automated optimization tool we developed is used and it exploits a genetic algorithm for fast design exploration. This paper shows that the time for the optimizing task can be reduced by more than 90% when the tool is used and, more significantly the task can be done without an expert's hand while ending up with a better solution.
Chulho Shin, Young-Taek Kim, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
DATE3
2002 Contents provider-assisted dynamic voltage scaling for low energy multimedia applications
abstract
This paper presents a new concept of DVS (Dynamic Voltage Scaling) for multimedia applications. Many multimedia applications have a periodic property, but each period shows a large variation in terms of its execution time. Exact estimation of such variation is a crucial factor for low energy software execution with DVS technique. Previous DVS techniques focused only on end users (client sites) and their quality heavily depends on the accurateness of the worst case execution time estimation. This paper proposes that contents providers (server sites) supply the information of the execution time variations in addition to the content itself. This makes it possible to perform DVS independent to worst case execution time estimation. The extra work required to the contents provider for this purpose is fully compensated by the benefits for the end users because single content is often provided to many users. Experimental results show that our method greatly reduces the energy consumption of client systems compared to previous DVS techniques.
Eui-Young Chung, Giovanni De Micheli, Luca Benini
ISLPED1
2002 Dynamic Power Management for Nonstationary Service Requests
abstract
Dynamic power management (DPM) is a design methodology aimed at reducing power consumption of electronic systems by performing selective shutdown of idle system resources. The effectiveness of a power management scheme depends critically on accurate modeling of service requests and on computation of the control policy. In this work, we present an online adaptive DPM scheme for systems that can be modeled as finite-state Markov chains. Online adaptation is required to deal with initially unknown or nonstationary workloads, which are very common in real-life systems. Our approach moves from exact policy optimization techniques in a known and stationary stochastic environment and extends optimum stationary control policies to handle the unknown and nonstationary stochastic environment for practical applications. We introduce two workload learning techniques based on sliding windows and study their properties. Furthermore, a two-dimensional interpolation technique is introduced to obtain adaptive policies from a precomputed look-up table of optimum stationary policies. The effectiveness of our approach is demonstrated by a complete DPM implementation on a laptop computer with a power-manageable hard disk that compares very favorably with existing DPM schemes.
Eui-Young Chung, Luca Benini, Alessandro Bogliolo, Yung-Hsiang Lu, Giovanni De Micheli
IEEE Trans. Computers1
2002 Value-sensitive automatic code specialization for embedded software
abstract
The objective of this work is to create a framework for the optimization of embedded software. We present algorithms and a tool flow to reduce the computational effort of programs, using value profiling and partial evaluation. Such a reduction translates into both energy savings and average-case performance improvement, while preserving a tolerable increase of worst case performance and code size. Our tool reduces the computational effort by specializing frequently executed procedures for the most common values of their parameters. The most effective specializations are automatically searched and identified, and the code is transformed through partial evaluation. Experimental results show that their technique improves both energy consumption and performance of the source code up to more than a factor of two, in average about 35% over the original program. Also, their automatic search engine greatly reduces code optimization time with respect to exhaustive search.
Eui-Young Chung, Luca Benini, Giovanni De Micheli, Gabriele Luculli, Marco Carilli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2001 Automatic source code specialization for energy reduction
abstract
This paper presents a framework to reduce the computational eort of software programs, using value pro ling and partial evaluation. Our tool reduces computational eort by specializing a program for highly expected situations and such a reduction translates into both energy and performance improvement. Procedure calls executed frequently with same parameter values are de ned as highly expected situations (common cases). The choice of the best transformation of common cases is achieved by solving three search problems. The rst identi es eective common cases to be specialized, the second searches for an optimal solution for eective common case, and the third examines the interplay among the specialized cases. Our technique improves both energy consumption and performance of the source code up to more than twice and in average about 25% over the original program. Also, our pruning techniques reduce the searching time by 80% compared to exhaustive approach.
Eui-Young Chung, Luca Benini, Giovanni De Micheli
ISLPED1
2000 Quantitative Comparison of Power Management Algorithms
abstract
Dynamic power management saves power by shutting down idle devices. Several management algorithms have been proposed and demonstrated to be effective in certain applications. We quantitatively compare the power saving and performance impact of these algorithms on hard disks of a desktop and notebook computers. This paper has three contributions. First, we build a framework in Windows NT to implement power managers running realistic workloads and directly interacting with users. Second, we define performance degradation that reflects user perception. Finally, we compare power saving and performance of existing algorithms and analyze the difference.
Yung-Hsiang Lu, Eui-Young Chung, Tajana Rosing, Giovanni De Micheli, Luca Benini
DATE2
1999 Dynamic Power Management for non-stationary service requests
abstract
Dynamic power management is a design methodology aiming at reducing power consumption of electronic systems, by performing selective shutdown of the idle system resources. The effectiveness of a power management scheme depends critically on an accurate modeling of the environment, and on the computation of the control policy. This paper presents two methods for characterizing nonstationary service requests by means of a prediction scheme based on sliding windows. Moreover; it describes how control policies for nonstationary models can be derived.
Eui-Young Chung, Luca Benini, Alessandro Bogliolo, Giovanni De Micheli
DATE1
1999 Dynamic power management using adaptive learning tree
abstract
Dynamic power management (DPM) is a technique to reduce the power consumption of electronic systems by selectively shutting down idle components. The quality of the shutdown control algorithm (the power management policy) mostly depends on knowledge of the user's behavior, which in many cases is initially unknown or non-stationary. For this reason, DPM policies should be capable of adapting to changes in user behavior. In this paper, we present a novel DPM scheme based on idle period clustering and adaptive learning trees. We also provide a design guide for applying our technique to components with multiple sleep states. Experimental results show that our technique outperforms other advanced DPM schemes as well as simple time-out policies. The proposed approach shows little deviation of efficiency for various workloads having different characteristics, while other policies show that their efficiency changes drastically depending on the trace data characteristics. Furthermore, experimental evidence indicates that our workload learning algorithm is stable and has fast convergence.
Eui-Young Chung, Luca Benini, Giovanni De Micheli
ICCAD1