EDBT 2026 Demo / reviewers in the wild / expert
Hyung Gyu Lee
dblp:10/2814
· DBLP profile ↗
33ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0003-2596-7130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Memory systems · 37% Energy-efficient computing · 30% Embedded and real-time systems · 10% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 100% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
energy harvesting |
0.6 | 2 | 2019 | REAP: Runtime Energy-Accuracy Optimization for Energy Harvesting IoT Devices · DAC 2019 Storage-Less and Converter-Less Photovoltaic Energy Harvesting With Maximum Power Point Tracking for Internet of Things · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.4 | 2 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 ECM: Effective Capacity Maximizer for high-performance compressed caching · HPCA 2013 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.4 | 1 | 2019 | REAP: Runtime Energy-Accuracy Optimization for Energy Harvesting IoT Devices · DAC 2019 |
Embedded and real-time systems › energy harvesting systems
energy harvesting embedded systems |
0.4 | 1 | 2019 | Tumbler: Energy Efficient Task Scheduling for Dual-Channel Solar-Powered Sensor Nodes · DAC 2019 |
Cloud and datacenter computing › job scheduling
reinforcement-learning-based scheduling |
0.4 | 1 | 2019 | Tumbler: Energy Efficient Task Scheduling for Dual-Channel Solar-Powered Sensor Nodes · DAC 2019 |
Internet of things and sensor networks › energy efficiency
iot energy management |
0.2 | 1 | 2016 | Storage-Less and Converter-Less Photovoltaic Energy Harvesting With Maximum Power Point Tracking for Internet of Things · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Energy-efficient computing › energy harvesting
solar energy harvesting |
0.2 | 1 | 2016 | Storage-Less and Converter-Less Photovoltaic Energy Harvesting With Maximum Power Point Tracking for Internet of Things · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Memory systems
cache |
0.2 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Memory systems › memory compression
cache compression |
0.2 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Memory systems › cache management
cache replacement |
0.2 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Memory systems
cache management |
0.2 | 1 | 2013 | ECM: Effective Capacity Maximizer for high-performance compressed caching · HPCA 2013 |
Interconnection networks and networks-on-chip
flow control |
0.2 | 1 | 2013 | TornadoNoC: A lightweight and scalable on-chip network architecture for the many-core era · ACM Trans. Archit. Code Optim. 2013 |
Interconnection networks and networks-on-chip
router architecture |
0.2 | 1 | 2013 | TornadoNoC: A lightweight and scalable on-chip network architecture for the many-core era · ACM Trans. Archit. Code Optim. 2013 |
Energy-efficient computing
power management |
0.1 | 2 | 2019 | REAP: Runtime Energy-Accuracy Optimization for Energy Harvesting IoT Devices · DAC 2019 Energy exploration and reduction of SDRAM memory systems · DAC 2002 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.1 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Electronic design automation › design space exploration
communication architecture exploration |
0.1 | 1 | 2006 | Design space exploration and prototyping for on-chip multimedia applications · DAC 2006 |
Electronic design automation
design space exploration |
0.1 | 1 | 2006 | Design space exploration and prototyping for on-chip multimedia applications · DAC 2006 |
Performance modeling and evaluation › simulation
cache simulation |
0.0 | 1 | 2013 | ECM: Effective Capacity Maximizer for high-performance compressed caching · HPCA 2013 |
Memory systems
DRAM |
0.0 | 1 | 2002 | Energy exploration and reduction of SDRAM memory systems · DAC 2002 |
Energy-efficient computing
energy characterization |
0.0 | 1 | 2002 | Energy exploration and reduction of SDRAM memory systems · DAC 2002 |
Memory systems › DRAM › DRAM architecture
SDRAM |
0.0 | 1 | 2002 | Energy exploration and reduction of SDRAM memory systems · DAC 2002 |
Energy-efficient computing
memory energy efficiency |
0.0 | 1 | 2002 | Energy exploration and reduction of SDRAM memory systems · DAC 2002 |
Methods — techniques the papers use, named apart from their topics
design-point switching · 0.8co-optimization · 0.8nonvolatile microprocessor · 0.5dynamic power management · 0.5solar energy prediction · 0.4reinforcement learning · 0.4trace-driven simulation · 0.2full-system simulation · 0.2size-aware insertion · 0.2dynamic threshold adjustment · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | REAP: Runtime Energy-Accuracy Optimization for Energy Harvesting IoT DevicesabstractThe use of wearable and mobile devices for health and activity monitoring is growing rapidly. These devices need to maximize their accuracy and active time under a tight energy budget imposed by battery and form-factor constraints. This paper considers energy harvesting devices that run on a limited energy budget to recognize user activities over a given period. We propose a technique to co-optimize the accuracy and active time by utilizing multiple design points with different energy-accuracy trade-offs. The proposed technique switches between these design points at runtime to maximize a generalized objective function under tight harvested energy budget constraints. We evaluate our approach experimentally using a custom hardware prototype and 14 user studies. It achieves 46% higher expected accuracy and 66% longer active time compared to the highest performance design point. Ganapati Bhat, Kunal Bagewadi, Hyung Gyu Lee, Ümit Y. Ogras |
DAC | 3 |
| 2019 | Tumbler: Energy Efficient Task Scheduling for Dual-Channel Solar-Powered Sensor NodesabstractEnergy harvesting technology has been popularly adopted in embedded systems. However, unstable energy source results in unsteady operation. In this paper, we devise a long-term energy efficient task scheduling targeting for solar-powered sensor nodes. The proposed method exploits a reinforcement learning with a solar energy prediction method to maximize the energy efficiency, which finally enhances the long-term quality of services (QoS) of the sensor nodes. Experimental results show that the proposed scheduling improves the energy efficiency by 6.0%, on average and achieves the better QoS level by 54.0%, compared with a state-of-the-art task scheduling algorithm. Hyung Gyu Lee, Yujuan Tan, Yu Wu 0016, Xianzhang Chen, Liang Liang 0002, Lei Qiao 0002, Duo Liu 0002 |
DAC | 2 |
| 2019 | An Ultra-Low Energy Human Activity Recognition Accelerator for Wearable Health ApplicationsabstractHuman activity recognition (HAR) has recently received significant attention due to its wide range of applications in health and activity monitoring. The nature of these applications requires mobile or wearable devices with limited battery capacity. User surveys show that charging requirement is one of the leading reasons for abandoning these devices. Hence, practical solutions must offer ultra-low power capabilities that enable operation on harvested energy. To address this need, we present the first fully integrated custom hardware accelerator (HAR engine) that consumes 22.4 μJ per operation using a commercial 65 nm technology. We present a complete solution that integrates all steps of HAR , i.e., reading the raw sensor data, generating features, and activity classification using a deep neural network (DNN). It achieves 95% accuracy in recognizing 8 common human activities while providing three orders of magnitude higher energy efficiency compared to existing solutions. Ganapati Bhat, Yigit Tuncel, Sizhe An, Hyung Gyu Lee, Ümit Y. Ogras |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2019 | A Task Failure Rate Aware Dual-Channel Solar Power System for Nonvolatile Sensor NodesabstractIn line with the rapid development of the Internet of Things (IoT), the maintenance of on-board batteries for a trillion sensor nodes has become prohibitive both in time and costs. Energy harvesting is a promising solution to this problem. However, conventional energy-harvesting systems with storage suffer from low efficiency because of conversion loss and storage leakage. Direct supply systems without energy buffer provide higher efficiency, but fail to satisfy quality of service (QoS) due to mismatches between input power and workloads. Recently, a novel dual-channel photovoltaic power system has paved the way to achieve both high energy efficiency and QoS guarantee. This article focuses on the design-time and run-time co-optimization of the dual-channel solar power system. At the design stage, we develop a task failure rate estimation framework to balance design costs and failure rate. At run-time, we propose a task failure rate aware QoS tuning algorithm to further enhance energy efficiency. Through the experiments on both a simulation platform and a prototype board, this study demonstrates a 27% task failure rate reduction compared with conventional architectures with identical design costs. And the proposed online QoS tuning algorithm brings up to 30% improvement in energy efficiency with nearly zero failure rate penalty. Fang Su, Yongpan Liu, Xiao Sheng, Hyung Gyu Lee, Naehyuck Chang, Huazhong Yang |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Puppet: Energy Efficient Task Mapping For Storage-Less and Converter-Less Solar-Powered Non-Volatile Sensor NodesabstractSolar powered sensor nodes have been adopted in many applications, but unstable energy source and high energy loss are hindrances to their wide spreading. Storage-less and converter-less solar powered non-volatile sensor nodes reduce the energy loss to a great extent. However, without energy buffers, sensor nodes become more sensitive to solar variations. Making full use of harvested energy to provide better quality of services (QoS) to guarantee stable operations under this circumstance is crucial. In this paper, we devise an energy efficient task mapping strategy for storage-less and converter-less solar powered non-volatile sensor nodes. The proposed strategy, Puppet uses a reinforcement learning to make nodes achieve higher energy utilization and finally enhance the QoS. Experimental results show that the proposed strategy reduces the deadline miss ratio (DMR) in Puppet by 22% while increases energy utilization and effective energy utilization by 11% and 25%, on the average, respectively. Hyung Gyu Lee, Xianzhang Chen, Duo Liu 0002, Liang Liang 0002 |
ICCD | 2 |
| 2017 | Flexible PV-cell Modeling for Energy Harvesting in Wearable IoT ApplicationsabstractWearable devices with sensing, processing and communication capabilities have become feasible with the advances in internet-of-things (IoT) and low power design technologies. Energy harvesting is extremely important for wearable IoT devices due to size and weight limitations of batteries. One of the most widely used energy harvesting sources is photovoltaic cell (PV-cell) owing to its simplicity and high output power. In particular, flexible PV-cells offer great potential for wearable applications. This paper models, for the first time , how bending a PV-cell significantly impacts the harvested energy. Furthermore, we derive an analytical model to quantify the harvested energy as a function of the radius of curvature. We validate the proposed model empirically using a commercial PV-cell under a wide range of bending scenarios, light intensities and elevation angles. Finally, we show that the proposed model can accelerate maximum power point tracking algorithms and increase the harvested energy by up to 25.0%. Jaehyun Park 0005, Hitesh Joshi, Hyung Gyu Lee, Sayfe Kiaei, Ümit Y. Ogras |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | HoPE: Hot-Cacheline Prediction for Dynamic Early Decompression in Compressed LLCsabstractData compression plays a pivotal role in improving system performance and reducing energy consumption, because it increases the logical effective capacity of a compressed memory system without physically increasing the memory size. However, data compression techniques incur some cost, such as non-negligible compression and decompression overhead. This overhead becomes more severe if compression is used in the cache. In this article, we aim to minimize the read-hit decompression penalty in compressed Last-Level Caches (LLCs) by speculatively decompressing frequently used cachelines. To this end, we propose a Hot-cacheline Prediction and Early decompression (HoPE) mechanism that consists of three synergistic techniques: Hot-cacheline Prediction (HP), Early Decompression (ED), and Hit-history-based Insertion (HBI). HP and HBI efficiently identify the hot compressed cachelines, while ED selectively decompresses hot cachelines, based on their size information. Unlike previous approaches, the HoPE framework considers the performance balance/tradeoff between the increased effective cache capacity and the decompression penalty. To evaluate the effectiveness of the proposed HoPE mechanism, we run extensive simulations on memory traces obtained from multi-threaded benchmarks running on a full-system simulation framework. We observe significant performance improvements over compressed cache schemes employing the conventional Least-Recently Used (LRU) replacement policy, the Dynamic Re-Reference Interval Prediction (DRRIP) scheme, and the Effective Capacity Maximizer (ECM) compressed cache management mechanism. Specifically, HoPE exhibits system performance improvements of approximately 11%, on average, over LRU, 8% over DRRIP, and 7% over ECM by reducing the read-hit decompression penalty by around 65%, over a wide range of applications. Jaehyun Park 0005, Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Vinson Young, Junghee Lee 0004, Jongman Kim |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Accurate personal ultraviolet dose estimation with multiple wearable sensorsabstractWearable devices begin to integrate into the daily lives along with recent technology development. One of such important applications is to accurately monitor ultraviolet (UV) radiation received by the human body. To compensate for the localized monitoring area of existing personal UV monitoring devices, this paper proposes a reconstruction method to estimate the UV dose over the entire body based on multiple discrete wearable UV sensor nodes. Ambient factors and individual factors are both considered in this paper. The proposed estimation method is validated by a range of UV data collection experiments in realistic scenarios. Experimental results show that the proposed method reduces 68.3% estimation errors on average compared with existing single sensor based methods. Jinyang Li 0002, Yongpan Liu, Hehe Li, Chun Jason Xue, Hyung Gyu Lee, Huazhong Yang |
BSN | 6 |
| 2016 | SATS: An Ultra-Low Power Time Synchronization for Solar Energy Harvesting WSNsabstractReliable and ultra-low power time synchronization becomes more and more important with the popularity of energy harvesting sensor nodes. This paper proposes an untethered and probabilistic ultra-lower power time synchronization method for energy intermittent sensor network. It avoids the frequent RF communications with the assistance of a solar clock. The SATS system consists of two main parts: the synchronizer, a low power solar clock module for time synchronization, and the S3-Mapping, an offline sequence matching algorithm. Furthermore, we develop an improved version of S3-Mapping, which reduces the computation complexity from exponential to linear using the redundancy models and the onion peeling method. The SATS system is validated by both simulations and a prototype, which shows that the second level synchronization precision can be achieved under reasonable probability. What's more, the energy consumption of time synchronization is reduced by over 1 ~ 2 magnitudes compared with the up-to-date low power time synchronization protocol. Tongda Wu, Yongpan Liu, Hehe Li, Chun Jason Xue, Hyung Gyu Lee, Huazhong Yang |
ISLPED | 5 |
| 2016 | Storage-Less and Converter-Less Photovoltaic Energy Harvesting With Maximum Power Point Tracking for Internet of ThingsabstractEnergy harvesting from natural environment gives range of benefits for the Internet of things. Scavenging energy from photovoltaic (PV) cells is one of the most practical solutions in terms of power density among existing energy harvesting sources. PV power systems mandate the maximum power point tracking (MPPT) to scavenge the maximum possible solar energy. In general, a switching-mode power converter, an MPPT charger, controls the charging current to the energy storage element (a battery or equivalent), and the energy storage element provides power to the load device. The mismatch between the maximum power point (MPP) current and the load current is managed by the energy storage element. However, such architecture causes significant energy loss (typically over 20%) and a significant weight/volume and a high cost due to the cascaded power converters and the energy storage element. This paper pioneers a converter-less PV power system with the MPPT that directly supplies power to the load without the power converters or the energy storage element. The proposed system uses a nonvolatile microprocessor to enable an extremely fine-grain dynamic power management in a few hundred microseconds. This makes it possible to match the load current with the MPP current. We present detailed modeling, simulation, and optimization of the proposed energy harvesting system including the radio frequency transceiver. Experiments show that the proposed setup achieves an 87.1% of overall system efficiency during a day, 30.6% higher than the conventional MPPT methods in actual measurements, and thus a significantly higher duty cycle under a weak solar irradiance. Yongpan Liu, Xiao Sheng, Hyung Gyu Lee, Naehyuck Chang, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2015 | Powering the IoT: Storage-less and converter-less energy harvestingabstractWide spread of Internet of Things (IoTs) still have huddles in cost and maintenance. Energy harvesting is a promising option to mitigate battery replacement, but the current energy harvesting methods still rely on batteries or equivalent and power converters for the maximum power point tracking (MPPT). Unfortunately, batteries are subject to wear and tear, which is a primary factor to prevent from being maintenance free. Power converters are expensive, heavy and lossy as well. In this paper, we introduce a novel energy harvesting and management technique to power the IoT, which does not require any long-term energy storages nor voltage converters unlike traditional energy harvesting systems. Extensive simulations and measurements from our prototype demonstrate that the proposed method harvests 8% more energy and extends the operation time of the device 60% more during a day. This paper also demonstrates a UV (ultraviolet) level meter for skin protect, named SmartPatch, using the proposed energy harvesting method. The proposed method is not limited to photovoltaic energy harvesting but applicable to most energy harvesting IoT power supplies that require impedance tracking. Hyung Gyu Lee, Naehyuck Chang |
ASP-DAC | 1 |
| 2015 | Prefetch-based dynamic row buffer management for LPDDR2-NVM devicesabstractLPDDR2-NVM has been announced as an industry standard to efficiently interface with non-volatile memory devices such as phase change memory (PCM). This standard interface has been adopted in most commercial PCM devices. In this paper, we devise a prefetch-based dynamic row buffer management that targets the LPDDR2-NVM devices for enhancing performance with almost negligible implementation overhead. Our extensive simulations with timing parameters from the industry's commercial PCM devices demonstrate that the proposed method enhances the performance of memory systems up to 11.3% when compared with the static optimum configuration with fairly low-cost overheads. Jaehyun Park 0005, Donghwa Shin, Hyung Gyu Lee |
VLSI-SoC | 3 |
| 2015 | Design space exploration of row buffer architecture for phase change memory with LPDDR2-NVM interfaceabstractPhase change memory (PCM) is an attractive candidate for the future memory, but it still has several limitations to overcome such as write latency and long-term endurance. A large body of literature has been dedicated to solving these problems. However, almost all of the previous studies did not consider an important practical aspect of the PCM - an interface. The LPDDR2-NVM standard interface recently introduced by JEDEC is widely adopted by the manufacturers of commercial PCM these days. The LPDDR2-NVM standard allows a more flexible use of row buffers compared to the conventional DRAM interface. In this paper, we explore the design space of row buffer architecture in the PCM with LPDDR2-NVM interface. The effect of row buffer architecture on memory performance is investigated in terms of unit size and number of RDBs, and its management policy. We use the timing parameters from industry prototype PCM and analyze the result from the perspective of Pareto's optimum. The experimental results show that a properly-designed row buffer architecture enhances system-level performance up to 44.2% even at the same cost. Jaehyun Park 0005, Donghwa Shin, Hyung Gyu Lee |
VLSI-SoC | 3 |
| 2015 | Size-Aware Cache Management for Compressed Cache ArchitecturesabstractA practical way to increase the effective capacity of a microprocessor's cache, without physically increasing the cache size, is to employ data compression. Last-Level Caches (LLC) are particularly amenable to such compression schemes, since the primary purpose of the LLC is to minimize the miss rate, i.e., it directly benefits from a larger logical capacity. In compressed LLCs, the cacheline size varies depending on the achieved compression ratio. Our observations indicate that this size information gives useful hints when managing the cache (e.g., when selecting a victim), which can lead to increased cache performance. However, there are currently no replacement policies tailored to compressed LLCs; existing techniques focus primarily on locality information. This article introduces the concept of size-aware cache management as a way to maximize the performance of compressed caches. Upon analyzing the benefits of considering size information in the management of compressed caches, we propose a novel mechanism-called Effective Capacity Maximizer (ECM)-to further enhance the performance and energy consumption of compressed LLCs. The proposed technique revolves around four fundamental principles: ECM Insertion (ECM-I), ECM Promotion (ECM-P), ECM Eviction Scheduling (ECM-ES), and ECM Replacement (ECM-R). Extensive simulations with memory traces from real applications running on a full-system simulator demonstrate significant improvements compared to compressed cache schemes employing conventional locality-aware cache replacement policies. Specifically, our ECM shows an average effective capacity increase of 18.4 percent over the Least-Recently Used (LRU) policy, and 23.9 percent over the Dynamic Re-Reference Interval Prediction (DRRIP) [1] scheme. This translates into average system performance improvements of 7.2 percent over LRU and 4.2 percent over DRRIP. Moreover, the average energy consumption is also reduced by 5.9 percent over LRU and 3.8 percent over DRRIP. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Junghee Lee 0004, Jongman Kim |
IEEE Trans. Computers | 2 |
| 2014 | Storage-less and converter-less maximum power point tracking of photovoltaic cells for a nonvolatile microprocessorabstractThis paper pioneers the maximum power point tracking (MPPT) of photovoltaic (PV) cells that directly supply power to a microprocessor without an energy storage element (a battery or a large-size capacitor) nor power converters. The maximum power point tracking is conventionally performed by an MPPT charger that stores in the energy storage element, and a voltage regulator (typically a DC-DC converter) produces a proper voltage level for the microprocessor. The energy storage element is an energy buffer and makes it possible to perform MPPT of the PV cells and power management of the microprocessor independently. However, the energy storage element, MPPT charger and DC-DC converter cause seriously limited lifetime (when a typical battery is adopted), significant energy loss (typically over 20%), increased weight/volume and high cost, etc. The proposed method enables extremely fine-grain dynamic power management (DPM) in every a few hundred microseconds and performs the MPPT without using an MPPT charger and a DC-DC converter as well as an energy storage element. We achieve 84.5% of energy harvesting efficiency using the proposed setup with huge reduction in cost, weight and volume, and extended lifetime, which is not even numerically comparable with conventional MPPT methods. Naehyuck Chang, Younghyun Kim 0001, Sangyoung Park, Yongpan Liu, Hyung Gyu Lee, Huazhong Yang |
ASP-DAC | 6 |
| 2014 | Designing Hybrid DRAM/PCM Main Memory Systems Utilizing Dual-Phase CompressionabstractThe last few years have witnessed the emergence of a promising new memory technology, namely Phase-Change Memory (PCM). Due to its inherent ability to scale deeply into the nanoscale regime and its low power consumption, PCM is increasingly viewed as an attractive alternative for the memory subsystem of future microprocessor architectures. However, PCM is marred by a duo of potentially show-stopping deficiencies, that is, poor write performance (especially when compared to the prevalent and ubiquitous DRAM technology) and limited durability. These weaknesses have urged designers to develop various supporting architectural techniques to aid and complement the operation of the PCM while mitigating its innate flaws. One promising such solution is the deployment of hybridized memory architectures that fuse DRAM and PCM, in order to combine the best attributes of each technology. In this article, we introduce a novel Dual-Phase Compression (DPC) scheme and its architectural design aimed at DRAM/PCM hybrids, which caters to the limitations of PCM technology while optimizing memory performance. The DPC technique is specifically optimized for PCM-based environments and is transparent to the operation of the remaining components of the memory subsystem. Furthermore, the proposed architecture is imbued with a multifaceted wear-leveling technique to enhance the durability and prolong the lifetime of the PCM. Extensive simulations with traces from real applications running on a full-system simulator demonstrate 20.4% performance improvement and 46.9% energy reduction, on average, as compared to a baseline DRAM/PCM hybrid implementation. Additionally, the multifaceted wear-leveling technique is shown to significantly prolong the lifetime of the PCM. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Jongman Kim |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2013 | ECM: Effective Capacity Maximizer for high-performance compressed cachingabstractCompressed Last-Level Cache (LLC) architectures have been proposed to enhance system performance by efficiently increasing the effective capacity of the cache, without physically increasing the cache size. In a compressed cache, the cacheline size varies depending on the achieved compression ratio. We observe that this size information gives a useful hint when selecting a victim, which can lead to increased cache performance. However, no replacement policy tailored to compressed LLCs has been investigated so far. This paper introduces the notion of size-aware compressed cache management as a way to maximize the performance of compressed caches. Toward this end, the Effective Capacity Maximizer (ECM) scheme is introduced, which targets compressed LLCs. The proposed mechanism revolves around three fundamental principles: Size-Aware Insertion (SAI), a Dynamically Adjustable Threshold Scheme (DATS), and Size-Aware Replacement (SAR). By adjusting the eviction criteria, based on the compressed data size, one may increase the effective cache capacity and minimize the miss penalty. Extensive simulations with memory traces from real applications running on a full-system simulator demonstrate significant improvements compared to compressed cache schemes employing the conventional Least-Recently Used (LRU) and Dynamic Re-Reference Interval Prediction (DRRIP) [11] replacement policies. Specifically, ECM shows an average effective capacity increase of 15% over LRU and 18.8% over DRRIP, an average cache miss reduction of 9.4% over LRU and 3.9% over DRRIP, and an average system performance improvement of 6.2% over LRU and 3.3% over DRRIP. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Junghee Lee 0004, Jongman Kim |
HPCA | 2 |
| 2013 | Sharded Router: A novel on-chip router architecture employing bandwidth sharding and stealing
Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim |
Parallel Comput. | 3 |
| 2013 | TornadoNoC: A lightweight and scalable on-chip network architecture for the many-core eraabstractThe rapid emergence of Chip Multi-Processors (CMP) as the de facto microprocessor archetype has highlighted the importance of scalable and efficient on-chip networks. Packet-based Networks-on-Chip (NoC) are gradually cementing themselves as the medium of choice for the multi-/many-core systems of the near future, due to their innate scalability. However, the prominence of the debilitating power wall requires the NoC to also be as energy efficient as possible. To achieve these two antipodal requirements—scalability and energy efficiency—we propose TornadoNoC, an interconnect architecture that employs a novel flow control mechanism. To prevent livelocks and deadlocks, a sequence numbering scheme and a dynamic ring inflation technique are proposed, and their correctness formally proven. The primary objective of TornadoNoC is to achieve substantial gains in (a) scalability to many-core systems and (b) the area/power footprint, as compared to current state-of-the-art router implementations. The new router is demonstrated to provide better scalability to hundreds of cores than an ideal single-cycle wormhole implementation and other scalability-enhanced low-cost routers. Extensive simulations using both synthetic traffic patterns and real applications running in a full-system simulator corroborate the efficacy of the proposed design. Finally, hardware synthesis analysis using commercial 65nm standard-cell libraries indicates that the area and power budgets of the new router are reduced by up to 53% and 58%, respectively, as compared to existing state-of-the-art low-cost routers. Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim |
ACM Trans. Archit. Code Optim. | 3 |
| 2013 | IsoNet: Hardware-Based Job Queue Management for Many-Core ArchitecturesabstractImbalanced distribution of workloads across a chip multiprocessor (CMP) constitutes wasteful use of resources. Most existing load distribution and balancing techniques employ very limited hardware support and rely predominantly on software for their operation. This paper introduces IsoNet, a hardware-based conflict-free dynamic load distribution and balancing engine. IsoNet is a lightweight job queue manager responsible for administering the list of jobs to be executed, and maintaining load balance among all CMP cores. By exploiting a micro-network of load-balancing modules, the proposed mechanism is shown to effectively reinforce concurrent computation in many-core environments. Detailed evaluation using a full-system simulation framework indicates that IsoNet significantly outperforms existing techniques and scales efficiently to as many as 1024 cores. Furthermore, to assess its feasibility, the IsoNet design is synthesized, placed, and routed in 45-nm VLSI technology. Analysis of the resulting low-level implementation shows that IsoNet's area and power overhead are almost negligible. Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Shreepad Panth, Sung Kyu Lim, Jongman Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | A dual-phase compression mechanism for hybrid DRAM/PCM main memory architecturesabstractPhase-Change Memory (PCM) is emerging as a promising new memory technology, due to its inherent ability to scale deeply into the nanoscale regime. However, PCM is still marred by a duet of potentially show-stopping deficiencies: poor write performance and limited durability. These weaknesses have urged designers to develop various supporting architectural techniques to aid and complement the operation of the PCM, while mitigating its innate flaws. One promising such solution is the deployment of hybridized memory architectures that fuse DRAM and PCM, in order to combine the best attributes of each technology. In this paper, we introduce a Dual-Phase Compression (DPC) scheme specifically optimized for DRAM/PCM hybrid environments. Extensive simulations with traces from real applications running on a full-system simulator of a multicore system demonstrate 35.1% performance improvement and 29.3% energy reduction, on average, as compared to a baseline DRAM/PCM hybrid implementation. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Jongman Kim |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | A High-Performance and Energy-Efficient Virtually Tagged Stack Cache Architecture for Multi-core EnvironmentsabstractVirtually tagged caches possess a key attribute that renders them more attractive than Physically Indexed, Physically Tagged (PIPT) caches, they operate natively within the virtual address space, hence taking full advantage of the original intention of a virtual memory implementation: the illusion of a contiguous address space. Consequently, virtually tagged caches eliminate Translation Look-aside Buffer (TLB) references, yielding both energy and performance improvements. On the other hand, virtually tagged caches incur substantial overhead in resolving homonym/synonym issues, which is a fairly complicated process in contemporary multicore environments. In this paper, we aim to markedly alleviate this overhead through the use of a new virtually tagged stack cache design specifically targeting multi-core environments. It will be demonstrated that special level-one virtually tagged stack caches can significantly boost the performance of a system running a heavy, dominating, multi-threaded workload -- among other applications -- while actually reducing its energy consumption. This scheme is aimed at modern server environments that run a single, dedicated multi-threaded application workload per server. The proposed virtually tagged stack cache for multi-core processors minimizes the overhead incurred in resolving virtual-tag-related artifacts, by granting exclusive access to only one multi-threaded workload at a time. In other words, said virtually tagged cache filters the Virtual Address (VA) spaces and subsequently handles only the stack areas of the selected virtual address space. A cost effective way to implement the proposed stack cache for multi-core systems is also presented, yielding average performance improvements of around 20%. Suk Chan Kang, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim |
HPCC | 3 |
| 2011 | An energy- and performance-aware DRAM cache architecture for hybrid DRAM/PCM main memory systemsabstractThe last few years have witnessed the emergence of a promising new memory technology. Phase-Change Memory (PCM) is increasingly viewed as an attractive alternative for the memory sub-system of future microprocessor architectures, mainly because of its inherent ability to scale deeply into the nanoscale regime, and its low power consumption. However, PCM's write performance is its Achilles' heel, especially when compared to the prevalent DRAM technology. This weakness necessitates the deployment of hybridized solutions that fuse DRAM and PCM, in order to attain high overall system performance. In this paper, we set out to explore how various DRAM/PCM hybrid configurations affect system performance and energy consumption, and then proceed with the presentation of a novel architecture that maximizes performance without adversely affecting power efficiency. An energy-delay product improvement of 42.2%, on average, over conventional hybrid structures, is demonstrated. Hyung Gyu Lee, Seungcheol Baek, Chrysostomos Nicopoulos, Jongman Kim |
ICCD | 1 |
| 2011 | Hardware-Based Job Queue Management for Manycore Architectures and OpenMP EnvironmentsabstractThe seemingly interminable dwindle of technology feature sizes well into the nano-scale regime has afforded computer architects with an abundance of computational resources on a single chip. The Chip Multi-Processor (CMP) paradigm is now seen as the de facto architecture for years to come. However, in order to efficiently exploit the increasing number of on-chip processing cores, it is imperative to achieve and maintain efficient utilization of the resources at run time. Uneven and skewed distribution of workloads misuses the CMP resources and may even lead to such undesired effects as traffic and temperature hotspots. While existing techniques rely mostly on software for the undertaking of load balancing duties and exploit hardware mainly for synchronization, we will demonstrate that there are wider opportunities for hardware support of load balancing in CMP systems. Based on this fact, this paper proposes IsoNet, a conflict-free dynamic load distribution engine that exploits hardware aggressively to reinforce massively parallel computation in many core settings. Moreover, the proposed architecture provides extensive fault-tolerance against both CPU faults and intra-IsoNet faults. The hardware takes charge of both (1) the management of the list of jobs to be executed, and (2) the transfer of jobs between processing elements to maintain load balance. Experimental results show that, unlike the existing popular techniques of blocking and job stealing, IsoNet is scalable with as many as 1024 processing cores. Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim |
IPDPS | 4 |
| 2008 | A PRAM and NAND flash hybrid architecture for high-performance embedded storage subsystemsabstractNAND flash-based storage is widely used in embedded systems due to its numerous benefits: low cost, high density, small form factor and so on. However, NAND flash-based storage is still suffering from serious performance degradation for random or small size write access. This degradation mainly comes from the physical constraints of NAND flash: erase-before-program and different unit size of erase and program operations. To overcome these constraints, we propose to use PRAM (Phase-change RAM) which supports advanced features: fast byte access capability and no requirement for erase-before-program. Jin Kyu Kim, Hyung Gyu Lee, Shinho Choi, Kyoung Il Bahng |
EMSOFT | 2 |
| 2007 | On-chip communication architecture exploration: A quantitative evaluation of point-to-point, bus, and network-on-chip approachesabstractTraditionally, design-space exploration for systems-on-chip (SoCs) has focused on the computational aspects of the problem at hand. However, as the number of components on a single chip and their performance continue to increase, a shift from computation-based to communication-based design becomes mandatory. As a result, the communication architecture plays a major role in the area, performance, and energy consumption of the overall system. This article presents a comprehensive evaluation of three on-chip communication architectures targeting multimedia applications. Specifically, we compare and contrast the network-on-chip (NoC) with point-to-point (P2P) and bus-based communication architectures in terms of area, performance, and energy consumption. As the main contribution, we present complete P2P, bus-, and NoC-based implementations of a real multimedia application (i. e. the MPEG-2 encoder), and provide direct measurements using an FPGA prototype and actual video clips, rather than simulation and synthetic workloads. We also support the experimental findings through a theoretical analysis. Both experimental and analysis results show that the NoC architecture scales very well in terms of area, performance, energy, and design effort, while the P2P and bus-based architectures scale poorly on all accounts except for performance and area, respectively. Hyung Gyu Lee, Naehyuck Chang, Ümit Y. Ogras, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | Design space exploration and prototyping for on-chip multimedia applicationsabstractTraditionally, design space exploration for Systems-on-Chip (SoCs) has focused on the computational aspects of the problem at hand. However, as the number of components on a single chip and their performance continue to increase, a shift from computation-bound to communication-bound design becomes mandatory. Towards this end, this paper presents a comprehensive evaluation of two communication architectures targeting multimedia applications. Specifically, we compare and contrast the Network-on-Chip (NoC) and Point-to-Point (P2P) communication architectures in terms of power, performance, and area. As the main contribution, we present complete P2P and NoC-based implementations of a real multimedia application (MPEG-2 encoder), and provide direct measurements using a FPGA prototype and actual video clips, rather than simulation and synthetic workload. From an experi-mental standpoint, we show that the NoC architecture scales very well in terms of area, performance, power and design effort, while the P2P architecture scales poorly on all accounts except performance. Hyung Gyu Lee, Ümit Y. Ogras, Radu Marculescu, Naehyuck Chang |
DAC | 1 |
| 2006 | Communication architecture optimization: making the shortest path shorter in regular networks-on-chipabstractNetwork-on-chip (NoC)-based communication represents a promising solution to complex on-chip communication problems. Due to their regular structure, mesh-like NoC architectures have become very popular recently. However, they have poor topological properties such as long inter-node distances. In this paper, we address this very issue and explore the potential of partial NoC customization to improve both static and dynamic properties of the network significantly, while minimally affecting its regularity. Precise energy measurements on an FPGA prototype show that the improvement in network properties is achieved without a significant penalty in area and communication energy consumption. Ümit Y. Ogras, Radu Marculescu, Hyung Gyu Lee, Naehyuck Chang |
DATE | 3 |
| 2003 | Energy-aware memory allocation in heterogeneous non-volatile memory systemsabstractMemory systems consume a significant portion of power in hand-held embedded systems. So far, low-power memory techniques have addressed the power consumption when the system is turned on. In this paper, we consider data retention energy during the power-off period. For this purpose, we first characterize the data retention energy and cycle-accurate active mode energy of the non-volatile memory systems. Next, we present energy-aware memory allocation for a given task set taking into account arrival rate, execution time, code size, user data size and the number of memory transactions by the use of trace-driven simulation. Experiments demonstrate that our optimal configuration can save up to 26% of the memory system energy compared with traditional allocation schemes. Hyung Gyu Lee, Naehyuck Chang |
ISLPED | 1 |
| 2003 | Low-energy off-chip SDRAM memory systems for embedded applicationsabstractMemory systems are dominant energy consumers, and thus many energy reduction techniques for memory buses and devices have been proposed. For practical energy reduction practices, we have to take into account the interaction between a processor and cache memories together with application programs. Furthermore, energy characterization of memory systems must be accurate enough to justify various techniques. In this article, we build an in-house energy simulator for memory systems that is accelerated by special hardware support while maintaining accuracy. We explore energy behavior of memory systems for various values of the processor and memory clock frequencies and cache configuration. Each experiment is performed with 24M instruction steps of real application programs to guarantee accuracy.The simulator is based on precise energy characterization of memory systems including buses, bus drivers, and memory devices by a cycle-accurate energy measurement technique. We characterize energy consumption of each component by an energy state machine whose states and transitions are associated with the dynamic and static energy costs, respectively. Our approach easily characterizes the energy consumption of complex SDRAMs. We divide and quantify energy components of main memory systems for high-level reduction. The energy simulator enables us to devise practical energy reduction schemes by providing the actual amount of reduction out of the total energy consumption in main memory systems. We introduce several practical energy reduction techniques for SDRAM memory systems and demonstrate energy reduction ratio over the SDRAM memory systems with commercial SDRAM controller chipsets. We classify the SDRAM memory systems into high-performance and mid-performance classes and achieve suitable system configurations for each class. For instance, a typical high-performance 32-bit, 64 MB SDRAM memory system consumes 19.6 mJ, 33.8 mJ, 35.4 mJ, and 37.0 mJ for 24M instructions of an MP3 decoder, a JPEG compressor, a JPEG decompressor, and an MPEG4 decoder, respectively. Our reduction scheme saves 12.7 mJ, 15.1 mJ, 15.5 mJ, and 14.8 mJ, and the reduction ratios are 64.8%, 44.6%, 43.8%, and 40.1%, respectively, without compromising execution speed. Hojun Shim, Yongsoo Joo, Yongseok Choi, Hyung Gyu Lee, Naehyuck Chang |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2002 | Energy exploration and reduction of SDRAM memory systemsabstractIn this paper, we introduce a precise energy characterization of SDRAM main memory systems and explore the amount of energy associated with design parameters, leading to energy reduction techniques that we are able to recommend for practical use. Yongsoo Joo, Yongseok Choi, Hojun Shim, Hyung Gyu Lee, Kwanho Kim, Naehyuck Chang |
DAC | 4 |
| 2002 | Cycle-accurate energy measurement and characterization with a case study of the ARM7TDMI [microprocessors]abstractEnergy characterization is the basis for high-level energy reduction. Measurement-based characterization is accurate and independent of model availability and is thus suitable for commercial off-the-shelf (COTS) components, but conventional measurement equipment has serious limitations in this context. We introduce a new technique for the energy characterization of a microprocessor using a cycle-accurate energy measurement system based on charge transfer which is robust to spiky noise and is able to collect a range of energy consumption profiles in real time. It measures the energy variation of the CPU core by changing the instruction-level energy-sensitive factors such as opcodes (operations), instruction fetch addresses, register numbers, register values, data fetch addresses and immediate operand values at each pipeline stage. Using the ARM7TDMI RISC processor as a case study, we observe that the energy contributions of most instruction-level energy-sensitive factors are orthogonal to the operations. We are able to characterize the energy variation, preserving all the effects of the energy-sensitive factors for various software methods of energy reduction. We also demonstrate applications of our measurement and characterization techniques. Naehyuck Chang, Kwanho Kim, Hyung Gyu Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | Cycle-accurate energy consumption measurement and analysis: case study of ARM7TDMIabstractWe introduce an energy consumption analysis of complex digital systems through a case study of ARM7TDMI RISC processor by using a new energy measurement technique. We developed a cycle-accurate energy consumption measurement system based on charge transfer which is robust to spiky noise and is capable of collecting a range of power consumption profiles in real time. The relative energy variation of the RISC core is measured by changing the opcode, the instruction fetch address, the register number, in each pipeline stage, respectively. We demonstrated energy characterization of a pipelined RISC processor for high-level power reduction. Naehyuck Chang, Kwanho Kim, Hyung Gyu Lee |
ISLPED | 3 |