Junghee Lee 0004

dblp:63/4391-4 · DBLP profile ↗
← Back
34ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0003-0733-0136ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 8 first-author · 5 since 2021Security and privacy · 5 · 4 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Implementation study of cost-effective verification for Pietrzak's VDF in Ethereum smart contract
abstract
Verifiable Delay Function (VDF) is a cryptographic concept that ensures a minimum delay before output through sequential processing, which is resistant to parallel computing. One of the significant VDF protocols academically reviewed is the VDF protocol proposed by Pietrzak. However, for the blockchain environment, the Pietrzak VDF has drawbacks including long proof size and recursive protocol computation. In this paper, we present an implementation study of Pietrzak VDF verification on Ethereum Virtual Machine (EVM). We found that the discussion in the Pietrzak's original paper can help a clear optimization in EVM where the costs of computation are predefined as the specific amounts of gas. In our results, the cost of VDF verification can be reduced from 4M to 2M gas, and the proof length can be generated under 8 KB with the 2048-bit RSA key length, which is much smaller than the previous expectation.
Suhyeon Lee 0001, Euisin Gee, Junghee Lee 0004
Blockchain Res. Appl.3
2026 Network breaker: Practical countermeasure to denial of service attacks on smart grid
Hyuk Kwon, Wonjun Lee 0001, Junghee Lee 0004, Yoojin Kwon, No-Gil Myoung, Myunghye Park, Jae-ju Song
Comput. Secur.4
2025 GAE4HT: Detecting Hardware Trojans with Graph Autoencoder-Trained on Golden Model Data Flow Graphs
Daehyeon Lee, Junghee Lee 0004
AsiaCCS2
2025 ByteExpress: A High-Performance and Traffic-Efficient Inline Transfer of Small Payloads over NVMe
abstract
Recent computational storage devices enable host-side tasks such as SQL filtering and key-value operations to be offloaded to the device. However, these tasks often involve small payloads, typically a few dozen to hundreds of bytes, which are inefficiently handled by the conventional NVMe protocol due to its page-based DMA mechanism. Even tiny payloads incur 4 KB PCIe transfers, leading to severe bandwidth waste and increased latency. Prior approaches either break NVMe compatibility or are only effective for very small payloads on the order of a few dozen bytes. This paper presents ByteExpress, a new mechanism that efficiently transmits small payloads by placing them inline in 64-byte chunks directly into the NVMe submission queue, immediately following the NVMe command. ByteExpress requires only slight modifications to the NVMe driver and controller logic, while preserving full compatibility with existing APIs and SSD architectures. We implemented ByteExpress on the Linux NVMe driver and OpenSSD, demonstrating up to 98% reduction in PCIe traffic and 40% and 39% lower latency compared to PRP and a state-of-the-art approach, respectively, for sub-page payloads.
Junhyeok Park 0002, Junghee Lee 0004, Youngjae Kim 0001
HotStorage2
2025 MultiFile View: File-View-Based Isolation in a Single-User Environment to Protect User Data Files
abstract
Isolation technology is often used to reduce the impact of cyber attacks, and it is mainly used in multi-user environments. Representative examples of said technology include access-control mechanisms and virtual machines. In a single-user environment, virtual addressing and a trusted execution environment isolate applications. However, the focus of such techniques is usually only on isolation, while the sharing of files has not been given much attention. In a single-user environment, users have the ability to access the same file through multiple applications. In this paper, we introduce the concept offile view, and propose file isolation based on the notion ofviews. Under the proposedMultiFile Viewmechanism, files within a view can be accessed by multiple applications when that particular view is activated. In other words, files appear only if their associated view is activated. The proposed technique is effective in protecting files from attacks on user data files, such as ransomware, wiper, and evil maid attacks. We also develop three models to describe how to assign views to applications. The proposed technique is prototyped in Windows 10. Through extensive experiments, we demonstrate that the new technique’s performance overhead does not noticeably affect the overall user experience.
Jione Choi, Junghee Lee 0004, Jaegwan Yu, Aran Park, Chrysostomos Nicopoulos
IEEE Trans. Dependable Secur. Comput.2
2024 BandSlim: A Novel Bandwidth and Space-Efficient KV-SSD with an Escape-from-Block Approach
abstract
The Key-Value Solid State Drive (KV-SSD) represents a significant evolution in storage device interfaces by accommodating non-page-aligned key-value pairs, a departure from conventional models. However, KV-SSDs encounter challenges as their specialized data transfer and packing requirements conflict with established storage protocols like NVMe, which are designed around fixed memory page units. This discord leads to inefficient data movement and increased NAND page write I/Os, which in turn escalates network traffic and degrades both performance and NAND efficiency. To tackle these challenges, this paper introduces BandSlim, a novel solution equipped with two methods to streamline bandwidth during I/O transmission: (i) a fine-grained inline value transfer utilizing NVMe commands for bandwidth-efficient value transfer, and (ii) a selective value packing strategy combined with a backfilling policy to reduce NAND page write I/Os. We integrated BandSlim on a state-of-the-art FPGA-based LSM-tree KV-SSD, utilizing the Cosmos+ OpenSSD platform. Our comprehensive evaluations illustrate that BandSlim achieves a remarkable reduction in PCIe traffic of up to 97.9% and NAND page write counts by up to 98.1% compared to the NVMe-based KV-SSD without employing BandSlim.
Junhyeok Park 0002, Chang-Gyu Lee, Soon Hwang, Soonyeal Yang, Jungki Noh, Woosuk Chung, Junghee Lee 0004, Youngjae Kim 0001
ICPP7
2024 Robust Hardware Trojan Detection Method by Unsupervised Learning of Electromagnetic Signals
abstract
This article explores the threat posed by Hardware Trojans (HTs), malicious circuits clandestinely embedded in hardware akin to software backdoors. Activation by attackers renders these Trojans capable of inducing malfunctions or leaking confidential information by manipulating the hardware’s normal operation. Despite robust software security, detecting and ensuring normal hardware operation becomes challenging in the presence of malicious circuits. This issue is particularly acute in weapon systems, where HTs can present a significant threat, potentially leading to immediate disablement in adversary countries. Given the severe risks associated with HTs, detection becomes imperative. The study focuses on demonstrating the efficacy of deep learning-based HT detection by comparing and analyzing methods using deep learning with existing approaches. This article proposes utilizing the deep support vector data description (Deep SVDD) model for HT detection. The proposed method outperforms existing methods when detecting untrained HTs. It achieves 92.87% of accuracy on average, which is higher than that of an existing method, 50.00%. This finding contributes valuable insights to the field of hardware security and lays the foundation for practical applications of Deep SVDD in real-world scenarios.
Daehyeon Lee, Junghee Lee 0004, Janghyuk Kauh, Taigon Song
IEEE Trans. Very Large Scale Integr. Syst.2
2022 A Content-Based Ransomware Detection and Backup Solid-State Drive for Ransomware Defense
abstract
Ransomware is a growing concern in business and government because it causes immediate financial damages or loss of important data. There is a way to detect and block ransomware in advance, but evolved ransomware can still attack while avoiding detection. Another alternative is to back up the original data. However, existing backup solutions can be under the control of ransomware and backup copies can be destroyed by ransomware. Moreover, backup methods incur storage and performance overhead. In this article, we propose AMOEBA, a device-level backup solution that does not require additional storage for backup. AMOEBA is armed with: 1) a hardware accelerator to run content-based detection algorithms for ransomware detection at high speed and 2) a fine-grained backup control mechanism to minimize space overhead for data backup. For evaluations, we not only implemented AMOEBA using the Microsoft solid-state drive (SSD) simulator but also prototyped it on the OpenSSD-platform. Our extensive evaluations with real ransomware workloads show that AMOEBA has high ransomware detection accuracy with negligible performance overhead.
Donghyun Min, Yungwoo Ko, Ryan Walker, Junghee Lee 0004, Youngjae Kim 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 A Hardware-Assisted Heartbeat Mechanism for Fault Identification in Large-Scale IoT Systems
abstract
With increased inter-connectivity among disparate devices, such as Internet-of-Things (IoT) devices, including those deployed in a nation’s critical infrastructure, there is a need to ensure that any failure in the deployed devices can be detected. The capability to automatically detect device failures is particularly crucial in a large-scale, complex IoT system, since it can be very time-consuming and challenging to investigate a large number of geographically-dispersed devices that are also of different makes and types. In this paper, we present a faulty-device identification technique that is designed to achieve lightweight processor-level architectural support. Specifically, a hardware-based monitoring agent is incorporated within a processor and connected to a separate monitoring program when an examination is required. By analyzing information collected by the agent, the monitoring program determines whether the device being monitored is functioning. Findings from our detailed evaluation demonstrate that the proposed approach can detect around 90 percent of the failures with minimal hardware overhead of approximately 5k gates. This area overhead is reasonable and amounts to 7.69 percent of the ARM Cortex-M4 – a lightweight IoT-class processor – that has a total area (excluding optional caches and scratch-pad memory) of 65k gates.
Mandrita Banerjee, Carlo Borges, Kim-Kwang Raymond Choo, Junghee Lee 0004, Chrysostomos Nicopoulos
IEEE Trans. Dependable Secur. Comput.4
2021 Multicast-enabled network-on-chip routers leveraging partitioned allocation and switching
Dimitris Konstantinou, Chrysostomos Nicopoulos, Junghee Lee 0004, Giorgos Dimitrakopoulos
Integr.3
2021 IoTCop: A Blockchain-Based Monitoring Framework for Detection and Isolation of Malicious Devices in Internet-of-Things Systems
abstract
Unlike conventional servers housed in a centralized and secured indoor environment (e.g., data centers), Internet-of-Things (IoT) devices such as sensor/actuator are geographically distributed and may be closely located to the physical systems where IoT devices are utilized. However, the resource-constrained nature of IoT devices limits their capacity to deploy sophisticated security solutions. The proposed approach assumes that a device can be compromised and hence, the need to be able to automatically isolate the compromised device(s). In order to enforce security policies even when devices are compromised, we propose using blockchain in the monitoring framework. Unlike existing centralized or distributed security solutions (which do not consider the possibility that the solutions themselves can be compromised), the proposed blockchain-based framework can enforce the security policies as long as a majority of the devices are not compromised. By employing the permissioned blockchain (Hyperledger Fabric) and add-on hardware modules, the proposed framework offers significantly lower latency and overhead compared to permissionless blockchain frameworks (e.g., Ethereum) and allows existing IoT devices to join the framework without modification.
Sreenivas Sudarshan Seshadri, Mukunda Subedi, Kim-Kwang Raymond Choo, Qian Chen 0019, Junghee Lee 0004
IEEE Internet Things J.7
2020 DISKSHIELD: A Data Tamper-Resistant Storage for Intel SGX
abstract
With the increasing importance of data, the threat of malware which destroys data has been increasing. If malware acquires the highest software privilege, any attempt to detect and remove malware can be disabled. In this paper, we propose DISKSHIELD, a secure storage framework. DISKSHIELD uses Intel SGX to provide Trusted Execution Environment (TEE) to the host, implements the file system into SSD firmware that provides a Trusted Computing Base (TCB), and uses a two-way authentication mechanism to securely transfer data from the host TEE to the SSD TCB against data tampering attacks. This design frees DISKSHIELD from attacks to the kernel. To show the efficacy of DISKSHIELD, we prototyped a DISKSHIELD system by modifying Intel IPFS and developing a device file system on the Jasmine OpenSSD Platform in a Linux environment. Our results show that DISKSHIELD provides strong data tamper resistance the throughput of read and write is on average to 28%, 19% lower than IPFS.
Jinwoo Ahn, Junghee Lee 0004, Yungwoo Ko, Donghyun Min, Jiyun Park, Sungyong Park, Youngjae Kim 0001
AsiaCCS2
2020 Position: SGX-SSD: A Policy-based Versioning SSD with Intel SGX
Jinwoo Ahn, Jinhoon Lee, Yungwoo Ko, Donghyun Min, Junghee Lee 0004, Youngjae Kim 0001
HotStorage6
2020 SmartFork: Partitioned Multicast Allocation and Switching in Network-on-Chip Routers
abstract
Multicast on-chip communication is encountered in various cache-coherence protocols targeting multi-core processors, and its pervasiveness is increasing due to the proliferation of machine learning accelerators. In-network handling of multicast traffic imposes additional switching-level restrictions to guarantee deadlock freedom, while it stresses the allocation efficiency of Network-on-Chip (NoC) routers. In this work, we propose a novel NoC router microarchitecture, called SmartFork, which employs a versatile and cost-efficient multicast packet replication scheme that allows the design of high-throughput and low-cost NoCs. The design is adapted to the average branch splitting observed in real-world multicast routing algorithms. Compared to state-of-the-art NoC multicast approaches, SmartFork is demonstrated to yield higher performance in terms of latency and throughput, while still offering a cost-effective implementation.
Dimitris Konstantinou, Chrysostomos Nicopoulos, Junghee Lee 0004, Georgios Ch. Sirakoulis, Giorgos Dimitrakopoulos
ISCAS3
2018 Blockchain-Based Security Layer for Identification and Isolation of Malicious Things in IoT: A Conceptual Design
abstract
Internet-of-Things (IoT) is increasingly becoming the norm in both civilian and military settings. In this paper, we present a comprehensive security abstraction layer for IoT systems based on blockchain, which provides us a logical view of a system that comprises trusted devices. The goal of the proposed layer is to detect and isolate untrusted devices. The proposed abstraction layer provides three services, namely: authorization, authentication, and auditing by using blockchain and smart contract-based approaches. We adopt a hardware based approach, where dedicated hardware modules are used to monitor the behavior of the firmware without incurring excessive performance overhead.
Mandrita Banerjee, Junghee Lee 0004, Qian Chen 0019, Kim-Kwang Raymond Choo
ICCCN2
2017 Vulnerability Analysis of On-Chip Access-Control Memory
Chintan Chavda, Ethan C. Ahn, Yu-Sheng Chen, Youngjae Kim 0001, Kalidas Ganesh, Junghee Lee 0004
HotStorage6
2017 LAWC: Optimizing Write Cache Using Layout-Aware I/O Scheduling for All Flash Storage
abstract
Flash memory-based SSD-RAIDs are swiftly replacing conventional hard disk drives by exhibiting improved performance and stability, especially in I/O-intensive environments. However, the variations in latency and throughput occurring due to uncoordinated internal garbage collection cripples further boosting of performance. In addition, the unwanted variations in each SSD can influence the overall performance of the entire flash storage adversely. This performance bottleneck can be essentially reduced by an internal write cache in the RAID controller designed prudently by considering the crucial device characteristics. The state-of-the-art cache write for the RAID controller fails to incorporate device characteristics of flash memory-based SSDs and mitigates the performance gain. In this paper, we propose a novel cache design namely Layout-Aware Write Cache (LAWC) to overcome the performance barrier inculcated by independent garbage collections. LAWC implements (i) improved I/O scheduling for logically partitioned write caches, (ii) a destage write synchronization mechanism to allow individual write caches to flush write blocks into the SSD array in a coordinated manner, and (iii) a two-level hybrid cache algorithm utilizing small front level cache for the improved write cache efficiency. LAWC shows significant reduction in response time by 82.39 percent on RAID-0 and 68.51 percent on RAID-5 types of SSDs when compared with state-of-theart write cache algorithms.
Kalidas Ganesh, Youngjae Kim 0001, Monobrata Debnath, Sungyong Park, Junghee Lee 0004
IEEE Trans. Computers5
2017 HoPE: Hot-Cacheline Prediction for Dynamic Early Decompression in Compressed LLCs
abstract
Data compression plays a pivotal role in improving system performance and reducing energy consumption, because it increases the logical effective capacity of a compressed memory system without physically increasing the memory size. However, data compression techniques incur some cost, such as non-negligible compression and decompression overhead. This overhead becomes more severe if compression is used in the cache. In this article, we aim to minimize the read-hit decompression penalty in compressed Last-Level Caches (LLCs) by speculatively decompressing frequently used cachelines. To this end, we propose a Hot-cacheline Prediction and Early decompression (HoPE) mechanism that consists of three synergistic techniques: Hot-cacheline Prediction (HP), Early Decompression (ED), and Hit-history-based Insertion (HBI). HP and HBI efficiently identify the hot compressed cachelines, while ED selectively decompresses hot cachelines, based on their size information. Unlike previous approaches, the HoPE framework considers the performance balance/tradeoff between the increased effective cache capacity and the decompression penalty. To evaluate the effectiveness of the proposed HoPE mechanism, we run extensive simulations on memory traces obtained from multi-threaded benchmarks running on a full-system simulation framework. We observe significant performance improvements over compressed cache schemes employing the conventional Least-Recently Used (LRU) replacement policy, the Dynamic Re-Reference Interval Prediction (DRRIP) scheme, and the Effective Capacity Maximizer (ECM) compressed cache management mechanism. Specifically, HoPE exhibits system performance improvements of approximately 11%, on average, over LRU, 8% over DRRIP, and 7% over ECM by reducing the read-hit decompression penalty by around 65%, over a wide range of applications.
Jaehyun Park 0005, Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Vinson Young, Junghee Lee 0004, Jongman Kim
ACM Trans. Design Autom. Electr. Syst.6
2016 Minimizing CMT Miss Penalty in Selective Page-Level Address Mapping Table
abstract
Flash Translation Layer (FTL) performs virtual-to-physical address translations and hides the erase-before-write characteristics of Flash. Pure page mapped FTL, which maintains page-level address mappings, is known as the most efficient FTL. However, its huge SRAM requirement to load the entire mapping table limited adoption of its use. In order to reduce SRAM space utilization while maintaining comparable performance, we can selectively cache page-level address mappings into a small SRAM. However, the performance of this approach is limited by miss ratio of cached mapping table (CMT) on SRAM. In this paper, we propose a replica approach of the page-mapped FTL on flash, called Replica to minimize the performance penalty of CMT miss.
Ronnie Mativenga, Joon-Young Paik, Junghee Lee 0004, Tae-Sun Chung, Youngjae Kim 0001
CLUSTER3
2016 A Low-Power Network-on-Chip Architecture for Tile-based Chip Multi-Processors
abstract
Technology scaling of tiled-based CMPs reduces the physical size of each tile and increases the number of tiles per die. This trend directly impacts the on-chip interconnect; even though the tile population increases, the inter-tile link distances scale down proportionally to the tile dimensions. The decreasing inter-tile wire lengths can be exploited to enable swift link traversal between neighboring tiles, after appropriate wire engineering. Building on this premise, we propose a technique to rapidly transfer its between adjacent routers in half a clock cycle, by utilizing both edges of the clock during the sending and receiving operations. Half-cycle link traversal enables, for the first time, substantial reductions in (a) link power, irrespective of the data switching profile, and (b) buffer power (through buffer-size reduction), without incurring any latency/throughput loss. In fact, the proposed architecture also yields some latency improvements over a baseline NoC. Detailed hardware analysis using placed-and-routed designs, and cycle-accurate full-system simulations corroborate the significant power and latency improvements.
Anastasios Psarras, Junghee Lee 0004, Pavlos M. Mattheakis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
ACM Great Lakes Symposium on VLSI2
2016 PhaseNoC: Versatile Network Traffic Isolation Through TDM-Scheduled Virtual Channels
abstract
As multi/many-core architectures evolve, the demands on the network-on-chip (NoC) are amplified. In addition to high performance and physical scalability, the NoC is increasingly required to also provide specialized functionality, such as network virtualization, flow isolation, and quality-of-service. Although traditional architectures supporting virtual channels (VCs) offer the resources for flow partitioning and isolation, an adversarial workload can still interfere and degrade the performance of other workloads that are active in a different set of VCs. In this paper, we present PhaseNoC, a truly noninterfering VC-based architecture that adopts time-division multiplexing at the VC level. Distinct flows, or application domains, mapped to disjoint sets of VCs are isolated, both inside the router's pipeline and at the network level. Any latency overhead is minimized by appropriate scheduling of flows in separate phases of operation, irrespective of the chosen topology. When strict isolation is not required, the proposed architecture can employ opportunistic bandwidth stealing. This novel mechanism works synergistically with the baseline PhaseNoC techniques to improve the overall latency/throughput characteristics of the NoC, while still preserving performance isolation. Experimental results corroborate that-with lower cost than state-of-the-art NoC architectures, and with minimum latency overhead-PhaseNoC removes any flow interference and allows for efficient network traffic isolation.
Anastasios Psarras, Junghee Lee 0004, Ioannis Seitanidis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 Size-Aware Cache Management for Compressed Cache Architectures
abstract
A practical way to increase the effective capacity of a microprocessor's cache, without physically increasing the cache size, is to employ data compression. Last-Level Caches (LLC) are particularly amenable to such compression schemes, since the primary purpose of the LLC is to minimize the miss rate, i.e., it directly benefits from a larger logical capacity. In compressed LLCs, the cacheline size varies depending on the achieved compression ratio. Our observations indicate that this size information gives useful hints when managing the cache (e.g., when selecting a victim), which can lead to increased cache performance. However, there are currently no replacement policies tailored to compressed LLCs; existing techniques focus primarily on locality information. This article introduces the concept of size-aware cache management as a way to maximize the performance of compressed caches. Upon analyzing the benefits of considering size information in the management of compressed caches, we propose a novel mechanism-called Effective Capacity Maximizer (ECM)-to further enhance the performance and energy consumption of compressed LLCs. The proposed technique revolves around four fundamental principles: ECM Insertion (ECM-I), ECM Promotion (ECM-P), ECM Eviction Scheduling (ECM-ES), and ECM Replacement (ECM-R). Extensive simulations with memory traces from real applications running on a full-system simulator demonstrate significant improvements compared to compressed cache schemes employing conventional locality-aware cache replacement policies. Specifically, our ECM shows an average effective capacity increase of 18.4 percent over the Least-Recently Used (LRU) policy, and 23.9 percent over the Dynamic Re-Reference Interval Prediction (DRRIP) [1] scheme. This translates into average system performance improvements of 7.2 percent over LRU and 4.2 percent over DRRIP. Moreover, the average energy consumption is also reduced by 5.9 percent over LRU and 3.8 percent over DRRIP.
Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Junghee Lee 0004, Jongman Kim
IEEE Trans. Computers4
2014 Coordinating Garbage Collectionfor Arrays of Solid-State Drives
abstract
Although solid-state drives (SSDs) offer significant performance improvements over hard disk drives (HDDs) for a number of workloads, they can exhibit substantial variance in request latency and throughput as a result of garbage collection (GC). When GC conflicts with an I/O stream, the stream can make no forward progress until the GC cycle completes. GC cycles are scheduled by logic internal to the SSD based on several factors such as the pattern, frequency, and volume of write requests. When SSDs are used in a RAID with currently available technology, the lack of coordination of the SSD-local GC cycles amplifies this performance variance. We propose a global garbage collection (GGC) mechanism to improve response times and reduce performance variability for a RAID of SSDs. We include a high-level design of SSD-aware RAID controller and GGC-capable SSD devices and algorithms to coordinate the GGC cycles. We develop reactive and proactive GC coordination algorithms and evaluate their I/O performance and block erase counts for various workloads. Our simulations show that GC coordination by a reactive scheme improves average response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write-dominated workloads, response time was improved by 69 percent and performance variability was reduced by 71 percent. We show that a proactive GC coordination algorithm can further improve the I/O response times by up to 9 percent and the performance variability by up to 15 percent. We also observe that it could increase the lifetimes of SSDs with some workloads (e.g., Financial) by reducing the number of block erase counts by up to 79 percent relative to a reactive algorithm for write-dominant enterprise workloads.
Youngjae Kim 0001, Junghee Lee 0004, Sarp Oral, David Dillow, Feiyi Wang, Galen M. Shipman
IEEE Trans. Computers2
2013 ECM: Effective Capacity Maximizer for high-performance compressed caching
abstract
Compressed Last-Level Cache (LLC) architectures have been proposed to enhance system performance by efficiently increasing the effective capacity of the cache, without physically increasing the cache size. In a compressed cache, the cacheline size varies depending on the achieved compression ratio. We observe that this size information gives a useful hint when selecting a victim, which can lead to increased cache performance. However, no replacement policy tailored to compressed LLCs has been investigated so far. This paper introduces the notion of size-aware compressed cache management as a way to maximize the performance of compressed caches. Toward this end, the Effective Capacity Maximizer (ECM) scheme is introduced, which targets compressed LLCs. The proposed mechanism revolves around three fundamental principles: Size-Aware Insertion (SAI), a Dynamically Adjustable Threshold Scheme (DATS), and Size-Aware Replacement (SAR). By adjusting the eviction criteria, based on the compressed data size, one may increase the effective cache capacity and minimize the miss penalty. Extensive simulations with memory traces from real applications running on a full-system simulator demonstrate significant improvements compared to compressed cache schemes employing the conventional Least-Recently Used (LRU) and Dynamic Re-Reference Interval Prediction (DRRIP) [11] replacement policies. Specifically, ECM shows an average effective capacity increase of 15% over LRU and 18.8% over DRRIP, an average cache miss reduction of 9.4% over LRU and 3.9% over DRRIP, and an average system performance improvement of 6.2% over LRU and 3.3% over DRRIP.
Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Junghee Lee 0004, Jongman Kim
HPCA4
2013 Hardware-Assisted Intrusion Detection by Preserving Reference Information Integrity
Junghee Lee 0004, Chrysostomos Nicopoulos, Gi-Hwan Oh, Sang-Won Lee 0001, Jongman Kim
ICA3PP (1)1
2013 Sharded Router: A novel on-chip router architecture employing bandwidth sharding and stealing
Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim
Parallel Comput.1
2013 TornadoNoC: A lightweight and scalable on-chip network architecture for the many-core era
abstract
The rapid emergence of Chip Multi-Processors (CMP) as the de facto microprocessor archetype has highlighted the importance of scalable and efficient on-chip networks. Packet-based Networks-on-Chip (NoC) are gradually cementing themselves as the medium of choice for the multi-/many-core systems of the near future, due to their innate scalability. However, the prominence of the debilitating power wall requires the NoC to also be as energy efficient as possible. To achieve these two antipodal requirements—scalability and energy efficiency—we propose TornadoNoC, an interconnect architecture that employs a novel flow control mechanism. To prevent livelocks and deadlocks, a sequence numbering scheme and a dynamic ring inflation technique are proposed, and their correctness formally proven. The primary objective of TornadoNoC is to achieve substantial gains in (a) scalability to many-core systems and (b) the area/power footprint, as compared to current state-of-the-art router implementations. The new router is demonstrated to provide better scalability to hundreds of cores than an ideal single-cycle wormhole implementation and other scalability-enhanced low-cost routers. Extensive simulations using both synthetic traffic patterns and real applications running in a full-system simulator corroborate the efficacy of the proposed design. Finally, hardware synthesis analysis using commercial 65nm standard-cell libraries indicates that the area and power budgets of the new router are reduced by up to 53% and 58%, respectively, as compared to existing state-of-the-art low-cost routers.
Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim
ACM Trans. Archit. Code Optim.1
2013 Preemptible I/O Scheduling of Garbage Collection for Solid State Drives
abstract
Unlike hard disks, flash devices use out-of-place updates operations and require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in an I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semipreemptible GC (PGC) scheme that allows GC processing to be preempted while pending I/O requests in the queue are serviced. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-PGC scheme with realistic workloads demonstrates both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-PGC scheme. In addition, we explore opportunities of a new NAND flash device that supports suspend/resume commands for read, write, and erase operations for fully PGC (F-PGC). Our experiments with an F-PGC enabled flash device show that request response time can be improved by up to 14.57% compared to semi-PGC.
Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Jongman Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 IsoNet: Hardware-Based Job Queue Management for Many-Core Architectures
abstract
Imbalanced distribution of workloads across a chip multiprocessor (CMP) constitutes wasteful use of resources. Most existing load distribution and balancing techniques employ very limited hardware support and rely predominantly on software for their operation. This paper introduces IsoNet, a hardware-based conflict-free dynamic load distribution and balancing engine. IsoNet is a lightweight job queue manager responsible for administering the list of jobs to be executed, and maintaining load balance among all CMP cores. By exploiting a micro-network of load-balancing modules, the proposed mechanism is shown to effectively reinforce concurrent computation in many-core environments. Detailed evaluation using a full-system simulation framework indicates that IsoNet significantly outperforms existing techniques and scales efficiently to as many as 1024 cores. Furthermore, to assess its feasibility, the IsoNet design is synthesized, placed, and routed in 45-nm VLSI technology. Analysis of the resulting low-level implementation shows that IsoNet's area and power overhead are almost negligible.
Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Shreepad Panth, Sung Kyu Lim, Jongman Kim
IEEE Trans. Very Large Scale Integr. Syst.1
2011 Hardware-Based Job Queue Management for Manycore Architectures and OpenMP Environments
abstract
The seemingly interminable dwindle of technology feature sizes well into the nano-scale regime has afforded computer architects with an abundance of computational resources on a single chip. The Chip Multi-Processor (CMP) paradigm is now seen as the de facto architecture for years to come. However, in order to efficiently exploit the increasing number of on-chip processing cores, it is imperative to achieve and maintain efficient utilization of the resources at run time. Uneven and skewed distribution of workloads misuses the CMP resources and may even lead to such undesired effects as traffic and temperature hotspots. While existing techniques rely mostly on software for the undertaking of load balancing duties and exploit hardware mainly for synchronization, we will demonstrate that there are wider opportunities for hardware support of load balancing in CMP systems. Based on this fact, this paper proposes IsoNet, a conflict-free dynamic load distribution engine that exploits hardware aggressively to reinforce massively parallel computation in many core settings. Moreover, the proposed architecture provides extensive fault-tolerance against both CPU faults and intra-IsoNet faults. The hardware takes charge of both (1) the management of the list of jobs to be executed, and (2) the transfer of jobs between processing elements to maintain load balance. Experimental results show that, unlike the existing popular techniques of blocking and job stealing, IsoNet is scalable with as many as 1024 processing cores.
Junghee Lee 0004, Chrysostomos Nicopoulos, Hyung Gyu Lee, Jongman Kim
IPDPS1
2011 A semi-preemptive garbage collector for solid state drives
abstract
NAND flash memory is a preferred storage media for various platforms ranging from embedded systems to enterprise-scale systems. Flash devices do not have any mechanical moving parts and provide low-latency access. They also require less power compared to rotating media. Unlike hard disks, flash devices use out-of-update operations and they require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semi-preemptive GC scheme that can preempt on-going GC processing and service pending I/O requests in the queue. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-preemptive GC sheme with realistic workloads demonstrate both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-preemptive GC scheme.
Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Feiyi Wang, Jongman Kim
ISPASS1
2011 Harmonia: A globally coordinated garbage collector for arrays of Solid-State Drives
abstract
Solid-State Drives (SSDs) offer significant performance improvements over hard disk drives (HDD) on a number of workloads. The frequency of garbage collection (GC) activity is directly correlated with the pattern, frequency, and volume of write requests, and scheduling of GC is controlled by logic internal to the SSD. SSDs can exhibit significant performance degradations when garbage collection (GC) conflicts with an ongoing I/O request stream. When using SSDs in a RAID array, the lack of coordination of the local GC processes amplifies these performance degradations. No RAID controller or SSD available today has the technology to overcome this limitation. This paper presents Harmonia, a Global Garbage Collection (GGC) mechanism to improve response times and reduce performance variability for a RAID array of SSDs. Our proposal includes a high-level design of SSD-aware RAID controller and GGC-capable SSD devices, as well as algorithms to coordinate the global GC cycles. Our simulations show that this design improves response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write dominant workloads response time was improved by 69% while performance variability was reduced by 71%.
Youngjae Kim 0001, Sarp Oral, Galen M. Shipman, Junghee Lee 0004, David Dillow, Feiyi Wang
MSST4
2006 Cycle error correction in asynchronous clock modeling for cycle-based simulation
abstract
As the complexity of SoCs is increasing, hardware/software co-verification becomes an important part of system verification. C-level cycle-based simulation could be an efficient methodology for system verification because of its fast simulation speed. The cycle-based simulation has a limitation in using asynchronous clocks that causes inherent cycle errors. In order to reuse the output of a C-level cycle-based simulation for the verification of a lower level model, the C-level model should be cycle-accurate with respect to the lower level model. In this paper, a cycle error correction technique is presented for two asynchronous clock models. An example design is devised to show the effectiveness of the proposed method. Our experimental results show that the fast speed of cycle-based simulation can be fully exploited without sacrificing the cycle accuracy.
Junghee Lee 0004, Joonhwan Yi
ASP-DAC1
2003 Memory access pattern analysis and stream cache design for multimedia applications
abstract
Memory system is a major performance and power bottleneck in embedded systems especially for multimedia applications. Most multimedia applications access stream type of data structures with regular access patterns. It is observed that conventional caches behave poorly for stream-type data structure. Therefore, prediction-based prefetching techniques have been extensively researched to exploit the regular access patterns. Prefetching, however, may pollute the cache if the prediction is not accurate and needs extra hardware prediction logic. To overcome these problems, we propose a novel hardware prefetching technique that is assisted by static analysis of data access pattern with stream caches. With the proposed stream cache architecture, we could achieve significant performance improvement compared with the conventional cache architecture.
Junghee Lee 0004, Chanik Park, Soonhoi Ha
ASP-DAC1