VLDB 2026 Research / reviewers in the wild / expert
Jihong Kim 0001
dblp:74/2683
· DBLP profile ↗
86ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-7977-9883ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 71 · 20 since 2021Software engineering, systems software and programming languages · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9Databases, data management, data science and information retrieval · 6 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorComputer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STRAW: Stress-Aware WL-Based Read Disturbance Management for High-Density NAND Flash MemoryabstractWhile NAND flash memory has continuously increased its storage density over decades, this progress has exacerbated the read-disturbance problem. In this work, we identify two fundamental limitations of existing read-disturbance management techniques that trigger read reclaim (RR) at the block granularity: (i) they overlook the heterogeneous reliability impact of read disturbance across individual wordlines (WLs), leading to unnecessary RR in many cases; and (ii) they address read disturbance only after disturbance-induced errors have already accumulated, which forces substantial RR-induced copy overheads in read-disturbance-prone modern NAND flash memory. To address these limitations, we propose STRAW (STRess-Aware Wordline-based read-disturbance management), a new technique that minimizes RR overheads through two key ideas: (i) stress-aware WL-based read reclaim, which monitors the accumulated read-disturbance effect on each WL and reclaims only heavily disturbed WLs, and (ii) stress-reduced read, which mitigates disturbance on valid WLs during each read operation by scaling pass-through voltages based on WL validity. Our experimental results using a modern SSD emulator show that STRAW reduces RR-induced page-copy overhead by 88.6% on average compared with the state-of-the-art technique. Myoungjun Chun, Jae Yong Lee 0004, Inhyuk Choi, Jisung Park 0001, Myungsuk Kim, Jihong Kim 0001 |
ASPLOS (2) | 6 |
| 2026 | P2Cache: Enhancing Data-Centric Applications via Application-Guided Management of OS Page CachesabstractData-centric applications perform tasks that require intensive data processing and ample memory resources. These tasks have varying I/O access patterns, significantly impacted by the OS cache. Therefore, it is desirable to enable application-specific cache management without compromising memory efficiency. However, infusing user-level policies into the OS cache management is challenging because it is difficult to communicate application-level I/O semantics and access patterns to general-purpose OSs. This article addresses the challenge by enabling applications to safely convey their I/O semantics to OSs via eBPF, allowing for more application-specific control over the OS cache. To this end, we introduce P2Cache , a programmable OS page cache. P2Cache extends the Linux page cache with three new probe points (i.e., eviction , prefetching , and swapping ) that can support application-directed custom policies on OS cache management using eBPF programs. Our experimental results showed that P2Cache significantly enhanced the performance of an LLM inference, a graph processing application, and a database by up to 230%, 49%, and 18%, respectively, with minimal effort. Dusol Lee, Inhyuk Choi, Hyungsoo Jung 0001, Jihong Kim 0001 |
ACM Trans. Storage | 5 |
| 2026 | Processing-in-Memory Architecture for In-Storage Information Retrieval on 3-D NAND Flash Solid-State DrivesabstractRecent advances in retrieval-augmented generation (RAG) highlight retrieval latency, dominated by storage access, and large embedding-index size as the major bottlenecks. Although processing-in-memory (PIM) based on 3-Dnandflash and binary passage retriever (BPR) algorithm offer potential solutions for reducing latency and memory usage, naively integrating BPR algorithm into an SSD based on prior in-memory search (IMS) architectures results in negligible performance gain with substantial hardware overhead. This stems from the fact that prior IMS architectures are optimized only for search, leaving candidate readout misaligned with thenandpage direction and requiring excessive readcycles, while BPR's original top-$L$selection further requires high-precision ADCs and sorting logic. To address these limitations, we propose an algorithm–hardware co-designed in-storage information retrieval architecture. We introduce a hardware-efficient optimized BPR algorithm with a threshold-based selection and two-stage in-storage processing approach to minimize data transfer overheads. We further present an IMS architecture with page-aligned encoding that efficiently supports both search and candidate read operations, eliminating the read cycle penalty. We also present a lightweight in situ error monitoring scheme that bypasses error-correcting code (ECC) decoding while preserving error-aware refresh triggering. Experiments show that our design reduces the amount of data read fromnandflash by 7.94–$1330.40\times $compared with host-side processing baselines and further reduces controller-to-host data transfer, while maintaining the retrieval accuracy. It also reduces total operationcycles by up to 96.93% compared with prior IMS architectures, and lowers error-detection power by 94.71% and total read energy by 17.22% compared with a conventional ECC decoder for error monitoring. Huiwon Yun, Inho Jeong, Kyoyun Lee, Jae Yong Lee 0004, Myoungjun Chun, Jihong Kim 0001, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware RandomizerabstractAlthough NAND flash memory has achieved continuous capacity improvements via advanced 3D stacking and multi-level cell technologies, these innovations introduce new reliability challenges, particularly lateral charge spreading (LCS), absent in low-capacity 2D flash memory. Since LCS significantly increases retention errors over time, addressing this problem is essential to ensure the lifetime of modern SSDs employing high-capacity 3D flash memory. In this paper, we propose a novel data randomizer, STate-Aware Randomizer (STAR), which proactively eliminates the majority of weak data patterns responsible for retention errors caused by LCS. Unlike existing techniques that target only specific worst-case patterns, STAR effectively removes a broad spectrum of weak patterns, significantly enhancing reliability against LCS. By employing several optimization schemes, STAR can be efficiently integrated into the existing I/O datapath of an SSD controller with negligible timing overhead. To evaluate the proposed STAR scheme, we developed a STAR-aware SSD emulator based on characterization results from 160 real 3D NAND flash chips. Experimental results demonstrate that STAR improves SSD lifetime by up to 2.3× and reduces read latency by an average of 50% on real-world traces compared to conventional SSDs. Omin Kwon, Kyungjun Oh, Myungsuk Kim, Jihong Kim 0001 |
ICCAD | 5 |
| 2025 | AiF: Accelerating On-Device LLM Inference Using In-Flash ProcessingabstractWhile large language models (LLMs) achieve remarkable performance across diverse application domains, their substantial memory demands present challenges, especially on personal devices with limited DRAM capacity.Recent LLM inference engines have introduced SSD offloading for model parameters to reduce memory footprint.However, the highly memory-bound nature of on-device LLMs makes inference speed heavily dependent on read bandwidth, leading to significant performance degradation due to the limited bandwidth of SSDs.In this paper, we propose an in-flash processing solution for on-device LLM, called Accelerator-in-Flash (AiF), which integrates matrix-vector multiplication (GEMV) operations directly into flash chips.By enabling in-flash GEMV operations, AiF leverages the high internal bandwidth of flash chips without being constrained by the limited external bandwidth.Building on this core structure, AiF employs two novel flash read techniques that were specifically optimized for reading LLM parameters stored in flash memory.AiF achieves a 4x boost in internal read bandwidth during inference with minimal implementation overhead, thanks to its streamlined error correction process.Evaluations on eight real-world LLMs reveal that AiF provides a 14.6x throughput improvement compared to baseline SSD offloading schemes.Furthermore, AiF surpasses in-memory inference, delivering 1.4x higher throughput with a significantly reduced memory footprint. Jae Yong Lee 0004, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001 |
ISCA | 6 |
| 2025 | DEAR: Improving Performance and Lifetime of SSDs Using Dynamic Error-Aware Refresh
Jae Yong Lee 0004, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001 |
MICRO | 5 |
| 2025 | Program context-assisted address translation for high-capacity SSDs
Xiaochang Li, Minjae Kim 0015, Sungjin Lee 0001, Zhengjun Zhai, Jihong Kim 0001 |
Future Gener. Comput. Syst. | 5 |
| 2025 | DeepPM: Predicting Performance and Energy Consumption of Program Binaries Using TransformersabstractAccurate estimation of performance and energy consumption is critical for optimizing application efficiency on diverse hardware platforms. Traditional methods often rely on profiling and measurements, requiring at least one execution, making them time-consuming and resource-intensive. This article introduces the Deep Power Meter (DeepPM) framework, leveraging deep learning, specifically the Transformer architecture, to predict performance and energy consumption of basic blocks directly from compiled binaries, eliminating the need for explicit measurement processes. The DeepPM model effectively learns the performance and energy consumption of basic blocks, enabling accurate predictions for each. Furthermore, the framework enhances applicability across different ISAs and microarchitectures, addressing limitations of state-of-the-art ML-based techniques restricted to specific processor architectures. Experimental results using the SPEC CPU 2017 benchmark suite show that DeepPM achieves significantly lower prediction errors compared to state-of-the-art ML-based techniques, with a 24% improvement in performance and an 18% improvement in energy consumption for x86 basic blocks, and similar gains for ARM processors. Fine-tuning with minimal data from the Phoronix Test Suite further validates DeepPM’s robustness, achieving an error of approximately 13.7%, close to the fully trained model’s 13.3% error. These findings demonstrate DeepPM’s ability to enhance the accuracy and efficiency of performance and energy consumption predictions, making it a valuable tool for optimizing computing systems across diverse hardware environments. Jun S. Shim, Hyeonji Chang, Yeseong Kim, Jihong Kim 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | RiF: Improving Read Performance of Modern SSDs Using an On-Die Early-Retry EngineabstractModern high-performance SSDs have multiple flash channels operating in parallel to achieve their high I/O bandwidth. However, when the effective bandwidth of these flash channels declines, the SSD's overall bandwidth is substantially impacted. In contemporary SSDs featuring high-density 3D NAND flash memory, frequent invocations of a read-retry procedure pose a significant challenge to fully utilizing the maximum I/O bandwidth of a flash channel. In this paper, we propose a novel read-retry optimization scheme, Retry-in-Flash (RiF), which proactively minimizes the amount of time wasted in conventional read-retry procedures. Unlike existing read-retry solutions that focus on identifying an optimal read-reference voltage for a sensed page, the RiF scheme focuses on determining early on whether a read-retry will be required for the sensed data. To know if a read-retry is needed or not at the earliest possible time, we propose a RiF-enabled flash chip with an on-die early-retry (ODEAR) engine. When the ODEAR engine determines that a sensed page requires a read-retry, a readreference voltage is immediately adjusted and the same page is re-read while ignoring the previously sensed page. By performing the key steps of a read-retry procedure inside a RiF flash chip without transferring the sensed uncorrectable page to an offchip controller, the RiF scheme prevents the read bandwidth of a flash channel from being wasted due to failed read data. To evaluate the RiF scheme, we developed a prototype RiF-enabled flash chip and constructed a RiF-aware SSD simulator using RiF flash chips. Our evaluation results show that the proposed RiF scheme improves the effective SSD bandwidth by 72.1% on average over a state-of-the-art read-retry solution at 2K P/E cycles with negligible power and area overheads. Myoungjun Chun, Jae Yong Lee 0004, Myungsuk Kim, Jisung Park 0001, Jihong Kim 0001 |
HPCA | 5 |
| 2024 | ReadGuard: Integrated SSD Management for Priority-Aware Read Performance DifferentiationabstractWhen multiple apps with different I/O priorities share a high-performance SSD, it is important to differentiate the I/O QoS level based on the I/O priority of each app. In this paper, we study how a modern flash-based SSD should be designed to support priority-aware read performance differentiation. From an in-depth evaluation study using 3D TLC SSDs, we observed that existing FTLs have several weaknesses that need to be improved for better read performance differentiation. In order to overcome the existing FTL weaknesses, we propose ReadGuard , a novel priority-aware SSD management technique that enables an FTL to manage its blocks in a fully read-latency-aware fashion. ReadGuard leverages a new read-latency-centric block quality marker that can accurately distinguish the read latency of a block and ensures that higher-quality blocks are used for higher-priority apps. ReadGuard extends an existing suspend/resume technique to handle collisions among reads. Our experimental results show that a ReadGuard -enabled SSD is effective in supporting differentiated read performance in modern 3D flash SSDs. Myoungjun Chun, Myungsuk Kim, Dusol Lee, Jisung Park 0001, Jihong Kim 0001 |
ACM Trans. Storage | 5 |
| 2023 | Integrated Host-SSD Mapping Table Management for Improving User Experience of Smartphones
Yoona Kim, Inhyuk Choi, Juhyung Park, Jaeheon Lee, Sungjin Lee 0001, Jihong Kim 0001 |
FAST | 6 |
| 2023 | P2Cache: An Application-Directed Page Cache for Improving Performance of Data-Intensive ApplicationsabstractWe propose P2Cache, an application-directed kernel-level page cache that allows an application developer to build a custom kernel-level page cache that matches the I/O characteristics of a target application. P2Cache extends a Linux kernel page cache by adding new probe points that are used to support application-programmable kernel page caches by eBPF programs. Our experimental results show that custom page caches implemented with our P2Cache achieve up to 32% performance improvement in data-intensive graph applications with little effort. Dusol Lee, Inhyuk Choi, Sungjin Lee 0001, Jihong Kim 0001 |
HotStorage | 5 |
| 2022 | TailCut: improving performance and lifetime of SSDs using pattern-aware state encodingabstractAlthough lateral charge spreading is considered as a dominant error source in 3D NAND flash memory, little is known about its detailed characteristics at the storage system level. From a device characterization study, we observed that lateral charge spreading strongly depends on vertically adjacent state patterns and a few specific patterns are responsible for a large portion of bit errors from lateral charge spreading. We propose a new state encoding scheme, called TailCut, which removes vulnerable state patterns by modifying encoded states. By removing vulnerable patterns, TailCut can improve the SSD lifetime and read latency by 80% and 25%, respectively. Jae Yong Lee 0004, Myungsuk Kim, Wonil Choi, Sanggu Lee, Jihong Kim 0001 |
DAC | 5 |
| 2022 | DeepPM: Transformer-based Power and Performance Prediction for Energy-Aware SoftwareabstractMany system-level management and optimization techniques need accurate estimates of power consumption and performance. Earlier research has proposed many high-level/source-level estimation modeling works, particularly for basic blocks. However, most of them still need to execute the target software at least once on a fine-grained simulator or real hardware to extract required features. This paper proposes a performance/power prediction framework, called Deep Power Meter (DeepPM), which estimates them accurately only using the compiled binary. Inspired by the deep learning techniques in natural language processing, we convert the program instructions in the form of vectors and predict the average power and performance of basic blocks based on a transformer model. In addition, unlike existing works based on a Long Short-Term Memory (LSTM) model structure, which only works for basic blocks with a small number of instructions, DeepPM provides highly accurate results for long basic blocks, which takes the majority of the execution time for actual application runs. In our evaluation conducted with SPEC2006 benchmark suite, we show that DeepPM can provide accurate prediction for performance and power consumption with 10.2% and 12.3% error, respectively. DeepPM also outperforms the LSTM-based model by up to 67.2% and 34.9% error for performance and power, respectively. Jun S. Shim, Bogyeong Han, Yeseong Kim, Jihong Kim 0001 |
DATE | 4 |
| 2022 | GuardedErase: Extending SSD Lifetimes by Protecting Weak Wordlines
Duwon Hong, Myungsuk Kim, Geonhee Cho, Dusol Lee, Jihong Kim 0001 |
FAST | 5 |
| 2022 | PiF: in-flash acceleration for data-intensive applicationsabstractTo minimize unnecessary data movements from storage to a host, processing-in-storage (PiS) techniques, which move a compute unit to storage, have been proposed. In this position paper, we propose an extreme version of PiS solutions, called a processing-in-flash (PiF) scheme, that moves computation inside flash chips where data are physically present. As a key building block of a PiF solution, we present a novel flash chip architecture, CoX. Using a prototype PiF SSD based on CoX chips, we demonstrate that PiF-based SSDs are promising in accelerating data-intensive applications. Myoungjun Chun, Jae Yong Lee 0004, Sanggu Lee, Myungsuk Kim, Jihong Kim 0001 |
HotStorage | 5 |
| 2022 | Alohomora: protecting files from ransomware attacks using fine-grained I/O whitelistingabstractWe propose a novel whitelist-based anti-ransomware solution called alohomora. Alohomora is based on our observation that an I/O activity of an application can be an effective abstraction level for managing I/O whitelisting. In alohomora, when a write request is sent to an SSD, its program context value (which is supported by a host CPU register) is passed to the SSD. The SSD checks if the request was pre-approved using the program context value, thus preventing ransomware from modifying files in the SSD. Our experimental results using a prototype alohomora system show that alohomora can achieve a strong security level against sophisticated ransomware attacks without degrading I/O performance. Sanggu Lee, Yoona Kim, Dusol Lee, Inhyuk Choi, Jihong Kim 0001 |
HotStorage | 5 |
| 2021 | Reducing solid-state drive read latency by optimizing read-retryabstract3D NAND flash memory with advanced multi-level cell techniques provides high storage density, but suffers from significant performance degradation due to a large number of read-retry operations. Although the read-retry mechanism is essential to ensuring the reliability of modern NAND flash memory, it can significantly in-crease the read latency of an SSD by introducing multiple retry steps that read the target page again with adjusted read-reference voltage values. Through a detailed analysis of the read mechanism and rigorous characterization of 160 real 3D NAND flash memory chips, we find new opportunities to reduce the read-retry latency by exploiting two advanced features widely adopted in modern NAND flash-based SSDs: 1) the CACHE READ command and 2) strong ECC engine. First, we can reduce the read-retry latency using the advanced CACHE READ command that allows a NAND flash chip to perform consecutive reads in a pipelined manner. Second, there exists a large ECC-capability margin in the final retry step that can be used for reducing the chip-level read latency. Based on our new findings, we develop two new techniques that effectively reduce the read-retry latency: 1) Pipelined Read-Retry (PR²) and 2) Adaptive Read-Retry (AR²). PR² reduces the latency of a read-retry operation by pipelining consecutive retry steps using the CACHE READ command. AR² shortens the latency of each retry step by dynamically reducing the chip-level read latency depending on the current operating conditions that determine the ECC-capability margin. Our evaluation using twelve real-world workloads shows that our proposal improves SSD response time by up to 31.5% (17% on average)over a state-of-the-art baseline with only small changes to the SSD controller. Jisung Park 0001, Myungsuk Kim, Myoungjun Chun, Lois Orosa 0001, Jihong Kim 0001, Onur Mutlu |
ASPLOS | 5 |
| 2021 | RealWear: Improving performance and lifetime of SSDs using a NAND aging marker
Myungsuk Kim, Myoungjun Chun, Duwon Hong, Yoona Kim, Geonhee Cho, Dusol Lee, Jihong Kim 0001 |
Perform. Evaluation | 7 |
| 2021 | Reparo: A Fast RAID Recovery Scheme for Ultra-large SSDsabstractA recent ultra-large SSD (e.g., a 32-TB SSD) provides many benefits in building cost-efficient enterprise storage systems. Owing to its large capacity, however, when such SSDs fail in a RAID storage system, a long rebuild overhead is inevitable for RAID reconstruction that requires a huge amount of data copies among SSDs. Motivated by modern SSD failure characteristics, we propose a new recovery scheme, called reparo , for a RAID storage system with ultra-large SSDs. Unlike existing RAID recovery schemes, reparo repairs a failed SSD at the NAND die granularity without replacing it with a new SSD, thus avoiding most of the inter-SSD data copies during a RAID recovery step. When a NAND die of an SSD fails, reparo exploits a multi-core processor of the SSD controller in identifying failed LBAs from the failed NAND die and recovering data from the failed LBAs. Furthermore, reparo ensures no negative post-recovery impact on the performance and lifetime of the repaired SSD. Experimental results using 32-TB enterprise SSDs show that reparo can recover from a NAND die failure about 57 times faster than the existing rebuild method while little degradation on the SSD performance and lifetime is observed after recovery. Duwon Hong, Keonsoo Ha, Minseok Ko, Myoungjun Chun, Yoona Kim, Sungjin Lee 0001, Jihong Kim 0001 |
ACM Trans. Storage | 7 |
| 2020 | Evanesco: Architectural Support for Efficient Data Sanitization in Modern Flash-Based Storage SystemsabstractAs data privacy and security rapidly become key requirements, securely erasing data from a storage system becomes as important as reliably storing data in the system. Unfortunately, in modern flash-based storage systems, it is challenging to irrecoverably erase (i.e., sanitize) a file without large performance or reliability penalties. In this paper, we propose Evanesco, a new data sanitization technique specifically designed for high-density 3D NAND flash memory. Unlike existing techniques that physically destroy stored data, Evanesco provides data sanitization by blocking access to stored data. By exploiting existing spare flash cells in the flash memory chip, Evanesco efficiently supports two new flash lock commands (pLock and bLock) that disable access to deleted data at both page and block granularities. Since the locked page (or block) can be unlocked only after its data is erased, Evanesco provides a strong security guarantee even against an advanced threat model. To evaluate our technique, we build SecureSSD, an Evanesco-enabled emulated flash storage system. Our experimental results show that SecureSSD can effectively support data sanitization with a small performance overhead and no reliability degradation. Myungsuk Kim, Jisung Park 0001, Genhee Cho, Yoona Kim, Lois Orosa 0001, Onur Mutlu, Jihong Kim 0001 |
ASPLOS | 7 |
| 2019 | RansomBlocker: a Low-Overhead Ransomware-Proof SSDabstractWe present a low-overhead ransomware-proof SSD, called RansomBlocker (RBlocker). RBlocker provides 100% full protections against all possible ransomware attacks by delaying every data deletion until no attack is guaranteed. To reduce storage overheads of the delayed deletion, RBlocker employs a time-out based backup policy. Based on the fact that ransomware must store encrypted version of target files, early deletions of obsolete data are allowed if no encrypted write was detected for a short interval. Otherwise, RBlocker keeps the data for an interval long enough to guarantee no attack condition. For an accurate in-line detection of encrypted writes, we leverages entropy- and CNN-based detectors in an integrated fashion. Our experimental results show that RBlocker can defend all types of ransomware attacks with negligible overheads. Jisung Park 0001, Youngdon Jung, Jonghoon Won, Minji Kang, Sungjin Lee 0001, Jihong Kim 0001 |
DAC | 6 |
| 2019 | Fully Automatic Stream Management for Multi-Streamed SSDs Using Program Contexts
Duwon Hong, Sangwook Shane Hahn, Myoungjun Chun, Sungjin Lee 0001, Joo Young Hwang, Jongyoul Lee, Jihong Kim 0001 |
FAST | 8 |
| 2019 | Exploiting Process Similarity of 3D Flash Memory for High Performance SSDsabstract3D NAND flash memory exhibits two contrasting process characteristics from its manufacturing process. While process variability between different horizontal layers are well known, little has been systematically investigated about strong process similarity (PS) within the horizontal layer. In this paper, based on an extensive characterization study using real 3D flash chips, we show that 3D NAND flash memory possesses very strong process similarity within a 3D flash block: the word lines (WLs) on the same horizontal layer of the 3D flash block exhibit virtually equivalent reliability characteristics. This strong process similarity, which was not previously utilized, opens simple but effective new optimization opportunities for 3D flash memory. In this paper, we focus on exploiting the process similarity for improving the I/O latency. By carefully reusing various flash operating parameters monitored from accessing the leading WL, the remaining WLs on the same horizontal layer can be quickly accessed, avoiding unnecessary redundant steps for subsequent program and read operations. We also propose a new program sequence, called mixed order scheme (MOS), for 3D NAND flash memory which can further reduce the program latency. We have implemented a PS-aware FTL, called cubeFTL, which takes advantage of the proposed techniques. Our evaluation results show that cubeFTL can improve the IOPS by up to 48% over an existing PS-unaware FTL. Youngseop Shim, Myungsuk Kim, Myoungjun Chun, Jisung Park 0001, Yoona Kim, Jihong Kim 0001 |
MICRO | 6 |
| 2019 | File Fragmentation in Mobile Devices: Measurement, Evaluation, and TreatmentabstractMobile devices, such as smartphones, have become a necessity in our daily life. However, users may notice that after being used for a longtime, mobile devices begin to exhibit a sluggish response. Based on an empirical study on a collection of aged smartphones, this work identified that file fragmentation is among the key factors that contribute to the progressive degradation of response time. This study takes a three-step approach: First, this study designed a set of reproducible file-system aging processes based on User-Interface (UI) script replay. Through the aging processes, it confirmed that file fragmentation quickly emerged, and SQLite files were among the most severely fragmented files. Second, based on the workloads of a selection of popular mobile applications, this study observed that file fragmentation did have an impact on user-perceived latencies. Specifically, the launching time of Chrome on an aged file system was 79 percent slower than it was on a pristine file system. Third, this study evaluated existing treatments of file fragmentation, including space preallocation, persistent journal, and file defragmentation to understand their efficacies and limitations. This study also evaluated a state-of-the-art copyless defragmenter, janusd, to show its advantage over the existing methods. Cheng Ji 0002, Li-Pin Chang, Sangwook Shane Hahn, Sungjin Lee 0001, Riwei Pan, Liang Shi 0001, Jihong Kim 0001, Chun Jason Xue |
IEEE Trans. Mob. Comput. | 7 |
| 2018 | SARO: A State-Aware Reliability Optimization Technique for High Density NAND Flash MemoryabstractRecent advances in flash technologies, such as scaling and multi-leveling schemes, have been successful to make flash denser and secure more storage spaces per die. Unfortunately, these technology advances significantly degrade flash's reliability due to a smaller cell geometry and a finer-grained cell state control. In this paper, we propose a state-aware reliability optimization technique SARO), new flash optimization that improves the flash reliability under diverse scaling and multi-leveling schemes. To this end, we first reveal that reliability-related flash errors are highly skewed among flash cell states, which was not captured by prior studies. The proposed SARO exploits then the different per-state error behavior in flash cell states by selecting the most error-prone flash states (for each error type) and by forming narrow threshold voltage distributions(for the selected states only). Furthermore, SARO is applied only when the program time gets shorter because of flash cell aging, thereby keeping the program latency unchanged. Our experimental results with real MLC and TLC flash devices show that SARO can reduce a significant number of flash bit errors, which can in turn reduce the read latency by 40%, on average. Myungsuk Kim, Youngsun Song, Myoungsoo Jung, Jihong Kim 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | PCStream: Automatic Stream Allocation Using Program Contexts
Sangwook Shane Hahn, Sungjin Lee 0001, Joo Young Hwang, Jongyoul Lee, Jihong Kim 0001 |
HotStorage | 6 |
| 2018 | FlashShare: Punching Through Server Storage Stack from Kernel to Firmware for Ultra-Low Latency SSDs
Jie Zhang 0048, Miryeong Kwon, Donghyun Gouk, Sungjoon Koh, Changlim Lee, Mohammad Alian, Myoungjun Chun, Mahmut T. Kandemir, Nam Sung Kim, Jihong Kim 0001, Myoungsoo Jung |
OSDI | 10 |
| 2018 | FastTrack: Foreground App-Aware I/O Management for Improving User Experience of Android Smartphones
Sangwook Shane Hahn, Sungjin Lee 0001, Inhyuk Yee, Donguk Ryu, Jihong Kim 0001 |
USENIX ATC | 5 |
| 2017 | Improving Performance and Lifetime of Large-Page NAND Storages Using Erase-Free Subpage ProgrammingabstractRecent NAND flash devices have large page sizes. Although large pages are useful in increasing the flash capacity, they can degrade both the performance and lifetime of flash storage systems when small writes are dominant. We propose a new NAND programming scheme, called erase-free sub-page programming (ESP), which allows the same page to be programmed multiple times for small writes. By avoiding internal fragmentation, the ESP scheme reduces the overhead of garbage collection for large-page NAND storages. Experimental results show that an ESP-aware FTL can improve the IOPS and lifetime by up to 74% and 177%, respectively. Myungsuk Kim, Sungjin Lee 0001, Jisung Park 0001, Jihong Kim 0001 |
DAC | 5 |
| 2017 | DAC: Dedup-assisted compression scheme for improving lifetime of NAND storage systemsabstractThanks to an aggressive scaling of semiconductor devices, the capacity of NAND flash-based solid-state-drives (SSDs) has increased greatly. However, this benefit comes at the expense of a serious degradation of NAND device's lifetime. In order to improve the lifetime of flash-based SSDs, various data reduction techniques, such as deduplication, lossless compression, and delta compression, are rapidly adopted to SSDs. Although each technique has been extensively studied, efficiently combining these techniques in a synergistic fashion was not thoroughly investigated. In this paper, we propose a novel dedup-assisted compression (DAC) scheme that integrates existing data reduction techniques so that potential benefits of individual ones can be maximized while overcoming their inherent limitations. By doing so, DAC greatly reduces the amount of write traffic sent to SSDs. DAC also requires negligible resources by utilizing existing hardware modules. Our evaluation results show that the proposed DAC decreases the amount of written data by up to 30% over a simple integration of deduplication and lossless compression. Jisung Park 0001, Sungjin Lee 0001, Jihong Kim 0001 |
DATE | 3 |
| 2017 | Improving File System Performance of Mobile Storage Systems Using a Decoupled Defragmenter
Sangwook Shane Hahn, Sungjin Lee 0001, Cheng Ji 0002, Li-Pin Chang, Inhyuk Yee, Liang Shi 0001, Chun Jason Xue, Jihong Kim 0001 |
USENIX ATC | 8 |
| 2017 | Dynamic Erase Voltage and Time Scaling for Extending Lifetime of NAND Flash-Based SSDsabstractThe decreasing lifetime of NAND flash memory, as a side effect of recent advanced semiconductor process scaling, is emerging as one of major barriers to the wide adoption of SSDs in high-performance computing systems. In this paper, we propose Dynamic Erase Voltage and Time Scaling (DeVTS), an integrated approach to extend the lifetime (particularly, endurance) of NAND flash memory. DeVTS is motivated by our key observation that erasing a NAND block with a lower voltage or at a slower speed can significantly improve NAND endurance. However, using a lower erase voltage causes adverse side effects on the write performance and retention capability of NAND flash memory. In order to improve NAND endurance without affecting the other NAND requirements, we take advantage of idle times between write requests and variations of the retention requirement when writing data to a NAND block erased with a lower voltage. We have implemented a DeVTS-aware FTL, called dvsFTL, which exploits the tradeoff relationship between the endurance and erase voltages/times by accurately predicting the write performance and retention requirements. Our experimental results show that dvsFTL can improve NAND endurance by 94 percent, on average, over an existing DeVTS-unaware FTL while all the NAND requirements are preserved. Jaeyong Jeong, Youngsun Song, Sangwook Shane Hahn, Sungjin Lee 0001, Jihong Kim 0001 |
IEEE Trans. Computers | 5 |
| 2016 | Improving performance and lifetime of NAND storage systems using relaxed program sequenceabstractWe propose a new system-level solution that improves both the performance and lifetime of NAND storage systems by exploiting the performance asymmetry of NAND devices. At the device level, we propose a new program sequence, called relaxed program sequence (RPS), which allows more flexible page allocations in a block without compromising NAND reliability. By combining RPS with per-block parity pages, we can improve the write bandwidth and eliminate expensive paired page backup operations. Experimental results show that the proposed technique can increase IOPS by up to 56% and reduce the number of block erasures by up to 30% over an existing RPS-oblivious FTL. Jisung Park 0001, Jaeyong Jeong, Sungjin Lee 0001, Youngsun Song, Jihong Kim 0001 |
DAC | 5 |
| 2016 | Application-Managed Flash
Sungjin Lee 0001, Sang Woo Jun, Shuotao Xu, Jihong Kim 0001, Arvind 0001 |
FAST | 5 |
| 2016 | Improving I/O Resource Sharing of Linux Cgroup for NVMe SSDs on Multi-core Systems
Sungyong Ahn, Kwanghyun La, Jihong Kim 0001 |
HotStorage | 3 |
| 2016 | Effective Lifetime-Aware Dynamic Throttling for NAND Flash-Based SSDsabstractNAND flash-based solid-state drives (SSDs) are increasingly popular in enterprise server systems because of their advantages over hard disk drives such as higher performance and lower power consumption. However, the decreasing write endurance and the unpredictable lifetime remains to be a serious obstacle to their wider adoption in enterprise systems. In this paper, we propose effective lifetime-aware dynamic throttling, called LADY, which guarantees the required storage lifetime by intentionally throttling the write performance of SSDs with consideration of the effective write endurance of NAND flash memory. Unlike existing static throttling, LADY makes throttling decisions based on the characteristics of a workload so that the required SSD lifetime can be guaranteed with less performance degradation. LADY also exploits the improvement on write endurance depending on the NAND program speed and the recovery effects of floating-gate transistors, thereby maximally utilizing the available write endurance of NAND flash while mitigating the decreasing write endurance problem. Our experimental results show that LADY improves write performance by 4.7x with small write response time variations over existing static throttling while guaranteeing the required SSD lifetime. Sungjin Lee 0001, Jihong Kim 0001 |
IEEE Trans. Computers | 2 |
| 2016 | An Integrated Approach for Managing Read Disturbs in High-Density NAND Flash MemoryabstractThe read-disturb problem is emerging as one of the main reliability issues in high-density NAND flash memory. A read-disturb error, which causes data loss, occurs to a page when a large number of reads are performed to its neighboring pages. In this paper, we propose a novel integrated approach for managing the read-disturb problem. Our approach is based on our key observations from the NAND physics that the read disturbance to neighboring pages is a function of the read voltage and the read time. Since the read disturbance has an exponential dependence on the read voltage, lowering the read voltage can improve the read-disturb resistance of a NAND block. By modifying NAND chips to support multiple read modes with different read voltages, our approach allows a flash translation layer module to exploit the tradeoff between the read disturbance and write speed. Since the read disturbance is also proportional to the read time, our approach exploits the difference in the read time among different NAND pages so that frequently read pages can be less intensively read-disturbed using fast page reads. By intelligently relocating read-intensive data to read-disturb resistant blocks and pages, our approach can reduce a large portion of the time overhead from managing read-disturb errors. We also propose a proactive data migration technique which is effective in reducing large variations in I/O response times of the existing on-demand read reclaim (RR) technique. Our experimental results show that our proposed techniques can reduce the execution time overhead by 73% over the existing read-disturb management technique while reducing I/O response time fluctuations during RR activations. Keonsoo Ha, Jaeyong Jeong, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | A Personalized Network Activity-Aware Approach to Reducing Radio Energy Consumption of SmartphonesabstractThe radio energy consumption takes a large portion of the total energy consumption in smartphones. However, a significant portion of radio energy is wasted in a special waiting interval, known as the tail time after a transmission is completed while waiting for a subsequent transmission. In order to reduce the wasted energy in the tail time, the fast dormancy feature allows a quick release of a radio connection in the tail time. For supporting the fast dormancy efficiently, it is important to accurately predict whether a subsequent transmission will occur in the tail time. In this paper, we show that there are strong personal characteristics on how user interacts with a radio network within the tail time. Based on these observations, we propose a novel personalized network activity-aware predictive dormancy technique, called Personalized Diapause (pD). By automatically identifying user-specific tail-time transmission characteristics for various network activities, our proposed technique takes advantages of personalized high-level network usage patterns in deciding when to release radio connections. Our experimental results using real network usage logs from 25 users show that pD can reduce the amount of the wasted tail time energy by 51 percent on average, thus saving the total radio energy consumption by 23 percent with less than 10 percent reconnection increase. Yeseong Kim, Boyeong Jeon, Jihong Kim 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2016 | Exploiting Sequential and Temporal Localities to Improve Performance of NAND Flash-Based SSDsabstractNAND flash-based Solid-State Drives (SSDs) are becoming a viable alternative as a secondary storage solution for many computing systems. Since the physical characteristics of NAND flash memory are different from conventional Hard-Disk Drives (HDDs), flash-based SSDs usually employ an intermediate software layer, called a Flash Translation Layer (FTL). The FTL runs several firmware algorithms for logical-to-physical mapping, I/O interleaving, garbage collection, wear-leveling, and so on. These FTL algorithms not only have a great effect on storage performance and lifetime, but also determine hardware cost and data integrity. In general, a hybrid FTL scheme has been widely used in mobile devices because it exhibits high performance and high data integrity at a low hardware cost. Recently, a demand-based FTL based on page-level mapping has been rapidly adopted in high-performance SSDs. The demand-based FTL more effectively exploits the device-level parallelism than the hybrid FTL and requires a small amount of memory by keeping only popular mapping entries in DRAM. Because of this caching mechanism, however, the demand-based FTL is not robust enough for power failures and requires extra reads to fetch missing mapping entries from NAND flash. In this article, we propose a new flash translation layer called LAST++. The proposed LAST++ scheme is based on the hybrid FTL, thus it has the inherent benefits of the hybrid FTL, including low resource requirements, strong robustness for power failures, and high read performance. By effectively exploiting the locality of I/O references, LAST++ increases device-level parallelism and reduces garbage collection overheads. This leads to a great improvement of I/O performance and makes it possible to overcome the limitations of the hybrid FTL. Our experimental results show that LAST++ outperforms the demand-based FTL by 27% for writes and 7% for reads, on average, while offering higher robustness against sudden power failures. LAST++ also improves write performance by 39%, on average, over the existing hybrid FTL. Sungjin Lee 0001, Dongkun Shin, Young-Jin Kim 0002, Jihong Kim 0001 |
ACM Trans. Storage | 4 |
| 2015 | To collect or not to collect: just-in-time garbage collection for high-performance SSDs with long lifetimesabstractFor NAND flash-based storage systems, managing garbage collection (GC) efficiently is a critical requirement to achieve both high performance and long lifetimes. In this paper, we propose a just-in-time GC technique, called JIT-GC, which invokes background GC operations only when necessary depending on future write demands. JIT-GC was motivated by our measurement study, which strongly suggested that deciding when to invoke background GC operations is a key parameter for efficient GC. By accurately estimating the amount of future SSD writes, JIT-GC can choose the best time to invoke a background GC operation. JIT-GC reserves necessary free space in advance so that high write performance can be achieved while it extends the SSD lifetime by preventing premature block erasures. Our evaluations on real SSDs show that JIT-GC can achieve both high performance and long lifetimes, thus overcoming the shortcomings of existing background GC invocation heuristics. Sangwook Shane Hahn, Jihong Kim 0001, Sungjin Lee 0001 |
DAC | 2 |
| 2014 | Lifetime improvement of NAND flash-based storage systems using dynamic program and erase scaling
Jaeyong Jeong, Sangwook Shane Hahn, Sungjin Lee 0001, Jihong Kim 0001 |
FAST | 4 |
| 2014 | Improving Performance and Capacity of Flash Storage Devices by Exploiting Heterogeneity of MLC Flash MemoryabstractThe multi-level cell (MLC) NAND flash memory technology enables multiple bits of information to be stored in a memory cell, thus making it possible to increase the density of flash memory without increasing the die size. In MLC NAND flash memory, each memory cell can be programmed as a single-level cell or a multi-level cell at runtime because of its performance/capacity asymmetric programming property, which is called flexible programming in this paper. Therefore, MLC flash memory has a potential to achieve the high performance of SLC flash memory while preserving its maximum capacity. In this paper, we present a flexible flash file system, called FlexFS, which takes advantage of flexible programming. FlexFS divides a flash memory medium into SLC and MLC regions, and then dynamically changes two different types of regions to provide an optimal storage solution to end-users in terms of performance and capacity. FlexFS also provides a reasonable storage lifetime by managing the wearing rate of NAND flash memory, which is accelerated by the use of flexible programming. Our implementation of FlexFS in the Linux 2.6 kernel shows that it achieves the I/O performance comparable to SLC flash memory while guaranteeing the capacity of MLC flash memory in various real-world workloads. Sungjin Lee 0001, Jihong Kim 0001 |
IEEE Trans. Computers | 2 |
| 2014 | Personalized optimization for android smartphonesabstractAs a highly personalized computing device, smartphones present a unique new opportunity for system optimization. For example, it is widely observed that a smartphone user exhibits very regular application usage patterns (although different users are quite different in their usage patterns). User-specific high-level app usage information, when properly managed, can provide valuable hints for optimizing various system design requirements. In this article, we describe the design and implementation of a personalized optimization framework for the Android platform that takes advantage of user's application usage patterns in optimizing the performance of the Android platform. Our optimization framework consists of two main components, the application usage modeling module and the usage model-based optimization module. We have developed two novel application usage models that correctly capture typical smartphone user's application usage patterns. Based on the application usage models, we have implemented an app-launching experience optimization technique which tries to minimize user-perceived delays, extra energy consumption, and state loss when a user launches apps. Our experimental results on the Nexus S Android reference phones show that our proposed optimization technique can avoid unnecessary application restarts by up to 78.4% over the default LRU-based policy of the Android platform. Wook Song, Yeseong Kim, Hakbong Kim, Jehun Lim, Jihong Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2013 | An integrated approach for managing the lifetime of flash-based SSDsabstractAs the semiconductor process is scaled down, the endurance of NAND flash memory greatly deteriorates. To overcome such a poor endurance characteristic and to provide a reasonable storage lifetime, system-level endurance enhancement techniques are rapidly adopted in recent NAND flash-based storage devices like solid-state drives (SSDs). In this paper, we propose an integrated lifetime management approach for SSDs. The proposed lifetime management technique combines several lifetime-enhancement schemes, including lossless compression, deduplication, and performance throttling, in an integrated fashion so that the lifetime of SSDs can be maximally extended. By selectively disabling less effective lifetime-enhancement schemes, the proposed technique achieves both high performance and high energy efficiency while meeting the required lifetime. Our evaluation results show that the proposed technique, over the SSDs with no lifetime management schemes, improves write performance by up to 55% and reduces energy consumption by up to 43% while satisfying a 5-year lifetime warranty. Sungjin Lee 0001, Jisung Park 0001, Jihong Kim 0001 |
DATE | 4 |
| 2013 | Improving NAND Endurance by Dynamic Program and Erase Scaling
Jaeyong Jeong, Sangwook Shane Hahn, Sungjin Lee 0001, Jihong Kim 0001 |
HotStorage | 4 |
| 2013 | SOS: Software-based out-of-order scheduling for high-performance NAND flash-based SSDsabstractWe propose an efficient software-based out-of-order scheduling technique, called SOS, for high-performance NAND flash-based SSDs. Unlike an existing hardware-based out-of-order technique, our proposed software-based solution, SOS, can make more efficient out-of-order scheduling decisions by exploiting various mapping information and I/O access characteristics obtained from the flash translation layer (FTL) software. Furthermore, SOS can avoid unnecessary hardware-level operations and manage I/O request rearrangements more efficiently, thus maximizing the multiple-chip parallelism of SSDs. Experimental results on a prototype SSD show that SOS is effective in improving the overall SSD performance, lowering the average I/O response time by up to 42% over a hardware-based out-of-order flash controller. Sangwook Shane Hahn, Sungjin Lee 0001, Jihong Kim 0001 |
MSST | 3 |
| 2013 | BAGC: Buffer-Aware Garbage Collection for Flash-Based Storage SystemsabstractNAND flash-based storage device is becoming a viable storage solution for mobile and desktop systems. Because of the erase-before-write nature, flash-based storage devices require garbage collection that causes significant performance degradation, incurring a large number of page migrations and block erasures. To improve I/O performance, therefore, it is important to develop an efficient garbage collection algorithm. In this paper, we propose a novel garbage collection technique, called buffer-aware garbage collection (BAGC), for flash-based storage devices. The BAGC improves the efficiency of two main steps of garbage collection, a block merge step and a victim block selection step, by taking account of the contents of a buffer cache, which is typically used to enhance I/O performance. The buffer-aware block merge (BABM) scheme eliminates unnecessary page migrations by evicting dirty data from a buffer cache during a block merge step. The buffer-aware victim block selection (BAVBS) scheme, on the other hand, selects a victim block so that the benefit of the buffer-aware block merge is maximized. Our experimental results show that BAGC improves I/O performance by up to 43 percent over existing buffer-unaware schemes for various benchmarks. Sungjin Lee 0001, Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Computers | 3 |
| 2013 | Exploiting Replicated Cache Blocks to Reduce L2 Cache Leakage in CMPsabstractModern chip multiprocessors (CMPs) employ large L2 caches to reduce the performance gap between processors and off-chip memory. However, as the size of an L2 cache increases, its leakage power consumption also becomes a major contributor to the total power dissipation. Managing the leakage power of L2 caches, therefore, is an important issue in realizing low-power CMPs. In CMPs with private L2 caches, each processor makes a copy of the data in its local cache in order to access the data faster, which is called replication. In this paper, we propose a novel leakage management technique that dynamically turns off replications in private L2 caches for leakage power reduction by exploiting two key observations: 1) the cost of an extra cache miss due to the turned-off replication is small because the same cache block exists in another on-chip cache and 2) turning off the replication incurs no extra cache miss if it is invalidated by other processors in order to maintain cache coherence. Since blindly turning off the frequently accessed replications can degrade performance, the proposed technique dynamically controls the number of turned-off replications. The proposed technique can be implemented by slightly modifying the MESI protocol with a new turned-off shared (TOS) coherence state. The TOS state indicates that the corresponding block is shared by other caches but turned off. Experiments on a four-processor CMP with private L2 caches show that the proposed technique reduces the energy consumption of the L2 caches and the main memory by 19.4% on average, with less than 1% performance loss over the existing cache leakage management technique. Hyunhee Kim, Jung Ho Ahn, Jihong Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Lifetime management of flash-based SSDs using recovery-aware dynamic throttling
Sungjin Lee 0001, Kyungho Kim, Jihong Kim 0001 |
FAST | 4 |
| 2012 | FlashBench: A workbench for a rapid development of flash-based storage devicesabstractAs the cell size of NAND flash memory is shrinking, its physical characteristics such as performance and lifetime are significantly degraded. As effective solutions of overcoming such poor physical characteristics, more cross-layer system-level approaches (such as compression and deduplication techniques) are expected to be developed. These system-level techniques typically employ intelligent software algorithms supported by specialized hardware accelerators. Using hardware accelerators combined with sophisticated software algorithms greatly increases the design complexity of flash-based storage devices. However, existing storage design environments are not adequate enough to handle this increased design complexity in a timely and efficient manner. To address this new challenge, we propose a novel storage development environment, called FlashBench, that helps developers to build high-complexity storage solutions quickly. FlashBench is designed to provide a generic framework for the rapid development and validation of storage software/hardware algorithms by supporting multi-level design environments, specifically optimized for seamless hardware/software cross-layer integrations. Our case study demonstrates that FlashBench enables developers to implement high-complexity flash devices with specialized optimization functions in a shorter development time over traditional design environments. Sungjin Lee 0001, Jisung Park 0001, Jihong Kim 0001 |
RSP | 3 |
| 2012 | ARC-H: Adaptive replacement cache management for heterogeneous storage devices
Young-Jin Kim 0002, Jihong Kim 0001 |
J. Syst. Archit. | 2 |
| 2011 | A leakage-aware L2 cache management technique for producer-consumer sharing in low-power chip multiprocessors
Hyunhee Kim, Jihong Kim 0001 |
J. Parallel Distributed Comput. | 2 |
| 2010 | Replication-aware leakage management in chip multiprocessors with private L2 cacheabstractPower dissipation has become a critical issue in modern chip multiprocessors (CMPs). Managing the leakage power of their L2 caches is particularly important in realizing low-power CMPs because most CMPs employ large L2 caches to hide the performance gap between processors and an off-chip memory while leakage power becomes a major portion in the total power dissipation of CMPs as process technology advances below 90 nm. We propose a replication-aware leakage management technique that selectively turns off a replicated block in a private L2 cache for leakage power reduction. Once a cache line is turned off, the data is lost, but its tag maintains the coherence state. The cost of an extra cache miss due to the turned-off replication is limited since the data of the cache line exists in another on-chip cache. Furthermore, the replicated block incurs no overhead if it is invalidated by other processors in order to maintain cache coherence. Our proposed technique can be implemented by slightly modifying the MESI protocol with a new turned-off shared coherence state. This state indicates that the corresponding block is shared by other caches but turned off. Experiments on a 4 processor CMP with private L2 caches show that the proposed technique reduces the energy consumption of the L2 caches and main memory by 20.0% on average without introducing significant performance loss over the existing cache leakage management technique. Hyunhee Kim, Jung Ho Ahn, Jihong Kim 0001 |
ISLPED | 3 |
| 2009 | FlexFS: A Flexible Flash File System for MLC NAND Flash Memory
Sungjin Lee 0001, Keonsoo Ha, Kangwon Zhang, Jihong Kim 0001 |
USENIX ATC | 4 |
| 2009 | Broadcast filtering: Snoop energy reduction in shared bus-based low-power MPSoCs
Chun-Mok Chung, Jihong Kim 0001 |
J. Syst. Archit. | 2 |
| 2009 | Reusability-aware cache memory sharing for chip multiprocessors with private L2 caches
Hyunhee Kim, Sungjun Youn, Jihong Kim 0001 |
J. Syst. Archit. | 3 |
| 2007 | Reducing snoop-energy in shared bus-based mpsocs by filtering useless broadcastsabstractIn shared bus-based multiprocessor system-on-a-chips (MPSoCs), snoop-based schemes are widely used to maintain cache coherency. However, many of broadcasts are useless because remote caches seldom have the matching blocks and their tag lookups do not supply data. From the energy perspective, such tag lookups consume unnecessary energy and make the system energy wasteful. In this paper, we propose a broadcast filtering technique to reduce snoop-energy in both of cache and bus. Broadcast filtering is achieved by help of snooping cache and split-bus. The snooping cache checks if matching blocks exist in remote caches before broad casting a coherency request. If no remote cache has the matching block, it eliminates the broadcast. If broadcasting is necessary, only a part of split-bus is used so that the request is selectively broadcasted only to the remote caches that have matching blocks. Simulation results show that our technique reduces 90%, 50%, and 30% of cache lookups, bus usage, and snoop-energy, respectively, with only 2% of degradation in performance. Our technique reduces more energy than other state-of-the-art techniques. Chun-Mok Chung, Jihong Kim 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | A reusability-aware cache memory sharing technique for high-performance low-power CMPs with private L2 cachesabstractChip multiprocessors (CMPs) emerge as a dominant architectural alternative in high-end embedded systems. Since off-chip accesses require a long latency and consume a large amount of power, CMPs are typically based on multiple levels of on-chip cache memories. To meet the performance demand and power budget, an efficient support for memory hierarchy is important. We propose an on-chip L2 cache organization which takes advantage of both a private L2 cache and a shared L2 cache to improve the performance and reduce energy consumption. Our L2 cache organization is based on a private L2 cache organization which has the short access latency. When a cache block in the private L2 cache is selected for an eviction, our proposed organization first evaluates the reusability of the cache block. If the cache block is likely to be reused, we save the evicted cache block in one of peer L2 caches which may have efficiently invalid blocks. By selectively writing evicted cache blocks to peer L2 caches, the proposed L2 cache organization can effectively simulate a shared L2 cache. Experimental results using a CMP simulator showed that the proposed L2 cache organization improved the average memory latency by up to 27% and reduced energy consumption by up to 16.6% over a 256KB private L2 cache organization for the SPLASH2 benchmark programs.. Sungjun Youn, Hyunhee Kim, Jihong Kim 0001 |
ISLPED | 3 |
| 2007 | Optimizing Intratask Voltage Scheduling Using Profile and Data-Flow InformationabstractIntratask dynamic-voltage scheduling (IntraDVS), which adjusts the supply voltage within an individual-task boundary, has been introduced as an effective technique for developing low-power single-task applications or low-power multitask applications, where a small number of tasks are dominant in total execution time. The original IntraDVS technique used the remaining worst case execution cycles, and the control-flow information to identify the voltage-scaling points (VSPs) of a program. In this paper, two kinds of improvement techniques enhancing the energy performance of the IntraDVS are proposed. One is to use profile information to optimize the voltage schedule for the remaining average-case execution path (RAEP-IntraDVS). The other is to use data-flow information to optimize the locations of VSPs [look-ahead IntraDVS (LaIntraDVS)]. The experimental results show that the RAEP-IntraDVS can reduce the energy consumption by 20% on average and the LaIntraDVS can reduce the energy consumption by 40%-45% compared with the original IntraDVS Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Energy-efficient file placement techniques for heterogeneous mobile storage systemsabstractWhile hard disk drives are the most common secondary storage devices, their high power consumption and low shock-resistance limit them as an ideal mobile storage solution. On the other hand, flash memory devices overcome the main problems of hard disk drives, but they are still more expensive in the cost per bit over hard disk drives and can only support a limited number of erase cycles. In this paper, we show that combining the merits of a hard disk and a flash memory device can produce an energy-efficient secondary storage solution for mobile platforms. We propose an energy-efficient file placement technique for such heterogeneous storage systems. The proposed technique adapts an existing data concentration technique by separating read and write I/O requests. Experimental results show that the proposed technique reduces the energy consumption by up to 74.5% when the combination of a 1.8" disk and a flash memory is used instead of a single 2.5" disk, at the cost of small increase in the average response time. Young-Jin Kim 0002, Kwon-Taek Kwon, Jihong Kim 0001 |
EMSOFT | 3 |
| 2006 | Operating System Support for Procedural Abstraction in Embedded SystemsabstractProcedural abstraction reduces code size by replacing repeated code fragments with call instructions to a subroutine that executes the repeated fragment. However, in order to build a subroutine, extra instructions are necessary to support the procedural call mechanism. In this paper, we present an operating system level technique which improves the space efficiency of a procedural abstraction-based code compaction technique. The call-related extra instructions are not used in the proposed technique because operating system routines implicitly supports the procedure call and return. The proposed technique consists of three execution modes including one applicable to ROM-based systems. The experimental results show the proposed technique reduces the code size significantly while increasing the execution time slightly Keun Soo Yim, Jeong-Joon Yoo, Jae Don Lee, Jihong Kim 0001 |
RTCSA | 4 |
| 2006 | Dynamic voltage scaling of mixed task sets in priority-driven systemsabstractThis paper describes dynamic voltage scaling (DVS) algorithms for real-time systems with both periodic and aperiodic tasks. Although many DVS algorithms have been developed for real-time systems with periodic tasks, none of them can be used for a system with both periodic and aperiodic tasks because of the arbitrary temporal behaviors of aperiodic tasks. This paper proposes off-line and on-line DVS algorithms that are based on existing DVS algorithms. The proposed algorithms utilize the execution behaviors of scheduling servers for aperiodic tasks. Since there is a tradeoff between the energy consumption and the response time of aperiodic tasks, the proposed algorithms focus on bounding the response time degradation of aperiodic tasks although they delay the response time by stretching the task execution to get high energy savings in mixed task sets. Experimental results show that the proposed algorithms reduce the energy consumption by 48% and 35% over the non-DVS scheme under rate monotonic (RM) scheduling and earliest deadline first (EDF) scheduling, respectively. Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Optimizing intra-task voltage scheduling using data flow analysisabstractIntra-task voltage scheduling (IntraDVS), which adjusts the supply voltage within an individual task boundary, is an effective technique for developing low-power applications. In IntraDVS, slack times are estimated by analyzing program's control flow information. In this paper, we propose an optimization technique for IntraDVS using data flow information. The proposed algorithm improves the energy efficiency by moving the voltage scaling points to earlier instructions based on the analysis results of program's data flow. The experimental results using an MPEG-4 encoder program show that the proposed algorithm reduces the energy consumption by 40-45% over the original IntraDVS algorithm. Dongkun Shin, Jihong Kim 0001 |
ASP-DAC | 2 |
| 2005 | Exploration of Memory-Aware Dynamic Voltage Scheduling for Soft Real-Time ApplicationsabstractDynamic voltage scaling (DVS) and dynamic power management (DPM) are widely-used techniques to reduce energy consumption in modern computing systems. Although combining these techniques can save more energy, there has not been much work focused on energy-optimal combination of these techniques under variable memory clock frequencies. In this paper, we explore system-wide energy-optimal frequency space for a memory-aware DVS technique based on a stochastic memory access model for systems with frequency-variable memory devices. In addition, we propose a simple but practical DVS method, which can be applied to an actual platform. Young-Jin Kim 0002, Jihong Kim 0001 |
RTCSA | 2 |
| 2005 | Intra-task voltage scheduling on DVS-enabled hard real-time systemsabstractThis paper proposes a novel intra-task dynamic voltage scheduling (IntraDVS) framework for low-energy hard real-time applications. Based on a static timing analysis technique, the proposed approach controls the supply voltage within an individual task boundary. By fully exploiting all the slack times, a scheduled program by the proposed technique always completes its execution near the deadline, thus achieving a high energy reduction ratio. The problem formulation of IntraDVS is first presented and two heuristics are proposed: one based on worst-case execution information and the other on average-case execution information. In order to validate the effectiveness of the proposed heuristics, a software tool that automatically converts a DVS-unaware program into an equivalent low-energy program was built. In an experiment on a DVS-enabled system, the low-energy version of a Moving Pictures Expert Group (MPEG)-4 encoder/decoder consumed only 35%-51% of the energy consumption of the original program running on a fixed-voltage system with a power-down mode. The energy efficiency of the IntraDVS algorithms was also compared with that of task-level voltage scheduling algorithms. The experimental results show that the IntraDVS algorithm can be useful in multitask environments as well. Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Dynamic voltage scaling of periodic and aperiodic tasks in priority-driven systems
Dongkun Shin, Jihong Kim 0001 |
ASP-DAC | 2 |
| 2004 | Power-Aware Scheduling of Mixed Task Sets in Priority-Driven Systems
Dongkun Shin, Jihong Kim 0001 |
EUC | 2 |
| 2004 | An Energy-Efficient Routing and Reporting Scheme to Exploit Data Similarities in Wireless Sensor Networks
Keun Soo Yim, Jihong Kim 0001, Kern Koh |
EUC | 2 |
| 2004 | Preemption-aware dynamic voltage scaling in hard real-time systemsabstractDynamic voltage scaling (DVS) is a well-known low-power design technique for embedded real-time systems. Because of its effectiveness on energy reduction, several variable voltage processors have been developed and many DVS algorithms targeting these processors have been proposed. However, most existing DVS algorithms focus on reducing the energy consumption of CPU only, ignoring their negative impacts on task scheduling and system wide energy consumption. In this paper, we address one of such side effects, an increase in task preemptions due to DVS. We present two preemption control techniques which can reduce the number of task preemptions of DVS algorithms. Experimental results show that the delayed-preemption technique is effective in reducing the number of preemptions incurred by DVS algorithms while achieving a high energy efficiency. Woonseok Kim, Jihong Kim 0001, Sang Lyul Min |
ISLPED | 2 |
| 2004 | A Space-Efficient On-Chip Compressed Cache Organization for High Performance Computing
Keun Soo Yim, Jang-Soo Lee, Jihong Kim 0001, Shin-Dug Kim, Kern Koh |
ISPA | 3 |
| 2003 | A low-power image convolution algorithm for variable voltage processorsabstractWe describe a low-power image convolution algorithm for variable voltage processors. The algorithm takes advantages of common properties of popular kernels. Unlike a direct algorithm of convolution operation where the dynamic voltage scaling (DVS) feature of variable voltage processors cannot be used, our algorithm modifies the sequence of computing convolution sums so that DVS can be effectively utilized. Our implementation on Itsy, a DVS research platform from Compaq, shows the energy saving of up to 71% over that of the direct algorithm without any performance degradation. Hyugjin Kwon, Jihong Kim 0001 |
ICASSP (2) | 2 |
| 2003 | Dynamic voltage scaling algorithm for fixed-priority real-time systems using work-demand analysisabstractDynamic Voltage Scaling (DVS), which adjusts the clock speed and supply voltage dynamically, is an effective technique in reducing the energy consumption of embedded real-time systems. Unlike dynamic-priority real-time scheduling for which highly effective DVS algorithms are available, existing fixed-priority DVS algorithms are less effective in energy efficiency because they are based on inefficient slack estimation methods. This paper describes an efficient on-line slack estimation heuristic for the rate-monotonic (RM) scheduling. The proposed heuristic estimates the slack times using the short term work-demand analysis. The DVS algorithm based on the proposed heuristic is also presented. Experimental results show that the proposed DVS algorithm reduces the energy consumption by 25∼42% over the existing rate-monotonic DVS algorithms. Woonseok Kim, Jihong Kim 0001, Sang Lyul Min |
ISLPED | 2 |
| 2003 | Power-aware scheduling of conditional task graphs in real-time multiprocessor systemsabstractWe propose a novel power-aware task scheduling algorithm for DVS-enabled real-time multiprocessor systems. Unlike the existing algorithms, the proposed DVS algorithm can handle conditional task graphs (CTGs) which model more complex precedence constraints. We first propose a condition-unaware task scheduling algorithm integrating the task ordering algorithm for CTGs and the task stretching algorithm for unconditional task graphs. We then describe a condition-aware task scheduling algorithm which assigns to each task the start time and the clock speed, taking account of the condition matching and task execution profiles. Experimental results show that the proposed condition-aware task scheduling algorithm can reduce the energy consumption by 50% on average over the non-DVS task scheduling algorithm. Dongkun Shin, Jihong Kim 0001 |
ISLPED | 2 |
| 2003 | Application-driven network capacity adaptation for energy efficient ad-hoc networksabstractIn this paper, we propose a novel technique for reducing the energy consumption in idle mode. Our proposed technique called the network capacity adaptation derives an effective network capacity required by a given application and logically slows down the network capacity when the traffic load is low. By slowing down the logical network capacity, the neighboring mobile stations can stay longer in the power-off mode, significantly saving the energy consumed in idle mode. Yongho Seok, Nakjung Choi, Yanghee Choi, Jihong Kim 0001 |
PIMRC | 4 |
| 2003 | On energy-optimal voltage scheduling for fixed-priority hard real-time systemsabstractWe address the problem of energy-optimal voltage scheduling for fixed-priority hard real-time systems, on which we present a complete treatment both theoretically and practically. Although most practical real-time systems are based on fixed-priority scheduling, there have been few research results known on the energy-optimal fixed-priority scheduling problem. First, we prove that the problem is NP-hard. Then, we present a fully polynomial time approximation scheme (FPTAS) for the problem. For any ε > 0, the proposed approximation scheme computes a voltage schedule whose energy consumption is at most (1 + ε) times that of the optimal voltage schedule. Furthermore, the running time of the proposed approximation scheme is bounded by a polynomial function of the number of input jobs and 1/ε. Given the NP-hardness of the problem, the proposed approximation scheme is practically the best solution because it can compute a near-optimal voltage schedule (i.e., provably arbitrarily close to the optimal schedule) in polynomial time. Experimental results show that the approximation scheme finds more efficient (almost optimal) voltage schedules faster than the best existing heuristic. Han-Saem Yun, Jihong Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2002 | A Dynamic Voltage Scaling Algorithm for Dynamic-Priority Hard Real-Time Systems Using Slack Time AnalysisabstractDynamic voltage scaling (DVS), which adjusts the clock speed and supply voltage dynamically, is an effective technique in reducing the energy consumption of embedded real-time systems. The energy efficiency of a DVS algorithm largely depends on the performance of the slack estimation method used in it. In this paper, we propose a novel DVS algorithm for periodic hard real-time tasks based on an improved slack estimation algorithm. Unlike the existing techniques, the proposed method takes full advantage of the periodic characteristics of the real-time tasks under priority-driven scheduling such as EDF. Experimental results show that the proposed algorithm reduces the energy consumption by 20/spl sim/40 % over the existing DVS algorithm. The experiment results also show that our algorithm based on the improved slack estimation method gives comparable energy savings to the DVS algorithm based on the theoretically optimal (but impractical) slack estimation method. Woonseok Kim, Jihong Kim 0001, Sang Lyul Min |
DATE | 2 |
| 2002 | Optimal software pipelining of loops with control flowsabstractSoftware pipelining is widely used as a compiler optimization technique to achieve high performance in machines that exploit instruction -level parallelism. However, surprisingly, there have been few theoretical or empirical results on optimal software pipelining of loops with control flows. In this paper, we present three new contributions for this under-investigated problem. First, we propose a necessary and sufficient condition for a loop with control flows to have an optimally software-pipelined program. We also present a decision procedure to compute the condition. Second, we present two software pipelining algorithms. The first algorithm computes an optimal solution for every loop satisfying the condition, but may run in exponential time. The second algorithm computes optimal solutions efficiently for most (but not all) loops satisfying the condition. Third, we present experimental results which strongly indicate that achieving the optimality in the software-pipelined programs is a viable goal in practice with realistic hardware support. Han-Saem Yun, Jihong Kim 0001, Soo-Mook Moon |
ICS | 2 |
| 2001 | A First Step Towards Time Optimal Software Pipelining of Loops with Control Flows
Han-Saem Yun, Jihong Kim 0001, Soo-Mook Moon |
CC | 2 |
| 2001 | Low-Energy Intra-Task Voltage Scheduling Using Static Timing AnalysisabstractWe propose an intra-task voltage scheduling algorithm for low-energy hard real-time applications. Based on a static timing analysis technique, the proposed algorithm controls the supply voltage within an individual task boundary. By fully exploiting all the slack times, a scheduled program by the proposed algorithm always complete its execution near the deadline, thus achieving a high energy reduction ratio. In order to validate the effectiveness of the proposed algorithm, we built a software tool that automatically converts a DVS-unaware program into an equivalent low-energy program. Experimental results show that the low-energy version of an MPEG-4 encoder/decoder (converted by the software tool) consumes less than 7$\sim$25% of the original program running on a fixed-voltage system with a power-down mode. Dongkun Shin, Jihong Kim 0001, Seongsoo Lee |
DAC | 2 |
| 2001 | An operation rearrangement technique for power optimization in VLIM instruction fetchabstractIn VLIW machines where a single instruction contains multiple operations, the power consumption during instruction fetches varies significantly depending on how the operations are arranged within the instruction. In this paper we describe a post-pass operation rearrangement method that reduces the power consumption from the instruction-fetch datapath. The proposed method modifies operation placement orders within VLIW instructions so that the switching activity between successive instruction fetches is minimized. Our experiment shows that the switching activity can be reduced by 34% on average for benchmark programs. Dongkun Shin, Jihong Kim 0001, Naehyuck Chang |
DATE | 2 |
| 2001 | A profile-based energy-efficient intra-task voltage scheduling algorithm for real-time applicationsabstractArticle Share on A profile-based energy-efficient intra-task voltage scheduling algorithm for real-time applications Authors: Dongkun Shin School of Computer Science and Engineering, Seoul National University School of Computer Science and Engineering, Seoul National UniversityView Profile , Jihong Kim School of Computer Science and Engineering, Seoul National University School of Computer Science and Engineering, Seoul National UniversityView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 271–274https://doi.org/10.1145/383082.383162Online:06 August 2001Publication History 24citation267DownloadsMetricsTotal Citations24Total Downloads267Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Dongkun Shin, Jihong Kim 0001 |
ISLPED | 2 |
| 2001 | Power-aware modulo scheduling for high-performance VLIW processorsabstractFor high-performance processors, the step power and peak power, which are closely related to the chip reliability, are important design constraints, often more than the average power. In VLIW processors where a single instruction may contain a variable number of operations, the step power and peak power vary significantly depending on the parallel schedule generated by a parallelizing compiler. In this paper, we propose a power-aware modulo scheduling algorithm for high-performance VLIW processors. The proposed algorithm reduces both the step power and peak power by producing a more balanced parallel schedule while not compromising performance. Experimental results show that the proposed scheduling technique significantly improves the power characteristics of highperformance processors over an existing power-unaware modulo scheduling technique. 1. Han-Saem Yun, Jihong Kim 0001 |
ISLPED | 2 |
| 1998 | Performance evaluation of register allocator for the advanced DSP of TMS320C80abstractPPCA is an assembly language-level register allocator and instruction compactor for the advanced DSPs (ADSPs) of the TMS320C80 digital signal processor. It was developed to help the implementation of time-critical ADSP assembly programs which heavily utilize powerful ADSP features optimized for multimedia and image computing applications for maximum efficiency. PPCA takes as an input ADSP assembly operations with symbolic variables. It then allocates the ADSP's physical registers to the symbolic variables and rearranges the operations into a highly-parallelized compact format. In this paper, we have evaluated the performance of a register allocation capability of PPCA using an extensive image computing library for the TMS320C80. We present the basic algorithm of the PPCA's register allocation module and describe the performance evaluation approach used. The result shows that PPCA essentially achieves optimal register allocation for the test cases based on the image computing library functions. Jihong Kim 0001, Graham Short |
ICASSP | 1 |
| 1998 | A Worst Case Timing Analysis Technique for Multiple-Issue MachinesabstractWe propose a worst case timing analysis technique for in-order multiple-issue machines. In the proposed technique, timing information for each program construct is represented by a directed acyclic graph (DAG) that shows dependences among instructions in the program construct. From this information, we derive for each pair of instructions the distance bounds between their issue times. Using these distance bounds, we identify the sets of instructions that can be issued at the same time. Deciding such instructions is an essential task in reasoning about the timing behavior of multiple-issue machines. In order to reduce the complexity of analysis, the distance bounds are progressively refined through a hierarchical analysis over the program syntax tree in a bottom-up fashion. Our experimental results show that the proposed technique can predict the worst case execution times for in-order multiple-issue machines as accurately as ones for simpler RISC processors. Sung-Soo Lim, Jung Hee Han, Jihong Kim 0001, Sang Lyul Min |
RTSS | 3 |
| 1995 | Efficient 2-D Convolution Algorithm with the Single-Data Multiple Kernel Approach
Jihong Kim 0001, Yongmin Kim 0001 |
CVGIP Graph. Model. Image Process. | 1 |