EDBT 2026 Demo / reviewers in the wild / expert
Changhee Jung
dblp:85/2308
· DBLP profile ↗
58ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0002-6422-6549ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 3 first-author · 27 since 2021Software engineering, systems software and programming languages · 16 · 2 first-author · 7 since 2021Security and privacy · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intermittence-Aware Cache CompressionabstractCache compressions are proven effective in improving the performance of caches in conventional processors. They compress data into a smaller size, allowing caches to accommodate more blocks. This helps reduce cache misses and expensive memory accesses, ultimately improving performance. However, conventional cache compression is less effective for energy harvesting systems (EHSs) which experience frequent power failure, as many compressed blocks end up not being used before their loss upon power outage. This wastes hard-won energy, which would otherwise be used for making more program progress. To address this issue, this paper introduces Kagura, an adaptive cache compression extension with frequent power failure in mind. Specifically, Kagura disables cache compression when it finds out that many cached blocks are unlikely to be reused before the next power outage. That way, Kagura avoids the energy waste on useless compressions/decompressions, and the resulting speedup is on par with the ideal intermittenceaware cache compressor. Experimental results show that when combined with an existing cache compressor, Kagura reduces the total energy consumption by an average of$\mathbf{4. 5 3 \%}$(up to 16.21 %) and improves the performance by an average of 4.74 % (up to$\mathbf{1 7. 8 7 \%}$) compared to the baseline EHS without cache compression. Gan Fang, Jianping Zeng 0001, Yuchen Zhou 0005, Changhee Jung |
HPCA | 4 |
| 2026 | Anchoring Whole-System Persistence and Resilience in CXLabstractThis paper presents ANCHOR, a novel CXL-based memory hierarchy that achieves both persistence and resilience without recompilation or core modifications. For performant whole-system persistence (WSP), ANCHOR introduces a dual-path store architecture in which every committed store is sent to both the L1 cache and a CXL-attached memory-semantic SSD. To reduce the SSD write traffic, ANCHOR employs an STT-RAM write-combining buffer that coalesces stores before they reach the SSD. In particular, the SSD-resident committed store data serves as a ground-truth anchor for error repair, where ANCHOR repurposes cache and DRAM ECCs as detection codes and repairs detected errors using the corresponding SSD-resident clean copy. This approach effectively upgrades cache-side end-to-end reliability from intrinsic 1-bit correction to 3-bit correction, while also extending DRAM repair coverage beyond the limits of standard ChipKill—all at no additional ECC cost. Evaluation across 42 benchmarks shows that ANCHOR incurs only a 4.36% average run-time overhead while ensuring both WSP and system-wide resilience. Yuchen Zhou 0005, Jianping Zeng 0001, Changhee Jung |
ICS | 3 |
| 2026 | Intermittence-Aware Speculative Page Coloring for Secure NVM
Jongouk Choi, Junyeong Park, Nicholas L'Heureux, Yan Solihin, Hyunwoo Joe, Changhee Jung |
ISCA | 6 |
| 2026 | Joint Optimization of Continuous Variables and Priority Assignments for Real-Time Systems with Black-Box Schedulability ConstraintsabstractIn real-time systems optimization, designers often face a challenging problem posed by the non-convex and non-continuous schedulability conditions, which may even lack an analytical form to understand their properties. To tackle this challenging problem, we treat the schedulability analysis as a black box that only returns true/false results. We propose a general and scalable framework to optimize real-time systems with continuous variables, named Numerical Optimizer with Real-Time Highlight (NORTH). NORTH is built upon the gradient-based active-set methods from the numerical optimization literature but with new methods to manage active constraints for the non-differentiable schedulability constraints. In addition, we also generalize NORTH to NORTH+ to collaboratively optimize priority assignments, a common type of discrete variables, with continuous variables based on numerical optimization algorithms. We demonstrate the algorithm performance with two example applications: energy minimization based on dynamic voltage and frequency scaling (DVFS), and optimization of control system performance. In these experiments, NORTH are 10 2 to 10 5 times faster than state-of-the-art methods while maintaining similar or better solution quality. NORTH+ outperforms NORTH by 30% with similar algorithm scalability. Both NORTH and NORTH+ support black-box schedulability analysis, ensuring broad applicability. Sen Wang 0014, Dong Li 0035, Shao-Yu Huang, Xuanliang Deng, Ashrarul H. Sifat, Changhee Jung, Ryan K. Williams, Haibo Zeng 0001 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | Rethinking Dead Block Prediction for Intermittent ComputingabstractExisting dead block predictors have proven to be effective in reducing cache leakage power of conventional systems. However, prior work is significantly less effective in energy harvesting systems in that it does not take into account their unique characteristic, i.e., frequent power failure during program execution. Even if some cache blocks are predicted to be live, they may not be used due to their loss upon power failure. In response, this paper introduces EDBP, an extension to existing dead block predictors, to enhance their performance in various energy harvesting environments. EDBP can identify and deactivate those cache blocks that are not reused before upcoming power failure-though they are considered live by the existing predictor-thereby lowering cache leakage and preserving more energy for forward execution progress. Experimental results show that for 20 applications from Mediabench and MiBench, EDBP alone reduces total energy consumption by 6.5% and improves performance by 6.9% compared to the baseline with no dead block predictor. When combined with a conventional dead block predictor (Cache Decay), EDBP achieves 9.8% reduction in total energy consumption—approaching the theoretical minimum—and 11.9% performance improvement. Gan Fang, Changhee Jung |
HPCA | 2 |
| 2025 | Rethinking Prefetching for Intermittent ComputingabstractPrefetching improves performance by reducing cache misses.However, conventional prefetchers are too aggressive to serve batteryless energy harvesting systems (EHSs) where energy efficiency is the utmost design priority due to weak input energy and the resulting frequent outages.To this end, this paper proposes IPEX, an Intermittence-aware Prefetching EXtension that can be integrated into existing prefetchers on EHSs.IPEX aims to avert useless prefetches by suppressing the prefetching of the cache blocks receiving no hit before their loss on power failure, which would otherwise waste harvested energy.At a proper moment before an upcoming outage, IPEX throttles the prefetch degree to target only those blocks that are likely to be used before the outage.That way IPEX saves energy and spends it on making further execution progress.Experimental results show that on average, IPEX reduces energy consumption by 7.86% (up to 21.64%) and improves performance by 8.96% (up to 23.49%) compared to a conventional prefetcher. Gan Fang, Jianping Zeng 0001, Aditya Gupta 0012, Changhee Jung |
ISCA | 4 |
| 2025 | WarmCache: Exploiting STT-RAM Cache for Low-Power Intermittent SystemsabstractThis paper introduces WarmCache, an optimized STT-RAM cache design with relaxed non-volatility, for an energy harvesting system (EHS) to avoid such compulsory misses across power failure.The key insight is that if the retention time of a cache is longer than a power outage period, the cache contents can be preserved, thereby preventing compulsory misses.Based on this insight, WarmCache leverages a STT-RAM cache with reduced thermal stability to preserve non-volatility during power outages while not persisting any data.To mitigate retention failure that may occur in the relaxed STT-RAM cache, WarmCache lets its compiler partition program into a series of regions and conducts region-level error correction.At each region boundary, WarmCache verifies the execution of the region by scrubbing updated cache lines and re-executes it if any multi-bit error is detected therein.For optimization, WarmCache compiler introduces a novel region formation technique that adjusts the size of each region to match the scrubbing interval.This is achieved through region stitching for combining shorter regions and region splitting for dividing longer regions.Our experiments demonstrate that WarmCache manages to avoid compulsory cache misses and improves the performance by 1.3∼1.4x on average compared to the state-of-the-art cache design for EHS. Noureldin Hassan, Byounguk Min, Changhee Jung, Yan Solihin, Jongouk Choi |
ISCA | 3 |
| 2025 | A Novel Efficient Crash Consistency Solution Enabling Rollback Recovery for Secure NVM in Low-Power Energy Harvesting SystemsabstractEnergy Harvesting Systems (EHSs) frequently suffer power failures and are particularly deployed in remote and open environments where physical access attacks on Non-volatile Memories (NVMs) are practical. However, prior crash consistency solutions for secure NVM were designed only for conventional power-rich systems with the assumption that enough power is steadily supplied. Moreover, the prior solutions rely on roll-forward recovery and cause a significant performance overhead in low-power EHSs. To achieve a low-cost and high-performance crash-consistent secure NVM working on low-power EHSs, this paper presents Milestone, the first efficient crash consistency solution that introduces a novel hybrid checkpoint mechanism to enable a rollback recovery for secure NVM working in frequent power failures.The hybrid checkpointing atomically (1) undo-logs data updates from program writes and (2) redo-logs the updates of security metadata associated with the data updates when an adaptive hardware timer expires. In particular, Milestone discovers an optimized eager update method for the security metadata that can be performed in parallel with the program writes to NVM by leveraging the rollback recovery. Our experimental results demonstrate that Milestone significantly outperforms the state-of-the-art roll-forward recovery-based solution for secure NVM running on low-power EHSs, achieving up to a 1.87x speedup, on average. Youngkwang Han, Jongouk Choi, Kazi Abu Zubair, Amro Awad, Changhee Jung, Brent ByungHoon Kang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Adaptive Computing in Memory Meets Conventional Batteryless PlatformsabstractComputing In-Memory (CIM) with emerging nonvolatile memory (NVM) technologies is promising for batteryless systems since it removes the need for explicit backup and energy-hungry data transfer between the processor and memory. However, existing CIM solutions are not effective in accelerating memory-bound inference tasks efficiently on batteryless systems. They operate at relatively low frequencies, complicate application development, and do not consider energy harvesting dynamics to optimize their throughput. To address the issues, this article presents a novel CIM-based batteryless computing platform, called Viadotto, that provides efficient and adaptive acceleration for memory-bound computing workloads. Viadotto meets adaptive CIM and microcontroller-based (MCU-based) conventional batteryless platforms for the first time. Basically, Viadotto exposes a programming model supported by its compiler and a pipelined memory controller, which hides low-level CIM operations from applications. Furthermore, its runtime issues CIM operations in an energy-efficient manner and optimizes throughput in a programmer-transparent way by adapting CIM parallelism to react to ambient power dynamics. Our evaluation shows that Viadotto outperforms existing CIM solutions for batteryless systems by 48%. Khakim Akhunov, Kasim Sinan Yildirim, Jongouk Choi, Changhee Jung |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | Hybrid Power Failure Recovery for Intermittent ComputingabstractEnergy harvesting systems rely on either rollback or roll-forward recovery to resume power-interrupted program correctly. However, both recovery schemes have their own inherent drawbacks. To this end, this paper presents RollSwitch, a hybrid power failure recovery scheme that can achieve low-cost yet high-performance intermittent computation for energy harvesting systems. According to the underlying energy harvesting condition, RollSwitch dynamically switches between rollback and roll-forward recovery modes to maximize the performance. In particular, RollSwitch leverages the level of available energy in the capacitor as a proxy for determining the appropriate recovery mode. For this purpose, RollSwitch devises a simple capacitor energy predictor whose outcome governs the recovery mode selection in the near future. The experimental results demonstrate that RollSwitch achieves 15.0% and 19.8% performance gain on average over the state-of-art rollback and roll-forward recovery schemes, respectively. Gan Fang, Jongouk Choi, Changhee Jung |
ICCAD | 3 |
| 2024 | Soft Error Resilience at Near-Zero CostabstractAmong existing schemes for soft error resilience, acoustic-sensor-based detection stands out owing to its ability to prevent silent data corruption at low hardware cost. However, the state-of-the-art work not only incurs a considerable run-time overhead but also complicates the processor pipeline with intrusive microarchitectural modifications, hindering its practical deployment in real silicon. To this end, this paper presents VeriPipe, a near-zero-cost compiler/architecture codesign scheme for soft error resilience. VeriPipe compiler partitions input program to a series of regions (epochs) statically, while VeriPipe hardware verifies if they are error-free dynamically. In particular, VeriPipe achieves a simple yet efficient region-level verification where each region goes through a three-stage (Execute, Verify, and Commit) verification pipeline to ensure the absence of soft errors before proceeding to the next region. In particular, VeriPipe hardware overlaps the Verify stage of each region with the Execute stage of the next region, thereby effectively hiding the Verify delay. Experiments with 43 applications from SPEC2006/2017/NPB-CPP/SPLASH3/DoE Mini-Apps highlight the negligible overheads of VeriPipe, i.e., an average of 1% run-time overhead and a storage overhead of only 3 registers and 1 countdown timer. Jianping Zeng 0001, Shao-Yu Huang, Jiuyang Liu, Changhee Jung |
ICS | 4 |
| 2024 | Compiler-Directed Whole-System PersistenceabstractNonvolatile memory (NVM) technologies have gained increasing attention thanks to their density and durability benefits. However, leveraging NVM can cause a crash consistency issue. For example, if a younger store is evicted (persisted) to NVM from volatile caches before an older one and power failure occurs in between, it might be impossible to correctly resume the interrupted program in the wake of the failure. Traditionally, addressing this issue involves expensive persist barriers for enforcing the original store order, which not only incurs a high run-time overhead but also places a significant burden on users due to the difficulty of persistent programming. To this end, this paper presents cWSP, compiler/architecture codesign for lightweight yet performant whole-system persistence (WSP). In particular, cWSP compiler partitions not only user applications but also OS and runtime libraries into a series of recoverable regions (epochs), thus enabling persistence and crash consistency for the entire software stack. To achieve high-performance crash consistency, cWSP leverages advanced compiler optimizations for checkpointing a minimal set of registers and proposes simple hardware support for expediting data persistence on the cheap. Experimental results with 37 applications from SPEC CPU2006/2017, DOE Mini-apps, SPLASH3, WHISPER, and STAMP, show that cWSP incurs an average runtime overhead of $6 \%$, outperforming the state-of-the-art work with a significant margin. Jianping Zeng 0001, Changhee Jung |
ISCA | 3 |
| 2024 | Defending Against EMI Attacks on Just-In-Time Checkpoint for Resilient Intermittent SystemsabstractEnergy harvesting systems have emerged as an alternative to battery-powered IoT devices. The systems utilize a just-in-time checkpoint protocol that stores volatile states when a power outage occurs, ensuring crash consistency. However, this paper uncovers a new security vulnerability in the checkpoint protocol, revealing its susceptibility to electromagnetic interference (EMI). If exploited, adversaries could cause denial of service or data corruption in victim devices. To defeat EMI attacks, this paper introduces GECKO, a compiler-directed countermeasure that operates on commodity platforms used in energy harvesting systems without requiring hardware support. Our experiments on real boards demonstrate that GECKO defeats the EMI attack with a trivial performance overhead by 6% on average. Jaeseok Choi, Hyunwoo Joe, Changhee Jung, Jongouk Choi |
MICRO | 3 |
| 2024 | LightWSP: Whole-System Persistence on the CheapabstractWhole-system persistence (WSP) has recently attracted more interest thanks to its transparency and performance benefits over partial-system persistence where users are not only burdened by complex persistent programming but also incapable of using DRAM as LLC. Nevertheless, existing WSP work either introduces high hardware cost or causes non-trivial performance overhead. To this end, this paper presents LightWSP, a compiler/architecture co-design scheme that can achieve WSP in a lightweight yet performant manner. LightWSP compiler partitions program into a series of recoverable regions (epochs) with their live-out registers checkpointed, while LightWSP hardware persists the stores of the regions-whose boundary serves as a power failure recovery point-enforcing crash consistency; LightWSP leverages the battery-backed write pending queue (WPQ) of a memory controller as a redo buffer, i.e., all stores are first buffered in WPQ and then persisted together in non-volatile memory (NVM) at each region end. In this way, no matter when power failure happens, NVM is never corrupted by the stores of the power-interrupted region, facilitating correct recovery. In particular, LightWSP supports multiple memory controllers on the cheap without costly speculation/misspeculation handling mechanisms used by prior work. The experimental results with 38 applications show that LightWSP incurs only an average of 9.0% run-time overhead. This is on par with the state-of-the-art work, that complicates the core microarchitecture significantly with its intrusive design for memory controller speculation, yet the hardware cost of LightWSP is near zero (0.5B per core). Yuchen Zhou 0005, Jianping Zeng 0001, Changhee Jung |
MICRO | 3 |
| 2024 | IntOS: Persistent Embedded Operating System and Language Support for Multi-threaded Intermittent Computing
Byounguk Min, Mohannad Ismail, Wenjie Xiong 0001, Changhee Jung |
OSDI | 5 |
| 2024 | Optimizing Logical Execution Time Model for Both Determinism and Low LatencyabstractThe Logical Execution Time (LET) programming model has recently received considerable attention, particularly because of its timing and dataflow determinism. In LET, task computation appears always to take the same amount of time (called the task's LET interval), and the task reads (resp. writes) at the beginning (resp. end) of the interval. Compared to other communication mechanisms, such as implicit communication and Dynamic Buffer Protocol (DBP), LET performs worse on many metrics, such as end-to-end latency (including reaction time and data age) and time disparity jitter. Compared with the default LET setting, the flexible LET (fLET) model shrinks the LET interval while still guaranteeing schedulability by introducing the virtual offset to defer the read operation and using the virtual deadline to move up the write operation. Therefore, fLET has the potential to significantly improve the end-to-end timing performance while keeping the benefits of deterministic behavior on timing and dataflow. To fully realize the potential of fLET, we consider the problem of optimizing the assignments of its virtual offsets and deadlines. We propose new abstractions to describe the task communication pattern and new optimization algorithms to explore the solution space efficiently. The algorithms leverage the linearizability of communication patterns and utilize symbolic operations to achieve efficient optimization while providing a theoretical guarantee. The framework supports optimizing multiple performance metrics, and guarantees bounded suboptimality when optimizing end-to-end latency. Experimental results show that our optimization algorithms improve upon the default LET and its existing extensions and significantly outperform implicit communication and DBP in terms of various metrics, such as end-to-end latency, time disparity, and its jitter. Sen Wang 0014, Dong Li 0035, Ashrarul H. Sifat, Shao-Yu Huang, Xuanliang Deng, Changhee Jung, Ryan K. Williams, Haibo Zeng 0001 |
RTAS | 6 |
| 2024 | Partitioned scheduling with safety-performance trade-offs in stochastic conditional DAG models
Xuanliang Deng, Ashrarul H. Sifat, Shao-Yu Huang, Sen Wang 0014, Jia-Bin Huang 0001, Changhee Jung, Ryan K. Williams, Haibo Zeng 0001 |
J. Syst. Archit. | 6 |
| 2024 | Caphammer: Exploiting Capacitor Vulnerability of Energy Harvesting SystemsabstractAn energy harvesting system (EHS) has emerged as an alternative to traditional battery-operated Internet of Things (IoT) devices. An EHS harnesses ambient energy and stores it in a small capacitor, enabling batteryless operation when sufficient energy is available. However, capacitors are susceptible to malicious charging/discharging and over-voltages, which can lead to a loss of capacitance. With the capacitor vulnerability in mind, this article introduces a capacitor hammering attack, simply Caphammer, that can undermine the security of every EHS. The idea is that Caphammer can degrade the capacitance by using frequent power outages. Once Caphammer degrades the capacitor of the victim EHS, it can suffer from denial of service, data corruption, data encryption failure, and abnormal termination. To defeat Caphammer, this article presents FanCap, a capacitor bank scheduling scheme that can dynamically transform energy storage organization, taking into account the capacitor vulnerability. The experimental results demonstrate that FanCap can successfully thwart Caphammer with a negligible run-time overhead. Jongouk Choi, Jaeseok Choi, Hyunwoo Joe, Changhee Jung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Time-Triggered Scheduling for Nonpreemptive Real-Time DAG Tasks Using 1-Opt Local SearchabstractModern real-time systems often involve numerous computational tasks characterized by intricate dependency relationships. Within these systems, data propagate through cause–effect chains from one task to another, making it imperative to minimize end-to-end latency to ensure system safety and reliability. In this article, we introduce innovative nonpreemptive scheduling techniques designed to reduce the worst-case end-to-end latency and/or time disparity for task sets modeled with directed acyclic graphs (DAGs). This is challenging because of the noncontinuous and nonconvex characteristics of the objective functions, hindering the direct application of standard optimization frameworks. Customized optimization frameworks aiming at achieving optimal solutions may suffer from scalability issues, while general heuristic algorithms often lack theoretical performance guarantees. To address this challenge, we incorporate the “1-opt” concept from the optimization literature (Essentially, 1-opt means that the quality of a solution cannot be improved if only one single variable can be changed) into the design of our algorithm. We propose a novel optimization algorithm that effectively balances the tradeoff between theoretical guarantees and algorithm scalability. By demonstrating its theoretical performance guarantees, we establish that the algorithm produces 1-opt solutions while maintaining polynomial run-time complexity. Through extensive large-scale experiments, we demonstrate that our algorithm can effectively reduce the latency metrics by 20% to 40%, compared to state-of-the-art methods. Sen Wang 0014, Dong Li 0035, Shao-Yu Huang, Xuanliang Deng, Ashrarul H. Sifat, Jia-Bin Huang 0001, Changhee Jung, Ryan K. Williams, Haibo Zeng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Write-Light Cache for Energy Harvesting SystemsabstractEnergy harvesting system has huge potential to enable battery-less Internet of Things (IoT) services. However, it has been designed without a cache due to the difficulty of crash consistency guarantee, limiting its performance. This paper introduces Write-Light Cache (WL-Cache), a specialized cache architecture with a new write policy for energy harvesting systems. WL-Cache combines benefits of a write-back cache and a write-through cache while avoiding their downsides. Unlike a write-through cache, WL-Cache does not access a non-volatile main memory (NVM) at every store but it holds dirty cache lines in a cache to exploit locality, saving energy and improving performance. Unlike a write-back cache, WL-Cache limits the number of dirty lines in a cache. When power is about to be cut off, WL-Cache flushes the bounded set of dirty lines to NVM in a failure-atomic manner by leveraging a just-in-time (JIT) checkpointing mechanism to achieve crash consistency across power failure. For optimization, WL-Cache interacts with a run-time system that estimates the quality of energy source during each power-on period, and adaptively reconfigures the possible number of dirty cache lines at boot time. Our experiments demonstrate that WL-Cache reduces hardware complexity and provides a significant speedup over the state-of-the-art volatile cache design with non-volatile backup. For two representative power outage traces, WL-Cache achieves 1.35x and 1.44x average speedups, respectively, across 23 benchmarks used in prior work. Jongouk Choi, Jianping Zeng 0001, Changwoo Min, Changhee Jung |
ISCA | 5 |
| 2023 | Persistent Processor ArchitectureabstractThis paper presents PPA (Persistent Processor Architecture), simple microarchitectural support for lightweight yet performant whole-system persistence. PPA offers fully transparent crash consistency to all sorts of program covering the entire computing stack and even legacy applications without any source code change or recompilation. As a basis for crash consistency, PPA leverages so-called store integrity that preserves store operands during program execution, persists them on impending power failure, and replays the stores when power comes back. In particular, PPA realizes the store integrity via hardware by keeping the operands in a physical register file (PRF), though the stores are committed. Such store integrity enforcement leads to region-level persistence, i.e., whenever PRF runs out, PPA starts a new region after ensuring that all stores of the prior region have already been written to persistent memory. To minimize the pipeline stall across regions, PPA writes back the stores of each region asynchronously, overlapping their persistence latency with the execution of other instructions in the region. The experimental results with 41 applications from SPEC CPU2006/2017, SPLASH3, STAMP, WHISPER, and DOE Mini-apps show that PPA incurs only a 2% average run-time overhead and a 0.005% areal cost, while the state-of-the-art work suffers a 26% overhead along with prohibitively high hardware and energy costs. Jianping Zeng 0001, Jungi Jeong, Changhee Jung |
MICRO | 3 |
| 2023 | SweepCache: Intermittence-Aware Cache on the CheapabstractThis paper presents SweepCache, a new compiler/architecture co-design scheme that can equip energy harvesting systems with a volatile cache in a performant yet lightweight way. Unlike prior just-in-time checkpointing designs that persists volatile data just before power failure and thus dedicates additional energy, SweepCache partitions program into a series of recoverable regions and persists stores at region granularity to fully utilize harvested energy for computation. In particular, SweepCache introduces persist buffer—as a redo buffer resident in nonvolatile memory (NVM)—to keep the main memory consistent across power failure while persisting region’s stores in a failure-atomic manner. Specifically, for writebacks during region execution, SweepCache saves their cachelines to the persist buffer. At each region end, SweepCache first flushes dirty cachelines to the buffer, allowing the next region to start with a clean cache, and then moves all buffered cachelines to the corresponding NVM locations. In this way, no matter when power failure occurs, the buffer contents or their memory locations always remain intact, which serves as a basis for correct recovery. To hide the persistence delay, SweepCache speculatively starts a region right after the prior region finishes its execution—as if its stores were already persisted—with the two regions having their own persist buffer, i.e., dual-buffering. This region-level parallelism helps SweepCache to achieve the full potential of a high-performance data cache. The experimental results show that compared to the original cache-free nonvolatile processor, SweepCache delivers speedups of 14.60x and 14.86x—outperforming the state-of-the-art work by 3.47x and 3.49x—for two representative energy harvesting power traces, respectively. Yuchen Zhou 0005, Jianping Zeng 0001, Jungi Jeong, Jongouk Choi, Changhee Jung |
MICRO | 5 |
| 2023 | RTailor: Parameterizing Soft Error Resilience for Mixed-Criticality Real-Time SystemsabstractEquipping real-time systems with soft error resilience can be challenging due to the tradeoff of the timing and failure requirements for mixed-criticality tasks. Violation of these requirements yields failed task scheduling in one way or another. However, not every task requires the same degree of soft error resilience. For example, low-criticality tasks can run with low or even no soft error resilience, whereas mid- or highcriticality tasks may require relatively high resilience depending on their inherent failure requirement. Unfortunately, existing soft error resilience schemes do not have the ability to control the degree of their resilience in a fine-grained way, i.e., they can only be turned on or off as a whole during task execution. To this end, this paper presents RTailor (Resilience Tailor), a compiler-directed parameterized soft error resilience scheme that achieves the desired level of soft error protection according to the demand of each task. The key idea is that for a given protection ratio, compilers can transform a hot loop such that the number of its iterations protected over the total iterations matches the ratio. Compared to full resilience protecting every iteration, RTailor's parameterized soft error resilience significantly reduces the performance overhead of tasks, thereby improving their real-time schedulability. The experimental results highlight that for four representative fault rates, RTailor achieves 15%~average schedulability improvements over the state-of-the-art work that lacks parameterized soft error resilience. Shao-Yu Huang, Jianping Zeng 0001, Xuanliang Deng, Sen Wang 0014, Ashrarul H. Sifat, Burhanuddin Bharmal, Jia-Bin Huang 0001, Ryan K. Williams, Haibo Zeng 0001, Changhee Jung |
RTSS | 10 |
| 2023 | DevFuzz: Automatic Device Model-Guided Device Driver FuzzingabstractThe security of device drivers is critical for the entire operating system’s reliability. Yet, it remains very challenging to validate if a device driver can properly handle potentially malicious input from a hardware device. Unfortunately, existing symbolic execution-based solutions often do not scale, while fuzzing solutions require real devices or manual device models, leaving many device drivers under-tested and insecure.This paper presents DevFuzz, a new model-guided device driver fuzzing framework that does not require a physical device. DevFuzz uses symbolic execution to automatically generate the probe model that can guide a fuzzer to properly initialize a device driver under test. DevFuzz also leverages both static and dynamic program analyses to construct MMIO, PIO, and DMA device models to improve the effectiveness of fuzzing further. DevFuzz successfully tested 191 device drivers of various bus types (PCI, USB, RapidIO, I2C) from different operating systems (Linux, FreeBSD, and Windows) and detected 72 bugs, 41 of which have been patched and merged into the mainstream. Changhee Jung |
SP | 3 |
| 2023 | Automatic Permission Check Analysis for Linux KernelabstractPermission checks play an essential role in operating system security by providing access control to privileged functionalities. However, it is challenging for kernel developers to scalably verify the soundness of existing checks due to the large codebase and complexity of the kernel. In fact, Linux kernel contains millions of lines of code with hundreds of permission checks, and even worse its complexity is fast-growing. This paper presents PeX, a static permission check error detector for Linux, which takes as input a kernel source code and reports any missing, inconsistent, and redundant permission checks. PeX uses KIRIN (Kernel InteRface based Indirect call aNalysis), a novel, precise, and scalable indirect call analysis technique. Over the interprocedural control flow graph built by KIRIN, PeX automatically identifies permission checks and infers the mappings between permission checks and privileged functions. For each privileged function, PeX examines all possible paths to the function to check if necessary permission checks are correctly enforced. We evaluated PeX on the latest stable Linux kernel v4.18.5 for three types of permission checks: Discretionary Access Controls (DAC), Capabilities, and Linux Security Modules (LSM). PeX reported 45 new permission check errors, 17 of which have been confirmed by the kernel developers. Jinmeng Zhou, Wenbo Shen, Changhee Jung, Ahmed M. Azab, Ruowen Wang, Peng Ning, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2022 | Capri: Compiler and Architecture Support for Whole-System PersistenceabstractThis paper investigates whole-system persistence (WSP) that ensures hassle-free crash consistency for all programs while simultaneously leveraging both advantages of the non-volatile memory technologies: high-density and in-memory persistence. Despite the promising characteristics, there are two challenges that must be addressed to make WSP a reality. First, programs must be able to resume the execution from where they had a failure. Second, failure recovery must be offered to any program including the OS in a transparent manner while minimizing persistence overheads. Jungi Jeong, Jianping Zeng 0001, Changhee Jung |
HPDC | 3 |
| 2022 | Featherweight Soft Error Resilience for GPUsabstractThis paper presents Flame, a hardware/software co-designed resilience scheme for protecting GPUs against soft errors. For low-cost yet high-performance resilience, Flame uses acoustic sensors and idempotent processing for error detection and recovery, respectively. That is, Flame seeks to correct any sensor-detected errors by re-executing the idempotent region where they occurred. To achieve this, it is essential for each idempotent region to ensure the absence of errors before moving on to the next region. This is so-called soft error verification that takes sensors’ worst-case detection latency (WCDL) to verify each region finished. Rather than waiting for WCDL at each region end, which incurs too much performance overhead, Flame proposes WCDL-aware warp scheduling that can hide the error verification delay (i.e., WCDL) with GPU’s inherent massive warp-level parallelism. When a warp hits each idempotent region boundary, Flame deschedules the warp and switches to one of the other ready warps—as if the region boundary were a regular long-latency operation triggering the warp switching. By leveraging GPU’s inherent ability for the latency hiding, Flame can completely eliminate the verification delay without significant hardware modification. The experimental results demonstrate that the performance overhead of Flame is near zero, i.e., 0.6% on average for 34 GPU benchmark applications. Changhee Jung |
MICRO | 2 |
| 2022 | Compiler-Directed High-Performance Intermittent Computation with Power Failure ImmunityabstractThis paper introduces power failure immunity (PFI), an essential program execution property for energy harvesting systems to achieve efficient intermittent computation. PFI ensures program code regions never fail more than once i.e., at most single in-region outage, during intermittent computation as if they are immunized after the first power outage. To enforce PFI automatically for such batteryless systems that use a tiny energy buffer instead, we present its compiler-directed enforcement. The compiler leverages a precise static analysis to partition the program into recoverable regions with the energy buffer size in mind so that their execution can be completed—using the full energy buffered in a single charge cycle—regardless of program execution paths. In this way, no matter how unstable the energy harvesting source is, no region fails more than once.In the virtue of PFI, this paper presents ROCKCLIMB, a high-performance and rollback-free intermittent computation scheme. It guarantees that PFI-enforced regions never fail, i.e., there is no in-region outage at all. To achieve this, ROCKCLIMB checks if the fully buffered energy is secured at each region boundary. If it is not secured, ROCKCLIMB waits until the energy buffer is fully charged, before executing the following region. In particular, the rollback-free nature of ROCKCLIMB obviates the need to log memory writes—required for rollback recovery—since no region is power-interrupted. As a result, PFI+ROCKCLIMB achieves rollback-free and memory-log-free intermittent computation, ensuring forward execution progress and maximizing it even in the presence of frequent power outages. Our real board experiments demonstrate that PFI+ROCKCLIMB outperforms the state-of-the-art work by 5%—550% on average in various energy harvesting conditions. Jongouk Choi, Larry Kittinger, Qingrui Liu, Changhee Jung |
RTAS | 4 |
| 2022 | CapOS: Capacitor Error Resilience for Energy Harvesting SystemsabstractEnergy harvesting systems have emerged as an alternative to battery-operated Internet of Things (IoT) devices. To deal with frequent power outages in the absence of battery, energy harvesting systems rely on a capacitor-backed checkpoint mechanism also known as just-in-time (JIT) checkpointing. It checkpoints volatile data in nonvolatile memory (NVM) just before a power outage occurs—using the energy buffered in the capacitor—and restores the checkpointed data from NVM in the wake of the outage. While the JIT checkpointing gives an illusion that volatile data survive a power outage as if they were nonvolatile, it turns out that due to capacitor degradation, energy harvesting systems can unexpectedly fail the JIT checkpointing, losing or corrupting data across the outage. To address the problem, this article presents an operating system-driven solution called CapOS. At a high level, CapOS diagnoses the capacitor in a reactive yet safe manner. When the JIT checkpoint failure occurs, CapOS detects the capacitor degradation without causing the data corruption. To recover from such a capacitor error, CapOS electrically isolates the degraded capacitor—so that it restores its original capacitance by itself with the help of capacitor’s resilient nature—and disables the JIT checkpointing. In case, power outages occur during the capacitor isolation, CapOS leverages undo logging with interval-based checkpointing for their recovery. Once the capacitor is fully recovered, CapOS gets back to the capacitor-based JIT checkpointing. The experimental results demonstrate that CapOS can effectively address the capacitor error of energy harvesting systems at a low run-time cost, without compromising the recovery of power outages. Jongouk Choi, Hyunwoo Joe, Changhee Jung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | PMEM-spec: persistent memory speculation (strict persistency can trump relaxed persistency)abstractPersistency models define the persist-order that controls the order in which stores update persistent memory (PM). As with memory consistency, the relaxed persistency models provide better performance than the strict ones by relaxing the ordering constraints. To support such relaxed persistency models, previous studies resort to APIs for annotating the persist-order in program and hardware implementations for enforcing the programmer-specified order. However, these approaches to supporting relaxed persistency impose costly burdens on both architects and programmers. In light of this, the goal of this study is to demonstrate that the strict persistency model can outperform the relaxed models with significantly less hardware complexity and programming difficulty. To achieve that, this paper presents PMEM-Spec that speculatively allows any PM accesses without stalling or buffering, detecting their ordering violation (e.g., misspeculation for PM loads and stores). PMEM-Spec treats misspeculation as power failure and thus leverages failure-atomic transactions to recover from misspeculation by aborting and restarting them purposely. Since the ordering violation rarely occurs, PMEM-Spec can accelerate persistent memory accesses without significant misspeculation penalty. Experimental results show that PMEM-Spec outperforms two epoch-based persistency models with Intel X86 ISA and the state-of-the-art hardware support by 27.2% and 10.6%, respectively. Jungi Jeong, Changhee Jung |
ASPLOS | 2 |
| 2021 | ReplayCache: Enabling Volatile Cachesfor Energy Harvesting SystemsabstractEnergy harvesting systems have shown their unique benefit of ultra-long operation time without maintenance and are expected to be more prevalent in the era of Internet of Things. However, due to the batteryless nature, they suffer unpredictable frequent power outages. They thus require a lightweight mechanism for crash consistency since saving/restoring checkpoints across the outages can limit forward progress by consuming hard-won energy. For the reason, energy harvesting systems have been designed with a non-volatile memory (NVM) only. The use of a volatile data cache has been assumed to be not viable or at least challenging due to the difficulty to ensure cacheline persistence. Jianping Zeng 0001, Jongouk Choi, Xinwei Fu, Ajay Paddayuru Shreepathi, Changwoo Min, Changhee Jung |
MICRO | 7 |
| 2021 | Turnpike: Lightweight Soft Error Resilience for In-Order CoresabstractAcoustic-sensor-based soft error resilience is particularly promising, since it can verify the absence of soft errors and eliminate silent data corruptions at a low hardware cost. However, the state-of-the-art work incurs a significant performance overhead for in-order cores due to frequent structural/data hazards during the verification. To address the problem, this paper presents Turnpike, a compiler/architecture co-design scheme that can achieve lightweight yet guaranteed soft error resilience for in-order cores. The key idea is that many of the data computed in the core can bypass the soft error verification without compromising the resilience. Along with simple microarchitectural support for realizing the idea, Turnpike leverages compiler optimizations to further reduce the performance overhead. Experimental results with 36 benchmarks demonstrate that Turnpike only incurs a 0-14% run-time overhead on average while the state-of-the-art incurs a 29-84% overhead when the worst-case latency of the sensor based error detection is 10-50 cycles. Jianping Zeng 0001, Hongjune Kim, Jaejin Lee, Changhee Jung |
MICRO | 4 |
| 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory SystemabstractPersistent memory programming requires failure atomicity. To achieve this in an efficient manner, recent proposals use hardware-based logging for atomic-durable updates and hardware transactional memory (HTM) for isolation. Although the unbounded HTMs are promising for both performance and programmability reasons, none of the previous studies satisfies the practical requirements. They either require unrealistic hard-ware overheads or do not allow transactions to exceed on-chip cache boundaries. Furthermore, it has never been possible to use both DRAM and NVM in HTM, though it is becoming a popular persistency model. To this end, this study proposes UHTM, unbounded hardware transactional memory for DRAM and NVM hybrid memory systems. UHTM combines the cache coherence protocol and address-signatures to detect conflicts in the entire memory space. This approach improves concurrency by significantly reducing the false-positive rates of previous studies. More importantly, UHTM allows both DRAM and NVM data to interact with each other in transactions without compromising the consistency guarantee. This is rendered possible by UHTM's hybrid version management that provides an undo-based log for DRAM and a redo-based log for NVM. The experimental results show that UHTM outperforms the state-of-the-art durable HTM, which is LLC-bounded, by 56% on average and up to 818%. Jungi Jeong, Jaewan Hong, Seung Ryoul Maeng, Changhee Jung, Youngjin Kwon |
MICRO | 4 |
| 2020 | Compiler-directed soft error resilience for lightweight GPU register file protectionabstractThis paper presents Penny, a compiler-directed resilience scheme for protecting GPU register files (RF) against soft errors. Penny replaces the conventional error correction code (ECC) based RF protection by using less expensive error detection code (EDC) along with idempotence based recovery. Compared to the ECC protection, Penny can achieve either the same level of RF resilience yet with significantly lower hardware costs or stronger resilience using the same ECC due to its ability to detect multi-bit errors when it is used solely for detection. In particular, to address the lack of store buffers in GPUs, which causes both checkpoint storage overwriting and the high cost of checkpointing stores, Penny provides several compiler optimizations such as storage coloring and checkpoint pruning. Across 25 benchmarks, Penny causes only ≈3% run-time overhead on average. Hongjune Kim, Jianping Zeng 0001, Qingrui Liu, Mohammad Abdel-Majeed, Jaejin Lee, Changhee Jung |
PLDI | 6 |
| 2019 | BOGO: Buy Spatial Memory Safety, Get Temporal Memory Safety (Almost) FreeabstractA memory safety violation occurs when a program has an out-of-bound (spatial safety) or use-after-free (temporal safety) memory access. Given its importance as a security vulnerability, recent Intel processors support hardware-accelerated bound checks, called Memory Protection Extensions (MPX). Unfortunately, MPX provides no temporal safety. This paper presents BOGO, a lightweight full memory safety enforcement scheme that transparently guarantees temporal safety on top of MPX's spatial safety. Instead of tracking separate metadata for temporal safety, BOGO reuses the bounds metadata maintained by MPX for both spatial and temporal safety. On free, BOGO scans the MPX bound tables to invalidate the bound of dangling pointers; any following use-after-free error can be detected by MPX as an out-of-bound error. Since scanning the entire MPX bound tables could be expensive, BOGO tracks a small set of hot MPX bound table pages to check on free, and relies on the page fault mechanism to detect any potentially missing dangling pointer, ensuring sound temporal safety protection. Our evaluation shows that BOGO provides full memory safety at 60% runtime overhead and at 36% memory overhead for SPEC CPU 2006 benchmarks. We also show that BOGO incurs reasonable 2.7x slowdown for the worst-case malloc-free intensive benchmarks; and moderate 1.34x overhead for real-world applications. Changhee Jung |
ASPLOS | 3 |
| 2019 | CoSpec: Compiler Directed Speculative Intermittent ComputationabstractEnergy harvesting systems have emerged as an alternative to battery-operated embedded devices. Due to the intermittent nature of energy harvesting, researchers equip the systems with nonvolatile memory (NVM) and crash consistency mechanisms. However, prior works require non-trivial hardware modifications, e.g., a voltage monitor, nonvolatile flip-flops/scratchpad, dependence tracking modules, etc., thereby causing significant area/power/manufacturing costs. Jongouk Choi, Qingrui Liu, Changhee Jung |
MICRO | 3 |
| 2019 | Achieving Stagnation-Free Intermittent Computation with Boundary-Free Adaptive ExecutionabstractThis paper presents ELASTIN, a stagnation-free intermittent computing system for energy-harvesting devices that ensures forward progress in the presence of frequent power outages without partitioning program into recoverable regions or tasks. ELASTIN leverages both timer-based checkpointing of volatile registers and copy-on-write mappings of nonvolatile memory pages to restore them in the wake of power failure. During each checkpoint interval, ELASTIN tracks memory writes on a per-page basis and backs up the original page using custom software-controlled memory protection without MMU or TLB. When a new interval starts at each timer expiration, ELASTIN clears the write permission of all the pages written during the previous interval and checkpoints all registers including a program counter as a recovery point. In particular, ELASTIN dynamically reconfigures both the checkpoint interval and the page size to achieve stagnation-free intermittent computation and maximize forward progress across power outages. The experiments on TI's MSP430 board with energy harvesting traces show that ELASTIN outperforms the state-of-the-art scheme by 3.5X on average (up to orders of magnitude speedup) and guarantees forward progress. Jongouk Choi, Hyunwoo Joe, Changhee Jung |
RTAS | 4 |
| 2019 | PeX: A Permission Check Analysis Framework for Linux Kernel
Wenbo Shen, Changhee Jung, Ahmed M. Azab, Ruowen Wang |
USENIX Security Symposium | 4 |
| 2018 | nAdroid: statically detecting ordering violations in Android applicationsabstractModern mobile applications use a hybrid concurrency model. In this model, events are handled sequentially by event loop(s), and long-running tasks are offloaded to other threads. Concurrency errors in this hybrid concurrency model can take multiple forms: traditional atomicity and ordering violations between threads, as well as ordering violations between event callbacks on a single event loop. Xinwei Fu, Changhee Jung |
CGO | 3 |
| 2018 | CommAnalyzer: automated estimation of communication cost and scalability on HPC clusters from sequential codeabstractTo deliver scalable performance to large-scale scientific and data analytic applications, HPC cluster architectures adopt the distributed-memory model. The performance and scalability of parallel applications on such systems are limited by the communication cost across compute nodes. Therefore, projecting the minimum communication cost and maximum scalability of the user applications plays a critical role in assessing the benefits of porting these applications to HPC clusters as well as developing efficient distributed-memory implementations. Unfortunately, this task is extremely challenging for end users, as it requires comprehensive knowledge of the target application and hardware architecture and demands significant effort and time for manual system analysis. Ahmed E. Helal, Changhee Jung, Wu-chun Feng, Yasser Y. Hanafy |
HPDC | 2 |
| 2018 | iDO: Compiler-Directed Failure Atomicity for Nonvolatile MemoryabstractThis paper presents iDO, a compiler-directed approach to failure atomicity with nonvolatile memory. Unlike most prior work, which instruments each store of persistent data for redo or undo logging, the iDO compiler identifies idempotent instruction sequences, whose re-execution is guaranteed to be side-effect-free, thereby eliminating the need to log every persistent store. Using an extension of prior work on JUSTDO logging, the compiler then arranges, during recovery from failure, to back up each thread to the beginning of the current idempotent region and re-execute to the end of the current failure-atomic section. This extension transforms JUSTDO logging from a technique of value only on hypothetical future machines with nonvolatile caches into a technique that also significantly outperforms state-of-the art lock-based persistence mechanisms on current hardware during normal execution, while preserving very fast recovery times. Qingrui Liu, Joseph Izraelevitz, Se Kwon Lee, Michael L. Scott, Sam H. Noh, Changhee Jung |
MICRO | 6 |
| 2018 | Sampler: PMU-Based Sampling to Detect Memory Errors Latent in Production SoftwareabstractDeployed software is still faced with numerous in-production memory errors. They can significantly affect system reliability and security, causing application crashes, erratic execution behavior, or security attacks. Unfortunately, existing tools cannot be deployed in the production environment, since they either impose significant performance/memory overhead, or can only detect partial errors. This paper presents Sampler, a library that employs the combination of hardware-based SAMPLing and novel heap allocator design to efficiently identify a range of memory ERrors, including buffer overflows, use-after-frees, invalid frees, and double-frees. Due to the stringent Quality of Service (QoS) requirement of production services, Sampler proposes to trade detection effectiveness for performance on each execution. Rather than inspecting every memory access, Sampler proposes the use of the Performance Monitoring Unit (PMU) hardware to sample memory accesses, and only checks the validity of sampled accesses. At the same time, Sampler proposes a novel dynamic allocator supporting fast metadata lookup, and a solution to prevent false alarms potentially caused by sampling. The sampling-based approach, although it may lead to reduced effectiveness on each execution, is suitable for in-production software, since software is generally employed by a large number of individuals, and may be executed many times or over a long period of time. By randomizing the start of the sampling, different executions may sample different sequences of memory accesses, working together to enable effective detection. Experimental results demonstrate that Sampler detects all known memory bugs inside real applications, without any false positive. Sampler only imposes negligible performance overhead (2.4% on average). Sampler is the first work that simultaneously satisfies efficiency, preciseness, completeness, accuracy, and transparency, making it a practical tool for in-production deployment. Sam Silvestro, Hongyu Liu 0005, Changhee Jung, Tongping Liu |
MICRO | 4 |
| 2017 | ProRace: Practical Data Race Detection for Production UseabstractThis paper presents ProRace, a dynamic data race detector practical for production runs. It is lightweight, but still offers high race detection capability. To track memory accesses, ProRace leverages instruction sampling using the performance monitoring unit (PMU) in commodity processors. Our PMU driver enables ProRace to sample more memory accesses at a lower cost compared to the state-of-the-art Linux driver. Moreover, ProRace uses PMU-provided execution contexts including register states and program path, and reconstructs unsampled memory accesses offline. This technique allows \ProRace to overcome inherent limitations of sampling and improve the detection coverage by performing data race detection on the trace with not only sampled but also reconstructed memory accesses. Experiments using racy production software including apache and mysql shows that, with a reasonable offline cost, ProRace incurs only 2.6% overhead at runtime with 27.5% detection probability with a sampling period of 10,000. Changhee Jung |
ASPLOS | 2 |
| 2017 | Compiler-Directed Soft Error Detection and Recovery to Avoid DUE and SDC via Tail-DMRabstractThis article presents Clover, a compiler-directed soft error detection and recovery scheme for lightweight soft error resilience. The compiler carefully generates soft-error-tolerant code based on idempotent processing without explicit checkpoints. During program execution, Clover relies on a small number of acoustic wave detectors deployed in the processor to identify soft errors by sensing the wave made by a particle strike. To cope with DUEs (detected unrecoverable errors) caused by the sensing latency of error detection, Clover leverages a novel selective instruction duplication technique called tail-DMR (dual modular redundancy) that provides a region-level error containment. Once a soft error is detected by either the sensors or the tail-DMR, Clover takes care of the error as in the case of exception handling. To recover from the error, Clover simply redirects program control to the beginning of the code region where the error is detected. The experimental results demonstrate that the average runtime overhead is only 26%, which is a 75% reduction compared to that of the state-of-the-art soft error resilience technique. In addition, this article evaluates an alternative technique called tail-wait, comparing it to Clover. According to the evaluation with the different processor configurations and the various error detection latencies, Clover turns out to be a superior technique, achieving 1.06 to 3.49 × speedup over the tail-wait. Qingrui Liu, Changhee Jung, Devesh Tiwari |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | BenchPrime: Effective Building of a Hybrid Benchmark SuiteabstractThis paper presents BenchPrime, an automated benchmark analysis toolset that is systematic and extensible to analyze the similarity and diversity of benchmark suites. BenchPrime takes multiple benchmark suites and their evaluation metrics as inputs and generates a hybrid benchmark suite comprising only essential applications. Unlike prior work, BenchPrime uses linear discriminant analysis rather than principal component analysis, as well as selects the best clustering algorithm and the optimized number of clusters in an automated and metric-tailored way, thereby achieving high accuracy. In addition, BenchPrime ranks the benchmark suites in terms of their application set diversity and estimates how unique each benchmark suite is compared to other suites. As a case study, this work for the first time compares the DenBench with the MediaBench and MiBench using four different metrics to provide a multi-dimensional understanding of the benchmark suites. For each metric, BenchPrime measures to what degree DenBench applications are irreplaceable with those in MediaBench and MiBench. This provides means for identifying an essential subset from the three benchmark suites without compromising the application balance of the full set. The experimental results show that the necessity of including DenBench applications varies across the target metrics and that significant redundancy exists among the three benchmark suites. Qingrui Liu, Larry Kittinger, Markus Levy, Changhee Jung |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | TxRace: Efficient Data Race Detection Using Commodity Hardware Transactional MemoryabstractDetecting data races is important for debugging shared-memory multithreaded programs, but the high runtime overhead prevents the wide use of dynamic data race detectors. This paper presents TxRace, a new software data race detector that leverages commodity hardware transactional memory (HTM) to speed up data race detection. TxRace instruments a multithreaded program to transform synchronization-free regions into transactions, and exploits the conflict detection mechanism of HTM for lightweight data race detection at runtime. However, the limitations of the current best-effort commodity HTMs expose several challenges in using them for data race detection: (1) lack of ability to pinpoint racy instructions, (2) false positives caused by cache line granularity of conflict detection, and (3) transactional aborts for non-conflict reasons (e.g., capacity or unknown). To overcome these challenges, TxRace performs lightweight HTM-based data race detection at first, and occasionally switches to slow yet precise data race detection only for the small fraction of execution intervals in which potential races are reported by HTM. According to the experimental results, TxRace reduces the average runtime overhead of dynamic data race detection from 11.68x to 4.65x with only a small number of false negatives. Changhee Jung |
ASPLOS | 3 |
| 2016 | Low-cost soft error resilience with unified data verification and fine-grained recovery for acoustic sensor based detectionabstractThis paper presents Turnstile, a hardware/software cooperative technique for low-cost soft error resilience. Leveraging the recent advance of acoustic sensor based soft error detection, Turnstile achieves guaranteed recovery by taking into account the bounded detection latency. The compiler forms verifiable regions and selectively inserts store instructions to checkpoint their register inputs so that Turnstile can verify the register/memory states with regard to a region boundary in a unified way without expensive register file protection. At runtime, for each region, Turnstile regards any stores (to both memory and register checkpoints) as unverified, and thus holds them in a store queue until the region ends and spends the time of the error detection latency. If no error is detected during the time, the verified stores are merged into memory systems, and registers are checkpointed. When all the stores including checkpointing stores prior to a region boundary are verified, the architectural and memory states with regard to the boundary are verified, thus it can serve as a recovery point. In this way, Turnstile contains the errors within the core without extra memory buffering. When an error is detected, Turnstile invalidates unverified entries in the store queue and restores the checkpointed register values to get the architectural and memory states back to what they were at the most recently verified region boundary. Then, Turnstile simply redirects program control to the verified region boundary and continues execution. The experimental results demonstrate that Turnstile can offer guaranteed soft error recovery with low performance overhead (<8% on average). Qingrui Liu, Changhee Jung, Devesh Tiwari |
MICRO | 2 |
| 2016 | Compiler-directed lightweight checkpointing for fine-grained guaranteed soft error recoveryabstractThis paper presents Bolt, a compiler-directed soft error recovery scheme, that provides fine-grained and guaranteed recovery without excessive performance and hardware overhead. To get rid of expensive hardware support, the compiler protects the architectural inputs during their entire liveness period by safely checkpointing the last updated value in idempotent regions. To minimize the performance overhead, Bolt leverages a novel compiler analysis that eliminates those checkpoints whose value can be reconstructed by other checkpointed values without compromising the recovery guarantee. As a result, Bolt incurs only 4.7% performance overhead on average which is 57% reduction compared to the state-of-the-art scheme that requires expensive hardware support for the same recovery guarantee as Bolt. Qingrui Liu, Changhee Jung, Devesh Tiwari |
SC | 2 |
| 2015 | Clover: Compiler Directed Lightweight Soft Error ResilienceabstractThis paper presents Clover, a compiler directed soft error detection and recovery scheme for lightweight soft error resilience. The compiler carefully generates soft error tolerant code based on idempotent processing without explicit checkpoint. During program execution, Clover relies on a small number of acoustic wave detectors deployed in the processor to identify soft errors by sensing the wave made by a particle strike. To cope with DUE (detected unrecoverable errors) caused by the sensing latency of error detection, Clover leverages a novel selective instruction duplication technique called tail-DMR (dual modular redundancy). Once a soft error is detected by either the sensor or the tail-DMR, Clover takes care of the error as in the case of exception handling. To recover from the error, Clover simply redirects program control to the beginning of the code region where the error is detected. The experiment results demonstrate that the average runtime overhead is only 26%, which is a 75% reduction compared to that of the state-of-the-art soft error resilience technique. Qingrui Liu, Changhee Jung, Devesh Tiwari |
LCTES | 2 |
| 2014 | Automated memory leak detection for production useabstractThis paper presents Sniper, an automated memory leak detection tool for C/C++ production software. To track the staleness of allocated memory (which is a clue to potential leaks) with little overhead (mostly <3%), Sniper leverages instruction sampling using performance monitoring units available in commodity processors. It also offloads the time- and space-consuming analyses, and works on the original software without modifying the underlying memory allocator; it neither perturbs the application execution nor increases the heap size. The Sniper can even deal with multithreaded applications with very low overhead. In particular, it performs a statistical analysis, which views memory leaks as anomalies, for automated and systematic leak determination. Consequently, it accurately detected real-world memory leaks with no false positive, and achieved an F-measure of 81% on average for 17 benchmarks stress-tested with various memory leaks. Changhee Jung, Sangho Lee 0005, Easwaran Raman, Santosh Pande |
ICSE | 1 |
| 2014 | Detecting memory leaks through introspective dynamic behavior modelling using machine learningabstractThis paper expands staleness-based memory leak detection by presenting a machine learning-based framework. The proposed framework is based on an idea that object staleness can be better leveraged in regard to similarity of objects; i.e., an object is more likely to have leaked if it shows significantly high staleness not observed from other similar objects with the same allocation context. Sangho Lee 0005, Changhee Jung, Santosh Pande |
ICSE | 2 |
| 2011 | Brainy: effective selection of data structuresabstractData structure selection is one of the most critical aspects of developing effective applications. By analyzing data structures' behavior and their interaction with the rest of the application on the underlying architecture, tools can make suggestions for alternative data structures better suited for the program input on which the application runs. Consequently, developers can optimize their data structure usage to make the application conscious of an underlying architecture and a particular program input. Changhee Jung, Silvius Rus, Brian P. Railing, Nathan Clark, Santosh Pande |
PLDI | 1 |
| 2010 | Adaptive execution techniques of parallel programs for multiprocessors
Jaejin Lee, Jung-Ho Park, Honggyu Kim, Changhee Jung, Daeseob Lim, Sang-Yong Han |
J. Parallel Distributed Comput. | 4 |
| 2009 | DDT: design and evaluation of a dynamic program analysis for optimizing data structure usageabstractData structures define how values being computed are stored and accessed within programs. By recognizing what data structures are being used in an application, tools can make applications more robust by enforcing data structure consistency properties, and developers can better understand and more easily modify applications to suit the target architecture for a particular application. Changhee Jung, Nathan Clark |
MICRO | 1 |
| 2009 | Prefetching with Helper Threads for Loosely Coupled Multiprocessor SystemsabstractThis paper presents a helper thread prefetching scheme that is designed to work on loosely coupled processors, such as in a standard chip multiprocessor (CMP) system or an intelligent memory system. Loosely coupled processors have an advantage in that resources such as processor and L1 cache resources are not contended by the application and helper threads, hence preserving the speed of the application. However, interprocessor communication is expensive in such a system. We present techniques to alleviate this. Our approach exploits large loop-based code regions and is based on a new synchronization mechanism between the application and helper threads. This mechanism precisely controls how far ahead the execution of the helper thread can be with respect to the application thread. We found that this is important in ensuring prefetching timeliness and avoiding cache pollution. To demonstrate that prefetching in a loosely coupled system can be done effectively, we evaluate our prefetching by simulating a standard unmodified CMP system and an intelligent memory system where a simple processor in memory executes the helper thread. Evaluating our scheme with nine memory-intensive applications with the memory processor in DRAM achieves an average speedup of 1.25. Moreover, our scheme works well in combination with a conventional processor-side sequential L1 prefetcher, resulting in an average speedup of 1.31. In a standard CMP, the scheme achieves an average speedup of 1.33. Using a real CMP system with a shared L2 cache between two cores, our helper thread prefetching plus hardware L2 prefetching achieves an average speedup of 1.15 over the hardware L2 prefetching for the subset of applications with high L2 cache misses per cycle. Jaejin Lee, Changhee Jung, Daeseob Lim, Yan Solihin |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2007 | Performance characterization of prelinking and preloadingfor embedded systemsabstractApplication launching times in embedded systems are more crucial than in general-purpose systems since the response times of embedded applications are significantly affected by the launching times. As general-purpose operating systems are increasingly used in embedded systems, reducing appli-cation launching times are one of the most influential factors for performance improvement. In order to reduce the application launching times, three factors should be considered at the same time: relocation time, symbol resolution time, and binary loading time. In this paper, we propose a new application execution model using a combination of prelinking and preloading to reduce the relocation, symbol resolution, and binary load overheads at the same time. Such application execution model is realized using fork and dlopen execution model instead of traditional fork and exec execution model. We evaluate the performance effect of the proposed fork and dlopen application execution model on a Linux-based embedded system using XScale processor. By applying the proposed application execution model using both prelinking and preloading, the application launching times are reduced up to 71% and relocation counts are reduced up to 91% in the benchmark programs we used. Changhee Jung, Duk-Kyun Woo, Kanghee Kim, Sung-Soo Lim |
EMSOFT | 1 |
| 2006 | Helper thread prefetching for loosely-coupled multiprocessor systemsabstractThis paper presents a helper thread prefetching scheme that is designed to work on loosely-coupled processors, such as in a standard chip multiprocessor (CMP) system or an intelligent memory system. Loosely-coupled processors have an advantage in that fine-grain resources, such as processor and L1 cache resources, are not contended by the application and helper threads, hence preserving the speed of the application. However, inter-processor communication is expensive in such a system. We present techniques to alleviate this. Our approach exploits large loop-based code regions and is based on a new synchronization mechanism between the application and helper threads. This mechanism precisely controls how far ahead the execution of the helper thread can be with respect to the application thread. We found that this is important in ensuring prefetching timeliness and avoiding cache pollution. To demonstrate that prefetching in a loosely-coupled system can be done effectively, we evaluate our prefetching in a standard, unmodified CMP system, and in an intelligent memory system where a simple processor in memory executes the helper thread. Evaluating our scheme with nine memory-intensive applications with the memory processor in DRAM achieves an average speedup of 1.25. Moreover, our scheme works well in combination with a conventional processor-side sequential L1 prefetcher, resulting in an average speedup of 1.31. In a standard CMP, the scheme achieves an average speedup of 1.33. Changhee Jung, Daeseob Lim, Jaejin Lee, Yan Solihin |
IPDPS | 1 |
| 2005 | Adaptive execution techniques for SMT multiprocessor architecturesabstractIn simultaneous multithreading (SMT) multiprocessors, using all the available threads (logical processors) to run a parallel loop is not always beneficial due to the interference between threads and parallel execution overhead. To maximize performance in an SMT multiprocessor, finding the optimal number of threads is important. This paper presents adaptive execution techniques to find the optimal execution mode for SMT multiprocessor architectures. A compiler preprocessor generates code that, based on dynamic feedback, automatically determines at run time the optimal number of threads for each parallel loop in the application. Using 10 standard numerical applications and running them with our techniques on an Intel 4-processor Hyper-Threading Xeon SMP with 8 logical processors, our code is, on average, about 2 and 18 times faster than the original code executed on 4 and 8 logical processors, respectively. Changhee Jung, Daeseob Lim, Jaejin Lee, Sang-Yong Han |
PPoPP | 1 |