EDBT 2026 Demo / reviewers in the wild / expert
Jaehyun Park 0005
dblp:92/6473-5
· DBLP profile ↗
18ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-2276-4998ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PLP-DVS: Adaptive Energy Scaling of Capacitor-Based Power Loss Protection in SSDsabstractA capacitor-based Power Loss Protection (PLP) has been widely adopted in modern enterprise-level solid-state drives (SSD) for data security and performance. It becomes challenging to guarantee lifetime data security as SSD capacity rapidly increases in limited physical volume and cost. In this work, we proposed a dynamic control of capacitor voltage to enhance the reliability and energy efficiency of the PLP capacitors. Unlike previous studies, we reduce the voltage charged in the capacitor according to the status of the volatile contents to be protected. The proposed method is implemented in an SSD platform prototype. We successfully verified the feasibility of synchronized control of the capacitor voltage and volatile contents. The experimental result shows that the proposed method can reduce the capacitor voltage by 33% on average, corresponding to a 2.27 times longer lifetime and 41% reduction of leakage-induced energy loss. Finally, we can reduce the over-design factor by 36% of the fixed voltage setup. Jaehyeok Cho, Sungyong Ahn, Jaehyun Park 0005, Donghwa Shin |
IEEE Trans. Computers | 4 |
| 2026 | MFIT : Multi-FIdelity Thermal Modeling for 2.5D and 3D Multi-Chiplet ArchitecturesabstractRapidly evolving artificial intelligence and machine learning applications require ever-increasing computational capabilities, while monolithic 2D design technologies approach their limits. 2.5D/3D heterogeneous integration of smaller chiplets using advanced packaging has emerged as a promising paradigm for addressing this limit and meeting performance demands. These approaches offer a significant cost reduction and higher manufacturing yield than monolithic 2D integrated circuits. However, the compact arrangement and high compute density of these systems exacerbate thermal management challenges, potentially compromising performance. Addressing these thermal modeling challenges is critical, especially as system sizes grow and different design stages require varying levels of accuracy and speed. Since no single thermal modeling technique meets all these needs, this article introduces MFIT, a range of multi-fidelity thermal models that effectively balance accuracy and speed. These multi-fidelity models can enable efficient design space exploration and runtime thermal management. Our extensive testing on systems with 16, 36, and 64 2.5D integrated chiplets and 16×3 3D integrated chiplets demonstrates that these models can reduce execution times from days to mere seconds and milliseconds with negligible loss in accuracy. Lukas Pfromm, Alish Kanani, Parth Solanki, Eric Tervo, Jaehyun Park 0005, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | eMamba: Efficient Acceleration Framework for Mamba Models in Edge ComputingabstractState Space Model (SSM)-based machine learning architectures have recently gained significant attention for processing sequential data. Mamba, a recent sequence-to-sequence SSM, offers competitive accuracy with superior computational efficiency compared to state-of-the-art transformer models. While this advantage makes Mamba particularly promising for resource-constrained edge devices, no hardware acceleration frameworks are currently optimized for deploying it in such environments. This article presents eMamba, a comprehensive end-to-end hardware acceleration framework explicitly designed for deploying Mamba models on edge platforms. eMamba maximizes computational efficiency by replacing complex normalization layers with lightweight hardware-aware alternatives and approximating expensive operations, such as SiLU activation and exponentiation, considering the target applications. Then, it performs an approximation-aware neural architecture search (NAS) to tune the learnable parameters used during approximation. Evaluations with Fashion-MNIST, CIFAR-10, and MARS, an open-source human pose estimation dataset, show eMamba achieves comparable accuracy to state-of-the-art techniques using 1.63–19.9× fewer parameters. In addition, it generalizes well to large-scale natural language tasks, demonstrating stable perplexity across varying sequence lengths on the WikiText2 dataset. We also quantize and implement the entire eMamba pipeline on an AMD ZCU102 FPGA and ASIC using GlobalFoundries (GF) 22 nm technology. Experimental results show 4.95–5.62× lower latency and 2.22–9.95× higher throughput, with 4.77× smaller area, 9.84× lower power, and 48.6× lower energy consumption than baseline solutions while maintaining competitive accuracy. Alish Kanani, Ümit Y. Ogras, Jaehyun Park 0005 |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2024 | Thermal Modeling and Management Challenges in Heterogenous Integration: 2.5D Chiplet Platforms and BeyondabstractHeterogeneous integration using 2.5D chiplet platforms provides a new avenue for compact scale-out implementations of emerging applications, such as deep learning (DL). Integrating multiple small chiplets using a Network-on-Interposer (NoI) offers not only significant cost reductions and higher manufacturing yield compared to 2D ICs but also better thermal efficiency than 3D ICs and easier heterogeneous integration. However, dense integration and substantial compute density exacerbate thermal design problems, threatening to undermine the potential performance and cost benefits. Due to the significant role of temperature in the operation and reliability of integrated systems, it is critical to understand the role of heat in this emerging design area. However, little work has considered the thermal consequences of closely packaging a large number of computational elements. This paper overviews the thermal modeling challenges for chiplet-based 2.5D platforms, overviews existing approaches, and discusses the opportunities enabled by fast and accurate thermal models. Jaehyun Park 0005, Alish Kanani, Lukas Pfromm, Parth Solanki, Eric Tervo, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras |
VTS | 1 |
| 2022 | ECO: Enabling Energy-Neutral IoT Devices Through Runtime Allocation of Harvested EnergyabstractEnergy harvesting offers an attractive and promising mechanism to power low-energy devices. However, it alone is insufficient to enable an energy-neutral operation, which can eliminate tedious battery charging and replacement requirements. Achieving an energy-neutral operation is challenging since the uncertainties in harvested energy undermine the quality of service requirements. To address this challenge, we present a runtime energy-allocation framework that optimizes the utility of the target device under energy constraints using a rollout algorithm, which is a sequential approach to solve dynamic optimization problems. The proposed framework uses an efficient iterative algorithm to compute initial energy allocations at the beginning of a day. The initial allocations are then corrected at every interval to compensate for the deviations from the expected energy harvesting pattern. We evaluate this framework using solar and motion energy harvesting modalities andAmerican Time Use Surveydata from 4772 different users. Compared to prior techniques, the proposed framework achieves up to 35% higher utility even under energy-limited scenarios. Moreover, measurements on a wearable device prototype show that the proposed framework has$1000\times $smaller energy overhead than iterative approaches with a negligible loss in utility. Yigit Tuncel, Ganapati Bhat, Jaehyun Park 0005, Ümit Y. Ogras |
IEEE Internet Things J. | 3 |
| 2020 | AxFTL: Exploiting Error Tolerance for Extending Lifetime of NAND Flash StorageabstractNAND flash storage has become a standard choice in consumer electronics and is gaining popularity in enterprise systems due to its superior performance and low-power consumption. While its cost disadvantage is rapidly fading thanks to multibit cell technologies and 3-D stacking architectures, the challenge of limited endurance is still lingering and is expected to become more daunting as bits-per-cell continues to increase. In this article, we propose a novel flash translation layer (FTL) design named AxFTL (Approximate FTL) that extends the lifetime of NAND flash storage for error-tolerant applications. For error-tolerant data, AxFTL adopts shallow erase that lowers erase voltage to reduce the erase-induced wearing at the cost of an increased error rate. AxFTL manages multiple groups of blocks by error rates and allocates them according to the error tolerance of write requests. The key components of AxFTL include error tolerance-aware garbage collection and wear leveling schemes that manage the blocks with different error rates with minimal overhead. We implement AxFTL in an SSD simulator for the evaluation of the lifetime improvement and the actual allocation of the blocks. For application-level evaluation, we apply AxFTL to compressed video storage and evaluate the quality of video playback. Our experimental results show that AxFTL greatly improves the lifetime of NAND flash storage by 61% while maintaining a high structural similarity (SSIM) of 0.86 as compared to the conventional FTL. Jaehyun Park 0005, Junhee Ryu, Younghyun Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Instinctive Assistive Indoor Navigation using Distributed IntelligenceabstractCyber-physical systems (CPS) and the Internet of Things (IoT) offer a significant potential to improve the effectiveness of assistive technologies for those with physical disabilities. Practical assistive technologies should minimize the number of inputs from users to reduce their cognitive and physical effort. This article presents an energy-efficient framework and algorithm for assistive indoor navigation with multi-modal user input. The goal of the proposed framework is to simplify the navigation tasks and make them more instinctive for the user. Our framework automates indoor navigation using only a few user commands captured through a wearable device. The proposed methodology is evaluated using both a virtual smart building and a prototype. The evaluations for three different floorplans show one order of magnitude reduction in user effort and communication energy required for navigation, when compared to conventional navigation methodologies that require continuous user inputs. Md Muztoba, Rohit Voleti, Fatih Karabacak, Jaehyun Park 0005, Ümit Y. Ogras |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2017 | Near-optimal energy allocation for self-powered wearable systemsabstractWearable internet of things (IoT) devices are becoming popular due to their small form factor and low cost. Potential applications include human health and activity monitoring by embedding sensors such as accelerometer, gyroscope, and heart rate sensor. However, these devices have severely limited battery capacity, which requires frequent recharging. Harvesting ambient energy and optimal energy allocation can make wearable IoT devices practical by eliminating the charging requirement. This paper presents a near-optimal runtime energy management technique by considering the harvested energy. The proposed solution maximizes the performance of the wearable device under minimum energy constraints. We show that the results of the proposed algorithm are, on average, within 3% of the optimal solution computed offline. Ganapati Bhat, Jaehyun Park 0005, Ümit Y. Ogras |
ICCAD | 2 |
| 2017 | Flexible PV-cell Modeling for Energy Harvesting in Wearable IoT ApplicationsabstractWearable devices with sensing, processing and communication capabilities have become feasible with the advances in internet-of-things (IoT) and low power design technologies. Energy harvesting is extremely important for wearable IoT devices due to size and weight limitations of batteries. One of the most widely used energy harvesting sources is photovoltaic cell (PV-cell) owing to its simplicity and high output power. In particular, flexible PV-cells offer great potential for wearable applications. This paper models, for the first time , how bending a PV-cell significantly impacts the harvested energy. Furthermore, we derive an analytical model to quantify the harvested energy as a function of the radius of curvature. We validate the proposed model empirically using a commercial PV-cell under a wide range of bending scenarios, light intensities and elevation angles. Finally, we show that the proposed model can accelerate maximum power point tracking algorithms and increase the harvested energy by up to 25.0%. Jaehyun Park 0005, Hitesh Joshi, Hyung Gyu Lee, Sayfe Kiaei, Ümit Y. Ogras |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2017 | HoPE: Hot-Cacheline Prediction for Dynamic Early Decompression in Compressed LLCsabstractData compression plays a pivotal role in improving system performance and reducing energy consumption, because it increases the logical effective capacity of a compressed memory system without physically increasing the memory size. However, data compression techniques incur some cost, such as non-negligible compression and decompression overhead. This overhead becomes more severe if compression is used in the cache. In this article, we aim to minimize the read-hit decompression penalty in compressed Last-Level Caches (LLCs) by speculatively decompressing frequently used cachelines. To this end, we propose a Hot-cacheline Prediction and Early decompression (HoPE) mechanism that consists of three synergistic techniques: Hot-cacheline Prediction (HP), Early Decompression (ED), and Hit-history-based Insertion (HBI). HP and HBI efficiently identify the hot compressed cachelines, while ED selectively decompresses hot cachelines, based on their size information. Unlike previous approaches, the HoPE framework considers the performance balance/tradeoff between the increased effective cache capacity and the decompression penalty. To evaluate the effectiveness of the proposed HoPE mechanism, we run extensive simulations on memory traces obtained from multi-threaded benchmarks running on a full-system simulation framework. We observe significant performance improvements over compressed cache schemes employing the conventional Least-Recently Used (LRU) replacement policy, the Dynamic Re-Reference Interval Prediction (DRRIP) scheme, and the Effective Capacity Maximizer (ECM) compressed cache management mechanism. Specifically, HoPE exhibits system performance improvements of approximately 11%, on average, over LRU, 8% over DRRIP, and 7% over ECM by reducing the read-hit decompression penalty by around 65%, over a wide range of applications. Jaehyun Park 0005, Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Vinson Young, Junghee Lee 0004, Jongman Kim |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2016 | Multi-objective design optimization for flexible hybrid electronicsabstractFlexible systems that can conform to any shape are desirable for wearable applications. Over the past decade, there have been tremendous advances in the domain of flexible electronics which enabled printing of devices, such as sensors on a flexible substrate. Despite these advances, pure flexible electronics systems are limited by poor performance and large feature sizes. Flexible hybrid electronics (FHE) is an emerging technology which addresses these issues by integrating high performance rigid integrated circuits and flexible devices. Yet, there are no system-level design flows and algorithms for the design of FHE systems. To this end, this paper presents a multi-objective design algorithm to implement a target application optimally using a library of rigid and flexible components. Our algorithm produces a set of Pareto frontiers that optimize the physical flexibility, energy per operation and area metrics. Simulation studies show a 32× range in area and 4× range in flexibility across the set of Pareto-optimal design points. Ganapati Bhat, Ujjwal Gupta, Jaehyun Park 0005, Sule Ozev, Ümit Y. Ogras |
ICCAD | 4 |
| 2015 | Prefetch-based dynamic row buffer management for LPDDR2-NVM devicesabstractLPDDR2-NVM has been announced as an industry standard to efficiently interface with non-volatile memory devices such as phase change memory (PCM). This standard interface has been adopted in most commercial PCM devices. In this paper, we devise a prefetch-based dynamic row buffer management that targets the LPDDR2-NVM devices for enhancing performance with almost negligible implementation overhead. Our extensive simulations with timing parameters from the industry's commercial PCM devices demonstrate that the proposed method enhances the performance of memory systems up to 11.3% when compared with the static optimum configuration with fairly low-cost overheads. Jaehyun Park 0005, Donghwa Shin, Hyung Gyu Lee |
VLSI-SoC | 1 |
| 2015 | Design space exploration of row buffer architecture for phase change memory with LPDDR2-NVM interfaceabstractPhase change memory (PCM) is an attractive candidate for the future memory, but it still has several limitations to overcome such as write latency and long-term endurance. A large body of literature has been dedicated to solving these problems. However, almost all of the previous studies did not consider an important practical aspect of the PCM - an interface. The LPDDR2-NVM standard interface recently introduced by JEDEC is widely adopted by the manufacturers of commercial PCM these days. The LPDDR2-NVM standard allows a more flexible use of row buffers compared to the conventional DRAM interface. In this paper, we explore the design space of row buffer architecture in the PCM with LPDDR2-NVM interface. The effect of row buffer architecture on memory performance is investigated in terms of unit size and number of RDBs, and its management policy. We use the timing parameters from industry prototype PCM and analyze the result from the perspective of Pareto's optimum. The experimental results show that a properly-designed row buffer architecture enhances system-level performance up to 44.2% even at the same cost. Jaehyun Park 0005, Donghwa Shin, Hyung Gyu Lee |
VLSI-SoC | 1 |
| 2013 | Accurate Modeling of the Delay and Energy Overhead of Dynamic Voltage and Frequency Scaling in Modern MicroprocessorsabstractDynamic voltage and frequency scaling (DVFS) has been studied for well over a decade. Nevertheless, existing DVFS transition overhead models suffer from significant inaccuracies; for example, by incorrectly accounting for the effect of DC-DC converters, frequency synthesizers, voltage, and frequency change policies on energy losses incurred during mode transitions. Incorrect and/or inaccurate DVFS transition overhead models prevent one from determining the precise break-even time and thus forfeit some of the energy saving that is ideally achievable. This paper introduces accurate DVFS transition overhead models for both energy consumption and delay. In particular, we redefine the DVFS transition overhead including the underclocking-related losses in a DVFS-enabled microprocessor, additional inductor IR losses, and power losses due to discontinuous-mode DC-DC conversion. We report the transition overheads for a desktop, a mobile and a low-power representative processor. We also present DVFS transition overhead macromodel for use by high-level DVFS schedulers. Sangyoung Park, Jaehyun Park 0005, Donghwa Shin, Yanzhi Wang 0001, Qing Xie 0001, Massoud Pedram, Naehyuck Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Control-theoretic cyber-physical system modeling and synthesis: A case study of an active direct methanol fuel cellabstractA joint optimization of the physical system and the cyber world is one of the key problems in the design of a cyber-physical system (CPS). The major mechanical forces and/or chemical reactions in a plant are commonly modified by actuators in the balance-of-plant (BOP) system. More powerful actuators requires more power, but generally increase the response of the physical system powered by the electrical energy generated by the physical system. To maximize the overall output of a power generating plant therefore requires joint optimization of the physical system and the cyber world, and this is a key factor in the design of a CPS. We introduce a systematic approach to the modeling and synthesis of a CPS that emphasize joint power optimization, using an active direct methanol fuel cell (DMFC) as a case study. Active DMFC systems are superior to passive DMFCs in terms of fuel efficiency thanks to their BOP system, which includes pumps, air blowers, and fans. However, designing a small-scale active DMFC with the best overall system efficiency requires the BOP system to be jointly optimized with the DMFC stack operation, because the BOP components are powered by the stack. Our approach to this synthesis problem involves i) BOP system characterization, ii) integrated DMFC system modeling, iii) configuring a system for the maximum net power output through design space exploration, iv) synthesis of feedback control tasks, and v) implementation. Donghwa Shin, Jaehyun Park 0005, Younghyun Kim 0001, Jaeam Seo, Naehyuck Chang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2010 | Accurate modeling and calculation of delay and energy overheads of dynamic voltage scaling in modern high-performance microprocessorsabstractDynamic voltage and frequency scaling (DVS) has been studied for well over a decade, and even commercial systems widely support DVS nowadays. Nevertheless, existing DVS transition overhead models do not accurately reflect modern DVS architectures including modern DC-DC converters, PLL (Phase Lock Loop), and voltage and frequency change policies. Incorrect DVS overhead models prevent one from achieving the maximum energy gain, by misleading the DVS control policies. This paper introduces an accurate DVS overhead model, in terms of both energy consumption and time penalty, through detailed observation of modern DVS setups and voltage and frequency change guidelines from vendors. We introduce new major contributors to the DVS overhead including the performance underdrive loss of the DVS-enabled microprocessor, additional inductor IR loss, and so on, as well as consideration of power efficiency from discontinuous-mode DC-DC conversion. Our DVS overhead model enhances the DVS overhead model accuracy from 86% to 238% for Intel Core2 Duo E6850 and LTC3733. Jaehyun Park 0005, Donghwa Shin, Naehyuck Chang, Massoud Pedram |
ISLPED | 1 |
| 2008 | Energy and Performance Optimization of Demand Paging With OneNAND FlashabstractNew fusion memory devices consisting of multiple heterogeneous memory components in a single die or package offer efficient ways to optimize embedded systems in terms of energy, performance, and cost. Samsung Electronics recently announced the OneNAND fusion memory, in which a NAND flash array is integrated with dual SRAM buffers to provide a nor-type I/O interface. OneNAND has the low cost and large capacity of a NAND flash but also permits eXecution-in-Place (XIP) like a nor flash. The deployment of such devices requires careful system-level resource management because of their impact on energy consumption and performance, and existing memory optimization techniques, such as the demand paging used with NAND flash, may no longer be appropriate for systems with a fusion memory. We introduce a new online demand paging scheme that fully exploits the XIP capability of OneNAND flash by classifying pages as load preferred (residing in the on-chip SRAM) and XIP preferred (accessed directly from the OneNAND flash and discarded after use). This achieves, on average, a 26% reduction in energy consumption and a 19% increase in performance, compared with conventional NAND flash demand paging. Yongsoo Joo, Yongseok Choi, Jaehyun Park 0005, Chanik Park, Sung Woo Chung, Eui-Young Chung, Naehyuck Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | An energy characterization platform for memory devices and energy-aware data compression for multilevel-cell flash memoryabstractMemory devices often consume more energy than microprocessors in current portable embedded systems, but their energy consumption changes significantly with the type of transaction, data values, and access timing, as well as depending on the total number of transactions. These variabilities mean that an innovative tool and framework are required to characterize modern memory devices running in embedded system architectures. We introduce an energy measurement and characterization platform for memory devices, and demonstrate an application to multilevel-cell (MLC) flash memories, in which we discover significant value-dependent programming energy variations. We introduce an energy-aware data compression method that minimizes the flash programming energy, rather than the size of the compressed data, which is formulated as an entropy coding with unequal bit-pattern costs. Deploying a probabilistic approach, we derive energy-optimal bit-pattern probabilities and expected values of the bit-pattern costs which are applicable to the large amounts of compressed data typically found in multimedia applications. Then we develop an energy-optimal prefix coding that uses integer linear programming, and construct a prefix-code table. From a consideration of Pareto-optimal energy consumption, we can make tradeoffs between data size and programming energy, such as a 41% energy savings for a 52% area overhead. Yongsoo Joo, Youngjin Cho, Donghwa Shin, Jaehyun Park 0005, Naehyuck Chang |
ACM Trans. Design Autom. Electr. Syst. | 4 |