EDBT 2026 Demo / reviewers in the wild / expert
Myeonggyun Han
dblp:187/4019
· DBLP profile ↗
13ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-1832-1032ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge DevicesabstractSpeculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps. Its effectiveness depends heavily on draft selection, motivating adaptive methods that exploit variation across inputs and generation stages. On memory-constrained edge devices, however, these methods often fail to improve end-to-end throughput due to the overhead of switching between draft models. We identify a key limitation in this setting: the mismatch between draft selection and draft availability under tight memory budgets. Eunjeong Kim, Yeong Jun Jeon, Myeonggyun Han |
LCTES | 3 |
| 2023 | MARF: A Memory-Aware CLFLUSH-Based Intra- and Inter-CPU Side-Channel Attack
Sowoong Kim, Myeonggyun Han, Woongki Baek |
ESORICS (3) | 2 |
| 2022 | DPrime+DAbort: A High-Precision and Timer-Free Directory-Based Side-Channel Attack in Non-Inclusive Cache Hierarchies using Intel TSXabstractRecent CPUs have begun to adopt non-inclusive cache hierarchies for more effective cache utilization. Non-inclusive cache hierarchies have an additional advantage in that they eliminate the vulnerability to cache-based side-channel attacks. In addition, precise timers are often disabled or added with noise to defeat timer-based side-channel attacks. With the combination of such countermeasures, existing cache- and directory-based side-channel attacks can robustly be defeated on commodity systems.In this work, we discover the vulnerability caused by the undocumented interactions between the coherence directories and Intel TSX transactions in latest Intel CPUs with non-inclusive cache hierarchies. Guided by the observation, we propose a high-precision and timer-free directory attack called DPrime+DAbort in non-inclusive cache hierarchies using Intel TSX, which nullifies the aforementioned countermeasures. Our quantitative evaluation conducted on real systems equipped with latest Intel CPUs in three different generations demonstrates the practicality of the DPrime+DAbort attack in that it can be used to attack cryptographic and genomesequencing applications. We also discuss potential countermeasures and evaluate the feasibility of an Intel TSX-based countermeasure against the DPrime+DAbort attack. Sowoong Kim, Myeonggyun Han, Woongki Baek |
HPCA | 2 |
| 2022 | SDRP: Safe, Efficient, and SLO-Aware Workload Consolidation Through Secure and Dynamic Resource PartitioningabstractWorkload consolidation is a widely-used technique to improve the resource utilization of services computing systems by consolidating latency-critical (LC) and batch workloads on the same physical server. The resource manager for workload consolidation dynamically allocates hardware resources (e.g., cores, caches) to the workloads to maximize the resource utilization while satisfying the service-level objective (SLO) of the LC workloads. Since security-critical hardware resources are dynamically allocated across consolidated workloads, information leakages can be created among workloads through microarchitectural side-channel (SC) attacks. Despite extensive prior works, it is yet to investigate efficient system software support for achieving high resource utilization without compromising the SLO and security of consolidated workloads. To bridge this gap, we propose SDRP, secure and dynamic resource partitioning for safe, efficient, and SLO-aware workload consolidation. As with the state-of-the-art techniques, SDRP dynamically allocates hardware resources to enhance the resource utilization and provide the SLO guarantees. In contrast to the state-of-the-art techniques, SDRP dynamically sanitizes security-critical hardware resources to robustly defeat microarchitectural SC attacks. Our quantitative evaluation demonstrates that SDRP achieves high resource sanitization quality, introduces low performance overheads, delivers high resource utilization with the SLO and security guarantees, and defeats the last-level cache (LLC)-based SC attack. Myeonggyun Han, Woongki Baek |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | HERTI: A Reinforcement Learning-Augmented System for Efficient Real-Time Inference on Heterogeneous Embedded SystemsabstractReal-time inference is the key technology that enables a variety of latency-critical intelligent services such as autonomous driving and augmented reality. Heterogeneous embedded systems that consist of various computing devices with widely-different architectural and system-level characteristics are emerging as a promising solution for real-time inference. Despite extensive prior works, it still remains unexplored to design and implement a practical system that enables efficient real-time inference on heterogeneous embedded systems. To bridge this gap, we propose HERTI, a reinforcement learning-augmented system for efficient real-time inference on heterogeneous embedded systems. HERTI efficiently explores the state space and robustly finds an efficient state that significantly improves the efficiency of the target inference workload while satisfying its deadline constraint through reinforcement learning. Our quantitative evaluation conducted on a real heterogeneous embedded system demonstrates the effectiveness of HERTI in that HERTI achieves high inference efficiency in multiple metrics (i.e., energy and energy-delay product) with a strong deadline guarantee in contrast to the state-of-the-art techniques, delivers larger gains as the inference deadline and the system heterogeneity increase, provides strong generality for hyper-parameter tuning, and significantly reduces the training time through its estimation-based approach across all the evaluated inference workloads and scenarios. Myeonggyun Han, Woongki Baek |
PACT | 1 |
| 2021 | PALM: Progress- and Locality-Aware Adaptive Task Migration for Efficient Thread PackingabstractThread packing (TP) is an effective and widely-used technique to significantly improve the efficiency of parallel systems by dynamically controlling the number of cores allocated to multithreaded applications based on their requirements such as performance and energy efficiency. Despite the extensive prior works on TP, little work has been done to investigate and address its performance inefficiencies that arise across various parallel systems and applications with different characteristics. To bridge this gap, we investigate the performance inefficiencies of TP using a wide range of parallel applications and system configurations and identify their root causes. Guided by the in-depth performance characterization results, we propose PALM, progress- and locality-aware adaptive task migration for efficient TP. Through quantitative evaluation, we demonstrate that PALM achieves significantly higher performance and lower energy consumption than TP across various synchronization-intensive applications and system configurations, provides the performance and energy consumption comparable with the thread reduction technique, and considerably improves the efficiency of dynamic server consolidation and the performance under power capping. Jinsu Park, Seongbeom Park, Myeonggyun Han, Woongki Baek |
IPDPS | 3 |
| 2021 | Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel ApplicationsabstractHeterogeneous multiprocessing (HMP) is an emerging technology for high-performance and energy-efficient computing. While task parallelism is widely used in various computing domains, such as embedded, big-data, and machine-learning computing domains, it still remains unexplored to investigate the efficient runtime support that effectively utilizes the criticality of the tasks of the target application and the heterogeneity of the underlying HMP system with full resource management. To bridge this gap, we propose CHRT, a criticality- and heterogeneity-aware runtime system for task-parallel applications. CHRT dynamically estimates the performance and power consumption of the target task-parallel application and robustly manages the full HMP system resources (i.e., core types, counts, and voltage/frequency levels) to maximize the overall efficiency. Our quantitative evaluation based on widely-used task parallel benchmarks and two full HMP systems (i.e., the XU3 and HiKey970 HMP systems) demonstrates the effectiveness of CHRT in that CHRT achieves significantly higher energy (e.g., 60.4 and 57.2 percent on average on the XU3 system) and energy-delay product (e.g., 52.2 and 44.0 percent on average on the HiKey970 system) efficiency than the baseline runtime system that employs the breadth-first scheduler and the state-of-the-art criticality-aware runtime system and incurs low performance overheads. Myeonggyun Han, Jinsu Park, Woongki Baek |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Hotness- and Lifetime-Aware Data Placement and Migration for High-Performance Deep Learning on Heterogeneous Memory SystemsabstractHeterogeneous memory systems that comprise memory nodes with disparate architectural characteristics (e.g., DRAM and high-bandwidth memory (HBM)) have surfaced as a promising solution in a variety of computing domains ranging from embedded to high-performance computing. Since deep learning (DL) is one of the most widely-used workloads in various computing domains, it is crucial to explore efficient memory management techniques for DL applications that execute on heterogeneous memory systems. Despite extensive prior works on system software and architectural support for efficient DL, it still remains unexplored to investigate heterogeneity-aware memory management techniques for high-performance DL on heterogeneous memory systems. To bridge this gap, we analyze the characteristics of representative DL workloads on a real heterogeneous memory system. Guided by the characterization results, we propose HALO, hotness- and lifetime-aware data placement and migration for high-performance DL on heterogeneous memory systems. Through quantitative evaluation, we demonstrate the effectiveness of HALO in that it significantly outperforms various memory management policies (e.g., 28.2 percent higher performance than the HBM-Preferred policy) supported by the underlying system software and hardware, achieves the performance comparable to the ideal case with infinite HBM, incurs small performance overheads, and delivers high performance across a wide range of application working-set sizes. Myeonggyun Han, Jihoon Hyun, Seongbeom Park, Woongki Baek |
IEEE Trans. Computers | 1 |
| 2019 | MOSAIC: Heterogeneity-, Communication-, and Constraint-Aware Model Slicing and Execution for Accurate and Efficient InferenceabstractHeterogeneous embedded systems have surfaced as a promising solution for accurate and efficient deep-learning inference on mobile devices. Despite extensive prior works, it still remains unexplored to investigate the system-software support that efficiently executes inference workloads by judiciously considering their performance and energy heterogeneity, communication overheads, and constraints. To bridge this gap, we propose MOSAIC, heterogeneity-, communication-, and constraint-aware model slicing and execution for accurate and efficient inference on heterogeneous embedded systems. MOSAIC generates the efficient model slicing and execution plan for the target inference workload through dynamic programming. MOSAIC significantly reduces inference latency and energy, exhibits high estimation accuracy, and incurs small overheads. Myeonggyun Han, Jihoon Hyun, Seongbeom Park, Jinsu Park, Woongki Baek |
PACT | 1 |
| 2019 | POSTER: The Performance Impact of Thread Packing on Synchronization-Intensive ApplicationsabstractThread packing (TP) is a widely-used technique to improve the efficiency of parallel systems. Despite extensive prior works, relatively little work has been done to investigate its performance inefficiencies. To bridge this gap, we quantify its performance impact on synchronization-intensive applications and identify the root causes of its performance inefficiencies. Jinsu Park, Seongbeom Park, Myeonggyun Han, Woongki Baek |
PACT | 3 |
| 2018 | Hypart: a hybrid technique for practical memory bandwidth partitioning on commodity serversabstractMemory bandwidth is a highly performance-critical shared resource on modern computer systems. To prevent the contention on memory bandwidth among the collocated workloads, prior works have investigated memory bandwidth partitioning techniques. Despite the extensive prior works, it still remains unexplored to characterize the widely-used memory bandwidth partitioning techniques based on various metrics and investigate a hybrid technique that employs multiple memory bandwidth partitioning techniques to improve the overall efficiency. Jinsu Park, Seongbeom Park, Myeonggyun Han, Jihoon Hyun, Woongki Baek |
PACT | 3 |
| 2018 | Secure and Dynamic Core and Cache Partitioning for Safe and Efficient Server ConsolidationabstractWith server consolidation, latency-critical and batch workloads are collocated on the same physical servers. The resource manager dynamically allocates the hardware resources to the workloads to maximize the overall throughput while providing the service-level objective (SLO) guarantees for the latency-critical workloads. As the hardware resources are dynamically allocated across the workloads on the same physical server, information leakage can be established, making them vulnerable to micro-architectural side-channel attacks. Despite extensive prior works, it remains unexplored to investigate the efficient design and implementation of the dynamic resource management system that maximizes resource efficiency without compromising the SLO and security guarantees. To bridge this gap, this work proposes SDCP, secure and dynamic core and cache partitioning for safe and efficient server consolidation. In line with the state-of-the-art dynamic server consolidation techniques, SDCP dynamically allocates the hardware resources (i.e., cores and caches) to maximize the resource utilization with the SLO guarantees. In contrast to the existing techniques, however, SDCP dynamically sanitizes the hardware resources to ensure that no micro-architectural side channel is established between different security domains. Our experimental results demonstrate that SDCP provides high resource sanitization quality, incurs small performance overheads, and achieves high resource efficiency with the SLO and security guarantees. Myeonggyun Han, Seongdae Yu, Woongki Baek |
CCGrid | 1 |
| 2017 | CHRT: A criticality- and heterogeneity-aware runtime system for task-parallel applicationsabstractHeterogeneous multiprocessing (HMP) is an emerging technology for high-performance and energy-efficient computing. While task parallelism is widely used in various computing domains from the embedded to machine-learning computing domains, relatively little work has been done to investigate the efficient runtime support that effectively utilizes the criticality of the tasks of the target application and the heterogeneity of the underlying HMP system with full resource management. To bridge this gap, we propose a criticality- and heterogeneity-aware runtime system for task-parallel applications (CHRT). CHRT dynamically estimates the performance and power consumption of the target task-parallel application and robustly manages the full HMP system resources (i.e., core types, counts, and voltage/frequency levels) to maximize the overall efficiency. Our experimental results show that CHRT achieves significantly higher energy efficiency than the baseline runtime system that employs the breadth-first scheduler and the state-of-the-art criticality-aware runtime system. Myeonggyun Han, Jinsu Park, Woongki Baek |
DATE | 1 |