Woongki Baek

dblp:37/6811 · DBLP profile ↗
← Back
42ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-1877-7307ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Harmonia: QoS-Aware and High-Throughput Generative Inference with a Single GPU
Sowoong Kim, Youngsam Shin, Yeongon Cho, Woongki Baek
Euro-Par (2)4
2026 Performance Analysis of Hardware-Accelerated Compressed Memory Swap
Eunyeong Sim, Woongki Baek
Euro-Par (1)2
2024 Activation Sequence Caching: High-Throughput and Memory-Efficient Generative Inference with a Single GPU
abstract
Generative artificial intelligence is widely used for various tasks such as language translation and art creation. Generative inference employs the key and value tensors to encode relational information among the tokens in the input and output sequences. Most of the existing generative inference frameworks use KV caching (KVC), which caches the key and value tensors (i.e., the KV cache) to avoid recomputing the key and value tensors for each of the processed tokens. Despite the widespread use of KVC, in-depth characterization of KVC with tensor offloading, which enables generative inference with a single GPU, remains yet to be explored.
Sowoong Kim, Eunyeong Sim, Youngsam Shin, Yeongon Cho, Woongki Baek
PACT5
2023 MARF: A Memory-Aware CLFLUSH-Based Intra- and Inter-CPU Side-Channel Attack
Sowoong Kim, Myeonggyun Han, Woongki Baek
ESORICS (3)3
2022 DPrime+DAbort: A High-Precision and Timer-Free Directory-Based Side-Channel Attack in Non-Inclusive Cache Hierarchies using Intel TSX
abstract
Recent CPUs have begun to adopt non-inclusive cache hierarchies for more effective cache utilization. Non-inclusive cache hierarchies have an additional advantage in that they eliminate the vulnerability to cache-based side-channel attacks. In addition, precise timers are often disabled or added with noise to defeat timer-based side-channel attacks. With the combination of such countermeasures, existing cache- and directory-based side-channel attacks can robustly be defeated on commodity systems.In this work, we discover the vulnerability caused by the undocumented interactions between the coherence directories and Intel TSX transactions in latest Intel CPUs with non-inclusive cache hierarchies. Guided by the observation, we propose a high-precision and timer-free directory attack called DPrime+DAbort in non-inclusive cache hierarchies using Intel TSX, which nullifies the aforementioned countermeasures. Our quantitative evaluation conducted on real systems equipped with latest Intel CPUs in three different generations demonstrates the practicality of the DPrime+DAbort attack in that it can be used to attack cryptographic and genomesequencing applications. We also discuss potential countermeasures and evaluate the feasibility of an Intel TSX-based countermeasure against the DPrime+DAbort attack.
Sowoong Kim, Myeonggyun Han, Woongki Baek
HPCA3
2022 SDRP: Safe, Efficient, and SLO-Aware Workload Consolidation Through Secure and Dynamic Resource Partitioning
abstract
Workload consolidation is a widely-used technique to improve the resource utilization of services computing systems by consolidating latency-critical (LC) and batch workloads on the same physical server. The resource manager for workload consolidation dynamically allocates hardware resources (e.g., cores, caches) to the workloads to maximize the resource utilization while satisfying the service-level objective (SLO) of the LC workloads. Since security-critical hardware resources are dynamically allocated across consolidated workloads, information leakages can be created among workloads through microarchitectural side-channel (SC) attacks. Despite extensive prior works, it is yet to investigate efficient system software support for achieving high resource utilization without compromising the SLO and security of consolidated workloads. To bridge this gap, we propose SDRP, secure and dynamic resource partitioning for safe, efficient, and SLO-aware workload consolidation. As with the state-of-the-art techniques, SDRP dynamically allocates hardware resources to enhance the resource utilization and provide the SLO guarantees. In contrast to the state-of-the-art techniques, SDRP dynamically sanitizes security-critical hardware resources to robustly defeat microarchitectural SC attacks. Our quantitative evaluation demonstrates that SDRP achieves high resource sanitization quality, introduces low performance overheads, delivers high resource utilization with the SLO and security guarantees, and defeats the last-level cache (LLC)-based SC attack.
Myeonggyun Han, Woongki Baek
IEEE Trans. Serv. Comput.2
2021 HERTI: A Reinforcement Learning-Augmented System for Efficient Real-Time Inference on Heterogeneous Embedded Systems
abstract
Real-time inference is the key technology that enables a variety of latency-critical intelligent services such as autonomous driving and augmented reality. Heterogeneous embedded systems that consist of various computing devices with widely-different architectural and system-level characteristics are emerging as a promising solution for real-time inference. Despite extensive prior works, it still remains unexplored to design and implement a practical system that enables efficient real-time inference on heterogeneous embedded systems. To bridge this gap, we propose HERTI, a reinforcement learning-augmented system for efficient real-time inference on heterogeneous embedded systems. HERTI efficiently explores the state space and robustly finds an efficient state that significantly improves the efficiency of the target inference workload while satisfying its deadline constraint through reinforcement learning. Our quantitative evaluation conducted on a real heterogeneous embedded system demonstrates the effectiveness of HERTI in that HERTI achieves high inference efficiency in multiple metrics (i.e., energy and energy-delay product) with a strong deadline guarantee in contrast to the state-of-the-art techniques, delivers larger gains as the inference deadline and the system heterogeneity increase, provides strong generality for hyper-parameter tuning, and significantly reduces the training time through its estimation-based approach across all the evaluated inference workloads and scenarios.
Myeonggyun Han, Woongki Baek
PACT2
2021 PALM: Progress- and Locality-Aware Adaptive Task Migration for Efficient Thread Packing
abstract
Thread packing (TP) is an effective and widely-used technique to significantly improve the efficiency of parallel systems by dynamically controlling the number of cores allocated to multithreaded applications based on their requirements such as performance and energy efficiency. Despite the extensive prior works on TP, little work has been done to investigate and address its performance inefficiencies that arise across various parallel systems and applications with different characteristics. To bridge this gap, we investigate the performance inefficiencies of TP using a wide range of parallel applications and system configurations and identify their root causes. Guided by the in-depth performance characterization results, we propose PALM, progress- and locality-aware adaptive task migration for efficient TP. Through quantitative evaluation, we demonstrate that PALM achieves significantly higher performance and lower energy consumption than TP across various synchronization-intensive applications and system configurations, provides the performance and energy consumption comparable with the thread reduction technique, and considerably improves the efficiency of dynamic server consolidation and the performance under power capping.
Jinsu Park, Seongbeom Park, Myeonggyun Han, Woongki Baek
IPDPS4
2021 Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel Applications
abstract
Heterogeneous multiprocessing (HMP) is an emerging technology for high-performance and energy-efficient computing. While task parallelism is widely used in various computing domains, such as embedded, big-data, and machine-learning computing domains, it still remains unexplored to investigate the efficient runtime support that effectively utilizes the criticality of the tasks of the target application and the heterogeneity of the underlying HMP system with full resource management. To bridge this gap, we propose CHRT, a criticality- and heterogeneity-aware runtime system for task-parallel applications. CHRT dynamically estimates the performance and power consumption of the target task-parallel application and robustly manages the full HMP system resources (i.e., core types, counts, and voltage/frequency levels) to maximize the overall efficiency. Our quantitative evaluation based on widely-used task parallel benchmarks and two full HMP systems (i.e., the XU3 and HiKey970 HMP systems) demonstrates the effectiveness of CHRT in that CHRT achieves significantly higher energy (e.g., 60.4 and 57.2 percent on average on the XU3 system) and energy-delay product (e.g., 52.2 and 44.0 percent on average on the HiKey970 system) efficiency than the baseline runtime system that employs the breadth-first scheduler and the state-of-the-art criticality-aware runtime system and incurs low performance overheads.
Myeonggyun Han, Jinsu Park, Woongki Baek
IEEE Trans. Parallel Distributed Syst.3
2021 Holistic VM Placement for Distributed Parallel Applications in Heterogeneous Clusters
abstract
In a heterogeneous cluster, virtual machine (VM) placement for a distributed parallel application is challenging due to numerous possible ways of placing the application and complexity of estimating the performance of the application. This study investigates a holistic VM placement technique for distributed parallel applications in a heterogeneous cluster, aiming to maximize the efficiency of the cluster and consequently reduce the costs for service providers and users. The proposed technique accommodates various factors that have an impact on performance in a combined manner. First, we analyze the effects of the heterogeneity of resources, different VM configurations, and interference between VMs on the performance of distributed parallel applications with a wide diversity of characteristics, including scientific and big data analytics applications. We then propose a placement technique that uses a machine learning algorithm to estimate the runtime of a distributed parallel application. To train a performance estimation model, a distributed parallel application is profiled against synthetic workloads that mostly utilize the dominant resource of the application, which strongly affects the application performance, reducing the profiling space dramatically. Through experimental and simulation studies, we show that the proposed placement technique can find good VM placement configurations for various workloads.
Seontae Kim, Nguyen Pham, Woongki Baek, Young-ri Choi
IEEE Trans. Serv. Comput.3
2020 Hotness- and Lifetime-Aware Data Placement and Migration for High-Performance Deep Learning on Heterogeneous Memory Systems
abstract
Heterogeneous memory systems that comprise memory nodes with disparate architectural characteristics (e.g., DRAM and high-bandwidth memory (HBM)) have surfaced as a promising solution in a variety of computing domains ranging from embedded to high-performance computing. Since deep learning (DL) is one of the most widely-used workloads in various computing domains, it is crucial to explore efficient memory management techniques for DL applications that execute on heterogeneous memory systems. Despite extensive prior works on system software and architectural support for efficient DL, it still remains unexplored to investigate heterogeneity-aware memory management techniques for high-performance DL on heterogeneous memory systems. To bridge this gap, we analyze the characteristics of representative DL workloads on a real heterogeneous memory system. Guided by the characterization results, we propose HALO, hotness- and lifetime-aware data placement and migration for high-performance DL on heterogeneous memory systems. Through quantitative evaluation, we demonstrate the effectiveness of HALO in that it significantly outperforms various memory management policies (e.g., 28.2 percent higher performance than the HBM-Preferred policy) supported by the underlying system software and hardware, achieves the performance comparable to the ideal case with infinite HBM, incurs small performance overheads, and delivers high performance across a wide range of application working-set sizes.
Myeonggyun Han, Jihoon Hyun, Seongbeom Park, Woongki Baek
IEEE Trans. Computers4
2019 MOSAIC: Heterogeneity-, Communication-, and Constraint-Aware Model Slicing and Execution for Accurate and Efficient Inference
abstract
Heterogeneous embedded systems have surfaced as a promising solution for accurate and efficient deep-learning inference on mobile devices. Despite extensive prior works, it still remains unexplored to investigate the system-software support that efficiently executes inference workloads by judiciously considering their performance and energy heterogeneity, communication overheads, and constraints. To bridge this gap, we propose MOSAIC, heterogeneity-, communication-, and constraint-aware model slicing and execution for accurate and efficient inference on heterogeneous embedded systems. MOSAIC generates the efficient model slicing and execution plan for the target inference workload through dynamic programming. MOSAIC significantly reduces inference latency and energy, exhibits high estimation accuracy, and incurs small overheads.
Myeonggyun Han, Jihoon Hyun, Seongbeom Park, Jinsu Park, Woongki Baek
PACT5
2019 POSTER: The Performance Impact of Thread Packing on Synchronization-Intensive Applications
abstract
Thread packing (TP) is a widely-used technique to improve the efficiency of parallel systems. Despite extensive prior works, relatively little work has been done to investigate its performance inefficiencies. To bridge this gap, we quantify its performance impact on synchronization-intensive applications and identify the root causes of its performance inefficiencies.
Jinsu Park, Seongbeom Park, Myeonggyun Han, Woongki Baek
PACT4
2019 CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers
abstract
Workload consolidation is a widely-used technique to maximize server resource utilization in cloud and datacenter computing. Recent commodity CPUs support last-level cache (LLC) and memory bandwidth partitioning functionalities that can be used to ensure the fairness of the consolidated workloads. While prior work has proposed a variety of resource partitioning techniques, it still remains unexplored to characterize the impact of LLC and memory bandwidth partitioning on the fairness of the consolidated workloads and investigate system software support to dynamically control LLC and memory bandwidth partitioning in a coordinated manner.
Jinsu Park, Seongbeom Park, Woongki Baek
EuroSys3
2019 Analyzing and optimizing the performance and energy efficiency of transactional scientific applications on large-scale NUMA systems with HTM support
Jinsu Park, Woongki Baek
J. Parallel Distributed Comput.2
2019 Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management
abstract
Hardware caches are widely employed in GPGPUs to achieve higher performance and energy efficiency. Incorporating hardware caches in GPGPUs, however, does not immediately guarantee enhanced performance and energy efficiency due to high cache contention and thrashing. To address the inefficiency of GPGPU caches, various adaptive techniques (e.g., warp limiting) have been proposed. However, relatively little work has been done in the context of creating an architectural framework that tightly integrates adaptive cache management techniques and investigating their effectiveness and interaction. To bridge this gap, we propose IACM, integrated adaptive cache management for high-performance and energy-efficient GPGPU computing. IACM integrates the state-of-the-art adaptive cache management techniques (i.e., cache indexing, bypassing, and warp limiting) in a unified architectural framework. Our quantitative evaluation demonstrates that IACM significantly improves the performance and energy efficiency of various GPGPU workloads over the baseline architecture (i.e., 98.1 and 61.9 percent on average, respectively), achieves considerably higher performance than the state-of-the-art technique (i.e., 361.4 percent at maximum and 7.7 percent on average), and delivers significant performance and energy-efficiency gains over the baseline GPGPU architecture enhanced with advanced architectural technologies.
Kyu Yeun Kim, Jinsu Park, Woongki Baek
IEEE Trans. Parallel Distributed Syst.3
2018 Hypart: a hybrid technique for practical memory bandwidth partitioning on commodity servers
abstract
Memory bandwidth is a highly performance-critical shared resource on modern computer systems. To prevent the contention on memory bandwidth among the collocated workloads, prior works have investigated memory bandwidth partitioning techniques. Despite the extensive prior works, it still remains unexplored to characterize the widely-used memory bandwidth partitioning techniques based on various metrics and investigate a hybrid technique that employs multiple memory bandwidth partitioning techniques to improve the overall efficiency.
Jinsu Park, Seongbeom Park, Myeonggyun Han, Jihoon Hyun, Woongki Baek
PACT5
2018 Secure and Dynamic Core and Cache Partitioning for Safe and Efficient Server Consolidation
abstract
With server consolidation, latency-critical and batch workloads are collocated on the same physical servers. The resource manager dynamically allocates the hardware resources to the workloads to maximize the overall throughput while providing the service-level objective (SLO) guarantees for the latency-critical workloads. As the hardware resources are dynamically allocated across the workloads on the same physical server, information leakage can be established, making them vulnerable to micro-architectural side-channel attacks. Despite extensive prior works, it remains unexplored to investigate the efficient design and implementation of the dynamic resource management system that maximizes resource efficiency without compromising the SLO and security guarantees. To bridge this gap, this work proposes SDCP, secure and dynamic core and cache partitioning for safe and efficient server consolidation. In line with the state-of-the-art dynamic server consolidation techniques, SDCP dynamically allocates the hardware resources (i.e., cores and caches) to maximize the resource utilization with the SLO guarantees. In contrast to the existing techniques, however, SDCP dynamically sanitizes the hardware resources to ensure that no micro-architectural side channel is established between different security domains. Our experimental results demonstrate that SDCP provides high resource sanitization quality, incurs small performance overheads, and achieves high resource efficiency with the SLO and security guarantees.
Myeonggyun Han, Seongdae Yu, Woongki Baek
CCGrid3
2018 RPPC: A Holistic Runtime System for Maximizing Performance Under Power Capping
abstract
Maximizing performance in power-constrained computing environments is highly important in cloud and datacenter computing. To achieve the best possible performance of parallel applications under power capping, it is crucial to execute them with the optimal concurrency level and cross-component power allocation between CPUs and memory. Despite extensive prior works, it still remains unexplored to investigate the efficient runtime support that maximizes the performance of parallel applications under power capping through the coordinated control of concurrency level and cross-component power allocation. To bridge this gap, this work proposes RPPC, a holistic runtime system for maximizing performance under power capping. In contrast to the state-of-the-art techniques, RPPC robustly controls the two performance-critical knobs (i.e., concurrency level and cross-component power allocation) in a coordinated manner to maximize the performance of parallel applications under power capping. RPPC dynamically identifies the characteristics of the target parallel application and explores the system state space to find an efficient system state. Our experimental results demonstrate that RPPC significantly outperforms the two state-of-the-art power-capping techniques, achieves the performance comparable with the static best version that requires extensive per-application offline profiling, incurs small performance overheads, and provides the re-adaptation mechanism to external events such as total power budget changes.
Jinsu Park, Seongbeom Park, Woongki Baek
CCGrid3
2018 CEML: a Coordinated Runtime System for Efficient Machine Learning on Heterogeneous Computing Systems
Jihoon Hyun, Jinsu Park, Kyu Yeun Kim, Seongdae Yu, Woongki Baek
Euro-Par5
2018 BLPP: Improving the Performance of GPGPUs with Heterogeneous Memory through Bandwidth- and Latency-Aware Page Placement
abstract
GPGPUs with heterogeneous memory have surfaced as a promising solution to improve the programmability and flexibility of GPGPU computing. Despite the extensive prior works, relatively little work has been done to investigate holistic system software support for heterogeneity-aware memory management. To bridge this gap, we propose bandwidth-and latency-aware page placement (BLPP) for GPGPUs with heterogeneous memory. BLPP dynamically places pages across the heterogeneous memory nodes by preserving the optimal allocation ratio computed based on their performance characteristics. Our experimental results show that BLPP considerably outperforms the state-of-the-art technique and performs similarly to the static-best version, which requires extensive offline profiling.
Kyu Yeun Kim, Woongki Baek
ICCD2
2018 Quantifying the Performance and Energy-Efficiency Impact of Hardware Transactional Memory on Scientific Applications on Large-Scale NUMA Systems
abstract
Hardware transactional memory (HTM) is supported by widely-used commodity processors. While the effectiveness of HTM has been evaluated based on small-scale multi-core systems, it still remains unexplored to quantify the performance and energy-efficiency of HTM for scientific workloads on large-scale NUMA systems, which have been increasingly adopted to high-performance computing. To bridge this gap, this work investigates the performance and energy-efficiency impact of HTM on scientific applications on large-scale NUMA systems. We first quantify the performance and energy efficiency of HTM for scientific workloads based on the widely-used CLOMP-TM benchmark. We then discuss a set of generic software optimizations that can be effectively used to improve the performance and energy efficiency of transactional scientific workloads on large-scale NUMA systems. Finally, we present case studies in which we apply a set of the optimizations to representative transactional scientific applications and significantly optimize their performance and energy efficiency on large-scale NUMA systems.
Jinsu Park, Woongki Baek
IPDPS2
2017 Failure-Atomic Slotted Paging for Persistent Memory
abstract
The slotted-page structure is a database page format commonly used for managing variable-length records. In this work, we develop a novel "failure-atomic slotted page structure" for persistent memory that leverages byte addressability and durability of persistent memory to minimize redundant write operations used to maintain consistency in traditional database systems. Failure-atomic slotted paging consists of two key elements: (i) in-place commit per page using hardware transactional memory and (ii) slot header logging that logs the commit mark of each page. The proposed scheme is implemented in SQLite and compared against NVWAL, the current state-of-the-art scheme. Our performance study shows that our failure-atomic slotted paging shows optimal performance for database transactions that insert a single record. For transactions that touch more than one database page, our proposed slot-header logging scheme minimizes the logging overhead by avoiding duplicating pages and logging only the metadata of the dirty pages. Overall, we find that our failure-atomic slotted-page management scheme reduces database logging overhead to 1/6 and improves query response time by up to 33% compared to NVWAL.
Jihye Seo, Wook-Hee Kim, Woongki Baek, Beomseok Nam, Sam H. Noh
ASPLOS3
2017 CHRT: A criticality- and heterogeneity-aware runtime system for task-parallel applications
abstract
Heterogeneous multiprocessing (HMP) is an emerging technology for high-performance and energy-efficient computing. While task parallelism is widely used in various computing domains from the embedded to machine-learning computing domains, relatively little work has been done to investigate the efficient runtime support that effectively utilizes the criticality of the tasks of the target application and the heterogeneity of the underlying HMP system with full resource management. To bridge this gap, we propose a criticality- and heterogeneity-aware runtime system for task-parallel applications (CHRT). CHRT dynamically estimates the performance and power consumption of the target task-parallel application and robustly manages the full HMP system resources (i.e., core types, counts, and voltage/frequency levels) to maximize the overall efficiency. Our experimental results show that CHRT achieves significantly higher energy efficiency than the baseline runtime system that employs the breadth-first scheduler and the state-of-the-art criticality-aware runtime system.
Myeonggyun Han, Jinsu Park, Woongki Baek
DATE3
2017 Machine-Learning Based Performance Estimation for Distributed Parallel Applications in Virtualized Heterogeneous Clusters
abstract
In a virtualized heterogeneous cluster, for a distributed parallel application which runs in multiple virtual machines (VMs) concurrently, there are a huge number of possible ways to place its VMs. This paper investigates a performance estimation technique for distributed parallel applications in virtualized heterogeneous clusters. We first analyze the effects of different VM configurations on the performance of various distributed parallel applications. We then present a machine-learning based performance model for a distributed parallel application. Using a heterogeneous cluster with two different types of nodes, we show that our machine-learning based models can estimate the runtimes of distributed parallel applications with modest error rates.
Seontae Kim, Nguyen Pham, Woongki Baek, Young-ri Choi
ICDCS3
2017 Design and implementation of bandwidth-aware memory placement and migration policies for heterogeneous memory systems
abstract
Heterogeneous memory systems that comprise memory nodes based on widely-different device technologies (e.g., DRAM and nonvolatile memory (NVM)) are emerging in various computing domains ranging from high-performance to embedded computing. Despite the extensive prior work on architectural and system software support for heterogeneous memory systems, relatively little work has been done to investigate the OS-level memory placement and migration policies that consider the bandwidth differences of heterogeneous memory nodes.
Seongdae Yu, Seongbeom Park, Woongki Baek
ICS3
2016 NVWAL: Exploiting NVRAM in Write-Ahead Logging
abstract
Emerging byte-addressable non-volatile memory is considered an alternative storage device for database logs that require persistency and high performance. In this work, we develop NVWAL (NVRAM Write-Ahead Logging) for SQLite. The contribution of NVWAL consists of three elements: (i) byte-granularity differential logging that effectively eliminates the excessive I/O overhead of filesystem-based logging or journaling, (ii) transaction-aware lazy synchronization that reduces cache synchronization overhead by two-thirds, and (iii) user-level heap management of the NVRAM persistent WAL structure, which reduces the overhead of managing persistent objects.
Wook-Hee Kim, Jinwoong Kim, Woongki Baek, Beomseok Nam, Youjip Won
ASPLOS3
2016 RMC: an integrated runtime system for adaptive many-core computing
abstract
Many-core computing has surfaced as a promising solution to satisfy the rapidly increasing computational needs for various areas ranging from embedded to datacenter computing. However, when allocated with an excessive number of cores, multithreaded applications may fail to achieve optimal performance and energy efficiency due to the contention on software and/or hardware resources. While previous research has proposed adaptive techniques such as thread packing (TP) and dynamic threading (DT), they often lead to suboptimal results because they are used in an isolated manner. To address this problem, we propose RMC, an integrated runtime system for adaptive many-core computing. Guided by the runtime information of parallel applications, RMC dynamically adapts their execution by combining the TP and DT techniques. We apply RMC to six PARSEC benchmarks that use representative parallelism models (i.e., fork-join, task, and pipeline). We demonstrate that RMC is easy to use, considerably outperforms the state-of-the-art techniques for three PARSEC benchmarks, and incurs a small overhead to the rest of the benchmarks.
Jinsu Park, Eunbi Cho, Woongki Baek
EMSOFT3
2016 HAP: A Heterogeneity-Conscious Runtime System for Adaptive Pipeline Parallelism
Jinsu Park, Woongki Baek
Euro-Par2
2016 IACM: Integrated adaptive cache management for high-performance and energy-efficient GPGPU computing
abstract
Hardware caches are widely employed in GPGPUs to achieve higher performance and energy efficiency. Incorporating hardware caches in GPGPUs, however, does not immediately guarantee enhanced performance and energy efficiency due to high cache contention and thrashing. To address the inefficiency of GPGPU caches, various adaptive techniques (e.g., warp limiting) have been proposed. However, relatively little work has been done in the context of creating an architectural framework that tightly integrates adaptive cache management techniques and investigating their effectiveness and interaction. To bridge this gap, we propose IACM, integrated adaptive cache management for high-performance and energy-efficient GPGPU computing. IACM integrates the state-of-the-art adaptive cache management techniques (i.e., cache indexing, bypassing, and warp limiting) in a unified architectural framework. Our quantitative evaluation demonstrates that IACM significantly improves the performance and energy efficiency of various GPGPU workloads over the baseline architecture (i.e., 98.1% and 61.9% on average).
Kyu Yeun Kim, Jinsu Park, Woongki Baek
ICCD3
2016 RCHC: A Holistic Runtime System for Concurrent Heterogeneous Computing
abstract
Concurrent heterogeneous computing (CHC) is rapidly emerging as a promising solution for high-performance and energy-efficient computing. The fundamental challenges for efficient CHC are how to partition the workload of the target application across the devices in the underlying CHC system and how to control the operating frequency of each device in order to maximize the overall efficiency. Despite the extensive prior work on the system software techniques for CHC, efficient runtime support for CHC that robustly supports both functional and performance heterogeneity without the need for extensive offline profiling still remains unexplored. To bridge this gap, we propose RCHC, a holistic runtime system for concurrent heterogeneous computing. RCHC dynamically profiles the target application and constructs the performance and power estimation models based on the runtime information. Guided by the estimation models, RCHC explores the system state space, determines the best system state that is expected to maximize the efficiency of the target application, and accordingly executes it. Our experimental results demonstrate that RCHC significantly outperforms the baseline version (e.g., 61.0% higher energy efficiency on average) that employs the GPU and achieves the efficiency comparable with that of the static best version, which requires extensive offline profiling.
Jinsu Park, Woongki Baek
ICPP2
2015 HARS: a heterogeneity-aware runtime system for self-adaptive multithreaded applications
abstract
Heterogeneous multi-processing (HMP) is rapidly emerging as a promising solution for high-performance and low-power computing. Despite extensive prior work, system-software support for self-adaptive multithreaded applications has been little explored in the context of HMP. To bridge this gap, we propose HARS, a heterogeneity-aware runtime system for self-adaptive multithreaded applications. HARS continuously monitors the application performance and dynamically adapts the system state to enhance the performance/watt of the target self-adaptive multithreaded applications on HMP systems, while satisfying the user-specified performance goal. We quantify the effectiveness of HARS by demonstrating that HARS achieves significantly higher efficiency than the baseline version with the Linux HMP scheduler and comparable efficiency with that of the static optimal version.
Jaeyoung Yun, Jinsu Park, Woongki Baek
DAC3
2010 Implementing and Evaluating a Model Checker for Transactional Memory Systems
abstract
Transactional Memory (TM) is a promising technique that addresses the difficulty of parallel programming. Since TM takes responsibility for all concurrency control, TM systems are highly vulnerable to subtle correctness errors. Due to the difficulty of fully proving the correctness of TM systems, many of them are used without any formal correctness guarantees. This paper presents ChkTM, a flexible model checking environment to verify the correctness of various TM systems. ChkTM aims to model TM systems close to the implementation level to reveal as many potential bugs as possible. For example, ChkTM accurately models the version control mechanism in timestamp-based software TMs (STMs). In addition, ChkTM can flexibly model TM systems that use additional hardware components or support nested parallelism. Using ChkTM, we model several TM systems including a widely-used industrial STM (TL2), a hybrid TM (SigTM) that uses hardware signatures, and an STM (NesTM) that supports nested parallel transactions. We then demonstrate how ChkTM can be used to find a previously unreported correctness bug in the current implementation of eager-versioning TL2. We also verify the serializability of TL2 and SigTM and strong isolation guarantees of SigTM. Finally, we quantitatively analyze ChkTM to understand the practical issues and motivate further research in model checking TM systems.
Woongki Baek, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun
ICECCS1
2010 Making nested parallel transactions practical using lightweight hardware support
abstract
Transactional Memory (TM) simplifies parallel programming by supporting parallel tasks that execute in an atomic and isolated way. To achieve the best possible performance, TM must support the nested parallelism available in real-world applications and supported by popular programming models. A few recent papers have proposed support for nested parallelism in software TM (STM) and hardware TM (HTM). However, the proposed designs are still impractical, as they either introduce excessive runtime overheads or require complex hardware structures.
Woongki Baek, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun
ICS1
2010 Green: a framework for supporting energy-conscious programming using controlled approximation
abstract
Energy-efficient computing is important in several systems ranging from embedded devices to large scale data centers. Several application domains offer the opportunity to tradeoff quality of service/solution (QoS) for improvements in performance and reduction in energy consumption. Programmers sometimes take advantage of such opportunities, albeit in an ad-hoc manner and often without providing any QoS guarantees.
Woongki Baek, Trishul M. Chilimbi
PLDI1
2010 Implementing and evaluating nested parallel transactions in software transactional memory
abstract
Transactional Memory (TM) is a promising technique that simplifies parallel programming for shared-memory applications. To date, most TM systems have been designed to efficiently support single-level parallelism. To achieve widespread use and maximize performance gains, TM must support nested parallelism available in many applications and supported by several programming models.
Woongki Baek, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun
SPAA1
2009 Fast memory snapshot for concurrent programmingwithout synchronization
abstract
The industry-wide turn toward chip-multiprocessors (CMPs) provides an increasing amount of parallel resources for commodity systems. However, it is still difficult to harness the available parallelism in user applications and system software code.
JaeWoong Chung, Woongki Baek, Christoforos E. Kozyrakis
ICS2
2008 Ased: availability, security, and debugging support usingtransactional memory
abstract
We propose ASeD that uses the hardware resources of transactional memory systems for non transactional memory purpose. We show that the hardware components for register checkpointing, data versioning, and conflict detection can be reused as basic building blocks for reliability, security, and debugging support.
JaeWoong Chung, Woongki Baek, Nathan Bronson, Jiwon Seo 0002, Christoforos E. Kozyrakis, Kunle Olukotun
SPAA2
2008 Improving software concurrency with hardware-assisted memory snapshot
abstract
We propose a hardware-assisted memory snapshot to improve software concurrency. It is built on top of the hardware resources for transactional memory and allows for easy development of system software modules such as concurrent garbage collector and dynamic profiler.
JaeWoong Chung, Jiwon Seo 0002, Woongki Baek, Chi Cao Minh, Austen McDonald, Christoforos E. Kozyrakis, Kunle Olukotun
SPAA3
2007 The OpenTM Transactional Application Programming Interface
Woongki Baek, Chi Cao Minh, Martin Trautmann, Christoforos E. Kozyrakis, Kunle Olukotun
PACT1
2007 A Scalable, Non-blocking Approach to Transactional Memory
abstract
Transactional memory (TM) provides mechanisms that promise to simplify parallel programming by eliminating the need for locks and their associated problems (deadlock, livelock, priority inversion, convoying). For TM to be adopted in the long term, not only does it need to deliver on these promises, but it needs to scale to a high number of processors. To date, proposals for scalable TM have relegated livelock issues to user-level contention managers. This paper presents the first scalable TM implementation for directory-based distributed shared memory systems that is livelock free without the need for user-level intervention. The design is a scalable implementation of optimistic concurrency control that supports parallel commits with a two-phase commit protocol, uses write-back caches, and filters coherence messages. The scalable design is based on transactional coherence and consistency (TCC), which supports continuous transactions and fault isolation. A performance evaluation of the design using both scientific and enterprise benchmarks demonstrates that the directory-based TCC design scales efficiently for NUMA systems up to 64 processors
Hassan Chafi, Jared Casper, Brian D. Carlstrom, Austen McDonald, Chi Cao Minh, Woongki Baek, Christoforos E. Kozyrakis, Kunle Olukotun
HPCA6
2007 Towards soft optimization techniques for parallel cognitive applications
Woongki Baek, JaeWoong Chung, Chi Cao Minh, Christoforos E. Kozyrakis, Kunle Olukotun
SPAA1