Jinsu Park

dblp:40/1015 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
2since 2021 · last 2021
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 9 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 46% Parallel and multicore computing · 21% Cloud and datacenter computing · 15%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling
0.512021
Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2021
Parallel and multicore computing › parallel programming runtimes
task-based runtime
0.512021
Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2021
Memory systems › cache management
adaptive cache management
0.412019
Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › cache management › cache insertion policy
cache bypassing
0.412019
Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › cache design
cache indexing
0.412019
Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems
cache management
0.412019
Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › cache management
cache partitioning
0.412019
CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers · EuroSys 2019
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.412019
CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers · EuroSys 2019
GPUs and heterogeneous computing
GPU computing
0.412019
Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › cache management › cache partitioning
last-level cache partitioning
0.412019
CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers · EuroSys 2019
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation
0.412019
CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers · EuroSys 2019
Processor architecture and microarchitecture › multiprocessor architecture
heterogeneous multiprocessing
0.422021
HARS: a heterogeneity-aware runtime system for self-adaptive multithreaded applications · DAC 2015
Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2021
Parallel and multicore computing
parallel programming runtimes
0.212015
HARS: a heterogeneity-aware runtime system for self-adaptive multithreaded applications · DAC 2015
Energy-efficient computing
power management
0.112021
Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2021
Cloud and datacenter computing › cluster resource management and scheduling
fairness-aware scheduling
0.112019
CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers · EuroSys 2019

Methods — techniques the papers use, named apart from their topics

performance and power estimation · 0.5criticality-aware scheduling · 0.5quantitative evaluation · 0.4performance monitoring · 0.2dynamic adaptation · 0.2
YearPublicationVenuePosition
2021 PALM: Progress- and Locality-Aware Adaptive Task Migration for Efficient Thread Packing
abstract
Thread packing (TP) is an effective and widely-used technique to significantly improve the efficiency of parallel systems by dynamically controlling the number of cores allocated to multithreaded applications based on their requirements such as performance and energy efficiency. Despite the extensive prior works on TP, little work has been done to investigate and address its performance inefficiencies that arise across various parallel systems and applications with different characteristics. To bridge this gap, we investigate the performance inefficiencies of TP using a wide range of parallel applications and system configurations and identify their root causes. Guided by the in-depth performance characterization results, we propose PALM, progress- and locality-aware adaptive task migration for efficient TP. Through quantitative evaluation, we demonstrate that PALM achieves significantly higher performance and lower energy consumption than TP across various synchronization-intensive applications and system configurations, provides the performance and energy consumption comparable with the thread reduction technique, and considerably improves the efficiency of dynamic server consolidation and the performance under power capping.
Jinsu Park, Seongbeom Park, Myeonggyun Han, Woongki Baek
IPDPS1
2021 Design and Implementation of a Criticality- and Heterogeneity-Aware Runtime System for Task-Parallel Applications
abstract
Heterogeneous multiprocessing (HMP) is an emerging technology for high-performance and energy-efficient computing. While task parallelism is widely used in various computing domains, such as embedded, big-data, and machine-learning computing domains, it still remains unexplored to investigate the efficient runtime support that effectively utilizes the criticality of the tasks of the target application and the heterogeneity of the underlying HMP system with full resource management. To bridge this gap, we propose CHRT, a criticality- and heterogeneity-aware runtime system for task-parallel applications. CHRT dynamically estimates the performance and power consumption of the target task-parallel application and robustly manages the full HMP system resources (i.e., core types, counts, and voltage/frequency levels) to maximize the overall efficiency. Our quantitative evaluation based on widely-used task parallel benchmarks and two full HMP systems (i.e., the XU3 and HiKey970 HMP systems) demonstrates the effectiveness of CHRT in that CHRT achieves significantly higher energy (e.g., 60.4 and 57.2 percent on average on the XU3 system) and energy-delay product (e.g., 52.2 and 44.0 percent on average on the HiKey970 system) efficiency than the baseline runtime system that employs the breadth-first scheduler and the state-of-the-art criticality-aware runtime system and incurs low performance overheads.
Myeonggyun Han, Jinsu Park, Woongki Baek
IEEE Trans. Parallel Distributed Syst.2
2020 Virtual Subcarrier Aided Channel Estimation Schemes for Tracking Rapid Time Variant Channels in IEEE 802.11p Systems
abstract
This paper investigates the channel estimation schemes for tracking rapid time variant channels in IEEE 802. 11p systems. To overcome the problems of the existing preamble based channel estimation, various data-pilot aided (DPA) algorithms have been developed. However, their performance is not sufficient to meet the advanced V2X use cases. In this paper, we propose a novel two step channel estimation scheme which outperforms the conventional DPA designs. For the first step, we develop an enhanced time-domain reliability test and frequency domain interpolation scheme by exploiting the Euclidean distance between the received signals and constellation points, and the channel information in the virtual subcarriers to improve the prediction accuracy at high signal-to-noise ratio (SNR). In the second step, we attenuate the noise components from the estimated channel frequency responses in the previous step to address the low SNR performance gain. To this end, we adopt the time domain least square estimation strategy. The proposed scheme exhibits dramatic performance gain over the SNR with marginal complexity increase compared to the conventional DPA algorithms. Finally, simulation results verify the efficiency of the proposed scheme.
Seungho Han, Jinsu Park, Changick Song
VTC Spring2
2019 MOSAIC: Heterogeneity-, Communication-, and Constraint-Aware Model Slicing and Execution for Accurate and Efficient Inference
abstract
Heterogeneous embedded systems have surfaced as a promising solution for accurate and efficient deep-learning inference on mobile devices. Despite extensive prior works, it still remains unexplored to investigate the system-software support that efficiently executes inference workloads by judiciously considering their performance and energy heterogeneity, communication overheads, and constraints. To bridge this gap, we propose MOSAIC, heterogeneity-, communication-, and constraint-aware model slicing and execution for accurate and efficient inference on heterogeneous embedded systems. MOSAIC generates the efficient model slicing and execution plan for the target inference workload through dynamic programming. MOSAIC significantly reduces inference latency and energy, exhibits high estimation accuracy, and incurs small overheads.
Myeonggyun Han, Jihoon Hyun, Seongbeom Park, Jinsu Park, Woongki Baek
PACT4
2019 POSTER: The Performance Impact of Thread Packing on Synchronization-Intensive Applications
abstract
Thread packing (TP) is a widely-used technique to improve the efficiency of parallel systems. Despite extensive prior works, relatively little work has been done to investigate its performance inefficiencies. To bridge this gap, we quantify its performance impact on synchronization-intensive applications and identify the root causes of its performance inefficiencies.
Jinsu Park, Seongbeom Park, Myeonggyun Han, Woongki Baek
PACT1
2019 CoPart: Coordinated Partitioning of Last-Level Cache and Memory Bandwidth for Fairness-Aware Workload Consolidation on Commodity Servers
abstract
Workload consolidation is a widely-used technique to maximize server resource utilization in cloud and datacenter computing. Recent commodity CPUs support last-level cache (LLC) and memory bandwidth partitioning functionalities that can be used to ensure the fairness of the consolidated workloads. While prior work has proposed a variety of resource partitioning techniques, it still remains unexplored to characterize the impact of LLC and memory bandwidth partitioning on the fairness of the consolidated workloads and investigate system software support to dynamically control LLC and memory bandwidth partitioning in a coordinated manner.
Jinsu Park, Seongbeom Park, Woongki Baek
EuroSys1
2019 Analyzing and optimizing the performance and energy efficiency of transactional scientific applications on large-scale NUMA systems with HTM support
Jinsu Park, Woongki Baek
J. Parallel Distributed Comput.1
2019 Improving the Performance and Energy Efficiency of GPGPU Computing through Integrated Adaptive Cache Management
abstract
Hardware caches are widely employed in GPGPUs to achieve higher performance and energy efficiency. Incorporating hardware caches in GPGPUs, however, does not immediately guarantee enhanced performance and energy efficiency due to high cache contention and thrashing. To address the inefficiency of GPGPU caches, various adaptive techniques (e.g., warp limiting) have been proposed. However, relatively little work has been done in the context of creating an architectural framework that tightly integrates adaptive cache management techniques and investigating their effectiveness and interaction. To bridge this gap, we propose IACM, integrated adaptive cache management for high-performance and energy-efficient GPGPU computing. IACM integrates the state-of-the-art adaptive cache management techniques (i.e., cache indexing, bypassing, and warp limiting) in a unified architectural framework. Our quantitative evaluation demonstrates that IACM significantly improves the performance and energy efficiency of various GPGPU workloads over the baseline architecture (i.e., 98.1 and 61.9 percent on average, respectively), achieves considerably higher performance than the state-of-the-art technique (i.e., 361.4 percent at maximum and 7.7 percent on average), and delivers significant performance and energy-efficiency gains over the baseline GPGPU architecture enhanced with advanced architectural technologies.
Kyu Yeun Kim, Jinsu Park, Woongki Baek
IEEE Trans. Parallel Distributed Syst.2
2018 Hypart: a hybrid technique for practical memory bandwidth partitioning on commodity servers
abstract
Memory bandwidth is a highly performance-critical shared resource on modern computer systems. To prevent the contention on memory bandwidth among the collocated workloads, prior works have investigated memory bandwidth partitioning techniques. Despite the extensive prior works, it still remains unexplored to characterize the widely-used memory bandwidth partitioning techniques based on various metrics and investigate a hybrid technique that employs multiple memory bandwidth partitioning techniques to improve the overall efficiency.
Jinsu Park, Seongbeom Park, Myeonggyun Han, Jihoon Hyun, Woongki Baek
PACT1
2018 RPPC: A Holistic Runtime System for Maximizing Performance Under Power Capping
abstract
Maximizing performance in power-constrained computing environments is highly important in cloud and datacenter computing. To achieve the best possible performance of parallel applications under power capping, it is crucial to execute them with the optimal concurrency level and cross-component power allocation between CPUs and memory. Despite extensive prior works, it still remains unexplored to investigate the efficient runtime support that maximizes the performance of parallel applications under power capping through the coordinated control of concurrency level and cross-component power allocation. To bridge this gap, this work proposes RPPC, a holistic runtime system for maximizing performance under power capping. In contrast to the state-of-the-art techniques, RPPC robustly controls the two performance-critical knobs (i.e., concurrency level and cross-component power allocation) in a coordinated manner to maximize the performance of parallel applications under power capping. RPPC dynamically identifies the characteristics of the target parallel application and explores the system state space to find an efficient system state. Our experimental results demonstrate that RPPC significantly outperforms the two state-of-the-art power-capping techniques, achieves the performance comparable with the static best version that requires extensive per-application offline profiling, incurs small performance overheads, and provides the re-adaptation mechanism to external events such as total power budget changes.
Jinsu Park, Seongbeom Park, Woongki Baek
CCGrid1
2018 CEML: a Coordinated Runtime System for Efficient Machine Learning on Heterogeneous Computing Systems
Jihoon Hyun, Jinsu Park, Kyu Yeun Kim, Seongdae Yu, Woongki Baek
Euro-Par2
2018 Quantifying the Performance and Energy-Efficiency Impact of Hardware Transactional Memory on Scientific Applications on Large-Scale NUMA Systems
abstract
Hardware transactional memory (HTM) is supported by widely-used commodity processors. While the effectiveness of HTM has been evaluated based on small-scale multi-core systems, it still remains unexplored to quantify the performance and energy-efficiency of HTM for scientific workloads on large-scale NUMA systems, which have been increasingly adopted to high-performance computing. To bridge this gap, this work investigates the performance and energy-efficiency impact of HTM on scientific applications on large-scale NUMA systems. We first quantify the performance and energy efficiency of HTM for scientific workloads based on the widely-used CLOMP-TM benchmark. We then discuss a set of generic software optimizations that can be effectively used to improve the performance and energy efficiency of transactional scientific workloads on large-scale NUMA systems. Finally, we present case studies in which we apply a set of the optimizations to representative transactional scientific applications and significantly optimize their performance and energy efficiency on large-scale NUMA systems.
Jinsu Park, Woongki Baek
IPDPS1
2017 CHRT: A criticality- and heterogeneity-aware runtime system for task-parallel applications
abstract
Heterogeneous multiprocessing (HMP) is an emerging technology for high-performance and energy-efficient computing. While task parallelism is widely used in various computing domains from the embedded to machine-learning computing domains, relatively little work has been done to investigate the efficient runtime support that effectively utilizes the criticality of the tasks of the target application and the heterogeneity of the underlying HMP system with full resource management. To bridge this gap, we propose a criticality- and heterogeneity-aware runtime system for task-parallel applications (CHRT). CHRT dynamically estimates the performance and power consumption of the target task-parallel application and robustly manages the full HMP system resources (i.e., core types, counts, and voltage/frequency levels) to maximize the overall efficiency. Our experimental results show that CHRT achieves significantly higher energy efficiency than the baseline runtime system that employs the breadth-first scheduler and the state-of-the-art criticality-aware runtime system.
Myeonggyun Han, Jinsu Park, Woongki Baek
DATE2
2016 RMC: an integrated runtime system for adaptive many-core computing
abstract
Many-core computing has surfaced as a promising solution to satisfy the rapidly increasing computational needs for various areas ranging from embedded to datacenter computing. However, when allocated with an excessive number of cores, multithreaded applications may fail to achieve optimal performance and energy efficiency due to the contention on software and/or hardware resources. While previous research has proposed adaptive techniques such as thread packing (TP) and dynamic threading (DT), they often lead to suboptimal results because they are used in an isolated manner. To address this problem, we propose RMC, an integrated runtime system for adaptive many-core computing. Guided by the runtime information of parallel applications, RMC dynamically adapts their execution by combining the TP and DT techniques. We apply RMC to six PARSEC benchmarks that use representative parallelism models (i.e., fork-join, task, and pipeline). We demonstrate that RMC is easy to use, considerably outperforms the state-of-the-art techniques for three PARSEC benchmarks, and incurs a small overhead to the rest of the benchmarks.
Jinsu Park, Eunbi Cho, Woongki Baek
EMSOFT1
2016 HAP: A Heterogeneity-Conscious Runtime System for Adaptive Pipeline Parallelism
Jinsu Park, Woongki Baek
Euro-Par1
2016 IACM: Integrated adaptive cache management for high-performance and energy-efficient GPGPU computing
abstract
Hardware caches are widely employed in GPGPUs to achieve higher performance and energy efficiency. Incorporating hardware caches in GPGPUs, however, does not immediately guarantee enhanced performance and energy efficiency due to high cache contention and thrashing. To address the inefficiency of GPGPU caches, various adaptive techniques (e.g., warp limiting) have been proposed. However, relatively little work has been done in the context of creating an architectural framework that tightly integrates adaptive cache management techniques and investigating their effectiveness and interaction. To bridge this gap, we propose IACM, integrated adaptive cache management for high-performance and energy-efficient GPGPU computing. IACM integrates the state-of-the-art adaptive cache management techniques (i.e., cache indexing, bypassing, and warp limiting) in a unified architectural framework. Our quantitative evaluation demonstrates that IACM significantly improves the performance and energy efficiency of various GPGPU workloads over the baseline architecture (i.e., 98.1% and 61.9% on average).
Kyu Yeun Kim, Jinsu Park, Woongki Baek
ICCD2
2016 RCHC: A Holistic Runtime System for Concurrent Heterogeneous Computing
abstract
Concurrent heterogeneous computing (CHC) is rapidly emerging as a promising solution for high-performance and energy-efficient computing. The fundamental challenges for efficient CHC are how to partition the workload of the target application across the devices in the underlying CHC system and how to control the operating frequency of each device in order to maximize the overall efficiency. Despite the extensive prior work on the system software techniques for CHC, efficient runtime support for CHC that robustly supports both functional and performance heterogeneity without the need for extensive offline profiling still remains unexplored. To bridge this gap, we propose RCHC, a holistic runtime system for concurrent heterogeneous computing. RCHC dynamically profiles the target application and constructs the performance and power estimation models based on the runtime information. Guided by the estimation models, RCHC explores the system state space, determines the best system state that is expected to maximize the efficiency of the target application, and accordingly executes it. Our experimental results demonstrate that RCHC significantly outperforms the baseline version (e.g., 61.0% higher energy efficiency on average) that employs the GPU and achieves the efficiency comparable with that of the static best version, which requires extensive offline profiling.
Jinsu Park, Woongki Baek
ICPP1
2015 HARS: a heterogeneity-aware runtime system for self-adaptive multithreaded applications
abstract
Heterogeneous multi-processing (HMP) is rapidly emerging as a promising solution for high-performance and low-power computing. Despite extensive prior work, system-software support for self-adaptive multithreaded applications has been little explored in the context of HMP. To bridge this gap, we propose HARS, a heterogeneity-aware runtime system for self-adaptive multithreaded applications. HARS continuously monitors the application performance and dynamically adapts the system state to enhance the performance/watt of the target self-adaptive multithreaded applications on HMP systems, while satisfying the user-specified performance goal. We quantify the effectiveness of HARS by demonstrating that HARS achieves significantly higher efficiency than the baseline version with the Linux HMP scheduler and comparable efficiency with that of the static optimal version.
Jaeyoung Yun, Jinsu Park, Woongki Baek
DAC2