Jonas H. Müller Korndörfer

dblp:247/6464 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0003-3014-3275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2023 How Do OS and Application Schedulers Interact? An Investigation with Multithreaded Applications
abstract
Abstract Scheduling is critical for achieving high performance for parallel applications executing on high performance computing (HPC) systems. Scheduling decisions can be taken at batch system, application, and operating system (OS) levels. In this work, we investigate the interaction between the Linux scheduler and various OpenMP scheduling options during the execution of three multithreaded codes on two types of computing nodes. When threads are unpinned, we found that OS scheduling events significantly interfere with the performance of compute-bound applications, aggravating their inherent load imbalance or overhead (by additional context switches). While the Linux scheduler balances system load in the absence of application-level load balancing, we also found it decreases performance via additional context switches and thread migrations. We observed that performing load balancing operations both at the OS and application levels is advantageous for the performance of concurrently executing applications. These results show the importance of considering the role of OS scheduling in the design of application scheduling techniques and vice versa. This work motivates further research into coordination of scheduling within multithreaded applications and the OS.
Jonas H. Müller Korndörfer, Ahmed Eleliemy, Osman Seckin Simsek, Thomas Ilsche, Robert Schöne, Florina M. Ciorba
Euro-Par1
2023 Automated Scheduling Algorithm Selection in OpenMP
abstract
Scientific and data analysis applications are increasingly complex, with evolving computational and memory requirements during execution. Conversely, modern high performance computing (HPC) systems are heterogeneous and offer significant parallelism at the node and core levels. Scheduling and load balancing techniques are essential for maximizing the performance of applications on HPC systems. Recent work has shown the importance and the need of bringing scheduling techniques from the literature into commonly used parallelization frameworks, such as OpenMP. While this results in a multitude of scheduling options, it renders challenging the offline or online selection of the most suitable scheduling technique for an application-system pair, in the context of evolving applications’ computational requirements and variable capacities of modern HPC systems. Therefore, approaches for automatic selection of scheduling algorithms are urgently needed. This is an instance of the algorithm selection problem, proposed by Rice [1]. This oral communication presents recent results and ongoing work on the topic of automated scheduling algorithm selection for improving the performance of OpenMP applications. We present and evaluate two selection approaches, expert-based and reinforcement learning-based, both implemented in the LB4OMP scheduling library [2] – an extension of the LLVM OpenMP runtime library. The results show that automatic scheduling algorithm selection in LB4OMP surpasses manual selection, and that depending on the case, reinforcement learning-based selection outperforms expert-based selection or viceversa. This work is part of our ongoing efforts in solving the multilevel scheduling problem [3]. Similar scheduling algorithm selection solutions are needed at other parallelism levels, i.e., process level (MPI) and batch level (SLURM), to adapt to unpredictable variations both in applications and resources (including failures) that may arise during execution.
Florina M. Ciorba, Ali Mohammed, Jonas H. Müller Korndörfer, Ahmed Eleliemy
ISPDC3
2022 LB4OMP: A Dynamic Load Balancing Library for Multithreaded Applications
abstract
Exascale computing systems will exhibit high degrees of hierarchical parallelism, with thousands of computing nodes and hundreds of cores per node. Efficiently exploiting hierarchical parallelism is challenging due to load imbalance that arises at multiple levels. OpenMP is the most widely-used standard for expressing and exploiting the ever-increasing node-level parallelism. The scheduling options in OpenMP are insufficient to address the load imbalance that arises during the execution of multithreaded applications. The limited scheduling options in OpenMP hinder research on novel scheduling techniques which require comparison with others from the literature. This work introduces LB4OMP, an open-source dynamic load balancing library that implements successful scheduling algorithms from the literature. LB4OMP is a research infrastructure designed to spur and support present and future scheduling research, for the benefit of multithreaded applications performance. Through an extensive performance analysis campaign, we assess the effectiveness and demystify the performance of all loop scheduling techniques in the library. We show that, for numerous applications-systems pairs, the scheduling techniques in LB4OMP outperform the scheduling options in OpenMP. Node-level load balancing using LB4OMP leads to reduced cross-node load imbalance and to improved MPI+OpenMP applications performance, which is critical for Exascale computing.
Jonas H. Müller Korndörfer, Ahmed Eleliemy, Ali Mohammed, Florina M. Ciorba
IEEE Trans. Parallel Distributed Syst.1
2022 Automated Scheduling Algorithm Selection and Chunk Parameter Calculation in OpenMP
abstract
Increasing node and cores-per-node counts in supercomputers render scheduling and load balancing critical for exploiting parallelism. OpenMP applications can achieve high performance via careful selection of schedulingkindandchunkparameters on a per-loop, per-application, and per-system basis from a portfolio of advanced scheduling algorithms (Korndörferet al., 2022). This selection approach is time-consuming, challenging, and may need to change during execution. We proposeAuto4OMP, a novel approach for automated load balancing of OpenMP applications. With Auto4OMP, we introduce three schedulingalgorithm selection methodsand anexpert-defined chunk parameterfor OpenMP'sscheduleclause'skindandchunk, respectively. Auto4OMP extends the OpenMPschedule(auto)andchunkparameter implementation in LLVM's OpenMP runtime library to automatically select a scheduling algorithm and calculate a chunk parameter during execution. Loop characteristics are inferred in Auto4OMP from the loop execution over the application's time-steps. The experiments performed in this work show that Auto4OMP improves applications performance by up to$11\%$compared to LLVM'sschedule(auto)implementation and outperforms manual selection. Auto4OMP improves MPI+OpenMP applications performance byexplicitlyminimizing thread- andimplicitlyreducing process-load imbalance.
Ali Mohammed, Jonas H. Müller Korndörfer, Ahmed Eleliemy, Florina M. Ciorba
IEEE Trans. Parallel Distributed Syst.2