EDBT 2026 Demo / reviewers in the wild / expert
Ahmed Eleliemy
dblp:210/3460
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-3258-1738ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for KubernetesabstractConverged HPC-Cloud computing is an emerging computing paradigm that aims to support increasingly complex and multi-tenant scientific workflows. These systems require reconciliation of the isolation requirements of native cloud workloads and the performance demands of HPC applications. In this context, networking hardware is a critical boundary component: it is the conduit for high-throughput, low-latency communication and enables isolation across tenants. HPE Slingshot is a high-speed network interconnect that provides up to 200 Gbps of throughput per port and targets high-performance computing (HPC) systems. The Slingshot host software, including hardware drivers and network middleware libraries, is designed to meet HPC deployments, which predominantly use singletenant access modes. Hence, the Slingshot stack is not suited for secure use in multi-tenant deployments, such as converged HPCCloud deployments. In this paper, we design and implement an extension to the Slingshot stack targeting converged deployments on the basis of Kubernetes. Our integration provides secure, container-granular, and multi-tenant access to Slingshot RDMA networking capabilities at minimal overhead. Philipp Friese, Ahmed Eleliemy, Utz-Uwe Haus, Martin Schulz 0001 |
CLUSTER | 2 |
| 2023 | How Do OS and Application Schedulers Interact? An Investigation with Multithreaded ApplicationsabstractAbstract Scheduling is critical for achieving high performance for parallel applications executing on high performance computing (HPC) systems. Scheduling decisions can be taken at batch system, application, and operating system (OS) levels. In this work, we investigate the interaction between the Linux scheduler and various OpenMP scheduling options during the execution of three multithreaded codes on two types of computing nodes. When threads are unpinned, we found that OS scheduling events significantly interfere with the performance of compute-bound applications, aggravating their inherent load imbalance or overhead (by additional context switches). While the Linux scheduler balances system load in the absence of application-level load balancing, we also found it decreases performance via additional context switches and thread migrations. We observed that performing load balancing operations both at the OS and application levels is advantageous for the performance of concurrently executing applications. These results show the importance of considering the role of OS scheduling in the design of application scheduling techniques and vice versa. This work motivates further research into coordination of scheduling within multithreaded applications and the OS. Jonas H. Müller Korndörfer, Ahmed Eleliemy, Osman Seckin Simsek, Thomas Ilsche, Robert Schöne, Florina M. Ciorba |
Euro-Par | 2 |
| 2023 | Automated Scheduling Algorithm Selection in OpenMPabstractScientific and data analysis applications are increasingly complex, with evolving computational and memory requirements during execution. Conversely, modern high performance computing (HPC) systems are heterogeneous and offer significant parallelism at the node and core levels. Scheduling and load balancing techniques are essential for maximizing the performance of applications on HPC systems. Recent work has shown the importance and the need of bringing scheduling techniques from the literature into commonly used parallelization frameworks, such as OpenMP. While this results in a multitude of scheduling options, it renders challenging the offline or online selection of the most suitable scheduling technique for an application-system pair, in the context of evolving applications’ computational requirements and variable capacities of modern HPC systems. Therefore, approaches for automatic selection of scheduling algorithms are urgently needed. This is an instance of the algorithm selection problem, proposed by Rice [1]. This oral communication presents recent results and ongoing work on the topic of automated scheduling algorithm selection for improving the performance of OpenMP applications. We present and evaluate two selection approaches, expert-based and reinforcement learning-based, both implemented in the LB4OMP scheduling library [2] – an extension of the LLVM OpenMP runtime library. The results show that automatic scheduling algorithm selection in LB4OMP surpasses manual selection, and that depending on the case, reinforcement learning-based selection outperforms expert-based selection or viceversa. This work is part of our ongoing efforts in solving the multilevel scheduling problem [3]. Similar scheduling algorithm selection solutions are needed at other parallelism levels, i.e., process level (MPI) and batch level (SLURM), to adapt to unpredictable variations both in applications and resources (including failures) that may arise during execution. Florina M. Ciorba, Ali Mohammed, Jonas H. Müller Korndörfer, Ahmed Eleliemy |
ISPDC | 4 |
| 2023 | DaphneSched: A Scheduler for Integrated Data Analysis PipelinesabstractDAPHNE is a new open-source software infrastructure designed to address the increasing demands of integrated data analysis (IDA) pipelines, comprising data management (DM), high performance computing (HPC), and machine learning (ML) systems. Efficiently executing IDA pipelines is challenging due to their diverse computing characteristics and demands. Therefore, IDA pipelines executed with the DAPHNE infrastructure require an efficient and versatile scheduler to support these demands. This work introduces DaphneSched, the task-based scheduler at the core of DAPHNE [1]. DaphneSched is versatile by incorporating eleven task partitioning and three task assignment techniques, bringing the state-of-the-art closer to the state-of-the-practice task scheduling. To showcase DaphneSched’s effectiveness in scheduling IDA pipelines, we evaluate its performance on two applications: a product recommendation system and training of a linear regression model. We conduct performance experiments on multicore platforms with 20 and 56 cores, respectively. The results show that the versatility of DaphneSched enabled combinations of scheduling strategies that outperform commonly used scheduling techniques by up to 13%. This work confirms the benefits of employing DaphneSched for the efficient execution of applications with IDA pipelines. Ahmed Eleliemy, Florina M. Ciorba |
ISPDC | 1 |
| 2022 | DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines
Patrick Damme, Marius Birkenbach, Constantinos Bitsakos, Matthias Boehm 0001, Philippe Bonnet, Florina M. Ciorba, Mark Dokter, Pawel Dowgiallo, Ahmed Eleliemy, Christian Färber, Georgios I. Goumas, Dirk Habich, Niclas Hedam, Marlies Hofer, Kevin Innerebner, Vasileios Karakostas, Roman Kern, Tomaz Kosar, Alexander Krause 0001, Daniel Krems, Andreas Laber, Wolfgang Lehner, Eric Mier, Marcus Paradies, Bernhard Peischl, Gabrielle Poerwawinata, Stratos Psomadakis, Tilmann Rabl, Piotr Ratuszniak, Pedro Silva 0011, Nikolai Skuppin, Andreas Starzacher, Benjamin Steinwender, Ilin Tolovski, Pinar Tözün, Wojciech Ulatowski, Yuanyuan Wang 0002, Izajasz P. Wrosz, Ales Zamuda, Ce Zhang 0001, Xiao Xiang Zhu 0001 |
CIDR | 9 |
| 2022 | LB4OMP: A Dynamic Load Balancing Library for Multithreaded ApplicationsabstractExascale computing systems will exhibit high degrees of hierarchical parallelism, with thousands of computing nodes and hundreds of cores per node. Efficiently exploiting hierarchical parallelism is challenging due to load imbalance that arises at multiple levels. OpenMP is the most widely-used standard for expressing and exploiting the ever-increasing node-level parallelism. The scheduling options in OpenMP are insufficient to address the load imbalance that arises during the execution of multithreaded applications. The limited scheduling options in OpenMP hinder research on novel scheduling techniques which require comparison with others from the literature. This work introduces LB4OMP, an open-source dynamic load balancing library that implements successful scheduling algorithms from the literature. LB4OMP is a research infrastructure designed to spur and support present and future scheduling research, for the benefit of multithreaded applications performance. Through an extensive performance analysis campaign, we assess the effectiveness and demystify the performance of all loop scheduling techniques in the library. We show that, for numerous applications-systems pairs, the scheduling techniques in LB4OMP outperform the scheduling options in OpenMP. Node-level load balancing using LB4OMP leads to reduced cross-node load imbalance and to improved MPI+OpenMP applications performance, which is critical for Exascale computing. Jonas H. Müller Korndörfer, Ahmed Eleliemy, Ali Mohammed, Florina M. Ciorba |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Automated Scheduling Algorithm Selection and Chunk Parameter Calculation in OpenMPabstractIncreasing node and cores-per-node counts in supercomputers render scheduling and load balancing critical for exploiting parallelism. OpenMP applications can achieve high performance via careful selection of schedulingkindandchunkparameters on a per-loop, per-application, and per-system basis from a portfolio of advanced scheduling algorithms (Korndörferet al., 2022). This selection approach is time-consuming, challenging, and may need to change during execution. We proposeAuto4OMP, a novel approach for automated load balancing of OpenMP applications. With Auto4OMP, we introduce three schedulingalgorithm selection methodsand anexpert-defined chunk parameterfor OpenMP'sscheduleclause'skindandchunk, respectively. Auto4OMP extends the OpenMPschedule(auto)andchunkparameter implementation in LLVM's OpenMP runtime library to automatically select a scheduling algorithm and calculate a chunk parameter during execution. Loop characteristics are inferred in Auto4OMP from the loop execution over the application's time-steps. The experiments performed in this work show that Auto4OMP improves applications performance by up to$11\%$compared to LLVM'sschedule(auto)implementation and outperforms manual selection. Auto4OMP improves MPI+OpenMP applications performance byexplicitlyminimizing thread- andimplicitlyreducing process-load imbalance. Ali Mohammed, Jonas H. Müller Korndörfer, Ahmed Eleliemy, Florina M. Ciorba |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | An approach for realistically simulating the performance of scientific applications on high performance computing systemsabstractScientific applications often contain large, computationally-intensive, and irregular parallel loops or tasks that exhibit stochastic behavior leading to load imbalance. Load imbalance often manifests during the execution of parallel scientific applications on large and complex high performance computing (HPC) systems. The extreme scale of HPC systems on the road to Exascale computing only exacerbates the loss in performance due to load imbalance. Dynamic loop self-scheduling (DLS) techniques are instrumental in improving the performance of scientific applications on HPC systems via load balancing. Selecting a DLS technique that results in the best performance for different problem and system sizes requires a large number of exploratory experiments. Currently, a theoretical model that can be used to predict the scheduling technique that yields the best performance for a given problem and system has not yet been identified. Therefore, simulation is the most appropriate approach for conducting such exploratory experiments in a reasonable amount of time. However, conducting realistic and trustworthy simulations of application performance under different configurations is challenging. This work devises an approach to realistically simulate computationally-intensive scientific applications that employ DLS and execute on HPC systems. The proposed approach minimizes the sources of uncertainty in the simulative experiments results by bridging the native and simulative experimental approaches. A new method is proposed to capture the variation of application performance between different native executions. Several approaches to represent the application tasks (or loop iterations) are compared to establish their influence on the simulative application performance. A novel simulation strategy is introduced that applies the proposed approach, which transforms a native application code into simulative code. The native and simulative performance of two computationally-intensive scientific applications that employ eight task scheduling techniques (static, nonadaptive dynamic, and adaptive dynamic) are compared to evaluate the realism of the proposed simulation approach. The comparison of the performance characteristics extracted from the native and simulative performance shows that the proposed simulation approach fully captured most of the performance characteristics of interest. This work shows and establishes the importance of simulations that realistically predict the performance of DLS techniques for different applications and system configurations. Ali Mohammed, Ahmed Eleliemy, Florina M. Ciorba, Franziska Kasielke, Ioana Banicescu |
Future Gener. Comput. Syst. | 2 |
| 2019 | Dynamic Loop Scheduling Using MPI Passive-Target Remote Memory AccessabstractScientific applications often contain large computationally-intensive parallel loops. Loop scheduling techniques aim to achieve load balanced executions of such applications. For distributed-memory systems, existing dynamic loop scheduling (DLS) libraries are typically MPI-based, and employ a master-worker execution model to assign variably-sized chunks of loop iterations. The master-worker execution model may adversely impact performance due to the master-level contention. This work proposes a distributed chunk-calculation approach that does not require the master-worker execution scheme. Moreover, it considers the novel features in the latest MPI standards, such as passive-target remote memory access, shared-memory window creation, and atomic read-modify-write operations. To evaluate the proposed approach, five well-known DLS techniques, two applications, and two heterogeneous hardware setups have been considered. The DLS techniques implemented using the proposed approach outperformed their counterparts implemented using the traditional master-worker execution model. Ahmed Eleliemy, Florina M. Ciorba |
PDP | 1 |
| 2018 | Experimental Verification and Analysis of Dynamic Loop Scheduling in Scientific ApplicationsabstractScientific applications are often irregular and characterized by large computationally-intensive parallel loops. Dynamic loop scheduling (DLS) techniques improve the performance of computationally-intensive scientific applications via load balancing of their execution on high-performance computing (HPC) systems. Identifying the most suitable choices of data distribution strategies, system sizes, and DLS techniques which improve the performance of a given application, requires intensive assessment and a large number of exploratory native experiments (using real applications on real systems), which may not always be feasible or practical due to associated time and costs. In such cases, simulative experiments are more appropriate for studying the performance of applications. This motivates the question of 'How realistic are the simulations of executions of scientific applications using DLS on HPC platforms?' In the present work, a methodology is devised to answer this question. It involves the experimental verification and analysis of the performance of DLS in scientific applications. The proposed methodology is employed for a computer vision application executing using four DLS techniques on two different HPC platforms, both via native and simulative experiments. The evaluation and analysis of the native and simulative results indicate that the accuracy of the simulative experiments is strongly influenced by the approach used to extract the computational effort of the application (FLOP-or time-based), the choice of application model representation into simulation (data or task parallel) and the available HPC subsystem models in the simulator (multi-core CPUs, memory hierarchy and network topology). The minimum and the maximum percent errors achieved between the native and the simulative experiments are 0.95% and 8.03%, respectively. Ali Mohammed, Ahmed Eleliemy, Florina M. Ciorba, Franziska Kasielke, Ioana Banicescu |
ISPDC | 2 |
| 2017 | Exploring the Relation between Two Levels of Scheduling Using a Novel Simulation ApproachabstractModern high performance computing (HPC) systems exhibit a rapid growth in size, both “horizontally” in the number of nodes, as well as “vertically” in the number of cores per node. As such, they offer additional levels of hardware parallelism. Each level requires and employs algorithms for appropriately scheduling the computational work at the respective level. The present work explores the relation between two scheduling levels: batch and application. To understand and explore this relation, a novel simulation approach is presented that bridges two existing simulators from the two scheduling levels. A novel two-level simulator that implements the proposed approach is introduced. The two-level simulator is used to simulate all combinations of three batch scheduling and four application scheduling algorithms from the literature. These combinations are considered for allocating resources and executing the parallel jobs from a workload of a production HPC system. The results of the scheduling experiments reveal the strong relation between decisions taken at the two scheduling levels and their mutual influence. Complementing the simulations, the two-level simulator produces abstract parallel execution traces, which can visually be examined and illustrate the execution of different jobs and, for each job, the execution of its tasks at node and core levels, respectively. Ahmed Eleliemy, Ali Mohammed, Florina M. Ciorba |
ISPDC | 1 |