VLDB 2026 Research / reviewers in the wild / expert
Florina M. Ciorba
dblp:15/2321
· DBLP profile ↗
45ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-2773-4499ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Autonomy Loop for Dynamic HPC Job Time Limit Adjustment
Thomas Jakobsche, Osman Seckin Simsek, Jim M. Brandt, Ann C. Gentile, Florina M. Ciorba |
Euro-Par (1) | 5 |
| 2025 | SIREN: Software Identification and Recognition in HPC SystemsabstractHPC systems use monitoring and operational data analytics to ensure efficiency, performance, and orderly operations. Application-specific insights are crucial for analyzing the increasing complexity and diversity of HPC workloads, particularly through the identification of unknown software and recognition of repeated executions, which facilitate system optimization and security improvements. However, traditional identification methods using job or file names are unreliable for arbitrary user-provided names. Fuzzy hashing the content of executables detects similarities despite different code versions or compilation approaches while preserving privacy and file integrity, overcoming these limitations. We introduce SIREN, a process-level data collection framework for software identification and recognition. SIREN improves observability in HPC job execution by enabling analysis of process metadata, environment information, and executable fuzzy hashes. Findings from an opt-in deployment campaign on LUMI show SIREN’s ability to provide insights into software usage, recognition of repeated executions of known applications, and similarity-based identification of unknown applications. Thomas Jakobsche, Fredrik Robertsén, Jessica R. Jones, Utz-Uwe Haus, Florina M. Ciorba |
SC | 5 |
| 2023 | How Do OS and Application Schedulers Interact? An Investigation with Multithreaded ApplicationsabstractAbstract Scheduling is critical for achieving high performance for parallel applications executing on high performance computing (HPC) systems. Scheduling decisions can be taken at batch system, application, and operating system (OS) levels. In this work, we investigate the interaction between the Linux scheduler and various OpenMP scheduling options during the execution of three multithreaded codes on two types of computing nodes. When threads are unpinned, we found that OS scheduling events significantly interfere with the performance of compute-bound applications, aggravating their inherent load imbalance or overhead (by additional context switches). While the Linux scheduler balances system load in the absence of application-level load balancing, we also found it decreases performance via additional context switches and thread migrations. We observed that performing load balancing operations both at the OS and application levels is advantageous for the performance of concurrently executing applications. These results show the importance of considering the role of OS scheduling in the design of application scheduling techniques and vice versa. This work motivates further research into coordination of scheduling within multithreaded applications and the OS. Jonas H. Müller Korndörfer, Ahmed Eleliemy, Osman Seckin Simsek, Thomas Ilsche, Robert Schöne, Florina M. Ciorba |
Euro-Par | 6 |
| 2023 | Automated Scheduling Algorithm Selection in OpenMPabstractScientific and data analysis applications are increasingly complex, with evolving computational and memory requirements during execution. Conversely, modern high performance computing (HPC) systems are heterogeneous and offer significant parallelism at the node and core levels. Scheduling and load balancing techniques are essential for maximizing the performance of applications on HPC systems. Recent work has shown the importance and the need of bringing scheduling techniques from the literature into commonly used parallelization frameworks, such as OpenMP. While this results in a multitude of scheduling options, it renders challenging the offline or online selection of the most suitable scheduling technique for an application-system pair, in the context of evolving applications’ computational requirements and variable capacities of modern HPC systems. Therefore, approaches for automatic selection of scheduling algorithms are urgently needed. This is an instance of the algorithm selection problem, proposed by Rice [1]. This oral communication presents recent results and ongoing work on the topic of automated scheduling algorithm selection for improving the performance of OpenMP applications. We present and evaluate two selection approaches, expert-based and reinforcement learning-based, both implemented in the LB4OMP scheduling library [2] – an extension of the LLVM OpenMP runtime library. The results show that automatic scheduling algorithm selection in LB4OMP surpasses manual selection, and that depending on the case, reinforcement learning-based selection outperforms expert-based selection or viceversa. This work is part of our ongoing efforts in solving the multilevel scheduling problem [3]. Similar scheduling algorithm selection solutions are needed at other parallelism levels, i.e., process level (MPI) and batch level (SLURM), to adapt to unpredictable variations both in applications and resources (including failures) that may arise during execution. Florina M. Ciorba, Ali Mohammed, Jonas H. Müller Korndörfer, Ahmed Eleliemy |
ISPDC | 1 |
| 2023 | DaphneSched: A Scheduler for Integrated Data Analysis PipelinesabstractDAPHNE is a new open-source software infrastructure designed to address the increasing demands of integrated data analysis (IDA) pipelines, comprising data management (DM), high performance computing (HPC), and machine learning (ML) systems. Efficiently executing IDA pipelines is challenging due to their diverse computing characteristics and demands. Therefore, IDA pipelines executed with the DAPHNE infrastructure require an efficient and versatile scheduler to support these demands. This work introduces DaphneSched, the task-based scheduler at the core of DAPHNE [1]. DaphneSched is versatile by incorporating eleven task partitioning and three task assignment techniques, bringing the state-of-the-art closer to the state-of-the-practice task scheduling. To showcase DaphneSched’s effectiveness in scheduling IDA pipelines, we evaluate its performance on two applications: a product recommendation system and training of a linear regression model. We conduct performance experiments on multicore platforms with 20 and 56 cores, respectively. The results show that the versatility of DaphneSched enabled combinations of scheduling strategies that outperform commonly used scheduling techniques by up to 13%. This work confirms the benefits of employing DaphneSched for the efficient execution of applications with IDA pipelines. Ahmed Eleliemy, Florina M. Ciorba |
ISPDC | 2 |
| 2023 | Investigating HPC Job Resource Requests and Job Efficiency ReportingabstractHigh Performance Computing (HPC) systems are ever-evolving, increasing in complexity and heterogeneity, while providing high computing power to scientific and data analysis applications. They are attracting a broad range of users that execute applications from computational chemistry, physics, digital humanities, life sciences, and machine learning, to name a few. Not all users possess the knowledge and expertise to harness the full performance capabilities of complex HPC systems. These users tend to submit inaccurate resource requests for their jobs, especially regarding time limits. This leads to inefficient scheduling of jobs, increased wait times for other users, overall inefficient system utilization, and wasted computing resources and energy, ultimately slowing down scientific discovery. This situation motivates the analysis of the accuracy of resource requests and its reporting to users to raise awareness about job efficiency, improve future job submissions, and reduce job wait times. Existing analyses of job wait times often neglect the connection between the causes of job wait times that can range from user-given and Quality of Service (QoS) limits, to unavailable resources, and approaches of reducing wait times that can be reported to users in an understandable way. In this work, we analyze almost 350’000 jobs collected over 2 months on a local university HPC cluster. Our analysis placed an emphasis on time limit accuracy, QoS characteristics, and reasons for long wait times, as well as how to engage users to improve resource requests. This work shows the importance of analyzing reasons for wait times, approaches to reduce wait times, and motivates further research into supporting users to improve resource requests and system utilization, and to reduce avoidable resource waste. Thomas Jakobsche, Nicolas Lachiche, Florina M. Ciorba |
ISPDC | 3 |
| 2023 | Hot-n-Cold: Mapping the Syscall Attack Surface Using Thermal Side ChannelsabstractAs we increasingly rely on digital technologies, cyber security is of paramount importance. While computing systems offer numerous advantages, they also introduce unwanted security risks. In High Performance Computing, security risks have largely been ignored in the name of high performance [1]. Nevertheless, ensuring security and privacy of computations and data is essential [2]. Linux operating systems, running on all Top500 HPC systems, use kernel and user modes to implement security, which prevents unauthorized access to critical kernel functions. System calls connect the two modes. Therefore, they were frequently attacked, as reported in over 100 Common Vulnerabilities and Exposures (CVEs) in the last 7 years [3]. Combining static and dynamic syscalls analysis [4] [5] has recently been shown to be imperative for creating a syscalls whitelist, to minimize the potential attack surface they introduce. This work introduces a novel dynamic analysis technique, Hot-n-Cold, to detect anomalies in the Linux commands’ behavior by monitoring the CPU temperature. We use Hot-n-Cold to map the syscall attack surface on a local HPC system. Hot-n-Cold can be extended and applied to detect, in real-time, an anomaly that may facilitate creation of an attack. The results on two Linux frequently used commands (ls & chmod) show a positive correlation of up to 80% between the original Linux command and a version augmented with syscalls from CVEs. This work shows the importance of security in HPC, and motivates further research into studying and designing security mechanisms that preserve high performance. Teodora Vasilas, Thomas Jakobsche, Florina M. Ciorba |
ISPDC | 3 |
| 2022 | First Experiences in Performance Benchmarking with the New SPEChpc 2021 SuitesabstractModern High Performance Computing (HPC) sys-tems are built with innovative system architectures and novel programming models to further push the speed limit of computing. The increased complexity poses challenges for performance portability and performance evaluation. The Standard Perfor-mance Evaluation Corporation (SPEC) has a long history of producing industry-standard benchmarks for modern computer systems. SPEC's newly released SPEChpc 2021 benchmark suites, developed by the High Performance Group, are a bold attempt to provide a fair and objective benchmarking tool designed for state-of-the-art HPC systems. With the support of multiple host and accelerator programming models, the suites are portable across both homogeneous and heterogeneous architectures. Different workloads are developed to fit system sizes ranging from a few compute nodes to a few hundred compute nodes. In this work we present our first experiences in performance benchmarking the new SPEChpc2021 suites and evaluate their portability and basic performance characteristics on various popular and emerging HPC architectures, including x86 CPU, NVIDIA GPU, and AMD GPU. This study provides a first-hand experience of executing the SPEChpc 2021 suites at scale on production HPC systems, discusses real-world use cases, and serves as an initial guideline for using the benchmark suites. Holger Brunst, Sunita Chandrasekaran, Florina M. Ciorba, Nick Hagerty, Robert Henschel, Guido Juckeland, Junjie Li 0003, Verónica G. Vergara Larrea, Sandra Wienke, Miguel Zavala |
CCGRID | 3 |
| 2022 | DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines
Patrick Damme, Marius Birkenbach, Constantinos Bitsakos, Matthias Boehm 0001, Philippe Bonnet, Florina M. Ciorba, Mark Dokter, Pawel Dowgiallo, Ahmed Eleliemy, Christian Färber, Georgios I. Goumas, Dirk Habich, Niclas Hedam, Marlies Hofer, Kevin Innerebner, Vasileios Karakostas, Roman Kern, Tomaz Kosar, Alexander Krause 0001, Daniel Krems, Andreas Laber, Wolfgang Lehner, Eric Mier, Marcus Paradies, Bernhard Peischl, Gabrielle Poerwawinata, Stratos Psomadakis, Tilmann Rabl, Piotr Ratuszniak, Pedro Silva 0011, Nikolai Skuppin, Andreas Starzacher, Benjamin Steinwender, Ilin Tolovski, Pinar Tözün, Wojciech Ulatowski, Yuanyuan Wang 0002, Izajasz P. Wrosz, Ales Zamuda, Ce Zhang 0001, Xiao Xiang Zhu 0001 |
CIDR | 6 |
| 2022 | Algorithmic and software development advances for next-generation heterogeneous platformsabstractHeterogeneity is emerging as one of the most profound and challenging characteristics of today's and tomorrow's parallel and distributed computing environments, presenting new and exciting opportunities for their development. Most modern computing systems are heterogeneous, either for organic reasons because components grew independently, as is the case of desktop grids, by design to leverage the strength of specific hardware, as is the case of accelerated systems, or both. The impact of heterogeneity on all forms of parallel and distributed computing is increasing rapidly. Traditional algorithms, programming environments, and tools designed for legacy homogeneous systems will at best achieve a small fraction of the efficiency and the potential performance expected from parallel computing in tomorrow's highly diversified and mixed architectures. Innovative ideas, fresh models, novel algorithms, and other specialized or unified programming environments and tools are needed to efficiently use these new and increasingly diverse computing systems—for accelerating scientific discovery and impactful innovation. The International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar) has been the premier forum over the last 20 years, bringing together researchers to discuss these challenges and the solutions. The wide range of topics includes achieving performance portability on heterogeneous architectures, advances in software environments that facilitate efficient use of heterogeneous systems, performance and energy optimization of numerical and machine learning algorithms on heterogeneous platforms, to name a few. The works presented at the HeteroPar'2020 workshop covered topics clearly exhibiting the significance and growth of the heterogeneous computing field. However, one general trend is apparent: the broad adoption of Graphics Processing Units (GPU) accelerators. Over the last decade, GPUs have been established as the main powerhouse in leadership supercomputers and an invaluable component to accelerate computations for a vast spectrum of applications—from numerical linear algebra libraries powering computational science to various machine learning workloads. This trend is evidenced by the increasing number of GPU-related publications submitted to HeterPar and supported by growing diversity within the GPU world, where AMD accelerator architectures start to compete with Nvidia's comprehensive solutions, along with the third GPU accelerator option—from Intel—available soon. This special issue of Concurrency and Computation: Practice and Experience contains six selected papers from the HeteroPar'2020 workshop. We hope you find them interesting and stimulating new ideas and future advancements for next-generation heterogeneous platforms. The 18th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2020) was held in Warsaw, Poland, on 25 August 2020. For the 11th time, this workshop was organized in conjunction with the Euro-Par annual series of international conferences. Because of the COVID-19 pandemic, HeteroPar'2020 was held as a virtual event. Sixteen articles were submitted for review, with authors from eight countries. Each paper secured at least three reviews from members of the program committee, whereas 12 submissions received at least four reviews. After a thorough peer-reviewing process that included discussion and agreement among reviewers whenever necessary, nine articles were selected for presentation at the workshop. The review process focused on the quality of the papers, their innovative ideas and applicability to heterogeneous computing. The topics addressed in the accepted papers include domain-specific languages for numerical algorithms, virtualization for CUDA applications, unified memory in CUDA, porting CUDA codes to AMD GPUs, management of heterogeneous cloud resources, GPU implementation of graph neural networks, GPU and CPU signal processing for a wildlife tracking system, parallelization of the k-means algorithm on CPU-GPU platforms, and a portable solver for systems of linear equations. An integral part of the workshop was two keynote talks given by Enrique S. Quintana-Orti (Technical University of Valencia, Spain) and Tal Ben-Nun (ETHZ Zurich, Switzerland) devoted to using approximate and transprecision computing in sparse linear solvers, and data-centric approach for performance portability on heterogeneous architecture, respectively. After the workshop, the program committee invited the authors of the presented works to submit revised and extended versions of their contributions as part of the papers submitted to this special issue. These new versions were reviewed independently again by at least three reviewers. Finally, six papers were accepted for publication in the special issue. They are summarized below. Aliaga et al.1 focus on optimizing the sparse matrix–vector product (SpMV), which dictates, to a large extent, the performance of a considerable variety of scientific applications. The proposed approach introduces a variant of the coordinate sparse matrix format that allows combining load-balancing with compressing both the indexing arrays and the numerical information to reduce the pressure on memory while using the available compute power of modern CPUs and GPUs efficiently. This approach is multi-platform, in the sense that the realizations are built upon common principles but differ in the implementation details, which are adapted either to avoid thread divergence in the GPU case or to maximize compression for multicore architectures. The evaluation on the two last generations of NVIDIA GPUs as well as Intel and AMD processors demonstrate the benefits of the new kernels compared with the optimized implementations of SpMV in Nvidia's cuSPARSE and Intel's MKL libraries. k-Means is a standard algorithm for clustering data used as the final step for high-quality spectral clustering. To overcome the scalability challenge when processing large datasets, the authors of paper2 propose to apply also the k-means algorithm as a preprocessing task to reduce the input data instances. Additionally, parallel optimization techniques are introduced to improve the efficiency of the k-means algorithm on CPU and GPU. Notably, a two-step summation method with package processing is used to handle the effect of rounding errors that may occur during the phase of updating cluster centroids. The extensive experiments on synthetic and real-world datasets containing millions of instances exhibit a speedup up to 7 for the k-means iteration time on GPU versus 20/40 CPU threads using AVX units while achieving double-precision accuracy with single-precision computations. Dmitruk et al.3 show how the OpenACC standard can be efficiently used to implement solvers for tridiagonal Toeplitz systems of linear equations for a variety of modern GPU-accelerated and multicore architectures. Two parallel algorithms are studied concerning particular assumptions about coefficient matrices. In the first case, a new, faster implementation of the divide and conquer method is proposed, while in the second one, a novel, vectorizable algorithm is introduced. Using both column-wise and row-wise matrix storage formats is studied, along with efficient conversion between them using cache memory to improve the overall performance. It is also shown how to tune the performance by predicting the best values of the methods' parameters. Numerical experiments performed on Intel CPUs and Nvidia GPUs confirm the excellent performance and accuracy of the developed implementations. Robust high-performance implementations of signal-processing tasks performed by a high-throughput wildlife tracking system are presented by Rubinpur et al.4 The system tracks radio transmitters attached to wild animals by estimating the time of arrival of radio packets to multiple receivers. The time-consuming estimation of wideband radio signals is a bottleneck that limits the system's throughput. A sequential high-performance CPU implementation has been developed first, and then a GPU implementation to overcome this bottleneck. The authors carefully evaluate the performance of these real-world codes. The evaluation indicates that the GPU version dramatically improves both performance and power-performance efficiency relative to a desktop CPU—a scenario typical for current base stations. Performance improves by more than 50 times on a high-end GPU and more than four times with a GPU platform that consumes almost five times less power than the CPU one. The desire to take advantage of virtualization in heterogeneous computing resources with GPU accelerators motivates Eiling et al.5 Currently, GPUs do not offer virtualization support that enables fine-grained control, increased flexibility, and fault tolerance. The authors present Cricket—a transparent and low-overhead solution to GPU virtualization that enables future research of various virtualization techniques, due to its open-source nature. Cricket supports remote execution and checkpoint/restart of CUDA applications. Both features allow the distribution of GPU tasks dynamically and flexibly across computing nodes and the multitenant usage of GPU resources, improving their flexibility and utilization in high-performance and cloud computing. Solving partial differential equations (PDEs) on unstructured grids is a cornerstone of engineering and scientific computing. Alhaddad et al.6 introduce the HighPerMeshes C++-embedded domain-specific language (DSL) that bridges the abstraction gap between the mathematical formulation of mesh-based algorithms for PDE problems and an increasing number of heterogeneous platforms with their various programming models. The HighPerMeshes DSL aims at higher productivity of the code development for multiple target platforms. For this aim, the OpenCL is used as a backend, targeting various GPUs and other heterogeneous architectures such as FPGAs. Apart from describing the basic structure of the DSL, its usage is demonstrated with three examples. The mapping of the abstract algorithmic description onto parallel hardware, including compute clusters, is also presented. Finally, the achievable performance and scalability are demonstrated for different example problems. The guest editors of this special issue wish to thank the authors of the submitted papers, the reviewers for the careful evaluation of the papers, and the valuable suggestions that helped the authors to improve their contributions. In addition, we would like to sincerely thank Prof. David W. Walker (Editor-in-Chief of Concurrency and Computation: Practice and Experience) for the opportunity to guest edit this special issue and for his guidance during this process. Data sharing is not applicable to this article as no datasets were generated or analyzed in this study. Roman Wyrzykowski, Florina M. Ciorba |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | LB4OMP: A Dynamic Load Balancing Library for Multithreaded ApplicationsabstractExascale computing systems will exhibit high degrees of hierarchical parallelism, with thousands of computing nodes and hundreds of cores per node. Efficiently exploiting hierarchical parallelism is challenging due to load imbalance that arises at multiple levels. OpenMP is the most widely-used standard for expressing and exploiting the ever-increasing node-level parallelism. The scheduling options in OpenMP are insufficient to address the load imbalance that arises during the execution of multithreaded applications. The limited scheduling options in OpenMP hinder research on novel scheduling techniques which require comparison with others from the literature. This work introduces LB4OMP, an open-source dynamic load balancing library that implements successful scheduling algorithms from the literature. LB4OMP is a research infrastructure designed to spur and support present and future scheduling research, for the benefit of multithreaded applications performance. Through an extensive performance analysis campaign, we assess the effectiveness and demystify the performance of all loop scheduling techniques in the library. We show that, for numerous applications-systems pairs, the scheduling techniques in LB4OMP outperform the scheduling options in OpenMP. Node-level load balancing using LB4OMP leads to reduced cross-node load imbalance and to improved MPI+OpenMP applications performance, which is critical for Exascale computing. Jonas H. Müller Korndörfer, Ahmed Eleliemy, Ali Mohammed, Florina M. Ciorba |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Automated Scheduling Algorithm Selection and Chunk Parameter Calculation in OpenMPabstractIncreasing node and cores-per-node counts in supercomputers render scheduling and load balancing critical for exploiting parallelism. OpenMP applications can achieve high performance via careful selection of schedulingkindandchunkparameters on a per-loop, per-application, and per-system basis from a portfolio of advanced scheduling algorithms (Korndörferet al., 2022). This selection approach is time-consuming, challenging, and may need to change during execution. We proposeAuto4OMP, a novel approach for automated load balancing of OpenMP applications. With Auto4OMP, we introduce three schedulingalgorithm selection methodsand anexpert-defined chunk parameterfor OpenMP'sscheduleclause'skindandchunk, respectively. Auto4OMP extends the OpenMPschedule(auto)andchunkparameter implementation in LLVM's OpenMP runtime library to automatically select a scheduling algorithm and calculate a chunk parameter during execution. Loop characteristics are inferred in Auto4OMP from the loop execution over the application's time-steps. The experiments performed in this work show that Auto4OMP improves applications performance by up to$11\%$compared to LLVM'sschedule(auto)implementation and outperforms manual selection. Auto4OMP improves MPI+OpenMP applications performance byexplicitlyminimizing thread- andimplicitlyreducing process-load imbalance. Ali Mohammed, Jonas H. Müller Korndörfer, Ahmed Eleliemy, Florina M. Ciorba |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | An Execution Fingerprint Dictionary for HPC Application RecognitionabstractApplications running on HPC systems waste time and energy if they: (a) use resources inefficiently, (b) deviate from allocation purpose (e.g. cryptocurrency mining), or (c) encounter errors and failures. It is important to know which applications are running on the system, how they use the system, and whether they have been executed before. To recognize known applications during execution on a noisy system, we draw inspiration from the way Shazam recognizes known songs playing in a crowded bar. Our contribution is an Execution Fingerprint Dictionary (EFD) that stores execution fingerprints of system metrics (keys) linked to application and input size information (values) as key-value pairs for application recognition. Related work often relies on extensive system monitoring (many system metrics collected over large time windows) and employs machine learning methods to identify applications. Our solution only uses the first 2 minutes and a single system metric to achieve F-scores above 95 percent, providing comparable results to related work but with a fraction of the necessary data and a straightforward mechanism of recognition. Thomas Jakobsche, Nicolas Lachiche, Aurélien Cavelan, Florina M. Ciorba |
CLUSTER | 4 |
| 2020 | SimAS: A simulation-assisted approach for the scheduling algorithm selection under perturbationsabstractSummary Many scientific applications consist of large and computationally intensive loops. Dynamic loop self‐scheduling (DLS) techniques are used to parallelize and to balance the load of such applications during execution. Load imbalance arises from variations in the loop iteration (or tasks) execution times, caused by problem, algorithmic, or systemic characteristics. Variations in systemic characteristics are referred to as perturbations. Our hypothesis is that no single DLS technique can achieve the absolute best performance under various perturbations on heterogeneous high‐performance computing (HPC) systems. Therefore, the selection of the most efficient DLS technique is critical to achieve the best application performance. The goal of this work is to solve the algorithm selection problem for the scheduling of computationally intensive applications under perturbations. Existing work only considers perturbations caused by variations in the delivered computational speed of the HPC systems. However, perturbations in available network bandwidth or latency are inevitable on production HPC systems. A simulation‐assisted scheduling algorithm selection (SimAS) approach is introduced herein as a novel control‐theoretic‐inspired approach to select DLS techniques dynamically that improve the performance of applications executing on heterogeneous HPC systems under perturbations. The present work examines the performance of seven applications on a heterogeneous HPC system under all the above system perturbations. SimAS is evaluated using native and simulative experiments. The performance results confirm the original hypothesis that motivates this work. The experimental evaluation shows that the SimAS‐based DLS selection identifies the most efficient technique and improves application performance in most cases. Ali Mohammed, Florina M. Ciorba |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | An approach for realistically simulating the performance of scientific applications on high performance computing systemsabstractScientific applications often contain large, computationally-intensive, and irregular parallel loops or tasks that exhibit stochastic behavior leading to load imbalance. Load imbalance often manifests during the execution of parallel scientific applications on large and complex high performance computing (HPC) systems. The extreme scale of HPC systems on the road to Exascale computing only exacerbates the loss in performance due to load imbalance. Dynamic loop self-scheduling (DLS) techniques are instrumental in improving the performance of scientific applications on HPC systems via load balancing. Selecting a DLS technique that results in the best performance for different problem and system sizes requires a large number of exploratory experiments. Currently, a theoretical model that can be used to predict the scheduling technique that yields the best performance for a given problem and system has not yet been identified. Therefore, simulation is the most appropriate approach for conducting such exploratory experiments in a reasonable amount of time. However, conducting realistic and trustworthy simulations of application performance under different configurations is challenging. This work devises an approach to realistically simulate computationally-intensive scientific applications that employ DLS and execute on HPC systems. The proposed approach minimizes the sources of uncertainty in the simulative experiments results by bridging the native and simulative experimental approaches. A new method is proposed to capture the variation of application performance between different native executions. Several approaches to represent the application tasks (or loop iterations) are compared to establish their influence on the simulative application performance. A novel simulation strategy is introduced that applies the proposed approach, which transforms a native application code into simulative code. The native and simulative performance of two computationally-intensive scientific applications that employ eight task scheduling techniques (static, nonadaptive dynamic, and adaptive dynamic) are compared to evaluate the realism of the proposed simulation approach. The comparison of the performance characteristics extracted from the native and simulative performance shows that the proposed simulation approach fully captured most of the performance characteristics of interest. This work shows and establishes the importance of simulations that realistically predict the performance of DLS techniques for different applications and system configurations. Ali Mohammed, Ahmed Eleliemy, Florina M. Ciorba, Franziska Kasielke, Ioana Banicescu |
Future Gener. Comput. Syst. | 3 |
| 2019 | Detection of Silent Data Corruptions in Smoothed Particle Hydrodynamics SimulationsabstractSilent data corruptions (SDCs) hinder the correctness of long-running scientific applications on large scale computing systems. Selective particle replication (SPR) is proposed herein as the first particle-based replication method for detecting SDCs in Smoothed particle hydrodynamics (SPH) simulations. SPH is a mesh-free Lagrangian method commonly used to perform hydrodynamical simulations in astrophysics and computational fluid dynamics. SPH performs interpolation of physical properties over neighboring discretization points (called SPH particles) that dynamically adapt their distribution to the mass density field of the fluid. When a fault (e.g., a bit-flip) strikes the computation or the data associated with a particle, the resulting error is silently propagated to all nearest neighbors through such interpolation steps. SPR replicates the computation and data of a few carefully selected SPH particles. SDCs are detected when the data of a particle differs, due to corruption, from its replicated counterpart. SPR is able to detect many DRAM SDCs as they propagate by ensuring that all particles have at least one neighbor that is replicated. The detection capabilities of SPR were assessed through a set of error-injection and detection experiments and the overhead of SPR was evaluated via a set of strong-scaling experiments conducted on a HPC system. The results show that SPR achieves detection rates of 91-99.9%, no false-positives, at an overhead of 1-10%. Aurélien Cavelan, Rubén M. Cabezón, Florina M. Ciorba |
CCGRID | 3 |
| 2019 | Algorithm-Based Fault Tolerance for Parallel Stencil ComputationsabstractThe increase in HPC systems size and complexity, together with increasing on-chip transistor density, power limitations, and number of components, render modern HPC systems subject to soft errors. Silent data corruptions (SDCs) are typically caused by such soft errors in the form of bit-flips in the memory subsystem and hinder the correctness of scientific applications. This work addresses the problem of protecting a class of iterative computational kernels, called stencils, against SDCs when executing on parallel HPC systems. Existing SDC detection and correction methods are in general either inaccurate, inefficient, or targeting specific application classes that do not include stencils. This work proposes a novel algorithm-based fault tolerance (ABFT) method to protect scientific applications that contain arbitrary stencil computations against SDCs. The ABFT method can be applied both online and offline to accurately detect and correct SDCs in 2D and 3D parallel stencil computations. We present a formal model for the proposed method including theorems and proofs for the computation of the associated check-sums as well as error detection and correction. We experimentally evaluate the use of the proposed ABFT method on a real 3D stencil-based application (HotSpot3D) via a fault-injection, detection, and correction campaign. Results show that the proposed ABFT method achieves less than 8% overhead compared to the performance of the unprotected stencil application. Moreover, it accurately detects and corrects SDCs. While the offline ABFT version corrects errors more accurately, it may incur a small additional overhead than its online counterpart. Aurélien Cavelan, Florina M. Ciorba |
CLUSTER | 2 |
| 2019 | Anomaly Detection in High Performance Computers: A Vicinity PerspectiveabstractIn response to the demand for higher computational power, the number of computing nodes in high performance computers (HPC) increases rapidly. Exascale HPC systems are expected to arrive by 2020. With drastic increase in the number of HPC system components, it is expected to observe a sudden increase in the number of failures which, consequently, poses a threat to the continuous operation of the HPC systems. Detecting failures as early as possible and, ideally, predicting them, is a necessary step to avoid interruptions in HPC systems operation. Anomaly detection is a well-known general purpose approach for failure detection, in computing systems. The majority of existing methods are designed for specific architectures, require adjustments on the computing systems hardware and software, need excessive information, or pose a threat to users' and systems' privacy. This work proposes a node failure detection mechanism based on a vicinity-based statistical anomaly detection approach using passively collected and anonymized system log entries. Application of the proposed approach on system logs collected over 8 months indicates an anomaly detection precision between 62% to 81%. Siavash Ghiasvand, Florina M. Ciorba |
ISPDC | 2 |
| 2019 | Exploring Loop Scheduling Enhancements in OpenMP: An LLVM Case StudyabstractOpenMP is the de-facto standard for parallel programming on shared-memory systems. The choice of scheduling methods in OpenMP work sharing parallel loops is a critical aspect for performance, especially for computationally-intensive and irregular parallel loops. In this work, we explore loop scheduling enhancements in OpenMP. Three loop scheduling choices are covered today in the OpenMP standard: static, guided, and dynamic. These are no longer sufficient to address the load imbalance that adversely affects the execution of computationally-intensive and irregular parallel loops. In this work, we present a generic methodology for exploring loop scheduling enhancements in OpenMP that allows the implementation, testing, and usage of additional (more advanced) loop scheduling choices in OpenMP runtime systems. We showcase the methodology by enhancing the LLVM OpenMP runtime with an additional dynamic loop self-scheduling (DLS) technique, known to offer superior load balancing over the existing OpenMP scheduling choices for computationally-intensive and irregular parallel loops. We analyze the overhead of the (existing and newly added) OpenMP loop scheduling methods and show that the proposed methodology incurs no additional overhead. We also study the performance of four benchmarks using the enhanced LLVM OpenMP runtime. The results show that, for the four benchmarks considered, no single loop scheduling strategy outperforms the others. The newly implemented DLS technique provides an additional opportunity for improved execution time with the LLVM OpenMP runtime, which was not possible before this study. Our newly implemented scheduling strategy is competitive with the best previous scheduling choices. This methodology for exploring loop scheduling enhancements in OpenMP lays the foundation for further loop scheduling additions and explorations in OpenMP. Franziska Kasielke, Ronny Tschüter, Christian Iwainsky, Markus Velten, Florina M. Ciorba, Ioana Banicescu |
ISPDC | 5 |
| 2019 | Dynamic Loop Scheduling Using MPI Passive-Target Remote Memory AccessabstractScientific applications often contain large computationally-intensive parallel loops. Loop scheduling techniques aim to achieve load balanced executions of such applications. For distributed-memory systems, existing dynamic loop scheduling (DLS) libraries are typically MPI-based, and employ a master-worker execution model to assign variably-sized chunks of loop iterations. The master-worker execution model may adversely impact performance due to the master-level contention. This work proposes a distributed chunk-calculation approach that does not require the master-worker execution scheme. Moreover, it considers the novel features in the latest MPI standards, such as passive-target remote memory access, shared-memory window creation, and atomic read-modify-write operations. To evaluate the proposed approach, five well-known DLS techniques, two applications, and two heterogeneous hardware setups have been considered. The DLS techniques implemented using the proposed approach outperformed their counterparts implemented using the traditional master-worker execution model. Ahmed Eleliemy, Florina M. Ciorba |
PDP | 2 |
| 2018 | Towards a Mini-App for Smoothed Particle Hydrodynamics at ExascaleabstractThe smoothed particle hydrodynamics (SPH) technique is a purely Lagrangian method, used in numerical simulations of fluids in astrophysics and computational fluid dynamics, among many other fields. SPH simulations with detailed physics represent computationally-demanding calculations. The parallelization of SPH codes is not trivial due to the absence of a structured grid. Additionally, the performance of the SPH codes can be, in general, adversely impacted by several factors, such as multiple time-stepping, long-range interactions, and/or boundary conditions. This work presents insights into the current performance and functionalities of three SPH codes: SPHYNX, ChaNGa, and SPH-flow. These codes are the starting point of an interdisciplinary co-design project, SPH-EXA, for the development of an Exascale-ready SPH mini-app. To gain such insights, a rotating square patch test was implemented as a common test simulation for the three SPH codes and analyzed on two modern HPC systems. Furthermore, to stress the differences with the codes stemming from the astrophysics community (SPHYNX and ChaNGa), an additional test case, the Evrard collapse, has also been carried out. This work extrapolates the common basic SPH features in the three codes for the purpose of consolidating them into a pure-SPH, Exascale-ready, optimized, mini-app. Moreover, the outcome of this serves as direct feedback to the parent codes, to improve their performance and overall scalability. Danilo Guerrera, Rubén M. Cabezón, Jean-Guillaume Piccinali, Aurélien Cavelan, Florina M. Ciorba, David Imbert, Lucio Mayer, Darren S. Reed |
CLUSTER | 5 |
| 2018 | Assessing Data Usefulness for Failure Analysis in Anonymized System LogsabstractSystem logs are a valuable source of information for the analysis and understanding of systems behavior for the purpose of improving their performance. Such logs contain various types of information, including sensitive information. Information deemed sensitive can either directly be extracted from system log entries by correlation of several log entries, or can be inferred from the combination of the (non-sensitive) information contained within system logs with other logs and/or additional datasets. The analysis of system logs containing sensitive information compromises data privacy. Therefore, various anonymization techniques, such as generalization and suppression have been employed, over the years, by data and computing centers to protect the privacy of their users, their data, and the system as a whole. Privacy-preserving data resulting from anonymization via generalization and suppression may lead to significantly decreased data usefulness, thus, hindering the intended analysis for understanding the system behavior. Maintaining a balance between data usefulness and privacy preservation, therefore, remains an open and important challenge. Irreversible encoding of system logs using collision-resistant hashing algorithms, such as SHAKE-128, is a novel approach previously introduced by the authors to mitigate data privacy concerns. The present work describes a study of the applicability of the encoding approach from earlier work on the system logs of a production high performance computing system. Moreover, a metric is introduced to assess the data usefulness of the anonymized system logs to detect and identify the failures encountered in the system. Siavash Ghiasvand, Florina M. Ciorba |
ISPDC | 2 |
| 2018 | Experimental Verification and Analysis of Dynamic Loop Scheduling in Scientific ApplicationsabstractScientific applications are often irregular and characterized by large computationally-intensive parallel loops. Dynamic loop scheduling (DLS) techniques improve the performance of computationally-intensive scientific applications via load balancing of their execution on high-performance computing (HPC) systems. Identifying the most suitable choices of data distribution strategies, system sizes, and DLS techniques which improve the performance of a given application, requires intensive assessment and a large number of exploratory native experiments (using real applications on real systems), which may not always be feasible or practical due to associated time and costs. In such cases, simulative experiments are more appropriate for studying the performance of applications. This motivates the question of 'How realistic are the simulations of executions of scientific applications using DLS on HPC platforms?' In the present work, a methodology is devised to answer this question. It involves the experimental verification and analysis of the performance of DLS in scientific applications. The proposed methodology is employed for a computer vision application executing using four DLS techniques on two different HPC platforms, both via native and simulative experiments. The evaluation and analysis of the native and simulative results indicate that the accuracy of the simulative experiments is strongly influenced by the approach used to extract the computational effort of the application (FLOP-or time-based), the choice of application model representation into simulation (data or task parallel) and the available HPC subsystem models in the simulator (multi-core CPUs, memory hierarchy and network topology). The minimum and the maximum percent errors achieved between the native and the simulative experiments are 0.95% and 8.03%, respectively. Ali Mohammed, Ahmed Eleliemy, Florina M. Ciorba, Franziska Kasielke, Ioana Banicescu |
ISPDC | 3 |
| 2017 | An Autonomic Approach for the Selection of Robust Dynamic Loop Scheduling TechniquesabstractParallel applications are highly irregular and high performance computing (HPC) infrastructures are very complex. The HPC applications of interest herein are timestepping scientific applications (TSSA). Often, TSSA involve the repeated execution of multiple parallel loops with thousands of iterations and irregular behavior. Dynamic loop scheduling (DLS) techniques were developed over time and have proven to be effective in scheduling parallel loops for achieving load balancing of TSSA. Using a single particular DLS technique throughout the entire execution of a time-step, or even over the entire application, does not guarantee optimal performance due to the unpredictable variations in problem and algorithmic characteristics as well as those of the infrastructure capabilities. For that reason, an autonomic selection of DLS techniques as function of the parallel loop execution time has shown to improve application performance. Recently, a robustness metric of DLS techniques, named "flexibility", has been proposed to estimate the capability of a DLS technique to resist to variations in the loop iterations execution time. To improve the performance of TSSA, we propose in this work an approach that involves the autonomic selection of DLS techniques as function of the flexibility of DLS techniques. The first major novelty of our approach lies in the use of state-of-the-art reinforcement learning (RL) algorithms as smart agents. The second novelty lies in the design of a modified flexibility metric. The third major novelty resides in using the new modified flexibility metric as a reward for the smart agents. The fourth novelty is the evaluation of the proposed approach within a simulated environment, in particular using the SimGrid-SMPI interface to execute DLS algorithms. We discuss the advantages and the limitations of the new proposed flexibility metric as a reward. Anthony Boulmier, Ioana Banicescu, Florina M. Ciorba, Nabil Abdennadher |
ISPDC | 3 |
| 2017 | Exploring the Relation between Two Levels of Scheduling Using a Novel Simulation ApproachabstractModern high performance computing (HPC) systems exhibit a rapid growth in size, both “horizontally” in the number of nodes, as well as “vertically” in the number of cores per node. As such, they offer additional levels of hardware parallelism. Each level requires and employs algorithms for appropriately scheduling the computational work at the respective level. The present work explores the relation between two scheduling levels: batch and application. To understand and explore this relation, a novel simulation approach is presented that bridges two existing simulators from the two scheduling levels. A novel two-level simulator that implements the proposed approach is introduced. The two-level simulator is used to simulate all combinations of three batch scheduling and four application scheduling algorithms from the literature. These combinations are considered for allocating resources and executing the parallel jobs from a workload of a production HPC system. The results of the scheduling experiments reveal the strong relation between decisions taken at the two scheduling levels and their mutual influence. Complementing the simulations, the two-level simulator produces abstract parallel execution traces, which can visually be examined and illustrate the execution of different jobs and, for each job, the execution of its tasks at node and core levels, respectively. Ahmed Eleliemy, Ali Mohammed, Florina M. Ciorba |
ISPDC | 3 |
| 2017 | Towards the Reproducibility of Using DLS Techniques in Scientific ApplicationsabstractReproducibility of the execution of scientific applications on parallel and distributed systems is a growing interest, underlying the trustworthiness of the experiments and the conclusions derived from experiments. Dynamic loop scheduling (DLS) techniques are an effective approach towards performance improvement of scientific applications via load balancing. These techniques address algorithmic and systemic sources of load imbalance by dynamically assigning tasks to processing elements. The DLS techniques have demonstrated their effectiveness when applied in real applications. Complementing native experiments, simulation is a powerful tool for studying the behavior of parallel and distributed applications. This work is a comprehensive reproducibility study of experiments using DLS techniques published in the earlier literature to verify their implementations into SimGrid-MSG [1]. The reproducibility study is carried out by comparing the performance of the SimGrid-MSG-based experiments with those reported in [2]. In earlier work [3] it was shown that a very detailed degree of information regarding the experiments to be reproduced is essential for successful reproducibility. This work concentrates on the reproducibility of experiments with variable application behavior and a high degree of parallelism . It is shown that reproducing measurements of applications with high variance is challenging, albeit feasible and useful. The success of the present reproducibility study denotes the fact that the implementation of the DLS techniques in SimGrid-MSG is verified for the considered applications and systems. Thus, it enables well-founded future research using the DLS techniques in simulation. Franziska Hoffeins, Florina M. Ciorba, Ioana Banicescu |
ISPDC | 2 |
| 2016 | Lessons Learned from Spatial and Temporal Correlation of Node Failures in High Performance ComputersabstractIn this paper we study the correlation of node failures in time and space. Our study is based on measurements of a production high performance computer over an 8-month time period. We draw possible types of correlations between node failures and show that, in many cases, there are direct correlations between observed node failures. The significance of such a study is twofold: achieving a clearer understanding of correlations between node failures and enabling failure detection as early as possible. The results of this study are aimed at helping the system administrators minimize (or even prevent) the destructive effects of correlated node failures. Siavash Ghiasvand, Florina M. Ciorba, Ronny Tschüter, Wolfgang E. Nagel |
PDP | 2 |
| 2015 | Investigating the Resilience of Dynamic Loop Scheduling in Heterogeneous Computing SystemsabstractTo improve the performance of complex scientific applications, dynamic loop scheduling(DLS) techniques are often employed for load balancing. However, it is a challenge to select the most resilient scheduling technique for guaranteeing optimized performance of scientific applications on large-scale computing systems. Such systems comprise widely distributed and highly heterogeneous resources, and often are prone to failures. Hence, in this work we perform a comprehensive study of resilience of DLS techniques. In our study, we employed Sim Grid-based simulations. The use of a simulation framework assists in overcoming the limits of quantifying the resilience and evaluating the performance of the DLS techniques on real test beds by allowing us to model, control, and reproduce large scale computing systems with irregular behaviour in order to analyze the resilience of DLS technique son computationally intensive scientific applications. The results are used to compare the resilience of scheduling techniques under different case scenarios comprising of variable problem sizes, system sizes, characteristics of the variations in the application task computation times, and those of the processor availabilities and failures. Nitin Sukhija, Ioana Banicescu, Florina M. Ciorba |
ISPDC | 3 |
| 2014 | Robustness Prediction and Evaluation of Divisible Load Scheduling on Computing Systems with Unpredictable VariationsabstractThis work addresses the problem of predicting and evaluating the robustness of divisible load scheduling of data parallel workloads (also called arbitrarily divisible workloads) onto high performance parallel and distributed computing systems with unpredictable variations. Divisible load scheduling is based on the divisible load theory (DLT) which offers a linear, deterministic, and tractable model for scheduling arbitrarily divisible workloads. High performance parallel and distributed computing systems operate in an environment characterized by unpredictable variations (or perturbations) such as system load or unexpected resource failures. In this work, we analytically evaluate and empirically determine the robustness of divisible load scheduling algorithms (called DLT algorithms) with respect to variations in processor availability via realistic simulation. The realism arises from modeling the characteristics of two applications from the NAS parallel benchmark suite, as well as from modeling the target system as a 3D torus topology, one of the most widely used interconnection networks. Extending prior related work, we conduct an analytical evaluation as well as a simulation-based study of the robustness of divisible load scheduling for scheduling the two NAS parallel benchmarks, namely, embarrassingly parallel (EP) which is computationally intensive and integer sort (IS) which is communication intensive. The simulation results indicate that the robustness observed via simulation of scheduling the EP benchmark is always within the analytically predicted range, and that the robustness observed via simulation of scheduling the IS benchmark is within the analytically predicted range in the best case, and within 6.56% on average in the worst case. Mahadevan Balasubramaniam, Ioana Banicescu, Florina M. Ciorba |
ISPDC | 3 |
| 2013 | Scheduling Data Parallel Workloads - A Comparative Study of Two Common Algorithmic ApproachesabstractThe dynamic loop scheduling (DLS) and the divisible load theory (DLT) are two common algorithmic approaches used in the scheduling of arbitrarily divisible workloads. Despite sharing the same goal of achieving load balancing via scheduling, they are fundamentally different. Specifically, the DLS approach is probabilistic and platform agnostic, whereas the DLT approach is deterministic and platform aware. To the best of our knowledge, this is the first work to conduct a comparative study and a performance analysis of the two approaches. The study is beneficial for identifying the application, algorithmic, and systemic characteristics that favor one approach over the other. In this work, we report the results of a comparative performance study of the two approaches. Simulations provide a greater flexibility and control over running experiments on a real computing platform. Hence, we employ simulations in this work to study the behavior of the DLS and the DLT approaches on two types of network topologies, namely a single level tree network and a linear array network. Various application, algorithmic, and systemic characteristics introduce load imbalance, and therefore we use a wide range of synthetically generated workloads and variable system conditions to inject load imbalance into the simulated platform. The simulation results demonstrate the effectiveness of applying the DLS approach, when the computing environments are mainly characterized by high load imbalance, and that of applying the DLT approach when they are mainly characterized by high communication costs. Mahadevan Balasubramaniam, Ioana Banicescu, Florina M. Ciorba |
ICPP | 3 |
| 2013 | Analyzing the Robustness of Scheduling Algorithms Using Divisible Load Theory on Heterogeneous SystemsabstractArbitrarily divisible workloads are present in a large class of scientific applications, such as N-body simulations, Monte Carlo simulations, CFD applications, and others. Divisible load theory (DLT) provides a tractable approach to the scheduling of arbitrarily divisible workloads. High performance parallel and distributed systems may operate in an unreliable environment, and a robust system is expected to deliver a certain level of performance when operating in such an environment. To the best of our knowledge, this is the first work to study and analyze the robustness of DLT algorithms. Using simulations, a study of the resiliency of DLT to variations in certain system features, such as the network latency, the network bandwidth, and the processor availability on a single level tree topology is presented. The simulation results demonstrate the robustness of the DLT algorithms under certain application and system characteristics. Mahadevan Balasubramaniam, Ioana Banicescu, Florina M. Ciorba |
ISPDC | 3 |
| 2013 | Predicting the Flexibility of Dynamic Loop Scheduling Using an Artificial Neural NetworkabstractIn this paper, an artificial neural network (ANN) model is proposed to predict the flexibility (or robustness against system load fluctuations in heterogeneous computing systems) of dynamic loop scheduling (DLS) methods. The multilayer perceptron (MLP) ANN model has been used to predict the degree of robustness of a DLS method, given specific values for the problem size, the system size, and the characteristics of the system load fluctuations as a compound effect of the variations in the application's iteration execution times and the processor availabilities. The developed MLP ANN model can be useful in an effective selection of the most robust DLS technique for scheduling a certain type of scientific application onto a given set of non-dedicated heterogeneous processors, when their system load is expected to fluctuate unpredictably during the application's runtime. Brandon M. Malone, Nitin Sukhija, Ioana Banicescu, Florina M. Ciorba |
ISPDC | 5 |
| 2012 | Analyzing the Robustness of Dynamic Loop Scheduling for Heterogeneous Computing SystemsabstractScheduling scientific applications in parallel on non-dedicated, heterogeneous systems, where the computing resources may differ in availability, is a challenging task, and requires efficient execution and robust scheduling methods. Dynamic loop scheduling methods provide means to achieve the desired robust performance. These methods are based on probabilistic analyses and are inherently robust. However, a methodology is required to measure the robustness of the dynamic loop scheduling methods that ensures their performance in unpredictably changing computing environments. In this paper, a methodology is proposed for performing robustness analysis of the dynamic loop scheduling techniques using a metric, formulated in earlier work, to measure their robustness in heterogeneous computing systems with uncertainties. The dynamic loop scheduling methods have been implemented in a simulation. The experimental results were used as an input to the proposed methodology, which in turn has been used to experimentally analyze the robustness of a number of dynamic loop scheduling methods on a heterogeneous system with variable availability. Nitin Sukhija, Ioana Banicescu, Florina M. Ciorba |
ISPDC | 4 |
| 2012 | Towards the optimal synchronization granularity for dynamic scheduling of pipelined computations on heterogeneous computing systemsabstractSUMMARY Loops are the richest source of parallelism in scientific applications. A large number of loop scheduling schemes have therefore been devised for loops with and without data dependencies (modeled as dependence distance vectors) on heterogeneous clusters. The loops with data dependencies require synchronization via cross‐node communication. Synchronization requires fine‐tuning to overcome the communication overhead and to yield the best possible overall performance. In this paper, a theoretical model is presented to determine the granularity of synchronization that minimizes the parallel execution time of loops with data dependencies when these are parallelized on heterogeneous systems using dynamic self‐scheduling algorithms. New formulas are proposed for estimating the total number of scheduling steps when a threshold for the minimum work assigned to a processor is assumed. The proposed model uses these formulas to determine the synchronization granularity that minimizes the estimated parallel execution time. The accuracy of the proposed model is verified and validated via extensive experiments on a heterogeneous computing system. The results show that the theoretically optimal synchronization granularity, as determined by the proposed model, is very close to the experimentally observed optimal synchronization granularity, with no deviation in the best case, and within 38.4% in the worst case. Copyright © 2012 John Wiley & Sons, Ltd. Ioannis Riakiotakis, Florina M. Ciorba, Theodore Andronikos, George K. Papakonstantinou, Anthony T. Chronopoulos |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Enhancing the Functionality of a GridSim-Based Scheduler for Effective Use with Large-Scale Scientific ApplicationsabstractThe performance of computationally intensive scientific applications on the underlying system can be maximized by providing application-level load balancing of loop iterates via the use of dynamic loop scheduling (DLS) algorithms. These DLS methods are based on probabilistic analyses, and therefore account for unpredictable variations of algorithmic, systemic and application level characteristics. A considerable number of DLS algorithms has been proposed in the last decade, and some of them have been effectively integrated into scientific and engineering applications, yielding significant performance improvements. However, scheduling scientific applications in large-scale distributed systems where, the chances of failure, such as processor or link failure, are high, makes the problem of achieving a load balanced execution even more challenging. Although real experiments are necessary to verify the benefits of using DLS, they prove to be very time consuming when every level of detail is required for the assessment of the execution of complex, data parallel and irregular scientific applications using DLS on a wide range of architectural platforms and computational environments. Thus, we propose the use of simulators which can give results that are not always experimentally measurable with the current technology. Simulations also help in studying the problem at various levels of abstraction and provide practical feedback. In this paper, we discuss the implementation of DLS techniques in Alea, a Grid Sim based scheduling simulator. Based on the simulation results, we further compare the load balancing characteristics of these methods in a simulated parallel and distributed computing environment. Ioana Banicescu, Florina M. Ciorba, Wolfgang E. Nagel |
ISPDC | 3 |
| 2011 | Distributed dynamic load balancing for pipelined computations on heterogeneous systems
Ioannis Riakiotakis, Florina M. Ciorba, Theodore Andronikos, George K. Papakonstantinou |
Parallel Comput. | 2 |
| 2010 | Studying the impact of synchronization frequency on scheduling tasks with dependencies in heterogeneous systems
Theodore Andronikos, Florina M. Ciorba, Ioannis Riakiotakis, George K. Papakonstantinou, Anthony T. Chronopoulos |
Perform. Evaluation | 2 |
| 2009 | Towards the Robustness of Dynamic Loop Scheduling on Large-Scale Heterogeneous Distributed SystemsabstractDynamic loop scheduling (DLS) algorithms provide application-level load balancing of loop iterates, with the goal of maximizing application performance on the underlying system. These methods use run-time information regarding the performance of the application's execution (for which irregularities change over time). Many DLS methods are based on probabilistic analyses, and therefore account for unpredictable variations of application and system related parameters. Scheduling scientific and engineering applications in large-scale distributed systems (possibly shared with other users) makes the problem of DLS even more challenging. Moreover, the chances of failure, such as processor or link failure, are high in such large-scale systems. In this paper, we employ the hierarchical approach for three DLS methods, and propose metrics for quantifying their robustness with respect to variations of two parameters (load and processor failures), for scheduling irregular applications in large-scale heterogeneous distributed systems. Ioana Banicescu, Florina M. Ciorba, Ricolindo Cariño |
ISPDC | 2 |
| 2008 | Enhancing self-scheduling algorithms via synchronization and weighting
Florina M. Ciorba, Ioannis Riakiotakis, Theodore Andronikos, George K. Papakonstantinou, Anthony T. Chronopoulos |
J. Parallel Distributed Comput. | 1 |
| 2008 | Cronus: A platform for parallel code generation based on computational geometry methods
Theodore Andronikos, Florina M. Ciorba, Panayiotis Theodoropoulos, Dimitris Kamenopoulos, George K. Papakonstantinou |
J. Syst. Softw. | 2 |
| 2007 | Studying the impact of synchronization frequency on scheduling tasks with dependencies in heterogeneous systems
Florina M. Ciorba, Ioannis Riakiotakis, George K. Papakonstantinou, Theodore Andronikos, Anthony T. Chronopoulos |
PACT | 1 |
| 2007 | Optimal synchronization frequency for dynamic pipelined computations on heterogeneous systemsabstractIn this paper we give a theoretical model for determining the synchronization frequency that minimizes the parallel execution time of loops with uniform dependencies dynamically scheduled on heterogeneous systems. Using this model we determine the synchronization frequency that minimizes the estimated parallel time. The accuracy of our method is validated through experiments on a heterogeneous cluster. The results show that the synchronization frequency minimizing the parallel time determined by our method, is very close to the synchronization frequency found experimentally. Florina M. Ciorba, Ioannis Riakiotakis, Theodore Andronikos, Anthony T. Chronopoulos, George K. Papakonstantinou |
CLUSTER | 1 |
| 2006 | Self-Adapting Scheduling for Tasks with Dependencies in Stochastic EnvironmentsabstractThis paper addresses dynamic load balancing algorithms for non-dedicated heterogeneous clusters of workstations. We propose an algorithm called self-adapting scheduling (SAS), targeted at nested loops with dependencies in a stochastic environment. This means that the load entering the system, not belonging to the parallel application under execution, follows an unpredictable pattern which can be modeled by a stochastic process. SAS takes into account the history of previous timing results and the load patterns in order to make accurate load balancing predictions. We study the performance of SAS in comparison with DTSS. We established in previous work that DTSS is the most efficient self-scheduling algorithm for loops with dependencies on heterogeneous clusters. We test our algorithm under the assumption that the interarrival times and life-times of incoming jobs are exponentially distributed. The experimental results show that SAS significantly outperforms DTSS especially with rapidly varying loads Ioannis Riakiotakis, Florina M. Ciorba, Theodore Andronikos, George K. Papakonstantinou |
CLUSTER | 2 |
| 2006 | Dynamic multi phase scheduling for heterogeneous clustersabstractDistributed computing systems are a viable and less expensive alternative to parallel computers. However, concurrent programming methods in distributed systems have not been studied as extensively as for parallel computers. Some of the main research issues are how to deal with scheduling and load balancing of such a system, which may consist of heterogeneous computers. In the past, a variety of dynamic scheduling schemes suitable for parallel loops (with independent iterations) on heterogeneous computer clusters have been obtained and studied. However, no study of dynamic schemes for loops with iteration dependencies has been reported so far. In this work we study the problem of scheduling loops with iteration dependencies for heterogeneous (dedicated and non-dedicated) clusters. The presence of iteration dependencies incurs an extra degree of difficulty and makes the development of such schemes quite a challenge. We extend three well known dynamic schemes (CSS, TSS and DTSS) by introducing synchronization points at certain intervals so that processors compute in pipelined fashion. Our scheme is called dynamic multi-phase scheduling (DMPS) and we apply it to loops with iteration dependencies. We implemented our new scheme on a network of heterogeneous computers and studied its performance. Through extensive testing on two real-life applications (the heat equation and the Floyd-Steinberg algorithm), we show that the proposed method is efficient for parallelizing nested loops with dependencies on heterogeneous systems. Florina M. Ciorba, Theodore Andronikos, Ioannis Riakiotakis, Anthony T. Chronopoulos, George K. Papakonstantinou |
IPDPS | 1 |
| 2005 | Reducing the Communication Cost via Chain Pattern SchedulingabstractThis paper deals with general nested loops and proposes a novel scheduling methodology for reducing the communication cost of parallel programs. General loops contain complex loop bodies (consisting of arbitrary program statements, such as assignments, conditions and repetitions) that exhibit uniform loop-carried dependencies. Therefore it is now possible to achieve efficient parallelization for a vast class of loops, mostly found in DSP, PDEs, signal and video coding. We use computational geometry methods, that exploit efficiently the regularity of nested loops index spaces, in order to significantly reduce the communication cost, which in most cases is the main drawback of parallel programs' performance. Through extensive testing, we show that the proposed method outperforms in all cases the classic cyclic mapping, succeeding to reduce the communication by 15%-35%. This significant reduction of the communication volume makes our method a promising candidate to be incorporated into existing automatic parallel code generation tools Florina M. Ciorba, Theodore Andronikos, Ioannis Drositis, George K. Papakonstantinou, Panayiotis Tsanakas |
NCA | 1 |