VLDB 2026 Research / reviewers in the wild / expert
Michael Gerndt
dblp:g/MichaelGerndt · also Hans Michael Gerndt
· DBLP profile ↗
77ranked-venue papers
19as first author
26since 2021 · last 2026
0000-0002-3210-5048ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 15 first-author · 10 since 2021Software engineering, systems software and programming languages · 10 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ElastiFlow: Elastic Resource Management for Iterative Scientific Workflows in Hybrid HPC-Cloud Infrastructure
Srishti Dasgupta, Kavitha Subramaniam, Michael Gerndt |
CCGrid | 3 |
| 2025 | VersaSlot: Efficient Fine-grained FPGA Sharing with Big.Little Slots and Live Migration in FPGA ClusterabstractAs FPGAs gain popularity for on-demand application acceleration in data center computing, dynamic partial reconfiguration (DPR) has become an effective fine-grained sharing technique for FPGA multiplexing. However, current FPGA sharing encounters partial reconfiguration contention and task execution blocking problems introduced by the DPR, which significantly degrade application performance. In this paper, we propose VersaSlot, an efficient spatio-temporal FPGA sharing system with novel Big.Little slot architecture that can effectively resolve the contention and task blocking while improving resource utilization. For the heterogeneous Big.Little architecture, we introduce an efficient slot allocation and scheduling algorithm, along with a seamless cross-board switching and live migration mechanism, to maximize FPGA multiplexing across the cluster. We evaluate the VersaSlot system on an FPGA cluster composed of the latest Xilinx UltraScale+ FPGAs (ZCU216) and compare its performance against four existing scheduling algorithms. The results demonstrate that VersaSlot achieves up to 13.66x lower average response time than the traditional temporal FPGA multiplexing, and up to $2.19 x$ average response time improvement over the state-of-the-art spatio-temporal sharing systems. Furthermore, VersaSlot enhances the LUT and FF resource utilization by 35% and 29% on average, respectively. Jianfeng Gu 0001, Xiaorang Guo, Martin Schulz 0001, Michael Gerndt |
DAC | 5 |
| 2025 | HAS-GPU: Efficient Hybrid Auto-scaling with Fine-Grained GPU Allocation for SLO-Aware Serverless Inferences
Jianfeng Gu 0001, Puxuan Wang, Isaac David Núñez Araya, Kai Huang 0001, Michael Gerndt |
Euro-Par (1) | 5 |
| 2025 | Deadline Miss Minimization Scheduling for License-Constrained CAE Jobs in Hybrid Cloud Infrastructure
Mohamed Noaman, Srishti Dasgupta, Michael Gerndt |
JSSPP | 3 |
| 2025 | Energy Budget Distribution in Mobile Edge Computing ApplicationsabstractAlthough the performance of mobile devices (MDs) has been steadily improving, advanced applications might suffer long latency for complex functions as well as from draining the battery when continuously used. Function offloading to cloud servers is a proposed technique to solve these challenges. Recently edge computing enables low latency access to resources and, in combination with Function-as-a-service (FaaS), is an obvious target for offloading from MDs. In this paper we introduce a framework for FaaS applications on MDs and function offloading to a FaaS Edge implementation. The scheduling of function invocations to local or remote resources is determined by function-specific offloading policies. The policies are selected periodically according to historic data and function specific performance and energy models. The periodic planning considers the distribution of the application's energy budget across a number of periods and adapts the functionspecific policies for a period. This enables more time consuming planning and fast decision making for individual function invocations. Our results validate the efficiency of the proposed approach, showing about 33 % reduction in energy consumption. Khadija Akherfi, Michael Gerndt |
WiMob | 2 |
| 2024 | Apodotiko: Enabling Efficient Serverless Federated Learning in Heterogeneous EnvironmentsabstractFederated Learning (FL) is an emerging machine learning paradigm that enables the collaborative training of a shared global model across distributed clients while keeping the data decentralized. Recent works on designing systems for efficient FL have shown that utilizing serverless computing technologies, particularly Function-as-a-Service (FaaS) for FL, can enhance resource efficiency, reduce training costs, and alleviate the complex infrastructure management burden on data holders. However, current serverless FL systems still suffer from the presence of stragglers, i.e., slow clients that impede the collaborative training process. While strategies aimed at mitigating stragglers in these systems have been proposed, they overlook the diverse hardware resource configurations among FL clients. To this end, we present Apodotiko, a novel asynchronous training strategy designed for serverless FL. Our strategy incorporates a scoring mechanism that evaluates each client’s hardware capacity and dataset size to intelligently prioritize and select clients for each training round, thereby minimizing the effects of stragglers on system performance. We comprehensively evaluate Apodotiko across diverse datasets, considering a mix of CPU and GPU clients, and compare its performance against five other FL training strategies. Results from our experiments demonstrate that Apodotiko outperforms other FL training strategies, achieving an average speedup of 2.75x and a maximum speedup of 7.03x. Furthermore, our strategy significantly reduces cold starts by a factor of four on average, demonstrating suitability in serverless environments. Mohak Chadha, Alexander Jensen, Jianfeng Gu 0001, Osama Abboud, Michael Gerndt |
CCGrid | 5 |
| 2024 | Economy-based Greedy Bidding for Resources for CAE Workflows in Hybrid Cloud InfrastructureabstractThe advent of generative design in the automotive sector, characterised by the automatic and iterative exploration of expansive solution spaces to discover optimal design configurations, has significantly increased the demand for computational resources to run intensive computer-aided engineering (CAE) simulations within constrained time frames. The inherent limitations of static high-performance computing (HPC) clusters have necessitated the adoption of cloud resources due to their flexible and elastic nature, thereby enhancing the capacity to accommodate the computational demands of these iterative workflows. These workflows, represented as Directed Acyclic Graphs (DAGs), involve the serial and parallel execution of tasks, which can dynamically share resources with other workflows during idle periods. In this paper, we propose an economy-based approach to exploit the gaps generated by these idle periods through a bidding system, thereby enabling more efficient resource utilisation and reducing the average wait time, makespan, cost and deadline miss by more than 40%, 6%, 13% and 45%respectively against certain infrastructures and baselines. Furthermore, we explore the potential for generating revenue by renting out idle resources in a hybrid cloud setup. This approach not only aims to optimise the use of computational resources but also seeks to provide cost-effective solutions to meet the escalating demands of generative design in the automotive sector. Srishti Dasgupta, Tähvend Uustalu, Michael Gerndt, Babak Gholami |
e-Science | 3 |
| 2024 | Leveraging Resource-Aware Application-Level Checkpointing and RDMA for Fault Tolerance and Data Distribution in Malleable MPI ApplicationsabstractDynamic resource management and application malleability present numerous opportunities in High-Performance Computing (HPC), enhancing both system-level services and application performance. Recent trends in malleability research, encompassing both application and system dynamism, are building a new era in HPC. Dynamic applications, particularly those leveraging malleable resources, require adaptive checkpointing systems to enhance performance and resource use. Applications can significantly benefit from checkpointing systems becoming dynamic, especially in handling data redistribution during resource changes. Consequently, checkpointing services should also become malleable (or adaptive). Therefore, we propose iCheck, an adaptive application-level checkpoint management system that caters to malleable MPI applications. iCheck aids these applications by dynamically reconfiguring checkpointing resources and offering robust checkpointing and data redistribution services. By leveraging Remote Direct Memory Access (RDMA) to support malleable applications, iCheck facilitates faster data transfers (up to 40 times improvement over the PFS-based solution), ensuring efficient performance. The system can dynamically adjust checkpointing processes based on metrics such as available memory, checkpoint frequency, and number of processes, maintaining or improving checkpoint performance amid resource changes. This adaptive approach supports fault tolerance as well as simplifies the development of malleable applications by effectively managing resource redistribution during resource changes. Jophin John, Michael Gerndt |
HPCC | 2 |
| 2024 | Elastic Workflows in Hybrid Cloud for CAE SimulationsabstractCompanies traditionally dependent on on-premise HPC clusters for simulations are increasingly migrating workloads to the cloud. Cloud computing offers greater flexibility in selecting processors, memory, network bandwidth, along with enhanced resource availability and scalability. Automotive companies rely on computationally intensive numerical simulation tools for CAE (Computer-Aided Engineering), particularly with the growing demand for generative design, which utilizes algorithms to automatically explore a large solution space. This work addresses the gap between the growing runtime demands of these simulations and the limitations of static HPC infrastructure by representing iterative workflows as Directed Acyclic Graphs (DAGs) and optimizing their scheduling. We propose a unified hybrid infrastructure that leverages the elasticity of cloud resources along with existing HPC clusters to maximize computational efficiency, ensure timely completion of simulations, and optimize resource utilization and costs. Srishti Dasgupta, Michael Gerndt, Babak Gholami |
IC2E | 2 |
| 2024 | gFaaS: Enabling Generic Functions in Serverless ComputingabstractWith the advent of AWS Lambda in 2014, Serverless Computing, particularly Function-as-a-Service (FaaS), has witnessed growing popularity across various application domains. FaaS enables an application to be decomposed into fine-grained functions that are executed on a FaaS platform. It offers several advantages such as no infrastructure management, a pay-per-use billing policy, and on-demand fine-grained autoscaling. However, despite its advantages, developers today encounter various challenges while adopting FaaS solutions that reduce productivity. These include FaaS platform lock-in, support for diverse function deployment parameters, and diverse interfaces for interacting with FaaS platforms. To address these challenges, we present gFaaS, a novel framework that facilitates the holistic development and management of functions across diverse FaaS platforms. Our framework enables the development of generic functions in multiple programming languages that can be seamlessly deployed across different platforms without modifications. Results from our experiments demonstrate that gFaaS functions perform similarly to native platform-specific functions across various scenarios. A video demonstrating the functioning of gFaaS is available from https://youtu.be/STbb6ykJFf0. Mohak Chadha, Paul Wieland, Michael Gerndt |
SANER | 3 |
| 2023 | An Optical Transceiver Reliability Study based on SFP Monitoring and OS-level Metric DataabstractThe increasing demand for cloud computing drives the expansion in scale of datacenters and their internal optical network, in a strive for increasing bandwidth, high reliability, and lower latency. Optical transceivers are essential elements of optical networks, whose reliability has not been well-studied compared to other hardware components. In this paper, we leverage high quantities of monitoring data from optical transceivers and OS-level metrics to provide statistical insights about the occurrence of optical transceiver failures. We estimate transceiver failure rates and normal operating ranges for monitored attributes, correlate early-observable patterns to known failure symptoms, and finally develop failure prediction models based on our analyses. Our results enable network administrators to deploy early-warning systems and enact predictive maintenance strategies, such as replacement or traffic re-routing, reducing the number of incidents and their associated costs. Paolo Notaro, Qiao Yu 0003, Soroush Haeri, Jorge Cardoso 0001, Michael Gerndt |
CCGrid | 5 |
| 2023 | FaST-GShare: Enabling Efficient Spatio-Temporal GPU Sharing in Serverless Computing for Deep Learning InferenceabstractServerless computing (FaaS) has been extensively utilized for deep learning (DL) inference due to the ease of deployment and pay-per-use benefits. However, existing FaaS platforms utilize GPUs in a coarse manner for DL inferences, without taking into account spatio-temporal resource multiplexing and isolation, which results in severe GPU under-utilization, high usage expenses, and SLO (Service Level Objectives) violation. There is an imperative need to enable an efficient and SLO-aware GPU-sharing mechanism in serverless computing to facilitate cost-effective DL inferences. In this paper, we propose FaST-GShare, an efficient FaaS-oriented Spatio-Temporal GPU Sharing architecture for deep learning inferences. In the architecture, we introduce the FaST-Manager to limit and isolate spatio-temporal resources for GPU multiplexing. In order to realize function performance, the automatic and flexible FaST-Profiler is proposed to profile function throughput under various resource allocations. Based on the profiling data and the isolation mechanism, we introduce the FaST-Scheduler with heuristic auto-scaling and efficient resource allocation to guarantee function SLOs. Meanwhile, FaST-Scheduler schedules function with efficient GPU node selection to maximize GPU usage. Furthermore, model sharing is exploited to mitigate memory contention. Our prototype implementation on the OpenFaaS platform and experiments on MLPerf-based benchmark prove that FaST-GShare can ensure resource isolation and function SLOs. Compared to the time sharing mechanism, FaST-GShare can improve throughput by 3.15x, GPU utilization by 1.34x, and SM (Streaming Multiprocessor) occupancy by 3.13x on average. Jianfeng Gu 0001, Puxuan Wang, Mohak Chadha, Michael Gerndt |
ICPP | 5 |
| 2023 | Exploring the Use of WebAssembly in HPCabstractContainerization approaches based on namespaces offered by the Linux kernel have seen an increasing popularity in the HPC community both as a means to isolate applications and as a format to package and distribute them. However, their adoption and usage in HPC systems faces several challenges. These include difficulties in unprivileged running and building of scientific application container images directly on HPC resources, increasing heterogeneity of HPC architectures, and access to specialized networking libraries available only on HPC systems. These challenges of container-based HPC application development closely align with the several advantages that a new universal intermediate binary format called WebAssembly (Wasm) has to offer. These include a lightweight userspace isolation mechanism and portability across operating systems and processor architectures. In this paper, we explore the usage of Wasm as a distribution format for MPI-based HPC applications. To this end, we present MPIWasm, a novel Wasm embedder for MPI-based HPC applications that enables high-performance execution of Wasm code, has low-overhead for MPI calls, and supports high-performance networking interconnects present on HPC systems. We evaluate the performance and overhead of MPIWasm on a production HPC system and AWS Graviton2 nodes using standardized HPC benchmarks. Results from our experiments demonstrate that MPIWasm delivers competitive native application performance across all scenarios. Moreover, we observe that Wasm binaries are 139.5x smaller on average as compared to the statically-linked binaries for the different standardized benchmarks. Mohak Chadha, Nils Krueger, Jophin John, Anshul Jindal, Michael Gerndt, Shajulin Benedict |
PPoPP | 5 |
| 2023 | LogRule: Efficient Structured Log Mining for Root Cause AnalysisabstractAccurate, timely Root Cause Analysis (RCA) is essential to successful IT operations as a primary step to incident remediation. RCA automation using data mining techniques in large heterogeneous systems is, however, a challenging task, because it requires correlating multimodal information across various data sources. An increasing number of services are migrating to structured logging to enable automated monitoring and debugging of complex large-scale systems. In this paper, we leverage structured logs and association rule mining (ARM) to automate RCA. We propose the LogRule algorithm, which automatically analyzes structured logs to generate a list of explanations for an event of interest. It achieves 0.921 F1-score for the diagnosis task, while computing results 37x faster compared to the state-of-the-art solution based on FP-growth, making it a time-efficient, accurate, and interpretable ARM-based RCA algorithm. Evaluation results show that LogRule enables RCA in complex multidimensional datasets, where the execution time of the current state-of-the-art algorithm is prohibitively large. Paolo Notaro, Soroush Haeri, Jorge Cardoso 0001, Michael Gerndt |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2022 | SLAM: SLO-Aware Memory Optimization for Serverless ApplicationsabstractServerless computing paradigm has become more ingrained into the industry, as it offers a cheap alternative for application development and deployment. This new paradigm has also created new kinds of problems for the developer, who needs to tune memory configurations for balancing cost and performance. Many researchers have addressed the issue of minimizing cost and meeting Service Level Objective (SLO) requirements for a single FaaS function, but there has been a gap for solving the same problem for an application consisting of many FaaS functions, creating complex application workflows.In this work, we designed a tool called SLAM to address the issue. SLAM uses distributed tracing to detect the relationship among the FaaS functions within a serverless application. By modeling each of them, it estimates the execution time for the application at different memory configurations. Using these estimations, SLAM determines the optimal memory configuration for the given serverless application based on the specified SLO requirements and user-specified objectives (minimum cost or minimum execution time). We demonstrate the functionality of SLAM on AWS Lambda by testing on four applications. Our results show that the suggested memory configurations guarantee that more than 95% of requests are completed within the predefined SLOs. Gor Safaryan, Anshul Jindal, Mohak Chadha, Michael Gerndt |
CLOUD | 4 |
| 2022 | FedLesScan: Mitigating Stragglers in Serverless Federated LearningabstractFederated Learning (FL) is a machine learning paradigm that enables the training of a shared global model across distributed clients while keeping the training data local. While most prior work on designing systems for FL has focused on using stateful always running components, recent work has shown that components in an FL system can greatly benefit from the usage of serverless computing and Function-as-a-Service technologies. To this end, distributed training of models with severless FL systems can be more resource-efficient and cheaper than conventional FL systems. However, serverless FL systems still suffer from the presence of stragglers, i.e., slow clients due to their resource and statistical heterogeneity. While several strategies have been proposed for mitigating stragglers in FL, most methodologies do not account for the particular characteristics of serverless environments, i.e., cold-starts, performance variations, and the ephemeral stateless nature of the function instances. Towards this, we propose FedLesScan, a novel clustering-based semi-asynchronous training strategy, specifically tailored for serverless F L. FedLesScan dynamically adapts to the behavior of clients and minimizes the effect of stragglers on the overall system. We implement our strategy by extending an open-source serverless FL system called FedLess. Moreover, we comprehensively evaluate our strategy using the 2ndgeneration Google Cloud Functions with four datasets and varying percentages of stragglers. Results from our experiments show that compared to other approaches FedLesScan reduces training time and cost by an average of 8% and 20% respectively while utilizing clients better with an average increase in the effective update ratio of 17.75%. Mohamed Elzohairy, Mohak Chadha, Anshul Jindal, Andreas Grafberger, Jianfeng Gu 0001, Michael Gerndt, Osama Abboud |
IEEE Big Data | 6 |
| 2022 | Scalable Infrastructure for Workload Characterization of Cluster TracesabstractIn the recent past, characterizing workloads has been attempted to gain a foothold in the emerging serverless cloud market, especially in the large production cloud clusters of Google, AWS, and so forth. While analyzing and characterizing real workloads from a large production cloud cluster benefits cloud providers, researchers, and daily users, analyzing the workload traces of these clusters has been an arduous task due to the heterogeneous nature of data. This article proposes a scalable infrastructure based on Google's dataproc for analyzing the workload traces of cloud environments. We evaluated the functioning of the proposed infrastructure using the workload traces of Google cloud cluster-usage-traces-v3. We perform the workload characterization on this dataset, focusing on the heterogeneity of the workload, the variations in job durations, aspects of resources consumption, and the overall availability of resources provided by the cluster. The findings reported in the paper will be beneficial for cloud infrastructure providers and users while managing the cloud computing resources, especially serverless platforms. Thomas van Loo, Anshul Jindal, Shajulin Benedict, Mohak Chadha, Michael Gerndt |
CLOSER | 5 |
| 2022 | FaDO: FaaS Functions and Data Orchestrator for Multiple Serverless Edge-Cloud ClustersabstractFunction-as-a-Service (FaaS) is an attractive cloud computing model that simplifies application development and deployment. However, current serverless compute platforms do not consider data placement when scheduling functions. With the growing demand for edge-cloud continuum, multi-cloud, and multi-serverless applications, this flaw means serverless technologies are still ill-suited to latency-sensitive operations like media streaming. This work proposes a solution by presenting a tool called FaDO: FaaS Functions and Data Orchestrator, designed to allow data-aware functions scheduling across multi-serverless compute clusters present at different locations, such as at the edge and in the cloud. FaDO works through header-based HTTP reverse proxying and uses three load-balancing algorithms: 1) The Least Connections, 2) Round Robin, and 3) Random for load balancing the invocations of the function across the suitable serverless compute clusters based on the set storage policies. FaDO further provides users with an abstraction of the serverless compute cluster’s storage, allowing users to interact with data across different storage services through a unified interface. In addition, users can configure automatic and policy-aware granular data replications, causing FaDO to spread data across the clusters while respecting location constraints. Load testing results show that it is capable of load balancing high-throughput workloads, placing functions near their data without contributing any significant performance overhead. Christopher Peter Smith, Anshul Jindal, Mohak Chadha, Michael Gerndt, Shajulin Benedict |
ICFEC | 4 |
| 2022 | iCheck: Leveraging RDMA and Malleability for Application-Level Checkpointing in HPC SystemsabstractThe estimate that the mean time between failures will be in minutes in exascale supercomputers should be alarming for application developers. The inherent system’s complexity, millions of components, and susceptibility to failures make checkpointing more relevant than ever. Since most high performance scientific applications contain an in-house checkpoint restart mechanism, their performance can be impacted by the contention of parallel file system resources. A shift in checkpointing strategies is needed to thwart this behavior. With iCheck, we present a novel checkpointing framework that supports malleable multilevel application-level checkpointing. We employ an RDMA enabled configurable multi-agent-based checkpoint transfer mechanism where minimal application resources are utilized for checkpointing. The high-level API of iCheck facilitates easy integration and malleability. We have added the iCheck library into the Is1 mardyn application providing performance improvement up to five thousand times over the in-house checkpointing mechanism. LULESH, Jacobi 2D heat simulation, and a synthetic application were also used for extensive analysis. Jophin John, Isaac David Núñez Araya, Michael Gerndt |
ICPADS | 3 |
| 2022 | Bunk8s: Enabling Easy Integration Testing of Microservices in KubernetesabstractMicroservice architecture is the common choice for cloud applications these days since each individual microservice can be independently modified, replaced, and scaled. However, the complexity of microservice applications requires automated testing with a focus on the interactions between the services. While this is achievable with end-to-end tests, they are error-prone, brittle, expensive to write, time-consuming to run, and require the entire application to be deployed. Integration tests are an alternative to end-to-end tests since they have a smaller test scope and require the deployment of a significantly fewer number of services. The de-facto standard for deploying microservice applications in the cloud is containers with Kubernetes being the most widely used container orchestration platform. To support the integration testing of microservices in Kubernetes, several tools such as Octopus, Istio, and Jenkins exist. However, each of these tools either lack crucial functionality or lead to a substantial increase in the complexity and growth of the tool landscape when introduced into a project. To this end, we present Bunk8s, a tool for integration testing of microservice applications in Kubernetes that overcomes the limitations of these existing tools. Bunk8s is independent of the test framework used for writing integration tests, independent of the used CI/CD infrastructure, and supports test result publishing. A video demonstrating the functioning of our tool is available from https://www.youtube.com/watch?v=e8wbS25O4Bo. Christoph Reile, Mohak Chadha, Valentin Hauner, Anshul Jindal, Benjamin Hofmann, Michael Gerndt |
SANER | 6 |
| 2021 | Architecture-Specific Performance Optimization of Compute-Intensive FaaS FunctionsabstractFaaS allows an application to be decomposed into functions that are executed on a FaaS platform. The FaaS platform is responsible for the resource provisioning of the functions. Recently, there is a growing trend towards the execution of compute-intensive FaaS functions that run for several seconds. However, due to the billing policies followed by commercial FaaS offerings, the execution of these functions can incur significantly higher costs. Moreover, due to the abstraction of underlying processor architectures on which the functions are executed, the performance optimization of these functions is challenging. As a result, most FaaS functions use pre-compiled libraries generic to x86-64 leading to performance degradation. In this paper, we examine the underlying processor architectures for Google Cloud Functions (GCF) and determine their prevalence across the 19 available GCF regions. We modify, adapt, and optimize three compute-intensive FaaS workloads written in Python using Numba, a JIT compiler based on LLVM, and present results wrt performance, memory consumption, and costs on GCF. Results from our experiments show that the optimization of FaaS functions can improve performance by 12.8x (geometric mean) and save costs by 73.4% on average for the three functions. Our results show that optimization of the FaaS functions for the specific architecture is very important. We achieved a maximum speedup of 1.79x by tuning the function especially for the instruction set of the underlying processor architecture. Mohak Chadha, Anshul Jindal, Michael Gerndt |
CLOUD | 3 |
| 2021 | FedLess: Secure and Scalable Federated Learning Using Serverless ComputingabstractThe traditional cloud-centric approach for Deep Learning (DL) requires training data to be collected and processed at a central server which is often challenging in privacy-sensitive domains like healthcare. Towards this, a new learning paradigm called Federated Learning (FL) has been proposed that brings the potential of DL to these domains while addressing privacy and data ownership issues. FL enables clients to learn a shared ML model while keeping the data local. However, conventional FL systems face challenges such as scalability, complex infrastructure management, and wasted compute and incurred costs due to idle clients. These challenges of FL systems closely align with the core problems that serverless computing and Function-as-a-Service (FaaS) platforms aim to solve. These include rapid scalability, no infrastructure management, automatic scaling to zero for idle clients, and a pay-per-use billing model. To this end, we present a novel system and framework for serverless FL, called FedLess. Our system supports multiple commercial and self-hosted FaaS providers and can be deployed in the cloud, on-premise in institutional data centers, and on edge devices. To the best of our knowledge, we are the first to enable FL across a large fabric of heterogeneous FaaS providers while providing important features like security and Differential Privacy. We demonstrate with comprehensive experiments that the successful training of DNNs for different tasks across up to 200 client functions and more is easily possible using our system. Furthermore, we demonstrate the practical viability of our methodology by comparing it against a traditional FL system and show that it can be cheaper and more resource-efficient. Andreas Grafberger, Mohak Chadha, Anshul Jindal, Jianfeng Gu 0001, Michael Gerndt |
IEEE BigData | 5 |
| 2021 | DeepEdgeBench: Benchmarking Deep Neural Networks on Edge DevicesabstractEdgeAI (Edge computing based Artificial Intelligence) has been most actively researched for the last few years to handle variety of massively distributed AI applications to meet up the strict latency requirements. Meanwhile, many companies have released edge devices with smaller form factors (low power consumption and limited resources) like the popular Raspberry Pi and Nvidia's Jetson Nano for acting as compute nodes at the edge computing environments. Although the edge devices are limited in terms of computing power and hardware resources, they are powered by accelerators to enhance their performance behavior. Therefore, it is interesting to see how AI-based Deep Neural Networks perform on such devices with limited resources. In this work, we present and compare the performance in terms of inference time and power consumption of the four SoCs: Asus Tinker Edge R, Raspberry Pi 4, Google Coral Dev Board, Nvidia Jetson Nano, and one microcontroller: Arduino Nano 33 BLE, on different deep learning models and frameworks. We also provide a method for measuring power consumption, inference time and accuracy for the devices, which can be easily extended to other devices. Our results showcase that, for Tensorflow based quantized model, the Google Coral Dev Board delivers the best performance, both for inference time and power consumption. For a low fraction of inference computation time, i.e. less than 29.3% of the time for MobileNetV2, the Jetson Nano performs faster than the other devices. Stephan Patrick Baller, Anshul Jindal, Mohak Chadha, Michael Gerndt |
IC2E | 4 |
| 2021 | Poster: Function Delivery Network: Extending Serverless to Heterogeneous ComputingabstractSeveral of today's cloud applications are spread over heterogeneous connected computing resources and are highly dynamic in their structure and resource requirements. However, serverless computing and Function-as-a-Service (FaaS) platforms are limited to homogeneous clusters and homogeneous functions. We introduce an extension of FaaS to heterogeneous computing and to support heterogeneous functions through a network of distributed heterogeneous target platforms called Function Delivery Network (FDN). A target platform is a combination of a cluster of a homogeneous computing system and a FaaS platform on top of it. FDN provides Function-Delivery-as-a-Service (FDaaS), delivering the function invocations to the right target platform. We showcase the opportunities such as collaborative execution between multiple target platforms and varied target platform's characteristics that the FDN offers in fulfilling two objectives: Service Level Objective (SLO) requirements and energy efficiency when scheduling functions invocations by evaluating over five distributed target platforms. Anshul Jindal, Mohak Chadha, Michael Gerndt, Julian Frielinghaus, Vladimir Podolskiy, Pengfei Chen 0002 |
ICDCS | 3 |
| 2021 | Function delivery network: Extending serverless computing for heterogeneous platformsabstractSummary Serverless computing has rapidly grown following the launch of Amazon's Lambda platform. Function‐as‐a‐Service (FaaS) a key enabler of serverless computing allows an application to be decomposed into simple, standalone functions that are executed on a FaaS platform. The FaaS platform is responsible for deploying and facilitating resources to the functions. Several of today's cloud applications spread over heterogeneous connected computing resources and are highly dynamic in their structure and resource requirements. However, FaaS platforms are limited to homogeneous clusters and homogeneous functions and do not account for the data access behavior of functions before scheduling. We introduce an extension of FaaS to heterogeneous clusters and to support heterogeneous functions through a network of distributed heterogeneous target platforms called Function Delivery Network (FDN). A target platform is a combination of a cluster of homogeneous nodes and a FaaS platform on top of it. FDN provides Function‐Delivery‐as‐a‐Service (FDaaS), delivering the function to the right target platform. We showcase the opportunities such as varied target platform's characteristics, possibility of collaborative execution between multiple target platforms, and localization of data that the FDN offers in fulfilling two objectives: Service Level Objective (SLO) requirements and energy efficiency when scheduling functions by evaluating over five distributed target platforms using the FDNInspector, a tool developed by us for benchmarking distributed target platforms. Scheduling functions on an edge target platform in our evaluation reduced the overall energy consumption by 17× without violating the SLO requirements in comparison to scheduling on a high‐end target platform. Anshul Jindal, Michael Gerndt, Mohak Chadha, Vladimir Podolskiy, Pengfei Chen 0002 |
Softw. Pract. Exp. | 2 |
| 2021 | A Survey of AIOps Methods for Failure ManagementabstractModern society is increasingly moving toward complex and distributed computing systems. The increase in scale and complexity of these systems challenges O&M teams that perform daily monitoring and repair operations, in contrast with the increasing demand for reliability and scalability of modern applications. For this reason, the study of automated and intelligent monitoring systems has recently sparked much interest across applied IT industry and academia. Artificial Intelligence for IT Operations (AIOps) has been proposed to tackle modern IT administration challenges thanks to Machine Learning, AI, and Big Data. However, AIOps as a research topic is still largely unstructured and unexplored, due to missing conventions in categorizing contributions for their data requirements, target goals, and components. In this work, we focus on AIOps for Failure Management (FM), characterizing and describing 5 different categories and 14 subcategories of contributions, based on their time intervention window and the target problem being solved. We review 100 FM solutions, focusing on applicability requirements and the quantitative results achieved, to facilitate an effective application of AIOps solutions. Finally, we discuss current development problems in the areas covered by AIOps and delineate possible future trends for AI-based failure management. Paolo Notaro, Jorge Cardoso 0001, Michael Gerndt |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2020 | Performance Evaluation of Container Runtimes
Lennart Espe, Anshul Jindal, Vladimir Podolskiy, Michael Gerndt |
CLOSER | 4 |
| 2020 | Microservices vs Serverless: A Performance Comparison on a Cloud-native Web ApplicationabstractA microservices architecture has gained higher popularity among enterprises due to its agility, scalability, and resiliency. However, serverless computing has become a new trendy topic when designing cloud-native applications. Compared to the monolithic and microservices, serverless architecture offloads management and server configuration from the user to the cloud provider and let the user focus only on the product development. Hence, there are debates regarding which deployment strategy to use. This research provides a performance comparison of a cloud-native web application in terms of scalability, reliability, cost, and latency when deployed using microservices and serverless deployment strategy. This research shows that neither the microservices nor serverless deployment strategy fits all the scenarios. The experimental results demonstrate that each type of deployment strategy has its advantages under different scenarios. The microservice deployment strategy has a cost advantag Chen-Fu Fan, Anshul Jindal, Michael Gerndt |
CLOSER | 3 |
| 2020 | Toward an End-to-End Auto-tuning Framework in HPC PowerStackabstractEfficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software stack to manage power and energy of high-performance computing systems and standardizes the interfaces between different components of the stack. This survey paper presents the findings of a working group focused on the end-to-end tuning of the PowerStack. First, we provide a background on the PowerStack layer-specific tuning efforts in terms of their high-level objectives, the constraints and optimization goals, layer-specific telemetry, and control parameters, and we list the existing software solutions that address those challenges. Second, we propose the PowerStack end-to-end auto-tuning framework, identify the opportunities in co-tuning different layers in the PowerStack, and present specific use cases and solutions. Third, we discuss the research opportunities and challenges for collective auto-tuning of two or more management layers (or domains) in the PowerStack. This paper takes the first steps in identifying and aggregating the important R&D challenges in streamlining the optimization efforts across the layers of the PowerStack. Xingfu Wu, Aniruddha Marathe, Siddhartha Jana, Ondrej Vysocky, Jophin John, Andrea Bartolini, Lubomir Riha, Michael Gerndt, Valerie Taylor 0001, Sridutt Bhalachandra |
CLUSTER | 8 |
| 2020 | Extending SLURM for Dynamic Resource-Aware Adaptive Batch SchedulingabstractWith the growing constraints on power budget and increasing hardware failure rates, the operation of future exascale systems faces several challenges. Towards this, resource awareness and adaptivity by enabling malleable jobs has been actively researched in the HPC community. Malleable jobs can change their computing resources at runtime and can significantly improve HPC system performance. However, due to the rigid nature of popular parallel programming paradigms such as MPI and lack of support for dynamic resource management in batch systems, malleable jobs have been largely unrealized. In this paper, we extend the SLURM batch system to support the execution and batch scheduling of malleable jobs. The malleable applications are written using a new adaptive parallel paradigm called Invasive MPI which extends the MPI standard to support resource-adaptivity at runtime. We propose two malleable job scheduling strategies to support performance-aware and power-aware dynamic reconfiguration decisions at runtime. We implement the strategies in SLURM and evaluate them on a production HPC system. Results for our performance-aware scheduling strategy show improvements in makespan, average system utilization, average response, and waiting times as compared to other scheduling strategies. Moreover, we demonstrate dynamic power corridor management using our power-aware strategy. Mohak Chadha, Jophin John, Michael Gerndt |
HiPC | 3 |
| 2020 | Windsurfing with APPA: Automating Computational Fluid Dynamics Simulations of Wind Flow using Cloud ComputingabstractComputational fluid dynamics (CFD) can serve as a complementary approach to conventional wind tunnel testing to assess the wind flow around tall buildings. Being a clear High Performance Computing (HPC) task, CFD simulations conventionally run on supercomputers and compute clusters using specialized software such as OpenFOAM. The limited availability and high maintenance costs of supercomputers and clusters force small and medium companies to search for the cost-efficient infrastructure to conduct their simulations with the appropriate performance. The on-demand offer of compute capacity by cloud service providers are well suited this task. However, engineers and researchers require extensive expertise and experience in working with cloud computing in order to benefit from running CFD simulations on a cloud.The contribution of the paper to the outlined problem is two-fold: 1) a unique Automated Parallel Processing Application (APPA) tool that hides the cloud management details from the wind engineer and provides an intuitive user interface; 2) the estimation of the optimal number of cores (vCPUs) for virtual machine instances provided by AWS and Google Cloud based on average run time and total cost metrics for a given number of cells of a CFD-simulation. n1-highcpu-96 Google Cloud VM met both goals: low cost and low runtime per timestep. For the number of vCPUs below 16, the c4.8xlarge AWS VM type has the least runtime per timestep in all the cases. Google Cloud instances with high vCPUs are recommended to run the simulations if budget is a big concern. Anshul Jindal, Benedikt Strahm, Vladimir Podolskiy, Michael Gerndt |
PDP | 4 |
| 2019 | Modelling DVFS and UFS for Region-Based Energy Aware Tuning of HPC ApplicationsabstractEnergy efliciency and energy conservation are one of the most crucial constraints for meeting the 20MW power envelope desired for exascale systems. Towards this, most of the research in this area has been focused on the utilization of user-controllable hardware switches such as per-core dynamic voltage frequency scaling (DVFS) and software controlled clock modulation at the application level. In this paper, we present a tuning plugin for the Periscope Tuning Framework which integrates line-grained autotuning at the region level with DVFS and uncore frequency scaling (UFS). The tuning is based on a feed-forward neural network which is formulated using Performance Monitoring Counters (PMC) supported by x86 systems and trained using standardized benchmarks. Experiments on live standardized hybrid benchmarks show an energy improvement of 16.1% on average when the applications are tuned according to our methodology as compared to 7.8% for static tuning. Mohak Chadha, Michael Gerndt |
IPDPS | 2 |
| 2019 | Performance Modeling for Cloud Microservice ApplicationsabstractMicroservices enable a fine-grained control over the cloud applications that they constitute and thus became widely-used in the industry. Each microservice implements its own functionality and communicates with other microservices through language- and platform-agnostic API. The resources usage of microservices varies depending on the implemented functionality and the workload. Continuously increasing load or a sudden load spike may yield a violation of a service level objective (SLO). To characterize the behavior of a microservice application which is appropriate for the user, we define a MicroService Capacity (MSC) as a maximal rate of requests that can be served without violating SLO. Anshul Jindal, Vladimir Podolskiy, Michael Gerndt |
ICPE | 3 |
| 2019 | Domain knowledge specification for energy tuningabstractSummary To overcome the challenges of energy consumption of HPC systems, the European Union Horizon 2020 READEX (Runtime Exploitation of Application Dynamism for Energy‐efficient Exascale computing) project uses an online auto‐tuning approach to improve energy efficiency of HPC applications. The READEX methodology pre‐computes optimal system configurations at design‐time, such as the CPU frequency, for instances of program regions and switches at runtime to the configuration given in the tuning model when the region is executed. READEX goes beyond previous approaches by exploiting dynamic changes of a region's characteristics by leveraging region and characteristic specific system configurations. While the tool suite supports an automatic approach, specifying domain knowledge such as the structure and characteristics of the application and application tuning parameters can significantly help to create a more refined tuning model. This paper presents the means available for an application expert to provide domain knowledge and presents tuning results for some benchmarks. Madhura Kumaraswamy, Anamika Chowdhury, Michael Gerndt, Zakaria Bendifallah, Othman Bouizi, Uldis Locans, Lubomir Riha, Ondrej Vysocky, Martin Beseda, Jan Zapletal |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | IaaS Reactive Autoscaling Performance ChallengesabstractThe main feature of a cloud application is its scalability. Major IaaS cloud services providers (CSP) employ autoscaling on the level of virtual machines (VM). Other virtualization solutions (e.g. containers, pods) can also scale. An application scales in response to change in observed metrics, e.g. in CPU utilization. Occasionally, cloud applications exhibit the inability to meet the Quality of Service (QoS) requirements during the scaling caused by the reactivity of autoscaling solutions. This paper provides the results of the autoscaling performance evaluation for two-layered virtualization (VMs and Kubernetes pods) conducted in the public clouds of AWS, Microsoft and Google using the approach and the Autoscaling Performance Measurement Tool developed by the authors. Vladimir Podolskiy, Anshul Jindal, Michael Gerndt |
IEEE CLOUD | 3 |
| 2018 | Practical Education in IoT through Collaborative Work on Open-Source Projects with Industry and Entrepreneurial OrganizationsabstractThis Innovative Practice Full Paper is devoted to the practices in the software engineering education that involve industrial and entrepreneurial partners. The particular focus of the paper is on the preparation, conduction and evaluation of the master practical course in the Internet of Things (IoT) open-source middleware. The motivation for the course is to train software engineers and architects that are able to solve the practical IoT challenges. The course was prepared and conducted during the winter term 2017-2018 at the Technical University of Munich. This course was created as a joint effort of the Chair for Computer Architecture and Parallel Systems, the manufacturing workshop UnternehmerTUM MakerSpace, and the entrepreneurship center UnternehmerTUM associated with Technical University of Munich. 14 participants of the course from different countries have developed an open-source IoT middleware, namely IoT Platform, that enables the retrieval of the data from sensors, distributed data storing and streaming, and platform services automatic detection. The middleware also incorporates well-documented application program interface (API). Performance monitoring facilities prepared by the team of students enable the Platform testing under different load. The IoT Platform was developed based on the 3D-printer management use-case from UnternehmerTUM MakerSpace and was presented at the UnternehmerTUM event Tech Challenge in a pitch format. Main contributions of the paper are in three areas: the methodology and technique to organize, conduct and evaluate the practical course in IoT that involves industrial and entrepreneurial partners; the pointers at particular technologies, solutions, and interconnections thereof for the development of new practically-oriented IoT courses; open-source IoT middleware that could be used for teaching and research purposes. Vladimir Podolskiy, Yesika M. Ramírez, Atakan Yenel, Shumail Mohyuddin, Hakan Uyumaz, Ali Naci Uysal, Moawiah Assali, Sergei Drugalev, Michael Gerndt, Matthias Friessnig, Anastasia Myasnichenko |
FIE | 9 |
| 2018 | A multi-aspect online tuning framework for HPC applications
Michael Gerndt, Siegfried Benkner, Eduardo César, Carmen B. Navarrete, Enes Bajrovic, Jirí Dokulil, Carla Guillén, Robert Mijakovic, Anna Sikora |
Softw. Qual. J. | 1 |
| 2017 | READEX: Linking two ends of the computing continuum to improve energy-efficiency in dynamic applicationsabstractIn both the embedded systems and High Performance Computing domains, energy-efficiency has become one of the main design criteria. Efficiently utilizing the resources provided in computing systems ranging from embedded systems to current petascale and future Exascale HPC systems will be a challenging task. Suboptimal designs can potentially cause large amounts of underutilized resources and wasted energy. In both domains, a promising potential for improving efficiency of scalable applications stems from the significant degree of dynamic behaviour, e.g., runtime alternation in application resource requirements and workloads. Manually detecting and leveraging this dynamism to improve performance and energy-efficiency is a tedious task that is commonly neglected by developers. However, using an automatic optimization approach, application dynamism can be analysed at design time and used to optimize system configurations at runtime. The European Union Horizon 2020 READEX (Runtime Exploitation of Application Dynamism for Energy-efficient eX-ascale computing) project will develop a tools-aided auto-tuning methodology inspired by the system scenario methodology used in embedded systems. Dynamic behaviour of HPC applications will be exploited to achieve improved energy-efficiency and performance. Driven by a consortium of European experts from academia, HPC resource providers, and industry, the READEX project aims at developing the first of its kind generic framework to split design time and runtime automatic tuning while targeting heterogeneous system at the Exascale level. This paper describes plans for the project as well as early results achieved during its first year. Furthermore, it is shown how project results will be brought back into the embedded systems domain. Per Gunnar Kjeldsberg, Andreas Gocht, Michael Gerndt, Lubomir Riha, Joseph Schuchart, Umbreen Sabir Mian |
DATE | 3 |
| 2016 | Infrastructure and API Extensions for Elastic Execution of MPI ApplicationsabstractDynamic Processes support was added to MPI in version 2.0 of the standard. This feature of MPI has not been widely used by application developers in part due to the performance cost and limitations of the spawn operation. In this paper, we propose an extension to MPI that consists of four new operations. These operations allow an application to be initialized in an elastic mode of execution and enter an adaptation window when necessary, where resources are incorporated into or released from the application's world communicator. A prototype solution based on the MPICH library and the SLURM resource manager is presented and evaluated alongside an elastic scientific application that makes use of the new MPI extensions. The cost of these new operations is shown to be negligible due mainly to the latency hiding design, leaving the application's time for data redistribution as the only significant performance cost. Isaías A. Comprés Ureña, Ao Mo-Hellenbrand, Michael Gerndt, Hans-Joachim Bungartz |
EuroMPI | 3 |
| 2016 | A Visualization Framework for ParallelizationabstractSince the advent of multicore processors, developers struggle with the parallelization of legacy software. Automatic methods are only appropriate to identify parallelism at instruction level or within simple loops. For most applications, however, a scalable redesign require profound comprehension of the underlying software architecture and its dynamic aspects. This leads to an increasing demand for interactive tools that foster parallelization at various granularity levels. To cope with this problem, we propose a visualization framework, and three tailored views for parallelism detection. The framework is part of Parceive, a tool that utilizes dynamic binary instrumentation to trace C/C++ and C# programs. The cooperative views allow identification and analysis of potential parallelism scenarios using seamless navigation, abstraction, and filtering. In this paper, we motivate our approach, illustrate the architecture of the visualization framework, and highlight the key features of the views. A case study demonstrates the usefulness of Parceive. Andreas Wilhelm, Victor Savu, Efe Amadasun, Michael Gerndt, Tobias Schüle |
VISSOFT | 4 |
| 2016 | Model-based MPI-IO tuning with Periscope tuning frameworkabstractSummary For many parallel applications, I/O performance is a major bottleneck. MPI‐IO, defined by the MPI forum, can help parallel applications overcome the performance and portability limitations of existing parallel I/O interfaces. Although autotuning has been used to improve the performance of computing kernels, MPI‐IO autotuning has rarely been studied. To automate MPI‐IO performance tuning, we designed and implemented an automatic tuner. The tuner relies on the Periscope tuning framework for transparently passing hints to the MPI‐IO library and for automatically collecting performance data. Unlike computational code, each MPI‐IO function takes a relatively long time to complete. Thus, exhaustively searching through the entire parameter space is impractical. So we developed a performance model that can direct us to shorten the tuning time. Copyright © 2015 John Wiley & Sons, Ltd. Michael Gerndt |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Topic 2: Performance Prediction and Evaluation - (Introduction)
Adolfy Hoisie, Michael Gerndt, Shajulin Benedict, Thomas Fahringer, Vladimir Getov, Scott Pakin |
Euro-Par | 2 |
| 2012 | An integrated simulation framework for invasive computing
Michael Gerndt, Frank Hannig, Andreas Herkersdorf, Andreas Hollmann, Marcel Meyer, Sascha Roloff, Josef Weidendorfer, Thomas Wild, Aurang Zaib |
FDL | 1 |
| 2012 | Invasive computing with iOMP
Michael Gerndt, Andreas Hollmann, Marcel Meyer, Martin Schreiber 0001, Josef Weidendorfer |
FDL | 1 |
| 2012 | Hands-on Practical Hybrid Parallel Application Performance Engineering
Markus Geimer, Michael Gerndt, Sameer Shende, Bert Wesarg, Brian J. N. Wylie |
EuroMPI | 2 |
| 2012 | Wait-Free Message Passing Protocol for Non-coherent Shared Memory Architectures
Isaías A. Comprés Ureña, Michael Gerndt, Carsten Trinitis |
EuroMPI | 2 |
| 2010 | Distributed Systems and Algorithms
Omer F. Rana, Giandomenico Spezzano, Michael Gerndt, Daniel S. Katz |
Euro-Par (1) | 3 |
| 2010 | Special Issue: Scalable Tools for High-end ComputingabstractCurrent high-end parallel systems consist of hundreds of thousands of compute cores arranged in a complex hierarchical structure; future systems will have millions of cores. Systems, such as the Altix 4700, Blue Gene, Roadrunner, and Cray XT5, deploy multiple compute cores (homogeneous or heterogeneous) with multiple levels of shared and private caches within a processor, clustered into SMP nodes and coupled via a communication network to large-scale distributed systems. The development of efficient programs is extremely complex since the architectural details are exposed to the programmer. Productive use of such machines requires highly scalable programming tools for debugging, performance analysis, and fault tolerance. In addition, new programming models might significantly facilitate the task of the programmer. This special issue of Concurrency and Computation: Practice and Experience is devoted to programming tools that facilitate the development of efficient programs for such large-scale architectures. It is a collection of the best papers submitted to the international workshop on Scalable Tools for High-end Computing (STHEC 2008) that was held in conjunction with the International Conference on Supercomputing on June 7th on the Greek Island Kos. The papers present state-of-the-art tools for performance analysis and checkpointing on those machines. Performance analysis tools use measurements gathered during the execution of the application to detect portions of the code that can be further improved. Thus, they have to be able to cope with the large number of processors. Tools for checkpointing provide the possibility to restart an application in the case of a system failure; they have to be able to handle large number of cores as well. The selected papers present different techniques for building tools that will scale to thousands of cores. HPCToolkit 1 is a profiling-based performance analysis environment presenting the data in close relation to the source code without requiring an instrumentation of the source code. Scalasca 2 performs a parallel replay of the execution on the application's processors to find performance bottlenecks automatically. The combination of TAU and MRNet 3 provides a scalable infrastructure to offload performance data. Establishing the overlay network requires no added support from the job manager or application. Periscope 4 is based on a network of analysis agents that performs an online analysis of the application's performance behavior. When the application is started, additional processors can be allocated for the analysis agents to scale the analysis. CPPC 5 is a tool for portable checkpointing of message-passing applications. It consists of a runtime library and a compiler that assists the user by performing time-consuming tasks, such as data flow and communications analyses as well as code instrumentation. We would like to thank the authors for their excellent contributions to this special issue. We hope that it inspires future research in tools that support programmers of high-end systems in the development of efficient programs. Michael Gerndt, Barton P. Miller |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Automatic performance analysis with periscopeabstractAbstract Performance analysis is essential to fully exploit the potential of high‐performance computers. With the imminence of petascale systems which will consist of ten thousands or even hundred thousands of processor cores, this task will increase in complexity. Hence, tools are required that automatically detect the performance bottlenecks and thus ease the performance analysis of an application. On large‐scale systems, collecting information about performance‐relevant events of an application can easily produce a huge amount of data whose analysis is very challenging. Aggregating the performance data during runtime and conducting the search for performance properties online allows users to distill essential performance bottlenecks without overwhelming the user with an uncontrollable load of data. In this paper we present the recent developments on Periscope, a highly scalable tool for the automatic distributed online search for the performance properties of large‐scale applications on high‐end computers. It allows for both detection of the performance bottlenecks limiting the scalability on parallel systems as well as pinpointing the issues concerning the single‐node performance of an application. Copyright © 2009 John Wiley & Sons, Ltd. Michael Gerndt, Michael Ott 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | Topic 2: Performance Prediction and Evaluation
Francisco Almeida, Michael Gerndt, Adolfy Hoisie, Martin Schulz 0001 |
Euro-Par | 2 |
| 2007 | On Using Incremental Profiling for the Performance Analysis of Shared Memory Parallel Applications
Karl Fürlinger, Michael Gerndt, Jack J. Dongarra |
Euro-Par | 2 |
| 2007 | Search Strategies for Automatic Performance Analysis Tools
Michael Gerndt, Edmond Kereku |
Euro-Par | 1 |
| 2007 | Specification and detection of performance problems with ASLabstractAbstract Performance analysis is an important step in tuning performance‐critical applications. It is a cyclic process of measuring and analyzing performance data, driven by the programmer's hypotheses on potential performance problems. Currently this process is controlled manually by the programmer. The goal of the work described in this article is to automate the performance analysis process based on a formal specification of performance properties. One result of the APART project is the APART Specification Language (ASL) for the formal specification of performance properties. Performance bottlenecks can then be identified based on the specification, since bottlenecks are viewed as performance properties with a large negative impact. We also present the overall design and an initial evaluation of the Periscope system which utilizes ASL specifications to automatically search for performance bottlenecks in a distributed manner. Copyright © 2006 John Wiley & Sons, Ltd. Michael Gerndt, Karl Fürlinger |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Special Issue: European-American Working Group on Automatic Performance Analysis (APART)abstractWorking Group on Automatic Performance Analysis (APART)This special issue is devoted to research undertaken by the European-American Working Group on Automatic Performance Analysis (APART).APART (http://www.fz-juelich.de/apart)was established in 1999 as an EU-funded working group with more than 15 partners from Europe and the U.S.A., plus more than 10 associate partners from academia Michael Gerndt, John R. Gurd |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | A test suite for parallel performance analysis toolsabstractAbstract Parallel performance analysis tools must be tested as to whether they perform their task correctly, which comprises at least three aspects. First, it must be ensured that the tools neither alter the semantics nor distort the run‐time behavior of the application under investigation. Next, it must be verified that the tools collect the correct performance data as required by their specification. Finally, it must be checked that the tools perform their intended tasks and detect relevant performance problems. Focusing on the latter (correctness) aspect, testing can be done using synthetic test functions with controllable performance properties, possibly complemented by real‐world applications with known performance behavior. A systematic test suite can be built from synthetic test functions and other components, possibly with the help of tools to assist the user in putting the pieces together into executable test programs. Clearly, such a test suite can be highly useful to builders of performance analysis tools. It is surprising that, up until now, no systematic effort has been undertaken to provide such a suite. In this paper we describe the APART Test Suite (ATS) for checking the correctness (in the above sense) of parallel performance analysis tools. In particular, we describe a collection of synthetic test functions which allows one to easily construct both simple and more complex test programs with desired performance properties. We briefly report on experience with MPI and OpenMP performance tools when applied to the test cases generated by ATS. Copyright © 2006 John Wiley & Sons, Ltd. Michael Gerndt, Bernd Mohr, Jesper Larsson Träff |
Concurr. Comput. Pract. Exp. | 1 |
| 2006 | The monitoring request interface (MRI)abstractIn this paper, we present MRI, a high level interface for selective monitoring of code regions and data structures in single and multiprocessor environments. MRI keeps transparent the available monitoring resources from the performance analysis tools and can electively generate monitoring results as online profile information, or as postmortem traces. MRI is the first step toward a standard monitoring interface which can be used by a broad range of performance analysis tools, from profiler tools, trace producers and visualizers, up to complex automatic performance analyzers. We also present an implementation of MRI for SMPs which transparently use a simulation backend and a PAPI backend to obtain performance data. Edmond Kereku, Michael Gerndt |
IPDPS | 2 |
| 2005 | Performance Cockpit: An Extensible GUI Platform for Performance Tools
Tianchao Li 0001, Michael Gerndt |
Euro-Par | 2 |
| 2005 | Performance Analysis of Shared-Memory Parallel Applications Using Performance Properties
Karl Fürlinger, Michael Gerndt |
HPCC | 2 |
| 2005 | SMART: A Simulation Tool for Analyzing Cache Access Behavior on SMPsabstractThis paper presents SMART - a simulation tool for analyzing the cache access behavior on SMP systems. SMART traps memory access events of multi-threaded applications, simulates the accesses in multiple levels of caches of multiple processors and the shared memory, emulates a novel hardware monitor that records events within given address ranges of interest, and presents the result as event counts or histogram in arbitrary granularity. Used independently or together with the advanced tools developed in the EP-Cache project, SMART can help evaluate the performance of multi-threaded applications with different hardware configurations and facilitate the application of effective code transformations for optimization. Tianchao Li 0001, Michael Gerndt |
MASCOTS | 2 |
| 2005 | Automatic performance analysis tools for the GridabstractAbstract Applications on Grids require scalable and online performance analysis tools. The execution environment of such applications includes a large number of processors. In addition, some of the resources such as the network will be shared with other applications. This requires applications to adapt dynamically to resource changes. The article presents the requirements of Grid application classes for performance analysis tools. It introduces a new analysis environment currently being developed for teraflop computers within the Peridot project at Technische Universität München. This environment applies a distributed automatic performance analysis approach. It is based on a formal specification of performance properties in the APART specification language. It uses a hierarchy of analysis agents that obtain performance data from a configurable monitoring system. This scalable design allows performance analysis for large Grid applications. Copyright © 2005 John Wiley & Sons, Ltd. Michael Gerndt |
Concurr. Pract. Exp. | 1 |
| 2005 | Monitoring cache behavior on parallel SMP architectures and related programming tools
Thomas Brandes, Helmut Schwamborn, Michael Gerndt, Jürgen Jeitner, Edmond Kereku, Martin Schulz 0001, Holger Brunst, Wolfgang E. Nagel, Reinhard Neumann, Ralph Müller-Pfefferkorn, Bernd Trenkler, Wolfgang Karl, Jie Tao 0001, Hans-Christian Hoppe |
Future Gener. Comput. Syst. | 3 |
| 2004 | Evaluating OpenMP Performance Analysis Tools with the APART Test Suite
Michael Gerndt, Bernd Mohr, Jesper Larsson Träff |
Euro-Par | 1 |
| 2004 | A Data Structure Oriented Monitoring Environment for Fortran OpenMP Programs
Edmond Kereku, Tianchao Li 0001, Michael Gerndt, Josef Weidendorfer |
Euro-Par | 3 |
| 2003 | Distributed Application Monitoring for Clustered SMP Architectures
Karl Fürlinger, Michael Gerndt |
Euro-Par | 2 |
| 2003 | Topic Introduction
Michael Gerndt, Chau-Wen Tseng, Michael F. P. O'Boyle, Markus Schordan |
Euro-Par | 1 |
| 2001 | Topic 01: Support Tools and Environments
Michael Gerndt |
Euro-Par | 1 |
| 2000 | Support Tools and Environments
Barton P. Miller, Michael Gerndt |
Euro-Par | 2 |
| 2000 | Specification of Performance Problems in MPI Programs with ASLabstractPerformance analysis is an important step in tuning performance critical applications. It is a cyclic process of measuring and analyzing performance data which is driven by the programmers hypotheses on potential performance problems. Currently this process is controlled manually by the programmer. The implicit knowledge applied in this cyclic process must be formalized in order to be reused in the automation of performance analysis tools. This article describes the performance property specification language ASL developed in the APART Esprit IV working group. ASL allows the specification of performance data via an object model and of performance properties via a specially designed notation. Performance bottlenecks can then be identified based on the specification since bottlenecks are viewed as performance properties with a huge negative impact. We present the ASL language in the context of MPI applications. Thomas Fahringer, Michael Gerndt, Graham D. Riley, Jesper Larsson Träff |
ICPP | 2 |
| 1998 | High-Level Programming of Massively Parallel Computers Based on Shared Virtual Memory
Michael Gerndt |
Parallel Comput. | 1 |
| 1997 | A Rule-based Approach for Automatic Bottleneck Detection in Programs on SharedabstractProgramming distributed memory multiprocessors requires program parallelization as well as program optimization with respect to data locality. SVM-Fortran is a programming language for shared virtual memory architectures with special language features for specifying the distribution of parallel tasks onto the processors. It is realized on top of a shared virtual memory implementation on Intel Paragon. A programming environment provides performance analysis tools helping the user in the optimization of data locality. The paper outlines the environment, describes the basic concepts of the performance analysis support, and presents a design for the automation of performance analysis. Michael Gerndt, Andreas Krumme |
HIPS | 1 |
| 1992 | Concurrent File Operations in a High Performance FORTRANabstractThe authors propose constructs to specify I/O (input/output) operations for distributed data structures in the context of Vienna FORTRAN. These operations can be used by the programmer to provide information which will allow the compiler and runtime environment to optimize the transfer of data to and from secondary storage. Although the language constructs presented have been proposed in the context of Vienna FORTRAN, they can be easily integrated into any other high-performance FORTRAN extension.> Peter Brezany, Michael Gerndt, Piyush Mehrotra, Hans P. Zima |
SC | 2 |
| 1991 | Work distribution in parallel programs for distributed memory multiprocessorsabstractA4424DANQAWlUNIl l Michael Gerndt |
ICS | 1 |
| 1990 | Updating Distributed Variables in Local ComputationsabstractAbstract This paper describes special aspects of MIMD parallelization in SUPERB. SUPERB is an interactive SIMD/MIMD parallelizing system for the SUPRENUM machine. The main topic of this paper is the updating of distributed variables in parallelized applications. The intended applications perform local computations on a large data domain. Michael Gerndt |
Concurr. Pract. Exp. | 1 |
| 1989 | Array distribution in SUPERBabstractThis paper describes MIMD parallelization in SUPERB. SUPERB is an interactive system for semi-automatic transformation of Fortran 77 programs into parallel programs for the SUPRENUM machine. The main topic of this paper is array distribution as a basis for transformations used in MIMD parallelization. Michael Gerndt |
ICS | 1 |
| 1988 | Advanced tools and techniques for automatic parallelization
Ulrich Kremer, Heinz-J. Bast, Michael Gerndt, Hans P. Zima |
Parallel Comput. | 3 |
| 1988 | SUPERB: A tool for semi-automatic MIMD/SIMD parallelization
Hans P. Zima, Heinz-J. Bast, Michael Gerndt |
Parallel Comput. | 3 |
| 1987 | MIMD-Parallelization for SUPENUM
Michael Gerndt, Hans P. Zima |
ICS | 1 |