Hartmut Kaiser

dblp:63/6379 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-8712-2806ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Locks Must Die: Composable Mutual Exclusion Implemented by Dynamic Resource Sharing on Task Graphs
abstract
Locks, the enduring de-facto standard for synchronization in parallel programs, suffer from non-composability and are conducive to a variety of hard-to-find bugs such as data races and deadlocks. In this paper, we propose a straightforward replacement of the traditional lock-based mutex with the Guard, a construct providing mutual exclusion through efficient, dynamic serialization of tasks in a dynamic task graph. Unlike with locks, programming with Guards is lock-free, composable, and immune to deadlock. Guards are compatible with arbitrary executor services, and can be implemented on any platform which supports atomic variables and compare-and-set operations. We will describe the general construction and API for Guards, introduce our reference implementation, and analyze through various benchmarks the performance of our reference implementation against the standard ReentrantLock class used in Java. The practice of writing multithreaded code will benefit from the safety, ease of use, and performance of Guards.
Max Morris, Steven R. Brandt, Hartmut Kaiser
eScience3
2024 Simulating stellar merger using HPX/Kokkos on A64FX on Supercomputer Fugaku
Patrick Diehl, Gregor Daiß, Kevin A. Huck, Dominic Marcello, Sagiv Shiber, Hartmut Kaiser, Dirk Pflüger
J. Supercomput.6
2023 Traveler: Navigating Task Parallel Traces for Performance Analysis
abstract
Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace-a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. To address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler, an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.
Sayef Azad Sakin, Alex Bigelow, R. Tohid, Connor Scully-Allison, Carlos Scheidegger, Steven R. Brandt, Kevin A. Huck, Hartmut Kaiser, Katherine E. Isaacs
IEEE Trans. Vis. Comput. Graph.9
2022 Towards superior software portability with SHAD and HPX C++ libraries
abstract
As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD's portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.
Nanmiao Wu, Vito Giovanni Castellana, Hartmut Kaiser
CF3
2021 Octo-Tiger's New Hydro Module and Performance Using HPX+CUDA on ORNL's Summit
abstract
Octo-Tiger is a code for modeling three-dimensional self-gravitating astrophysical fluids. It was particularly designed for the study of dynamical mass transfer between interacting binary stars. Octo-Tiger is parallelized for distributed systems using the asynchronous many-task runtime system, the C++ standard library for parallelism and concurrency (HPX) and utilizes CUDA for its gravity solver. Recently, we have remodeled Octo-Tiger’s hydro solver to use a three-dimensional reconstruction scheme. In addition, we have ported the hydro solver to GPU using CUDA kernels. We present scaling results for the new hydro kernels on ORNL’s Summit machine using a Sedov-Taylor blast wave problem. We also compare Octo-Tiger’s new hydro scheme with its old hydro scheme, using a rotating star as a test problem.
Patrick Diehl, Gregor Daiß, Dominic Marcello, Kevin A. Huck, Sagiv Shiber, Hartmut Kaiser, Juhan Frank, Geoffrey C. Clayton, Dirk Pflüger
CLUSTER6
2019 Runtime Adaptive Task Inlining on Asynchronous Multitasking Runtime Systems
abstract
As the era of high frequency, single core processors have come to a close, the new paradigm of many core processors has come to dominate. In response to these systems, asynchronous multitasking runtime systems have been developed as a promising solution to efficiently utilize these newly available hardware. Asynchronous multitasking runtime systems work by dividing a problem into a large number of fine grained tasks. However, as the number of tasks created increase, the overheads associated with task creation and management cannot be ignored. Task inlining, a method where the parent thread consumes a child thread, enables the runtime system to achieve the balance between parallelism and its overhead. As largely impacted by different processor architectures, the decision of task inlining is dynamic in nature. In this research, we present adaptive techniques for deciding, at runtime, whether a particular task should be inlined or not. We present two policies, a baseline policy that makes inlining decision based on a fixed threshold and an adaptive policy which decides the threshold dynamically at runtime. We also evaluate and justify the performance of these policies on different processor architectures. To the best of our knowledge, this is the first study of the impacts of adaptive policy at runtime for task inlining in an asynchronous multitasking runtime system on different processor architectures. From experimentation, we find that the baseline policy improves the execution time from 7.61% to 54.09%. Furthermore, the adaptive policy improves over the baseline policy by up to 74%.
Bibek Wagle, Mohammad Alaul Haque Monil, Kevin A. Huck, Allen D. Malony, Adrian Serio, Hartmut Kaiser
ICPP6
2019 From piz daint to the stars: simulation of stellar mergers using high-level abstractions
abstract
We study the simulation of stellar mergers, which requires complex simulations with high computational demands. We have developed Octo-Tiger, a finite volume grid-based hydrodynamics simulation code with Adaptive Mesh Refinement which is unique in conserving both linear and angular momentum to machine precision. To face the challenge of increasingly complex, diverse, and heterogeneous HPC systems, Octo-Tiger relies on high-level programming abstractions.
Gregor Daiß, Parsa Amini, John Biddiscombe, Patrick Diehl, Juhan Frank, Kevin A. Huck, Hartmut Kaiser, Dominic Marcello, David Pfander, Dirk Pflüger
SC7
2015 The Performance Implication of Task Size for Applications on the HPX Runtime System
abstract
As High Performance Computing moves toward Exascale, where parallel applications will be expected to run on millions of cores concurrently, every component of the computational model must perform optimally. One such component, the task scheduler, can potentially be optimized to runtime application requirements. We focus our study using a task-based runtime system, one possible solution towards Exascale computation. Based on task size and scheduler, the overheads associated with task scheduling vary. Therefore, to minimize overheads and optimize performance, either the task size or the scheduler must adapt. In this paper, we focus on adapting the task size, which can be easily done statically and potentially done dynamically. To this end, we first show how scheduling overheads change with task size or granularity. We then propose and execute a methodology to characterize these overheads and dynamically measure the effects of task granularity. The HPX runtime system [1] employs asynchronous fine-grained task scheduling and incorporates a dynamic performance modeling capability, providing an ideal experimental platform. Using the performance counter capabilities in HPX, we characterize task scheduling overheads and show metrics to determine optimal task size. This is the first step toward the goal of dynamically adapting task size to optimize parallel performance.
Patricia Grubel, Hartmut Kaiser, Jeanine E. Cook, Adrian Serio
CLUSTER2
2010 What Is the Price of Simplicity? - A Cross-Platform Evaluation of the SAGA API
Mathijs den Burger, Ceriel J. H. Jacobs, Thilo Kielmann, André Merzky, Ole Weidner, Hartmut Kaiser
Euro-Par (1)6
2009 Programming Abstractions for Data Intensive Computing on Clouds and Grids
abstract
MapReduce has emerged as an important data-parallel programming model for data-intensive computing - for Clouds and Grids. However most if not all implementations of MapReduce are coupled to a specific infrastructure. SAGA is a high-level programming interface which provides the ability to create distributed applications in an infrastructure independent way. In this paper, we show how MapReduce has been implemented using SAGA and demonstrate its interoperability across different distributed platforms - Grids, Cloud-like infrastructure and Clouds. We discuss the advantages of programmatically developing MapReduce using SAGA, by demonstrating that the SAGA-based implementation is infrastructure independent whilst still providing control over the deployment, distribution and runtime decomposition. The ability to control the distribution and placement of the computation units (workers) is critical in order to implement the ability to move computational work to the data. This is required to keep data network transfer low and in the case of commercial Clouds the monetary cost of computing the solution low. Using data-sets of size up to 10GB, and upto 10 workers, we provide detailed performance analysis of the SAGA-MapReduce implementation, and show how controllingthe distribution of computation and the payload per worker helps enhance performance.
Chris Miceli, Michael Miceli, Shantenu Jha, Hartmut Kaiser, André Merzky
CCGRID4
2008 Towards an integrated GIS-based coastal forecast workflow
abstract
Abstract The SURA Coastal Ocean Observing and Prediction (SCOOP) program is using geographical information system (GIS) technologies to visualize and integrate distributed data sources from across the United States and Canada. Hydrodynamic models are run at different sites on a developing multi‐institutional computational Grid. Some of these predictive simulations of storm surge and wind waves are triggered by tropical and subtropical cyclones in the Atlantic and the Gulf of Mexico. Model predictions and observational data need to be merged and visualized in a geospatial context for a variety of analyses and applications. A data archive at LSU aggregates the model outputs from multiple sources, and a data‐driven workflow triggers remotely performed conversion of a subset of model predictions to georeferenced data sets, which are then delivered to a Web Map Service located at Texas A&M University. Other nodes in the distributed system aggregate the observational data. This paper describes the use of GIS within the SCOOP program for the 2005 hurricane season, along with details of the data‐driven distributed dataflow and workflow, which results in geospatial products. We also focus on future plans related to the complimentary use of GIS and Grid technologies in the SCOOP program, through which we hope to provide a wider range of tools that can enhance the tools and capabilities of earth science research and hazard planning. Copyright © 2008 John Wiley & Sons, Ltd.
Gabrielle Allen, Philip Bogden, Gerry Creager, Chirag Dekate, Carola Jesch, Hartmut Kaiser, Jon MacLaren, Will Perrie, Gregory W. Stone, Xiongping Zhang
Concurr. Comput. Pract. Exp.6
2007 Design and Implementation of Network Performance Aware Applications Using SAGA and Cactus
abstract
This paper demonstrates the use of appropriate programming abstractions - SAGA and cactus - that facilitate the development of applications for distributed infrastructure. SAGA provides a high-level programming interface to Grid- functionality; Cactus is an extensible, component based framework for scientific applications. We show how SAGA can be integrated with cactus to develop simple, useful and easily extensible applications that can be deployed on a wide variety of distributed infrastructure, independent of the details of the resources. Our model application can gather and analyze network performance data and migrate across heterogeneous resources. We outline the architecture of our application and discuss how it imparts important features required of eScience applications. As a proof-of-concept, we present details of the successful deployment of our application over distinct and heterogeneous Grids and present the network performance data gathered. We also discuss several interesting use cases for such an application - which can be used either as stand-alone network diagnostic agent, or in conjunction with more complex scientific applications.
Shantenu Jha, Hartmut Kaiser, Yaakoub El Khamra, Ole Weidner
eScience2
2007 Grid Interoperability at the Application Level Using SAGA
abstract
SAGA is a high-level programming abstraction, which significantly facilitates the development and deployment of Grid-aware applications. The primary aim of this paper is to discuss how each of the three main components of the SAGA landscape - interface specification, specific implementation and the different adaptors for middleware distribution - facilitate application-level interoperability. We discuss SAGA in relation to the ongoing GIN Community Group efforts and show the consistency of the SAGA approach with the GIN Group efforts. We demonstrate how interoperability can be enabled by the use of SAGA, by discussing two simple, yet meaningful applications: in the first, SAGA enables applications to utilize interoperability and in the second example SAGA adaptors provide the basis for interoperability.
Shantenu Jha, Hartmut Kaiser, André Merzky, Ole Weidner
eScience2
2006 Poster reception - The SAGA C++ reference implementation: a milestone toward new high-level grid applications
abstract
In Grid computing multiple, incompatible middleware frameworks exist and are widely used in large research and production environments. Standard specifications for such Grid middleware are rarely available, and often still unstable. This hinders the ability of application programmers to write portable Grid application code.The Simple API for Grid Applications (SAGA) is an ongoing standardization effort within the Open Grid Forum (OGF). The SAGA API provides a simple, uniform interface for applications which utilize very dynamic and heterogeneous Grid environments. With SAGA, programmers of high-level Grid application do not have to learn about underlying Grid middleware layers and frameworks.Our newly developed SAGA C++ reference implementation makes this API available for real-world applications, providing a flexible, extensible, portable, and generic framework usable in dynamic environments.This poster describes the key features of SAGA, and provides examples of its use from the C++ reference implementation.
Hartmut Kaiser, André Merzky, Stephan Hirmer, Gabrielle Allen, Edward Seidel
SC1
2006 Poster reception - Utilizing grid computing technologies for advanced reservoir studies
abstract
Reservoir studies are crucial to obtain accurate assessments and predictions of reservoir performance. However, this is a challenging issue because 1) it relies on massive modeling-related, geographically distributed, terabyte or even petabyte sized datasets (seismic and well-logging data), 2) needs to rapidly perform hundreds or thousands of simulations, being identical runs with different reservoir models circulating the impacts of various uncertainty factors, 3) the lack of easy-to-use problem solving toolkits to assist the uncertainty analysis.The poster focuses on leveraging Grid computing technologies to address the challenging issue mentioned above. It describes a newly developed data archive tool based on metadata and replica services and high performance file transfer. Our task farming framework enables a large amount of parallel job runs across a Grid, a related Grid portal eases the management of advanced reservoir studies. Our solutions are being employed by other Grid applications.
Zhou Lei 0001, Gabrielle Allen, Dayong Huang, Hartmut Kaiser, Christopher D. White
SC4
2006 Distributed and collaborative visualization of large data sets using high-speed networks
Andrei Hutanu, Gabrielle Allen, Stephen David Beck, Petr Holub, Hartmut Kaiser, Archit Kulshrestha, Milos Liska, Jon MacLaren, Ludek Matyska, Ravi Paruchuri, Steffen Prohaska, Edward Seidel, Brygg Ullmer, Shalini Venkataraman
Future Gener. Comput. Syst.5
2005 The Grid Application Toolkit: Toward Generic and Easy Application Programming Interfaces for the Grid
abstract
Core Grid technologies are rapidly maturing, but there remains a shortage of real Grid applications. One important reason is the lack of a simple and high-level application programming toolkit, bridging the gap between existing Grid middleware and application-level needs. The Grid Application Toolkit (GAT), as currently developed by the EC-funded project GridLab, provides this missing functionality. As seen from the application, the GAT provides a unified simple programming interface to the Grid infrastructure, tailored to the needs of Grid application programmers and users. A uniform programming interface will be needed for application developers to create a new generation of "Grid-aware" applications. The GAT implementation handles both the complexity and the variety of existing Grid middleware services via so-called adaptors. Complementing existing Grid middleware, GridLab also provides high-level services to implement the GAT functionality. We present the GridLab software architecture, consisting of the GAT, environment-specific adaptors, and GridLab services. We elaborate the concepts underlying the GAT and outline the corresponding application programming interface. We present the functionality of GridLab's high-level services and demonstrate how a dynamic Grid application can easily benefit from the GAT. All GridLab software is open source and can be downloaded from the project Web site.
Gabrielle Allen, Kelly Davis, Tom Goodale, Andrei Hutanu, Hartmut Kaiser, Thilo Kielmann, André Merzky, Rob van Nieuwpoort, Alexander Reinefeld, Florian Schintke, Thorsten Schütt, Edward Seidel, Brygg Ullmer
Proc. IEEE5