EDBT 2026 Demo / reviewers in the wild / expert
David Böhme
dblp:89/5516
· DBLP profile ↗
13ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-4159-1519ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Performance modeling and evaluation · 47% High-performance computing · 26% Interconnection networks and networks-on-chip · 15% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
workload characterization |
0.7 | 1 | 2023 | Thicket: Seeing the Performance Experiment Forest for the Individual Run Trees · HPDC 2023 |
High-performance computing
application porting |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Parallel and multicore computing
programming models |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Interconnection networks and networks-on-chip › network topology › tree networks
fat-tree network |
0.3 | 1 | 2017 | Predicting the performance impact of different fat-tree configurations · SC 2017 |
Interconnection networks and networks-on-chip › network reconfiguration
network configuration |
0.3 | 1 | 2017 | Predicting the performance impact of different fat-tree configurations · SC 2017 |
Performance modeling and evaluation › simulation › communication system simulation
network simulation |
0.3 | 1 | 2017 | Predicting the performance impact of different fat-tree configurations · SC 2017 |
Performance modeling and evaluation
performance analysis tools |
0.2 | 1 | 2016 | Caliper: performance introspection for HPC software stacks · SC 2016 |
Performance modeling and evaluation
trace analysis |
0.2 | 1 | 2015 | Recovering logical structure from Charm++ event traces · SC 2015 |
High-performance computing
performance optimization |
0.2 | 1 | 2023 | Thicket: Seeing the Performance Experiment Forest for the Individual Run Trees · HPDC 2023 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Parallel and multicore computing › parallel programming runtimes
task-based runtime |
0.1 | 1 | 2015 | Recovering logical structure from Charm++ event traces · SC 2015 |
Methods — techniques the papers use, named apart from their topics
performance modeling · 0.7k-means clustering · 0.7exploratory data analysis · 0.7dimensionality reduction · 0.7TraceR-CODES simulation · 0.3performance data collection abstraction · 0.2task ordering · 0.2heuristics · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Mechanism to Generate Interception Based Tools for HPC Libraries
Bengisu Elis, David Böhme, Olga Pearce, Martin Schulz 0001 |
Euro-Par (1) | 2 |
| 2024 | Non-Blocking GPU-CPU Notifications to Enable More GPU-CPU ParallelismabstractGPUs are increasingly popular in HPC systems, and more applications are adopting GPUs each day. However, the control synchronization of GPUs with CPUs is suboptimal and only possible after GPU kernel termination points, resulting in serialized host and device tasks. In this paper, we propose a novel CPU-GPU notification method that enables non-blocking in-kernel control synchronization of device and host tasks in combination with persistent GPU kernels. Using this notification method, we increase the overlap of CPU and GPU execution and with that parallelism. We present the concept and structure of the proposed notification mechanism together with in-kernel GPU-CPU control synchronization, using halo-exchange as an example. We analyze the performance of the halo-exchange pattern using our new notification method, as well as the interference between CPU and GPU operations due to the execution overlap. Finally, we verify our results using a performance model covering the halo-exchange pattern with the new notification method. Bengisu Elis, Olga Pearce, David Böhme, Jason Burmark, Martin Schulz 0001 |
HPC Asia | 3 |
| 2023 | Thicket: Seeing the Performance Experiment Forest for the Individual Run TreesabstractThicket is an open-source Python toolkit for Exploratory Data Analysis (EDA) of multi-run performance experiments. It enables an understanding of optimal performance configuration for large-scale application codes. Most performance tools focus on a single execution (e.g., single platform, single measurement tool, single scale). Thicket bridges the gap to convenient analysis in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets by providing an interface for interacting with the performance data. Thicket has a modular structure composed of three components. The first component is a data structure for multi-dimensional performance data, which is composed automatically on the portable basis of call trees, and accommodates any subset of dimensions present in the dataset. The second is the metadata, enabling distinction and sub-selection of dimensions in performance data. The third is a dimensionality reduction mechanism, enabling analysis such as computing aggregated statistics on a given data dimension. Extensible mechanisms are available for applying analyses (e.g., top-down on Intel CPUs), data science techniques (e.g., K-means clustering from scikit-learn), modeling performance (e.g., Extra-P), and interactive visualization. We demonstrate the power and flexibility of Thicket through two case studies, first with the open-source RAJA Performance Suite on CPU and GPU clusters and another with a large physics simulation run on both a traditional HPC cluster and an AWS Parallel Cluster instance. Stephanie Brink, Michael McKinsey, David Böhme, Connor Scully-Allison, Ian Lumsden, W. Daryl Hawkins, Treece Burgess, Vanessa Lama, Jakob Lüttgau, Katherine E. Isaacs, Michela Taufer, Olga Pearce |
HPDC | 3 |
| 2020 | CodeSeer: input-dependent code variants selection via machine learningabstractIn high performance computing (HPC), scientific simulation codes are executed repeatedly with different inputs. The peak performance of these programs heavily depends on various compiler optimizations, which are often selected agnostically on program input or may be selected with sensitivity to just a single input. When subsequently executed, often with different inputs, performance may suffer for all or all but the one input tested, and for the latter potentially even compared to the O3 baseline. Tao Wang 0077, David Böhme, D. A. Beckingsale, Frank Mueller 0001, Todd Gamblin |
ICS | 3 |
| 2019 | FuncyTuner: Auto-tuning Scientific Applications With Per-loop CompilationabstractThe de facto compilation model for production software compiles all modules of a target program with a single set of compilation flags, typically 02 or 03. Such a per-program compilation strategy may yield sub-optimal executables since programs often have multiple hot loops with diverse code structures and may be better optimized with a per-region compilation model that assembles an optimized executable by combining the best per-region code variants. Tao Wang 0077, D. A. Beckingsale, David Böhme, Frank Mueller 0001, Todd Gamblin |
ICPP | 4 |
| 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous systemabstractProductivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience. Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz |
SC | 32 |
| 2017 | Flexible Data Aggregation for Performance ProfilingabstractAlmost all performance analysis tools in the HPC space perform some form of aggregation to compute summary information of a series of performance measurements, from summations to more complex operations like histograms. Aggregation not only reduces data volumes and consequently storage space requirements and overheads, but is also crucial to extract insights from recorded measurement data. In current tools, however, most aspects that control the aggregation, such as the data dimensions to be reduced, are hard-coded in the tool for a set of particular use cases identified by the tool developer and cannot be extended or modified by the user. This limits their flexibility and often results in users having to learn and use multiple tools with different aggregation options for their performance analysis needs.We present a novel approach for performance data aggregation based on a flexible key:value data model with user-defined attributes, where users can define custom aggregation schemes in a simple description language. This not only gives users the control to deploy the particular data aggregation they need, but also opens the door for aggregations along application-specific data dimensions that cannot be achieved with traditional profiling tools. We show how our approach can be applied for performance profiling at runtime, cross-process data aggregation, and interactive data analysis and demonstrate its functionality with several case studies driven by real world codes. David Böhme, D. A. Beckingsale, Martin Schulz 0001 |
CLUSTER | 1 |
| 2017 | Predicting the performance impact of different fat-tree configurationsabstractThe fat-tree topology is one of the most commonly used network topologies in HPC systems. Vendors support several options that can be configured when deploying fat-tree networks on production systems, such as link bandwidth, number of rails, number of planes, and tapering. This paper showcases the use of simulations to compare the impact of these design options on representative production HPC applications, libraries, and multi-job workloads. We present advances in the TraceR-CODES simulation framework that enable this analysis and evaluate its prediction accuracy against experiments on a production fat-tree network. In order to understand the impact of different network configurations on various anticipated scenarios, we study workloads with different communication patterns, computation-to-communication ratios, and scaling characteristics. Using multi-job workloads, we also study the impact of inter-job interference on performance and compare the cost-performance tradeoffs. Abhinav Bhatele, Louis H. Howell, David Böhme, Ian Karlin, Edgar A. León, Misbah Mubarak, Noah Wolfe, Todd Gamblin, Matthew L. Leininger |
SC | 4 |
| 2016 | Caliper: performance introspection for HPC software stacksabstractMany performance engineering tasks, from long-term performance monitoring to post-mortem analysis and online tuning, require efficient runtime methods for introspection and performance data collection. To understand interactions between components in increasingly modular HPC software, performance introspection hooks must be integrated into runtime systems, libraries, and application codes across the software stack. This requires an interoperable, cross-stack, general-purpose approach to performance data collection, which neither application-specific performance measurement nor traditional profile or trace analysis tools provide. With Caliper, we have developed a general abstraction layer to provide performance data collection as a service to applications, runtime systems, libraries, and tools. Individual software components connect to Caliper in independent data producer, data consumer, and measurement control roles, which allows them to share performance data across software stack boundaries. We demonstrate Caliper's performance analysis capbilities with two case studies of production scenarios. David Böhme, Todd Gamblin, D. A. Beckingsale, Peer-Timo Bremer, Alfredo Giménez, Matthew P. LeGendre, Olga Pearce, Martin Schulz 0001 |
SC | 1 |
| 2015 | Recovering logical structure from Charm++ event tracesabstractAsynchrony and non-determinism in Charm++ programs present a significant challenge in analyzing their event traces. We present a new framework to organize event traces of parallel programs written in Charm++. Our reorganization allows one to more easily explore and analyze such traces by providing context through logical structure. We describe several heuristics to compensate for missing dependencies between events that currently cannot be easily recorded. We introduce a new task ordering that recovers logical structure from the non-deterministic execution order. Using the logical structure, we define several metrics to help guide developers to performance problems. We demonstrate our approach through two proxy applications written in Charm++. Finally, we discuss the applicability of this framework to other task-based runtimes and provide guidelines for tracing to support this form of analysis. Katherine E. Isaacs, Abhinav Bhatele, Jonathan Lifflander, David Böhme, Todd Gamblin, Martin Schulz 0001, Bernd Hamann, Peer-Timo Bremer |
SC | 4 |
| 2013 | Understanding the formation of wait states in applications with one-sided communicationabstractTo better understand the formation of wait states in MPI programs and to support the user in finding optimization targets in the case of load imbalance, a major source of wait states, we added in our earlier work two new trace-analysis techniques to Scalasca, a performance analysis tool designed for large-scale applications. In this paper, we show how the two techniques, which were originally restricted to two-sided and collective MPI communication, are extended to cover also one-sided communication. We demonstrate our experiences with benchmark programs and a mini-application representing the core of the POP ocean model. Marc-André Hermanns, Manfred Miklosch, David Böhme, Felix Wolf 0001 |
EuroMPI | 3 |
| 2012 | Scalable Critical-Path Based Performance AnalysisabstractThe critical path, which describes the longest execution sequence without wait states in a parallel program, identifies the activities that determine the overall program runtime. Combining knowledge of the critical path with traditional parallel profiles, we have defined a set of compact performance indicators that help answer a variety of important performance-analysis questions, such as identifying load imbalance, quantifying the impact of imbalance on runtime, and characterizing resource consumption. By replaying event traces in parallel, we can calculate these performance indicators in a highly scalable way, making them a suitable analysis instrument for massively parallel programs with thousands of processes. Case studies with real-world parallel applications confirm that - in comparison to traditional profiles - our indicators provide enhanced insight into program behavior, especially when evaluating partitioning schemes of MPMD programs. David Böhme, Felix Wolf 0001, Bronis R. de Supinski, Martin Schulz 0001, Markus Geimer |
IPDPS | 1 |
| 2010 | Identifying the Root Causes of Wait States in Large-Scale Parallel ApplicationsabstractDriven by growing application requirements and accelerated by current trends in microprocessor design, the number of processor cores on modern supercomputers is increasing from generation to generation. However, load or communication imbalance prevents many codes from taking advantage of the available parallelism, as delays of single processes may spread wait states across the entire machine. Moreover, when employing complex point-to-point communication patterns, wait states may propagate along far-reaching cause-effect chains that are hard to track manually and that complicate an assessment of the actual costs of an imbalance. Building on earlier work by Meira Jr. et al., we present a scalable approach that identifies program wait states and attributes their costs in terms of resource waste to their original cause. By replaying event traces in parallel both in forward and backward direction, we can identify the processes and call paths responsible for the most severe imbalances even for runs with tens of thousands of processes. David Böhme, Markus Geimer, Felix Wolf 0001, Lukas Arnold |
ICPP | 1 |