EDBT 2026 Demo / reviewers in the wild / expert
Matthias S. Müller
dblp:13/1808
· DBLP profile ↗
41ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-2545-5258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Extending the SPMD IR for RMA Models and Static Data Race Detection
Semih Burak, Simon Schwitanski, Felix Tomski, Jens Domke, Matthias S. Müller |
EuroMPI | 5 |
| 2024 | RMASanitizer: Generalized Runtime Detection of Data Races in Remote Memory Access ApplicationsabstractRemote Memory Access (RMA) programming models enable processes running on a distributed-memory computer to access and manipulate the memory of other processes directly. Such one-sided communication has the benefit that the receiving process is not actively involved in the communication compared to the classical two-sided message-passing model. The three programming models MPI RMA, OpenSHMEM, and GASPI provide such a communication scheme. However, RMA models require the developer to synchronize the accesses with corresponding API calls correctly. Concurrent modifications of the same (remote) memory location due to wrong or missing synchronization lead to data races. Such data races are undefined behavior and may result in non-deterministic failures of the program execution. This paper presents RMASanitizer, an on-the-fly race detector for MPI RMA, OpenSHMEM, and GASPI applications. It relies on a generalized race detection model independent of the concrete RMA programming model. RMASanitizer combines a dynamic on-the-fly analysis with a static analysis at compile-time that detects and instruments only relevant memory accesses. It is implemented as part of the MPI correctness checking framework MUST which we extended with support for OpenSHMEM and GASPI. We show that RMASanitizer can detect races in MPI RMA, OpenSHMEM, and GASPI applications with an accuracy of over 95 percent by running it on the data race benchmark suite RMARaceBench. On proxy applications, the slowdown for the execution with up to 700 processes ranges from 1.1x to 30x, depending on the application, showing that our tool is applicable in practice. Simon Schwitanski, Yussur Mustafa Oraji, Cornelius Pätzold, Joachim Jenke, Felix Tomski, Matthias S. Müller |
ICPP | 6 |
| 2024 | SPMD IR: Unifying SPMD and Multi-value IR Showcased for Static Verification of Collectives
Semih Burak, Ivan R. Ivanov, Jens Domke, Matthias S. Müller |
EuroMPI | 4 |
| 2024 | An Experimental Setup to Evaluate RAPL Energy Counters for Heterogeneous MemoryabstractPower consumption of the main memory in modern heterogeneous high-performance computing (HPC) constitutes a significant part of the total power consumption of a node. This motivates energy-efficient solutions targeting the memory domain as well. Practitioners need reliable energy measurement techniques for analyzing energy and power consumption of applications and performance optimizations. Running Average Power Limit (RAPL) is a common choice, as it provides uncomplicated access to the energy measurements. While RAPL's accuracy has been studied and validated on homogeneous memory platforms, no work we are aware of investigated its accuracy on heterogeneous memory platforms, specifically with high-capacity memory (HCM). This paper describes the process of measuring the memory power consumption externally using riser cards in detail. We validate RAPL's accuracy by comparing results obtained from Intel's Ice Lake-SP system equipped with DDR4 DRAM and Intel Optane Persistent Memory Modules (PMM). In addition, we verify the accuracy of our instrumentation setup by comparing the results from an older Broadwell system with the results in the literature. We show that the RAPL values on a heterogeneous memory system report a higher offset from the reference measurements. The difference is more pronounced at lower memory load for all memory types. Also, we find that RAPL readings are inconsistent between multiple sockets and over time. Based on the evaluated scenarios, we conclude that RAPL overestimates the actual power consumption on heterogeneous memory systems and provide a discussion on the possible causes of this effect. Lukas Alt, Anara Kozhokanova, Thomas Ilsche, Christian Terboven, Matthias S. Müller |
ICPE | 5 |
| 2024 | Parallel Pattern Compiler for Automatic Global OptimizationsabstractHigh-performance computing (HPC) systems enable scientific advances through simulation and data processing. The heterogeneity in HPC hardware and software increases the application complexity and reduces its maintainability and productivity. This work proposes a prototype implementation for a parallel pattern-based source-to-source compiler to address these challenges. The prototype limits the complexity of parallelism and heterogeneous architectures to parallel patterns that are optimized towards a given target architecture. By applying high-level optimizations and a mapping between parallel patterns and execution units during compile time, portability between systems is achieved. The compiler can address architectures with shared memory, distributed memory, and accelerator offloading. The approach shows speedups for seven of the nine supported Rodinia benchmarks, reaching speedups of up to twelve times. Porting LULESH to the Parallel Pattern Language (PPL) shows a compression of code size by 65% (3.4 thousand lines of code) through a more concise expression and a higher level of abstraction. The tool’s limitations include dynamic algorithms that are challenging to analyze statically and overheads during the compile time optimization. This paper is an extended version of a previous PMAM publication (Schmitz et al., 2024). Adrian Schmitz, Semih Burak, Julian Miller, Matthias S. Müller |
Parallel Comput. | 4 |
| 2023 | RLP: Power Management Based on a Latency-Aware Roofline ModelabstractThe ever-growing power draw in high-performance computing (HPC) clusters and the rising energy costs enforce a pressing urge for energy-efficient computing. Consequently, advanced infrastructure orchestration is required to regulate power dissipation efficiently. In this work, we propose a novel approach for managing power consumption at runtime based on the well-known roofline model and call it Roofline Power (RLP) management. The RLP employs rigorously selected but generally available hardware performance events to construct rooflines, with minimal overheads. In particular, RLP extends the original roofline model to include the memory access latency metric for the first time. The extension identifies whether execution is bandwidth, latency, or compute-bound, and improves the modeling accuracy. We evaluated the RLP model on server-grade CPUs and a GPU with real-world HPC workloads in two scenarios: optimization with and without power capping. Compared to system default settings, RLP reduces the energy-to-solution up to 22% with negligible performance degradation. The other scenario accelerates the execution up to 14.7% under power capping. In addition, RLP outperforms other state-of-the-art techniques in generality and effectiveness. Anara Kozhokanova, Christian Terboven, Matthias S. Müller |
IPDPS | 4 |
| 2022 | MPI detach - Towards automatic asynchronous local completion
Joachim Jenke, Marc-André Hermanns, Matthias S. Müller, Van Man Nguyen, Julien Jaeger, Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
Parallel Comput. | 3 |
| 2020 | Operation-Aware Power Capping
Julian Miller, Christian Terboven, Matthias S. Müller |
Euro-Par | 4 |
| 2020 | MPI Detach - Asynchronous Local CompletionabstractWhen aiming for large scale parallel computing, waiting time due to network latency, synchronization, and load imbalance are the primary opponents of high parallel efficiency. A common approach to hide latency with computation is the use of non-blocking communication. In the presence of a consistent load imbalance, synchronization cost is just the visible symptom of the load imbalance. Tasking approaches as in OpenMP, TBB, OmpSs, or C++20 coroutines promise to expose a higher degree of concurrency, which can be distributed on available execution units and significantly increase load balance. Available MPI non-blocking functionality does not integrate seamlessly into such tasking parallelization. In this work, we present a slim extension of the MPI interface to allow seamless integration of non-blocking communication with available concepts of asynchronous execution in OpenMP and C++. Joachim Jenke, Marc-André Hermanns, Ali C. Demiralp, Matthias S. Müller, Torsten W. Kuhlen |
EuroMPI | 4 |
| 2020 | CHAMELEON: Reactive Load Balancing for Hybrid MPI+OpenMP Task-Parallel ApplicationsabstractMany applications in high performance computing are designed based on underlying performance and execution models. While these models could successfully be employed in the past for balancing load within and between compute nodes, modern software and hardware increasingly make performance predictability difficult if not impossible. Consequently, balancing computational load becomes much more difficult. Aiming to tackle these challenges in search for a general solution, we present a novel library for fine-granular task-based reactive load balancing in distributed memory based on MPI and OpenMP. With our approach, individual migratable tasks can be executed on any MPI rank. The actual executing rank is determined at run time based on online performance data. We evaluate our approach under an enforced power cap and under enforced clock frequency changes for a synthetic benchmark and show its robustness for work-induced imbalances for a realistic application. Our experiments demonstrate speedups of up to 1.31X. Jannis Klinkenberg, Philipp Samfass, Michael Bader, Christian Terboven, Matthias S. Müller |
J. Parallel Distributed Comput. | 5 |
| 2018 | Estimating the Impact of External Interference on Application Performance
Aamer Shah, Matthias S. Müller, Felix Wolf 0001 |
Euro-Par | 2 |
| 2018 | Thread-local concurrency: a technique to handle data race detection at programming model abstractionabstractWith greater adoption of various high-level parallel programming models to harness on-node parallelism, accurate data race detection has become more crucial than ever. However, existing tools have great difficulty spotting data races through these high-level models, as they primarily target low-level concurrent execution models (e.g., concurrency expressed at the level of POSIX threads). In this paper, we propose a novel technique to accurately detect those data races that can occur at higher levels of concurrent execution. The core idea of our technique is to introduce the general concept of Thread-Local Concurrency (TLC) as a new way to translate the concurrency expressed by a high-level programming paradigm into the low execution level understood by the existing tools. Specifically, we extend the definition of vector clocks to allow the existing state-of-the-art race detectors to recognize those races that occur at the higher level of concurrency with minor modifications to these tools. Our evaluation with our prototype implemented within ThreadSanitizer shows that TLC can allow the existing tool to detect these races accurately with only small additional analysis overheads. Joachim Jenke, Martin Schulz 0001, Dong H. Ahn, Matthias S. Müller |
HPDC | 4 |
| 2017 | Data Mining-Based Analysis of HPC Center OperationsabstractSize and complexity of contemporary High Performance Computing (HPC) systems increases permanently. While the reliability of a single component and compute node is high, the huge amount of components comprising these systems results in the fact that defects happen regularly. This drives the need to manage failure situations. Common issues are component failures or node soft lock-ups that typically lead to crashes of the user jobs that are scheduled on the affected node, and may cause undesired downtime. One approach to mitigate the impact of such problems is to predict node failures with a sufficient lead time in order to take proactive measures. However, accurate prediction is a challenging task.The literature describes several approaches that focus on gathering and analyzing system event logs in order to create prediction models. In this paper, we present a different approach by using descriptive statistics and supervised machine learning to create a prediction model from monitoring data. Our approach is based on the assumption, that features of a certain time frame before a critical event (i. e., a failure or soft lock-up) can serve as an indicator. Consequently, our model is trained with monitoring data from critical and healthy time frames. The evaluation with standard monitoring data collected from the HPC systems at RWTH Aachen University shows that our classifier is able to locate potentially failing nodes with a 10-fold cross precision of 98% and recall of 91 %. Jannis Klinkenberg, Christian Terboven, Stefan Lankes, Matthias S. Müller |
CLUSTER | 4 |
| 2016 | ARCHER: Effectively Spotting Data Races in Large OpenMP ApplicationsabstractOpenMP plays a growing role as a portable programming model to harness on-node parallelism, yet, existing data race checkers for OpenMP have high overheads and generate many false positives. In this paper, we propose the first OpenMP data race checker, ARCHER, that achieves high accuracy, low overheads on large applications, and portability. ARCHER incorporates scalable happens-before tracking, exploits structured parallelism via combined static and dynamic analysis, and modularly interfaces with OpenMP runtimes. ARCHER significantly outperforms TSan and Intel® Inspector XE, while providing the same or better precision. It has helped detect critical data races in the Hypre library that is central to many projects at Lawrence Livermore National Laboratory and elsewhere. Simone Atzeni, Ganesh Gopalakrishnan, Zvonimir Rakamaric, Dong H. Ahn, Ignacio Laguna, Martin Schulz 0001, Gregory L. Lee, Joachim Jenke, Matthias S. Müller |
IPDPS | 9 |
| 2016 | Development effort estimation in HPCabstractIn order to cover the ever increasing demands for computational power, while meeting electrical power and budget constraints, HPC systems are continuing to increase in hardware and software complexity. As a direct consequence, this also leads to increased development efforts to parallelize, tune or port applications. For an informed decision on how to spend available budgets, we therefore need quantitative metrics to estimate the development effort in HPC. While development effort estimation is widely used in software engineering, applying it to HPC, with its strong focus on performance, is not straightforward. In this paper, we first review existing approaches of effort estimation for general computing and then derive a novel methodology to estimate development effort specifically targeted at HPC. Further, we propose a concept to identify factors impacting development effort and encapsulate it in an effort log tool to collect data on development time. Sandra Wienke, Julian Miller, Martin Schulz 0001, Matthias S. Müller |
SC | 4 |
| 2015 | Event-Action Mappings for Parallel Tools Infrastructures
Tobias Hilbrich, Martin Schulz 0001, Holger Brunst, Joachim Jenke, Bronis R. de Supinski, Matthias S. Müller |
Euro-Par | 6 |
| 2014 | A Pattern-Based Comparison of OpenACC and OpenMP for Accelerator Computing
Sandra Wienke, Christian Terboven, James C. Beyer, Matthias S. Müller |
Euro-Par | 4 |
| 2013 | Assessing the Performance of OpenMP Programs on the Intel Xeon Phi
Dirk Schmidl, Tim Cramer, Sandra Wienke, Christian Terboven, Matthias S. Müller |
Euro-Par | 5 |
| 2013 | Intralayer Communication for Tree-Based Overlay NetworksabstractWhile various HPC tools use Tree-Based Overlay Networks (TBONs) to increase their scalability, some use cases do not map well to a tree-based hierarchy. We provide the concept of intralayer communication to improve this situation, where nodes in a specific hierarchy layer may exchange messages directly with each other. This concept targets data preprocessing that allows tool developers to avoid load imbalances in higher hierarchy levels. We implement intralayer communication within the Generic Tools Infrastructure (GTI) that provides TBON services, as well as a high-level abstraction to ease the creation of scalable runtime tools. An extension of GTI's abstractions allows simple and efficient use of intralayer communication. We demonstrate this capability with a runtime message matching tool for MPI's point-to-point communication, which we evaluate in an application study with up to 16,384 processes. Low overheads for two benchmark suites show the applicability of our approach, while a stress test demonstrates close to constant overheads across scales. The stress test measurements demonstrate that intralayer communication reduces application slowdown by two orders of magnitude at 2,048 processes, compared to a previous TBON-based implementation. Tobias Hilbrich, Joachim Jenke, Bronis R. de Supinski, Martin Schulz 0001, Matthias S. Müller, Wolfgang E. Nagel |
ICPP | 5 |
| 2013 | Runtime MPI collective checking with tree-based overlay networksabstractRuntime error detection tools detect many classes of MPI usage errors, including errors in collective communication calls. However, they often face scalability challenges. We present runtime checks for MPI collective operations that use a Tree-Based Overlay Network (TBON) for scalability and that provide full datatype matching. While we can use transitive correctness properties for most checks, some collective operations impose non-transitive correctness properties, e.g., MPI_Alltoallv, where we use an intralayer communication within the TBON to distribute datatype matching information. An overhead study with stress tests and two benchmark suites demonstrates applicability and scalability at 4,096, 2,048 and 16,384 processes respectively. Tobias Hilbrich, Bronis R. de Supinski, Fabian Hänsel, Matthias S. Müller, Martin Schulz 0001, Wolfgang E. Nagel |
EuroMPI | 4 |
| 2013 | Distributed wait state tracking for runtime MPI deadlock detectionabstractThe widely used Message Passing Interface (MPI) with its multitude of communication functions is prone to usage errors. Runtime error detection tools aid in the removal of these errors. We develop MUST as one such tool that provides a wide variety of automatic correctness checks. Its correctness checks can be run in a distributed mode, except for its deadlock detection. This limitation applies to a wide range of tools that either use centralized detection algorithms or a timeout approach. In order to provide scalable and distributed deadlock detection with detailed insight into deadlock situations, we propose a model for MPI blocking conditions that we use to formulate a distributed algorithm. This algorithm implements scalable MPI deadlock detection in MUST. Stress tests at up to 4,096 processes demonstrate the scalability of our approach. Finally, overhead results for a complex benchmark suite demonstrate an average runtime increase of 34% at 2,048 processes. Tobias Hilbrich, Bronis R. de Supinski, Wolfgang E. Nagel, Joachim Jenke, Christel Baier, Matthias S. Müller |
SC | 6 |
| 2013 | Performance and quality of service of data and video movement over a 100 Gbps testbedabstractDigital instruments and simulations are creating an ever-increasing amount of data. The need for institutions to acquire these data and transfer them for analysis, visualization, and archiving is growing as well. In parallel, networking technology is evolving, but at a much slower rate than our ability to create and store data. Single fiber 100 Gbps networking solutions have recently been deployed as national infrastructure. This article describes our experiences with data movement and video conferencing across a networking testbed, using the first commercially available single fiber 100 Gbps technology. The testbed is unique in its ability to be configured for a total length of 60, 200, or 400 km, allowing for tests with varying network latency. We performed low-level TCP tests and were able to use more than 99.9% of the theoretical available bandwidth with minimal tuning efforts. We used the Lustre file system to simulate how end users would interact with a remote file system over such a high performance link. We were able to use 94.4% of the theoretical available bandwidth with a standard file system benchmark, essentially saturating the wide area network. Finally, we performed tests with H.323 video conferencing hardware and quality of service (QoS) settings, showing that the link can reliably carry a full high-definition stream. Overall, we demonstrated the practicality of 100 Gbps networking and Lustre as excellent tools for data management. Michael Kluge, Stephen C. Simms, Thomas William, Robert Henschel, Andy Georgi, Christian Meyer 0004, Matthias S. Müller, Craig A. Stewart, Wolfgang Wünsch, Wolfgang E. Nagel |
Future Gener. Comput. Syst. | 7 |
| 2012 | GTI: A Generic Tools Infrastructure for Event-Based Tools in Parallel SystemsabstractRuntime detection of semantic errors in MPI applications supports efficient and correct large-scale application development. However, current approaches scale to at most one thousand processes and design limitations prevent increased scalability. The need for global knowledge for analyses such as type matching, and deadlock detection presents a major challenge. We present a scalable tool infrastructure - the Generic Tool Infrastructure (GTI) - that we will use to implement MPI runtime error detection tools and that applies to other use cases. GTI supports simple offloading of tool processing onto extra processes or threads and provides a tree based overlay network (TBON) for creating scalable tools that analyze global knowledge. We present its abstractions and code generation facilities that ease many hurdles in tool development, including wrapper generation, tool communication, trace reductions, and filters. GTI ultimately allows tool developers to focus on implementing tool functionality instead of the surrounding infrastructure. Further, we demonstrate that GTI supports scalable tool development through a lost message detector and a phase profiler. The former provides a more scalable implementation of important base functionality for MPI correctness checking, while the latter tool demonstrates that GTI can serve as the basis of further types of tools. Experiments with up to 2048 cores show that GTI's scalability features apply to both tools. Tobias Hilbrich, Matthias S. Müller, Bronis R. de Supinski, Martin Schulz 0001, Wolfgang E. Nagel |
IPDPS | 2 |
| 2012 | Holistic Debugging of MPI Derived DatatypesabstractThe Message Passing Interface (MPI) specifies an API that allows programmers to create efficient and scalable parallel applications. The standard defines multiple constraints for each function parameter. For performance reasons, no MPI implementation checks all of these constraints at runtime. Derived data types are an important concept of MPI and allow users to describe an application's data structures for efficient and convenient communication. Using existing infrastructure we present scalable algorithms to detect usage errors of basic and derived MPI data types. We detect errors that include constraints for construction and usage of derived data types, matching their type signatures in communication, and detecting erroneous overlaps of communication buffers. We implement these checks in the MUST runtime error detection framework. We provide a novel representation of error locations to highlight usage errors. Further, approaches to buffer overlap checking can cause unacceptable overheads for non-contiguous data types. We present an algorithm that uses patterns in derived MPI data types to avoid these overheads without losing precision. Application results for the benchmark suites SPEC MPI2007 and NAS Parallel Benchmarks for up to 2048 cores show that our approach applies to a broad range of applications and that our extended overlap check improves performance by two orders of magnitude. Finally, we augment our runtime error detection component with a debugger extension to support in-depth analysis of the errors that we find as well as semantic errors. This extension to gdb provides information about MPI data type handles and enables gdb -- and other debuggers based on gdb -- to display the content of a buffer as used in MPI communications. Joachim Jenke, Tobias Hilbrich, Andreas Knüpfer, Bronis R. de Supinski, Matthias S. Müller |
IPDPS | 5 |
| 2012 | MPI runtime error detection with MUST: advances in deadlock detectionabstractThe widely used Message Passing Interface (MPI) is complex and rich. As a result, application developers require automated tools to avoid and to detect MPI programming errors. We present the Marmot Umpire Scalable Tool (MUST) that detects such errors with significantly increased scalability. We present improvements to our graph-based deadlock detection approach for MPI, which cover future MPI extensions. Our enhancements also check complex MPI constructs that no previous graph-based detection approach handled correctly. Finally, we present optimizations for the processing of MPI operations that reduce runtime deadlock detection overheads. Existing approaches often require O(p) analysis time per MPI operation, for p processes. We empirically observe that our improvements lead to sub-linear or better analysis time per operation for a wide range of real world applications. Tobias Hilbrich, Joachim Jenke, Martin Schulz 0001, Bronis R. de Supinski, Matthias S. Müller |
SC | 5 |
| 2011 | Memory Performance and SPEC OpenMP Scalability on Quad-Socket x86_64 Systems
Daniel Molka, Robert Schöne, Daniel Hackenberg, Matthias S. Müller |
ICA3PP (1) | 4 |
| 2011 | Order Preserving Event Aggregation in TBONs
Tobias Hilbrich, Matthias S. Müller, Martin Schulz 0001, Bronis R. de Supinski |
EuroMPI | 2 |
| 2010 | SPEC MPI2007 - an application benchmark suite for parallel systems using MPIabstractAbstract The SPEC High‐Performance Group has developed the benchmark suite SPEC MPI2007 and its run rules over the last few years. The purpose of the SPEC MPI2007 benchmark and its run rules is to further the cause of fair and objective benchmarking of high‐performance computing systems. The rules help to ensure that the published results are meaningful, comparable to other results, and reproducible. MPI2007 includes 13 technical computing applications from the fields of computational fluid dynamics, molecular dynamics, electromagnetism, geophysics, ray tracing, and hydrodynamics. We describe the benchmark suite, and compare it with other benchmark suites. Copyright © 2009 John Wiley & Sons, Ltd. Matthias S. Müller, Matthijs van Waveren, Ron Lieberman, Brian Whitney, Hideki Saito 0001, Kalyan Kumaran, John Baron, William C. Brantley, Chris Parrott, Tom Elken, Huiyu Feng, Carl Ponder |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | PrefaceabstractWe are pleased to introduce the special issue of the best papers presented at the Scientific Day of the International Supercomputing Conference (ISC), during June 2008 in Dresden, Germany. ISC always has been an important platform for academia and industry to exchange ideas and to shape the future of HPC. Started in 1986 by Prof. Hans Werner Meuer in Mannheim, Germany, this conference has traditionally featured a diverse set of invited presentations. In 2008, the Scientific Day again offered supercomputing experts the opportunity to submit papers to be presented at the conference through a peer review process. A total of 40 papers were submitted for consideration and reviewed by leading experts in supercomputing and high-performance computing. A total of 18 presentations were given at the Scientific Day in Dresden. Six papers were accepted for publication in this special issue of Concurrency and Computation: Practice and Experience. As is typical of the ISC itself, the paper contributions cover a wide variety of important aspects of high-performance computing. The paper by Bouteiller, Bosilica, and Dongarra addresses the issue of fault tolerance by introducing a novel approach to message logging and implementing it in Open MPI. With this implementation, the authors are able to demonstrate significant performance improvements compared with other solutions. Performance measurement and analysis tools on large-scale systems are treated by Mohr, Wylie, and Wolf. The paper includes a comparison of tools scaling beyond 8000 processes and measurements using their own tool with up to 65 536 processes. This paper was selected as one of the two ISC'08 award winning papers. Performance portability is the topic of two papers. The first by Gabriel et al. describes an abstract data and communication library that provides runtime optimization of application-level communication patterns frequently found in scientific applications. The second by Turek et al. received the PRACE award and is focused on a finite element-based solver toolkit that executes efficiently on commodity clusters, vector supercomputers, and GPGPUs. Improvements for I/O are discussed by Balaji et al. Here, an environment for distributed I/O is presented and applied to gene sequence analysis with mpiBLAST. The analysis was distributed over nine sites and involved more than a petabyte of data. The paper was selected for the second ISC'08 award. The simulation of the strength and stiffness of human bones is covered by Bekas et al. The computation requires to scale up to thousands of processors; the paper contains a detailed performance and scaling study. We would like to thank the authors for contributing papers on their research in high-performance computing to this special publication, and all the reviewers for providing constructive reviews and helping to shape the contributions Finally, we would like to thank the co-editors of this journal—Professors Geoffrey Fox (Pervasive Technology Institute and School of Informatics, Indiana University) and Luc Moreau (School of Electronics and Computer Science, University of Southampton) for providing us the opportunity to create this special issue. With this special issue, we publish some of the best papers from the ISC Scientific Day, and we hope to contribute to the field of high-performance computing and aid the continued success of this aspect of the conference. Wolfgang E. Nagel, Matthias S. Müller |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Implementation, performance, and science results from a 30.7 TFLOPS IBM BladeCenter clusterabstractAbstract This paper describes Indiana University's implementation, performance testing, and use of a large high performance computing system. IU's Big Red, a 20.48 TFLOPS IBM e1350 BladeCenter cluster, appeared in the 27th Top500 list as the 23rd fastest supercomputer in the world in June 2006. In spring 2007, this computer was upgraded to 30.72 TFLOPS. The e1350 BladeCenter architecture, including two internal networks accessible to users and user applications and two networks used exclusively for system management, has enabled the system to provide good scalability on many important applications while being well manageable. Implementing a system based on the JS21 Blade and PowerPC 970MP processor within the US TeraGrid presented certain challenges, given that Intel‐compatible processors dominate the TeraGrid. However, the particular characteristics of the PowerPC have enabled it to be highly popular among certain application communities, particularly users of molecular dynamics and weather forecasting codes. A critical aspect of Big Red's implementation has been a focus on Science Gateways, which provide graphical interfaces to systems supporting end‐to‐end scientific workflows. Several Science Gateways have been implemented that access Big Red as a computational resource—some via the TeraGrid, some not affiliated with the TeraGrid. In summary, Big Red has been successfully integrated with the TeraGrid, and is used by many researchers locally at IU via grids and Science Gateways. It has been a success in terms of enabling scientific discoveries at IU and, via the TeraGrid, across the US. Copyright © 2009 John Wiley & Sons, Ltd. Craig A. Stewart, Matthew R. Link, D. Scott McCaulay, Greg Rodgers, George W. Turner, David Y. Hancock, Faisal Saied, Marlon E. Pierce, Ross Aiken, Matthias S. Müller, Matthias Jurenz, Matthias Lieber, Jenett Tillotson, Beth Plale |
Concurr. Comput. Pract. Exp. | 11 |
| 2009 | Memory Performance and Cache Coherency Effects on an Intel Nehalem Multiprocessor SystemabstractToday's microprocessors have complex memory subsystems with several cache levels. The efficient use of this memory hierarchy is crucial to gain optimal performance, especially on multicore processors. Unfortunately, many implementation details of these processors are not publicly available. In this paper we present such fundamental details of the newly introduced Intel Nehalem microarchitecture with its integrated memory controller, quick path interconnect, and ccNUMA architecture. Our analysis is based on sophisticated benchmarks to measure the latency and bandwidth between different locations in the memory subsystem. Special care is taken to control the coherency state of the data to gain insight into performance relevant implementation details of the cache coherency protocol. Based on these benchmarks we present undocumented performance data and architectural properties. Daniel Molka, Daniel Hackenberg, Robert Schöne, Matthias S. Müller |
PACT | 4 |
| 2009 | Pattern Matching and I/O Replay for POSIX I/O in Parallel Programs
Michael Kluge, Andreas Knüpfer, Matthias S. Müller, Wolfgang E. Nagel |
Euro-Par | 3 |
| 2009 | A graph based approach for MPI deadlock detectionabstractThe MPI standard defines several usage patterns that can lead to deadlock, some of which involve collective communications or non-deterministic operations such as wildcard receives. Further, some MPI programming deadlocks only occur for some MPI implementations or certain configurations. Many tools to detect MPI deadlocks exist; however, none precisely handles the increased complexity of deadlock detection created by the richness of the MPI standard, which requires a general deadlock model. Tobias Hilbrich, Bronis R. de Supinski, Martin Schulz 0001, Matthias S. Müller |
ICS | 4 |
| 2008 | Performance evaluation of supercomputers using HPCC and IMB Benchmarks
Subhash Saini, Robert Ciotti, Brian T. N. Gunney, Thomas E. Spelce, Alice E. Koniges, Don Dossa, Panagiotis A. Adamidis, Rolf Rabenseifner, Sunil Reddy Tiyyagura, Matthias S. Müller |
J. Comput. Syst. Sci. | 10 |
| 2007 | I/O Induced Scalability Limits of Bioinformatics ApplicationsabstractThe growing size of sequence, protein and other biological databases results in an increased computational complexity of the analysis process. Often parallelization is the only solution to limit the turnaround time within reasonable limits. Most scalability studies focus on the parallel algorithm and the resulting communication and synchronization patterns of the implementations. In this paper we examine to what extend I/O bottlenecks limit the scalability on current and future architectures. We study the behavior of two different bioinformatics applications (THREADER, HMMER) and show that these applications are representatives of two different classes with distinct I/O profiles and demands. Robert Henschel, Matthias S. Müller |
BIBE | 2 |
| 2007 | Quality Assurance for Clusters: Acceptance-, Stress-, and Burn-In Tests for General Purpose Clusters
Matthias S. Müller, Guido Juckeland, Matthias Jurenz, Michael Kluge |
HPCC | 1 |
| 2006 | Performance evaluation of supercomputers using HPCC and IMB benchmarksabstractThe HPC Challenge (HPCC) benchmark suite and the Intel MPI Benchmark (IMB) are used to compare and evaluate the combined performance of processor, memory subsystem and interconnect fabric of five leading supercomputers - SGI Altix BX2, Cray XI, Cray Opteron Cluster, Dell Xeon cluster, and NEC SX-8. These five systems use five different networks (SGI NUMALINK4, Cray network, Myrinet, InfiniBand, and NEC IXS). The complete set of HPCC benchmarks are run on each of these systems. Additionally, we present Intel MPI Benchmarks (IMB) results to study the performance of 11 MPI communication functions on these systems Subhash Saini, Robert Ciotti, Brian T. N. Gunney, Thomas E. Spelce, Alice E. Koniges, Don Dossa, Panagiotis A. Adamidis, Rolf Rabenseifner, Sunil Reddy Tiyyagura, Matthias S. Müller, Rod A. Fatoohi |
IPDPS | 10 |
| 2003 | Grid enabled MPI solutions for ClustersabstractDistributing an application onto several machines is one of the key aspects of Grid-computing. In the last few years several groups have developed solutions for the occurring communication problems. However, the focus on the machines used in distributed environments has changed over the time from Massively Parallel Processing Systems to Clusters. This paper presents the problems arising when coupling several cluster like systems and discusses possible solutions. Furthermore, we present performance results and performance drawbacks of the solutions discussed before. Matthias S. Müller, Matthias Hess, Edgar Gabriel |
CCGRID | 1 |
| 2003 | Towards Efficient Execution of MPI Applications on the Grid: Porting and Optimization Issues
Rainer Keller, Edgar Gabriel, Bettina Krammer, Matthias S. Müller, Michael M. Resch |
J. Grid Comput. | 4 |
| 2002 | A software development environment for Grid computingabstractAbstract Grid computing has become a popular concept in the last few years. While in the beginning the driving force was metacomputing, the focus has now shifted towards resource management issues and concepts like ubiquitous computing. For the High‐Performance Computing Center Stuttgart (HLRS) the key challenges of Grid computing have come from the demands of its users and customers. With high‐speed networks in place, programmers expect to be able to exploit the overall performance of several instruments and high‐speed systems for their applications. In order to meet these demands, HLRS has set out a research effort to provide these users with the necessary tools to develop and run their codes on clusters of supercomputers. This has resulted in the development of a basic Grid‐computing environment for technical and scientific computing. In this paper we describe the building blocks of this software development environment and focus specifically on communication and debugging. We present the Grid‐enabled MPI implementation PACX‐MPI and the MPI debugger MARMOT. Copyright © 2002 John Wiley & Sons, Ltd. Matthias S. Müller, Edgar Gabriel, Michael M. Resch |
Concurr. Comput. Pract. Exp. | 1 |
| 2001 | Metacomputing across intercontinental networks
Stephen Pickles, John M. Brooke, Fumie Costen, Edgar Gabriel, Matthias S. Müller, Michael M. Resch, Stephen M. Ord |
Future Gener. Comput. Syst. | 5 |