EDBT 2026 Demo / reviewers in the wild / expert
Brian J. N. Wylie
dblp:w/BrianJNWylie
· DBLP profile ↗
17ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0003-2770-2443ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reproducibility Report for SC25 Paper CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU ClustersabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU Clusters by James D. Trotter et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibilty Committee. Brian J. N. Wylie |
SC | 1 |
| 2025 | 15+ years of joint parallel application performance analysis/tools training with Scalasca/Score-P and Paraver/Extrae toolsetsabstractThe diverse landscape of distributed heterogeneous computer systems currently available and being created to address computational challenges with the highest performance requirements presents daunting complexity for application developers. They must effectively decompose and distribute their application functionality and data, efficiently orchestrating the associated communication and synchronisation, on multi/manycore CPU processors with multiple attached acceleration devices structured within compute nodes with interconnection networks of various topologies. Sophisticated compilers, runtime systems and libraries are (loosely) matched with debugging, performance measurement and analysis tools, with proprietary versions by integrators/vendors provided exclusively for their systems complemented by portable (primarily) open-source equivalents developed and supported by the international research community over many years. The Scalasca and Paraver toolsets are two widely employed examples of the latter, installed on personal notebook computers through to the largest leadership HPC systems. Over more than fifteen years their developers have worked closely together in numerous collaborative projects culminating in the creation of a universal parallel performance assessment and optimisation methodology focused on application execution efficiency and scalability, and the associated training and coaching of application developers (often in teams) in its productive use, reviewed in this article with lessons learnt therefrom. Brian J. N. Wylie, Judit Giménez, Christian Feld, Markus Geimer, Germán Llort, Sandra Méndez, Estanislao Mercadal, Anke Visser, Marta García-Gasulla |
Future Gener. Comput. Syst. | 1 |
| 2022 | Routing brain traffic through the von Neumann bottleneck: Efficient cache usage in spiking neural network simulation code on general purpose computersabstractSimulation is a third pillar next to experiment and theory in the study of complex dynamic systems such as biological neural networks. Contemporary brain-scale networks correspond to directed random graphs of a few million nodes, each with an in-degree and out-degree of several thousands of edges, where nodes and edges correspond to the fundamental biological units, neurons and synapses, respectively. The activity in neuronal networks is also sparse. Each neuron occasionally transmits a brief signal, called spike, via its outgoing synapses to the corresponding target neurons. In distributed computing these targets are scattered across thousands of parallel processes. The spatial and temporal sparsity represents an inherent bottleneck for simulations on conventional computers: irregular memory-access patterns cause poor cache utilization. Using an established neuronal network simulation code as a reference implementation, we investigate how common techniques to recover cache performance such as software-induced prefetching and software pipelining can benefit a real-world application. The algorithmic changes reduce simulation time by up to 50%. The study exemplifies that many-core systems assigned with an intrinsically parallel computational problem can alleviate the von Neumann bottleneck of conventional computer architectures. Jari Pronold, Jakob Jordan, Brian J. N. Wylie, Itaru Kitayama, Markus Diesmann, Susanne Kunkel |
Parallel Comput. | 3 |
| 2012 | Hands-on Practical Hybrid Parallel Application Performance Engineering
Markus Geimer, Michael Gerndt, Sameer Shende, Bert Wesarg, Brian J. N. Wylie |
EuroMPI | 5 |
| 2011 | Introduction
Shirley Moore, Derrick Kondo, Brian J. N. Wylie, Giuliano Casale |
Euro-Par (1) | 3 |
| 2011 | Reconciling Sampling and Direct Instrumentation for Unintrusive Call-Path Profiling of MPI ProgramsabstractWe can profile the performance behavior of parallel programs at the level of individual call paths through sampling or direct instrumentation. While we can easily control measurement dilation by adjusting the sampling frequency, the statistical nature of sampling and the difficulty of accessing the parameters of sampled events make it unsuitable for obtaining certain communication metrics, such as the size of message payloads. Alternatively, direct instrumentation, which is preferable for capturing message-passing events, can excessively dilate measurements, particularly for C++ programs, which often have many short but frequently called class member functions. Thus, we combine these techniques in a unified framework that exploits the strengths of each approach while avoiding their weaknesses: We use direct instrumentation to intercept MPI routines while we record the execution of the remaining code through low-overhead sampling. One of the main technical hurdles mastered was the inexpensive and portable determination of call-path information during the invocation of MPI routines. We show that the overhead of our implementation is sufficiently low to support substantial performance improvement of a C++ fluid-dynamics code. Zoltán Szebenyi, Todd Gamblin, Martin Schulz 0001, Bronis R. de Supinski, Felix Wolf 0001, Brian J. N. Wylie |
IPDPS | 6 |
| 2011 | Scaling Performance Tool MPI Communicator Management
Markus Geimer, Marc-André Hermanns, Christian Siebert, Felix Wolf 0001, Brian J. N. Wylie |
EuroMPI | 5 |
| 2010 | The Scalasca performance toolset architectureabstractAbstract Scalasca is a performance toolset that has been specifically designed to analyze parallel application execution behavior on large‐scale systems with many thousands of processors. It offers an incremental performance‐analysis procedure that integrates runtime summaries with in‐depth studies of concurrent behavior via event tracing, adopting a strategy of successively refined measurement configurations. Distinctive features are its ability to identify wait states in applications with very large numbers of processes and to combine these with efficiently summarized local measurements. In this article, we review the current toolset architecture, emphasizing its scalable design and the role of the different components in transforming raw measurement data into knowledge of application execution behavior. The scalability and effectiveness of Scalasca are then surveyed from experience measuring and analyzing real‐world applications on a range of computer systems. Copyright © 2010 John Wiley & Sons, Ltd. Markus Geimer, Felix Wolf 0001, Brian J. N. Wylie, Erika Ábrahám, Daniel Becker 0001, Bernd Mohr |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | Performance measurement and analysis tools for extremely scalable systemsabstractAbstract High‐performance computing systems continue to employ more and more processor cores. Current typical high‐end machines in industry, university, and government research laboratory computing centers feature thousands of computing cores. While these machines promise ever more compute power and memory capacity to tackle today's complex simulation problems, they force application developers to greatly enhance the scalability of their codes to be able to exploit it. To better support them in their porting and tuning process, many parallel‐tools research groups have already started to work on scaling their methods, techniques, and tools to extreme processor counts. In this paper, we survey existing profiling and tracing tools, report on our experience in using them in extreme scaling environments, review working and promising new methods and techniques, and discuss strategies for solving open issues and problems. Copyright © 2010 John Wiley & Sons, Ltd. Bernd Mohr, Brian J. N. Wylie, Felix Wolf 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Verifying Causality between Distant Performance Phenomena in Large-Scale MPI ApplicationsabstractIn message-passing applications, the temporal or spatial distance between cause and symptom of a performance problem constitutes a major difficulty in deriving helpful conclusions from performance data. Just knowing the locations of wait states in the program is often insufficient to understand the reason for their occurrence. We present a method for verifying hypotheses on causality between temporally or spatially distant performance phenomena in message-passing applications without altering the application itself. The verification is accomplished by modifying MPI event traces and using them to simulate the hypothetical message-passing behavior. By performing a parallel real-time reenactment of the communication to be simulated using the original execution configuration, we can achieve high scalability and good predictive accuracy in relation to the measured behavior. Not relying on a potentially complex model of the message-passing subsystem, our method is also platform independent. Marc-André Hermanns, Markus Geimer, Felix Wolf 0001, Brian J. N. Wylie |
PDP | 4 |
| 2009 | Space-efficient time-series call-path profiling of parallel applicationsabstractThe performance behavior of parallel simulations often changes considerably as the simulation progresses --- with potentially process-dependent variations of temporal patterns. While call-path profiling is an established method of linking a performance problem to the context in which it occurs, call paths reveal only little information about the temporal evolution of performance phenomena. However, generating call-path profiles separately for thousands of iterations may exceed available buffer space --- especially when the call tree is large and more than one metric is collected. In this paper, we present a runtime approach for the semantic compression of call-path profiles based on incremental clustering of a series of single-iteration profiles that scales in terms of the number of iterations without sacrificing important performance details. Our approach offers low runtime overhead by using only a condensed version of the profile data when calculating distances and accounts for process-dependent variations by making all clustering decisions locally. Zoltán Szebenyi, Felix Wolf 0001, Brian J. N. Wylie |
SC | 3 |
| 2009 | A scalable tool architecture for diagnosing wait states in massively parallel applications
Markus Geimer, Felix Wolf 0001, Brian J. N. Wylie, Bernd Mohr |
Parallel Comput. | 3 |
| 2007 | Automatic Trace-Based Performance Analysis of Metacomputing ApplicationsabstractThe processing power and memory capacity of independent and heterogeneous parallel machines can be combined to form a single parallel system that is more powerful than any of its constituents. However, achieving satisfactory application performance on such a metacomputer is hard because the high latency of inter-machine communication as well as differences in hardware of constituent machines may introduce various types of wait states. In our earlier work, we have demonstrated that automatic pattern search in event traces can identify the sources of wait states in parallel applications running on a single computer. In this article, we describe how this approach can be extended to metacomputing environments with special emphasis on performance problems related to inter-machine communication. In addition, we demonstrate the benefits of our solution using a real-world multi-physics application. Daniel Becker 0001, Felix Wolf 0001, Wolfgang Frings, Markus Geimer, Brian J. N. Wylie, Bernd Mohr |
IPDPS | 5 |
| 2003 | Memory Profiling using Hardware CountersabstractAlthough memory performance is often a limiting factor in application performance, most tools only show performance data relating to the instructions in the program, not to its data. In this paper, we describe a technique for directly measuring the memory profile of an application. We describe the tools and their user model, and then discuss a particular code, the MCFbenchmark from SPEC CPU 2000. We show performance data for the data structures and elements, and discuss the use of the data to improve program performance. Finally, we discuss extensions to the work to provide feedback to the compiler for prefetching and to generate additional reports from the data. Marty Itzkowitz, Brian J. N. Wylie, Christopher Aoki, Nicolai Kosche |
SC | 2 |
| 2002 | A callgraph-based search strategy for automated performance diagnosisabstractAbstract We introduce a new technique for automated performance diagnosis, using the program's callgraph. We discuss our implementation of this diagnosis technique in the Paradyn Performance Consultant. Our implementation includes the new search strategy and new dynamic instrumentation to resolve pointer‐based dynamic call sites at run‐time. We compare the effectiveness of our new technique to the previous version of the Performance Consultant for several sequential and parallel applications. Our results show that the new search method performs its search while inserting dramatically less instrumentation into the application, resulting in reduced application perturbation and consequently a higher degree of diagnosis accuracy. Copyright © 2002 John Wiley & Sons, Ltd. Harold W. Cain, Barton P. Miller, Brian J. N. Wylie |
Concurr. Comput. Pract. Exp. | 3 |
| 2000 | A Callgraph-Based Search Strategy for Automated Performance Diagnosis (Distinguished Paper)
Harold W. Cain, Barton P. Miller, Brian J. N. Wylie |
Euro-Par | 3 |
| 1994 | PARAMICS - moving vehicles on the connection machineabstractPARAMICS is a PARAllel MICroscopic Traffic Simulator which is, to our knowledge, the most powerful of its type in the world. The simulator can model around 200,000 vehicles on around 7,000 roads (taken from real road traffic network data) at faster than 'real-time' rates, making use of 16 K processor TMC Connection Machine CM-200 for the simulation aspect. The project aims to make available to road network planners a new range of tools, and demonstrates that use of high performance computing in real applications is possible and worthwhile, while yielding important and interesting research results.> Gordon Cameron, Brian J. N. Wylie, David McArthur |
SC | 2 |