VLDB 2026 Research / reviewers in the wild / expert
Allen D. Malony
dblp:35/4761
· DBLP profile ↗
125ranked-venue papers
24as first author
12since 2021 · last 2024
0000-0002-9598-7201ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 103 · 23 first-author · 10 since 2021Software engineering, systems software and programming languages · 11 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enabling Performance Observability for Heterogeneous HPC Workflows with SOMAabstractHeterogeneous workflows represent a promising approach for overcoming traditional application performance limitations and to accelerate scientific insight on high-performance computing (HPC) platforms. As HPC platforms grow in size and complexity, managing and optimizing workflow resources while maximizing scientific output assumes vital importance. Optimal workflow resource allocation requires high-quality and timely information about the state of the hardware resources, the status of the pending tasks, the performance of the tasks that have already been executed, and the current status of the workflow itself. A robust performance observability framework that captures and delivers this information can fundamentally improve the quality of decision-making within the workflow system, setting the stage for the adaptive execution of workflow tasks. We propose the use of SOMA, a service-based performance observability framework for such HPC workflows. With the RADICAL-Pilot runtime system as a development vehicle, SOMA demonstrates that service-based architectures coupled with an appropriate data model can serve the performance monitoring needs of large-scale ensemble workflows in a low-overhead fashion. Effective observability of workflow performance requires exporting, storing, and analyzing several types of performance data from across the application and workflow software stacks. Our study finds significant benefits in integrating observability frameworks as first-class citizens within an HPC workflow software stack. In this paper, we demonstrate how SOMA can simultaneously observe the performance states of the individual tasks, system hardware, and the workflow as a whole. Such information can then be employed to calculate better resource allocation and task configuration. Dewi Yokelson, Mikhail Titov, Srinivasan Ramesh, Ozgur O. Kilic, Matteo Turilli, Shantenu Jha, Allen D. Malony |
ICPP | 7 |
| 2024 | SOMA: Observability, monitoring, and in situ analytics for exascale applicationsabstractSummary With the rise of exascale systems and large, data‐centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service‐based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application‐specific diagnostic data and performance data addresses these needs. We present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi‐GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring. Dewi Yokelson, Oskar Lappi, Srinivasan Ramesh, Miikka S. Väisälä, Kevin A. Huck, Touko Puro, Boyana Norris, Maarit J. Korpi-Lagg, Keijo Heljanko, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 10 |
| 2022 | HPC Storage Service Autotuning Using Variational- Autoencoder -Guided Asynchronous Bayesian OptimizationabstractDistributed data storage services tailored to specific applications have grown popular in the high-performance computing (HPC) community as a way to address I/O and storage challenges. These services offer a variety of specific interfaces, semantics, and data representations. They also expose many tuning parameters, making it difficult for their users to find the best configuration for a given workload and platform. To address this issue, we develop a novel variational-autoencoder-guided asynchronous Bayesian optimization method to tune HPC storage service parameters. Our approach uses transfer learning to leverage prior tuning results and use a dynamically updated surrogate model to explore the large parameter search space in a systematic way. We implement our approach within the DeepHyper open-source framework, and apply it to the autotuning of a high-energy physics workflow on Argonne's Theta supercomputer. We show that our transfer-learning approach enables a more than 40 x search speedup over random search, compared with a 2.5 x to 10 x speedup when not using transfer learning. Additionally, we show that our approach is on par with state-of-the-art autotuning frameworks in speed and outperforms them in resource utilization and parallelization capabilities. Matthieu Dorier, Romain Egele, Prasanna Balaprakash, Jaehoon Koo, Sandeep Madireddy, Srinivasan Ramesh, Allen D. Malony, Robert B. Ross |
CLUSTER | 7 |
| 2022 | The Ghost of Performance Reproducibility PastabstractThe importance of ensemble computing is well established. However, executing ensembles at scale introduces interesting performance fluctuations that have not been well investigated. In this paper, we trace our experience uncovering performance fluctuations of ensemble applications (primarily constituting a workflow of GROMACS tasks), and unsuccessful attempts, so far, at trying to discern the underlying cause(s) of performance fluctuations. Is the failure to discern the causative or contributing factors a failure of capability? Or imagination? Do the fluctuations have their genesis in some inscrutable aspect of the system or software? Does it warrant a fundamental reassessment and rethinking of how we assume and conceptualize performance reproducibility? Answers to these questions are not straightforward, nor are they immediate or obvious. We conclude with a discussion about the performance of ensemble applications and ruminate over the implications for how we define and measure application performance. Srinivasan Ramesh, Mikhail Titov, Matteo Turilli, Shantenu Jha, Allen D. Malony |
e-Science | 5 |
| 2022 | MARTINI: The Little Match and Replace Tool for Automatic Application Rewriting with Code Examples
Alister Johnson, Camille Coti, Allen D. Malony, Johannes Doerfert |
Euro-Par | 3 |
| 2022 | Enabling Global MPI Process Addressing in MPI ApplicationsabstractDistributed software using MPI is now facing a complexity barrier. Indeed, given increasing intra-node parallelism, combined with the use of accelerator, programs’ states are becoming more intricate. A given code must cover several cases, generating work for multiple devices. Model mixing generally leads to increasingly large programs and hinders performance portability. In this paper, we pose the question of software composition, trying to split jobs in multiple services. In doing so, we advocate it would be possible to depend on more suitable units while removing the need for extensive runtime stacking (MPI+X+Y). For this purpose, we discuss what MPI shall provide and what is currently available to enable such software composition. After pinpointing (1) process discovery and (2) Remote Procedure Calls (RPCs) as facilitators in such infrastructure, we focus solely on the first aspect. We introduce an overlay-network providing whole-machine inter-job, discovery, and wiring at the level of the MPI runtime. MPI process Unique IDentifiers (UIDs) are then covered as a Unique Resource Locator (URL) leveraged as support for job interaction in MPI, enabling a more horizontal usage of the MPI interface. Eventually, we present performance results for large-scale wiring-up exchanges, demonstrating gains over PMIx in cross-job configurations. Jean-Baptiste Besnard, Sameer Shende, Allen D. Malony, Julien Jaeger, Marc Pérache |
EuroMPI | 3 |
| 2022 | SERVIZ: A Shared In Situ Visualization ServiceabstractInline and in transit visualization are popular in situ visualization models for high performance computing (HPC) applications. Inline visualization is invoked through a library call on the HPC application (simulation), while in transit methods invoke a visualization module running on in transit resources. In transit methods can offer better efficiency than inline by running the visualization at a lower concurrency level than the simulation. State-of-the-art in transit schemes are limited to employing a dedicated in transit resource for every simulation. The resulting idle time on the in transit resource can severely limit the cost savings over inline methods. This research proposes SERVIZ, an in transit visualization service that can be shared amongst multiple simulations to reduce idle time, thereby efficiently using in transit resources. SERVIZ achieves cost savings of up to 26% over inline and up to$4\mathbf{x}$reduction in idle time compared to a dedicated in transit implementation. Srinivasan Ramesh, Hank Childs, Allen D. Malony |
SC | 3 |
| 2021 | Dynamic and Adaptive Monitoring and Analysis for Many-task Ensemble ComputingabstractApplications are not what they used to be. Modern HPC applications are increasingly a mix of heterogeneous tasks and services – both internal and external. This imposes new constraints and requirements on application and performance monitoring, which is fundamentally different from single-task monitoring. Runtime decisions are needed, and critically rely on monitored information to determine what to do. Thus, online monitoring and analytics must be first-class requirements of all applications and the runtime environments that support their execution. This position paper explores the implications of novel application requirements on future monitoring and profiling subsystems, in the context of ensemble computing. Shantenu Jha, Allen D. Malony |
CLUSTER | 2 |
| 2021 | SYMBIOMON: A High-Performance, Composable Monitoring ServiceabstractHigh-performance computing (HPC) software is evolving to support an increasingly diverse set of applications and heterogeneous hardware architectures. As part of this evolution, the construction of scientific software has shifted from a traditional monolithic message passing interface executable model to a coupled, services-style model in which simulations run alongside a host of distributed HPC data services within the same batch job allocation. Microservices have emerged as a powerful new way to build these distributed data services through a composition model. However, performance analysis of composed microservices is a daunting challenge. It requires collecting, monitoring, aggre-gating, and exporting performance data from multiple sources. To be effective, the design of such a monitoring solution must allow for seamless integration into HPC applications and distributed services alike, be scalable, operate with a low overhead, and take advantage of the HPC platform. We propose SYMBIOMON, a monitoring service that is built by composing high-performance microservices. We describe its design and implementation within the context of the Mochi framework. SYMBIOMON combines a time-series data model with existing Mochi data services to collect, aggregate, and export performance metrics in a distributed manner. SYMBIOMON enables seamless, low-overhead monitoring and analysis of data services and HPC applications alike. Using HEPnOS, a production-quality Mochi data service, we demonstrate the use of SYMBIOMON to identify better service configurations. Srinivasan Ramesh, Robert B. Ross, Matthieu Dorier, Allen D. Malony, Philip H. Carns, Kevin A. Huck |
HiPC | 4 |
| 2021 | SYMBIOSYS: A Methodology for Performance Analysis of Composable HPC Data ServicesabstractMicroservices are a powerful new way of building, customizing, and deploying distributed services owing to their flexibility and maintainability. Several large-scale distributed platforms have emerged to serve the growing needs of data-centric workloads and services in commercial computing. Concurrently, high-performance computing (HPC) systems and software are rapidly evolving to meet the demands of diversified applications and heterogeneity. The interplay of hardware factors, software configuration parameters, and the flexibility offered with a microservice architecture makes it nontrivial to estimate the optimal service instantiation for a given application workload. Further, this problem is exacerbated when considering that these services operate in a dynamic and heterogeneous HPC environment. An optimally integrated service can be vastly more performant than a haphazardly integrated one. Existing performance tools for HPC either fail to understand the request-response model of communication inherent to microservices or they operate within a narrow scope, limiting the insight that can be gleaned from employing them in isolation.We propose a methodology for integrated performance analysis of HPC microservices frameworks and applications called SYMBIOSYS. We describe its design and implementation within the context of the Mochi framework. This integration is achieved by combining distributed callpath profiling and tracing with a performance data exchange strategy that collects fine-grained, low-level metrics from the RPC communication library and network layers. The result is a portable, low-overhead performance analysis setup that provides a holistic profile of the dependencies among microservices and how they interact with the Mochi RPC software stack. Using HEPnOS, a production-quality Mochi data service, we demonstrate the low-overhead operation of SYMBIOSYS at scale and use it to identify the root causes of poorly performing service configurations. Srinivasan Ramesh, Allen D. Malony, Philip H. Carns, Robert B. Ross, Matthieu Dorier, Jérome Soumagne, Shane Snyder |
IPDPS | 2 |
| 2021 | DiPOSH: A portable OpenSHMEM implementation for short API-to-network pathabstractSummary In this article, we introduce DiPOSH, a multi‐network, distributed implementation of the OpenSHMEM standard. The core idea behind DiPOSH is to have an API‐to‐network software stack as slim as possible, in order to minimize the software overhead. Following the heritage of its non‐distributed parent POSH, DiPOSH's communication engine is organized around the processes' shared heaps, and remote communications are moving data from and to these shared heaps directly. This article presents its architecture and several communication drivers, including one that takes advantage of a helper process, called the Hub, for inter‐process communications. This architecture allows use to explore different options for implementing the communication drivers, from using high‐level, portable, optimized libraries to low‐level, close to the hardware communication routines. We present the perspectives opened by this additional component in terms of communication scheduling between and on the nodes. DiPOSH is available at https://github.com/coti/DiPOSH . Camille Coti, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | Optimization with the OpenACC-to-FPGA framework on the Arria 10 and Stratix 10 FPGAs
Jacob Lambert 0002, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
Parallel Comput. | 4 |
| 2020 | MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-OffsabstractIntegrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed. Mohammad Alaul Haque Monil, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
PACT | 5 |
| 2020 | On-the-fly Optimization of Parallel Computation of Symbolic Symplectic InvariantsabstractGroup invariants are used in high energy physics to define quantum field theory interactions. In this paper, we present the parallel algebraic computation of special invariants called symplectic and focus on one particular invariant that finds recent interest in physics. Our results will export to other invariants. The cost of performing basic computations on the multivariate polynomials evolves during the computation, as the polynomials get larger and/or have increasing numbers of terms. Interestingly, in some cases, they stay small. Traditionally, high-performance software is optimized by running experiments with sample data sets in order to profile and optimize expected behavior of workloads in practice. Since the (communication and computation) costs depend on the changing behavior of the symplectic invariant calculations, the standard optimization approach is insufficient. Thus, it is necessary to implement online performance tuning methods that can track the algorithm's progress and state, evaluate performance data in situ, and control the parallel resources during execution. Joseph Ben Geloun, Camille Coti, Allen D. Malony |
ISPDC | 3 |
| 2020 | CCAMP: an integrated translation and optimization framework for OpenACC and OpenMPabstractHeterogeneous computing and exploration into specialized accelerators are inevitable in current and future supercomputers. Although this diversity of devices is promising for performance, the array of architectures presents programming challenges. High-level programming strategies have emerged to face these challenges, such as the OpenMP offloading model and OpenACC. However, the varying levels of support for these standards within vendor-specific and open-source tools, as well as the lack of performance portability across devices, have prevented the standards from achieving their goals. To address these shortcomings, we present CCAMP, an OpenMP and OpenACC interoperable framework. CCAMP provides two primary facilities: language translation between the two standards and device-specific directive optimization within each standard. We show that by using the CCAMP framework, programmers can easily transplant non-portable code into new ecosystems for new architectures. Additionally, by using CCAMP's device-specific directive optimizations, users can achieve optimized performance across architectures using a single source code. Jacob Lambert 0002, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
SC | 4 |
| 2019 | Scalable Performance Awareness for In Situ Scientific ApplicationsabstractPart of the promise of exascale computing and the next generation of scientific simulation codes is the ability to bring together time and spatial scales that have traditionally been treated separately. This enables creating complex coupled simulations and in situ analysis pipelines, encompassing such things as "whole device" fusion models or the simulation of cities from sewers to rooftops. Unfortunately, the HPC analysis tools that have been built up over the preceding decades are ill suited to the debugging and performance analysis of such computational ensembles. In this paper, we present a new vision for performance measurement and understanding of HPC codes, MonitoringAnalytics (MONA). MONA is designed to be a flexible, high performance monitoring infrastructure that can perform monitoring analysis in place or in transit by embedding analytics and characterization directly into the data stream, without relying upon delivering all monitoring information to a central database for post-processing. It addresses the trade-offs between the prohibitively expensive capture of all performance characteristics and not capturing enough to detect the features of interest. We demonstrate several uses of MONA; capturing and indexing multi-executable performance profiles to enable later processing, extraction of performance primitives to enable the generation of customizable benchmarks and performance skeletons, and extracting communication and application behaviors to enable better control and placement for the current and future runs of the science ensemble. Relevant performance information based on a system for MONA built from ADIOS and SOSflow technologies is provided for DOE science applications and leadership machines. Matthew Wolf, Julien Dominski, Gabriele Merlo, Jong Choi 0001, Greg Eisenhauer, Stéphane Ethier, Kevin A. Huck, Scott Klasky, Jeremy Logan, Allen D. Malony, Chad Wood |
eScience | 10 |
| 2019 | A Plugin Architecture for the TAU Performance SystemabstractSeveral robust performance systems have been created for parallel machines with the ability to observe diverse aspects of application execution on different hardware platforms. All of these are designed with the objective to support measurement methods that are efficient, portable, and scalable. For these reasons, the performance measurement infrastructure is tightly embedded with the application code and runtime execution environment. As parallel software and systems evolve, especially towards more heterogeneous, asynchronous, and dynamic operation, it is expected that the requirements for performance observation and awareness will change. For instance, heterogeneous machines introduce new types of performance data to capture and performance behaviors to characterize. Furthermore, there is a growing interest in interacting with the performance infrastructure for in situ analytics and policy-based control. The problem is that an existing performance system architecture could be constrained in its ability to evolve to meet these new requirements. The paper reports our research efforts to address this concern in the context of the TAU Performance System. In particular, we consider the use of a powerful plugin model to both capture existing capabilities in TAU and to extend its functionality in ways it was not necessarily conceived originally. The TAU plugin architecture supports three types of plugin paradigms: EVENT, TRIGGER, and AGENT. We demonstrate how each operates under several different scenarios. Results from larger-scale experiments are shown to highlight the fact that efficiency and robustness can be maintained, while new flexibility and programmability can be offered that leverages the power of the core TAU system while allowing significant and compelling extensions to be realized. Allen D. Malony, Srinivasan Ramesh, Kevin A. Huck, Nicholas Chaimov, Sameer Shende |
ICPP | 1 |
| 2019 | Runtime Adaptive Task Inlining on Asynchronous Multitasking Runtime SystemsabstractAs the era of high frequency, single core processors have come to a close, the new paradigm of many core processors has come to dominate. In response to these systems, asynchronous multitasking runtime systems have been developed as a promising solution to efficiently utilize these newly available hardware. Asynchronous multitasking runtime systems work by dividing a problem into a large number of fine grained tasks. However, as the number of tasks created increase, the overheads associated with task creation and management cannot be ignored. Task inlining, a method where the parent thread consumes a child thread, enables the runtime system to achieve the balance between parallelism and its overhead. As largely impacted by different processor architectures, the decision of task inlining is dynamic in nature. In this research, we present adaptive techniques for deciding, at runtime, whether a particular task should be inlined or not. We present two policies, a baseline policy that makes inlining decision based on a fixed threshold and an adaptive policy which decides the threshold dynamically at runtime. We also evaluate and justify the performance of these policies on different processor architectures. To the best of our knowledge, this is the first study of the impacts of adaptive policy at runtime for task inlining in an asynchronous multitasking runtime system on different processor architectures. From experimentation, we find that the baseline policy improves the execution time from 7.61% to 54.09%. Furthermore, the adaptive policy improves over the baseline policy by up to 74%. Bibek Wagle, Mohammad Alaul Haque Monil, Kevin A. Huck, Allen D. Malony, Adrian Serio, Hartmut Kaiser |
ICPP | 4 |
| 2019 | Understanding the Impact of Dynamic Power Capping on Application ProgressabstractElectrical power has become an important design constraint in high-performance computing (HPC) systems. On future HPC machines, power is likely to be a budgeted resource and thus managed dynamically. Power management software needs to reliably measure application performance at runtime in order to respond effectively to changes in application behavior. Execution time tells us little about how the science in the application is progressing toward an application-defined end goal. To the best of our knowledge, no study has defined or categorized online application progress in the context of power management. Based on semi-structured interviews with HPC application-specialists, we define an online notion of progress-an application-specific metric that can be monitored at runtime to provide a sense of the rate at which application science is being performed. Using instrumentation, we characterize and categorize the progress of various production scientific applications and benchmarks. We propose a model of the impact of dynamic power capping on application progress. By experimental evaluation, we show that our model accurately captures the general behavior of the progress of different classes of applications under a power cap. We believe that such a model is an important first step toward the design of more dynamic power management policies for HPC systems. Srinivasan Ramesh, Swann Perarnau, Sridutt Bhalachandra, Allen D. Malony, Pete Beckman |
IPDPS | 4 |
| 2019 | Mixing ranks, tasks, progress and nonblocking collectivesabstractSince the beginning, MPI has defined the rank as an implicit attribute associated with the MPI process' environment. In particular, each MPI process generally runs inside a given UNIX process and is associated with a fixed identifier in its WORLD communicator. However, this state of things is about to change with the rise of new abstractions such as MPI Sessions. In this paper, we propose to outline how such evolution could enable optimizations which were previously linked to specific MPI runtimes executing MPI processes in shared memory (e.g. thread-based MPI). By implementing runtime-level work-sharing through what we define as MPI tasks, enabling the ability to progress indifferently from stream context we show that there is potential for improved asynchronous progress. In the absence of a Session implementation, this assumption is validated in the context of a thread-based MPI where nonblocking Collective (NBC) were implemented on top of Extended Generic Requests progressed by any rank on the node thanks to an MPI extension enabling threads to dynamically share their MPI context. Jean-Baptiste Besnard, Julien Jaeger, Allen D. Malony, Sameer Shende, Hugo Taboada, Marc Pérache, Patrick Carribault |
EuroMPI | 3 |
| 2019 | Checkpoint/restart approaches for a thread-based MPI runtime
Julien Adam, Maxime Kermarquer, Jean-Baptiste Besnard, Leonardo Arturo Bautista-Gomez, Marc Pérache, Patrick Carribault, Julien Jaeger, Allen D. Malony, Sameer Shende |
Parallel Comput. | 8 |
| 2018 | Coupling Exascale Multiphysics Applications: Methods and Lessons LearnedabstractWith the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration. Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood |
eScience | 25 |
| 2018 | Directive-Based, High-Level Programming and Optimizations for High-Performance Computing with FPGAsabstractReconfigurable architectures like Field Programmable Gate Arrays (FPGAs) have been used for accelerating computations from several domains because of their unique combination of flexibility, performance, and power efficiency. However, FPGAs have not been widely used for high-performance computing, primarily because of their programming complexity and difficulties in optimizing performance. In this paper, we present a directive-based, high-level optimization framework for high-performance computing with FPGAs, built on top of an OpenACC-to-FPGA translation framework called OpenARC. We propose directive extensions and corresponding compile-time optimization techniques to enable the compiler to generate more efficient FPGA hardware configuration files. Empirical evaluation of the proposed framework on an Intel Stratix V with five OpenACC benchmarks from various application domains shows that FPGA-specific optimizations can lead to significant increases in performance across all tested applications. We also demonstrate that applying these high-level directive-based optimizations can allow OpenACC applications to perform similarly to lower-level OpenCL applications with hand-written FPGA-specific optimizations, and offer runtime and power performance benefits compared to CPUs and GPUs. Jacob Lambert 0002, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
ICS | 5 |
| 2018 | Stingray-HPC: A Scalable Parallel Seismic Raytracing SystemabstractThe Stingray raytracer was developed for marine seismology to compute minimum travel time from all sources in an earth model to determine the 3D geophysical structure below the ocean floor. The original sequential implementation of Stingray used Dijkstra's single-source, shortest-path (SSSP) algorithm. A data parallel version of Stingray was developed based on the Bellman-Ford-Moore iterative SSSP algorithm. Single node experiments demonstrated performance improvements from parallelization with multicore (using OpenMP) and manycore processors (using CUDA). Calculating seismic ray paths for larger earth models requires distributed, multi-node algorithms utilizing domain decomposition methods. Preliminary 2D decomposition strategies show promising scaling results. However, a general 3D decomposition methodology is needed to handle any seismic raytracing problem on any HPC computing platform. In this paper, we present Stingray-HPC, a framework for scalable seismic raytracing which can automatically decompose a 3D earth model across nodes in a distributed environment, allocate ghost cell regions for iterative updates, coordinate ghost cell communications, and test for global convergence. Stingray-HPC is implemented with MPI and either OpenMP or CUDA for node- level calculations. Our results validate Stingray-HPC's ability to handle large models (over a billion points) and to solve these models efficiently at scale up to 512 GPU nodes. Mohammad Alaul Haque Monil, Allen D. Malony, Douglas Toomey, Kevin A. Huck |
PDP | 2 |
| 2018 | Transparent High-Speed Network Checkpoint/Restart in MPIabstractFault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occur more often on systems gathering millions of computing units. Moreover, the larger jobs are, the more computing hours would be wasted by a crash. In this paper, we describe the work done in our MPI runtime to enable transparent checkpointing mechanism. Unlike the MPI 4.0 User-Level Failure Mitigation (ULFM) interface, our work targets solely Checkpoint/Restart (C/R) and ignores wider features such as resiliency. We show how existing transparent checkpointing methods can be practically applied to MPI implementations given a sufficient collaboration from the MPI runtime. Our C/R technique is then measured on MPI benchmarks such as IMB and Lulesh relying on Infiniband high-speed network, demonstrating that the chosen approach is sufficiently general and that performance is mostly preserved. We argue that enabling fault-tolerance without any modification inside target MPI applications is possible, and show how it could be the first step for more integrated resiliency combined with failure mitigation like ULFM. Julien Adam, Jean-Baptiste Besnard, Allen D. Malony, Sameer Shende, Marc Pérache, Patrick Carribault, Julien Jaeger |
EuroMPI | 3 |
| 2018 | MPI performance engineering with the MPI tool interface: The integration of MVAPICH and TAU
Srinivasan Ramesh, Aurèle Mahéo, Sameer Shende, Allen D. Malony, Hari Subramoni, Amit Ruhela, Dhabaleswar K. Panda 0001 |
Parallel Comput. | 4 |
| 2018 | The Long and Winding Road Toward Efficient High-Performance ComputingabstractThe major challenge to Exaflop computing, and more generally, efficient high-end computing, is in finding the best “matches” between advanced hardware capabilities and the software used to program applications, so that top performance will be achieved. Several benchmarks show very disappointing performance progress over the last decade, clearly indicating a mismatch between hardware and software. To remedy this problem, it is important that key performance enablers at the software level-autotuning, performance analysis tools, full application optimization-are understood. For each area, we highlight major limitations and most promising approaches to reaching better performance and energy levels. Finally, we conclude by analyzing hardware and software design, trying to pave the way for more tightly integrated hardware and software codesign. William Jalby, David J. Kuck, Allen D. Malony, Michel Masella, Abdelhafid Mazouz, Mihail Popov |
Proc. IEEE | 3 |
| 2017 | QoS-Aware Virtual Machine Consolidation in Cloud DatacenterabstractWith the rapid growth of the cloud industry in recent years, energy consumption of warehouse-scale datacenters has become a major concern. Energy-aware Virtual Machine consolidation has proven to be one of the most effective solutions for tackling this problem. Among the sub-problems of VM consolidation, VM placement is the trickiest and can be treated as a bin packing problem which is NP-hard, hence, it is logical to apply a heuristic approach. The main challenge of VM consolidation is to achieve a balance between energy consumption and quality of service (QoS). In this research, we evaluate this problem and design a combined strategy using best fit decreasing bin packing method and multi-pass optimization in VM placement for an efficient VM consolidation. We have used CloudSim toolkit to simulate our experiments. To evaluate the performance of the proposed algorithms we used real-world workload traces from thousand VMs. Results demonstrate that our proposed methods outperform other existing methods. Mohammad Alaul Haque Monil, Allen D. Malony |
IC2E | 2 |
| 2017 | Autotuning GPU Kernels via Static and Predictive AnalysisabstractOptimizing the performance of GPU kernels is challenging for both human programmers and code generators. For example, CUDA programmers must set thread and block parameters for a kernel, but might not have the intuition to make a good choice. Similarly, compilers can generate working code, but may miss tuning opportunities by not targeting GPU models or performing code transformations. Although empirical autotuning addresses some of these challenges, it requires extensive experimentation and search for optimal code variants. This research presents an approach for tuning CUDA kernels based on static analysis that considers fine-grained code structure and the specific GPU architecture features. Notably, our approach does not require any program runs in order to discover near-optimal parameter settings. We demonstrate the applicability of our approach in enabling code autotuners such as Orio to produce competitive code variants comparable with empirical-based methods, without the high cost of experiments. Robert V. Lim, Boyana Norris, Allen D. Malony |
ICPP | 3 |
| 2017 | Performance Analysis of Applications in the Context of Architectural RooflinesabstractIntuitive visual representations of architecture capabilities and the performance of applications are critical to enabling effective performance analysis, which in turn guides optimizations. The Roofline Model and its derivatives provide such an intuitive representation of the best achievable performance on a given architecture. The Roofline Toolkit project is a collaboration among researchers at Argonne National Laboratory, Lawrence Berkeley National Laboratory, and the University of Oregon and consists of three principal components: hardware characterization, software characterization, and data manipulation, which includes a visualization interface. These components address the different aspects of performance data acquisition and manipulation required for performance analysis, modeling and optimization of applications. In this paper we introduce an implementation of the third component, a system for visualizing roofline charts and managing roofline performance analysis data. We demonstrate analysis of an application use case within this framework and outline future directions for this type of performance analysis and visualization. Boyana Norris, Wyatt Spear, Allen D. Malony |
ICPE | 3 |
| 2016 | ARCS: Adaptive Runtime Configuration Selection for Power-Constrained OpenMP ApplicationsabstractPower is the most critical resource for the exascale high performance computing. In the future, system administrators might have to pay attention to the power consumption of the machine under different work loads. Hence, each application may have to run with an allocated power budget. Thus, achieving the best performance on future machines requires optimal performance subject to a power constraint. This additional performance requirement should not be the responsibility of HPC~(High Performance Computing) application developers. Optimizing the performance for a given power budget should be the responsibility of high-performance system software stack. Modern machines allow power capping of CPU and memory to implement power budgeting strategy. Finding the best runtime environment for a node at a given power level is important to get the best performance. This paper presents ARCS (Adaptive Runtime Configuration Selection) frameworkthat automatically selects the best runtime configuration for each OpenMPparallel region at a given power level. The framework uses OMPT (OpenMP Tools) API, APEX(Autonomic Performance Environment for eXascale), and Active Harmony frameworksto explore configuration search space and selects the best number of threads, scheduling policy, and chunk size for a given power level at run-time. We test ARCS using the NAS Parallel Benchmark, and proxy application LULESH with Intel Sandybridge, and IBM Power multi-core architectures. We show that for a given power level, efficient OpenMP runtime parameter selection can improve the execution time and energy consumption of an application up to 40% and 42% respectively. Md Abdullah Shahneous Bari, Nicholas Chaimov, Abid Muslim Malik, Kevin A. Huck, Barbara M. Chapman, Allen D. Malony, Osman Sarood |
CLUSTER | 6 |
| 2016 | Scaling Spark on HPC SystemsabstractWe report our experiences porting Spark to large production HPC systems. While Spark performance in a data center installation (with local disks) is dominated by the network, our results show that file system metadata access latency can dominate in a HPC installation using Lustre: it determines single node performance up to 4x slower than a typical workstation. We evaluate a combination of software techniques and hardware configurations designed to address this problem. For example, on the software side we develop a file pooling layer able to improve per node performance up to 2.8x. On the hardware side we evaluate a system with a large NVRAM buffer between compute nodes and the backend Lustre file system: this improves scaling at the expense of per-node performance. Overall, our results indicate that scalability is currently limited to O(102) cores in a HPC installation with Lustre and default Spark. After careful configuration combined with our pooling we can scale up to O(10^4). As our analysis indicates, it is feasible to observe much higher scalability in the near future. Nicholas Chaimov, Allen D. Malony, Shane Canon, Costin Iancu, Khaled Z. Ibrahim, Jay Srinivasan |
HPDC | 2 |
| 2016 | A Hartree-Fock Application Using UPC++ and the New DArray LibraryabstractThe Hartree-Fock (HF) method is the fundamental first step for incorporating quantum mechanics into many-electron simulations of atoms and molecules, and it is an important component of computational chemistry toolkits like NWChem. The GTFock code is an HF implementation that, while it does not have all the features in NWChem, represents crucial algorithmic advances that reduce communication and improve load balance by doing an up-front static partitioning of tasks, followed by work stealing whenever necessary. To enable innovations in algorithms and exploit next generation exascale systems, it is crucial to support quantum chemistry codes using expressive and convenient programming models and runtime systems that are also efficient and scalable. This paper presents an HF implementation similar to GTFock using UPC++, a partitioned global address space model that includes flexible communication, asynchronous remote computation, and a powerful multidimensional array library. UPC++ offers runtime features that are useful for HF such as active messages, a rich calculus for array operations, hardware-supported fetch-and-add, and functions for ensuring asynchronous runtime progress. We present a new distributed array abstraction, DArray, that is convenient for the kinds of random-access array updates and linear algebra operations on block-distributed arrays with irregular data ownership. We analyze the performance of atomic fetch-and-add operations (relevant for load balancing) and runtime attentiveness, then compare various techniques and optimizations for each. Our optimized implementation of HF using UPC++ and the DArrays library shows up to 20% improvement over GTFock with Global Arrays at scales up to 24,000 cores. David Ozog, Amir Kamil, Yili Zheng, Paul Hargrove, Jeff R. Hammond, Allen D. Malony, Wibe de Jong, Katherine A. Yelick |
IPDPS | 6 |
| 2016 | The UA?CG Workflow: High Performance Molecular Dynamics of Coarse-Grained PolymersabstractOur analytically based technique for coarse-graining (CG) polymer simulations dramatically improves spatial and temporal scaling while preserving thermodynamic quantities and bulk properties. The purpose of CG codes is to run more efficient molecular dynamics simulations, yet the research field generally lacks thorough analysis of how such codes scale with respect to full-atom representations. This paper conducts an in-depth performance study of highly realistic polymer melts on modern supercomputing systems. We also present a workflow that integrates our analytical solution for calculating CG forces with new high-performance techniques for mapping back and forth between the atomistic and CG descriptions in LAMMPS. The workflow benefits from the performance of CG, while maintaining full-atom accuracy. Our results show speedups up to 12x faster than atomistic simulations. David Ozog, Allen D. Malony, Marina Guenza |
PDP | 2 |
| 2016 | Concurrency in electrical neuroinformatics: parallel computation for studying the volume conduction of brain electrical fields in human head tissuesabstractSummary Advances in human brain neuroimaging for high‐temporal and high‐spatial resolutions will depend on localization of electroencephalography (EEG) signals to their cortex sources. The source localization inverse problem is inherently ill‐posed and depends critically on the modeling of human head electromagnetics. We present a systematic methodology to analyze the main factors and parameters that affect the EEG source‐mapping accuracy. These factors are not independent, and their effect must be evaluated in a unified way. To do so requires significant computational capabilities to explore the problem landscape, quantify uncertainty effects, and evaluate alternative algorithms. Bringing high‐performance computing to this domain is necessary to open new avenues for neuroinformatics research. The head electromagnetics forward problem is the heart of the source localization inverse. We present two parallel algorithms to address tissue inhomogeneity and impedance anisotropy. Highly accurate head modeling environments will enable new research and clinical neuroimaging applications. Cortex‐localized dense‐array EEG analysis is the next‐step in neuroimaging domains such as early childhood reading, understanding of resting‐state brain networks, and models of full brain function. Therapeutic treatments based on neurostimulation will also depend significantly on high‐performance computing integration. Copyright © 2015 John Wiley & Sons, Ltd. Adnan Salman, Allen D. Malony, Sergei Turovets, Vasily Volkov, David Ozog, Don M. Tucker |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | POW: System-wide Dynamic Reallocation of Limited Power in HPCabstractCurrent trends for high-performance systems are leading us towards hardware over-provisioning where it is no longer possible to run each component at peak power without exceeding a system or facility wide power bound. In such scenarios, the power consumed by individual components must be artificially limited to guarantee system operation under a given power bound. In this paper, we present the design of a power scheduler capable of enforcing such a bound using dynamic system-wide power reallocation in an application-agnostic manner. Our scheduler is expected to achieve better job runtimes than a naive power scheduling approach without requiring a priori knowledge of application power behavior. Daniel A. Ellsworth, Allen D. Malony, Barry Rountree, Martin Schulz 0001 |
HPDC | 2 |
| 2015 | Through the Looking-Glass: From Performance Observation to Dynamic AdaptationabstractSince the beginning of ``high-performance'' parallel computing, observing and analyzing performance for purposes of finding bottlenecks and identifying opportunities for improvement has been at the heart of delivering the performance potential of next-generation scalable systems. Interestingly, it is the ever-changing parallel computing landscape that is the main driver of requirements for parallel performance technology and the improvements necessary beyond the current state-of-the-art. Indeed, the development and application of our TAU Performance System over many years largely follows an evolutionary path of addressing measurement and analysis problems in new parallel machines and programming environments. However, the outlook to future parallel systems with high degrees of concurrency, heterogeneous components, dynamic runtime environments, asynchronous execution, and power constraints suggests a new perspective will be needed on the role of performance observation and analysis in respect to tool technology integration and performance optimization methods. The reliance on post-mortem analysis of application-level ("1st person") performance measurements is prohibitive for exascale-class machines because of the performance data volume, the primitive basis for performance data attribution, and the fundamental problem of performance variation that will exist. Instead, it will be important to provide introspection support across the exascale software stack to understand how system ("3rd person") resources are used during execution. Furthermore, the opportunity to couple a global performance introspection capability (a "performance backplane") with online performance decision analytics inspires the concept of an autonomic performance system that can feed back policy-based decisions to guide the computation to better states of execution. The talk will explore these issues by giving a brief retrospective on performance tool evolution, setting the stage for current research projects where a new performance perspective is being pursued. It will also speculate on what might be included in next-generation parallel systems hardware, specifically to make the exascale machines more performance-aware and dynamically-adaptive. Allen D. Malony |
HPDC | 1 |
| 2015 | A Performance Analysis of SIMD Algorithms for Monte Carlo Simulations of Nuclear Reactor CoresabstractA primary characteristic of history-based Monte Carlo neutron transport simulation is the application of MIMD-style parallelism: the path of each neutron particle is largely independent of all other particles, so threads of execution perform independent instructions with respect to other threads. This conflicts with the growing trend of HPC vendors exploiting SIMD hardware, which accomplishes better parallelism and more FLOPS per watt. Event-based neutron transport suits vectorization better than history-based transport, but it is difficult to implement and complicates data management and transfer. However, the Intel Xeon Phi architecture supports the familiar ×86 instruction set and memory model, mitigating difficulties in vector zing neutron transport codes. This paper compares the event-based and history-based approaches for exploiting SIMD in Monte Carlo neutron transport simulations. For both algorithms, we analyze performance using the three different execution models provided by the Xeon Phi (offload, native, and symmetric) within the full-featured OpenVMS framework. A representative micro-benchmark of the performance bottleneck computation shows about 10x performance improvement using the event-based method. In an optimized history-based simulation of a full-physics nuclear reactor core in OpenVMS, the MIC shows a calculation rate 1.6x higher than a modern 16-core CPU, 2.5x higher when balancing load between the CPU and 1 MIC, and 4x higher when balancing load between the CPU and 2 Macs. As far as we are aware, our calculation rate per node on a high fidelity benchmark (17, 098 particles/second) is higher than any other Monte Carlo neutron transport application. Furthermore, we attain 95% distributed efficiency when using MPI and up to 512 concurrent MIC devices. David Ozog, Allen D. Malony, Andrew R. Siegel |
IPDPS | 2 |
| 2015 | An MPI Halo-Cell Implementation for Zero-Copy AbstractionabstractIn the race for Exascale, the advent of many-core processors will bring a shift in parallel computing architectures to systems of much higher concurrency, but with a relatively smaller memory per thread. This shift raises concerns for the adaptability of HPC software, for the current generation to the brave new world. In this paper, we study domain splitting on an increasing number of memory areas as an example problem where negative performance impact on computation could arise. We identify the specific parameters that drive scalability for this problem, and then model the halo-cell ratio on common mesh topologies to study the memory and communication implications. Such analysis argues for the use of shared-memory parallelism, such as with OpenMP, to address the performance problems that could occur. In contrast, we propose an original solution based entirely on MPI programming semantics, while providing the performance advantages of hybrid parallel programming. Our solution transparently replaces halo-cells transfers with pointer exchanges when MPI tasks are running on the same node, effectively removing memory copies. The results we present demonstrate gains in terms of memory and computation time on Xeon Phi (compared to OpenMP-only and MPI-only) using a representative domain decomposition benchmark. Jean-Baptiste Besnard, Allen D. Malony, Sameer Shende, Marc Pérache, Patrick Carribault, Julien Jaeger |
EuroMPI | 2 |
| 2015 | Dynamic power sharing for higher job throughputabstractCurrent trends for high-performance systems are leading towards hardware overprovisioning where it is no longer possible to run all components at peak power without exceeding a system- or facility-wide power bound. The standard practice of static power scheduling is likely to lead to inefficiencies with over- and under-provisioning of power to components at runtime. In this paper we investigate the performance and scalability of an application agnostic runtime power scheduler (POWsched) that is capable of enforcing a system-wide power limit. Our experimental results show POWsched is robust, has negligible overhead, and can take advantage of opportunities to shift wasted power to more power-intensive applications, improving overall workload runtime by as much as 14% without job scheduler integration or application specific profiling. In addition, we conduct scalability studies to determine POWsched's overhead for large node counts. Lastly, we contribute a model and simulator (POWsim) for investigating dynamic power scheduling behavior and enforcement at scale. Daniel A. Ellsworth, Allen D. Malony, Barry Rountree, Martin Schulz 0001 |
SC | 2 |
| 2014 | Particle advection performance over varied architectures and workloadsabstractParticle advection is a foundational operation for many flow visualization techniques, including streamlines, Finite-Time Lyapunov Exponents (FTLE) calculation, and stream surfaces. The workload for particle advection problems varies greatly, including significant variation in computational requirements. With this study, we consider the performance impacts from hardware architecture on this problem, studying distributed-memory systems with CPUs with varying amounts of cores per node, and with nodes with one to three GPUs. Our goal was to explore which architectures were best suited to which workloads, and why. While the results of this study will help inform visualization scientists which architectures they should use when solving certain flow visualization problems, it is also informative for the larger HPC community, since many simulation codes will soon incorporate visualization via in situ techniques. Hank Childs, Scott Biersdorff, David Poliakoff, David Camp, Allen D. Malony |
HiPC | 5 |
| 2014 | Toward multi-target autotuning for acceleratorsabstractProducing high-performance implementations from simple, portable computation specifications is a challenge that compilers have tried to address for several decades. More recently, a relatively stable architectural landscape has evolved into a set of increasingly diverging and rapidly changing CPU and accelerator designs, with the main common factor being dramatic increases in the levels of parallelism available. The growth of architectural heterogeneity and parallelism, combined with the very slow development cycles of traditional compilers, has motivated the development of autotuning tools that can quickly respond to changes in architectures and programming models, and enable very specialized optimizations that are not possible or likely to be provided by mainstream compilers. In this paper we describe the new OpenCL code generator and autotuner OrCL and the introduction of detailed performance measurement into the autotuning process. OrCL is implemented within the Orio autotuning framework, which enables the rapid development of experimental languages and code optimization strategies aimed at achieving good performance on new platforms without rewriting or hand-optimizing critical kernels. The combination of the new OpenCL autotuning and TAU measurement capabilities enables users to consistently evaluate autotuning effectiveness across a range of architectures, including several NVIDIA and AMD accelerators and Intel Xeon Phi processors, and to compare the OpenCL and CUDA code generation capabilities. We present results of autotuning several numerical kernels that typically dominate the execution time of iterative sparse linear system solution and key computations from a 3-D parallel simulation of solid fuel ignition. Nicholas Chaimov, Boyana Norris, Allen D. Malony |
ICPADS | 3 |
| 2014 | WorkQ: A many-core producer/consumer execution model applied to PGAS computationsabstractPartitioned global address space (PGAS) applications, such as the Tensor Contraction Engine (TCE) in NWChem, often apply a one-process-per-core mapping in which each process iterates through the following work-processing cycle: (1) determine a work-item dynamically, (2) get data via one-sided operations on remote blocks, (3) perform computation on the data locally, (4) put (or accumulate) resultant data into an appropriate remote location, and (5) repeat the cycle. However, this simple flow of execution does not effectively hide communication latency costs despite the opportunities for making asynchronous progress. Utilizing nonblocking communication calls is not sufficient unless care is taken to efficiently manage a responsive queue of outstanding communication requests. This paper presents a new runtime model and its library implementation for managing tunable “work queues” in PGAS applications. Our runtime execution model, called WorkQ, assigns some number of on-node “producer” processes to primarily do communication (steps 1, 2, 4, and 5) and the other “consumer” processes to do computation (step 3); but processes can switch roles dynamically for the sake of performance. Load balance, synchronization, and overlap of communication and computation are facilitated by a tunable nodewise FIFO message queue protocol. Our WorkQ library implementation enables an MPI+X hybrid programming model where the X comprises SysV message queues and the user's choice of SysV, POSIX, and MPI shared memory. We develop a simplified software mini-application that mimics the performance behavior of the TCE at arbitrary scale, and we show that the WorkQ engine outperforms the original model by about a factor of 2. We also show performance improvement in the TCE coupled cluster module of NWChem. David Ozog, Allen D. Malony, Jeff R. Hammond, Pavan Balaji |
ICPADS | 2 |
| 2014 | From MultiTask to MultiCore: Design and Implementation Using an RTOSabstractPractice has shown that programming a new multicore system is a greater challenge than previously thought. The challenge is to produce the resulting system in a way, which is as easy as sequential programming. This new trend has changed the way we think about the whole development process. The aim of this work is to show that it is possible to develop a multicore embedded system application using existing tools, while at the same time, obtaining reuse. This process is carried out in a cyclic and increasing manner, generating a more refined version of the application at each iteration. The development process consists of five phases: Multitask Modelling, Code Generation, Test/Debugging, Mapping Tasks to Cores and Tuning the Application. The three initial ones are carried out using the Visual RTXC tool, whereas the last two use the performance tool TAU. Using a small application, a Case Study shows how the proposed development process works and the steps involved in the implementation of an embedded system. Célio Estevan Morón, Antonio Ideguchi, Marcio Merino Fernandes, Allen D. Malony |
ISPDC | 4 |
| 2014 | General Hybrid Parallel ProfilingabstractA hybrid parallel measurement system offers the potential to fuse the principal advantages of probe-based tools, with their exact measures of performance and ability to capture event semantics, and sampling-based tools, with their ability to observe performance detail with less overhead. Creating a hybrid profiling solution is challenging because it requires new mechanisms for integrating probe and sample measurements and calculating profile statistics during execution. In this paper, we describe a general hybrid parallel profiling tool that has been implemented in the TAU Performance System. Its generality comes from the fact that all of the features of the individual methods are retained and can be flexibly controlled when combined to address the measurement requirements for a particular parallel application. The design of the hybrid profiling approach is described and the implementation of the prototype in TAU presented. We demonstrate hybrid profiling functionality first on a simple sequential program and then show its use for several OpenMP parallel codes from the NAS Parallel Benchmark. These experiments also highlight the improvements in overhead efficiency made possible by hybrid profiling. A large-scale ocean modeling code based on OpenMP and MPI, MPAS-Ocean, is used to show how the TAU hybrid profiling tool can be effective at exposing performance-limiting behavior that would be difficult to identify otherwise. Allen D. Malony, Kevin A. Huck |
PDP | 1 |
| 2013 | MIL: A language to build program analysis tools through static binary instrumentationabstractAs software complexity increases, the analysis of code behavior during its execution is becoming more important. Instrumentation techniques, through the insertion of code directly into binaries, are essential for program analyses used in debugging, runtime profiling, and performance evaluation. In the context of high-performance parallel applications, building an instrumentation framework is quite challenging. One of the difficulties is due to the necessity to capture both coarse-grain behavior, such as the execution time of different functions, as well as finer-grain actions, in order to pinpoint performance issues. In this paper, we propose a language, MIL, for the development of program analysis tools based on static binary instrumentation. The key feature of MIL is to ease the integration of static, global program analysis with instrumentation. We will show how this enables both a precise targeting of the code regions to analyze and a better understanding of the optimized program behavior. Andres Charif Rubial, Denis Barthou, Cédric Valensi, Sameer Shende, Allen D. Malony, William Jalby |
HiPC | 5 |
| 2013 | Inspector-Executor Load Balancing Algorithms for Block-Sparse Tensor ContractionsabstractDeveloping effective yet scalable load-balancing methods for irregular computations is critical to the successful application of simulations in a variety of disciplines at petascale and beyond. This paper explores a set of static and dynamic scheduling algorithms for block-sparse tensor contractions within the NWChem computational chemistry code for different degrees of sparsity (and therefore load imbalance). In this particular application, a relatively large amount of task information can be obtained at minimal cost, which enables the use of static partitioning techniques that take the entire task list as input. However, fully static partitioning is incapable of dealing with dynamic variation of task costs, such as from transient network contention or operating system noise, so we also consider hybrid schemes that utilize dynamic scheduling within subgroups. These two schemes, which have not been previously implemented in NWChem or its proxies (i.e. quantum chemistry mini-apps) are compared to the original centralized dynamic load-balancing algorithm as well as improved centralized scheme. In all cases, we separate the scheduling of tasks from the execution of tasks into an inspector phase and an executor phase. The impact of these methods upon the application is substantial on a large InfiniBand cluster: execution time is reduced by as much as 50% at scale. The technique is applicable to any scientific application requiring load balance where performance models or estimations of kernel execution times are available. David Ozog, Jeff R. Hammond, James Dinan, Pavan Balaji, Sameer Shende, Allen D. Malony |
ICPP | 6 |
| 2013 | Inspector/executor load balancing algorithms for block-sparse tensor contractionsabstractDeveloping effective yet scalable load-balancing methods for irregular computations is critical to the successful application of simulations in a variety of disciplines at petascale and beyond. This paper explores a set of static and dynamic scheduling algorithms for block-sparse tensor contractions within the NWChem computational chemistry code for different degrees of sparsity (and therefore load imbalance). In this particular application, a relatively large amount of task information can be obtained at minimal cost, which enables the use of static partitioning techniques that take the entire task list as input. However, fully static partitioning is incapable of dealing with dynamic variation of task costs, such as from transient network contention or operating system noise, so we also consider hybrid schemes that utilize dynamic scheduling within subgroups. These two schemes, which have not been previously implemented in NWChem or its proxies (i.e. quantum chemistry mini-apps) are compared to the original centralized dynamic load-balancing algorithm as well as improved centralized scheme. In all cases, we separate the scheduling of tasks from the execution of tasks into an inspector phase and an executor phase. The impact of these methods upon the application is substantial on a large InfiniBand cluster: execution time is reduced by as much as 50% at scale. The technique is applicable to any scientific application requiring load balance where performance models or estimations of kernel execution times are available. David Ozog, Sameer Shende, Allen D. Malony, Jeff R. Hammond, James Dinan, Pavan Balaji |
ICS | 3 |
| 2012 | A Type-Based Approach to Separating Protocol from Application Logic - A Case Study in Hybrid Computer Programming
Geoffrey C. Hulette, Matthew J. Sottile, Allen D. Malony |
Euro-Par | 3 |
| 2012 | Topic 2: Performance Prediction and Evaluation
Allen D. Malony, Helen D. Karatza, William J. Knottenbelt, Sally A. McKee |
Euro-Par | 1 |
| 2012 | Composing typemaps in TwigabstractTwig is a language for writing typemaps, programs which transform the type of a value while preserving its underlying meaning. Typemaps are typically used by tools that generate code, such as multi-language wrapper generators, to automatically convert types as needed. Twig builds on existing typemap tools in a few key ways. Twig's typemaps are composable so that complex transformations may be built from simpler ones. In addition, Twig incorporates an abstract, formal model of code generation, allowing it to output code for different target languages. We describe Twig's formal semantics and show how the language allows us to concisely express typemaps. Then, we demonstrate Twig's utility by building an example typemap. Geoffrey C. Hulette, Matthew J. Sottile, Allen D. Malony |
GPCE | 3 |
| 2012 | Incorporating anatomical connectivity into EEG source estimation via sparse approximation with cortical graph waveletsabstractThe source estimation problem for EEG consists of estimating cortical activity from measurements of electrical potential on the scalp surface. This is a underconstrained inverse problem as the dimensionality of cortical source currents far exceeds the number of sensors. We develop a novel regularization for this inverse problem which incorporates knowledge of the anatomical connectivity of the brain, measured by diffusion tensor imaging. We construct an overcomplete wavelet frame, termed cortical graph wavelets, by applying the recently developed spectral graph wavelet transform to this anatomical connectivity graph. Our signal model is formed by assuming that the desired cortical currents have a sparse representation in these cortical graph wavelets, which leads to a convex ℓ1-regularized least squares problem for the coefficients. On data from a simple motor potential experiment, the proposed method shows improvement over the standard minimum-norm regularization. David K. Hammond, Benoit Scherrer, Allen D. Malony |
ICASSP | 3 |
| 2012 | Performance characterization of global address space applications: a case study with NWChemabstractSUMMARY The use of global address space languages and one‐sided communication for complex applications is gaining attention in the parallel computing community. However, lack of good evaluative methods to observe multiple levels of performance makes it difficult to isolate the cause of performance deficiencies and to understand the fundamental limitations of system and application design for future improvement. NWChem is a popular computational chemistry package, which depends on the Global Arrays/Aggregate Remote Memory Copy Interface suite for partitioned global address space functionality to deliver high‐end molecular modeling capabilities. A workload characterization methodology was developed to support NWChem performance engineering on large‐scale parallel platforms. The research involved both the integration of performance instrumentation and measurement in the NWChem software, as well as the analysis of one‐sided communication performance in the context of NWChem workloads. Scaling studies were conducted for NWChem on Blue Gene/P and on two large‐scale clusters using different generation Infiniband interconnects and x86 processors. The performance analysis and results show how subtle changes in the runtime parameters related to the communication subsystem could have significant impact on performance behavior. The tool has successfully identified several algorithmic bottlenecks, which are already being tackled by computational chemists to improve NWChem performance. Copyright © 2011 John Wiley & Sons, Ltd. Jeff R. Hammond, Sriram Krishnamoorthy, Sameer Shende, Nichols A. Romero, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 5 |
| 2011 | Development of embedded multicore systemsabstractThe concepts involved in the programming process of multicore systems have been quite well known for decades. The problem is to produce it in a form as easy as sequential programming. This new trend will change the way we think about the whole development process. We will show that it is possible to develop a multicore embedded system application using existing tools and the model-driven development process proposed. To do this, two tools will be used: VisualRTXC (available at www.quadrosbrasil.com.br) for generating the multithread communication/synchronization structures and a performance tool called TAU (available at http://www.cs.uoregon.edu/research/tau/home.php) for the tuning of the final implementation. Célio Estevan Morón, Allen D. Malony |
ETFA | 2 |
| 2011 | Parallel Performance Measurement of Heterogeneous Parallel Systems with GPUsabstractThe power of GPUs is giving rise to heterogeneous parallel computing, with new demands on programming environments, runtime systems, and tools to deliver high-performing applications. This paper studies the problems associated with performance measurement of heterogeneous machines with GPUs. A heterogeneous computation model and alternative host-GPU measurement approaches are discussed to set the stage for reporting new capabilities for heterogeneous parallel performance measurement in three leading HPC tools: PAPI, Vampir, and the TAU Performance System. Our work leverages the new CUPTI tool support in NVIDIA's CUDA device library. Heterogeneous benchmarks from the SHOC suite are used to demonstrate the measurement methods and tool support. Allen D. Malony, Scott Biersdorff, Sameer Shende, Heike Jagode, Stanimire Tomov, Guido Juckeland, Robert Dietrich, Duncan Poole, Christopher Lamb 0001 |
ICPP | 1 |
| 2010 | Design and Implementation of a Hybrid Parallel Performance Measurement SystemabstractModern parallel performance measurement systems collect performance information either through probes inserted in the application code or via statistical sampling. Probe-based techniques measure performance metrics directly using calls to a measurement library that execute as part of the application. In contrast, sampling-based systems interrupt program execution to sample metrics for statistical analysis of performance. Although both measurement approaches are represented by robust tool frameworks in the performance community, each has its strengths and weaknesses. In this paper, we investigate the creation of a hybrid measurement system, the goal being to exploit the strengths of both systems and mitigate their weaknesses. We show how such a system can be used to provide the application programmer with a more complete analysis of their application. Simple example and application codes are used to demonstrate its capabilities. We also show how the hybrid techniques can be combined to provide real cross-language performance evaluation of an uninstrumented run for mixed compiled/interpreted execution environments (e.g., Python and C/C++/Fortran). Alan Morris, Allen D. Malony, Sameer Shende, Kevin A. Huck |
ICPP | 2 |
| 2010 | An experimental approach to performance measurement of heterogeneous parallel applications using CUDAabstractHeterogeneous parallel systems using GPU devices for application acceleration have garnered significant attention in the supercomputing community. However, to realize the full potential of GPU computing, application developers will require tools to measure and analyze accelerator performance with respect to the parallel execution as a whole. A performance measurement technology for the NVIDIA CUDA platform has been developed and integrated with the TAU parallel performance system. The design of the TAUcuda package is based on an experimental NVIDIA CUDA driver and associated runtime and device libraries. In any environment where the CUDA experimental driver is installed, TAUcuda can provide detailed performance information regarding the execution of GPU kernels and the interactions with the parallel program without any modification to the program source or executable code. The paper describes the TAUcuda technology and how it is integrated with the TAU measurement framework to provide integrated performance views. Various examples of TAUcuda use are presented, including CUDA SDK examples, a GPU version of the Linpack benchmark, and a scalable molecular dynamics application, NAMD. Allen D. Malony, Scott Biersdorff, Wyatt Spear, Shangkar Mayanglambam |
ICS | 1 |
| 2010 | A framework for scalable, parallel performance monitoringabstractAbstract Performance monitoring of HPC applications offers opportunities for adaptive optimization based on the dynamic performance behavior, unavailable in purely post‐mortem performance views. However, a parallel performance monitoring system must have low overhead and high efficiency to make these opportunities tangible. We describe a scalable parallel performance monitor calledTAUoverMRNet (ToM), created from the integration of the TAU performance system and the Multicast Reduction Network (MRNet). The integration is achieved through a plug‐in architecture in TAU that allows the selection of different transport substrates to offload the online performance data. A method to establish the transport overlay structure of the monitor from within TAU, one that requires no added support from the job manager or application, is presented. We demonstrate the distribution of performance analysis from the sink to the overlay nodes and the reduction in the large‐scale profile data that could, otherwise, overwhelm any single sink. The results show low perturbation and significant savings accrued from reduction at large processor‐counts. Copyright © 2009 John Wiley & Sons, Ltd. Aroon Nataraj, Allen D. Malony, Alan Morris, Dorian C. Arnold, Barton P. Miller |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Integrated Performance Views in Charm++: Projections Meets TAUabstractThe Charm++ parallel programming system provides a modular performance interface that can be used to extend its performance measurement and analysis capabilities. The interface exposes execution events of interest representing Charm++ scheduling operations, application methods/routines, and communication events for observation by alternative performance modules configured to implement different measurement features. The paper describes the Charm++'s performance interface and how the Charm++ Projections tool and the TAU Performance System can provide integrated trace-based and profile-based performance views. These two tools are complementary, providing the user with different performance perspectives on Charm++ applications based on performance data detail and temporal and spatial analysis. How the tools work in practice is demonstrated in a parallel performance analysis of NAMD, a scalable molecular dynamics code that applies many of Charm++'s unique features. Scott Biersdorff, Chee Wai Lee, Allen D. Malony, Laxmikant V. Kalé |
ICPP | 3 |
| 2008 | In search of sweet-spots in parallel performance monitoringabstractParallel performance monitoring extends parallel measurement systems with infrastructure and interfaces for online performance data access, communication, and analysis. At the same time it raises concerns for the impact on application execution from monitor overhead. The application monitoring scheme parameterized by performance events to monitor, access frequency and the type of data analysis operation defines a set of monitoring requirements. The monitoring infrastructure presents its own choices, particularly the amount and configuration of resources devoted explicitly to monitoring. The key to scalable, low-overhead parallel performance monitoring is to match the application monitoring demands to the effective operating range of the monitoring system (or vice-versa). A poor match can result in over-provisioning (wasted resources) or in under-provisioning (lack of scalability, high overheads and poor quality of performance data). We present a methodology and evaluation framework to determine the sweet-spots for performance monitoring using TAU and MRNet. Aroon Nataraj, Allen D. Malony, Allen Morris, Dorian C. Arnold, Barton P. Miller |
CLUSTER | 2 |
| 2008 | WOOL: A Workflow Programming LanguageabstractWorkflows offer scientists a simple but flexible programming model at a level of abstraction closer to the domain-specific activities that they seek to perform. However, languages for describing workflows tend to be highly complex, or specialized towards a particular domain, or both. WOOL is an abstract workflow language with human-readable syntax, intuitive semantics, and a powerful abstract type system. WOOL workflows can be targeted to almost any kind of runtime system supporting data-flow computation. This paper describes the design of the WOOL language and the implementation of its compiler, along with a simple example runtime. We demonstrate its use in an image-processing workflow. Geoffrey C. Hulette, Matthew J. Sottile, Allen D. Malony |
eScience | 3 |
| 2008 | Observing Performance Dynamics Using Parallel Profile Snapshots
Alan Morris, Wyatt Spear, Allen D. Malony, Sameer Shende |
Euro-Par | 3 |
| 2008 | Capturing performance knowledge for automated analysisabstractAutomating the process of parallel performance experimentation, analysis, and problem diagnosis can enhance environments for performance-directed application development, compilation, and execution. This is especially true when parametric studies, modeling, and optimization strategies require large amounts of data to be collected and processed for knowledge synthesis and reuse. This paper describes the integration of the PerfExplorer performance data mining framework with the OpenUH compiler infrastructure. OpenUH provides auto-instrumentation of source code for performance experimentation and PerfExplorer provides automated and reusable analysis of the performance data through a scripting interface. More importantly, PerfExplorer inference rules have been developed to recognize and diagnose performance characteristics important for optimization strategies and modeling. Three case studies are presented which show our success with automation in OpenMP and MPI code tuning, parametric characterization, Pand power modeling. The paper discusses how the integration supports performance knowledge engineering across applications and feedback-based compiler optimization in general. Kevin A. Huck, Oscar R. Hernandez, Van Bui, Sunita Chandrasekaran, Barbara M. Chapman, Allen D. Malony, Lois C. McInnes, Boyana Norris |
SC | 6 |
| 2007 | TAUoverSupermon : Low-Overhead Online Parallel Performance Monitoring
Aroon Nataraj, Matthew J. Sottile, Alan Morris, Allen D. Malony, Sameer Shende |
Euro-Par | 4 |
| 2007 | Automatic Performance Diagnosis of Parallel Computations with Compositional ModelsabstractPerformance tuning involves a diagnostic process to locate and explain sources of program inefficiency. A performance diagnosis system can leverage knowledge of performance causes and symptoms that come from expertise with parallel computational models. This paper extends our model-based performance diagnosis approach to programs with multiple models. We study two types of model compositions (nesting and restructuring) and demonstrate how the Hercule performance diagnosis framework can automatically discover and interpret performance problems due to model nesting in the FLASH application. Li Li 0020, Allen D. Malony |
IPDPS | 2 |
| 2007 | Development of NeuroElectroMagnetic ontologies(NEMO): a framework for mining brainwave ontologiesabstractEvent-related potentials (ERP) are brain electrophysiological patterns created by averaging electroencephalographic (EEG) data, time-locking to events of interest (e.g., stimulus or response onset). In this paper, we propose a generic framework for mining anddeveloping domain ontologies and apply it to mine brainwave (ERP) ontologies. The concepts and relationships in ERP ontologies can be mined according to the following steps: pattern decomposition, extraction of summary metrics for concept candidates, hierarchical clustering of patterns for classes and class taxonomies, and clustering-based classification and association rules mining for relationships (axioms) of concepts. We have applied this process to several dense-array (128-channel) ERP datasets. Results suggest good correspondence between mined concepts and rules, on the one hand, and patterns and rules that were independently formulated by domain experts, on the other. Data mining results also suggest ways in which expert-defined rules might be refined to improve ontologyrepresentation and classification results. The next goal of our ERP ontology mining framework is to address some long-standing challenges in conducting large-scale comparison and integration of results across ERP paradigms and laboratories. In a more general context, this work illustrates the promise of an interdisciplinary research program, which combines data mining, neuroinformatics andontology engineering to address real-world problems. Dejing Dou, Gwen A. Frishkoff, Jiawei Rong, Robert M. Frank, Allen D. Malony, Don M. Tucker |
KDD | 5 |
| 2007 | The ghost in the machine: observing the effects of kernel operation on parallel application performanceabstractThe performance of a parallel application on a scalable HPC system is determined by user-level execution of the application code and system-level (OS kernel) operations. To understand the influences of system-level factors on application performance, the measurement of OS kernel activities is key. We describe a technology to observe kernel actions and make this information available to application-level performance measurement tools. The benefits of merged application and OS performance information and its use in parallel performance analysis are demonstrated, both for profiling and tracing methodologies. In particular, we focus on the problem of kernel noise assessment as a stress test of the approach. We show new results for characterizing noise and introduce new techniques for evaluating noise interference and its effects on application execution. Our kernel measurement and noise analysis technologies are being developed as part of Linux OS environments for scalable parallel systems. Aroon Nataraj, Alan Morris, Allen D. Malony, Matthew J. Sottile, Pete Beckman |
SC | 3 |
| 2007 | Knowledge engineering for automatic parallel performance diagnosisabstractAbstract Scientific parallel programs often undergo significant performance tuning before meeting their performance expectation. Performance tuning naturally involves a diagnostic process—locating performance bugs that make a program inefficient and explaining them in terms of high‐level program design. We present a systematic approach to generating performance knowledge for automatically diagnosing parallel programs. Our approach exploits program semantics and parallelism found in parallel programming patterns to search for and define bugs. The approach addresses how to extract the expert knowledge required for performance diagnosis from parallel patterns and represents this knowledge in a manner such that the diagnostic process can be automated. We demonstrate the effectiveness of our knowledge‐engineering approach through a case study. Our experience diagnosing divide‐and‐conquer programs shows that pattern‐based performance knowledge can provide effective guidance for locating and defining performance bugs at a high level of program abstraction. Copyright © 2006 John Wiley & Sons, Ltd. Li Li 0020, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Performance modeling of component assembliesabstractAbstract A parallel component environment places constraints on performance measurement and modeling. For instance, it must be possible to instrument the application without access to the source code. In addition, a component may admit multiple implementations, based on the choice of algorithm, data structure, parallelization strategy, etc., posing the user with the problem of having to choose the ‘correct’ implementation and achieve an optimal (fastest) component assembly. Under the assumption that an empirical performance model exists for each implementation of each component, simply choosing the optimal implementation of each component does not guarantee an optimal component assembly since components interact with each other. An optimal solution may be obtained by evaluating the performance of all of the possible realizations of a component assembly given the components and all of their implementations, but the exponential complexity renders the approach unfeasible as the number of components and their implementations rise. This paper describes a non‐intrusive, coarse‐grained performance monitoring system that allows the user to gather performance data through the use of proxies. In addition, a simple optimization library that identifies a nearly optimal configuration is proposed. Finally, some experimental results are presented that illustrate the measurement and optimization strategies. Copyright © 2006 John Wiley & Sons, Ltd. Nick Trebon, Allen Morris, Jaideep Ray, Sameer Shende, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 5 |
| 2006 | Kernel-Level Measurement for Integrated Parallel Performance Views: the KTAU ProjectabstractThe effect of the operating system on application performance is an increasingly important consideration in high performance computing. OS kernel measurement is key to understanding the performance influences and the interrelationship of system and user-level performance factors. The KTAU (Kernel TAU) methodology and Linux-based framework provides parallel kernel performance measurement from both a kernel-wide and process-centric perspective. The first characterizes overall aggregate kernel performance for the entire system. The second characterizes kernel performance when it runs in the context of a particular process. KTAU extends the TAU performance system with kernel-level monitoring, while leveraging TAU's measurement and analysis capabilities. We explain the rational and motivations behind our approach, describe the KTAU design and implementation, and show working examples on multiple platforms demonstrating the versatility of KTAU in integrated system/application monitoring Aroon Nataraj, Allen D. Malony, Sameer Shende, Alan Morris |
CLUSTER | 2 |
| 2006 | Model-Based Performance Diagnosis of Master-Worker Parallel Computations
Li Li 0020, Allen D. Malony |
Euro-Par | 2 |
| 2006 | Early Experiences with KTAU on the IBM BG/L
Aroon Nataraj, Allen D. Malony, Alan Morris, Sameer Shende |
Euro-Par | 2 |
| 2006 | Model-Based Relative Performance Diagnosis of Wavefront Parallel Computations
Li Li 0020, Allen D. Malony, Kevin A. Huck |
HPCC | 2 |
| 2006 | Integrating TAU with Eclipse: A Performance Analysis System in an Integrated Development Environment
Wyatt Spear, Allen D. Malony, Alan Morris, Sameer Shende |
HPCC | 2 |
| 2006 | Parallel ICA methods for EEG neuroimagingabstractHiPerSAT, a C++ library and tools, processes EEG data sets with ICA (independent component analysis) methods. HiPerSAT uses BLAS, LAPACK, MPI and OpenMP to achieve a high performance solution that exploits parallel hardware. ICA is a class of methods for analyzing a large set of data samples and extracting independent components that explain the observed data. ICA is used in EEG research for data cleaning and separation of spatiotemporal patterns that may reflect different underlying neural processes. We present two ICA implementations (FastICA and Info-max) that exploit parallelism to provide an EEG component decomposition solution of higher performance and data capacity than current MATLAB-based implementations. Experimental results and the methodology used to obtain them are presented. Integrating HiPerSAT with EEGLAB (A. Delorme and S. Makeig, 2004) is described, as well as future plans for this research. Dan B. Keith, Christian C. Hoge, Robert M. Frank, Allen D. Malony |
IPDPS | 4 |
| 2006 | Open trace - The open trace format (OTF) and open tracing for HPCabstractThe Open Trace Format (OTF) is a DOE-sponsored initiative to help deliver open, scalable performance tracing tools for HPC systems. OTF is an open specification of trace information to provide a target for trace generation and to allow trace analysis and visualization tools to operate efficiently at large scale. The Technical University of Dresden and ParaTools, Inc. developed the first version OTF with support from Lawrence Livermore National Laboratory.The BOF has two goals. First, we will report on the current status of OTF. This will include a review of the OTF specification and a description of the OTF reader/writer library to be released into open source at SC06. We will also report on recent ports of the library to HPC platforms.The second goal is to invite community involvement in the OTF initiative and to establish a working group chartered with evolving the OTF specification in the future. Allen D. Malony, Wolfgang E. Nagel |
SC | 1 |
| 2006 | Bridging the language gap in scientific computing: the Chasm approachabstractAbstract Chasm is a toolkit providing seamless language interoperability between Fortran 95 and C++. Language interoperability is important to scientific programmers because scientific applications are predominantly written in Fortran, while software tools are mostly written in C++. Two design features differentiate Chasm from other related tools. First, we avoid the common‐denominator type systems and programming models found in most Interface Definition Language (IDL)‐based interoperability systems. Chasm uses the intermediate representation generated by a compiler front‐end for each supported language as its source of interface information instead of an IDL. Second, bridging code is generated for each pairwise language binding, removing the need for a common intermediate data representation and multiple levels of indirection between the caller and callee. These features make Chasm a simple system that performs well, requires minimal user intervention and, in most instances, bridging code generation can be performed automatically. Chasm is also easily extensible and highly portable. Copyright © 2005 John Wiley & Sons, Ltd. Craig Edward Rasmussen, Matthew J. Sottile, Sameer Shende, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 4 |
| 2005 | Topic 2 - Performance Prediction and Evaluation
Allen D. Malony, Thomas Fahringer, Allan Snavely, Luís Silva |
Euro-Par | 1 |
| 2005 | Models for On-the-Fly Compensation of Measurement Overhead in Parallel Performance Profiling
Allen D. Malony, Sameer Shende |
Euro-Par | 1 |
| 2005 | Trace-Based Parallel Performance Overhead Compensation
Felix Wolf 0001, Allen D. Malony, Sameer Shende, Alan Morris |
HPCC | 2 |
| 2005 | Design and Implementation of a Parallel Performance Data Management FrameworkabstractEmpirical performance evaluation of parallel systems and applications can generate significant amounts of performance data and analysis results from multiple experiments as performance is investigated and problems diagnosed. Hence, the management of performance information is a core component of performance analysis tools. To better support tool integration, portability; and reuse, there is a strong motivation to develop performance data management technology that can provide a common foundation for performance data storage, access, merging, and analysis. This paper presents the design and implementation of the performance data management framework (PerfDMF). PerfDMF addresses objectives of performance tool integration, interoperation, and reuse by providing common data storage, access, and analysis infrastructure for parallel performance profiles. PerfDMF includes an extensible parallel profile data schema and relational database schema, a profile query and analysis programming interface, and an extendible toolkit for profile import/export and standard analysis. We describe the PerfDMF objectives and architecture, give detailed explanation of the major components, and show examples of PerfDMF application. Kevin A. Huck, Allen D. Malony, Robert Bell, Alan Morris |
ICPP | 2 |
| 2005 | PerfExplorer: A Performance Data Mining Framework For Large-Scale Parallel ComputingabstractParallel applications running on high-end computer systems manifest a complexity of performance phenomena. Tools to observe parallel performance attempt to capture these phenomena in measurement datasets rich with information relating multiple performance metrics to execution dynamics and parameters specific to the application-system experiment. However, the potential size of datasets and the need to assimilate results from multiple experiments makes it a daunting challenge to not only process the information, but discover and understand performance insights. In this paper, we present PerfExplorer, a framework for parallel performance data mining and knowledge discovery. The framework architecture enables the development and integration of data mining operations that will be applied to large-scale parallel performance profiles. PerfExplorer operates as a client-server system and is built on a robust parallel performance database (PerfDMF) to access the parallel profiles and save its analysis results. Examples are given demonstrating these techniques for performance analysis of ASCI applications. Kevin A. Huck, Allen D. Malony |
SC | 2 |
| 2005 | Performance technology for parallel and distributed component softwareabstractAbstract This work targets the emerging use of software component technology for high‐performance scientific parallel and distributed computing. While component software engineering will benefit the construction of complex science applications, its use presents several challenges to performance measurement, analysis, and optimization. The performance of a component application depends on the interaction (possibly nonlinear) of the composed component set. Furthermore, a component is a ‘binary unit of composition’ and the only information users have is the interface the component provides to the outside world. A performance engineering methodology and development approach is presented to address evaluation and optimization issues in high‐performance component environments. We describe a prototype implementation of a performance measurement infrastructure for the Common Component Architecture (CCA) system. A case study demonstrating the use of this technology for integrated measurement, monitoring, and optimization in CCA component‐based applications is given. Copyright © 2005 John Wiley & Sons, Ltd. Allen D. Malony, Sameer Shende, Nick Trebon, Jaideep Ray, Robert C. Armstrong, Craig Edward Rasmussen, Matthew J. Sottile |
Concurr. Pract. Exp. | 1 |
| 2004 | Topic 1: Support Tools and Environments
José C. Cunha, Allen D. Malony, Arndt Bode, Dieter Kranzlmüller |
Euro-Par | 2 |
| 2004 | Overhead Compensation in Performance Profiling
Allen D. Malony, Sameer Shende |
Euro-Par | 1 |
| 2004 | Performance Measurement and Modeling of Component Applications in a High Performance Computing Environment: A Case StudyabstractSummary form only given. We present a case study of performance measurement and modeling of a CCA (common component architecture) component-based application in a high performance computing environment. Component-based HPC applications allow the possibility of creating component-level performance models and synthesizing them into application performance models. However, they impose the restriction that performance measurement/monitoring needs to be done in a nonintrusive manner and at a fairly coarse-grained level. We propose a performance measurement infrastructure for HPC based loosely on recent work done for grid environments. A prototypical implementation of the infrastructure is used to collect data for three components in a scientific application and construct their performance models. Both computational and message-passing performance are addressed. Jaideep Ray, Nick Trebon, Robert C. Armstrong, Sameer Shende, Allen D. Malony |
IPDPS | 5 |
| 2003 | A Distributed Performance Analysis Architecture for ClustersabstractThe use of a cluster for distributed performance analysis of parallel trace data is discussed. We propose an analysis architecture that uses multiple cluster nodes as a server to execute analysis operations in parallel and communicate to remote clients where performance visualization and user interactions occur. The client-server system developed, VNG, is highly configurable and is shown to perform well for traces of large size, when compared to leading trace visualization systems. Holger Brunst, Wolfgang E. Nagel, Allen D. Malony |
CLUSTER | 3 |
| 2003 | ParaProf: A Portable, Extensible, and Scalable Tool for Parallel Performance Profile Analysis
Robert Bell, Allen D. Malony, Sameer Shende |
Euro-Par | 2 |
| 2003 | Performance Evaluation and Prediction
Jeffrey K. Hollingsworth, Allen D. Malony, Jesús Labarta, Thomas Fahringer |
Euro-Par | 2 |
| 2003 | Integration and application of TAU in parallel Java environmentsabstractAbstract Parallel Java environments present challenging problems for performance tools because of Java's rich language system and its multi‐level execution platform combined with the integration of native‐code application libraries and parallel run‐time software. In addition to the desire to provide robust performance measurement and analysis capabilities for the Java language itself, the coupling of different software execution contexts under a uniform performance model needs careful consideration of how events of interest are observed and how cross‐context parallel execution information is linked. This paper relates our experience in extending the TAU (Tuning and Analysis Utilities) performance system to a parallel Java environment based on mpiJava. We describe the complexities of the instrumentation model used, how performance measurements are made, and the overhead incurred. A parallel Java application simulating the game of Life is used to show the performance system's capabilities. Copyright © 2003 John Wiley & Sons, Ltd. Sameer Shende, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 2 |
| 2002 | Design and Prototype of a Performance Tool Interface for OpenMP
Bernd Mohr, Allen D. Malony, Sameer Shende, Felix Wolf 0001 |
J. Supercomput. | 2 |
| 2001 | Topic 02: Performance Evaluation and Prediction
Allen D. Malony, Graham D. Riley, Bernd Mohr, J. Mark Bull, Tomàs Margalef |
Euro-Par | 1 |
| 2001 | On using SCALEA for performance analysis of distributed and parallel programsabstractIn this paper we give an overview of SCALEA, which is a new performance analysis tool for OpenMP, MPI, HPF, and mixed parallel/distributed programs. SCALEA instruments, executes and measures programs and computes a variety of performance overheads based on a novel overhead classification. Source code and HW-profiling is combined in a single system which significantly extends the scope of possible overheads that can be measured and examined, ranging from HW-counters, such as the number of cache misses or floating point operations, to more complex performance metrics, such as control or loss of parallelism. Moreover, SCALEA uses a new representation of code regions, called the dynamic code region call graph, which enables detailed overhead analysis for arbitrary code regions. An instrumentation description file is used to relate performance information to code regions of the input program and to reduce instrumentation overhead. Several experiments with realistic codes that cover MPI, OpenMP, HPF, and mixed OpenMP/MPI codes demonstrate the usefulness of SCALEA. Hong Linh Truong 0001, Thomas Fahringer, Georg Madsen, Allen D. Malony, Hans Moritsch, Sameer Shende |
SC | 4 |
| 2001 | Performance data mining: Automated diagnosis, adaption, and optimization
Alois Ferscha, Allen D. Malony |
Future Gener. Comput. Syst. | 2 |
| 2001 | A theory and architecture for automating performance diagnosis
Allen D. Malony, B. Robert Helm |
Future Gener. Comput. Syst. | 1 |
| 2000 | A Tool Framework for Static and Dynamic Analysis of Object-Oriented Software with TemplatesabstractThe developers of high-performance scientific applications often work in complex computing environments that place heavy demands on program analysis tools. The developers need tools that interoperate, are portable across machine architectures, and provide source-level feedback. In this paper, we describe a tool framework, the Program Database Toolkit (PDT), that supports the development of program analysis tools meeting these requirements. PDT uses compile-time information to create a complete database of high-level program information that is structured for well-defined and uniform access by tools and applications. PDT’s current applications make heavy use of advanced features of C++, in particular, templates. We describe the toolkit, focussing on its most important contribution -- its handling of templates -- as well as its use in existing applications. Kathleen A. Lindlan, Janice E. Cuny, Allen D. Malony, Sameer Shende, Bernd Mohr, Reid D. Rivenburgh, Craig Edward Rasmussen |
SC | 3 |
| 2000 | Computational experiments using distributed tools in a web-based electronic notebook environment
Allen D. Malony, Janice E. Cuny, Jenifer L. Skidmore, Matthew J. Sottile |
Future Gener. Comput. Syst. | 1 |
| 1999 | INTERLACE: An Interoperation and Linking Architecture for Computational Engines
Matthew J. Sottile, Allen D. Malony |
Euro-Par | 2 |
| 1999 | SMARTS: exploiting temporal locality and parallelism through vertical executionabstractArticle SMARTS: exploiting temporal locality and parallelism through vertical execution Share on Authors: Suvas Vajracharya Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Steve Karmesin Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Peter Beckman Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , James Crotinger Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Allen Malony Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NM Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NMView Profile , Sameer Shende Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NM Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NMView Profile , Rod Oldehoeft Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Stephen Smith Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 302–310https://doi.org/10.1145/305138.305207Online:01 May 1999Publication History 16citation354DownloadsMetricsTotal Citations16Total Downloads354Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Suvas Vajracharya, Steve Karmesin, Pete Beckman, James Crotinger, Allen D. Malony, Sameer Shende, R. R. Oldehoeft, Stephen Smith 0002 |
International Conference on Supercomputing | 5 |
| 1998 | Performance Evaluation and Prediction
Allen D. Malony, Rajeev Alur |
Euro-Par | 1 |
| 1998 | A Framework for Interacting with Distributed Programs and DataabstractThe Distributed Array Query and Visualization (DAQV) project aims to develop systems and tools that facilitate interacting with distributed programs and data structures. Arrays distributed across the processes of a parallel or distributed application are made available to external clients via well defined interfaces and protocols. Our design considers the broad issues of language targets, models of interaction, and abstractions for data access, while our implementation attempts to provide a general framework that can be adapted to a range of application scenarios. The paper describes the second generation of DAQV work and places it in the context of the more general distributed array access problem. Current applications and future work are also described. Steven T. Hackstadt, Christopher W. Harrop, Allen D. Malony |
HPDC | 3 |
| 1998 | Supporting Runtime Tool Interaction for Parallel SimulationsabstractScientists from many disciplines now routinely use modeling and simulation techniques to study physical and biological phenomena. Advances in high-performance architectures and networking have made it possible to build complex simulations with parallel and distributed interacting components. Unfortunately, the software needed to support such complex simulations has lagged behind hardware developments. We focus here on one aspect of such support: runtime program interaction. We have developed a runtime interaction framework and we have implemented a specific instance of it for an application in seismic tomography. That instance, called TierraLab, extends the geoscientists' existing (legacy) tomography code with runtime interaction capabilities which they access through a MATLAB interface. The scientist can stop a program, retrieve data, analyze and visualize that data with existing MATLAB routines, modify the data, and resume execution. They can do this all within a familiar MATLAB-like environment without having to be concerned with any of the low- level details of parallel or distributed data distribution. Data distribution is handled transparently by the Distributed Array Query and Visualization (DAQV) system. Our framework allows scientists to construct and maintain their own customized runtime interaction system. Christopher W. Harrop, Steven T. Hackstadt, Janice E. Cuny, Allen D. Malony, Laura S. Magde |
SC | 4 |
| 1998 | A Prototype Notebook-Based Environment for Computational Tools Computational ToolsabstractThe Virtual Notebook Environment (ViNE) is a platform-independent, web-based interface designed to support a range of scientific activities across distributed, heterogeneous computing platforms. ViNE provides scientists with a web-based version of the common paper-based lab notebook, but in addition, it provides support for collaboration and management of computational experiments. Collaboration is supported with the web-based approach, which makes notebook material generally accessible and with a hierarchy of security mechanisms that screen that access. ViNE provides uniform, system-transparent access to data, tools, and programs throughout the scientist's computing infrastructure. Computational experiments can be launched from ViNE using a visual specification language. The scientist is freed from concerns about inter-tool connectivity, data distribution, or data management details. ViNE also provides support for dynamically linking analysis results back into the notebook content. In this paper we present the ViNE system architecture and a case study of its use in neuropsychology research at the University of Oregon. Our case study with the Brain Electrophysiology Laboratory (BEL) addresses their need for data security and management, collaborative support, and distributed analysis processes. The current version of ViNE is a prototype system being tested with this and other scientific applications. Jenifer L. Skidmore, Matthew J. Sottile, Janice E. Cuny, Allen D. Malony |
SC | 4 |
| 1998 | DAQV: Distributed Array Query and Visualization Framework
Steven T. Hackstadt, Allen D. Malony |
Theor. Comput. Sci. | 2 |
| 1995 | Performance Extrapolation of Parallel Programs
Allen D. Malony, Kesavan Shanmugam |
ICPP (2) | 1 |
| 1995 | Data Interpretation and Experiment Planning in Performance Tools (Panel)abstractThe parallel scientific computing community is placing increasing emphasis on portability and scalability of programs, languages, and architectures. This creates new challenges for developers of parallel performance analysis tools, who will have to deal with increasing volumes of performance data drawn from diverse platforms. One way to meet this challenge is to incorporate sophisticated facilities for data interpretation and experiment planning within the tools themselves, giving them increased flexibility and autonomy in gathering and selecting performance data. This panel discussion brings together four research groups that have made advances in this direction. Allen D. Malony, B. Robert Helm, Jeffrey K. Hollingsworth, Barton P. Miller, Karsten Schwan |
SIGMETRICS | 1 |
| 1994 | Stochastic Modeling of Scaled Parallel ProgramsabstractTesting the performance scalability of parallel programs can be a time consuming task, involving many performance runs for different computer configurations, processor numbers, and problem sizes. Ideally, scalability issues would be addressed during parallel program design, but tools are not presently available that allow program developers to study the impact of algorithmic choices under different problem and system scenarios. Hence, scalability analysis is often reserved to existing (and available) parallel machines as well as implemented algorithms. In this paper we propose techniques for analyzing scaled parallel programs using stochastic modeling approaches. Although allowing more generality and flexibility in analysis, stochastic modeling of large parallel Allen D. Malony, Vassilis Mertsiotakis, Andreas Quick |
ICPADS | 1 |
| 1993 | Perturbation Analysis of High Level Instrumentation for SPMD ProgramsabstractThe process of instrumenting a program to study its behavior can lead to perturbations in the program's execution. These perturbations can become severe for large parallel systems or problem sizes, even when one captures only high level events. In this paper, we address the important issue of eliminating execution perturbations caused by high-level instrumentation of SPMD programs. We will describe perturbation analysis techniques for common computation and communication measurements, and show examples which demonstrate the effectiveness of these techniques in practice. Sekhar R. Sarukkai, Allen D. Malony |
PPoPP | 2 |
| 1993 | Implementing a parallel C++ runtime system for scalable parallel systemsabstractNo abstract available. François Bodin, Pete Beckman, Dennis Gannon, Shelby X. Yang, S. Kesavan, Allen D. Malony, Bernd Mohr |
SC | 6 |
| 1993 | Common runtime support for high-performance parallel languagesabstractNo abstract available. Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick |
SC | 4 |
| 1993 | Supercomputing around the world (Mini symposium)abstractArticle Supercomputing around the world (Mini symposium) Share on Authors: D. X. Kahaner Office of Naval Research Asia, 23-17, 7-chome, roppingi, Minato-ku, Tokyo 106 Japan Office of Naval Research Asia, 23-17, 7-chome, roppingi, Minato-ku, Tokyo 106 JapanView Profile , A. D. Malony Dept. of Comp. and Info. Science, University of Oregon, Eugene, Oregon Dept. of Comp. and Info. Science, University of Oregon, Eugene, OregonView Profile Authors Info & Claims Supercomputing '93: Proceedings of the 1993 ACM/IEEE conference on SupercomputingDecember 1993 Pages 874–876https://doi.org/10.1145/169627.169854Online:01 December 1993Publication History 0citation132DownloadsMetricsTotal Citations0Total Downloads132Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access David K. Kahaner, Allen D. Malony |
SC | 2 |
| 1992 | Supercomputing Around the WorldabstractIn keeping with the 'Voyages of Discovery' theme of the Supercomputing 1992 conference, representatives of supercomputing endeavours from around the world have met to speak on national and international supercomputing activities. This minisymposium brings together international representatives from five areas of the world to discuss supercomputing activities in countries that have been underrepresented at the Supercomputing conferences in the past. The topics discussed are high-performance computing and networking in Europe, the supercomputing environment in Taiwan, supercomputing in Australia, India's initiative in massively parallel supercomputing, and supercomputing in Brazil.> Allen D. Malony |
SC | 1 |
| 1992 | Performance Measurement Intrusion and Perturbation AnalysisabstractThe authors study the instrumentation perturbations of software event tracing on the Alliant FX/80 vector multiprocessor in sequential, vector, concurrent, and vector-concurrent modes. Based on experimental data, they derive a perturbation model that can approximate true performance from instrumented execution. They analyze the effects of instrumentation coverage, (i.e., the ratio of instrumented to executed statements), source level instrumentation, and hardware interactions. The results show that perturbations in execution times for complete trace instrumentations can exceed three orders of magnitude. With appropriate models of performance perturbation, these perturbations in execution time can be reduced to less than 20% while retaining the additional information from detailed traces. In general, it is concluded that it is possible to characterize perturbations through simple models. This permits more detailed, accurate instrumentation than traditionally believed possible.> Allen D. Malony, Daniel A. Reed, Harry A. G. Wijshoff |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1991 | Event-Based Performance Perturbation: A Case StudyabstractDeterminingthe performance behavior of parallel computations requires some form of intrusive tracing measurement.The greater the need for detailed performance data, the more intrusion the measurement will cause.Recovering actual execution performance jfrom perturbed performance measurements using eventbased perturbation analysis is the topic of this paper.We show that the measurement and subsequent analysis of synchronization operations (particularly, advance and await) can produce, in practice, accurate approximations to actual performance behavior.We use as testcases three Lawrence Livermore loops that execute as parallel DOACROSS loops on an Alliant FX/80.The results of our experiments suggest that a systematic application of performance perturbation analysis techniques will allow more detailed, accurate instrumentation than traditionally believed possible.1 Allen D. Malony |
PPoPP | 1 |
| 1991 | Tracing application program execution on the CRAY X-MP and CRAY-2
Allen D. Malony, John L. Larson, Daniel A. Reed |
J. Supercomput. | 1 |
| 1990 | A hardware-based performance monitor for the Intel iPSC/2 hypercubeabstractThe complexity of parallel computer systems makes a priori performance prediction difficult and experimental performance analysis crucial. A complete characterization of software and hardware dynamics, needed to understand the performance of high-performance parallel systems, requires execution time performance instrumentation. Although software recording of performance data suffices for low frequency events, capture of detailed, high-frequency performance data ultimately requires hardware support if the performance instrumentation is to remain efficient and unobtrusive. Allen D. Malony, Daniel A. Reed |
ICS | 1 |
| 1990 | Tracing application program execution on the Cray X-MP and Cray 2abstractA tracing library for the Cray X-MP and Cray 2 supercomputers has been developed that supports the low-overhead capture of execution events for sequential and multitasked programs. This library has been extended to use the automatic instrumentation facilities on these machines, allowing trace data from routine entry and exit, and other program segments to be captured. To assess the utility of the trace-based tools, three of the Perfect Benchmark codes have been tested in scalar and vector models with the tracing instrumentation. In addition to computing summary execution statistics from the traces, interesting execution dynamics appear when studying the trace histories. It is also possible to compare codes across the two architectures by correlating the event traces. It is concluded that adding tracing support in Cray supercomputers can have significant returns in improved performance characterization and evaluation.> Allen D. Malony, John L. Larson, Daniel A. Reed |
SC | 1 |
| 1990 | Run-time monitoring of concurrent programs on the Cedar multiprocessorabstractA prototype run-time performance monitoring environment for the Cedar multiprocessor has been developed. The authors describe how the Cedar performance monitoring infrastructure is designed to achieve run-time monitoring and distributed communication of parallel program execution data. The underlying tracing procedures used in Cedar to gather execution data are described, and the procedures by which traces are dynamically off-loaded from the Cedar system and sent to remote processes are explained. The application of these tools to capture and visualize matrix data from an application at run-time is discussed. Finally, results regarding the performance of the run-time monitoring system in in terms of execution data bandwidth and program perturbation are presented.> Allen D. Malony, Michael W. Berry, Priyamvada Sinvhal-Sharma |
SC | 2 |
| 1990 | Experimentally Characterizing the Behavior of Multiprocessor Memory Systems. A Case StudyabstractIt is demonstrated how the behavior of a cache-based multi-vector-processor memory system can be systematically characterized and its performance experimentally correlated with key features of the address stream. The approach is based on the definition of a family of parameterized kernels used to explore specific aspects of the memory system's performance. The empirical results from this kernel suite provide the data from which architectural or algorithmic characteristics can be studied. The results of applying the approach to an Alliant FX/8 are presented and evaluated.> Kyle A. Gallivan, Dennis Gannon, William Jalby, Allen D. Malony, Harry A. G. Wijshoff |
IEEE Trans. Software Eng. | 4 |
| 1989 | Performance prediction of loop constructs on multiprocessor hierarchical-memory systemsabstractIn this paper we discuss the performance prediction of Fortran constructs commonly found in numerical scientific computing. Although the approach is applicable to multi-processors in general, within the scope of the paper we will concentrate on the Alliant FX/8 multiprocessor. The techniques proposed involve a combination of empirical observations, architectural models and analytical techniques, and exploits earlier work on data locality analysis and empirical characterization of the behavior of memory systems. The Lawrence Livermore Loops are used as a test-case to verify the approach. Kyle A. Gallivan, William Jalby, Allen D. Malony, Harry A. G. Wijshoff |
ICS | 3 |
| 1989 | Behavioral Characterization of Multiprocessor Memory Systems: A Case StudyabstractThe speed and efficiency of the memory system is a key limiting factor in the performance of supercomputers. Consequently, one of the major concerns when developing a high-performance code, either manually or automatically, is determining and characterizing the influence of the memory system on performance in terms of algorithmic parameters. Unfortunately, the performance data available to an algorithm designer such as various benchmarks and, occasionally, manufacturer-supplied information, e.g. instruction timings and architecture component characteristics, are rarely sufficient for this task. In this paper, we discuss a systematic methodology for probing the performance characteristics of a memory system via a hierarchy of data-movement kernels. We present and analyze the results obtained by such a methodology on a cache-based multi-vector processor (Alliant FX/8). Finally, we indicate how these experimental results can be used for predicting the performance of simple Fortran codes by a combination of empirical observations, architectural models and analytical techniques. Kyle A. Gallivan, Dennis Gannon, William Jalby, Allen D. Malony, Harry A. G. Wijshoff |
SIGMETRICS | 4 |
| 1988 | Parallel Discrete Event Simulation Using Shared MemoryabstractWith traditional event-list techniques, evaluating a detailed discrete event simulation-model can often require hours or even days of computation time. By eliminating the event list and maintaining only sufficient synchronization to ensure causality, parallel simulation can potentially provide speedups that are linear in the numbers of processors. A set of shared-memory experiments using the Chandy-Misra distributed simulation algorithm, to simulate networks of queues is presented. Parameters of the study include queueing network topology and routing probabilities, number of processors, and assignment of network nodes to processors. These experiments show that Chandy-Misra distributed simulation is a questionable alternative to sequential simulation of most queuing network models.> Daniel A. Reed, Allen D. Malony, Bradley D. McCredie |
IEEE Trans. Software Eng. | 2 |
| 1987 | MPF: A Portable Message Passing Facility for Shared Memory Multiprocessors
Daniel A. Reed, Allen D. Malony, Patrick J. McGuire |
ICPP | 2 |
| 1987 | Parallel Discrete Event Simulation: A Shared Memory ApproachabstractWith traditional event list techniques, evaluating a detailed discrete event simulation model can often require hours or even days of computation time. Parallel simulation mimics the interacting servers and queues of a real system by assigning each simulated entity to a processor. By eliminating the event list and maintaining only sufficient synchronization to insure causality, parallel simulation can potentially provide speedups that are linear in the number of processors. A set of shared memory experiments is presented using the Chandy-Misra distributed simulation algorithm to simulate networks of queues. Parameters include queueing network topology and routing probabilities, number of processors, and assignment of network nodes to processors. These experiments show that Chandy-Misra distributed simulation is a questionable alternative to sequential simulation of most queueing network models. Daniel A. Reed, Allen D. Malony, Bradley D. McCredie |
SIGMETRICS | 2 |
| 1986 | Vector Processing on the Alliant FX/8 Multiprocessor
Walid A. Abu-Sufah, Allen D. Malony |
ICPP | 2 |