Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Boyana Norris

dblp:26/2418 · DBLP profile ↗
← Back
30ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0001-5811-9731ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 37% Parallel and multicore computing · 21% GPUs and heterogeneous computing · 21%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 50% Computational geometry · 50%
Software engineering, system software, and programming languages
3 papers
Operating systems · 86% Software maintenance and evolution · 11% Compilers and program optimization · 3%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › kernel › kernel design
microkernel
0.812024
Manipulative Interference Attacks · CCS 2024
GPUs and heterogeneous computing › GPU graph processing
GPU graph algorithms
0.612022
A Parallel Algorithm Template for Updating Single-Source Shortest Paths in Large-Scale Dynamic Networks · IEEE Trans. Parallel Distributed Syst. 2022
Parallel and multicore computing
parallel graph algorithms
0.612022
A Parallel Algorithm Template for Updating Single-Source Shortest Paths in Large-Scale Dynamic Networks · IEEE Trans. Parallel Distributed Syst. 2022
Computational geometry › geometric data structures › shortest path queries
dynamic shortest paths
0.612022
A Parallel Algorithm Template for Updating Single-Source Shortest Paths in Large-Scale Dynamic Networks · IEEE Trans. Parallel Distributed Syst. 2022
Graph algorithms and graph theory
shortest path
0.612022
A Parallel Algorithm Template for Updating Single-Source Shortest Paths in Large-Scale Dynamic Networks · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing › performance optimization
auto-tuning
0.312018
Autotuning in High-Performance Computing Applications · Proc. IEEE 2018
High-performance computing
performance optimization
0.312018
Autotuning in High-Performance Computing Applications · Proc. IEEE 2018
High-performance computing › performance engineering
performance portability
0.312018
Autotuning in High-Performance Computing Applications · Proc. IEEE 2018
Embedded and real-time systems › real-time scheduling › mixed-criticality scheduling
mixed-criticality systems
0.212024
Manipulative Interference Attacks · CCS 2024
Distributed systems
dynamic network
0.212022
A Parallel Algorithm Template for Updating Single-Source Shortest Paths in Large-Scale Dynamic Networks · IEEE Trans. Parallel Distributed Syst. 2022
Performance modeling and evaluation
automated performance analysis
0.112008
Capturing performance knowledge for automated analysis · SC 2008
Performance modeling and evaluation
performance diagnosis
0.112008
Capturing performance knowledge for automated analysis · SC 2008
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.012008
Capturing performance knowledge for automated analysis · SC 2008

Methods — techniques the papers use, named apart from their topics

formal verification · 1.5rooted tree data structure · 1.1parallel algorithm template · 1.1performance modeling · 0.8empirical measurement · 0.7data mining · 0.2
YearPublicationVenuePosition
2024 Manipulative Interference Attacks
abstract
A μ-kernel is an operating system (OS) paradigm that facilitates a strong cybersecurity posture for embedded systems. Unlike a monolithic OS such as Linux, a μ-kernel reduces overall system privilege by deploying most OS functionality within isolated, userspace protection domains. Moreover, a μ-kernel ensures confidentiality and integrity between protection domains (i.e., spatial isolation), and offers timing predictability for real-time tasks in mixed-criticality systems (i.e., temporal isolation). One popular μ-kernel is seL4 which offers extensive formal guarantees of implementation correctness and flexible temporal budgeting mechanisms.
Samuel Mergendahl, Stephen Fickas, Boyana Norris, Richard Skowyra
CCS3
2024 SOMA: Observability, monitoring, and in situ analytics for exascale applications
abstract
Summary With the rise of exascale systems and large, data‐centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service‐based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application‐specific diagnostic data and performance data addresses these needs. We present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi‐GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring.
Dewi Yokelson, Oskar Lappi, Srinivasan Ramesh, Miikka S. Väisälä, Kevin A. Huck, Touko Puro, Boyana Norris, Maarit J. Korpi-Lagg, Keijo Heljanko, Allen D. Malony
Concurr. Comput. Pract. Exp.7
2023 A Distributed Algorithm for Identifying Strongly Connected Components on Incremental Graphs
abstract
Incremental graphs that change over time capture the changing relationships of different entities. Given that many real-world networks are extremely large, it is often necessary to partition the network over many distributed systems and solve a complex graph problem over the partitioned network. This paper presents a distributed algorithm for identifying strongly connected components (SCC) on incremental graphs. We propose a two-phase asynchronous algorithm that involves storing the intermediate results between each iteration of dynamic updates in a novel meta-graph storage format for efficient recomputation of the SCC for successive iterations. To the best of our knowledge, this is the first attempt at identifying SCC for incremental graphs across distributed compute nodes. Our experimental analysis on real and synthesized graphs shows up to 2.8x performance improvement over the state-of-the-art by reducing the overall memory utilized and improving the communication bandwidth.
Arindam Khanda, Sajal K. Das 0001, Sanjukta Bhowmick, Boyana Norris
SBAC-PAD7
2022 A Shared-Memory Algorithm for Updating Tree-Based Properties of Large Dynamic Networks
abstract
This paper presents a network-based template for analyzing large-scale dynamic data. Specifically, we propose a novel shared-memory parallel algorithm for updating tree-based structures or properties, such as connected components (CC) and minimum spanning trees (MST), on dynamic networks. The underlying idea is to update the information in a rooted tree data structure that stores the edges of the network that are most relevant to the analysis. Extensive experiments on real-world and synthetic networks demonstrate that, with the exception of the inherently sequential component for creating the rooted tree, our proposed updatiing algorithm is scalable and, in most cases, also requires significantly less memory, energy, and time than recomputing-from-scratch algorithm. To the best of our knowledge, this is the first parallel algorithm for updating MST on weighted dynamic networks. The rooted-tree based framework that we propose in this paper can be extended for updating other weighted and unweighted tree-based properties such as single source shortest path and betweenness and closeness centrality.
Sriram Srinivasan 0001, Samuel Pollard, Boyana Norris, Sajal K. Das 0001, Sanjukta Bhowmick
IEEE Trans. Big Data3
2022 A Parallel Algorithm Template for Updating Single-Source Shortest Paths in Large-Scale Dynamic Networks
abstract
The Single Source Shortest Path (SSSP) problem is a classic graph theory problem that arises frequently in various practical scenarios; hence, many parallel algorithms have been developed to solve it. However, these algorithms operate on static graphs, whereas many real-world problems are best modeled as dynamic networks, where the structure of the network changes with time. This gap between the dynamic graph modeling and the assumed static graph model in the conventional SSSP algorithms motivates this work. We present a novel parallel algorithmic framework for updating the SSSP in large-scale dynamic networks and implement it on the shared-memory and GPU platforms. The basic idea is to identify the portion of the network affected by the changes and update the information in a rooted tree data structure that stores the edges of the network that are most relevant to the analysis. Extensive experimental evaluations on real-world and synthetic networks demonstrate that our proposed parallel updating algorithm is scalable and, in most cases, requires significantly less execution time than the state-of-the-art recomputing-from-scratch algorithms.
Arindam Khanda, Sriram Srinivasan 0001, Sanjukta Bhowmick, Boyana Norris, Sajal K. Das 0001
IEEE Trans. Parallel Distributed Syst.4
2021 Empirical Investigation of Code Quality Rule Violations in HPC Applications
abstract
In large, collaborative open-source projects, developers must follow good coding standards to ensure the quality and sustainability of the resulting software. This is especially a challenge in high-performance computing projects, which admit a diverse set of contributions over decades of development. Some successful projects, such as the Portable, Extensible Toolkit for Scientific Computation (PETSc), have created comprehensive developer documentation, including specific code quality rules, which should be followed by contributors. However, none of the widely used and highly active open-source HPC projects have a way to automatically check whether these rules, typically expressed informally in English, are being violated. Hence, compliance checking is labor-intensive and difficult to ensure. To address this issue, we propose an automated method for detecting rule violations in HPC applications based on the PETSc development rules. In our empirical study, we consider 46 PETSc-based applications and assess the violations of two C-usage rules. The experimental results demonstrate the efficacy of the proposed method in identifying PETSc rule violations, which can be broadened to other HPC frameworks and extended by us and others in the community to include more rules.
Shahid Hussain 0001, Kaley Chicoine, Boyana Norris
EASE3
2018 Single-Source Shortest Path Tree for Big Dynamic Graphs
abstract
Computing single-source shortest paths (SSSP) is one of the fundamental problems in graph theory. There are many applications of SSSP including finding routes in GPS systems and finding high centrality vertices for effective vaccination. In this paper, we focus on calculating SSSP on big dynamic graphs, which change with time. We propose a novel distributed computing approach, SSSPIncJoint, to update SSSP on big dynamic graphs using GraphX. Our approach considerably speeds up the recomputation of the SSSP tree by reducing the number of map-reduce operations required for implementing SSSP in the gather-apply- scatter programming model used by GraphX.
Sara Riazi, Sriram Srinivasan 0001, Sajal K. Das 0001, Sanjukta Bhowmick, Boyana Norris
IEEE BigData5
2018 A Shared-Memory Parallel Algorithm for Updating Single-Source Shortest Paths in Large Dynamic Networks
abstract
Computing the single-source shortest path (SSSP) is one of the fundamental graph algorithms, and is used in many applications. Here, we focus on computing SSSP on large dynamic graphs, i.e. graphs whose structure evolves with time. We posit that instead of recomputing the SSSP for each set of changes on the dynamic graphs, it is more efficient to update the results based only on the region of change. To this end, we present a novel two-step shared-memory algorithm for updating SSSP on weighted large-scale graphs. The key idea of our algorithm is to identify changes, such as vertex/edge addition and deletion, that affect the shortest path computations and update only the parts of the graphs affected by the change. We provide the proof of correctness of our proposed algorithm. Our experiments on real and synthetic networks demonstrate that our algorithm is as much as 4X faster compared to computing SSSP with Galois, a state-of-the-art parallel graph analysis software for shared memory architectures. We also demonstrate how increasing the asynchrony can lead to even faster updates. To the best of our knowledge, this is one of the first practical parallel algorithms for updating networks on shared-memory systems, that is also scalable to large networks.
Sriram Srinivasan 0001, Sara Riazi, Boyana Norris, Sajal K. Das 0001, Sanjukta Bhowmick
HiPC3
2018 Autotuning in High-Performance Computing Applications
abstract
Autotuning refers to the automatic generation of a search space of possible implementations of a computation that are evaluated through models and/or empirical measurement to identify the most desirable implementation. Autotuning has the potential to dramatically improve the performance portability of petascale and exascale applications. To date, autotuning has been used primarily in high-performance applications through tunable libraries or previously tuned application code that is integrated directly into the application. This paper draws on the authors' extensive experience applying autotuning to high-performance applications, describing both successes and future challenges. If autotuning is to be widely used in the HPC community, researchers must address the software engineering challenges, manage configuration overheads, and continue to demonstrate significant performance gains and portability across architectures. In particular, tools that configure the application must be integrated into the application build process so that tuning can be reapplied as the application and target architectures evolve.
Prasanna Balaprakash, Jack J. Dongarra, Todd Gamblin, Mary W. Hall, Jeffrey K. Hollingsworth, Boyana Norris, Richard W. Vuduc
Proc. IEEE6
2017 Mira: A Framework for Static Performance Analysis
abstract
The performance model of an application can provide understanding about its runtime behavior on particular hardware. Such information can be analyzed by developers for performance tuning. However, model building and analyzing is frequently ignored during software development until performance problems arise because they require significant expertise and can involve many time-consuming application runs. In this paper, we propose a fast, accurate, flexible and user-friendly tool, Mira, for generating performance models by applying static program analysis, targeting scientific applications running on supercomputers. We parse both the source code and binary to estimate performance attributes with better accuracy than considering just source or just binary code. Because our analysis is static, the target program does not need to be executed on the target architecture, which enables users to perform analysis on available machines instead of conducting expensive experiments on potentially expensive resources. Moreover, statically generated models enable performance prediction on nonexistent or unavailable architectures. In addition to flexibility, because model generation time is significantly reduced compared to dynamic analysis approaches, our method is suitable for rapid application performance analysis and improvement. We present empirical validation results to demonstrate the current capabilities of our approach on small benchmarks and a mini application.
Kewen Meng, Boyana Norris
CLUSTER2
2017 A Comparison of Parallel Graph Processing Implementations
abstract
The rapidly growing number of large network analysis problems has led to the emergence of many parallel and distributed graph processing systems-one survey in 2014 identified over 80. Determining the best approach for a given problem is infeasible for most developers. We present an approach and associated software for analyzing the performance and scalability of parallel, open-source graph libraries. We demonstrate our approach on five graph processing packages: GraphMat, Graph500, Graph Algorithm Platform Benchmark Suite, GraphBIG, and PowerGraph using synthetic and real-world datasets. We examine previously overlooked aspects of parallel graph processing performance such as phases of execution and energy usage for three algorithms: breadth first search, single source shortest paths, and PageRank.
Samuel Pollard, Boyana Norris
CLUSTER2
2017 Autotuning GPU Kernels via Static and Predictive Analysis
abstract
Optimizing the performance of GPU kernels is challenging for both human programmers and code generators. For example, CUDA programmers must set thread and block parameters for a kernel, but might not have the intuition to make a good choice. Similarly, compilers can generate working code, but may miss tuning opportunities by not targeting GPU models or performing code transformations. Although empirical autotuning addresses some of these challenges, it requires extensive experimentation and search for optimal code variants. This research presents an approach for tuning CUDA kernels based on static analysis that considers fine-grained code structure and the specific GPU architecture features. Notably, our approach does not require any program runs in order to discover near-optimal parameter settings. We demonstrate the applicability of our approach in enabling code autotuners such as Orio to produce competitive code variants comparable with empirical-based methods, without the high cost of experiments.
Robert V. Lim, Boyana Norris, Allen D. Malony
ICPP2
2017 Performance Analysis of Applications in the Context of Architectural Rooflines
abstract
Intuitive visual representations of architecture capabilities and the performance of applications are critical to enabling effective performance analysis, which in turn guides optimizations. The Roofline Model and its derivatives provide such an intuitive representation of the best achievable performance on a given architecture. The Roofline Toolkit project is a collaboration among researchers at Argonne National Laboratory, Lawrence Berkeley National Laboratory, and the University of Oregon and consists of three principal components: hardware characterization, software characterization, and data manipulation, which includes a visualization interface. These components address the different aspects of performance data acquisition and manipulation required for performance analysis, modeling and optimization of applications. In this paper we introduce an implementation of the third component, a system for visualizing roofline charts and managing roofline performance analysis data. We demonstrate analysis of an application use case within this framework and outline future directions for this type of performance analysis and visualization.
Boyana Norris, Wyatt Spear, Allen D. Malony
ICPE1
2016 GraphFlow: Workflow-based big graph processing
abstract
We introduce GraphFlow, a big graph framework that is able to encode complex data science experiments as a set of high-level workflows. GraphFlow combines the Spark big data processing platform and the Galaxy workflow management system to offer a set of components for graph processing using a novel interaction model for creating and using complex workflows. GraphFlow contributes an easy-to-use interface and scalable algorithms for big graph analytics. We demonstrate GraphFlow use in large social network analysis with several case studies.
Sara Riazi, Boyana Norris
IEEE BigData2
2015 WattProf: A Flexible Platform for Fine-Grained HPC Power Profiling
abstract
The ability to monitor power and energy accurately and efficiently is becoming increasingly important in all areas of computing. In this paper, we introduce WattProf, a portable and flexible power monitoring platform that allows for fine-grained monitoring of power and energy consumption in various hardware components of HPC clusters. The WattProf system consist of a programmable monitoring PCIe expansion card, sensors for individual system components, host runtime system, and a high-level application programming interface. We demonstrate the use of WattProf on two benchmarks.
Mohammad J. Rashti, Gerald Sabin, David Vansickle, Boyana Norris
CLUSTER4
2015 Generating Efficient Tensor Contractions for GPUs
abstract
Many scientific and numerical applications, including quantum chemistry modeling and fluid dynamics simulation, require tensor product and tensor contraction evaluation. Tensor computations are characterized by arrays with numerous dimensions, inherent parallelism, moderate data reuse and many degrees of freedom in the order in which to perform the computation. The best-performing implementation is heavily dependent on the tensor dimensionality and the target architecture. In this paper, we map tensor computations to GPUs, starting with a high-level tensor input language and producing efficient CUDA code as output. Our approach is to combine tensor-specific mathematical transformations with a GPU decision algorithm, machine learning and auto tuning of a large parameter space. Generated code shows significant performance gains over sequential and Open MP parallel code, and a comparison with Open ACC shows the importance of auto tuning and other optimizations in our framework for achieving efficient results.
Thomas Nelson, Axel Rivera, Prasanna Balaprakash, Mary W. Hall, Paul D. Hovland, Elizabeth R. Jessup, Boyana Norris
ICPP7
2015 Reliable Generation of High-Performance Matrix Algebra
abstract
Scientific programmers often turn to vendor-tuned Basic Linear Algebra Subprograms (BLAS) to obtain portable high performance. However, many numerical algorithms require several BLAS calls in sequence, and those successive calls do not achieve optimal performance. The entire sequence needs to be optimized in concert. Instead of vendor-tuned BLAS, a programmer could start with source code in Fortran or C (e.g., based on the Netlib BLAS) and use a state-of-the-art optimizing compiler. However, our experiments show that optimizing compilers often attain only one-quarter of the performance of hand-optimized code. In this article, we present a domain-specific compiler for matrix kernels, the Build to Order BLAS (BTO), that reliably achieves high performance using a scalable search algorithm for choosing the best combination of loop fusion, array contraction, and multithreading for data parallelism. The BTO compiler generates code that is between 16% slower and 39% faster than hand-optimized code.
Thomas Nelson, Geoffrey Belter, Jeremy G. Siek, Elizabeth R. Jessup, Boyana Norris
ACM Trans. Math. Softw.5
2014 Toward multi-target autotuning for accelerators
abstract
Producing high-performance implementations from simple, portable computation specifications is a challenge that compilers have tried to address for several decades. More recently, a relatively stable architectural landscape has evolved into a set of increasingly diverging and rapidly changing CPU and accelerator designs, with the main common factor being dramatic increases in the levels of parallelism available. The growth of architectural heterogeneity and parallelism, combined with the very slow development cycles of traditional compilers, has motivated the development of autotuning tools that can quickly respond to changes in architectures and programming models, and enable very specialized optimizations that are not possible or likely to be provided by mainstream compilers. In this paper we describe the new OpenCL code generator and autotuner OrCL and the introduction of detailed performance measurement into the autotuning process. OrCL is implemented within the Orio autotuning framework, which enables the rapid development of experimental languages and code optimization strategies aimed at achieving good performance on new platforms without rewriting or hand-optimizing critical kernels. The combination of the new OpenCL autotuning and TAU measurement capabilities enables users to consistently evaluate autotuning effectiveness across a range of architectures, including several NVIDIA and AMD accelerators and Intel Xeon Phi processors, and to compare the OpenCL and CUDA code generation capabilities. We present results of autotuning several numerical kernels that typically dominate the execution time of iterative sparse linear system solution and key computations from a 3-D parallel simulation of solid fuel ignition.
Nicholas Chaimov, Boyana Norris, Allen D. Malony
ICPADS2
2013 Exascale workload characterization and architecture implications
abstract
Emerging exascale architectures bring forth new challenges related to heterogeneous systems power, energy, cost, and resilience. These new challenges require a shift from conventional paradigms in understanding how to best exploit and optimize these features and limitations. Our objective is to identify the top few dominant characteristics in a set of applications. Understanding these characteristics will allow the community to build and exploit customized architectures and tools best suited to optimize each dominant characteristic in the application domain. Every application will typically be composed of multiple characteristics and thus will use several of the customized accelerators and tools during its execution phases, with the eventual goal of using the entire system efficiently. In this poster, we describe a hybrid methodology, based on binary instrumentation, for characterizing scientific applications such as instruction mix and memory access patterns. We apply our methodology to proxy applications that are representative of a broad range of DOE scientific applications. With this empirical basis, we develop and validate statistical models that extrapolate application properties as a function of problem size. These models are then used to project the first quantitative characterization of an exascale computing workload, including computing and memory requirements. We evaluate the potential benefit of processor under memory, a radical new exascale architecture customization and understand how these new customization can impact applications.
Prasanna Balaprakash, Darius Buntinas, Apala Guha, Rinku Gupta, Sri Hari Krishna Narayanan, Andrew A. Chien, Paul D. Hovland, Boyana Norris
ISPASS9
2012 Autotuning Stencil-Based Computations on GPUs
abstract
Finite-difference, stencil-based discretization approaches are widely used in the solution of partial differential equations describing physical phenomena. Newton-Krylov iterative methods commonly used in stencil-based solutions generate matrices that exhibit diagonal sparsity patterns. To exploit these structures on modern GPUs, we extend the standard diagonal sparse matrix representation and define new matrix and vector data types in the PETSc parallel numerical toolkit. We create tunable CUDA implementations of the operations associated with these types after identifying a number of GPU-specific optimizations and tuning parameters for these operations. We discuss our implementation of GPU auto tuning capabilities in the Orio framework and present performance results for several kernels, comparing them with vendor-tuned library implementations.
Azamat Mametjanov, Daniel Lowell, Ching-Chen Ma, Boyana Norris
CLUSTER4
2009 Parametric multi-level tiling of imperfectly nested loops
abstract
Tiling is a crucial loop transformation for generating high performance code on modern architectures. Efficient generation of multi-level tiled code is essential for maximizing data reuse in systems with deep memory hierarchies. Tiled loops with parametric tile sizes (not compile-time constants) facilitate runtime feedback and dynamic optimizations used in iterative compilation and automatic tuning. Previous parametric multi-level tiling approaches have been restricted to perfectly nested loops, where all assignment statements are contained inside the innermost loop of a loop nest. Previous solutions to tiling for imperfect loop nests have only handled fixed tile sizes. In this paper, we present an approach to parametric multi-level tiling of imperfectly nested loops. The tiling technique generates loops that iterate over full rectangular tiles, making them amenable to compiler optimizations such as register tiling. Experimental results using a number of computational benchmarks demonstrate the effectiveness of the developed tiling approach.
Albert Hartono, Muthu Manikandan Baskaran, Cédric Bastoul, Albert Cohen 0001, Sriram Krishnamoorthy, Boyana Norris, J. Ramanujam, P. Sadayappan
ICS6
2009 Annotation-based empirical performance tuning using Orio
abstract
For many scientific applications, significant time is spent in tuning codes for a particular high-performance architecture. Tuning approaches range from the relatively nonintrusive (e.g., by using compiler options) to extensive code modifications that attempt to exploit specific architecture features. Intrusive techniques often result in code changes that are not easily reversible, and can negatively impact readability, maintainability, and performance on different architectures. We introduce an extensible annotation-based empirical tuning system called Orio that is aimed at improving both performance and productivity. It allows software developers to insert annotations in the form of structured comments into their source code to trigger a number of low-level performance optimizations on a specified code fragment. To maximize the performance tuning opportunities, the annotation processing infrastructure is designed to support both architecture-independent and architecture-specific code optimizations. Given the annotated code as input, Orio generates many tuned versions of the same operation and empirically evaluates the alternatives to select the best performing version for production use. We have also enabled the use of the Pluto automatic parallelization tool in conjunction with Orio to generate efficient OpenMP-based parallel code. We describe our experimental results involving a number of computational kernels, including dense array and sparse matrix operations.
Albert Hartono, Boyana Norris, P. Sadayappan
IPDPS2
2008 Capturing performance knowledge for automated analysis
abstract
Automating the process of parallel performance experimentation, analysis, and problem diagnosis can enhance environments for performance-directed application development, compilation, and execution. This is especially true when parametric studies, modeling, and optimization strategies require large amounts of data to be collected and processed for knowledge synthesis and reuse. This paper describes the integration of the PerfExplorer performance data mining framework with the OpenUH compiler infrastructure. OpenUH provides auto-instrumentation of source code for performance experimentation and PerfExplorer provides automated and reusable analysis of the performance data through a scripting interface. More importantly, PerfExplorer inference rules have been developed to recognize and diagnose performance characteristics important for optimization strategies and modeling. Three case studies are presented which show our success with automation in OpenMP and MPI code tuning, parametric characterization, Pand power modeling. The paper discusses how the integration supports performance knowledge engineering across applications and feedback-based compiler optimization in general.
Kevin A. Huck, Oscar R. Hernandez, Van Bui, Sunita Chandrasekaran, Barbara M. Chapman, Allen D. Malony, Lois C. McInnes, Boyana Norris
SC8
2007 Component Specification for Parallel Coupling Infrastructure
Jay Walter Larson, Boyana Norris
ICCSA (3)2
2005 Making automatic differentiation truly automatic: coupling PETSc with ADIC
Paul D. Hovland, Boyana Norris, Barry Smith 0002
Future Gener. Comput. Syst.2
2004 Faster PDE-based simulations using robust composite linear solvers
Sanjukta Bhowmick, Padma Raghavan, Lois C. McInnes, Boyana Norris
Future Gener. Comput. Syst.4
2003 The Role of Multi-method Linear Solvers in PDE-based Simulations
Sanjukta Bhowmick, Lois C. McInnes, Boyana Norris, Padma Raghavan
ICCSA (1)3
2002 Implementation of automatic differentiation tools
abstract
Automatic differentiation is a semantic transformation that applies the rules of differential calculus to source code. It thus transforms a computer program that computes a mathematical function into a program that computes the function and its derivatives. Derivatives play an important role in a wide variety of scientific computing applications, including optimization, solution of nonlinear equations, sensitivity analysis, and nonlinear inverse problems. We describe a simple component architecture for developing tools for automatic differentiation and other mathematically oriented semantic transformations of scientific software. This architecture consists of a compiler-based, language-specific front-end for source transformation, loosely coupled with one or more language-independent "plug-in" transformation modules. The coupling mechanism between the front-end and transformation modules is provided by the XML Abstract Interface Form (XAIF). XAIF provides an abstract, language-independent representation of language constructs common in imperative languages, such as C and Fortran. We describe the use of this architecture in constructing tools for automatic differentiation of Fortran 77 and ANSI C, and we discuss how access to compiler optimization techniques can enable more efficient derivative augmentation.
Christian H. Bischof, Paul D. Hovland, Boyana Norris
PEPM3
2002 Parallel components for PDEs and optimization: some issues and experiences
Boyana Norris, Satish Balay, Steven Benson, Lori A. Diachin, Paul D. Hovland, Lois C. McInnes, Barry Smith 0002
Parallel Comput.1
2001 A Distributed Application Server for Automatic Differentiation
Boyana Norris, Paul D. Hovland
IPDPS1