Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Timothy G. Mattson

dblp:56/794 · also Tim Mattson · DBLP profile ↗
← Back
30ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0002-6106-8717ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 9 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorArtificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 56% Storage systems · 21% Distributed systems · 11%
Software engineering, system software, and programming languages
2 papers
Debugging and program repair · 56% Software testing · 28% Empirical software engineering · 8%
Databases, data mining, and information retrieval
2 papers
Data integration and cleaning · 83% Data stream processing · 8% Distributed and cloud data management · 8%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data pipeline debugging
0.412020
Debugging Large-Scale Data Science Pipelines using Dagger · Proc. VLDB Endow. 2020
Parallel and multicore computing
parallel programming models
0.422018
The Ongoing Evolution of OpenMP · Proc. IEEE 2018
Parallel programming: can we PLEASE get it right this time? · DAC 2008
Debugging and program repair
fault localization
0.412019
A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions · NeurIPS 2019
Debugging and program repair › performance debugging
performance bug diagnosis
0.412019
A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions · NeurIPS 2019
Software testing › regression testing
performance regression testing
0.412019
A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions · NeurIPS 2019
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.312018
The Ongoing Evolution of OpenMP · Proc. IEEE 2018
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.312018
The Ongoing Evolution of OpenMP · Proc. IEEE 2018
Distributed systems › transaction processing
atomicity
0.212016
The TileDB Array Data Storage Manager · Proc. VLDB Endow. 2016
Storage systems › data layout
multidimensional array storage
0.212016
The TileDB Array Data Storage Manager · Proc. VLDB Endow. 2016
Storage systems
storage reliability
0.212016
The TileDB Array Data Storage Manager · Proc. VLDB Endow. 2016
Data integration and cleaning › interoperability › database interoperability
polystore
0.212015
A Demonstration of the BigDAWG Polystore System · Proc. VLDB Endow. 2015
Processor architecture and microarchitecture
many-core architecture
0.222010
The 48-core SCC Processor: the Programmer's View · SC 2010
Programming the Intel 80-core network-on-a-chip terascale processor · SC 2008
Program analysis
performance analysis tools
0.112018
The Ongoing Evolution of OpenMP · Proc. IEEE 2018
Parallel and multicore computing › parallel programming models
message passing
0.112008
Programming the Intel 80-core network-on-a-chip terascale processor · SC 2008
High-performance computing
scientific data management
0.112016
The TileDB Array Data Storage Manager · Proc. VLDB Endow. 2016
Data stream processing
streaming analytics
0.112015
A Demonstration of the BigDAWG Polystore System · Proc. VLDB Endow. 2015
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.012010
The 48-core SCC Processor: the Programmer's View · SC 2010
Energy-efficient computing
power management
0.012008
Programming the Intel 80-core network-on-a-chip terascale processor · SC 2008

Methods — techniques the papers use, named apart from their topics

zero-positive learning · 0.4hardware telemetry · 0.4autoencoder · 0.4data visualization · 0.2cross-storage-system queries · 0.2
YearPublicationVenuePosition
2025 What Quantum Can Learn from Classical Computer Engineering
abstract
Quantum computing represents a paradigm shift requiring reconceptualization of algorithms, architectures, and software. Although much is new, there is much that quantum computing can learn from traditional classical computer engineering. In this Special Issue, we focus on examples of quantum research inspired by classical computer engineering, ranging from quantum machine learning and image processing to the extension of the “multi-core” concept to quantum computing. Hardware/software co-design is difficult when communities have their unique jargon and disjoint problem domains. For parallel programming on classical computers, an approach for communicating application developer needs emerged around identifying key bottlenecks that limited software. It was difficult to guess features of future applications, but experienced developers knew features of a computer system that would limit software performance. For each feature, a small kernel was developed with the idea that if the kernel ran well, a system would be an effective system for application developers. This set of kernels was called the Parallel Research Kernels . They have been useful to guide design of parallel computers. We hypothesize that an analogous tool for quantum computing, called Quantum Research Kernels , may aid hardware/software co-design of quantum computing systems. In this article, we motivate the Quantum Research Kernels and provide working examples.
Anne Y. Matsuura, Timothy G. Mattson
ACM Trans. Quantum Comput.2
2024 Distributed Ranges: A Model for Distributed Data Structures, Algorithms, and Views
abstract
Data structures and algorithms are essential building blocks for programs, and distributed data structures, which automatically partition data across multiple memory locales, are essential to writing high-level parallel programs. While many projects have designed and implemented C++ distributed data structures and algorithms, there has not been widespread adoption of an interoperable model allowing algorithms and data structures from different libraries to work together. This paper introduces distributed ranges, which is a model for building generic data structures, views, and algorithms. A distributed range extends a C++ range, which is an iterable sequence of values, with a concept of segmentation, thus exposing how the distributed range is partitioned over multiple memory locales. Distributed data structures provide this distributed range interface, which allows them to be used with a collection of generic algorithms implemented using the distributed range interface. The modular nature of the model allows for the straightforward implementation of distributed views, which are lightweight objects that provide a lazily evaluated view of another range. Views can be composed together recursively and combined with algorithms to implement computational kernels using efficient, flexible, and high-level standard C++ primitives. We evaluate the distributed ranges model by implementing a set of standard concepts and views as well as two execution runtimes, a multi-node, MPI-based runtime and a single-process, multi-GPU runtime. We demonstrate that high-level algorithms implemented using generic, high-level distributed ranges can achieve performance competitive with highly-tuned, expert-written code.
Benjamin Brock, Robert S. Cohn, Suyash Bakshi, Tuomas Kärnä, Jeongnim Kim, Mateusz Nowak 0001, Lukasz Slusarczyk, Kacper Stefanski, Timothy G. Mattson
ICS9
2022 Self-Organizing Data Containers
Samuel Madden 0001, Jialin Ding 0001, Tim Kraska, Sivaprasad Sudhir, David E. Cohen, Timothy G. Mattson, Nesime Tatbul
CIDR6
2020 Debugging Large-Scale Data Science Pipelines using Dagger
abstract
Data pipelines are the new code. Consequently, data scientists need new tools to support the often time-consuming process of debugging their pipelines. We introduce Dagger , an end-to-end system to debug and mitigate data-centric errors in data pipelines, such as a data transformation gone wrong or a classifier underperforming due to noisy training data. Dagger supports inter-module debugging, where the pipeline blocks are treated as black boxes, as well as intra-module debugging, where users can debug data objects in Python scripts (e.g., DataFrames). In this demo, we will walk the audience through a rich, real-world business intelligence use case from our industrial collaborators at Intel, to highlight how Dagger enables data scientists to productively identify and mitigate data-centric problems at different stages of pipeline development.
El Kindi Rezig, Ashrita Brahmaroutu, Nesime Tatbul, Mourad Ouzzani, Nan Tang 0001, Timothy G. Mattson, Samuel Madden 0001, Michael Stonebraker
Proc. VLDB Endow.6
2019 Super-Node SLP: Optimized Vectorization for Code Sequences Containing Operators and Their Inverse Elements
abstract
SLP Auto-vectorization converts straight-line code into vector code. It scans input code for groups of instructions that can be combined into vectors and replaces them with their corresponding vector instructions. This work introduces Super-Node SLP (SN-SLP), a new SLP-style algorithm, optimized for expressions that include a commutative operator (such as addition) and its corresponding inverse element (subtraction). SN-SLP uses the algebraic properties of commutative operators and their inverse elements to enable additional transformations that extend auto-vectorization to cases difficult for state-of-the-art auto-vectorizing compilers. We implemented SN-SLP in LLVM. Our evaluation on a real system demonstrates considerable performance improvements of benchmark code with no significant change in compilation time.
Vasileios Porpodas, Rodrigo Caetano Rocha, Evgueni Brevnov, Fabrício Góes, Timothy G. Mattson
CGO5
2019 T2S-Tensor: Productively Generating High-Performance Spatial Hardware for Dense Tensor Computations
abstract
We present a language and compilation framework for productively generating high-performance systolic arrays for dense tensor kernels on spatial architectures, including FPGAs and CGRAs. It decouples a functional specification from a spatial mapping, allowing programmers to quickly explore various spatial optimizations for the same function. The actual implementation of these optimizations is left to a compiler. Thus, productivity and performance are achieved at the same time. We used this framework to implement several important dense tensor kernels. We implemented dense matrix multiply for an Arria-10 FPGA and a research CGRA, achieving 88% and 92% of the performance of manually written, and highly optimized expert (ninja") implementations in just 3% of their engineering time. Three other tensor kernels, including MTTKRP, TTM and TTMc, were also implemented with high performance and low design effort, and for the first time on spatial architectures."
Nitish Kumar Srivastava, Hongbo Rong, Prithayan Barua, Guanyu Feng, Huanqi Cao, Zhiru Zhang, David H. Albonesi, Vivek Sarkar, Paul Petersen, Geoff Lowney, Adam Herr, Christopher J. Hughes, Timothy G. Mattson, Pradeep Dubey
FCCM14
2019 A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions
abstract
The field of machine programming (MP), the automation of the development of software, is making notable research advances. This is, in part, due to the emergence of a wide range of novel techniques in machine learning. In this paper, we apply MP to the automation of software performance regression testing. A performance regression is a software performance degradation caused by a code change. We present AutoPerf – a novel approach to automate regression testing that utilizes three core techniques: (i) zero-positive learning, (ii) autoencoders, and (iii) hardware telemetry. We demonstrate AutoPerf’s generality and efficacy against 3 types of performance regressions across 10 real performance bugs in 7 benchmark and open-source programs. On average, AutoPerf exhibits 4% profiling overhead and accurately diagnoses more performance bugs than prior state-of-the-art approaches. Thus far, AutoPerf has produced no false negatives.
Mejbah Alam, Justin Emile Gottschlich, Nesime Tatbul, Javier Turek, Timothy G. Mattson, Abdullah Muzahid
NeurIPS5
2018 The Ongoing Evolution of OpenMP
abstract
This paper presents an overview of the past, present and future of the OpenMP application programming interface (API). While the API originally specified a small set of directives that guided shared memory fork-join parallelization of loops and program sections, OpenMP now provides a richer set of directives that capture a wide range of parallelization strategies that are not strictly limited to shared memory. As we look toward the future of OpenMP, we immediately see further evolution of the support for that range of parallelization strategies and the addition of direct support for debugging and performance analysis tools. Looking beyond the next major release of the specification of the OpenMP API, we expect the specification eventually to include support for more parallelization strategies and to embrace closer integration into its Fortran, C and, in particular, C++ base languages, which will likely require the API to adopt additional programming abstractions.
Bronis R. de Supinski, Thomas Scogland, Alejandro Duran, Michael Klemm, Sergi Mateo, Stephen Olivier, Christian Terboven, Timothy G. Mattson
Proc. IEEE8
2017 Enabling query processing across heterogeneous data models: A survey
abstract
Modern applications often need to manage and analyze widely diverse datasets that span multiple data models [1], [2], [3], [4], [5]. Warehousing the data through Extract-Transform-Load (ETL) processes can be expensive in such scenarios. Transforming disparate data into a single data model may degrade performance. Further, curating diverse datasets and maintaining the pipeline can prove to be labor intensive. As a result, an emerging trend is to shift the focus to federating specialized data stores and enabling query processing across heterogeneous data models [6]. This shift can bring many advantages: First, systems can natively leverage multiple data models, which can translate to maximizing the semantic expressiveness of underlying interfaces and leveraging the internal processing capabilities of component data stores. Second, federated architectures support query-specific data integration with just-in-time transformation and migration, which has the potential to significantly reduce the operational complexity and overhead. Projects that focus on developing systems in this research area stem from various backgrounds and address diverse concerns, which could make it difficult to form a consistent view of the work in this area. In this survey, we introduce a taxonomy for describing the state of the art and propose a systematic evaluation framework conducive to understanding of query-processing characteristics in the relevant systems. We use the framework to assess four representative implementations: BigDAWG [7], [8], CloudMdsQL [9], [10], Myria [11], [12], and Apache Drill [13].
Ran Tan, Rada Chirkova, Vijay Gadepally, Timothy G. Mattson
IEEE BigData4
2017 Demonstrating the BigDAWG Polystore System for Ocean Metagenomics Analysis
Timothy G. Mattson, Vijay Gadepally, Zuohao She, Adam Dziedzic, Jeff Parkhurst
CIDR1
2016 Design and Implementation of a Parallel Research Kernel for Assessing Dynamic Load-Balancing Capabilities
abstract
The Parallel Research Kernels (PRK) are a tool to study parallel architectures and runtime systems from an application perspective. It provides paper and pencil specifications and reference implementations of elementary operations covering a broad range of parallel application patterns. The current PRK are trivially statically load-balanced. Future large-scale systems will require dynamic load balancing for unsteady workloads and for handling system/network fluctuations and non-uniformities. We present a new PRK that requires dynamic load balancing, and provides knobs for controlling workload behavior. It is inspired by Particle-In-Cell (PIC) applications and captures one of the computational patterns in such codes. We give a detailed specification of the new PRK, highlighting the challenges and corresponding design choices that make it compact, arbitrarily scalable and self-verifying. We also present implementations of the PIC PRK in MPI, with and without application-specific load balancing, and show an implementation with runtime-assisted load balancing provided by Adaptive MPI features. Our experimental results provide an illustrative example of how PIC can be used to assess the load-balancing capabilities of modern parallel runtimes.
Evangelos Georganas, Rob F. Van der Wijngaart, Timothy G. Mattson
IPDPS3
2016 The TileDB Array Data Storage Manager
abstract
We present a novel storage manager for multi-dimensional arrays that arise in scientific applications, which is part of a larger scientific data management system called TileDB. In contrast to existing solutions, TileDB is optimized for both dense and sparse arrays. Its key idea is to organize array elements into ordered collections called fragments. Each fragment is dense or sparse, and groups contiguous array elements into data tiles of fixed capacity. The organization into fragments turns random writes into sequential writes, and, coupled with a novel read algorithm, leads to very efficient reads. TileDB enables parallelization via multi-threading and multi-processing, offering thread-/process-safety and atomicity via lightweight locking. We show that TileDB delivers comparable performance to the HDF5 dense array storage manager, while providing much faster random writes. We also show that TileDB offers substantially faster reads and writes than the SciDB array database system with both dense and sparse arrays. Finally, we demonstrate that TileDB is considerably faster than adaptations of the Vertica relational column-store for dense array storage management, and at least as fast for the case of sparse arrays.
Stavros Papadopoulos 0001, Kushal Datta, Samuel Madden 0001, Timothy G. Mattson
Proc. VLDB Endow.4
2015 Special Issue on Architectures and Algorithms for Irregular Applications (AAIA) - Guest editors' introduction
Antonino Tumeo, John Feo, Oreste Villa, Simone Secchi, Timothy G. Mattson
J. Parallel Distributed Comput.5
2015 A Demonstration of the BigDAWG Polystore System
abstract
This paper presents BigDAWG, a reference implementation of a new architecture for "Big Data" applications. Such applications not only call for large-scale analytics, but also for real-time streaming support, smaller analytics at interactive speeds, data visualization, and cross-storage-system queries. Guided by the principle that "one size does not fit all", we build on top of a variety of storage engines, each designed for a specialized use case. To illustrate the promise of this approach, we demonstrate its effectiveness on a hospital application using data from an intensive care unit (ICU). This complex application serves the needs of doctors and researchers and provides real-time support for streams of patient data. It showcases novel approaches for querying across multiple storage engines, data visualization, and scalable real-time analytics.
Aaron J. Elmore, Jennie Rogers, Michael Stonebraker, Magdalena Balazinska, Ugur Çetintemel, Vijay Gadepally, Jeffrey Heer, Bill Howe, Jeremy Kepner, Tim Kraska, Samuel Madden 0001, David Maier 0001, Timothy G. Mattson, Stavros Papadopoulos 0001, Jeff Parkhurst, Nesime Tatbul, Manasi Vartak, Stanley B. Zdonik
Proc. VLDB Endow.13
2012 Analysis of streaming social networks and graphs on multicore architectures
abstract
Analyzing static snapshots of massive, graph-structured data cannot keep pace with the growth of social networks, financial transactions, and other valuable data sources. We introduce a framework, STING (Spatio-Temporal Interaction Networks and Graphs), and evaluate its performance on multicore, multisocket Intel®-based platforms. STING achieves rates of around 100 000 edge updates per second on large, dynamic graphs with a single, general data structure. We achieve speedups of up to 1000× over parallel static computation, improve monitoring a dynamic graph's connected components, and show an exact algorithm for maintaining local clustering coefficients performs better on Intel-based platforms than our earlier approximate algorithm.
E. Jason Riedy, Henning Meyerhenke, David A. Bader, David Ediger, Timothy G. Mattson
ICASSP5
2012 Programming many-core architectures - a case study: dense matrix computations on the Intel single-chip cloud computer processor
abstract
SUMMARY A message passing, distributed‐memory parallel computer on a chip is one possible design for future, many‐core architectures. We discuss initial experiences with the Intel Single‐chip Cloud Computer research processor, which is a prototype architecture that incorporates 48 cores on a single die that can communicate via a small, shared, on‐die buffer. The experiment is to port a state‐of‐the‐art, distributed‐memory, dense matrix library, Elemental, to this architecture and gain insight from the experience. We show that programmability addressed by this library, especially the proper abstraction for collective communication, greatly aids the porting effort. This enables us to support a wide range of functionality with limited changes to the library code. Copyright © 2011 John Wiley & Sons, Ltd.
Bryan Marker, Ernie Chan, Jack Poulson, Robert A. van de Geijn, Rob F. Van der Wijngaart, Timothy G. Mattson, Theodore E. Kubaska
Concurr. Comput. Pract. Exp.6
2010 The 48-core SCC Processor: the Programmer's View
abstract
The number of cores integrated onto a single die is expected to climb steadily in the foreseeable future. This move to many-core chips is driven by a need to optimize performance per watt. How best to connect these cores and how to program the resulting many-core processor, however, is an open research question. Designs vary from GPUs to cache-coherent shared memory multiprocessors to pure distributed memory chips. The 48-core SCC processor reported in this paper is an intermediate case, sharing traits of message passing and shared memory architectures. The hardware has been described elsewhere. In this paper, we describe the programmer's view of this chip. In particular we describe RCCE: the native message passing model created for the SCC processor.
Timothy G. Mattson, Michael Riepen, Thomas Lehnig, Paul Brett, Werner Haas 0002, Patrick Kennedy, Jason Howard, Sriram R. Vangal, Nitin Borkar, Gregory Ruhl, Saurabh Dighe
SC1
2009 OpenCL*, heterogeneous computing, and the CPU
Timothy G. Mattson
Hot Chips Symposium1
2008 Parallel programming: can we PLEASE get it right this time?
abstract
The computer industry has a problem. As Moore's law marches on, it will be exploited to double cores, not frequencies. But all those cores, growing to 8, 16 and beyond over the next several years, are of little value without parallel software. Where will this come from? With few exceptions, only graduate students and other strange people write parallel software. Even for numerically intensive applications, where parallel algorithms are well understood, professional software engineers almost never write parallel software.
Timothy G. Mattson, Michael Wrinn
DAC1
2008 Programming the Intel 80-core network-on-a-chip terascale processor
abstract
Intel's 80-core terascale processor was the first generally programmable microprocessor to break the Teraflops barrier. The primary goal for the chip was to study power management and on-die communication technologies. When announced in 2007, it received a great deal of attention for running a stencil kernel at 1.0 single precision TFLOPS while using only 97 Watts. The literature about the chip, however, focused on the hardware, saying little about the software environment or the kernels used to evaluate the chip. This paper completes the literature on the 80-core terascale processor by fully defining the chip's software environment. We describe the instruction set, the programming environment, the kernels written for the chip, and our experiences programming this microprocessor. We close by discussing the lessons learned from this project and what it implies for future message passing, network-on-a-chip processors.
Timothy G. Mattson, Rob F. Van der Wijngaart, Michael A. Frumkin
SC1
2007 Reengineering for Parallelism: an entry point into PLPP for legacy applications
abstract
Abstract Many parallel programs begin as legacy sequential code that is later reengineered to take advantage of parallel hardware. This paper presents a pattern called Reengineering for Parallelism to help with this task. The new pattern is intended to be used in conjunction with PLPP (Pattern Language for Parallel Programming), described in our book (Mattson TG, Sanders BA, Massingill BL. Patterns for Parallel Programming. Addison‐Wesley: Reading, MA, 2004). PLPP contains a structured collection of patterns and embodies a methodology for developing parallel programs in which the programmer starts with a good understanding of the problem, works through a sequence of patterns, and finally ends up with the code. Most of the patterns in PLPP are also applicable when reengineering legacy code, but it is not always clear how to get started. Reengineering for Parallelism provides an alternate point of entry into PLPP and addresses particular issues that arise when dealing with legacy code. Copyright © 2006 John Wiley & Sons, Ltd.
Berna L. Massingill, Timothy G. Mattson, Beverly A. Sanders
Concurr. Comput. Pract. Exp.2
2006 S08 - Introduction to OpenMP
abstract
The OpenMP Application Programming Interface (API) defines compiler directives and library routines that make it relatively easy to create parallel programs. It first appeared in 1997 and has become the de facto standard for programming shared memory computers. Recent advances in OpenMP technology are expanding its reach to distributed memory systems (e.g. clusters) as well.In this tutorial, we will provide a comprehensive introduction to OpenMP. By dedicating a half-day tutorial to the API itself, we will be able to cover every construct within the language and show how OpenMP is used to program shared memory multiprocessor computers, multi-core CPUs and clusters.
Timothy G. Mattson
SC1
2001 An Introduction to OpenMP
abstract
OpenMP is the industry standard technology for writing explicitly parallel programs for shared memory computers. OpenMP is easy to use with most of the constructs consisting of simple compiler directives. These directives create threads, split work between them, and manage how data is shared. In this tutorial, I will provide a comprehensive introduction to OpenMP. We will discuss the origins of OpenMP, the programming model behind it, and the basic syntax of the language. While we assume the students know Fortran or C in a Unix environment, we do not require any background in parallel programming. Finally, at this point, I can't predict how the content of the tutorial will be divided between advanced and introductory material. I will determine this at the beginning of the tutorial based on the backgrounds of the students. My goal is to make the experience valuable regardless of the parallel computing experience of the students. Proceedings of the 1st International Symposium on Cluster Computing and the Grid (CCGRID ’01) 0-7695-1010-8/01 $10.00 © 2001 IEEE
Timothy G. Mattson
CCGRID1
2001 High Performance Computing at Intel: The OSCAR Software Solution Stack for Cluster Computing
abstract
This is an exciting time in high performance computing (HPC). Radical change has become the norm as clusters of commercial off the shelf (COTS) computers have come to dominate HPC. The hardware for HPC is in good shape and is steadily getting better. What about the software? The software for HPC is in poor shape. Its great for organizations that can accept the "people-time" required to keep clusters up and running, but for the commercial side of technical computing, the software solutions available today just are not up to the task. We believe that the best way to solve our HPC software problems is to work together. Intel has worked to create cross-industry groups dedicated to solving HPC software problems. On the shared memory front, we are active in the OpenMP Architecture Review Board. Working with this group, we have helped standardize API's for shared memory programming. For clusters, we have helped found the Open Cluster Group. The Open Cluster Group is an informal consortium of commercial and research organizations interested in cluster-computing. Our mission is to make clusters a realistic option for organizations that use high performance computing to accomplish their goals. The first project from the Open Cluster Group is OSCAR: a fully integrated easy to install bundle of software designed to make it easy to build and use a cluster for high performance computing. Everything you need to build, maintain, and use a modest sized Linux cluster is included in OSCAR. We introduce OSCAR and provide the background information you need to use it.
Timothy G. Mattson
CCGRID1
2001 Parallel programming with a pattern language
Berna L. Massingill, Timothy G. Mattson, Beverly A. Sanders
Int. J. Softw. Tools Technol. Transf.2
2000 Tutorial D: An Introduction to OpenMP and Its Use on Clusters
Timothy G. Mattson
CLUSTER1
2000 Cluster Computing at Intel
Timothy G. Mattson
CLUSTER1
2000 A Pattern Language for Parallel Application Programs (Research Note)
Berna L. Massingill, Timothy G. Mattson, Beverly A. Sanders
Euro-Par2
1994 The Linda® Alternative to Message-Passing Systems
Nicholas Carriero, David Gelernter, Timothy G. Mattson, Andrew H. Sherman
Parallel Comput.3
1990 Adventures in portable parallel programming: STRAND88 with embedded Fortran and C
abstract
STRAND/sup 88/ is a high-level language for writing programs that are portable across a broad range of parallel computers. The author introduces STRAND/sup 88/ and discusses how it is used to port sequential applications onto parallel computers. A key language feature making this possible is its interface to C and Fortran. Case studies are presented in weather modeling and protein structure prediction.>
Timothy G. Mattson
ICSM1