Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

John Feo

dblp:58/5212 · also John T. Feo · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0001-6546-8948ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Graph data management · 50% Data stream processing · 50%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Processor architecture and microarchitecture · 38% High-performance computing · 25% Performance modeling and evaluation · 19%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management
graph query processing
0.212013
StreamWorks: a system for dynamic graph search · SIGMOD Conference 2013
High-performance computing
fast fourier transform
0.012000
A Framework for the Design and Implementation of FFT Permutation Algorithms · IEEE Trans. Parallel Distributed Syst. 2000
Algorithms and data structures › sequence algorithms
sorting
0.012000
A Framework for the Design and Implementation of FFT Permutation Algorithms · IEEE Trans. Parallel Distributed Syst. 2000
Processor architecture and microarchitecture
memory latency tolerance
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Processor architecture and microarchitecture
multithreading
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Performance modeling and evaluation
parallel performance evaluation
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Programming languages and type systems
functional programming
0.012000
A Framework for the Design and Implementation of FFT Permutation Algorithms · IEEE Trans. Parallel Distributed Syst. 2000
Parallel and multicore computing › parallel computing
parallel programming languages
0.011990
SISAL versus FORTRAN: a comparison using the Livermore loops · SC 1990
Interconnection networks and networks-on-chip › switching network
multistage interconnection network
0.011985
Dynamic, Distributed Resource Configuration on SW-Banyans · ISCA 1985
Cloud and datacenter computing › resource management
resource configuration
0.011985
Dynamic, Distributed Resource Configuration on SW-Banyans · ISCA 1985
Compilers and program optimization › parallelization
automatic parallelization
0.011990
SISAL versus FORTRAN: a comparison using the Livermore loops · SC 1990
Distributed systems
distributed resource management
0.011985
Dynamic, Distributed Resource Configuration on SW-Banyans · ISCA 1985

Methods — techniques the papers use, named apart from their topics

incremental graph search · 0.2amdahl's law analysis · 0.0NAS benchmarks · 0.0performance benchmarking · 0.0distributed resource configuration · 0.0
YearPublicationVenuePosition
2025 DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems
abstract
We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms.We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores.Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges).By overcoming these limitations, ICS '25, June 08-11, 2025, Salt Lake City, UT, USA Minutoli et al.we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges.This two ordersof-magnitude improvement over the previous state-of-theart is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application's memory requirements.We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds.We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.
Marco Minutoli, Reece Neff, Naw Safrin Sattar, Hao Lu 0001, John Feo, Henning S. Mortveit, Anil Vullikanti, Dawen Xie, Mandy L. Wilson, Gregor von Laszewski, Parantapa Bhattacharya, S. M. Ferdous, Anantharaman Kalyanaraman, Michela Becchi, Madhav V. Marathe, Mahantesh Halappanavar
ICS5
2024 Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems
abstract
The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globally-consistent meta-data. In this paper, we propose a novel data-structure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.
Vito Giovanni Castellana, Burcu O. Mutlu, Ian Di Dio Lavore, Jesun Sahariar Firoz, Katherine E. Wolf, Marco Minutoli, John Feo
IEEE Big Data7
2023 AGILE Workflows and Graphs
abstract
IARPA's AGILE program aspires to develop innovative, efficient, scalable computer architecture designs capable of addressing the challenges of next-generation large-scale data-analytic applications, especially involving very large graphs. Given the complexity of these applications, AGILE has adopted a three-tier approach to modeling performance and scalability. The top tier is a set of 4 end-to-end analytic workflows comprising a broad set of methods and data structures. The middle tier are kernels of the workflows capturing critical algorithms and data construction operations. The bottom tier are small, stand-alone kernels such as breadth-first search, triangle counting, and sparse matrix multiply for measuring the speed-and-feeds of individual hardware components.
Marco Minutoli, John Feo
CF2
2019 A Parallel Graph Environment for Real-World Data Analytics Workflows
abstract
Economic competitiveness and national security depend increasingly on the insightful analysis of large data sets. The diversity of real-world data sources and analytic workflows impose challenging hardware and software requirements for parallel graph platforms. The irregular nature of graph methods is not supported well by the deep memory hierarchies of conventional distributed systems, requiring new processor and runtime system designs to tolerate memory and synchronization latencies. Moreover, the efficiency of relational table operations and matrix computations are not attainable when data is stored in common graph data structures. In this paper, we present HAGGLE, a high-performance, scalable data analytics platform. The platform's hybrid data model supports a variety of distributed, thread-safe data structures, parallel programming constructs, and persistent and streaming data. An abstract runtime layer enables us to map the stack to conventional, distributed computer systems with accelerators. The runtime uses multithreading, active messages, and data aggregation to hide memory and synchronization latencies on large-scale systems.
Vito Giovanni Castellana, Maurizio Drocco, John Feo, Jesun Sahariar Firoz, Thejaka Amila Kanewala, Andrew Lumsdaine, Joseph B. Manzano, Andrés Márquez 0001, Marco Minutoli, Joshua Suetterlein, Antonino Tumeo, Marcin Zalewski
DATE3
2019 Special Issue on: Systems for Learning, Inferencing, and Discovering (SLID)
Antonino Tumeo, John Feo, Oreste Villa
J. Parallel Distributed Comput.2
2018 A hybrid data model for large-scale analytics
abstract
Modern data analytic workflows are a complex mix of heterogeneous data and methods. No one size fits all. Consequently, analytic software platforms must support relational, graph, math, and machine learning methods. These methods use a variety of different data structures---tables, graphs, sets, matrices, and networks---and embody a variety of parallel execution models. At PNNL, we are developing HAGGLE, a hybrid, attributed analytic framework. In this talk, I will describe HAGGLEs integrated hybrid data layer capable of storing heterogeneous data in tables, graphs, or matrices. I will discuss different data views and the analytic methods they support. I will show the benefits of HAGGLE multiple data structure approach over platforms that support a single data structure type.
John Feo
CF1
2016 Special Issue on Theory and Practice of Irregular Applications (TaPIA)
Antonino Tumeo, John Feo, Oreste Villa
Parallel Comput.2
2015 Locality aware concurrent start for stencil applications
abstract
Stencil computations are at the heart of many physical simulations used in scientific codes. Thus, there exists a plethora of optimization efforts for this family of computations. Among these techniques, tiling techniques that allow concurrent start have proven to be very efficient in providing better performance for these critical kernels. Nevertheless, with many core designs being the norm, these optimization techniques might not be able to fully exploit locality (both spatial and temporal) on multiple levels of the memory hierarchy without compromising parallelism. It is no longer true that the machine can be seen as a homogeneous collection of nodes with caches, main memory and an interconnect network. New architectural designs exhibit complex grouping of nodes, cores, threads, caches and memory connected by an ever evolving network-on-chip design. These new designs may benefit greatly from carefully crafted schedules and groupings that encourage parallel actors (i.e. threads, cores or nodes) to be aware of the computational history of other actors in close proximity. In this paper, we provide an efficient tiling technique that allows hierarchical concurrent start for memory hierarchy aware tile groups. Each execution schedule and tile shape exploit the available parallelism, load balance and locality present in the given applications. We demonstrate our technique on the Intel Xeon Phi architecture with selected and representative stencil kernels. We show improvement ranging from 5.58% to 31.17% over existing state-of-the-art techniques.
Sunil Shrestha, Guang R. Gao, Joseph B. Manzano, Andrés Márquez 0001, John Feo
CGO5
2015 High-Performance, Distributed Dictionary Encoding of RDF Datasets
abstract
In this work we propose a novel approach for RDF (Resource Description Framework) dictionary encoding that employs a parallel RDF parser and a distributed dictionary data structure, exploiting RDF-specific optimizations. In contrast with previous solutions, this approach exploits the Partitioned Global Address Space (PGAS) programming model combined with active messages. We evaluate the performance of our dictionary encoder in our RDF database, GEMS (Graph Engine for Multithreaded Systems), and provide an empirical comparison against previous approaches. Our comparison shows that our dictionary encoder scales significantly better and achieves higher performance than the current state of the art, providing a key element for the realization of a more efficient RDF database.
Alessandro Morari, Jesse Weaver, Oreste Villa, David J. Haglin, Antonino Tumeo, Vito Giovanni Castellana, John Feo
CLUSTER7
2015 A Selectivity based approach to Continuous Pattern Detection in Streaming Graphs
Sutanay Choudhury, Lawrence B. Holder, George Chin, Khushbu Agarwal, John Feo
EDBT5
2015 Special Issue on Architectures and Algorithms for Irregular Applications (AAIA) - Guest editors' introduction
Antonino Tumeo, John Feo, Oreste Villa, Simone Secchi, Timothy G. Mattson
J. Parallel Distributed Comput.2
2014 Toward a data scalable solution for facilitating discovery of science resources
Jesse Weaver, Vito Giovanni Castellana, Alessandro Morari, Antonino Tumeo, Sumit Purohit, Alan R. Chappell, David J. Haglin, Oreste Villa, Sutanay Choudhury, Karen Schuchardt, John Feo
Parallel Comput.11
2013 Accelerating semantic graph databases on commodity clusters
abstract
We are developing a full software system for accelerating semantic graph databases on commodity cluster that scales to hundreds of nodes while maintaining constant query throughput. Our framework comprises a SPARQL to C++ compiler, a library of parallel graph methods and a custom multithreaded runtime layer, which provides a Partitioned Global Address Space (PGAS) programming model with fork/join parallelism and automatic load balancing over a commodity clusters. We present preliminary results for the compiler and for the runtime.
Alessandro Morari, Vito Giovanni Castellana, David J. Haglin, John Feo, Jesse Weaver, Antonino Tumeo, Oreste Villa
IEEE BigData4
2013 StreamWorks: a system for dynamic graph search
abstract
Acting on time-critical events by processing ever growing social media, news or cyber data streams is a major technical challenge. Many of these data sources can be modeled as multi-relational graphs. Mining and searching for subgraph patterns in a continuous setting requires an efficient approach to incremental graph search. The goal of our work is to enable real-time search capabilities for graph databases. This demonstration will present a dynamic graph query system that leverages the structural and semantic characteristics of the underlying multi-relational graph.
Sutanay Choudhury, Lawrence B. Holder, George Chin, Abhik Ray, Sherman Beus, John Feo
SIGMOD Conference6
2012 A Novel Multithreaded Algorithm for Extracting Maximal Chordal Subgraphs
abstract
Chordal graphs are triangulated graphs where any cycle larger than three is bisected by a chord. Many combinatorial optimization problems such as computing the size of the maximum clique and the chromatic number are NP-hard on general graphs but have polynomial time solutions on chordal graphs. In this paper, we present a novel multithreaded algorithm to extract a maximal chordal sub graph from a general graph. We develop an iterative approach where each thread can asynchronously update a subset of edges that are dynamically assigned to it per iteration and implement our algorithm on two different multithreaded architectures - Cray XMT, a massively multithreaded platform, and AMD Magny-Cours, a shared memory multicore platform. In addition to the proof of correctness, we present the performance of our algorithm using a test set of synthetical graphs with up to half-a-billion edges and real world networks from gene correlation studies and demonstrate that our algorithm achieves high scalability for all inputs on both types of architectures.
Mahantesh Halappanavar, John Feo, Kathryn Dempsey, Hesham Ali 0001, Sanjukta Bhowmick
ICPP2
2012 Scalable Triadic Analysis of Large-Scale Graphs: Multi-core vs. Multi-processor vs. Multi-threaded Shared Memory Architectures
abstract
Triadic analysis encompasses a useful set of graph mining methods that are centered on the concept of a triad, which is a sub graph of three nodes. Such methods are often applied in the social sciences as well as many other diverse fields. Triadic methods commonly operate on a triad census that counts the number of triads of every possible edge configuration in a graph. Like other graph algorithms, triadic census algorithms do not scale well when graphs reach tens of millions to billions of nodes. To enable the triadic analysis of large-scale graphs, we developed and optimized a triad census algorithm to efficiently execute on shared memory architectures. We then conducted performance evaluations of the parallel triad census algorithm on three specific systems: CrayXMT, HP Superdome, and AMD multi-core NUMA machine. These three systems have shared memory architectures but with markedly different hardware capabilities to manage parallelism.
George Chin, Andrés Márquez 0001, Sutanay Choudhury, John Feo
SBAC-PAD4
2012 Graph coloring algorithms for multi-core and massively multithreaded architectures
Ümit V. Çatalyürek, John Feo, Assefaw Hadish Gebremedhin, Mahantesh Halappanavar, Alex Pothen
Parallel Comput.2
2010 A novel application of parallel betweenness centrality to power grid contingency analysis
abstract
In Energy Management Systems, contingency analysis is commonly performed for identifying and mitigating potentially harmful power grid component failures. The exponentially increasing combinatorial number of failure modes imposes a significant computational burden for massive contingency analysis. It is critical to select a limited set of high-impact contingency cases within the constraint of computing power and time requirements to make it possible for real-time power system vulnerability assessment. In this paper, we present a novel application of parallel betweenness centrality to power grid contingency selection. We cross-validate the proposed method using the model and data of the western US power grid, and implement it on a Cray XMT system - a massively multithreaded architecture - leveraging its advantages for parallel execution of irregular algorithms, such as graph analysis. We achieve a speedup of 55 times (on 64 processors) compared against the single-processor version of the same code running on the Cray XMT. We also compare an OpenMP-based version of the same code running on an HP Superdome shared-memory machine. The performance of the Cray XMT code shows better scalability and resource utilization, and shorter execution time for large-scale power grids. This proposed approach has been evaluated in PNNL's Electricity Infrastructure Operations Center (EIOC). It is expected to provide a quick and efficient solution to massive contingency selection problems to help power grid operators to identify and mitigate potential widespread cascading power grid failures in real time.
Shuangshuang Jin, Zhenyu Huang 0001, Yousu Chen, Daniel G. Chavarría-Miranda, John Feo, Pak Chung Wong
IPDPS5
2009 A High-Performance Hybrid Computing Approach to Massive Contingency Analysis in the Power Grid
abstract
Operating the electrical power grid to prevent power black-outs is a complex task. An important aspect of this is contingency analysis, which involves understanding and mitigating potential failures in power grid elements such as transmission lines. When taking into account the potential for multiple simultaneous failures (known as the N-x contingency problem), contingency analysis becomes a massively computational task. In this paper we describe a novel hybrid computational approach to contingency analysis. This approach exploits the unique graph processing performance of the Cray XMT in conjunction with a conventional massively parallel compute cluster to identify likely simultaneous failures that could cause widespread cascading power failures that have massive economic and social impact on society. The approach has the potential to provide the first practical and scalable solution to the N-x contingency problem. When deployed in power grid operations, it will increase the grid operator’s ability to deal effectively with outages and failures with power grid components while preserving stable and safe operation of the grid. The paper describes the architecture of our solution and presents preliminary performance results that validate the efficacy of our approach.
Ian Gorton, Zhenyu Huang 0001, Yousu Chen, Benson Kalahar, Shuangshuang Jin, Daniel G. Chavarría-Miranda, Douglas J. Baxter, John Feo
eScience8
2009 Factors affecting the performance of parallel mining of minimal unique itemsets on diverse architectures
abstract
Abstract Three parallel implementations of a divide‐and‐conquer search algorithm (called SUDA2) for finding minimal unique itemsets (MUIs) are compared in this paper. The identification of MUIs is used by national statistics agencies for statistical disclosure assessment. The first parallel implementation adapts SUDA2 to a symmetric multi‐processor cluster using the message passing interface (MPI), which we call an MPI cluster; the second optimizes the code for the Cray MTA2 (a shared‐memory, multi‐threaded architecture) and the third uses a heterogeneous ‘group’ of workstations connected by LAN. Each implementation considers the parallel structure of SUDA2, and how the subsearch computation times and sequence of subsearches affect load balancing. All three approaches scale with the number of processors, enabling SUDA2 to handle larger problems than before. For example, the MPI implementation is able to achieve nearly two orders of magnitude improvement with 132 processors. Performance results are given for a number of data sets. Copyright © 2009 John Wiley & Sons, Ltd.
David J. Haglin, Kenneth R. Mayes, Anna M. Manning, John Feo, John R. Gurd, Mark J. Elliot, John A. Keane
Concurr. Comput. Pract. Exp.4
2007 Probability Convergence in a Multithreaded Counting Application
abstract
The problem of counting specified combinations of a given set of variables arises in many statistical and data mining applications. To solve this problem, we introduce the PDtree data structure, which avoids exponential time and space complexity associated with prior work by allowing user specification of the tree structure. A straightforward parallelization approach using a Cray MTA-2 provides a speedup that is linear in the number of processors, but introduces nondeterminism into probability estimates. We prove a general convergence result that bounds the non-deterministic deviation of probability estimates relative to a sequential implementation. Beyond PDtrees, this convergence result applies to any counting application that takes a multithreaded streaming approach.
Chad Scherrer, Nathaniel Beagley, Jarek Nieplocha, Andrés Márquez 0001, John Feo, Daniel G. Chavarría-Miranda
IPDPS5
2005 On the Architectural Requirements for Efficient Execution of Graph Algorithms
abstract
Combinatorial problems such as those from graph theory pose serious challenges for parallel machines due to non-contiguous, concurrent accesses to global data structures with low degrees of locality. The hierarchical memory systems of symmetric multiprocessor (SMP) clusters optimize for local, contiguous memory accesses, and so are inefficient platforms for such algorithms. Few parallel graph algorithms outperform their best sequential implementation on SMP clusters due to long memory latencies and high synchronization costs. In this paper, we consider the performance and scalability of two graph algorithms, list ranking and connected components, on two classes of shared-memory computers: symmetric multiprocessors such as the Sun Enterprise servers and multithreaded architectures (MTA) such as the Cray MTA-2. While previous studies have shown that parallel graph algorithms can speedup on SMPs, the systems' reliance on cache microprocessors limits performance. The MTA's latency tolerant processors and hardware support for fine-grain synchronization makes performance a function of parallelism. Since parallel graph algorithms have an abundance of parallelism, they perform and scale significantly better on the MTA. We describe and give a performance model for each architecture. We analyze the performance of the two algorithms and discuss how the features of each architecture affects algorithm development, ease of programming, performance, and scalability.
David A. Bader, Guojing Cong, John Feo
ICPP3
2000 A Framework for the Design and Implementation of FFT Permutation Algorithms
abstract
We propose an algebraic framework for the design and implementation of a large class of data-sorting procedures, including all index-digit permutations used in FFTs. We discuss both old and new algorithms in terms of this framework. We show that the algebraic formulation of the new algorithms can be easily encoded using a functional programming language, and that the resulting code introduces no inefficiencies. We present performance results for implementations of three new algorithms for mixed-radix digit-reversal on a Cray C-90 and on a Sun Sparc 5.
Jaime Seguel, Dorothy Bollman, John Feo
IEEE Trans. Parallel Distributed Syst.3
1998 Multi-processor Performance on the Tera MTA
abstract
The Tera MTA is a revolutionary commercial computer based on a multithreaded processor architecture. In contrast to many other parallel architectures, the Tera MTA can effectively use high amounts of parallelism on a single processor. By running multiple threads on a single processor, it can tolerate memory latency and to keep the processor saturated. If the computation is sufficiently large, it can benefit from running on multiple processors. A primary architectural goal of the MTA is that it provide scalable performance over multiple processors. This paper is a preliminary investigation of the first multi-processor Tera MTA. In a previous paper [1] we reported that on the kernel NAS 2 benchmarks [2], a single-processor MTA system running at the architected clock speed would be similar in performance to a single processor of the Cray T90. We found that the compilers of both machines were able to find the necessary threads or vector operations, after making standard changes to the random number generator. In this paper we update the single-processor results in two ways: we use only actual clock speeds, and we report improvements given by further tuning of the MTA codes. We then investigate the performance of the best single-processor codes when run on a two-processor MTA, making no further tuning effort. The parallel efficiency of the codes range from 77% to 99%. An analysis shows that the "serial bottlenecks" -- unparallelized code sections and the cost of allocating and freeing the parallel hardware resources -- account for less than a percent of the runtimes. Thus, Amdahl's Law needn't take effect on the NAS benchmarks until there are hundreds of processors running thousands of threads. Instead, the major source of inefficiency appears to be an imperfect network connecting the processors to the memory. Ideally, the network can support one memory reference per instruction. The current hardware has defects that reduce the throughput to about 85% of this rate. Except for the EP benchmark, the tuned codes issue memory references at nearly the peak rate of one per instruction. Consequently, the network can support the memory references issued by one, but not two, processors. As a result, the parallel efficiency of EP is near- perfect, but the others are reduced accordingly. Another reason for imperfect speedup pertains to the compiler. While the definition of a thread in a single processor or multi-processor mode is essentially the same, there is a different implementation and an associated overhead with running on multiple processors. We characterize the overhead of running "frays" (a collection of threads running on a single processor) and "crews" (a collection of frays, one per processor.)
Allan Snavely, Larry Carter, Jay Boisseau, Amitava Majumdar 0001, Kang Su Gatlin, Nick Mitchell, John Feo, Brian D. Koblenz
SC7
1993 Program Partitioning for NUMA Multiprocessor Computer Systems
Richard Wolski, John Feo
J. Parallel Distributed Comput.2
1990 SISAL versus FORTRAN: a comparison using the Livermore loops
abstract
The authors compare the performance of SISAL, an application language for parallel numerical computations, and Fortran. The intent is to show that applicative programs, when compiled using a set of powerful yet simple optimization techniques, can achieve sequential execution speeds comparable to Fortran, and automatically utilize conventional shared memory multiprocessors.>
David C. Cann, John Feo
SC2
1990 A Report on the Sisal Language Project
John Feo, David C. Cann, R. R. Oldehoeft
J. Parallel Distributed Comput.1
1988 An analysis of the computational and parallel complexity of the Livermore Loops
John Feo
Parallel Comput.1
1985 Dynamic, Distributed Resource Configuration on SW-Banyans
abstract
article Free Access Share on Dynamic, distributed resource configuration on SW-banyans Authors: John Feo Departments of Computer Sciences and Computer and Electrical Engineering, The University of Texas at Austin, Austin, Texas Departments of Computer Sciences and Computer and Electrical Engineering, The University of Texas at Austin, Austin, TexasView Profile , Roy Jenevein Departments of Computer Sciences and Computer and Electrical Engineering, The University of Texas at Austin, Austin, Texas Departments of Computer Sciences and Computer and Electrical Engineering, The University of Texas at Austin, Austin, TexasView Profile , J. C. Browne Departments of Computer Sciences and Computer and Electrical Engineering, The University of Texas at Austin, Austin, Texas Departments of Computer Sciences and Computer and Electrical Engineering, The University of Texas at Austin, Austin, TexasView Profile Authors Info & Claims ACM SIGARCH Computer Architecture NewsVolume 13Issue 3June 1985 pp 268–275https://doi.org/10.1145/327070.327233Published:01 June 1985Publication History 0citation156DownloadsMetricsTotal Citations0Total Downloads156Last 12 Months7Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
John Feo, Roy M. Jenevein, James C. Browne
ISCA1