Rod A. Fatoohi

dblp:94/6269 · also Rod Fatoohi · DBLP profile ↗
← Back
21ranked-venue papers
15as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 10 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
High-performance computing · 30% Cloud and datacenter computing · 21% Performance modeling and evaluation · 16%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed job execution
0.012000
An Object-Oriented Job Execution Environment · SC 2000
Cloud and datacenter computing
job scheduling
0.012000
An Object-Oriented Job Execution Environment · SC 2000
High-performance computing › scientific computing systems
computational fluid dynamics
0.021994
Performance evaluation of three distributed computing environments for scientific applications · SC 1994
NAS experiences with a prototype cluster of workstations · SC 1994
High-performance computing
cluster computing
0.011994
Performance evaluation of three distributed computing environments for scientific applications · SC 1994
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
High-performance computing › cluster computing
network of workstations
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
High-performance computing › scientific computing
scientific computing application
0.011994
Performance evaluation of three distributed computing environments for scientific applications · SC 1994
Performance modeling and evaluation
benchmarking
0.021994
The NAS parallel benchmarks - summary and preliminary results · SC 1991
NAS experiences with a prototype cluster of workstations · SC 1994
Performance modeling and evaluation › benchmarking › parallel benchmark suites
NAS parallel benchmarks
0.021994
The NAS parallel benchmarks - summary and preliminary results · SC 1991
NAS experiences with a prototype cluster of workstations · SC 1994
Performance modeling and evaluation › benchmarking
parallel benchmark suites
0.011991
The NAS parallel benchmarks - summary and preliminary results · SC 1991
Processor architecture and microarchitecture
vector processor
0.011989
Vector performance analysis of three supercomputers: Cray 2, Cray Y-MP, and ETA 10-Q · SC 1989
High-performance computing
scientific computing systems
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
High-performance computing › supercomputing
supercomputer performance evaluation
0.011991
The NAS parallel benchmarks - summary and preliminary results · SC 1991

Methods — techniques the papers use, named apart from their topics

java · 0.0design patterns · 0.0UML · 0.0CORBA · 0.0performance evaluation · 0.0performance model · 0.0
YearPublicationVenuePosition
2008 Performance evaluation of NSF application benchmarks on parallel systems
abstract
The National Science Foundation (NSF) recently released a set of application benchmarks that would be a key factor in selecting the next-generation high- performance computing environment. These benchmarks are designed to capture the salient attributes of those science and engineering applications placing the most stringent demands on the system to be provisioned. The application benchmarks consist of six codes that require large amount of memory and work with large data sets. In this work, we study the complexity, performance, and scalability of these codes on four machines: a 512-processor SGI Altix 3700, a 512-processor SGI Altix 3700/BX2, a 512-processor dual-core based SGI Altix 4700, and a 128-processor Cray Opteron cluster interconnected by the Myrinet network. We evaluated these codes for two different problem sizes using different numbers of processors. Our results show that per processor the SGI machines, using the Intel Itanium-2 processor, are faster than the Cray cluster, using the AMD Opteron processor, by a factor of up to three. Also, we found out that some of these codes scale up very well as we increase the number of processors while others scaled up poorly. In addition, one of the codes achieved about 2/3 of the peak rate of an SGI Altix processor. Moreover, the dual-core based system achieved comparable performance results to the single-core based system. Finally, we provide some limitations and concluding remarks.
Rod A. Fatoohi
IPDPS1
2007 Performance Evaluation of the Dual-Core Based SGI Altix 4700
abstract
The newest SGI Altix system is the 4000 series that uses dual-core Itanium-2 p9000 series processors. It differs from the Altix 3000 series in the processor architecture and in the type and number of the component modules. Here we compare and contrast between the Altix 4700 and Altix 3700/BX2 using a set of communication benchmarks, kernel benchmarks, and NSF application benchmarks. The communication benchmarks show that the 4700 has a higher effective bandwidth than the 3700/BX2 while the computation benchmarks show the two systems perform at comparable rates except where the application was able to take advantage of a larger L2 cache and faster memory bus on the 4700.
Rod A. Fatoohi
SBAC-PAD1
2006 Interconnect performance evaluation of SGI Altix 3700 BX2, Cray XI, Cray Opteron Cluster, and Dell PowerEdge
abstract
We study the performance of inter-process communication on four high-speed multiprocessor systems using a set of communication benchmarks. The goal is to identify certain limiting factors and bottlenecks with the interconnect of these systems as well as to compare these interconnects. We measured network bandwidth using different numbers of communicating processors and communication patterns - such as point-to-point communication, collective communication, and dense communication patterns. The four platforms are: a 512-processor SGIAltix 3700 BX2 shared-memory machine with 3.2 GB/s links; a 64-processor (single-streaming) Cray XI shared-memory machine with 32 1.6 GB/s links; a 128-processor Cray Opteron cluster using a Myrinet network; and a 1280-node Dell PowerEdge cluster with an InfiniBand network. Our results show the impact of the network bandwidth and topology on the overall performance of each interconnect
Rod A. Fatoohi, Subhash Saini, Robert Ciotti
IPDPS1
2006 Performance evaluation of supercomputers using HPCC and IMB benchmarks
abstract
The HPC Challenge (HPCC) benchmark suite and the Intel MPI Benchmark (IMB) are used to compare and evaluate the combined performance of processor, memory subsystem and interconnect fabric of five leading supercomputers - SGI Altix BX2, Cray XI, Cray Opteron Cluster, Dell Xeon cluster, and NEC SX-8. These five systems use five different networks (SGI NUMALINK4, Cray network, Myrinet, InfiniBand, and NEC IXS). The complete set of HPCC benchmarks are run on each of these systems. Additionally, we present Intel MPI Benchmarks (IMB) results to study the performance of 11 MPI communication functions on these systems
Subhash Saini, Robert Ciotti, Brian T. N. Gunney, Thomas E. Spelce, Alice E. Koniges, Don Dossa, Panagiotis A. Adamidis, Rolf Rabenseifner, Sunil Reddy Tiyyagura, Matthias S. Müller, Rod A. Fatoohi
IPDPS11
2006 Performance evaluation of high-speed interconnects using dense communication patterns
Rod A. Fatoohi, Ken Kardys, Sumy Koshy, Soundarya Sivaramakrishnan, Jeffrey S. Vetter
Parallel Comput.1
2005 iJob: an Internet-based job execution environment using asynchronous messaging
Rod A. Fatoohi, Nihar Gokhale, Suja Viswesan
Inf. Softw. Technol.1
2005 Optimizing transmission time of scalable coded images in peer-to-peer networks
Xiao Su 0006, Rod A. Fatoohi
Multim. Syst.2
2003 Scalable coded image transmissions over peer-to-peer networks
abstract
In this paper, we study the transmission of scalable coded images over peer-to-peer networks. Scalable coded images share common prefix of their resulted bit streams even when coded using different bit rates. This property implies two important consequences on the peer-to-peer system when compared to transmission of non-scalable coded images: (1) there exists a many-to-one relationship between supplying and requesting peers as multiple peers with the code images in different bit rates become eligible as supplying peers; and (2) the set of supplying peers is dynamic over time as the peers in the supplying set may finish transmission at different times. When we transmit the requested image from multiple supplying peers to a requesting peer, it is very important to design optimal peer assignment algorithms to minimize the overall transmission time for the requesting peer. For this purpose, we first establish a sufficient property for the optimal peer assignment vector, and then design an optimal media segmentation algorithm based on the sufficient property. Finally, we compare the performance of the proposed optimal media segmentation algorithm with two heuristics and verify its superior performance.
Xiao Su 0006, Rod A. Fatoohi
ICME2
2003 Migration of DCE applications into CORBA and SOAP environments
abstract
Abstract Legacy applications based on the Distributed Computing Environment (DCE) are subject to several significant limitations. As the development and support of DCE wanes, object‐oriented development becomes more desirable, and transmission over HTTP is established as the preferred protocol over the Internet, DCE application managers and developers are pressed to find extensions and alternatives to DCE. This paper briefly discusses several alternative targets for migration of DCE systems, then proceeds to detail, compare, and contrast two preferred candidates: CORBA and SOAP. Although we have found that developing a general migration solution for legacy DCE applications to CORBA or SOAP to be a non‐trivial, long‐term project, developing specific solutions that are based on a general architecture is feasible. Given a short list of reasonable premises, many DCE applications may be ported to technologies such as CORBA and SOAP. Copyright © 2002 John Wiley & Sons, Ltd.
Rod A. Fatoohi, D. Jensen
Softw. Pract. Exp.1
2000 Performance evaluation of middleware bridging technologies
abstract
This paper provides a state-of-the-art study of bridging between different middleware technologies. Two DCOM-CORBA bridges, IONA OrbixCOMet and Visual Edge ObjectBridge, as well as a DCE-CORBA bridge by Inprise are tested and evaluated. Several configurations, depending on the number of machines and location of the bridge, are employed and two languages (C++ and Java) are used. The results show that the three bridges perform reasonably well for different configurations and language mappings.
Rod A. Fatoohi, Vandana Gunwani, Charlton Zheng
ISPASS1
2000 An Object-Oriented Job Execution Environment
abstract
This is a project for developing a distributed job execution environment for highly iterative jobs. An iterative job is one where the same binary code is run hundreds of times with incremental changes in the input values for each run. An execution environment is a set of resources on a computing platform that can be made available to run the job and hold the output until it is collected. The goal is to design a complete, object-oriented scheduling system that will run a variety of jobs with minimal changes. Areas of code that are unique to one specific type of job are decoupled from the rest. The system allows for fine-grained job control, timely status notification and dynamic registration and deregistration of execution platforms depending on resources available. Several objected-oriented technologies are employed: Java, CORBA, UML, and software design patterns. The environment has been tested using a CFD code, INS2D.
Lance Smith, Rod A. Fatoohi
SC2
1995 Performance evaluation of communication networks for distributed computing
abstract
We present performance results for several high-speed networks in distributed computing environments. These networks are: HiPPI, ATM, Fibre Channel, IBM Allnode switch, FDDI, and Ethernet. These networks are parts of two testbeds: DaVinci-a cluster of 16 SGI R8000 workstations at NASA Ames-and LACE-a cluster of 96 IBM RS6000 workstations at NASA Lewis. Also, an IBM SP2 machine is considered for comparison. Several communication tests are performed and the results are presented for two programming levels: BSD socket programming interface using the program ttcp and PVM message passing library. These results show that the emerging network technologies can achieve reasonable performance under certain conditions. However, the achievable performance is still far behind the theoretical peak rates.
Rod A. Fatoohi
ICCCN1
1994 NAS experiences with a prototype cluster of workstations
abstract
This paper discusses the year-long activity at NAS to implement a large, loose cluster of workstations from the existing Silicon Graphics, Inc. (SGI) pool of systems. Issues related to establishing a loosely coupled cluster of workstations are presented. Included are steps needed to resolve system management issues intended to provided reasonable cycle recovery from these systems without disrupting the primary system users. Performance evaluation tests were run based on the NAS Parallel Benchmarks (NPB) and other codes, including OVERFLOW-PVM, a full-fledged computational fluid dynamics (CFD) application. This paper summarizes the activities related to the prototype cluster and identifies areas that need improvement, development, and research in order to make workstation clusters a viable computing environment for solving aeroscience problems.>
Karen Castagnera, Doreen Cheng, Rod A. Fatoohi, Edward Hook, William T. Kramer, Craig Manning, John Musch, Charles Niggley, William Saphir, Douglas Sheppard, Merritt Smith, Ian Stockdale, Shaun Welch, Rita Williams, David Yip
SC3
1994 Performance evaluation of three distributed computing environments for scientific applications
abstract
Presents performance results for three distributed computing environments using the three simulated computational fluid dynamics applications in the NAS Parallel Benchmark suite. These environments are the Distributed Computing Facility (DCF) cluster, the LACE cluster, and an Intel iPSC/860 machine. The DCF is a prototypic cluster of loosely-coupled SGI R3000 machines connected by Ethernet. The LACE cluster is a tightly-coupled cluster of 92 IBM RS6000/560 machines connected by Ethernet as well as by either FDDI or an IBM Allnode switch. Results of several parallel algorithms for the three simulated applications are presented and analyzed, based on the interplay between the communication requirements of an algorithm and the characteristics of the communication network of a distributed system.>
Rod A. Fatoohi, Sisira Weeratunga
SC1
1994 Adapting a Navier-Stokes solver for three parallel machines
Rod A. Fatoohi
J. Supercomput.1
1993 Performance Analysis of Four SIMD Machines
abstract
This paper presents the results of an experiment to study the performance of four SIMD machines. The objectives of this study are to analyze the cost of regular communication on several SIMD machines and study its impact on the performance of two kernels. The machines are: a 32k processor CM2, a 16k processor MPP, a 16k processor MasPar MP-1, and a 4k processor DAP 610C. Regular communication is exemplified, in this study, by the shift operation where all elements of an array are shifted some number of positions along an array dimension. The cost of shift operations on the four machines is measured and analyzed for several two-dimensional arrays. The study shows that shift cost varies significantly from one machine to another, and depends on several factors including network topology, communication bandwidth, and the compilation partitioning scheme. Results also show that the communication overhead is quite significant on some of these machines even for nearest neighbor communication. Finally, results from this study are useful for obtaining a rough estimate of the communication overhead for many algorithms on these machines.
Rod A. Fatoohi
International Conference on Supercomputing1
1991 The NAS parallel benchmarks - summary and preliminary results
abstract
Article Free Access Share on The NAS parallel benchmarks—summary and preliminary results Authors: D. H. Bailey Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , E. Barszcz Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , J. T. Barton Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , D. S. Browning Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , R. L. Carter Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , L. Dagum Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , R. A. Fatoohi Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , P. O. Frederickson Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , T. A. Lasinski Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , R. S. Schreiber Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , H. D. Simon Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , V. Venkatakrishnan Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , S. K. Weeratunga Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile Authors Info & Claims Supercomputing '91: Proceedings of the 1991 ACM/IEEE conference on SupercomputingAugust 1991 Pages 158–165https://doi.org/10.1145/125826.125925Published:01 August 1991Publication History 405citation1,162DownloadsMetricsTotal Citations405Total Downloads1,162Last 12 Months161Last 6 weeks25 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
David H. Bailey, Eric Barszcz, John T. Barton, D. S. Browning, Robert L. Carter, Leonardo Dagum, Rod A. Fatoohi, Paul O. Frederickson, T. A. Lasinski, Robert Schreiber, Horst D. Simon, V. Venkatakrishnan, Sisira Weeratunga
SC7
1990 Vector performance analysis of the NEC SX-2
Rod A. Fatoohi
ICS1
1989 Vector performance analysis of three supercomputers: Cray 2, Cray Y-MP, and ETA 10-Q
abstract
This paper presents the results of a series of experiments to study the single processor performance of three supercomputers: Cray-2, Cray Y-MP, and ETA10-Q. The main object of this study is to determine the impact of certain architectural features on the performance of modern supercomputers. Features such as clock period, memory links, memory organization, multiple functional units, and chaining are considered here. A simple performance model is used to examine the impact of these features on the performance of a set of basic operations. The results of implementing this set on these machines for three vector lengths and three memory strides are presented and compared. For unit stride operations, the Cray Y-MP outperformed the Cray-2 by as much as three times and the ETA10-Q by as much as four times for these operations. Moreover, unlike the Cray-2 and ETA10-Q, even-numbered strides do not cause a major performance degradation on the Cray Y-MP. Two numerical algorithms are also used for comparison. For three problem sizes of both algorithms, the Cray Y-MP outperformed the Cray-2 by 43% to 68% and the ETA10-Q by four to eight times.
Rod A. Fatoohi
SC1
1989 Multitasking a Navier-Stokes algorithm on the CRAY-2
Rod A. Fatoohi
J. Supercomput.1
1987 Implementation of a Four Color Cell Relaxation Scheme on the MPP, Flex/32 and CRAY/2
Rod A. Fatoohi, Chester E. Grosch
ICPP1