Piyush Mehrotra

dblp:33/354 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 6 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Performance modeling and evaluation · 32% GPUs and heterogeneous computing · 31% Parallel and multicore computing · 22%
Software engineering, system software, and programming languages
6 papers
Compilers and program optimization · 73% Runtime systems and virtual machines · 17% Programming languages and type systems · 10%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.212013
An early performance evaluation of many integrated core architecture based SGI rackable computing system · SC 2013
Parallel and multicore computing
parallel programming models
0.171994
On the design of Chant: a talking threads package · SC 1994
Dynamic data distributions in Vienna Fortran · SC 1993
High Performance Fortran Without Templates: An Alternative Model for Distribution and Alignment · PPoPP 1993
High-performance computing
performance optimization at scale
0.012013
An early performance evaluation of many integrated core architecture based SGI rackable computing system · SC 2013
Parallel and multicore computing › parallel programming models
message passing
0.021994
On the design of Chant: a talking threads package · SC 1994
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991
Data integration and cleaning
interoperability
0.011995
SmartFiles: An OO Approach to Data File Interoperabilty · OOPSLA 1995
Parallel and multicore computing
data distribution
0.011993
Dynamic data distributions in Vienna Fortran · SC 1993
Parallel and multicore computing › parallel programming models
data-parallel language
0.011993
High Performance Fortran Without Templates: An Alternative Model for Distribution and Alignment · PPoPP 1993
Parallel and multicore computing › parallel programming models › data-parallel language
high performance fortran
0.011993
High Performance Fortran Without Templates: An Alternative Model for Distribution and Alignment · PPoPP 1993
Distributed systems
distributed data structures
0.011992
Concurrent File Operations in a High Performance FORTRAN · SC 1992
High-performance computing
parallel i/o
0.011992
Concurrent File Operations in a High Performance FORTRAN · SC 1992
Compilers and program optimization
parallelizing compiler
0.011991
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991
Parallel and multicore computing › parallel computing
distributed execution
0.011991
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991
Memory systems › shared memory
distributed shared memory
0.011990
Supporting Shared Data Structures on Distributed Memory Architectures · PPoPP 1990
Parallel and multicore computing
shared data structures
0.011990
Supporting Shared Data Structures on Distributed Memory Architectures · PPoPP 1990
Performance modeling and evaluation
numerical algorithms
0.011989
Parallel language constructs for tensor product computations on loosely coupled architectures · SC 1989
Computational science and engineering
scientific data management
0.011995
SmartFiles: An OO Approach to Data File Interoperabilty · OOPSLA 1995
Compilers and program optimization › parallel language compilation
data-parallel compilation
0.011993
Dynamic data distributions in Vienna Fortran · SC 1993
Compilers and program optimization
dynamic optimization
0.011992
Concurrent File Operations in a High Performance FORTRAN · SC 1992
Programming languages and type systems
language design
0.011982
Language Concepts for Distributed Processing of Large Arrays · PODC 1982
High-performance computing
distributed memory systems
0.011989
Parallel language constructs for tensor product computations on loosely coupled architectures · SC 1989

Methods — techniques the papers use, named apart from their topics

microbenchmarks · 0.2OpenMP · 0.2NAS Parallel Benchmarks · 0.2MPI · 0.2object-oriented methodology · 0.0message passing · 0.0lightweight thread scheduling · 0.0data distribution · 0.0i/o constructs · 0.0data transfer optimization · 0.0compile-time analysis · 0.0run-time code generation · 0.0program analysis · 0.0message-passing transformation · 0.0abstract data types · 0.0
YearPublicationVenuePosition
2016 Performance evaluation of Amazon Elastic Compute Cloud for NASA high-performance computing applications
abstract
Abstract Cloud computing environments are now widely available and are being increasingly utilized for technical computing. They are also being touted for high‐performance computing (HPC) applications in science and engineering. For example, Amazon Elastic Compute Cloud (EC2) Services offers specialized Cluster Compute instance types to run HPC applications. In this paper, we compare the performance characteristics of two Amazon EC2 HPC instance types with that of National Aeronautics and Space Administration's (NASA) Pleiades supercomputer, an SGI® ICE™ cluster. For this study, we utilized the HPC Challenge kernels and the NAS Parallel Benchmarks along with four full‐scale applications from the repertoire of codes that are being used by NASA scientists and engineers. We compare the total runtime of these codes for varying number of cores. We also break out the computation and communication times for a subset of these applications to explore the effect of interconnect differences on the two systems. In general, the single node performance of the two platforms is equivalent. However, for most of the codes when scaling to larger core counts, the performance of the EC2 HPC instances generally lags that of Pleiades because of worse network performance of the former. In addition to analyzing application performance, we also briefly touch upon the overhead due to virtualization and the usability of cloud environments such as Amazon EC2. Published 2013. This article is a U.S. Government work and is in the public domain in the U.S.A.
Piyush Mehrotra, M. Jahed Djomehri, Steve Heistand, Robert Hood, Haoqiang Jin, Arthur Lazanoff, Subhash Saini, Rupak Biswas
Concurr. Comput. Pract. Exp.1
2013 An early performance evaluation of many integrated core architecture based SGI rackable computing system
abstract
Intel recently introduced the Xeon Phi coprocessor based on the Many Integrated Core architecture featuring 60 cores with a peak performance of 1.0 Tflop/s. NASA has deployed a 128-node SGI Rackable system where each node has two Intel Xeon E2670 8-core Sandy Bridge processors along with two Xeon Phi 5110P coprocessors. We have conducted an early performance evaluation of the Xeon Phi. We used microbenchmarks to measure the latency and bandwidth of memory and interconnect, I/O rates, and the performance of OpenMP directives and MPI functions. We also used OpenMP and MPI versions of the NAS Parallel Benchmarks along with two production CFD applications to test four programming modes: offload, processor native, coprocessor native and symmetric (processor plus coprocessor). In this paper we present preliminary results based on our performance evaluation of various aspects of a Phi-based system.
Subhash Saini, Haoqiang Jin, Dennis C. Jespersen, Huiyu Feng, M. Jahed Djomehri, William Arasin, Robert Hood, Piyush Mehrotra, Rupak Biswas
SC8
2012 I/O performance characterization of Lustre and NASA applications on Pleiades
abstract
In this paper we study the performance of the Lustre file system using five scientific and engineering applications representative of NASA workload on large-scale supercomputing systems such as NASA's Pleiades. In order to facilitate the collection of Lustre performance metrics, we have developed a software tool that exports a wide variety of client and server-side metrics using SGI's Performance Co-Pilot (PCP), and generates a human readable report on key metrics at the end of a batch job. These performance metrics are (a) amount of data read and written, (b) number of files opened and closed, and (c) remote procedure call (RPC) size distribution (4 KB to 1024 KB, in powers of 2) for I/O operations. RPC size distribution measures the efficiency of the Lustre client and can pinpoint problems such as small write sizes, disk fragmentation, etc. These extracted statistics are useful in determining the I/O pattern of the application and can assist in identifying possible improvements for users' applications. Information on the number of file operations enables a scientist to optimize the I/O performance of their applications. Amount of I/O data helps users choose the optimal stripe size and stripe count to enhance I/O performance. In this paper, we demonstrate the usefulness of this tool on Pleiades for five production quality NASA scientific and engineering applications. We compare the latency of read and write operations under Lustre to that with NFS by tracing system calls and signals. We also investigate the read and write policies and study the effect of page cache size on I/O operations. We examine the performance impact of Lustre stripe size and stripe count along with performance evaluation of file per process and single shared file accessed by all the processes for NASA workload using parameterized IOR benchmark.
Subhash Saini, Jason Rappleye, Johnny Chang, David Barker, Piyush Mehrotra, Rupak Biswas
HiPC5
2011 The impact of hyper-threading on processor resource utilization in production applications
abstract
Intel provides Hyper-Threading (HT) in processors based on its Pentium and Nehalem micro-architecture such as the Westmere-EP. HT enables two threads to execute on each core in order to hide latencies related to data access. These two threads can execute simultaneously, filling unused stages in the functional unit pipelines. To aid better understanding of HT-related issues, we collect Performance Monitoring Unit (PMU) data (instructions retired; unhalted core cycles; L2 and L3 cache hits and misses; vector and scalar floating-point operations, etc.). We then use the PMU data to calculate a new metric of efficiency in order to quantify processor resource utilization and make comparisons of that utilization between single-threading (ST) and HT modes. We also study performance gain using unhalted core cycles, code efficiency of using vector units of the processor, and the impact of HT mode on various shared resources like L2 and L3 cache. Results using four full-scale, production-quality scientific applications from computational fluid dynamics (CFD) used by NASA scientists indicate that HT generally improves processor resource utilization efficiency, but does not necessarily translate into overall application performance gain.
Subhash Saini, Haoqiang Jin, Robert Hood, David Barker, Piyush Mehrotra, Rupak Biswas
HiPC5
2011 Performance Analysis of CFD Application Cart3D Using MPInside and Performance Monitor Unit Data on Nehalem and Westmere Based Supercomputers
abstract
Cart3D is a computational fluid dynamics (CFD) application aimed at conceptual and preliminary design of aerospace vehicles with complex geometries. It is widely used by design engineers at NASA, Department of Defense and aerospace companies in the USA. We present detailed performance analysis of Cart3D using two tools SGI MPInside and op_scope that collects hardware counter data from Intel Performance Monitoring Unit (PMU) on supercomputers based on Nehalem micro-architecture. Using these tools, we have done dynamic profiling of Cart3D (compute time, communication time and I/O time), along with dynamic profiling of MPI functions (MPI_Sendrecv, MPI_Bcast, MPI_Isend, MPI_Irecv, MPI_Allreduce, MPI_Barrier, etc.) with respect to message size of each rank and time consumed by each function. MPI communication is further analyzed by studying the performance of MPI functions used in this application as a function of message size and number of cores. Using these tools we have also studied efficiency of the processor to measure its effective utilization, efficiency of the floating-point units, percentage of vectorization and percentage of data coming from L2 cache, L3 cache, and main memory. This study was performed on two computing sub-systems based on quad-core Nehalem-EP and hex-core West mere-EP processors that are part of Pleiades an SGI Altix ICE at NASA Ames Research Center.
Subhash Saini, Piyush Mehrotra, Kenichi Taylor, Michael J. Aftosmis, Rupak Biswas
HPCC2
2011 High performance computing using MPI and OpenMP on multi-core parallel systems
Haoqiang Jin, Dennis C. Jespersen, Piyush Mehrotra, Rupak Biswas, Lei Huang 0006, Barbara M. Chapman
Parallel Comput.3
2010 Performance Analysis of Scientific and Engineering Applications Using MPInside and TAU
abstract
In this paper, we present performance analysis of two NASA applications using performance tools like Tuning and Analysis Utilities (TAU) and SGI MP Inside. MITgcmUV and OVERFLOW are two production-quality applications used extensively by scientists and engineers at NASA. MITgcmUV is a global ocean simulation model, developed by the Estimating the Circulation and Climate of the Ocean (ECCO) Consortium, for solving the fluid equations of motion using the hydrostatic approximation. OVERFLOW is a general-purpose Navier-Stokes solver for computational fluid dynamics (CFD) problems. Using these tools, we analyze the MPI functions (MPI_Sendrecv, MPI_Bcast, MPI_Reduce, MPI_Allreduce, MPI_Barrier, etc.) with respect to message size of each rank, time consumed by each function, and how ranks communicate. MPI communication is further analyzed by studying the performance of MPI functions used in these two applications as a function of message size and number of cores. Finally, we present the compute time, communication time, and I/O time as a function of the number of cores.
Subhash Saini, Piyush Mehrotra, Kenichi Taylor, Sameer Shende, Rupak Biswas
HPCC2
2010 Performance impact of resource contention in multicore systems
abstract
Resource sharing in commodity multicore processors can have a significant impact on the performance of production applications. In this paper we use a differential performance analysis methodology to quantify the costs of contention for resources in the memory hierarchy of several multicore processors used in high-end computers. In particular, by comparing runs that bind MPI processes to cores in different patterns, we can isolate the effects of resource sharing. We use this methodology to measure how such sharing affects the performance of four applications of interest to NASA-OVERFLOW, MITgcm, Cart3D, and NCC. We also use a subset of the HPCC benchmarks and hardware counter data to help interpret and validate our findings. We conduct our study on high-end computing platforms that use four different quad-core microprocessors-Intel Clovertown, Intel Harpertown, AMD Barcelona, and Intel Nehalem-EP. The results help further our understanding of the requirements these codes place on their production environments and also of each computer's ability to deliver performance.
Robert Hood, Haoqiang Jin, Piyush Mehrotra, Johnny Chang, M. Jahed Djomehri, Sharad Gavali, Dennis C. Jespersen, Kenichi Taylor, Rupak Biswas
IPDPS3
2008 The development of a geospatial data Grid by integrating OGC Web services with Globus-based Grid technology
abstract
Abstract Geospatial science is the science and art of acquiring, archiving, manipulating, analyzing, communicating, modeling with, and utilizing spatially explicit data for understanding physical, chemical, biological, and social systems on the Earth's surface or near the surface. In order to share distributed geospatial resources and facilitate the interoperability, the Open Geospatial Consortium (OGC), an industry–government–academia consortium, has developed a set of widely accepted Web‐based interoperability standards and protocols. Grid is the technology enabling resource sharing and coordinated problem solving in dynamic, multi‐institutional virtual organizations. Geospatial Grid is an extension and application of Grid technology in the geospatial discipline. This paper discusses problems associated with directly using Globus‐based Grid technology in the geospatial disciplines, the needs for geospatial Grids, and the features of geospatial Grids. Then, the paper presents a research project that develops and deploys a geospatial Grid through integrating Web‐based geospatial interoperability standards and technology developed by OGC with Globus‐based Grid technology. The geospatial Grid technology developed by this project makes the interoperable, personalized, on‐demand data access and services a reality at large geospatial data archives. Such a technology can significantly reduce problems associated with archiving, manipulating, analyzing, and utilizing large volumes of geospatial data at distributed locations. Copyright © 2008 John Wiley & Sons, Ltd.
Liping Di, Aijun Chen, Wenli Yang 0002, Yang Liu 0051, Yaxing Wei, Piyush Mehrotra, Chaumin Hu, Dean N. Williams
Concurr. Comput. Pract. Exp.6
2006 ScyFlow: an environment for the visual specification and execution of scientific workflows
abstract
Abstract With the advent of Grid technologies, scientists and engineers are building more complex applications to utilize distributed Grid resources. Core Grid services provide a path for accessing and utilizing these resources in a secure and seamless fashion. However, what the scientists need is an environment that will allow them to specify their application runs at a high organizational level, and then will support efficient execution across any given set or sets of distributed resources. We have been designing and implementing ScyFlow, a dual‐interface architecture, both Graphical User Interface (GUI) and Application Programming Interface (API), that addresses this problem. The scientist/user specifies the application tasks along with the necessary control and data flow, and then monitors and manages the execution of the resulting workflow across the distributed resources. In this paper, we utilize two scenarios to provide the details of the two modules of the project, the visual editor and the runtime workflow engine. Published in 2005 by John Wiley & Sons, Ltd.
Karen M. McCann, Maurice Yarrow, Adrian De Vivo, Piyush Mehrotra
Concurr. Comput. Pract. Exp.4
2002 A Resource Brokering Infrastructure for Computational Grids
Ahmed Al-Theneyan, Piyush Mehrotra, Mohammad Zubair
HiPC2
2002 XML-based visual specification of multidisciplinary applications
Ahmed Al-Theneyan, Amol Jakatdar, Piyush Mehrotra, Mohammad Zubair
Future Gener. Comput. Syst.3
2001 XML-Based Visual Specification of Multidisciplinary Applications
abstract
The advancements in the Internet and Web technologies have fueled a growing interest in developing a Web-based distributed computing environment. We have designed and developed Arcade, a Web-based environment for designing, executing, monitoring and controlling distributed heterogeneous applications, which is easy to use and access, portable, and provides support through all phases of the application development and execution. A major focus of the environment is the specification of heterogeneous multidisciplinary applications. We focus on the visual and script-based specification interface of Arcade. The Web/browser-based visual interface is designed to be intuitive to use and can also be used for visual monitoring during execution. The script specification is based on XML to: make it portable across different frameworks; and make the development of our tools easier by using the existing freely available XML parsers and editors. There is a one-to-one correspondence between the visual and script-based interfaces allowing users to go back and forth between the two. To support this we have developed translators that translate a script-based specification to a visual-based specification and vice-versa. These translators are integrated with our tools and are transparent to users.
Ahmed Al-Theneyan, Amol Jakatdar, Mohammad Zubair, Piyush Mehrotra
CCGRID4
2001 High Performance Fortran for aerospace applications
Piyush Mehrotra, Hans P. Zima
Parallel Comput.1
2000 On the implementation of the Opus coordination language
abstract
Opus is a new programming language designed to assist in coordinating the execution of multiple, independent program modules. With the help of Opus, coarse grained task parallelism between data parallel modules can be expressed in a clean and structured way. In this paper we address the problems of how to build a compilation and runtime support system that can efficiently implement the Opus constructs. Our design considers the often-conflicting goals of efficiency and modular construction through software re-use. In particular, we present the system requirements for an efficient Opus implementation, the Opus runtime system, and describe how they work together to provide the underlying services that the Opus compiler needs for a broad class of machines. Copyright © 2000 John Wiley & Sons, Ltd.
Erwin Laure, Matthew Haines, Piyush Mehrotra, Hans P. Zima
Concurr. Pract. Exp.3
1999 Compiling Data Parallel Tasks for Coordinated Execution
Erwin Laure, Matthew Haines, Piyush Mehrotra, Hans P. Zima
Euro-Par3
1998 OpenMP and HPF: Integrating Two Paradigms
Barbara M. Chapman, Piyush Mehrotra
Euro-Par2
1998 High-level Management of Communication Schedules in HPF-like Languages
abstract
The goal of High Performance Fortran (HPF) is to "address the problems of writing data parallel programs where the distribution of data affects performance",
Siegfried Benkner, Piyush Mehrotra, John Van Rosendale, Hans P. Zima
International Conference on Supercomputing2
1998 High Performance Fortran: History, Status and Future
Piyush Mehrotra, John Van Rosendale, Hans P. Zima
Parallel Comput.1
1997 Web-based Framework for Distributed Computing
abstract
Parallel and distributed computing on a cluster of workstations is being increasingly applied to a variety of large size computational problems. Several software systems have been developed that make distributed computing available to an application programmer. However, these systems either are not Web-based or lack a collaborative environment. The increasing use of Web technology for Internet and Intranet applications is making the Web an attractive framework for solving distributed applications, in particular, because the interface can be made platform-independent. In this paper we describe JAVADC, a Web–Java-based framework for the execution of parallel SPMD applications which use PVM, pPVM and MPI. We also discuss the design of a collaborative, distributed computing environment, Arcade, focused on more general programming paradigms for multidisciplinary applications. Arcade is a Web-based integrated environment which provides support in all phases of the development of general multidisciplinary applications including the design, execution, monitoring and control of such applications. © 1997 John Wiley & Sons, Ltd.
Kurt Maly, Piyush Mehrotra, Praveen K. Vangala, Mohammad Zubair
Concurr. Pract. Exp.3
1996 Special Issue on Multithreading for Multiprocessors: Guest Editors' Introduction
Matthew Haines, Piyush Mehrotra
J. Parallel Distributed Comput.2
1995 On the Utility of Threads for Data Parallel Programming
abstract
Threads provide a useful programming model for asynchronous behavior because of their ability to encapsulate units of work that can then be scheduled for execution at runtime, based on the dynamic state of a system.Recently, the threaded model has been applied to the domain of data parallel scientific codes, and initial reports indicate that the threaded model can produce performance gains over non-threaded approaches, primarily through the use of overlapping useful computation with communicant ion latency.However, overlapping computation with communication is possible without the benefit of threads if the communication system supports asynchronous primitives, and this comparison has not been made in previous papers.This paper provides a critical look at the utility of lightweight threads as applied to data parallel scientific programming.
Thomas Fahringer, Matthew Haines, Piyush Mehrotra
International Conference on Supercomputing3
1995 SmartFiles: An OO Approach to Data File Interoperabilty
abstract
Data files for scientific and engineering codes typically consist of a series of raw data values whose description is buried in the programs that interact with these files. In this situation, making even minor changes in the file structure or sharing files between programs (interoperability) can only be done after careful examination of the data files and the I/O statements of the programs interacting with this file. In short, scientific data files lack self-description, and other self-describing data techniques are not always appropriate or useful for scientific data files. By applying an object-oriented methodology to data files, we can add the intelligence required to improve data interoperability and provide an elegant mechanism for supporting complex, evolving, or multidisciplinary applications, while still supporting legacy codes. As a result, scientists and engineers should be able to share datasets with far greater ease, simplifying multidisciplinary applications and greatly facilitating remote collaboration between scientists.
Matthew Haines, Piyush Mehrotra, John Van Rosendale
OOPSLA2
1995 High-Level Languages for Parallel Scientific Computing
Barbara M. Chapman, Piyush Mehrotra, Hans P. Zima
SOFSEM2
1995 High Performance Fortran Languages: Advanced applications and their implementation
Barbara M. Chapman, Piyush Mehrotra, Hans P. Zima
Future Gener. Comput. Syst.2
1994 Extending Vienna Fortran with Task Parallelism
abstract
Vienna Fortran supports a wide range of data-parallel numerical problems. However, a significant number of scientific and engineering applications are of a multi-disciplinary and heterogeneous nature and thus do not fit well into the data parallel paradigm. In this paper we present new language extensions to fill this gap. Tasks can be spawned as asynchronous activities in a homogeneous or heterogeneous computing environment; they interact by sharing access to Shared Data Abstractions (SDAs). SDAs are an extension of Fortran 90 modules, representing a pool of common data, together with a set of methods for controlled access to these data and a mechanism for providing persistent storage. These extensions support the integration of data and task parallelism and can be used to express task parallel applications in a natural and efficient way.
Barbara M. Chapman, Piyush Mehrotra, John Van Rosendale, Hans P. Zima
ICPADS2
1994 On the design of Chant: a talking threads package
abstract
Lightweight threads are becoming increasingly useful for supporting parallelism and asynchronous control structures in applications and language implementations. However, lightweight thread packages for distributed memory systems have received little attention. We introduce the design of a runtime interface, called Chant, that supports communicating threads in a distributed memory environment. In particular, Chant is layered atop standard message passing and lightweight thread libraries, and supports efficient point-to-point and remote service request communication primitives. We examine the design issues of Chant, the efficiency of its point-to-point communication layer, and the evaluation of scheduling policies to poll for the presence of incoming messages.>
Matthew Haines, David Cronk, Piyush Mehrotra
SC3
1993 High Performance Fortran Without Templates: An Alternative Model for Distribution and Alignment
abstract
article High performance Fortran without templates: an alternative model for distribution and alignment. Share on Authors: Barbara M. Chapman View Profile , Piyush Mehrotra View Profile , Hans P. Zima View Profile Authors Info & Claims ACM SIGPLAN NoticesVolume 28Issue 7July 1993 pp 92–101https://doi.org/10.1145/173284.155342Online:01 July 1993Publication History 8citation192DownloadsMetricsTotal Citations8Total Downloads192Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Barbara M. Chapman, Piyush Mehrotra, Hans P. Zima
PPoPP2
1993 Dynamic data distributions in Vienna Fortran
abstract
No abstract available.
Barbara M. Chapman, Piyush Mehrotra, Hans Moritsch, Hans P. Zima
SC2
1992 Concurrent File Operations in a High Performance FORTRAN
abstract
The authors propose constructs to specify I/O (input/output) operations for distributed data structures in the context of Vienna FORTRAN. These operations can be used by the programmer to provide information which will allow the compiler and runtime environment to optimize the transfer of data to and from secondary storage. Although the language constructs presented have been proposed in the context of Vienna FORTRAN, they can be easily integrated into any other high-performance FORTRAN extension.>
Peter Brezany, Michael Gerndt, Piyush Mehrotra, Hans P. Zima
SC3
1991 Programming data parallel algorithms on distributed memory using Kali
abstract
Current languages for distributed memory machines tend to directly reflect the underlying hardware and thus provide little support for implementing data parallel algorithms.This paper describes a programming environment, Kali, which provides a global name space and allows direct access to remote data values.In order to retain efficiency, Kali provides a system of annotations allowing the user to control aspects of the program critical to performance, such as data distribution and load balancing.We present a series of examples showing how Kali can easily express data parallel algorithms.We also discuss some of the issues raised in translating such programs for execution on distributed memory systems and show performance results for well-known numerical algorithms written in Kali.
Charles Koelbel, Piyush Mehrotra
ICS2
1991 Performance of Hashed Cache Data Migration Schemes on Multicomputers
Seema Hiranandani, Joel H. Saltz, Piyush Mehrotra, Harry Berryman
J. Parallel Distributed Comput.3
1991 Compiling Global Name-Space Parallel Loops for Distributed Execution
abstract
Compiler support required to allow programmers to express their algorithms using a global name-space is discussed. A general method for the analysis of a high-level source program and its translation into a set of independently executing tasks that communicate using messages is presented. It is shown that if the compiler has enough information, the translation can be carried out at compile time; otherwise; run-time code is generated to implement the required data movement. The analysis required in both situations is described, and the performance of the generated code on the Intel iPSC/2 hypercube is presented.>
Charles Koelbel, Piyush Mehrotra
IEEE Trans. Parallel Distributed Syst.2
1990 Supporting Shared Data Structures on Distributed Memory Architectures
abstract
Programming nonshared memory systems is more difficult than programming shared memory systems, since there is no support for shared data structures. Current programming languages for distributed memory architectures force the user to decompose all data structures into separate pieces, with each piece “owned” by one of the processors in the machine, and with all communication explicitly specified by low-level message-passing primitives. This paper presents a new programming environment for distributed memory architectures, providing a global name space and allowing direct access to remote parts of data values. We describe the analysis and program transformations required to implement this environment, and present the efficiency of the resulting code on the NCUBE/7 and IPSC/2 hypercubes.
Charles Koelbel, Piyush Mehrotra, John Van Rosendale
PPoPP2
1989 Parallel language constructs for tensor product computations on loosely coupled architectures
abstract
Distributed memory architectures offer high levels of performance and flexibility, but have proven awkward to program. Current languages for nonshared memory architectures provide a relatively low-level programming environment, and are poorly suited to modular programming, and to the construction of libraries. This paper describes a set of language primitives designed to allow the specification of parallel numerical algorithms at a higher level. We focus here on tensor product array computations, a simple but important class of numerical algorithms. We consider first the problem of programming one dimensional “kernel” routines, such as parallel tridiagonal solvers, and after that look at how such parallel kernels can be combined to form parallel tensor product algorithms.
Piyush Mehrotra, John Van Rosendale
SC1
1987 Semi-Automatic Domain Decomposition in BLAZE
Charles Koelbel, Piyush Mehrotra, John Van Rosendale
ICPP2
1987 The BLAZE language: A parallel language for scientific programming
Piyush Mehrotra, John Van Rosendale
Parallel Comput.1
1983 The FEM-2 Design Method
Terrence W. Pratt, Loyce M. Adams, Piyush Mehrotra, John Van Rosendale, Robert G. Voigt, Merrell L. Patrick
ICPP3
1982 Language Concepts for Distributed Processing of Large Arrays
abstract
A large array is an array whose storage is distributed among primary and secondary storage and whose processing may be distributed among several tasks in a distributed system. This paper presents a semantic model (set of language concepts) for representing large arrays in a distributed system in such a way that the performance realities inherent in the distributed storage and processing can be adequately represented. An implementation of the large array concept as an ADA package (abstract data type) is described, as well as a particular tailoring of the concept for the NASA Finite Element Machine. An example application program using the package is also described.
Piyush Mehrotra, Terrence W. Pratt
PODC1