Meenakshi A. Kandaswamy

dblp:08/2861 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2002
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 51% Storage systems · 41% Performance modeling and evaluation · 6%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
i/o optimization
0.132002
An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002
An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002
Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION · SC 1997
High-performance computing
parallel i/o
0.132002
An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002
An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002
Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION · SC 1997
High-performance computing › data-intensive computing
data-intensive applications
0.122002
An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002
An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002
Compilers and program optimization
loop transformation
0.012000
A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000
Storage systems › file systems › file organization
file layout optimization
0.012000
A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000
Storage systems
file systems
0.012000
A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000
High-performance computing
scientific computing systems
0.011997
Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION · SC 1997
Parallel and multicore computing › parallel programming models › message passing
distributed-memory message passing
0.012000
A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000

Methods — techniques the papers use, named apart from their topics

file layout optimization · 0.1collective i/o · 0.1iteration space transformation · 0.1data space transformation · 0.1prefetching · 0.0buffering · 0.0
YearPublicationVenuePosition
2002 An Experimental Evaluation of I/O Optimizations on Different Applications
abstract
Many large-scale applications have significant I/O requirements as well as computational and memory requirements. Unfortunately, the limited number of I/O nodes provided in a typical configuration of the modern message-passing distributed-memory architectures such as the Intel Paragon and the IBM SP-2 limits the I/O performance of these applications severely. In this paper, we examine some software optimization techniques and evaluate their effects in five different I/O-intensive codes from both small and large application domains. Our goals in this study are twofold. First, we want to understand the behavior of large-scale data-intensive applications and the impact of I/O subsystems on their performance and vice versa. Second, and more importantly, we strive to determine the solutions for improving the applications' performance by a mix of software techniques. Our results reveal that different applications can benefit from different optimizations. For example, we found that some applications benefit from file layout optimizations, whereas others take advantage of collective I/O. A combination of architectural and software solutions is normally needed to obtain good I/O performance. For example, we show that with a limited number of I/O resources, it is possible to obtain good performance by using appropriate software optimizations. We also show that beyond a certain level, imbalance in the architecture results in performance degradation even when using optimized software, thereby indicating the necessity of an increase in I/O resources.
Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt
IEEE Trans. Parallel Distributed Syst.1
2002 An Experimental Evaluation of I/O Optimizations on Different Applications
abstract
Many large scale applications have significant I/O requirements as well as computational and memory requirements. Unfortunately, the limited number of I/O nodes provided in a typical configuration of the modern message-passing distributed-memory architectures such as Intel Paragon and IBM SP-2 limits the I/O performance of these applications severely. We examine some software optimization techniques and evaluate their effects in five different I/O-intensive codes from both small and large application domains. Our goals in this study are twofold. First, we want to understand the behavior of large-scale data-intensive applications and the impact of I/O subsystems on their performance and vice versa. Second, and more importantly, we strive to determine the solutions for improving the applications' performance by a mix of software techniques. Our results reveal that different applications can benefit from different optimizations. For example, we found that some applications benefit from file layout optimizations whereas others take advantage of collective I/O. A combination of architectural and software solutions is normally needed to obtain good I/O performance. For example, we show that with a limited number of I/O resources, it is possible to obtain good performance by using appropriate software optimizations. We also show that beyond a certain level, imbalance in the architecture results in performance degradation even when using optimized software, thereby indicating the necessity of an increase in I/O resources.
Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt
IEEE Trans. Parallel Distributed Syst.1
2000 A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations
abstract
This paper presents a unified framework that optimizes out-of-core programs by exploiting locality and parallelism, and reducing communication overhead. For out-of-core problems where the data set sizes far exceed the size of the available in-core memory, it is particularly important to exploit the memory hierarchy by optimizing the I/O accesses. We present algorithms that consider both iteration space (loop) and data space (file layout) transformations in a unified framework. We show that the performance of an out-of-core loop nest containing references to out-of-core arrays can be improved by using a suitable combination of file layout choices and loop restructuring transformations. Our approach considers array references one-by-one and attempts to optimize each reference for parallelism and locality. When there are references for which parallelism optimizations do not work, communication is vectorized so that data transfer can be performed before the innermost loop. Results from hand-compiles on IBM SP-2 and Inter Paragon distributed-memory message-passing architectures show that this approach reduces the execution times and improves the overall speedups. In addition, we extend the base algorithm to work with file layout constraints and show how it is useful for optimizing programs that consist of multiple loop nests.
Mahmut T. Kandemir, Alok N. Choudhary, J. Ramanujam, Meenakshi A. Kandaswamy
IEEE Trans. Parallel Distributed Syst.4
1998 Performance Implications of Architectural and Software Techniques on I/O-Intensive Applications
abstract
Many large scale applications, have significant I/O requirements as well as computational and memory requirements. Unfortunately, limited number of I/O nodes provided by the contemporary message-passing distributed-memory architectures such as Intel Paragon and IBM SP-2 limits the I/O performance of these applications severely. In this paper, we examine some software optimization techniques and architectural scalability and evaluate the effect of them in five I/O intensive applications from both small and large application domains. Our goals in this study are twofold: First, we want to understand the behavior of large-scale data intensive applications and the impact of I/O subsystem on their performance and vice-versa. Second, and more importantly, we strive to determine the solutions for improving the applications' performance by a mix of architectural and software solutions. Our results reveal that the different applications can benefit from different optimizations. For example, we found that some applications benefit from file layout optimizations whereas some others benefit from collective I/O. A combination of architectural and software solutions is normally needed to obtain good I/O performance. For example, we show that with limited number of I/O resources, it is possible to obtain good performance by using appropriate software optimizations. We also show that beyond a certain level, imbalance in the architecture results in performance degradation even when using optimized software, thereby indicating the necessity of increase in I/O resources.
Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt
ICPP1
1997 Global I/O optimizations for out-of-core computations
abstract
The use of parallel machines to solve large-scale computational problems in science and engineering has increased considerably in recent times. Many of these problems have computational requirements which stretch the capabilities of even the fastest machine available today. In addition to requiring a great deal of computational power, these problems usually deal with large quantities of data up to a few terabytes. The main memory sizes of current parallel machines do not even come close to matching these requirements; hence data needs to be stored on disks and fetched during the execution of the program. Unfortunately, current optimizing compilers for parallel machines provide support only for in-core computations in which the data sets can fit into memory. This limitation severely affects the performance of programs which depend on disk-resident data. Our previous research demonstrated that file layout optimizations are extremely important for optimizing such programs. In this paper, we investigate solutions to the global I/O optimization problem for out-of-core computations. Since the general problem is NP-complete, we present fast heuristics that can result in near-optimal solutions for the programs encountered in practice. Preliminary results provide encouraging evidence that our algorithms can be successful in optimizing out-of-core programs.
Mahmut T. Kandemir, Meenakshi A. Kandaswamy, Alok N. Choudhary
HiPC2
1997 Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION
abstract
Parallel machines are an important part of the scientific application developer's tool box and the processing demands placed on these machines are rapidly increasing. Many scientific applications tend to perform high volume data storage, data retrieval and data processing, which demands high performance from the I/O subsystem. In this paper, we conduct an experimental study of the I/O performed by the Hartree-Fock (HF) method, as implemented using a fully distributed data approach in the NWChem parallel computational chemistry package. We use PASSION, a parallel and scalable I/O library to improve the I/O performance of the application and present extensive experimental results. The effects of both application-related factors and system-related factors on the application's I/O performance are studied in detail. We rank the optimizations based on the significance and impact on the performance of HF's I/O phase as: I. efficient interface to the file system, II. prefetching, and III. buffering. The results show that within the limits of our experimental framework, application-related factors are more effective on the overall I/O behavior of this application. We obtained up to 95% improvement in I/O time and 43% improvement in the overall application performance with the optimizations.
Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt
SC1