VLDB 2026 Research / reviewers in the wild / expert
Meenakshi A. Kandaswamy
dblp:08/2861
· DBLP profile ↗
6ranked-venue papers
4as first author
0since 2021 · last 2002
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 51% Storage systems · 41% Performance modeling and evaluation · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
i/o optimization |
0.1 | 3 | 2002 | An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002 An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002 Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION · SC 1997 |
High-performance computing
parallel i/o |
0.1 | 3 | 2002 | An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002 An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002 Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION · SC 1997 |
High-performance computing › data-intensive computing
data-intensive applications |
0.1 | 2 | 2002 | An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002 An Experimental Evaluation of I/O Optimizations on Different Applications · IEEE Trans. Parallel Distributed Syst. 2002 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 2000 | A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000 |
Storage systems › file systems › file organization
file layout optimization |
0.0 | 1 | 2000 | A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000 |
Storage systems
file systems |
0.0 | 1 | 2000 | A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1997 | Optimization and Evaluation of Hartree-Fock Application's I/O with PASSION · SC 1997 |
Parallel and multicore computing › parallel programming models › message passing
distributed-memory message passing |
0.0 | 1 | 2000 | A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core Computations · IEEE Trans. Parallel Distributed Syst. 2000 |
Methods — techniques the papers use, named apart from their topics
file layout optimization · 0.1collective i/o · 0.1iteration space transformation · 0.1data space transformation · 0.1prefetching · 0.0buffering · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2002 | An Experimental Evaluation of I/O Optimizations on Different ApplicationsabstractMany large-scale applications have significant I/O requirements as well as computational and memory requirements. Unfortunately, the limited number of I/O nodes provided in a typical configuration of the modern message-passing distributed-memory architectures such as the Intel Paragon and the IBM SP-2 limits the I/O performance of these applications severely. In this paper, we examine some software optimization techniques and evaluate their effects in five different I/O-intensive codes from both small and large application domains. Our goals in this study are twofold. First, we want to understand the behavior of large-scale data-intensive applications and the impact of I/O subsystems on their performance and vice versa. Second, and more importantly, we strive to determine the solutions for improving the applications' performance by a mix of software techniques. Our results reveal that different applications can benefit from different optimizations. For example, we found that some applications benefit from file layout optimizations, whereas others take advantage of collective I/O. A combination of architectural and software solutions is normally needed to obtain good I/O performance. For example, we show that with a limited number of I/O resources, it is possible to obtain good performance by using appropriate software optimizations. We also show that beyond a certain level, imbalance in the architecture results in performance degradation even when using optimized software, thereby indicating the necessity of an increase in I/O resources. Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2002 | An Experimental Evaluation of I/O Optimizations on Different ApplicationsabstractMany large scale applications have significant I/O requirements as well as computational and memory requirements. Unfortunately, the limited number of I/O nodes provided in a typical configuration of the modern message-passing distributed-memory architectures such as Intel Paragon and IBM SP-2 limits the I/O performance of these applications severely. We examine some software optimization techniques and evaluate their effects in five different I/O-intensive codes from both small and large application domains. Our goals in this study are twofold. First, we want to understand the behavior of large-scale data-intensive applications and the impact of I/O subsystems on their performance and vice versa. Second, and more importantly, we strive to determine the solutions for improving the applications' performance by a mix of software techniques. Our results reveal that different applications can benefit from different optimizations. For example, we found that some applications benefit from file layout optimizations whereas others take advantage of collective I/O. A combination of architectural and software solutions is normally needed to obtain good I/O performance. For example, we show that with a limited number of I/O resources, it is possible to obtain good performance by using appropriate software optimizations. We also show that beyond a certain level, imbalance in the architecture results in performance degradation even when using optimized software, thereby indicating the necessity of an increase in I/O resources. Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2000 | A Unified Framework for Optimizing Locality, Parallelism, and Communication in Out-of-Core ComputationsabstractThis paper presents a unified framework that optimizes out-of-core programs by exploiting locality and parallelism, and reducing communication overhead. For out-of-core problems where the data set sizes far exceed the size of the available in-core memory, it is particularly important to exploit the memory hierarchy by optimizing the I/O accesses. We present algorithms that consider both iteration space (loop) and data space (file layout) transformations in a unified framework. We show that the performance of an out-of-core loop nest containing references to out-of-core arrays can be improved by using a suitable combination of file layout choices and loop restructuring transformations. Our approach considers array references one-by-one and attempts to optimize each reference for parallelism and locality. When there are references for which parallelism optimizations do not work, communication is vectorized so that data transfer can be performed before the innermost loop. Results from hand-compiles on IBM SP-2 and Inter Paragon distributed-memory message-passing architectures show that this approach reduces the execution times and improves the overall speedups. In addition, we extend the base algorithm to work with file layout constraints and show how it is useful for optimizing programs that consist of multiple loop nests. Mahmut T. Kandemir, Alok N. Choudhary, J. Ramanujam, Meenakshi A. Kandaswamy |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 1998 | Performance Implications of Architectural and Software Techniques on I/O-Intensive ApplicationsabstractMany large scale applications, have significant I/O requirements as well as computational and memory requirements. Unfortunately, limited number of I/O nodes provided by the contemporary message-passing distributed-memory architectures such as Intel Paragon and IBM SP-2 limits the I/O performance of these applications severely. In this paper, we examine some software optimization techniques and architectural scalability and evaluate the effect of them in five I/O intensive applications from both small and large application domains. Our goals in this study are twofold: First, we want to understand the behavior of large-scale data intensive applications and the impact of I/O subsystem on their performance and vice-versa. Second, and more importantly, we strive to determine the solutions for improving the applications' performance by a mix of architectural and software solutions. Our results reveal that the different applications can benefit from different optimizations. For example, we found that some applications benefit from file layout optimizations whereas some others benefit from collective I/O. A combination of architectural and software solutions is normally needed to obtain good I/O performance. For example, we show that with limited number of I/O resources, it is possible to obtain good performance by using appropriate software optimizations. We also show that beyond a certain level, imbalance in the architecture results in performance degradation even when using optimized software, thereby indicating the necessity of increase in I/O resources. Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt |
ICPP | 1 |
| 1997 | Global I/O optimizations for out-of-core computationsabstractThe use of parallel machines to solve large-scale computational problems in science and engineering has increased considerably in recent times. Many of these problems have computational requirements which stretch the capabilities of even the fastest machine available today. In addition to requiring a great deal of computational power, these problems usually deal with large quantities of data up to a few terabytes. The main memory sizes of current parallel machines do not even come close to matching these requirements; hence data needs to be stored on disks and fetched during the execution of the program. Unfortunately, current optimizing compilers for parallel machines provide support only for in-core computations in which the data sets can fit into memory. This limitation severely affects the performance of programs which depend on disk-resident data. Our previous research demonstrated that file layout optimizations are extremely important for optimizing such programs. In this paper, we investigate solutions to the global I/O optimization problem for out-of-core computations. Since the general problem is NP-complete, we present fast heuristics that can result in near-optimal solutions for the programs encountered in practice. Preliminary results provide encouraging evidence that our algorithms can be successful in optimizing out-of-core programs. Mahmut T. Kandemir, Meenakshi A. Kandaswamy, Alok N. Choudhary |
HiPC | 2 |
| 1997 | Optimization and Evaluation of Hartree-Fock Application's I/O with PASSIONabstractParallel machines are an important part of the scientific application developer's tool box and the processing demands placed on these machines are rapidly increasing. Many scientific applications tend to perform high volume data storage, data retrieval and data processing, which demands high performance from the I/O subsystem. In this paper, we conduct an experimental study of the I/O performed by the Hartree-Fock (HF) method, as implemented using a fully distributed data approach in the NWChem parallel computational chemistry package. We use PASSION, a parallel and scalable I/O library to improve the I/O performance of the application and present extensive experimental results. The effects of both application-related factors and system-related factors on the application's I/O performance are studied in detail. We rank the optimizations based on the significance and impact on the performance of HF's I/O phase as: I. efficient interface to the file system, II. prefetching, and III. buffering. The results show that within the limits of our experimental framework, application-related factors are more effective on the overall I/O behavior of this application. We obtained up to 95% improvement in I/O time and 43% improvement in the overall application performance with the optimizations. Meenakshi A. Kandaswamy, Mahmut T. Kandemir, Alok N. Choudhary, David E. Bernholdt |
SC | 1 |