VLDB 2026 Research / reviewers in the wild / expert
Manojkumar Krishnan
dblp:46/6779
· DBLP profile ↗
13ranked-venue papers
4as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 92% High-performance computing · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational science and engineering · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
0.1 | 2 | 2006 | M12 - Overview of the global arrays parallel software development toolkit · SC 2006 Poster reception - Component architectures for quantum chemistry: forging new capabilities and insights · SC 2006 |
Computational science and engineering › computational chemistry
quantum chemistry |
0.1 | 1 | 2006 | Poster reception - Component architectures for quantum chemistry: forging new capabilities and insights · SC 2006 |
Parallel and multicore computing › parallel programming models › distributed memory programming models
global address space programming |
0.1 | 1 | 2006 | M12 - Overview of the global arrays parallel software development toolkit · SC 2006 |
Computational science and engineering
computational chemistry |
0.1 | 1 | 2005 | Multilevel Parallelism in Computational Chemistry using Common Component Architecture and Global Arrays · SC 2005 |
Parallel and multicore computing › parallelization strategies
multi-level parallelism |
0.0 | 1 | 2006 | Poster reception - Component architectures for quantum chemistry: forging new capabilities and insights · SC 2006 |
Parallel and multicore computing › parallel computing › parallel software engineering
parallel application development |
0.0 | 1 | 2006 | M12 - Overview of the global arrays parallel software development toolkit · SC 2006 |
Methods — techniques the papers use, named apart from their topics
common component architecture · 0.2global arrays · 0.1x10 · 0.1coarray fortran · 0.1UPC · 0.1MPI · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Tutorial StatementabstractThis tutorial will provide an overview of the Global Arrays (GA) programming toolkit and describe its capabilities, performance, and the use of GA in high performance computing applications. The tutorial will review basic concepts in onesided communication and Global Address Space languages and will discuss basic setup and elementary communication using GA. More advanced topics, including global counters, non-blocking communication, and sparse data structures will also be discussed. If time permits, the tutorial will also present new/advanced capabilities and ongoing research activities. Bruce J. Palmer, Manojkumar Krishnan, Abhinav Vishnu |
IPDPS | 2 |
| 2011 | Noncollective Communicator Creation in MPI
James Dinan, Sriram Krishnamoorthy, Pavan Balaji, Jeff R. Hammond, Manojkumar Krishnan, Vinod Tipparaju, Abhinav Vishnu |
EuroMPI | 5 |
| 2010 | Efficient On-Demand Connection Management Mechanisms with PGAS Models over InfiniBandabstractIn the last decade or so, clusters have observed a tremendous rise in popularity due to the excellent price to performance ratio. A variety of Interconnects have been proposed during this period, with InfiniBand leading the way due to its high performance and open standard. At the same time, multiple programming models have emerged in order to meet the requirements of various applications and their programming models. To support requirements of multiple programming models, InfiniBand provides multiple transport semantics, ranging from unreliable connectionless to reliable connected characteristics. Among them, the reliable connection (RC) semantics is being widely used due to its high performance and support for novel features like Remote Direct Memory Acesss (RDMA), hardware atomics and Network Fault Tolerance. However, the pair wise connection oriented nature of the RC transport semantics limits its scalability and usage at the increasing processor counts. In this paper, we design and implement on-demand connection management approaches in the context of Partitioned Global Address Space (PGAS) programming models, which provided shared memory abstraction and one-sided communication semantics, leading to the development of multiple languages (UPC, X10, Chapel) and libraries (Global Arrays, MPI-RMA). Using Global Arrays as the research vehicle, we implement this approach with Aggregate Remote Memory Copy Interface (ARMCI), the runtime system of Global Arrays. We evaluate our approach, ARMCI-On Demand Connection Management (ARMCI-ODCM) using various micro benchmarks and benchmarks (LU Factorization, Random-Access and Lennard Jones simulation) and application (Subsurface transport over multiple phases (STOMP)). With the performance evaluation for up to 4096 processors, we are able to have a multi-fold reduction in connection memory with a negligible degradation in performance. Using STOMP at 4096 processors, reduces the overall connection memory by 66 times with no performance degradation. To the best of our knowledge, this is the first design, implementation and evaluation of on-demand connection management with InfiniBand using PGAS models. Abhinav Vishnu, Manojkumar Krishnan |
CCGRID | 2 |
| 2009 | An efficient hardware-software approach to network fault tolerance with InfiniBandabstractIn the last decade or so, clusters have observed a tremendous rise in popularity due to excellent price to performance ratio. A variety of Interconnects have been proposed during this period, with InfiniBand leading the way due to its high performance and open standard. Increasing size of the InfiniBand clusters has reduced the mean time between failures of various components of these clusters tremendously. In this paper, we specifically focus on the network component failure and propose a hybrid hardware-software approach to handling network faults. The hybrid approach leverages the user-transparent network fault detection and recovery using Automatic Path Migration (APM), and the software approach is used in the wake of APM failure. Using Global Arrays as the programming model, we implement this approach with Aggregate Remote Memory Copy Interface (ARMCI), the runtime system of Global Arrays. We evaluate our approach using various benchmarks (siosi7, pentane, h2o7 and siosi3) with NWChem, a very popular ab initio quantum chemistry application. Using the proposed approach, the applications run to completion without restart on emulated network faults and acceptable overhead for benchmarks executing for a longer period of time. Abhinav Vishnu, Manojkumar Krishnan, Dhabaleswar K. Panda 0001 |
CLUSTER | 2 |
| 2007 | Scalable Visual Analytics of Massive Textual DatasetsabstractThis paper describes the first scalable implementation of a text processing engine used in visual analytics tools. These tools aid information analysts in interacting with and understanding large textual information content through visual interfaces. By developing a parallel implementation of the text processing engine, we enabled visual analytics tools to exploit cluster architectures and handle massive datasets. The paper describes key elements of our parallelization approach and demonstrates virtually linear scaling when processing multi-gigabyte data sets such as Pubmed. This approach enables interactive analysis of large datasets beyond capabilities of existing state-of-the art visual analytics tools. Manojkumar Krishnan, S. Bohn, W. Cowley, Vernon L. Crow, Jarek Nieplocha |
IPDPS | 1 |
| 2007 | Using the GA and TAO toolkits for solving large-scale optimization problems on parallel computersabstractChallenges in the scalable solution of large-scale optimization problems include the development of innovative algorithms and efficient tools for parallel data manipulation. This article discusses two complementary toolkits from the collection of Advanced CompuTational Software (ACTS), namely, Global Arrays (GA) for parallel data management and the Toolkit for Advanced Optimization (TAO), which have been integrated to support large-scale scientific applications of unconstrained and bound constrained minimization problems. Most likely to benefit are minimization problems arising in classical molecular dynamics, free energy simulations, and other applications where the coupling among variables requires dense data structures. TAO uses abstractions for vectors and matrices so that its optimization algorithms can easily interface to distributed data management and linear algebra capabilities implemented in the GA library. The GA/TAO interfaces are available both in the traditional library mode and as components compliant with the Common Component Architecture (CCA). We highlight the design of each toolkit, describe the interfaces between them, and demonstrate their use. Steven Benson, Manojkumar Krishnan, Lois C. McInnes, Jarek Nieplocha, Jason Sarich |
ACM Trans. Math. Softw. | 2 |
| 2006 | Poster reception - Component architectures for quantum chemistry: forging new capabilities and insightsabstractWe review the use of the Common Component Architecture approach within the quantum chemistry domain to tackle the software engineering challenges which arise as advanced algorithms are adopted and growing numbers of software packages are integrated to study complex, coupled physical phenomena. The development of common interfaces has allowed the adoption of advanced optimization solvers and high-level interchangeability of quantum chemistry packages. Components have been created which manage multiple levels of parallelism, providing much more efficient usage of parallel machines. Early efforts towards low-level integration of chemistry packages are examined. The ability to share intermediate data expands the capabilities available to any one software package, thereby enabling the rapid development of advanced methods. New methods for the study of reactions involving heavy elements, which depend on our component environment, are highlighted. Joseph P. Kenny, Curtis L. Janssen, Ida M. B. Nielsen, Manojkumar Krishnan, Vidhya Gurumoorthi, Edward F. Valeev, Theresa L. Windus |
SC | 4 |
| 2006 | M12 - Overview of the global arrays parallel software development toolkitabstractThe Global Arrays (GA) toolkit provides a global address space programming model to MPI applications. GA library allows programmers to distribute data while maintaining the type of global index space and programming syntax similar to what is available when programming on a single processor. The goal of GA is to free the programmers from the low level management of communication and allow them to deal with their problems at the level at which they were originally formulated. The compatibility of GA with MPI makes it possible to use distributed and global views of the data in the same application. The variety of applications that have been implemented using Global Arrays attests to the attractiveness of using higher level abstractions to write parallel code. The tutorial will provide an overview of GA toolkit, its applications, and compare GA to related programming models such as UPC, Co-Array Fortran, and X10. Jarek Nieplocha, Bruce J. Palmer, Manojkumar Krishnan, P. Sadayappan |
SC | 3 |
| 2005 | Multilevel Parallelism in Computational Chemistry using Common Component Architecture and Global ArraysabstractThe development of complex scientific applications for high-end systems is a challenging task. Addressing complexity of the involved software and algorithms is becoming increasingly difficult and requires appropriate software engineering approaches to address interoperability, maintenance, and software composition challenges. At the same time, the requirements for performance and scalability to thousand processor configurations magnifies the level of difficulties facing the scientific programmer due to the variable levels of parallelism available in different algorithms or functional modules of the application. This paper demonstrates how the Common Component Architecture (CCA) and Global Arrays (GA) can be used in context of computational chemistry to express and manage multi-level parallelism through the use of processor groups. For example, the numerical Hessian calculation using three levels of parallelism in NWChem computational chemistry package outperformed the original version of the NWChem code based on single level parallelism by a factor of 90% when running on 256 processors. Manojkumar Krishnan, Yuri Alexeev, Theresa L. Windus, Jarek Nieplocha |
SC | 1 |
| 2004 | Optimizing Parallel Multiplication Operation for Rectangular and Transposed Matrices
Manojkumar Krishnan, Jarek Nieplocha |
ICPADS | 1 |
| 2004 | SRUMMA: A Matrix Multiplication Algorithm Suitable for Clusters and Scalable Shared Memory SystemsabstractSummary form only given. We describe a novel parallel algorithm that implements a dense matrix multiplication operation with algorithmic efficiency equivalent to that of Cannon's algorithm. It is suitable for clusters and scalable shared memory systems. The current approach differs from the other parallel matrix multiplication algorithms by the explicit use of shared memory and remote memory access (RMA) communication rather than message passing. The experimental results on clusters (IBM SP, Linux-Myrinet) and shared memory systems (SGI Altix, Cray XI) demonstrate consistent performance advantages over pdgemm from the ScaLAPACK/PBBLAS suite, the leading implementation of the parallel matrix multiplication algorithms used today. In the best case on the SGI Altix, the new algorithm performs 20 times better than pdgemm for a matrix size of 1000 on 128 processors. The impact of zero-copy nonblocking RMA communications and shared memory communication on matrix multiplication performance on clusters are investigated. Manojkumar Krishnan, Jarek Nieplocha |
IPDPS | 1 |
| 2003 | Optimizing Mechanisms for Latency Tolerance in Remote Memory Access Communication on ClustersabstractThis paper describes the design and implementation of mechanisms for latency tolerance in the remote memory access communication on clusters equipped with high-performance networks such as Myrinet. It discusses strategies that bridge the gap between user-level requirements and network-specific communication interfaces while attempting to increase opportunities for latency hiding. Mechanisms for overlapping communication with computation and coalescing small messages (trading latency for bandwidth) are explored. The effectiveness of these techniques is evaluated using microbenchmarks and application kernels including the NAS parallel benchmark suite. The microbenchmark results showed a better degree of overlap for nonblocking operations in ARMCI as compared to MPI. Application results showed up 30% to 45% improvement over MPI on using nonblocking operations. The aggregation of small messages yielded performance improvement of up to 78% over non-aggregated communication. Jarek Nieplocha, Vinod Tipparaju, Manojkumar Krishnan, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda 0001 |
CLUSTER | 3 |
| 2003 | Exploiting Non-blocking Remote Memory Access Communication in Scientific Benchmarks
Vinod Tipparaju, Manojkumar Krishnan, Jarek Nieplocha, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda 0001 |
HiPC | 2 |