EDBT 2026 Demo / reviewers in the wild / expert
Abhinav Vishnu
dblp:92/2581
· DBLP profile ↗
57ranked-venue papers
17as first author
3since 2021 · last 2023
0000-0002-0593-4780ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 17 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | ADARNet: Deep Learning Predicts Adaptive Mesh RefinementabstractDeep Learning (DL) algorithms have gained popularity for super-resolution tasks - reconstructing a high-resolution (HR) output from its low-resolution (LR) counterpart. However, current DL approaches, both in computer vision and computational fluid dynamics (CFD), perform spatially uniform super-resolution. Therefore, DL for CFD approaches often over-resolve regions of the LR input that are already accurate at low numerical precision. This hardware over-utilization limits their scalability. To address this limitation, we propose ADARNet, a DL-based adaptive mesh refinement (AMR) framework. ADARNet takes a LR image as input and outputs its non-uniform HR counterpart, predicting HR only in areas that require higher numerical accuracy. As a result, ADARNet predicts the target 1024 × 1024 solution 7 − 28.5 × faster than state-of-the-art DL methods and reduces the memory usage by 4.4 − 7.65 × while maintaining the same level of accuracy. Moreover, unlike traditional AMR solvers that refine the mesh iteratively, ADARNet is a one-shot method that accelerates it by 2.6 − 4.5 ×. Octavi Obiols-Sales, Abhinav Vishnu, Nicholas Malaya, Aparna Chandramowlishwaran |
ICPP | 2 |
| 2023 | A Research Retrospective on AMD's Exascale Computing JourneyabstractThe pace of advancement of the top-end supercomputers historically followed an exponential curve similar to (and driven in part by) Moore's Law. Shortly after hitting the petaflop mark, the community started looking ahead to the next milestone: Exascale. However, many obstacles were already looming on the horizon, such as the slowing of Moore's Law, and others like the end of Dennard Scaling had already arrived. Anticipating significant challenges for the overall high-performance computing (HPC) community to achieve the next 1000x improvement, the U.S. Department of Energy (DOE) launched the Exascale Computing Program to enable and accelerate fundamental research across the many technologies needed to achieve exascale computing. Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang 0004, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, René van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov |
ISCA | 65 |
| 2021 | SURFNet: Super-Resolution of Turbulent Flows with Transfer Learning using Small DatasetsabstractDeep Learning (DL) algorithms are emerging as a key alternative to computationally expensive CFD simulations. However, state-of-the-art DL approaches require large and high-resolution training data to learn accurate models. The size and availability of such datasets are a major limitation for the development of next-generation data-driven surrogate models for turbulent flows. This paper introduces SURFNet, a transfer learning-based super-resolution flow network. SURFNet primarily trains the DL model on low-resolution datasets and transfer learns the model on a handful of high-resolution flow problems-accelerating the traditional numerical solver independent of the input size. We propose two approaches to transfer learning for the task of super-resolution, namely one-shot and incremental learning. Both approaches entail transfer learning on only one geometry to account for fine-grid flow fields requiring 15× less training data on high-resolution inputs compared to the tiny resolution ($64\times 256$) of the coarse model significantly, reducing the time for both data collection and training. We empirically evaluate SURFNet's performance by solving the Navier-Stokes equations in the turbulent regime on input resolutions up to 256× larger than the coarse model. On four test geometries and eight flow configurations unseen during training, we observe a consistent 2–2.1× speedup over the OpenFOAM physics solver independent of the test geometry and the resolution size (up to$2048 \times 2048$), demonstrating both resolution-invariance and generalization capabilities. Moreover, compared to the baseline model (aka oracle) that collects large training data at$256 \times 256$and$512 \times 512$grid resolutions, SURFNet achieves the same performance gain while reducing the combined data collection and training time by 3.6× and 10.2×, respectively. Our approach addresses the challenge of reconstructing high-resolution solutions from coarse grid models trained using low-resolution inputs (i.e., super-resolution) without loss of accuracy and requiring limited computational resources. Octavi Obiols-Sales, Abhinav Vishnu, Nicholas Malaya, Aparna Chandramowlishwaran |
PACT | 2 |
| 2020 | CFDNet: a deep learning-based accelerator for fluid simulationsabstractCFD is widely used in physical system design and optimization, where it is used to predict engineering quantities of interest, such as the lift on a plane wing or the drag on a motor vehicle. However, many systems of interest are prohibitively expensive for design optimization, due to the expense of evaluating CFD simulations. Octavi Obiols-Sales, Abhinav Vishnu, Nicholas Malaya, Aparna Chandramowlishwaran |
ICS | 2 |
| 2020 | Scaling Deep Learning workloads: NVIDIA DGX-1/Pascal and Intel Knights Landing
Nitin Gawande, Jeff Daily, Charles Siegel, Nathan R. Tallent, Abhinav Vishnu |
Future Gener. Comput. Syst. | 5 |
| 2019 | Kleio: A Hybrid Memory Page Scheduler with Machine IntelligenceabstractThe increasing demand of big data analytics for more main memory capacity in datacenters and exascale computing environments is driving the integration of heterogeneous memory technologies. The new technologies exhibit vastly greater differences in access latencies, bandwidth and capacity compared to the traditional NUMA systems. Leveraging this heterogeneity while also delivering application performance enhancements requires intelligent data placement. We present Kleio, a page scheduler with machine intelligence for applications that execute across hybrid memory components. Kleio is a hybrid page scheduler that combines existing, lightweight, history-based data tiering methods for hybrid memory, with novel intelligent placement decisions based on deep neural networks. We contribute new understanding toward the scope of benefits that can be achieved by using intelligent page scheduling in comparison to existing history-based approaches, and towards the choice of the deep learning algorithms and their parameters that are effective for this problem space. Kleio incorporates a new method for prioritizing pages that leads to highest performance boost, while limiting the resulting system resource overheads. Our performance evaluation indicates that Kleio reduces on average 80% of the performance gap between the existing solutions and an oracle with knowledge of future access pattern. Kleio provides hybrid memory systems with fast and effective neural network training and prediction accuracy levels, which bring significant application performance improvements with limited resource overheads, so as to lay the grounds for its practical integration in future systems. Thaleia Dimitra Doudali, Sergey Blagodurov, Abhinav Vishnu, Sudhanva Gurumurthi, Ada Gavrilovska |
HPDC | 3 |
| 2019 | Foreword to the special issue for the Workshop on Parallel Programming Models and Systems Software for High-End Computing (P2S2 2017)
Pavan Balaji, Abhinav Vishnu, Yong Chen 0001 |
Parallel Comput. | 2 |
| 2019 | Parallel programming models and systems software for high-end computing (P2S2 2018)
Min Si, Abhinav Vishnu, Yong Chen 0001 |
Parallel Comput. | 2 |
| 2019 | Guest Editor's Introduction: P2S2: SI 2016
Abhinav Vishnu, Pavan Balaji, Yong Chen 0001 |
Parallel Comput. | 1 |
| 2018 | Desh: deep learning for system health prediction of lead times to failure in HPCabstractToday's large-scale supercomputers encounter faults on a daily basis. Exascale systems are likely to experience even higher fault rates due to increased component count and density. Triggering resilience-mitigating techniques remains a challenge due to the absence of well defined failure indicators. System logs consist of unstructured text that obscures essential system health information contained within. In this context, efficient failure prediction via log mining can enable proactive recovery mechanisms to increase reliability. Anwesha Das 0001, Frank Mueller 0001, Charles Siegel, Abhinav Vishnu |
HPDC | 4 |
| 2018 | Using Rule-Based Labels for Weak Supervised Learning: A ChemNet for Transferable Chemical Property PredictionabstractWith access to large datasets, deep neural networks (DNN) have achieved human-level accuracy in image and speech recognition tasks. However, in chemistry data is inherently small and fragmented. In this work, we develop an approach of using rule-based knowledge for training ChemNet, a transferable and generalizable deep neural network for chemical property prediction that learns in a weak-supervised manner from large unlabeled chemical databases. When coupled with transfer learning approaches to predict other smaller datasets for chemical properties that it was not originally trained on, we show that ChemNet's accuracy outperforms contemporary DNN models that were trained using conventional supervised learning. Furthermore, we demonstrate that the ChemNet pre-training approach is equally effective on both CNN (Chemception) and RNN (SMILES2vec) models, indicating that this approach is network architecture agnostic and is effective across multiple data modalities. Our results indicate a pre-trained ChemNet that incorporates chemistry domain knowledge and enables the development of generalizable neural networks for more accurate prediction of novel chemical properties. Garrett B. Goh, Charles Siegel, Abhinav Vishnu, Nathan Oken Hodas |
KDD | 3 |
| 2018 | How Much Chemistry Does a Deep Neural Network Need to Know to Make Accurate Predictions?abstractThe meteoric rise of deep learning models in computer vision research, having achieved human-level accuracy in image recognition tasks is firm evidence of the impact of representation learning of deep neural networks. In the chemistry domain, recent advances have also led to the development of similar CNN models, such as Chemception, that is trained to predict chemical properties using images of molecular drawings. In this work, we investigate the effects of systematically removing and adding localized domain-specific information to the image channels of the training data. By augmenting images with only 3 additional basic information, and without introducing any architectural changes, we demonstrate that an augmented Chemception (AugChemception) outperforms the original model in the prediction of toxicity, activity, and solvation free energy. Then, by altering the information content in the images, and examining the resulting model's performance, we also identify two distinct learning patterns in predicting toxicity/activity as compared to solvation free energy. These patterns suggest that Chemception is learning about its tasks in the manner that is consistent with established knowledge. Thus, our work demonstrates that advanced chemical knowledge is not a pre-requisite for deep learning models to accurately predict complex chemical properties. Garrett B. Goh, Charles Siegel, Abhinav Vishnu, Nathan Oken Hodas, Nathan A. Baker |
WACV | 3 |
| 2018 | ColdRoute: effective routing of cold questions in stack exchange sites
Jiankai Sun, Abhinav Vishnu, Aniket Chakrabarti, Charles Siegel, Srinivasan Parthasarathy 0001 |
Data Min. Knowl. Discov. | 2 |
| 2018 | NUMA-Caffe: NUMA-Aware Deep Learning Neural NetworksabstractConvolution Neural Networks (CNNs), a special subcategory of Deep Learning Neural Networks (DNNs), have become increasingly popular in industry and academia for their powerful capability in pattern classification, image processing, and speech recognition. Recently, they have been widely adopted in High Performance Computing (HPC) environments for solving complex problems related to modeling, runtime prediction, and big data analysis. Current state-of-the-art designs for DNNs on modern multi- and many-core CPU architectures, such as variants of Caffe, have reported promising performance in speedup and scalability, comparable with the GPU implementations. However, modern CPU architectures employ Non-Uniform Memory Access (NUMA) technique to integrate multiple sockets, which incurs unique challenges for designing highly efficient CNN frameworks. Without a careful design, DNN frameworks can easily suffer from long memory latency due to a large number of memory accesses to remote NUMA domains, resulting in poor scalability. To address this challenge, we propose NUMA-aware multi-solver-based CNN design, named NUMA-Caffe , for accelerating deep learning neural networks on multi- and many-core CPU architectures. NUMA-Caffe is independent of DNN topology, does not impact network convergence rates, and provides superior scalability to the existing Caffe variants. Through a thorough empirical study on four contemporary NUMA-based multi- and many-core architectures, our experimental results demonstrate that NUMA-Caffe significantly outperforms the state-of-the-art Caffe designs in terms of both throughput and scalability. Probir Roy, Shuaiwen Song, Sriram Krishnamoorthy, Abhinav Vishnu, Dipanjan Sengupta, Xu Liu 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2017 | A Learning Framework for Control-Oriented Modeling of BuildingsabstractBuildings consume almost 40\% of energy in the US. In order to optimize the operation of buildings, models that describe the relationship between energy consumption and control knobs such as set-points with high predictive capability are required. Data driven modeling techniques have been investigated to a somewhat limited extent for optimizing the operation and control of buildings. In this context, deep learning techniques such as Recurrent Neural Networks (RNNs) hold promise, empowered by advanced computational capabilities and big data opportunities. This paper investigates the use of deep learning for modeling the power consumption of building heating, ventilation and air-conditioning (HVAC) systems. A preliminary analysis of the performance of the methodology for different architectures is conducted. Results show that the proposed methodology outperforms other data driven modeling techniques significantly. Javier Rubio-Herrero, Vikas Chandan, Charles Siegel, Abhinav Vishnu, Draguna L. Vrabie |
ICMLA | 4 |
| 2017 | Enabling scalability-sensitive speculative parallelization for FSM computationsabstractFinite state machines (FSMs) are the backbone of many applications, but are difficult to parallelize due to their inherent dependencies. Speculative FSM parallelization has shown promise on multicore machines with up to eight cores. However, as hardware parallelism grows (e.g., Xeon Phi has up to 288 logical cores), a fundamental question raises: How does the speculative FSM parallelization scale as the number of cores increases? Without answering this question, existing methods for speculative FSM parallelization simply choose to use all available cores, which might not only waste computing resources, but also result in suboptimal performance. Junqiao Qiu, Zhijia Zhao 0001, Bo Wu 0002, Abhinav Vishnu, Shuaiwen Song |
ICS | 4 |
| 2017 | Generating Performance Models for Irregular ApplicationsabstractMany applications have irregular behavior - e.g., input-dependent solvers, irregular memory accesses, or unbiased branches - that cannot be captured using today's automated performance modeling techniques. We describe new hierarchical critical path analyses for the Palm model generation tool. To obtain a good tradeoff between model accuracy, generality, and generation cost, we combine static and dynamic analysis. To create a model's outer structure, we capture tasks along representative MPI critical paths. We create a histogram of critical tasks with parameterized task arguments and instance counts. To model each task, we identify hot instruction-level paths and model each path based on data flow, data locality, and microarchitectural constraints. We describe application models that generate accurate predictions for strong scaling when varying CPU speed, cache and memory speed, microarchitecture, and (with supervision) input data class. Our models' errors are usually below 8%; and always below 13%. Ryan D. Friese, Nathan R. Tallent, Abhinav Vishnu, Darren J. Kerbyson, Adolfy Hoisie |
IPDPS | 3 |
| 2016 | Adaptive neuron apoptosis for accelerating deep learning on large scale systemsabstractWe present novel techniques to accelerate the convergence of Deep Learning algorithms by conducting low overhead removal of redundant neurons - apoptosis of neurons - which do not contribute to model learning, during the training phase itself. We provide in-depth theoretical underpinnings of our heuristics (bounding accuracy loss and handling apoptosis of several neuron types), and present the methods to conduct adaptive neuron apoptosis. Specifically, we are able to improve the training time for several datasets by 2-3x, while reducing the number of parameters by up to 30× (4-5× on average) on datasets such as ImageNet classification. For the Higgs Boson dataset, our implementation improves the accuracy (measured by Area Under Curve (AUC)) for classification from 0.88/1 to 0.94/1, while reducing the number of parameters by 3x in comparison to existing literature. The proposed methods achieve a 2.44x speedup in comparison to the default (no apoptosis) algorithm. Charles Siegel, Jeff Daily, Abhinav Vishnu |
IEEE BigData | 3 |
| 2016 | Fault Tolerant Frequent Pattern MiningabstractFP-Growth algorithm is a Frequent Pattern Mining (FPM) algorithm that has been extensively used to study correlations and patterns in large scale datasets. While several researchers have designed distributed memory FP-Growth algorithms, it is pivotal to consider fault tolerant FP-Growth, which can address the increasing fault rates in large scale systems. In this work, we propose a novel parallel, algorithm-level fault-tolerant FP-Growth algorithm. We leverage algorithmic properties and MPI advanced features to guarantee an O(1) space complexity, achieved by using the dataset memory space itself for checkpointing. We also propose a recovery algorithm that can use in-memory and disk-based checkpointing, though in many cases the recovery can be completed without any disk access, and incurring no memory overhead for checkpointing. We evaluate our FT algorithm on a large scale InfiniBand cluster with several large datasets using up to 2K cores. Our evaluation demonstrates excellent efficiency for checkpointing and recovery in comparison to the disk-based approach. We have also observed 20x average speed-up in comparison to Spark, establishing that a well designed algorithm can easily outperform a solution based on a general fault-tolerant programming model. Sameh Shohdy, Abhinav Vishnu, Gagan Agrawal |
HiPC | 2 |
| 2016 | Accelerating Deep Learning with Shrinkage and RecallabstractDeep Learning is a very powerful machine learning model. Deep Learning trains a large number of parameters for multiple layers and is very slow when data is in large scale and the architecture size is large. Inspired from the shrinking technique used in accelerating computation of Support Vector Machines (SVM) algorithm and screening technique used in LASSO, we propose a shrinking Deep Learning with recall (sDLr) approach to speed up deep learning computation. We experiment shrinking Deep Learning with recall (sDLr) using Deep Neural Network (DNN), Deep Belief Network (DBN) and Convolution Neural Network (CNN) on 4 data sets. Results show that the speedup using shrinking Deep Learning with recall (sDLr) can reach more than 2.0 while still giving competitive classification performance. Shuai Zheng 0002, Abhinav Vishnu, Chris Ding |
ICPADS | 2 |
| 2016 | Fault Tolerant Support Vector MachinesabstractSupport Vector Machines (SVM) is a popular Machine Learning algorithm, which is used for building classifiers and models. Parallel implementations of SVM, which can run on large scale supercomputers, are becoming commonplace. However, these supercomputers -- designed under constraints of data movement -- frequently observe faults in compute devices. Many device faults manifest as permanent process/node failures. In this paper, we present several approaches for designing fault tolerant SVM algorithms. First, we present an in-depth analysis to identify the critical data structures, and build baseline algorithms that simply periodically checkpoint these data structures. Next, we propose a novel algorithm, which requires no inter-node data movement for checkpointing, and only O(n2/p2) recovery time -- a small fraction of the expected O(n3/p) time-complexity of SVM. We implement these algorithms and evaluate them on a large scale cluster. Our evaluation indicates that the overall data movement for checkpointing in the baseline algorithm can be up to 100x the dataset size!, while the proposed novel algorithm is completely communication-free of checkpointing. In addition, it saves up to 20x space, while providing better (by an average of 5.5x speedup on 256 cores) recovery time than the baseline algorithm with different number of checkpoints. The experiments also show that our communication avoiding algorithm outperforms Spark MLLib SVM implementation by an average of 6.4x with 256 cores in the case of failure. Sameh Shohdy, Abhinav Vishnu, Gagan Agrawal |
ICPP | 2 |
| 2016 | Fault Modeling of Extreme Scale Applications Using Machine LearningabstractFaults are commonplace in large scale systems. These systems experience a variety of faults such as transient, permanent and intermittent. Multi-bit faults are typically not corrected by the hardware resulting in an error. This paper attempts to answer an important question: Given a multi-bit fault in main memory, will it result in an application error - and hence a recovery algorithm should be invoked - or can it be safely ignored? We propose an application fault modeling methodology to answer this question. Given a fault signature (a set of attributes comprising of system and application state), we use machine learning to create a model which predicts whether a multi-bit permanent/transient main memory fault will likely result in error. We present the design elements such as the fault injection methodology for covering important data structures, the application and system attributes which should be used for learning the model, the supervised learning algorithms (and potentially ensembles), and important metrics. We use three applications - NWChem, LULESH and SVM - as examples for demonstrating the effectiveness of the proposed fault modeling methodology. Abhinav Vishnu, Huub J. J. Van Dam, Nathan R. Tallent, Darren J. Kerbyson, Adolfy Hoisie |
IPDPS | 1 |
| 2016 | Performance and power for highly parallel systemsabstractThis special issue is the result of an open call for papers initiated after the minisymposium ‘Analysis and Modeling: Techniques and Tools’, conducted at the Society for Industrial and Applied Mathematics (SIAM) Conference on Parallel Processing in Scientific Computing in Savannah, GA, in February 2012. The minisymposium brought together tool developers, performance and power modeling experts, and application analysts to present the state of the art on performance analysis and modeling techniques. This unique combination of expertise is highly needed in a time where complex, hierarchical architectures are the standard for all highly parallel computer systems. Without a good grip on relevant performance limitations, any ptimization attempt is just a shot in the dark. Hence, it is crucial to fully understand the performance properties and bottlenecks that come about with clustered multicore/many-core, multisocket nodes. Another aspect of modern systems is the complicated interplay between power constraints and the need for compute performance, which leads to complicated trade-offs. The challenges ahead are many-fold as systems scale in size. While parallelism is increasing, memory systems, interconnection networks, storage, and uncertainties in programming models all add to the complexities. More rapid realization of energy savings will require significant increases in measurement resolution and optimization techniques. This special issue is focused on how performance and power properties of modern highly parallel systems can be analyzed using state-of-the-art modeling and analysis techniques and real-world applications and tools. G. Hager, J. Treibig, J. Habich, and G. Wellein 1 introduce simple but insightful analytic models for execution performance and energy consumption of multicore CPUs. Automatic dynamic voltage and frequency scaling is leveraged by the “Green Queue” framework presented by J. Peraza, A. Tiwari, M. Laurenzano, L. Carrington, and A. E. Snavely 2 in their article. They show that significant energy savings at low performance loss are in reach if dynamic voltage and frequency scaling is used in an application-aware manner. A. D. Breslow, L. Porter, A. Tiwari, M. Laurenzano, L. Carrington, D. M. Tullsen, and A. E. Snavely 3 investigate the potential of job striping, a technique for co-locating HPC workloads with different characteristics on the same CPU chip, and demonstrate increased throughput and energy efficiency for a mix of typical simulation codes on a production cluster. The problem of how to deal with coarse-grained power measurements is tackled in the paper by H. Servat, G. Llort, J. Giménez, and J. Labarta 4. They present a tool that can derive fine-grained power and performance data for code with quickly alternating phases. The power usage and power variability of workloads on production supercomputers at Los Alamos National Laboratory are studied by S. Pakin, C. Storlie, M. Lang, R. E. Fields, E. E. Romero, C. Idler, S. Michalak, H. Greenberg, J. Loncaric, R. Rheinheimer, G. Grider, and J. Wendelberger in their paper 5. One of their central findings is that real power dissipation under real-world workloads is significantly lower than what the power infrastructure can handle, which opens interesting possibilities for saving cost via power capping. We think that this selection of papers is unique in providing several very different views on the problem of performance and power efficiency on present-day parallel machines from the core to the computing center level. Georg Hager, Darren J. Kerbyson, Abhinav Vishnu, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Performance analysis of data intensive cloud systems based on data management and replication: a survey
Saif Ur Rehman Malik, Samee Ullah Khan, Sam J. Ewen, Nikos Tziritas, Joanna Kolodziej, Albert Y. Zomaya, Sajjad Ahmad Madani, Nasro Min-Allah, Lizhe Wang 0001, Cheng-Zhong Xu 0001, Qutaibah M. Malluhi, Johnatan E. Pecero, Pavan Balaji, Abhinav Vishnu, Rajiv Ranjan 0001, Sherali Zeadally, Hongxiang Li 0001 |
Distributed Parallel Databases | 14 |
| 2016 | Special Issue on Parallel Programming Models and Systems Software for High-End Computing
Pavan Balaji, Abhinav Vishnu, Yong Chen 0001 |
Parallel Comput. | 2 |
| 2016 | Editorial of the Special issue: SI: E2SC
Abhinav Vishnu, Andrés Márquez 0001, Dimitrios S. Nikolopoulos |
Parallel Comput. | 1 |
| 2015 | Large Scale Frequent Pattern Mining Using MPI One-Sided ModelabstractIn this paper, we propose a work-stealing runtime -- Library for Work Stealing LibWS -- using MPI one-sided model for designing scalable FP-Growth -- de facto frequent pattern mining algorithm -- on large scale systems. LibWS provides locality efficient and highly scalable work-stealing techniques for load balancing on a variety of data distributions. We also propose a novel communication algorithm for FP-growth data exchange phase, which reduces the communication complexity from state-of-the-art Θ(p) to Θ(f + p/f), for p processes and f frequent attributed-ids. FP-Growth is implemented using LibWS and evaluated on several work distributions and support counts. An experimental evaluation of the FP-Growth on LibWS using 4096 processes on an InfiniBand Cluster demonstrates excellent efficiency for several work distributions (91% efficiency for Power-law and 93% for Poisson). The proposed distributed FP-Tree merging algorithm provides 38x communication speedup on 4096 cores. Abhinav Vishnu, Khushbu Agarwal |
CLUSTER | 1 |
| 2015 | Fast and Accurate Support Vector Machines on Large Scale SystemsabstractSupport Vector Machines (SVM) is a supervised Machine Learning and Data Mining (MLDM) algorithm, which has become ubiquitous largely due to its high accuracy and obliviousness to dimensionality. The objective of SVM is to find an optimal boundary -- also known as hyperplane -- which separates the samples (examples in a dataset) of different classes by a maximum margin. Usually, very few samples contribute to the definition of the boundary. However, existing parallel algorithms use the entire dataset for finding the boundary, which is sub-optimal for performance reasons. In this paper, we propose a novel distributed memory algorithm to eliminate the samples which do not contribute to the boundary definition in SVM. We propose several heuristics, which range from early (aggressive) to late (conservative) elimination of the samples, such that the overall time for generating the boundary is reduced considerably. In a few cases, a sample may be eliminated (shrunk) pre-emptively -- potentially resulting in an incorrect boundary. We propose a scalable approach to synchronize the necessary data structures such that the proposed algorithm maintains its accuracy. We consider the necessary trade-offs of single/multiple synchronization using in-depth time-space complexity analysis. We implement the proposed algorithm using MPI and compare it with libsvm -- de facto sequential SVM software -- which we enhance with OpenMP for multi-core/many-core parallelism. Our proposed approach shows excellent efficiency using up to 4096 processes on several large datasets such as UCI HIGGS Boson dataset and Offending URL dataset. Abhinav Vishnu, Jeyanthi Narasimhan, Lawrence B. Holder, Darren J. Kerbyson, Adolfy Hoisie |
CLUSTER | 1 |
| 2015 | Diagnosing the causes and severity of one-sided message contentionabstractTwo trends suggest network contention for one-sided messages is poised to become a performance problem that concerns application developers: an increased interest in one-sided programming models and a rising ratio of hardware threads to network injection bandwidth. Often it is difficult to reason about when one-sided tasks decrease or increase network contention. We present effective and portable techniques for diagnosing the causes and severity of one-sided message contention. To detect that a message is affected by contention, we maintain statistics representing instantaneous network resource demand. Using lightweight measurement and modeling, we identify the portion of a message's latency that is due to contention and whether contention occurs at the initiator or target. We attribute these metrics to program statements in their full static and dynamic context. We characterize contention for an important computational chemistry benchmark on InfiniBand, Cray Aries, and IBM Blue Gene/Q interconnects. We pinpoint the sources of contention, estimate their severity, and show that when message delivery time deviates from an ideal model, there are other messages contending for the same network links. With a small change to the benchmark, we reduce contention by 50% and improve total runtime by 20%. Nathan R. Tallent, Abhinav Vishnu, Huub J. J. Van Dam, Jeff Daily, Darren J. Kerbyson, Adolfy Hoisie |
PPoPP | 2 |
| 2015 | A case for application-oblivious energy-efficient MPI runtimeabstractPower has become a major impediment in designing large scale high-end systems. Message Passing Interface (MPI) is the de facto communication interface used as the back-end for designing applications, programming models and runtime for these systems. Slack --- the time spent by an MPI process in a single MPI call---provides a potential for energy and power savings, if an appropriate power reduction technique such as core-idling/Dynamic Voltage and Frequency Scaling (DVFS) can be applied without affecting the application's performance. Existing techniques that exploit slack for power savings assume that application behavior repeats across iterations/executions. However, an increasing use of adaptive and data-dependent workloads combined with system factors (OS noise, congestion) negates this assumption. Akshay Venkatesh, Abhinav Vishnu, Khaled Hamidouche, Nathan R. Tallent, Dhabaleswar K. Panda 0001, Darren J. Kerbyson, Adolfy Hoisie |
SC | 2 |
| 2015 | A work stealing based approach for enabling scalable optimal sequence homology detectionabstractSequence homology detection is central to a number of bioinformatics applications including genome sequencing and protein family characterization. Given millions of sequences, the goal is to identify all pairs of sequences that are highly similar (or “homologous”) on the basis of alignment criteria. While there are optimal alignment algorithms to compute pairwise homology, their deployment for large-scale is currently not feasible; instead, heuristic methods are used at the expense of quality. Here, we present the design and evaluation of a parallel implementation for conducting optimal homology detection on distributed memory supercomputers. Our approach uses a combination of techniques from asynchronous load balancing (viz. work stealing, dynamic task counters), data replication, and exact-matching filters to achieve homology detection at scale. Results for 2.56 M sequences on up to 8K cores show parallel efficiencies of ∼75%–100%, a time-to-solution of 33 s, and a rate of ∼2.0M alignments per second. Jeff Daily, Anantharaman Kalyanaraman, Sriram Krishnamoorthy, Abhinav Vishnu |
J. Parallel Distributed Comput. | 4 |
| 2014 | On the suitability of MPI as a PGAS runtimeabstractPartitioned Global Address Space (PGAS) models are emerging as a popular alternative to MPI models for designing scalable applications. At the same time, MPI remains a ubiquitous communication subsystem due to its standardization, high performance, and availability on leading platforms. In this paper, we explore the suitability of using MPI as a scalable PGAS communication subsystem. We focus on the Remote Memory Access (RMA) communication in PGAS models which typically includes get, put, and atomic memory operations. We perform an in-depth exploration of design alternatives based on MPI. These alternatives include using a semantically-matching interface such as MPI-RMA, as well as not-so-intuitive interfaces such as MPI two-sided with a combination of multi-threading and dynamic process management. With an in-depth exploration of these alternatives and their shortcomings, we propose a novel design which is facilitated by the data-centric view in PGAS models. This design leverages a combination of highly tuned MPI two-sided semantics and an automatic, user-transparent split of MPI communicators to provide asynchronous progress. We implement the asynchronous progress ranks approach and other approaches within the Communication Runtime for Exascale which is a communication subsystem for Global Arrays. Our performance evaluation spans pure communication benchmarks, graph community detection and sparse matrix-vector multiplication kernels, and a computational chemistry application. The utility of our proposed PR-based approach is demonstrated by a 2.17x speedup on 1008 processors over the other MPI-based designs. Jeff Daily, Abhinav Vishnu, Bruce J. Palmer, Huub J. J. Van Dam, Darren J. Kerbyson |
HiPC | 2 |
| 2014 | A performance comparison of current HPC systems: Blue Gene/Q, Cray XE6 and InfiniBand systems
Darren J. Kerbyson, Kevin J. Barker, Abhinav Vishnu, Adolfy Hoisie |
Future Gener. Comput. Syst. | 3 |
| 2013 | Special issue on programming models, systems software, and tools for High-End Computing
Yong Chen 0001, Pavan Balaji, Abhinav Vishnu |
Parallel Comput. | 3 |
| 2013 | A survey on resource allocation in high performance distributed computing systems
Hameed Hussain, Saif Ur Rehman Malik, Abdul Hameed, Samee Ullah Khan, Gage Bickler, Nasro Min-Allah, Muhammad Bilal Qureshi, Yongji Wang 0002, Nasir Ghani, Joanna Kolodziej, Albert Y. Zomaya, Cheng-Zhong Xu 0001, Pavan Balaji, Abhinav Vishnu, Frédéric Pinel, Johnatan E. Pecero, Dzmitry Kliazovich, Pascal Bouvry, Hongxiang Li 0001, Lizhe Wang 0001, Dan Chen 0001, Ammar Rayes |
Parallel Comput. | 15 |
| 2013 | Guest Editors' introduction
Abhinav Vishnu, Pavan Balaji, Yong Chen 0001 |
J. Supercomput. | 1 |
| 2013 | Designing energy efficient communication runtime systems: a view from PGAS models
Abhinav Vishnu, Shuaiwen Song, Andrés Márquez 0001, Kevin J. Barker, Darren J. Kerbyson, Kirk W. Cameron, Pavan Balaji |
J. Supercomput. | 1 |
| 2012 | Global Futures: A Multithreaded Execution Model for Global Arrays-based ApplicationsabstractWe present Global Futures (GF), an execution model extension to Global Arrays, which is based on a PGAS-compatible active message-based paradigm. We describe the design and implementation of Global Futures and illustrate its use in a computational chemistry application benchmark (Hartree-Fock matrix construction using the Self-Consistent Field method). Our results show how we used GF to increase the scalability of the Hartree-Fock matrix build to 6,144 cores of an Infiniband cluster. We also show how GF's multithreaded execution has comparable performance to the traditional process-based SPMD model. Daniel G. Chavarría-Miranda, Sriram Krishnamoorthy, Abhinav Vishnu |
CCGRID | 3 |
| 2012 | Designing scalable PGAS communication subsystems on cray gemini interconnectabstractThe Cray Gemini Interconnect has been recently introduced as a next generation network architecture for building multi-petaflop supercomputers. Cray XE6 systems including LANL Cielo, NERSC Hopper, and the proposed NCSA Blue-Waters, as well as the Cray XK6 ORNL Titan leverage the Gemini Interconnect as their primary Interconnection network. At the same time, programming models such as the Message Passing Interface (MPI) and Partitioned Global Address Space (PGAS) models such as Unified Parallel C (UPC) and Co-Array Fortran (CAF) have become available on these systems. Global Arrays is a popular PGAS model used in a variety of application domains including hydrodynamics, chemistry and visualization. Global Arrays uses Aggregate Remote Memory Copy Interface (ARMCI) as the communication runtime system for Remote Memory Access (RMA) communication. This paper presents a design, implementation and performance evaluation of scalable and high performance communication ARMCI on Cray Gemini. The design space is explored and time-space complexities of communication protocols for one-sided communication primitives such as contiguous and uniformly non-contiguous datatypes, atomic memory operations (AMOs) and memory synchronization is presented. An implementation of the proposed design (referred as ARMCI-Gemini) demonstrates the efficacy on communication primitives, application kernels such as LU decomposition and applications such as Smooth Particle Hydrodynamics (SPH). Abhinav Vishnu, Jeff Daily, Bruce J. Palmer |
HiPC | 1 |
| 2012 | Comparing the Performance of Blue Gene/Q with Leading Cray XE6 and InfiniBand SystemsabstractThree types of systems dominate the current High Performance Computing landscape: the Cray XE6, the IBM Blue Gene, and commodity clusters using InfiniBand. These systems have quite different characteristics making the choice for a particular deployment difficult. The XE6 uses Cray's proprietary Gemini 3-D torus interconnect with two nodes at each network endpoint. The latest IBM Blue Gene/Q uses a single socket integrating processor and communication in a 5-D torus network. InfiniBand provides the flexibility of using nodes from many vendors connected in many possible topologies. The performance characteristics of each vary vastly along with their utilization model. In this work we compare the performance of these three systems using a combination of micro-benchmarks and a set of production applications. We also discuss the causes of variability in performance across the systems and quantify where performance is lost using a combination of measurements and models. Our results show that significant performance can be lost in normal production operation of the Cray XE6 and InfiniBand Clusters in comparison to Blue Gene/Q. Darren J. Kerbyson, Kevin J. Barker, Abhinav Vishnu, Adolfy Hoisie |
ICPADS | 3 |
| 2011 | Energy Templates: Exploiting Application Information to Save EnergyabstractIn this work we consider a novel application centric approach for saving energy on large-scale parallel systems. By using a priori information on the expected application behavior we identify points at which processor-cores will wait for incoming data and thus may be placed in a low power state to save energy. The approach is general and complements many of the existing approaches that rely on saving energy at points of global synchronization. We capture the expected application behavior into an Energy Template whose purpose is to identify when cores are expected to be in an idle state and allow the runtime to use the template information and change the power state of the core. We prototype an Energy Template for a wave front algorithm that contains an complex processing pattern in which cores wait for incoming data before processing local data and whose wait-time varies from phase to phase. The implementation uses PMPI and requires minimal changes to the application code. Using a power instrumented cluster we demonstrate that using an Energy Template for the wave front application lowers the power requirements by 8% when using 216 cores, from the system maximum of 23%, and the energy requirements by 4%. We also show that the wave front's inherent parallel activity will lead to increased savings on larger systems. Darren J. Kerbyson, Abhinav Vishnu, Kevin J. Barker |
CLUSTER | 2 |
| 2011 | Tutorial StatementabstractThis tutorial will provide an overview of the Global Arrays (GA) programming toolkit and describe its capabilities, performance, and the use of GA in high performance computing applications. The tutorial will review basic concepts in onesided communication and Global Address Space languages and will discuss basic setup and elementary communication using GA. More advanced topics, including global counters, non-blocking communication, and sparse data structures will also be discussed. If time permits, the tutorial will also present new/advanced capabilities and ongoing research activities. Bruce J. Palmer, Manojkumar Krishnan, Abhinav Vishnu |
IPDPS | 3 |
| 2011 | Iso-Energy-Efficiency: An Approach to Power-Constrained Parallel ComputationabstractFuture large scale high performance supercomputer systems require high energy efficiency to achieve exaflops computational power and beyond. Despite the need to understand energy efficiency in high-performance systems, there are few techniques to evaluate energy efficiency at scale. In this paper, we propose a system-level iso-energy-efficiency model to analyze, evaluate and predict energy-performance of data intensive parallel applications with various execution patterns running on large scale power-aware clusters. Our analytical model can help users explore the effects of machine and application dependent characteristics on system energy efficiency and isolate efficient ways to scale system parameters (e.g. processor count, CPU power/frequency, workload size and network bandwidth) to balance energy use and performance. We derive our iso-energy-efficiency model and apply it to the NAS Parallel Benchmarks on two power-aware clusters. Our results indicate that the model accurately predicts total system energy consumption within 5% error on average for parallel applications with various execution and communication patterns. We demonstrate effective use of the model for various application contexts and in scalability decision-making. Shuaiwen Song, Chun-Yi Su, Rong Ge 0002, Abhinav Vishnu, Kirk W. Cameron |
IPDPS | 4 |
| 2011 | Noncollective Communicator Creation in MPI
James Dinan, Sriram Krishnamoorthy, Pavan Balaji, Jeff R. Hammond, Manojkumar Krishnan, Vinod Tipparaju, Abhinav Vishnu |
EuroMPI | 7 |
| 2010 | Efficient On-Demand Connection Management Mechanisms with PGAS Models over InfiniBandabstractIn the last decade or so, clusters have observed a tremendous rise in popularity due to the excellent price to performance ratio. A variety of Interconnects have been proposed during this period, with InfiniBand leading the way due to its high performance and open standard. At the same time, multiple programming models have emerged in order to meet the requirements of various applications and their programming models. To support requirements of multiple programming models, InfiniBand provides multiple transport semantics, ranging from unreliable connectionless to reliable connected characteristics. Among them, the reliable connection (RC) semantics is being widely used due to its high performance and support for novel features like Remote Direct Memory Acesss (RDMA), hardware atomics and Network Fault Tolerance. However, the pair wise connection oriented nature of the RC transport semantics limits its scalability and usage at the increasing processor counts. In this paper, we design and implement on-demand connection management approaches in the context of Partitioned Global Address Space (PGAS) programming models, which provided shared memory abstraction and one-sided communication semantics, leading to the development of multiple languages (UPC, X10, Chapel) and libraries (Global Arrays, MPI-RMA). Using Global Arrays as the research vehicle, we implement this approach with Aggregate Remote Memory Copy Interface (ARMCI), the runtime system of Global Arrays. We evaluate our approach, ARMCI-On Demand Connection Management (ARMCI-ODCM) using various micro benchmarks and benchmarks (LU Factorization, Random-Access and Lennard Jones simulation) and application (Subsurface transport over multiple phases (STOMP)). With the performance evaluation for up to 4096 processors, we are able to have a multi-fold reduction in connection memory with a negligible degradation in performance. Using STOMP at 4096 processors, reduces the overall connection memory by 66 times with no performance degradation. To the best of our knowledge, this is the first design, implementation and evaluation of on-demand connection management with InfiniBand using PGAS models. Abhinav Vishnu, Manojkumar Krishnan |
CCGRID | 1 |
| 2010 | Fault-tolerant communication runtime support for data-centric programming modelsabstractThe largest supercomputers in the world today consist of hundreds of thousands of processing cores and many more other hardware components. At such scales, hardware faults are a commonplace, necessitating fault-resilient software systems. While different fault-resilient models are available, most focus on allowing the computational processes to survive faults. On the other hand, we have recently started investigating fault resilience techniques for data-centric programming models such as the partitioned global address space (PGAS) models. The primary difference in data-centric models is the decoupling of computation and data locality. That is, data placement is decoupled from the executing processes, allowing us to view process failure (a physical node hosting a process is dead) separately from data failure (a physical node hosting data is dead). In this paper, we take a first step toward data-centric fault resilience by designing and implementing a fault-resilient, one-sided communication runtime framework using Global Arrays and its communication system, ARMCI. The framework consists of a fault-resilient process manager; low-overhead and network-assisted remote-node fault detection module; non-data-moving collective communication primitives; and failure semantics and err or codes for one-sided communication runtime systems. Our performance evaluation indicates that the framework incurs little overhead compared to state-of-the-art designs and provides a fundamental framework of fault resiliency for PGAS models. Abhinav Vishnu, Huub J. J. Van Dam, Wibe de Jong, Pavan Balaji, Shuaiwen Song |
HiPC | 1 |
| 2009 | An efficient hardware-software approach to network fault tolerance with InfiniBandabstractIn the last decade or so, clusters have observed a tremendous rise in popularity due to excellent price to performance ratio. A variety of Interconnects have been proposed during this period, with InfiniBand leading the way due to its high performance and open standard. Increasing size of the InfiniBand clusters has reduced the mean time between failures of various components of these clusters tremendously. In this paper, we specifically focus on the network component failure and propose a hybrid hardware-software approach to handling network faults. The hybrid approach leverages the user-transparent network fault detection and recovery using Automatic Path Migration (APM), and the software approach is used in the wake of APM failure. Using Global Arrays as the programming model, we implement this approach with Aggregate Remote Memory Copy Interface (ARMCI), the runtime system of Global Arrays. We evaluate our approach using various benchmarks (siosi7, pentane, h2o7 and siosi3) with NWChem, a very popular ab initio quantum chemistry application. Using the proposed approach, the applications run to completion without restart on emulated network faults and acceptable overhead for benchmarks executing for a longer period of time. Abhinav Vishnu, Manojkumar Krishnan, Dhabaleswar K. Panda 0001 |
CLUSTER | 1 |
| 2009 | Topology agnostic hot-spot avoidance with InfiniBandabstractAbstract InfiniBand has become a very popular interconnect due to its advanced features and open standard. Large‐scale InfiniBand clusters are becoming very popular, as reflected by the TOP 500 supercomputer rankings. However, even with popular topologies such as constant bi‐section bandwidth Fat Tree, hot‐spots may occur with InfiniBand due to inappropriate configuration of network paths, presence of other jobs in the network and un‐availability of adaptive routing. In this paper, we present a hot‐spot avoidance layer (HSAL) for InfiniBand, which provides hot‐spot avoidance using path bandwidth estimation and multi‐pathing using LMC mechanism, without taking the network topology into account. We propose an adaptive striping policy with batch‐based striping and sorting approach, for efficient utilization of disjoint network paths. Integration of HSAL with MPI, thede factoprogramming model of clusters, shows promising results with collective communication primitives and MPI applications. Copyright © 2008 John Wiley & Sons, Ltd. Abhinav Vishnu, Matthew J. Koop, Adam Moody, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | High Performance Distributed Lock Management Services using Network-based Remote Atomic OperationsabstractThere has been a massive increase in computing requirements for parallel applications. These parallel applications and supporting cluster services often need to share system-wide resources. The coordination of these applications is typically managed by a distributed lock manager. The performance of the lock manager is extremely critical for application performance. Researchers have shown that the use of two sided communication protocols, like TCP/IP (used by current generation lock managers), can have significant impact on the scalability of distributed lock managers. In addition, existing one sided communication based locking designs support locking in exclusive access mode only and can pose significant scalability limitations on applications that need both shared and exclusive access modes like cooperative/file-system caching. Hence the utility of these existing designs in high performance scenarios can be limited. In this paper, we present a novel protocol, for distributed locking services, utilizing the advanced network-level one-sided atomic operations provided by InfiniBand. Our approach augments existing approaches by eliminating the need for two sided communication protocols in the critical locking path. Further, we also demonstrate that our approach provides significantly higher performance in scenarios needing both shared and exclusive mode access to resources. Our experimental results show 39% improvement in basic locking latencies over traditional send/receive based implementations. Further, we also observe a significant (up to 317% for 16 nodes) improvement over existing RDMA based distributed queuing schemes for shared mode locking scenarios. Sundeep Narravula, A. Marnidala, Abhinav Vishnu, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda 0001 |
CCGRID | 3 |
| 2007 | Hot-Spot Avoidance With Multi-Pathing Over InfiniBand: An MPI PerspectiveabstractLarge scale InfiniBand clusters are becoming increasingly popular, as reflected by the TOP 500 supercomputer rankings. At the same time, fat tree has become a popular interconnection topology for these clusters, since it allows multiple paths to be available in between a pair of nodes. However, even with fat tree, hot-spots may occur in the network depending upon the route configuration between end nodes and communication pattern(s) in the application. To make matters worse, the deterministic routing nature of InfiniBand limits the application from effective use of multiple paths transparently and avoid the hot-spots in the network. Simulation based studies for switches and adapters to implement congestion control have been proposed in the literature. However, these studies have focussed on providing congestion control for the communication path, and not on utilizing multiple paths in the network for hot-spot avoidance. In this paper, we design an MPI functionality, which provides hot-spot avoidance for different communications, without a priori knowledge of the pattern. We leverage LMC (LID mask count) mechanism of InfiniBand to create multiple paths in the network and present the design issues (scheduling policies, selecting number of paths, scalability aspects) of our design. We implement our design and evaluate it with Pallas collective communication and MPI applications. On an InfiniBand cluster with 48 processes, MPI All-to-all personalized shows an improvement of 27%. Our evaluation with NAS parallel benchmarks on 64 processes shows significant improvement in execution time with this functionality. Abhinav Vishnu, Matthew J. Koop, Adam Moody, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda 0001 |
CCGRID | 1 |
| 2007 | High Performance MPI over iWARP: Early ExperiencesabstractModern interconnects and corresponding high performance MPIs have been feeding the surge in the popularity of compute clusters and computing applications. Recently with the introduction of the iWARP (Internet wide area RDMA protocol) standard, RDMA and zero-copy data transfer capabilities have been introduced and standardized for Ethernet networks. While traditional Ethernet networks had largely been limited to the traditional kernel based TCP/IP stacks and hence their limitations, iWARP capabilities of the newer GigE and 10 GigE adapters have broken this barrier and thereby exposing the available potential performance. In order to enable applications to harness the performance benefits of iWARP and to study the quantitative extent of such improvements, we present MPI- iWARP, a high performance MPI implementation over the open fabrics verbs. Our preliminary results with Chelsio T3B adapters show an improvement of up to 37% in bandwidth, 75% in latency and 80% in MPI all reduce as compared to MPICH2 over TCP/IP. To the best of our knowledge, this is the first design, implementation and evaluation of a high performance MPI over the iWARP standard. Sundeep Narravula, Amith R. Mamidala, Abhinav Vishnu, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda 0001 |
ICPP | 3 |
| 2007 | High Performance MPI on IBM 12x InfiniBand ArchitectureabstractInfiniBand is becoming increasingly popular in the area of cluster computing due to its open standard and high performance. I/O interfaces like PCI-express and GX+ are being introduced as next generation technologies to drive InfiniBand with very high throughput. HCAs with throughput of 8x on PCI-express have become available. Recently, support for HCAs with 12x throughput on GX+ has been announced. In this paper, we design a message passing interface (MPI) on IBM 12x dual-port HCAs, which consist of multiple send/recv engines per port. We propose and study the impact of various communication scheduling policies (binding, striping and round robin). Based on this study, we present a new policy, EPC (enhanced point-to-point and collective), which incorporates different kinds of communication patterns; point-to-point (blocking, non-blocking) and collective communication, for data transfer. We implement our design and evaluate it with micro-benchmarks, collective communication and NAS parallel benchmarks. Using EPC on a 12x InfiniBand cluster with one HCA and one port, we can improve the performance by 41% with pingpong latency test and 63-65% with the unidirectional and bi-directional bandwidth tests, compared with the default single-rail MPI implementation. Our evaluation on NAS parallel benchmarks shows an improvement of 7-13% in execution time for integer sort and Fourier transform. Abhinav Vishnu, Brad Benton, Dhabaleswar K. Panda 0001 |
IPDPS | 1 |
| 2007 | Automatic Path Migration over InfiniBand: Early ExperiencesabstractHigh computational power of commodity PCs combined with the emergence of low latency and high bandwidth interconnects has escalated the trends of cluster computing. Clusters with InfiniBand are being deployed, as reflected in the TOP 500 Supercomputer rankings. However, increasing scale of these clusters has reduced the mean time between failures (MTBF) of components. Network component is one such component of clusters, where failure of network interface cards (NICs), cables and/or switches breaks existing path(s) of communication. InfiniBand provides a hardware mechanism, automatic path migration (APM), which allows user transparent detection and recovery from network fault(s), without application restart. In this paper, we design a set of modules; which work together for providing network fault tolerance for user level applications leveraging the APM feature. Our performance evaluation at the MPI layer shows that APM incurs negligible overhead in the absence of faults in the system. In the presence of network faults, APM incurs negligible overhead for reasonably long running applications. Abhinav Vishnu, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda 0001 |
IPDPS | 1 |
| 2007 | On using connection-oriented vs. connection-less transport for performance and scalability of collective and one-sided operations: trade-offs and impactabstractCommunication subsystem plays a pivotal role in achieving scalable performance in clusters. The communication semantics employed are dictated by the programming model used by the application such as MPI, UPC, etc. Out of the gamut of communication primitives, collective and one-sided operations are especially significant and have to be designed harnessing the capabilities and features exposed by the underlying networks. In some cases, there is a direct match between the semantics of the operations and the underlying network primitives. InfiniBand provides two transport modes: (i)Connection-oriented Reliable connection (RC) supporting Memory and Channel semantics and (ii) Connection-less Unreliable Datagram (UD) supporting Channel semantics. Achieving good performance and scalability needs careful analysis and design of communication primitives based on these options. Amith R. Mamidala, Sundeep Narravula, Abhinav Vishnu, Gopalakrishnan Santhanaraman, Dhabaleswar K. Panda 0001 |
PPoPP | 3 |
| 2006 | Scalable systems software - A software based approach for providing network fault tolerance in clusters with uDAPL interface: MPI level design and performance evaluationabstractIn the arena of cluster computing, MPI has emerged as the de facto standard for writing parallel applications. At the same time, introduction of high speed RDMA-enabled interconnects like InfiniBand, Myrinet, Quadrics, RDMA-enabled Ethernet has escalated the trends in cluster computing. Network APIs like uDAPL (user direct access provider library) are being proposed to provide a network-independent interface to different RDMA-enabled interconnects. Clusters with combination(s) of these interconnects are being deployed to leverage their unique features, and network failover in wake of transmission errors. In this paper, we design a network fault tolerant MPI using uDAPL interface, making this design portable for existing and upcoming interconnects. Our design provides failover to available paths, asynchronous recovery of the previous failed paths and recovery from network partitions without application restart. In addition, the design is able to handle network heterogeneity, making it suitable for the current state of the art clusters. We implement our design and evaluate it with micro-benchmarks and applications. Our performance evaluation shows that the proposed design provides significant performance benefits to both homogeneous and heterogeneous clusters. Using a heterogeneous combinations of IBA and Ammasso-GigE, we are able to improve the performance by 10-15% for different NAS parallel benchmarks on 8 times 1 configuration. For simple micro-benchmarks on a homogeneous configuration, we are able to achieve an improvement of 15-20% in throughput. In addition, experiments with simple MPI micro-benchmarks and NAS applications reveal that network fault tolerance modules incur negligible overhead and provide optimal performance in wake of network partitions Abhinav Vishnu, Prachi Gupta, Amith R. Mamidala, Dhabaleswar K. Panda 0001 |
SC | 1 |
| 2005 | Supporting MPI-2 One Sided Communication on Multi-rail InfiniBand Clusters: Design Challenges and Performance Benefits
Abhinav Vishnu, Gopalakrishnan Santhanaraman, Wei Huang 0003, Hyun-Wook Jin, Dhabaleswar K. Panda 0001 |
HiPC | 1 |
| 2004 | Building Multirail InfiniBand Clusters: MPI-Level Design and Performance EvaluationabstractIn the area of cluster computing, InfiniBand is becoming increasingly popular due to its open standard and high performance. However, even with InfiniBand, network bandwidth can still become the performance bottleneck for some of today’s most demanding applications. In this paper, we study the problem of how to overcome the bandwidth bottleneck by using multirail networks. We present different ways of setting up multirail networks with InfiniBand and propose a unified MPI design that can support all these approaches. We have also discussed various important design issues and provided in-depth discussions of different policies of using multirail networks, including an adaptive striping scheme that can dynamically change the striping parameters based on current system condition. We have implemented our design and evaluated it using both microbenchmarks and applications. Our performance results show that multirail networks can significant improve MPI communication performance. With a two rail InfiniBand cluster, we have achieved almost twice the bandwidth and half the latency for large messages compared with the original MPI. At the application level, the multirail MPI can significantly reduce communication time as well as running time depending on the communication pattern. We have also shown that the adaptive striping scheme can achieve excellent performance without a priori knowledge of the bandwidth of each rail. Jiuxing Liu, Abhinav Vishnu, Dhabaleswar K. Panda 0001 |
SC | 2 |