EDBT 2026 Demo / reviewers in the wild / expert
Sathish S. Vadhiyar
dblp:v/SathishSVadhiyar · also Sathish Vadhiyar
· DBLP profile ↗
50ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-5476-8328ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Scalable System for Visual Analysis of Ocean DataabstractAbstract Oceanographers rely on visual analysis to interpret model simulations, identify events and phenomena, and track dynamic ocean processes. The ever increasing resolution and complexity of ocean data due to its dynamic nature and multivariate relationships demands a scalable and adaptable visualization tool for interactive exploration. We introduce pyParaOcean, a scalable and interactive visualization system designed specifically for ocean data analysis. pyParaOcean offers specialized modules for common oceanographic analysis tasks, including eddy identification and salinity movement tracking. These modules seamlessly integrate with ParaView as filters, ensuring a user‐friendly and easy‐to‐use system while leveraging the parallelization capabilities of ParaView and a plethora of inbuilt general‐purpose visualization functionalities. The creation of an auxiliary dataset stored as a Cinema database helps address I/O and network bandwidth bottlenecks while supporting the generation of quick overview visualizations. We present a case study on the Bay of Bengal to demonstrate the utility of the system and scaling studies to evaluate the efficiency of the system. Toshit Jain, Upkar Singh, Varun Singh, Vijay Kumar Boda, Ingrid Hotz, Sathish S. Vadhiyar, P. N. Vinayachandran, Vijay Natarajan |
Comput. Graph. Forum | 6 |
| 2023 | Strategies for Fast I/O Throughput in Large-Scale Climate Modeling ApplicationsabstractLarge-scale HPC applications are highly data-intensive with significant times spent in I/O operations. Many large-scale scientific applications do not adequately optimize the I/O operations, leading to overall poor performance. In this work, we have developed two main strategies for providing fast I/O throughput for an important climate modeling application, namely, Regional Ocean Modeling System (ROMS) that uses NetCDF for I/O operations. The strategies include load balancing the I/O operations and selective writing of data. We have also implemented file striping to improve I/O performance. Our experiments with up to 1440 processor cores and 5 days of simulations showed that our load balancing strategy resulted in about 27 % decrease in execution times over the default executions, our selective writing strategy resulted in a further decrease of about 30 % and the optimized file striping resulted in a further decrease of about 12 % in execution times. All the strategies combined together improved the overall performance of the application by about 70 %. Koushik Sen, Sathish S. Vadhiyar, P. N. Vinayachandran |
HiPC | 2 |
| 2022 | Dynamic Strategies for High Performance Training of Knowledge Graph EmbeddingsabstractKnowledge graph embeddings (KGEs) are the low dimensional representations of entities and relations between the entities. They can be used for various downstream tasks such as triple classification, link prediction, knowledge base completion, etc. Training these embeddings for a large dataset takes a huge amount of time. This work proposes strategies to make the training of KGEs faster in a distributed memory parallel environment. The first strategy is to choose between either an all-gather or an all-reduce operation based on the sparsity of the gradient matrix. The second strategy focuses on selecting those gradient vectors which significantly contribute to the reduction in the loss. The third strategy employs gradient quantization to reduce the number of bits to be communicated. The fourth strategy proposes to split the knowledge graph triples based on relations so that inter-node communication for the gradient matrix corresponding to the relation embedding matrix is eliminated. The fifth and last strategy is to select the negative triple which the model finds difficult to classify. Anwesh Panda, Sathish S. Vadhiyar |
ICPP | 2 |
| 2022 | Scalable multi-node multi-GPU Louvain community detection algorithm for heterogeneous architecturesabstractSummary Community detection is an important problem that is widely applied for finding cluster patterns in brain, social, biological, and many other kinds of networks. In this work, we have developed a multi‐node multi‐GPU Louvain community detection algorithm, simultaneously harnessing the CPU and GPU cores of the devices. The algorithm partitions a given graph across multiple nodes and devices in the nodes and performs independent computations of Louvain algorithm on the parts on the devices. The independently formed communities in the devices are refined by identification of doubtful vertices and migrating them to the other processors. The communities are merged using a hierarchical merging algorithm that ensures that at any point the merged component can be accommodated within a processor. Our experiments show that our algorithm is highly scalable with increasing number of devices and provides large‐scale performance for BigData graphs. Anwesha Bhowmick, Sathish S. Vadhiyar, Varun PV |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Pipelined Preconditioned Conjugate Gradient Methods for real and complex linear systems for distributed memory architectures
Manasi Tiwari, Sathish S. Vadhiyar |
J. Parallel Distributed Comput. | 2 |
| 2021 | Pipelined Preconditioned s-step Conjugate Gradient Methods for Distributed Memory SystemsabstractPreconditioned Conjugate Gradient (PCG) method is a widely used iterative method for solving large linear systems of equations. Pipelined variants of PCG present independent computations in the PCG method and overlap these computations with non-blocking allreduces. We have developed a novel pipelined PCG algorithm called PIPE-sCG (Pipelined s-step Conjugate Gradient) that provides a large overlap of global communication and computations at higher number of cores in distributed memory CPU systems. Our method achieves this overlap by introducing new recurrence computations. We have also developed a preconditioned version of PIPE-sCG. The advantages of our methods are that they do not introduce any extra preconditioner or sparse matrix vector product kernels in order to provide the overlap and can work with preconditioned, unpreconditioned and natural norms of the residual, as opposed to the state-of-the-art methods. We compare our method with other pipelined CG methods for Poisson problems and demonstrate that our method gives the least runtimes. Our method gives up to 2.9x speedup over PCG method, 2.15x speedup over PIPECG method and 1.2x speedup over PIPECG-OATI method at large number of cores. Manasi Tiwari, Sathish S. Vadhiyar |
CLUSTER | 2 |
| 2020 | Fast Scalable Approximate Nearest Neighbor Search for High-dimensional DataabstractK-Nearest Neighbor (k-NN) search is one of the most commonly used approaches for similarity search. It finds extensive applications in machine learning and data mining. This era of big data warrants efficiently scaling k-NN search algorithms for billion-scale datasets with high dimensionality. In this paper, we propose a solution towards this end where we use vantage point trees for partitioning the dataset across multiple processes and exploit an existing graph-based sequential approximate k-NN search algorithm called HNSW (Hierarchical Navigable Small World) for searching locally within a process. Our hybrid MPI-OpenMP solution employs techniques including exploiting MPI one-sided communication for reducing communication times and partition replication for better load balancing across processes. We demonstrate computation of k-NN for 10,000 queries in the order of seconds using our approach on ~8000 cores on a dataset with billion points in an 128-dimensional space. We also show 10X speedup over a completely k-d tree-based solution for the same dataset, thus demonstrating better suitability of our solution for high dimensional datasets. Our solution shows almost linear strong scaling. K. G. Renga Bashyam, Sathish S. Vadhiyar |
CLUSTER | 2 |
| 2020 | Pipelined Preconditioned Conjugate Gradient Methods for Distributed Memory SystemsabstractPreconditioned Conjugate Gradient (PCG) method has been one of the widely used methods for solving linear systems of equations for sparse problems. Pipelined PCG (PIPECG) attempts to eliminate the dependencies in the computations in the PCG algorithm and overlap non-dependent computations by reorganizing the traditional PCG code and using non-blocking allreduces. We have developed a novel pipelined PCG algorithm called PIPECG-OATI (One Allreduce per Two Iterations) that provides large overlap of global communication and computations at higher number of cores in distributed memory CPU systems. Our method achieves this overlapping by using iteration combination and by introducing new non-recurrence computations. We compare our method with other pipelined CG methods on a variety of problems and demonstrate that our method always gives the least runtimes. Our method gives up to 3x speedup over PCG method and 1.73x speedup over PIPECG method at large number of cores. Manasi Tiwari, Sathish S. Vadhiyar |
HiPC | 2 |
| 2019 | HyDetect: A Hybrid CPU-GPU Algorithm for Community DetectionabstractCommunity detection is an important problem that is widely applied for finding cluster patterns in brain, social, biological and many other kinds of networks. In this work, we propose a divide-and-conquer community detection algorithm for hybrid CPU-GPU systems. The graph representing a network is partitioned among the CPU and GPU devices of a node, and independent community detection using Louvain's algorithm is carried out in both the parts. The communities are iteratively refined by a novel strategy for identifying and moving "doubtful" vertices between the devices. The resulting accuracy is found comparable with the single device parallel Louvain algorithms. Our hybrid algorithm helped to explore large graphs that cannot be accommodated in a single device. By harnessing the power of GPUs, our hybrid algorithm is able to provide 42-73% smaller execution times over state-of-art CPU-only algorithms. Anwesha Bhowmik, Sathish S. Vadhiyar |
HiPC | 2 |
| 2019 | Fast and Accurate Learning of Knowledge Graph Embeddings at ScaleabstractKnowledge Graph Embedding (KGE) is used to represent the entities and relations of a KG in a low dimensional vector space. KGE can then be used in a downstream task such as entity classification, link prediction and knowledge base completion. Training on large KG datasets takes a considerable amount of time. This work proposes three strategies which lead to faster training in distributed setting. The first strategy is a reduced communication approach which decreases the All-Gather size by sparsifying the Sparse Gradient Matrix (SGM). The second strategy is a variable margin approach that takes advantage of reduced communication for lower margins but retains the accuracy as obtained by the best fixed margin. The third strategy is called DistAdam which is a distributed version of the popular Adam optimization algorithm. Combining the three strategies results in reduction of training time for the FB250K dataset from twenty-seven hours on one processing node to under one hour on thirty-two nodes with each node consisting of twenty-four cores. Sathish S. Vadhiyar |
HiPC | 2 |
| 2019 | HyPar: A divide-and-conquer model for hybrid CPU-GPU graph processing
Rintu Panja, Sathish S. Vadhiyar |
J. Parallel Distributed Comput. | 2 |
| 2018 | MND-MST: A Multi-Node Multi-Device Parallel Boruvka's MST AlgorithmabstractEfficient processing of large-scale graph applications on heterogeneous CPU-GPU systems require effectively harnessing the combined power of both the CPU and GPU devices. Finding minimum spanning tree (MST) is an important graph application and is used in different domains. When applying MST algorithms for large-scale graphs across multiple nodes (or machines), the existing approaches use BSP (bulk synchronous parallel) model involving large-scale communications. In this paper, we propose a multi-node multi-device algorithm for MST, MND-MST, that uses a divide-and-conquer approach by partitioning the input graph across multiple nodes and devices and performing independent Boruvka's MST computations on the devices. The results from the different nodes are merged using a novel hybrid merging algorithm that ensures that the combined results on a node never exceeds it memory capacity. The algorithm also simultaneously harnesses both CPU and GPU devices. In our experiments, we show that our proposed algorithm shows 24-88% performance improvements over an existing BSP approach. We also show that the algorithm exhibits almost linear scalability, and the use of GPUs result in upto 23% improvement in performance over multi-node CPU-only performance. Rintu Panja, Sathish S. Vadhiyar |
ICPP | 2 |
| 2018 | Metascheduling of HPC Jobs in Day-Ahead Electricity MarketsabstractHigh performance grid computing is a key enabler of large scale collaborative computational science. With the promise of exascale computing, high performance grid systems are expected to incur electricity bills that grow super-linearly overtime. In order to achieve cost effectiveness in these systems, it is essential for the scheduling algorithms to exploit electricity price variations, both in space and time, that are prevalent in the dynamic electricity price markets. In this paper, we present a metascheduling algorithm to optimize the placement of jobs in a compute grid which consumes electricity from the day-ahead wholesale market. We formulate the scheduling problem as a Minimum Cost Maximum Flow problem and leverage queue waiting time and electricity price predictions to accurately estimate the cost of job execution at a system. Using trace based simulation with real and synthetic workload traces, and real electricity price data sets, we demonstrate our approach on two currently operational grids, XSEDE and NorduGrid. Our experimental setup collectively constitute more than 433K processors spread across 58 compute systems in 17 geographically distributed locations. Experiments show that our approach simultaneously optimizes the total electricity cost and the average response time of the grid, without being unfair to users of the local batch systems. Prakash Murali, Sathish S. Vadhiyar |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Adaptive Hybrid Queue Configuration for Supercomputer SystemsabstractSupercomputers have batch queues to which parallel jobs with specific requirements are submitted. Commercial schedulers come with various configurable parameters for the queues which can be adjusted based on the requirements of the system. The employed configuration affects both system utilization and job response times. Often times, choosing an optimal configuration with good performance is not straightforward and requires good knowledge of the system behavior to various kinds of workloads. In this paper, we propose a dynamic scheme for setting queue configurations, namely, the number of queues, partitioning of the processor space and the mapping of the queues to the processor partitions, and the processor size and execution time limits corresponding to the queues based on the historical workload patterns. We use a novel non-linear programming formulation for partitioning and mapping of nodes to the queues for homogeneous HPC systems. We also propose a novel hybrid partitioned-nonpartitioned scheme for allocating processors to the jobs submitted to the queues. Our simulation results for a supercomputer system with 35,000+ CPU cores show that our hybrid scheme gives up to 74% reduction in queue waiting times and up to 12% higher utilizations than static queue configurations. Vineetha Kondameedi, Sathish S. Vadhiyar |
CCGrid | 2 |
| 2017 | Asynchronous and synchronous models of executions on Intel® Xeon Phi™ coprocessor systems for high performance of long wave radiation calculations in atmosphere models
Amlesh Kashyap, Sathish S. Vadhiyar, Ravi S. Nanjundiah, P. N. Vinayachandran |
J. Parallel Distributed Comput. | 2 |
| 2016 | High Performance Horizontal Diffusion Calculations in Ocean Models on Intel® Xeon Phi™ Coprocessor SystemsabstractAccelerators and co-processors are widely prevalent and have been used to provide high performance for many scientific applications. Intel® Xeon Phi™ coprocessors have been gaining ground to provide speedups for advanced scientific applications. However, the use and demonstration of these coprocessors for climate modeling are limited. In this work, we have developed a comprehensive set of novel techniques for efficient use of Intel Xeon Phi coprocessors for ocean modeling. In particular, we focus on one of the most time consuming routines, namely, horizontal diffusion in tracers (hdifft). Our techniques include explicit and implicit fusion for data locality and vectorization, choice of coarse-grained over fine-grained parallelism, offloading hdifft function for asynchronous and simultaneous executions on CPU and accelerator cores, and effective data management to minimize CPU-accelerator data transfer overheads. Our comprehensive set of techniques has resulted in about 17-23% improvement in simulation throughput of the entire ocean model code. Our optimization strategies exhibit good scaling with the use of more Intel Xeon Phi processors. Among our optimization techniques, the use of our novel look-ahead asynchronous execution strategy on Intel Xeon Phis resulted in maximum benefit yielding about 36% improvement in performance. T. M. Aketh, Sathish S. Vadhiyar, P. N. Vinayachandran, Ravi S. Nanjundiah |
HiPC | 2 |
| 2016 | Qespera: an adaptive framework for prediction of queue waiting times in supercomputer systemsabstractSummary Production parallel systems are space‐shared, and resource allocation on such systems is usually performed using a batch queue scheduler. Jobs submitted to the batch queue experience a variable delay before the requested resources are granted. Predicting this delay can assist users in planning experiment time‐frames and choosing sites with less turnaround times and can also help meta‐schedulers make scheduling decisions. In this paper, we present an integrated adaptive framework, Qespera, for prediction of queue waiting times on parallel systems. We propose a novel algorithm based on spatial clustering for predictions using history of job submissions and executions. The framework uses adaptive set of strategies for choosing either distributions or summary of features to represent the system state and to compare with history jobs, varying the weights associated with the features for each job prediction, and selecting a particular algorithm dynamically for performing the prediction depending on the characteristics of the target and history jobs. Our experiments with real workload traces from different production systems demonstrate up to 22% reduction in average absolute error and up to 56% reduction in percentage prediction error over existing techniques. We also report prediction errors of less than 1 h for a majority of the jobs. Copyright © 2015 John Wiley & Sons, Ltd. Prakash Murali, Sathish S. Vadhiyar |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Metascheduling of HPC Jobs in Day-Ahead Electricity MarketsabstractHigh performance grid computing is a key enabler of large scale collaborative computational science. With the promise of exascale computing, high performance grid systems are expected to incur electricity bills that grow super-linearly over time. In order to achieve cost effectiveness in these systems, it is essential for the scheduling algorithms to exploit electricity price variations, both in space and time, that are prevalent in the dynamic electricity price markets. In this paper, we present a metascheduling algorithm to optimize the placement of jobs in a compute grid which consumes electricity from the day-ahead wholesale market. We formulate the scheduling problem as a Minimum Cost Maximum Flow problem and leverage queue waiting time and electricity price predictions to accurately estimate the cost of job execution at a system. Using trace based simulation with real and synthetic workload traces, and real electricity price data sets, we demonstrate our approach on two currently operational grids, XSEDE and NorduGrid. Our experimental setup collectively constitute more than 433K processors spread across 58 compute systems in 17 geographically distributed locations. Experiments show that our approach simultaneously optimizes the total electricity cost and the average response time of the grid, without being unfair to users of the local batch systems. Prakash Murali, Sathish S. Vadhiyar |
HiPC | 2 |
| 2015 | Matching Application Signatures for Performance Predictions Using a Single ExecutionabstractPerformance predictions for large problem sizes and processors using limited small scale runs are useful for a variety of purposes including scalability projections, and help in minimizing the time taken for constructing training data for building performance models. In this paper, we present a prediction framework that matches execution signatures for performance predictions of HPC applications using a single small scale application execution. Our framework extracts execution signatures of applications and performs automatic phase identification of different application phases. Application signatures of the different phases are matched with the execution profiles of reference kernels stored in a kernel database. The performance of the reference kernels are then used to predict the performance of the application phases. For phases that do not match significantly, our framework performs static analysis of loops and functions in the application to provide prediction ranges. We demonstrate this integrated set of techniques in our framework with three large scale applications, including GTC, a Particle-in-Cell code for turbulence simulation, Sweep3d, a 3D neutron transport application and SMG2000, a multigrid solver. We show that our prediction ranges are accurate in most cases. Anirudh Jayakumar, Prakash Murali, Sathish S. Vadhiyar |
IPDPS | 3 |
| 2015 | Fault Tolerance on Large Scale Systems using Adaptive Process ReplicationabstractExascale systems of the future are predicted to have mean time between failures (MTBF) of less than one hour. At such low MTBFs, employing periodic checkpointing alone will result in low efficiency because of the high number of application failures resulting in large amount of lost work due to rollbacks. In such scenarios, it is highly necessary to have proactive fault tolerance mechanisms that can help avoid significant number of failures. In this work, we have developed a mechanism for proactive fault tolerance using partial replication of a set of application processes. Our fault tolerance framework adaptively changes the set of replicated processes periodically based on failure predictions to avoid failures. We have developed an MPI prototype implementation, PAREP-MPI that allows changing the replica set. We have shown that our strategy involving adaptive process replication significantly outperforms existing mechanisms providing up to 20 percent improvement in application efficiency even for exascale systems. Cijo George, Sathish S. Vadhiyar |
IEEE Trans. Computers | 2 |
| 2014 | Prediction of Queue Waiting Times for Metascheduling on Parallel Batch Systems
Rajath Kumar, Sathish S. Vadhiyar |
JSSPP | 2 |
| 2013 | GPU-enabled efficient executions of radiation calculations in climate modelingabstractIn this paper, we discuss the acceleration of a climate model known as the Community Earth System Model (CESM). The use of Graphics Processor Units (GPUs) to accelerate scientific applications that are computationally intensive is well known. This work attempts to extract the performance of GPUs to enable faster execution of CESM and obtain better model throughput. We focus on two major routines that consume the largest amount of time namely, radabs and radcswmx, which compute parameters related to the long wave (infra-red) and short wave (visible and ultra-violet) radiations respectively. We propose a novel asynchronous execution strategy in which the results computed by the GPU for the current time step are used by the CPU in the subsequent time step. Such a technique effectively hides computational effort on the GPU. By exploiting the parallelism offered by the GPU and using asynchronous executions on the CPU and GPU, we obtain a speed-up of about 26× for the routine radabs and about 5.6× for routine radcswmx. Sai Kiran Korwar, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
HiPC | 2 |
| 2013 | Efficient homology computations on multicore and manycore systemsabstractHomology computations form an important step in topological data analysis that helps to identify connected components, holes, and voids in multi-dimensional data. Our work focuses on algorithms for homology computations of large simplicial complexes on multicore machines and on GPUs. This paper presents two parallel algorithms to compute homology. A core component of both algorithms is the algebraic reduction of a cell with respect to one of its faces while preserving the homology of the original simplicial complex. The first algorithm is a parallel version of an existing sequential implementation using OpenMP. The algorithm processes and reduces cells within each partition of the complex in parallel while minimizing sequential reductions on the partition boundaries. Cache misses are reduced by ensuring data locality for data in the same partition. We observe a linear speedup on algebraic reductions and an overall speedup of up to 4.9× with 16 cores over sequential reductions. The second algorithm is based on a novel approach for homology computations on manycore/GPU architectures. This GPU algorithm is memory efficient and capable of extremely fast computation of homology for simplicial complexes with millions of simplices. We observe up to 40× speedup in runtime over sequential reductions and up to 4.5× speedup over REDHOM library, which includes the sequential algebraic reductions together with other advanced homology engines supported in the software. N. Anurag Murty, Vijay Natarajan, Sathish S. Vadhiyar |
HiPC | 3 |
| 2013 | A Diffusion-Based Processor Reallocation Strategy for Tracking Multiple Dynamically Varying Weather PhenomenaabstractMany meteorological phenomena occur at different locations simultaneously. These phenomena vary temporally and spatially. It is essential to track these multiple phenomena for accurate weather prediction. Efficient analysis require high-resolution simulations which can be conducted by introducing finer resolution nested simulations, nests at the locations of these phenomena. Simultaneous tracking of these multiple weather phenomena requires simultaneous execution of the nests on different subsets of the maximum number of processors for the main weather simulation. Dynamic variation in the number of these nests require efficient processor reallocation strategies. In this paper, we have developed strategies for efficient partitioning and repartitioning of the nests among the processors. As a case study, we consider an application of tracking multiple organized cloud clusters in tropical weather systems. We first present a parallel data analysis algorithm to detect such clouds. We have developed a tree-based hierarchical diffusion method which reallocates processors for the nests such that the redistribution cost is less. We achieve this by a novel tree reorganization approach. We show that our approach exhibits up to 25% lower redistribution cost and 53% lesser hop-bytes than the processor reallocation strategy that does not consider the existing processor allocation. Preeti Malakar, Vijay Natarajan, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
ICPP | 3 |
| 2013 | G-Charm: an adaptive runtime system for message-driven parallel applications on hybrid systemsabstractThe effective use of GPUs for accelerating applications depends on a number of factors including effective asynchronous use of heterogeneous resources, reducing memory transfer between CPU and GPU, increasing occupancy of GPU kernels, overlapping data transfers with computations, reducing GPU idling and kernel optimizations. Overcoming these challenges require considerable effort on the part of the application developers and most optimization strategies are often proposed and tuned specifically for individual applications. In this paper, we present G-Charm, a generic framework with an adaptive runtime system for efficient execution of message-driven parallel applications on hybrid systems. The framework is based on Charm++, a message-driven programming environment and runtime for parallel applications. The techniques in our framework include dynamic scheduling of work on CPU and GPU cores, maximizing reuse of data present in GPU memory, data management in GPU memory, and combining multiple kernels. We have presented results using our framework on Tesla S1070 and Fermi C2070 systems using three classes of applications: a highly regular and parallel 2D Jacobi solver, a regular dense matrix Cholesky factorization representing linear algebra computations with dependencies among parallel computations and highly irregular molecular dynamics simulations. With our generic framework, we obtain 1.5 to 15 times improvement over previous GPU-based implementation of Charm++. We also obtain about 14\% improvement over an implementation of Cholesky factorization with a static work-distribution scheme. R. Vasudevan, Sathish S. Vadhiyar, Laxmikant V. Kalé |
ICS | 2 |
| 2013 | Efficient asynchronous executions of AMR computations and visualization on a GPU system
Hari K. Raghavan, Sathish S. Vadhiyar |
J. Parallel Distributed Comput. | 2 |
| 2012 | Identifying Quick Starters: Towards an Integrated Framework for Efficient Predictions of Queue Waiting Times of Batch Parallel Jobs
Rajath Kumar, Sathish S. Vadhiyar |
JSSPP | 2 |
| 2012 | A divide and conquer strategy for scaling weather simulations with multiple regions of interestabstractAccurate and timely prediction of weather phenomena, such as hurricanes and flash floods, require high-fidelity compute intensive simulations of multiple finer regions of interest within a coarse simulation domain. Current weather applications execute these nested simulations sequentially using all the available processors, which is sub-optimal due to their sub-linear scalability. In this work, we present a strategy for parallel execution of multiple nested domain simulations based on partitioning the 2-D processor grid into disjoint rectangular regions associated with each domain. We propose a novel combination of performance prediction, processor allocation methods and topology-aware mapping of the regions on torus interconnects. Experiments on IBM Blue Gene systems using WRF show that the proposed strategies result in performance improvement of up to 33% with topology-oblivious mapping and up to additional 7% with topology-aware mapping over the default sequential strategy. Preeti Malakar, Thomas George, Sameer Kumar 0001, Rashmi Mittal, Vijay Natarajan, Yogish Sabharwal, Vaibhav Saxena, Sathish S. Vadhiyar |
SC | 8 |
| 2012 | Large improvements in application throughput of long-running multi-component applications using batch gridsabstractSUMMARY Computational grids with multiple batch systems (batch grids) can be powerful infrastructures for executing long‐running multi‐component parallel applications. In this paper, we evaluate the potential improvements in throughput of long‐running multi‐component applications when the different components of the applications are executed on multiple batch systems of batch grids. We compare the multiple batch executions with executions of the components on a single batch system without increasing the number of processors used for executions. We perform our analysis with a foremost long‐running multi‐component application for climate modeling, the Community Climate System Model (CCSM). We have built a robust simulator that models the characteristics of both the multi‐component application and the batch systems. By conducting large number of simulations with different workload characteristics and queuing policies of the systems, processor allocations to components of the application, distributions of the components to the batch systems and inter‐cluster bandwidths, we show that multiple batch executions lead to 55% average increase in throughput over single batch executions for long‐running CCSM. We also conducted real experiments with a practical middleware infrastructure and showed that multi‐site executions lead to effective utilization of batch systems for executions of CCSM and give higher simulation throughput than single‐site executions. Copyright © 2011 John Wiley & Sons, Ltd. Sivagama Sundari Murugavel, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Adaptive Executions of Multi-Physics Coupled Applications on Batch Grids
Sivagama Sundari Murugavel, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
J. Grid Comput. | 2 |
| 2011 | Strategies for Rescheduling Tightly-Coupled Parallel Applications in Multi-Cluster Grids
H. A. Sanjay 0001, Sathish S. Vadhiyar |
J. Grid Comput. | 2 |
| 2010 | Morco: middleware framework for long-running multi-component applications on batch gridsabstractWhile computational grids with multiple batch systems (batch grids) have been used for efficient executions of loosely-coupled and workflow-based parallel applications, they can also be powerful infrastructures for executing long-running multi-component parallel applications. In this paper, we have constructed a generic middleware framework for executing long-running multi-component applications with execution times much greater than execution time limits of batch queues. Our framework coordinates the distribution, execution, migration and restart of the components of the application on the multiple queues, where the component jobs of the different queues can have different queue waiting and startup times. We have used our framework with a foremost long-running multi-component application for climate modeling, the Community Climate System Model (CCSM). We have performed real multiple-site CCSM runs for 6.5 days of wallclock time spanning three sites with four queues and emulated external workloads. Our experiments indicate that multi-site executions can lead to good throughput of application execution. Sundari M. Sivagama, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
HPDC | 2 |
| 2010 | An Adaptive Framework for Simulation and Online Remote Visualization of Critical Climate Applications in Resource-constrained EnvironmentsabstractCritical climate applications like cyclone tracking and earthquake modeling require high-performance simulations and online visualization simultaneously performed with the simulations for timely analysis. Remote visualization of critical climate events enables joint analysis by geographically distributed climate science community. However, resource constraints including limited storage and slow networks can limit the effectiveness of such online visualization. In this work, we have developed an adaptive framework that simultaneously performs numerical simulations and online remote visualization of critical climate applications in resource-constrained environments. Our framework considers both application and resource dynamics to adapt various application and resource parameters including simulation resolutions, resource configurations and amount of data for visualization. We have developed two algorithms for processor allocation for simulations and the frequency of data for visualization. We show that our optimization method is able to provide about 30% higher simulation rate and consumes about 25-50% lesser storage space than the greedy approach. Preeti Malakar, Vijay Natarajan, Sathish S. Vadhiyar |
SC | 3 |
| 2010 | Grids with multiple batch systems for performance enhancement of multi-component and parameter sweep parallel applications
Sundari M. Sivagama, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
Future Gener. Comput. Syst. | 2 |
| 2009 | Phylogenetic Predictions on GridsabstractA phylogenetic or evolutionary tree is constructed from a set of species or DNA sequences and depicts the relatedness between the sequences. Predictions of future sequences in a phylogenetic tree are important for a variety of applications including drug discovery, pharmaceutical research and disease control. In this work, we predict future DNA sequences in a phylogenetic tree using cellular automata. Cellular automata are used for modeling neighbor-dependent mutations from an ancestor to a progeny in a branch of the phylogenetic tree. Since the number of possible ways of transformations from an ancestor to a progeny is huge, we use computational grids and middleware techniques to explore the large number of cellular automata rules used for the mutations. We use the popular and recurring neighbor-based transitions or mutations to predict the progeny sequences in the phylogenetic tree. We performed predictions for three types of sequences, namely, triose phosphate isomerase, pyruvate kinase, and polyketide synthase sequences, by obtaining cellular automata rules on a grid consisting of 29 machines in 4 clusters located in 4 countries, and compared the predictions of the sequences using our method with predictions by random methods. We found that in all cases, our method gave about 40% better predictions than the random methods. Priyank Raj Katariya, Sathish S. Vadhiyar |
eScience | 2 |
| 2009 | A strategy for scheduling tightly coupled parallel applications on clustersabstractAbstract Although various strategies have been developed for scheduling parallel applications with independent tasks, very little work exists for scheduling tightly coupled parallel applications on cluster environments. In this paper, we compare four different strategies based on performance models of tightly coupled parallel applications for scheduling the applications on clusters. In addition to algorithms based on existing popular optimization techniques, we also propose a new algorithm called Box Elimination that searches the space of performance model parameters to determine the best schedule of machines. By means of real and simulation experiments, we evaluated the algorithms on single cluster and multi‐cluster setups. We show that our Box Elimination algorithm generates up to 80% more efficient schedules than other algorithms. We also show that the execution times of the schedules produced by our algorithm are more robust against the performance modeling errors. Copyright © 2009 John Wiley & Sons, Ltd. H. A. Sanjay 0001, Sathish S. Vadhiyar |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Analysis of DNA sequence transformations on grids
Yadnyesh Joshi, Sathish S. Vadhiyar |
J. Parallel Distributed Comput. | 2 |
| 2008 | Efficient reuse of replicated parallel data segments in computational grids
Sandip Tikar, Sathish S. Vadhiyar |
Future Gener. Comput. Syst. | 2 |
| 2008 | Performance modeling of parallel applications for grid scheduling
H. A. Sanjay 0001, Sathish S. Vadhiyar |
J. Parallel Distributed Comput. | 2 |
| 2007 | An efficient MPI_allgather for gridsabstractAllgather is an important MPI collective communication. Most of the algorithms for allgather have been designed for homogeneous and tightly coupled systems. The existing algorithms for allgather on Gridsystems do not efficiently utilize the bandwidths available on slow wide-area links of the grid. In this paper, we present an algorithm for allgather on grids that efficiently utilizes wide-area bandwidths and is also wide-area optimal. Our algorithm is also adaptive to gridload dynamics since it considers transient network characteristics for dividing the nodes into clusters. Our experiments on a real-grid setup consisting of 3 sites show that our algorithm gives an average performance improvement of 52% over existing strategies. Rakhi Gupta, Sathish S. Vadhiyar |
HPDC | 2 |
| 2006 | Performance Modeling based on Multidimensional Surface Learning for Performance Predictions of Parallel Applications in Non-Dedicated EnvironmentsabstractModeling the performance behavior of parallel applications to predict the execution times of the applications for larger problem sizes and number of processors has been an active area of research for several years. The existing curve fitting strategies for performance modeling utilize data from experiments that are conducted under uniform loading conditions. Hence the accuracy of these models degrade when the load conditions on the machines and network change. In this paper, we analyze a curve fitting model that attempts to predict execution times for any load conditions that may exist on the systems during application execution. Based on the experiments conducted with the model for a parallel eigenvalue problem, we propose a multi-dimensional curve-fitting model based on rational polynomials for performance predictions of parallel applications in non-dedicated environments. We used the rational polynomial based model to predict execution times for 2 other parallel applications on systems with large load dynamics. In all the cases, the model gave good predictions of execution times with average percentage prediction errors of less than 20% Jay Yagnik, H. A. Sanjay 0001, Sathish S. Vadhiyar |
ICPP | 3 |
| 2006 | Application-oriented adaptive MPI_Bcast for gridsabstractDue to the importance of collective communications in scientific parallel applications, many strategies have been devised for optimizing collective communications for different kinds of parallel environments. There has been an increasing interest to evolve efficient broadcast algorithms for computational grids. In this paper, we present application-oriented adaptive techniques that take into account resource characteristics as well as the application's usage of broadcasts for deriving efficient broadcast trees. In particular, we consider two broadcast parameters used in the application, namely, the broadcast message sizes and the time interval between the broadcasts. The results indicate that our adaptive strategies can provide 20% average improvement in performance over the popular MPICH-G2's MPI/spl I.bar/Bcast implementation for loaded network conditions. Rakhi Gupta, Sathish S. Vadhiyar |
IPDPS | 2 |
| 2005 | Self adaptivity in Grid computingabstractAbstract Optimizing a given software system to exploit the features of the underlying system has been an area of research for many years. Recently, a number of self‐adapting software systems have been designed and developed for various computing environments. In this paper, we discuss the design and implementation of a software system that dynamically adjusts the parallelism of applications executing on computational Grids in accordance with the changing load characteristics of the underlying resources. The migration framework implemented by our software system is aimed at performance‐oriented Grid systems and implements tightly coupled policies for both suspension and migration of executing applications. The suspension and migration policies consider both the load changes on systems as well as the remaining execution times of the applications thereby taking into account both system load and application characteristics. The main goal of our migration framework is to improve the response times for individual applications. We also present some results that demonstrate the usefulness of our migration framework. Published in 2005 by John Wiley & Sons, Ltd. Sathish S. Vadhiyar, Jack J. Dongarra |
Concurr. Pract. Exp. | 1 |
| 2004 | GrADSolve a grid-based RPC system for parallel computing with application-level scheduling
Sathish S. Vadhiyar, Jack J. Dongarra |
J. Parallel Distributed Comput. | 1 |
| 2003 | A Performance Oriented Migration Framework For The GridabstractAt least three factors in the existing migration frameworks make them less suitable in Grid systems especially when the goal is to improve the response times for individual applications. These factors are the separate policies for suspension and migration of executing applications employed by these migration frameworks, the use of pre-defined conditions for suspension and migration and the lack of knowledge of the remaining execution time of the applications. In this paper we describe a migration framework for performance oriented Grid systems that implements tightly coupled policies for both suspension and migration of executing applications and takes into account both system load and application characteristics. The main goal of our migration framework is to improve the response times for individual applications. We also present some results that demonstrate the usefulness of our migration framework. Sathish S. Vadhiyar, Jack J. Dongarra |
CCGRID | 1 |
| 2003 | GrADSolve - RPC for High Performance Computing on the Grid
Sathish S. Vadhiyar, Jack J. Dongarra, Asim YarKhan |
Euro-Par | 1 |
| 2002 | A Metascheduler For The GridabstractWith the advent of Grid computing, scheduling strategies for distributed heterogeneous systems have either become irrelevant or have to be extended significantly to support Grid dynamics. In this paper, we describe a metascheduling architecture for a Grid system that takes into account both the application and system level considerations. Results are presented to demonstrate the usefulness of the metascheduler. Sathish S. Vadhiyar, Jack J. Dongarra |
HPDC | 1 |
| 2002 | Middleware for the use of storage in communication
Micah D. Beck, Dorian C. Arnold, Alessandro Bassi, Francine Berman, Henri Casanova, Jack J. Dongarra, Terry Moore, Graziano Obertelli, James S. Plank, D. Martin Swany, Sathish S. Vadhiyar, Richard Wolski |
Parallel Comput. | 11 |
| 2001 | Numerical libraries and the grid: the GrADS experiments with ScaLAPACKabstractThis paper describes an overall framework for the design of numerical libraries on a computational Grid of processors where the processors may be geographically distributed and under the control of a Grid-based scheduling system. A set of experiments are presented in the context of solving systems of linear equations using routines from the ScaLAPACK software collection along with various grid service components, such as Globus, NWS, and Autopilot. Antoine Petitet, L. Susan Blackford, Jack J. Dongarra, Brett Ellis, Graham E. Fagg, Kenneth Roche, Sathish S. Vadhiyar |
SC | 7 |
| 2000 | Automatically Tuned Collective CommunicationsabstractThe performance of the MPI's collective communications is critical in most MPI-based applications. A general algorithm for a given collective communication operation may not give good performance on all systems due to the differences in architectures, network parameters and the storage capacity of the underlying MPI implementation. In this paper, we discuss an approach in which the collective communications are tuned for a given system by conducting a series of experiments on the system. We also discuss a dynamic topology method that uses the tuned static topology shape, but re-orders the logical addresses to compensate for changing run time variations. A series of experiments were conducted comparing our tuned collective communication operations to various native vendor MPI implementations. The use of the tuned collective communications resulted in about 30%-650% improvement in performance over the native MPI implelementations. Sathish S. Vadhiyar, Graham E. Fagg, Jack J. Dongarra |
SC | 1 |