Shantenu Jha

dblp:52/4407 · DBLP profile ↗
← Back
77ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-5040-026XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 42 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 21 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva
Future Gener. Comput. Syst.12
2025 Pilot-Quantum: A Middleware for Quantum-HPC Resource, Workload and Task Management
abstract
As quantum hardware advances, integrating quantum processing units (QPUs) into HPC environments and managing diverse infrastructure and software stacks becomes increasingly essential. Pilot-Quantum addresses these challenges as a middleware designed to provide unified application-level management of resources and workloads across hybrid quantumclassical environments. It is built on a rigorous analysis of existing quantum middleware systems and application execution patterns. It implements the Pilot Abstraction conceptual model, originally developed for HPC, to manage resources, workloads, and tasks. It is designed for quantum applications that rely on task parallelism, including (i) Hybrid algorithms, such as variational approaches, and (ii) Circuit cutting systems, used to partition and execute large quantum circuits. Pilot-Quantum facilitates seamless integration of QPUs, classical CPUs, and GPUs, while supporting high-level programming frameworks like Qiskit and Pennylane. This enables users to efficiently design and execute hybrid workflows across diverse computing resources. The capabilities of Pilot-Quantum are demonstrated through mini-apps - simplified yet representative kernels focusing on critical performance bottlenecks. We demonstrate the capabilities of Pilot-Quantum through multiple mini-apps, including different circuit execution (e.g., using IBM's Eagle QPU and simulators), circuit-cutting, and quantum machine learning scenarios.
Pradeep Kumar Mantha, Florian J. Kiwit, Nishant Saurabh, Shantenu Jha, André Luckow
CCGrid4
2025 Pareto Prompt Optimization
abstract
Natural language prompt optimization, or prompt engineering, has emerged as a powerful technique to unlock the potential of Large Language Models (LLMs) for various tasks. While existing methods primarily focus on maximizing a single task-specific performance metric for LLM outputs, real-world applications often require considering trade-offs between multiple objectives. In this work, we address this limitation by proposing an effective technique for multi-objective prompt optimization for LLMs. Specifically, we propose **ParetoPrompt**, a reinforcement learning~(RL) method that leverages dominance relationships between prompts to derive a policy model for prompts optimization using preference-based loss functions. By leveraging multi-objective dominance relationships, ParetoPrompt enables efficient exploration of the entire Pareto front without the need for a predefined scalarization of multiple objectives. Our experimental results show that ParetoPrompt consistently outperforms existing algorithms that use specific objective values. ParetoPrompt also yields robust performances when the objective metrics differ between training and testing.
Byung-Jun Yoon, Gilchan Park, Shantenu Jha, Shinjae Yoo, Xiaoning Qian
ICLR4
2025 Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
Arup Kumar Sarker, Aymen Alsaadi, Alexander James Halpern, Prabhath Tangella, Mikhail Titov, Niranda Perera, Mills Staylor, Gregor von Laszewski, Shantenu Jha, Geoffrey C. Fox
JSSPP9
2024 An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry
abstract
Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.
Tianle Wang 0001, Jorge Ramirez, Cristina Garcia-Cardona, Thomas Proffen, Shantenu Jha, Sudip K. Seal
IEEE Big Data5
2024 Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific Workflows
abstract
Workflows are critical for scientific discovery. However, the sophistication, heterogeneity, and scale of workflows make building, testing, and optimizing them increasingly challenging. Furthermore, their complexity and heterogeneity make performance reproducibility hard. In this paper, we propose workflow mini-apps as a tool to address the challenges in building and testing workflows while controlling the fidelity of representing real-world workflows. Workflow mini-apps are deployed and run on various HPC systems and architectures without workflow-specific constraints. We offer insight into their design and implementation, providing an analysis of their performance and reproducibility. Workflow mini-apps thus advance the science of workflows by providing simple, portable, and managed (fidelity) representations of otherwise complex and difficult-to-control real workflows.
Ozgur O. Kilic, Tianle Wang 0001, Matteo Turilli, Mikhail Titov, André Merzky, Line C. Pouchard, Shantenu Jha
CCGrid7
2024 Enabling Performance Observability for Heterogeneous HPC Workflows with SOMA
abstract
Heterogeneous workflows represent a promising approach for overcoming traditional application performance limitations and to accelerate scientific insight on high-performance computing (HPC) platforms. As HPC platforms grow in size and complexity, managing and optimizing workflow resources while maximizing scientific output assumes vital importance. Optimal workflow resource allocation requires high-quality and timely information about the state of the hardware resources, the status of the pending tasks, the performance of the tasks that have already been executed, and the current status of the workflow itself. A robust performance observability framework that captures and delivers this information can fundamentally improve the quality of decision-making within the workflow system, setting the stage for the adaptive execution of workflow tasks. We propose the use of SOMA, a service-based performance observability framework for such HPC workflows. With the RADICAL-Pilot runtime system as a development vehicle, SOMA demonstrates that service-based architectures coupled with an appropriate data model can serve the performance monitoring needs of large-scale ensemble workflows in a low-overhead fashion. Effective observability of workflow performance requires exporting, storing, and analyzing several types of performance data from across the application and workflow software stacks. Our study finds significant benefits in integrating observability frameworks as first-class citizens within an HPC workflow software stack. In this paper, we demonstrate how SOMA can simultaneously observe the performance states of the individual tasks, system hardware, and the workflow as a whole. Such information can then be employed to calculate better resource allocation and task configuration.
Dewi Yokelson, Mikhail Titov, Srinivasan Ramesh, Ozgur O. Kilic, Matteo Turilli, Shantenu Jha, Allen D. Malony
ICPP6
2024 Radical-Cylon: A Heterogeneous Data Pipeline for Scientific Computing
Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera, Mills Staylor, Gregor von Laszewski, Matteo Turilli, Ozgur O. Kilic, Mikhail Titov, André Merzky, Shantenu Jha, Geoffrey C. Fox
JSSPP10
2024 Quantum-centric supercomputing for materials science: A perspective on challenges and future directions
Yuri Alexeev, Maximilian Amsler, Marco Antonio Barroca, Sanzio Bassini, Torey Battelle, Daan Camps, David Casanova, Young Jay Choi, Fred Chong, Charles Chung, Christopher Codella, Antonio D. Córcoles, James Cruise, Alberto Di Meglio, Ivan Duran, Thomas Eckl, Sophia E. Economou, Stephan J. Eidenbenz, Bruce Elmegreen, Clyde Fare, Ismael Faro, Cristina Sanz Fernández, Rodrigo Neumann Barros Ferreira, Keisuke Fuji, Bryce Fuller, Laura Gagliardi, Giulia Galli, Jennifer R. Glick, Isacco Gobbi, Pranav Gokhale, Salvador de la Puente Gonzalez, Johannes Greiner, William Gropp, Michele Grossi, Emanuel Gull, Burns Healy, Matthew R. Hermes, Benchen Huang, Travis S. Humble, Nobuyasu Ito, Artur F. Izmaylov, Ali Javadi-Abhari, Douglas M. Jennewein, Shantenu Jha, Bert de Jong, Petar Jurcevic, William M. Kirby, Stefan Kister, Masahiro Kitagawa, Joel Klassen, Katherine Klymko, Kwangwon Koh, Masaaki Kondo, Doga Murat Kürkçüoglu, Krzysztof Kurowski, Teodoro Laino, Ryan Landfield, Matthew L. Leininger, Vicente Leyton-Ortega, Ang Li 0006, Meifeng Lin, Junyu Liu, Nicolás Lorente, André Luckow, Simon Martiel, Francisco Martín-Fernández, Margaret Martonosi, Claire Marvinney, Arcesio Castañeda Medina, Dirk Merten, Antonio Mezzacapo, Kristel Michielsen, Abhishek Mitra, Tushar Mittal, Kyungsun Moon, Joel Moore, Sarah Mostame, Mario Motta, Young-Hye Na, Yunseong Nam, Prineha Narang, Yu-ya Ohnishi, Daniele Ottaviani, Matthew Otten, Scott Pakin, Vincent R. Pascuzzi, Edwin Pednault, Tomasz Piontek, Jed W. Pitera, Patrick Rall, Gokul Subramanian Ravi, Niall Robertson, Matteo A. C. Rossi, Piotr Rydlichowski, Hoon Ryu, Georgy Samsonidze, Mitsuhisa Sato, Nishant Saurabh, Kunal Sharma, Soyoung Shin, George Slessman, Mathias Steiner, Iskandar Sitdikov, In-Saeng Suh, Eric D. Switzer, Joel Thompson, Synge Todo, Minh C. Tran, Dimitar Trenev, Christian Trott, Huan-Hsin Tseng, Norm M. Tubman, Esin Tureci, David García Valiñas, Sofia Vallecorsa, Christopher Wever, Konrad W. Wojciechowski, Xiaodi Wu 0001, Shinjae Yoo, Nobuyuki Yoshioka, Victor Wen-zhe Yu, Seiji Yunoki, Sergiy Zhuk, Dmitry Zubarev
Future Gener. Comput. Syst.44
2023 PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs
abstract
It is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead.
Mihael Hategan, André Merzky, Nicholson T. Collier, Ketan Maheshwari, Jonathan Ozik, Matteo Turilli, Andreas Wilke, Justin M. Wozniak, Kyle Chard, Ian T. Foster, Rafael Ferreira da Silva, Shantenu Jha, Daniel E. Laney
e-Science12
2023 Building the I (Interoperability) of FAIR for Performance Reproducibility of Large-Scale Composable Workflows in RECUP
abstract
Scientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment.
Bogdan Nicolae, Tanzima Z. Islam, Robert B. Ross, Huub J. J. Van Dam, Kevin Assogba, Polina Shpilker, Mikhail Titov, Matteo Turilli, Tianle Wang 0001, Ozgur O. Kilic, Shantenu Jha, Line C. Pouchard
e-Science11
2023 Asynchronous Execution of Heterogeneous Tasks in ML-Driven HPC Workflows
Vincent R. Pascuzzi, Ozgur O. Kilic, Matteo Turilli, Shantenu Jha
JSSPP4
2022 RAPTOR: Ravenous Throughput Computing
abstract
We describe the design, implementation and performance of the RADICAL-Pilot task overlay (RAPTOR). RAPTOR enables the execution of heterogeneous tasks-i.e., functions and executables with arbitrary duration-on HPC platforms, pro-viding high throughput and high resource utilization. RAPTOR supports the high throughput virtual screening requirements of DOE's National Virtual Biotechnology Laboratory effort to find therapeutic solutions for COVID-19. RAPTOR has been used on 8300 compute nodes to sustain 144M/hour docking hits, and to screen 1011 ligands. To the best of our knowledge, both the throughput rate and aggregated number of executed tasks are a factor of two greater than previously reported in literature. RAPTOR represents important progress towards improvement of computational drug discovery, in terms of size of libraries screened, and for the possibility of generating training data fast enough to serve the last generation of docking surrogate models.
André Merzky, Matteo Turilli, Shantenu Jha
CCGRID3
2022 The Ghost of Performance Reproducibility Past
abstract
The importance of ensemble computing is well established. However, executing ensembles at scale introduces interesting performance fluctuations that have not been well investigated. In this paper, we trace our experience uncovering performance fluctuations of ensemble applications (primarily constituting a workflow of GROMACS tasks), and unsuccessful attempts, so far, at trying to discern the underlying cause(s) of performance fluctuations. Is the failure to discern the causative or contributing factors a failure of capability? Or imagination? Do the fluctuations have their genesis in some inscrutable aspect of the system or software? Does it warrant a fundamental reassessment and rethinking of how we assume and conceptualize performance reproducibility? Answers to these questions are not straightforward, nor are they immediate or obvious. We conclude with a discussion about the performance of ensemble applications and ruminate over the implications for how we define and measure application performance.
Srinivasan Ramesh, Mikhail Titov, Matteo Turilli, Shantenu Jha, Allen D. Malony
e-Science4
2022 Coupling streaming AI and HPC ensembles to achieve 100-1000× faster biomolecular simulations
abstract
Machine learning (ML)-based steering can improve the performance of ensemble-based simulations by allowing for online selection of more scientifically meaningful computations. We present DeepDriveMD, a framework for ML-driven steering of scientific simulations that we have used to achieve orders-of-magnitude improvements in molecular dynamics (MD) performance via effective coupling of ML and HPC on large parallel computers. We discuss the design of DeepDriveMD and characterize its performance. We demonstrate that DeepDriveMD can achieve between 100-1000× acceleration for protein folding simulations relative to other methods, as measured by the amount of simulated time performed, while covering the same conformational landscape as quantified by the states sampled during a simulation. Experiments are performed on leadership-class platforms on up to 1020 nodes. The results establish DeepDriveMD as a high-performance framework for ML-driven HPC simulation scenarios, that supports diverse MD simulation and ML back-ends, and which enables new scientific insights by improving the length and time scales accessible with current computing capacity.
Alex Brace, Igor Yakushin, Anda Trifan, Todd S. Munson, Ian T. Foster, Arvind Ramanathan, Hyungro Lee, Matteo Turilli, Shantenu Jha
IPDPS10
2022 RADICAL-Pilot and PMIx/PRRTE: Executing Heterogeneous Workloads at Large Scale on Partitioned HPC Resources
Mikhail Titov, Matteo Turilli, André Merzky, Thomas J. Naughton, Wael R. Elwasif, Shantenu Jha
JSSPP6
2022 Design and Performance Characterization of RADICAL-Pilot on Leadership-Class Platforms
abstract
Many extreme scale scientific applications have workloads comprised of a large number of individual high-performance tasks. The Pilot abstraction decouples workload specification, resource management, and task execution via job placeholders and late-binding. As such, suitable implementations of the Pilot abstraction can support the collective execution of large number of tasks on supercomputers. We introduce RADICAL-Pilot (RP) as a portable, modular and extensible pilot-enabled runtime system. We describe RP's design, architecture and implementation. We characterize its performance and show its ability to scalably execute workloads comprised of tens of thousands heterogeneous tasks on DOE and NSF leadership-class HPC platforms. Specifically, we investigate RP's weak/strong scaling with CPU/GPU, single/multi core, (non)MPI tasks and Python functions when using most of ORNL Summit and TACC Frontera. RADICAL-Pilot can be used stand-alone, as well as the runtime for third-party workflow systems.
André Merzky, Matteo Turilli, Mikhail Titov, Aymen Alsaadi, Shantenu Jha
IEEE Trans. Parallel Distributed Syst.5
2021 Dynamic and Adaptive Monitoring and Analysis for Many-task Ensemble Computing
abstract
Applications are not what they used to be. Modern HPC applications are increasingly a mix of heterogeneous tasks and services – both internal and external. This imposes new constraints and requirements on application and performance monitoring, which is fundamentally different from single-task monitoring. Runtime decisions are needed, and critically rely on monitored information to determine what to do. Thus, online monitoring and analytics must be first-class requirements of all applications and the runtime environments that support their execution. This position paper explores the implications of novel application requirements on future monitoring and profiling subsystems, in the context of ensemble computing.
Shantenu Jha, Allen D. Malony
CLUSTER1
2021 Exploring Task Placement for Edge-to-Cloud Applications using Emulation
abstract
A vast and growing number of IoT applications connect physical devices, such as scientific instruments, technical equipment, machines, and cameras, across heterogenous infrastructure from the edge to the cloud to provide responsive, intelligent services while complying with privacy and security requirements. However, the integration of heterogeneous IoT, edge, and cloud technologies and the design of end-to-end applications that seamlessly work across multiple layers and types of infrastructures is challenging. A significant issue is resource management and the need to ensure that the right type and scale of resources is allocated on every layer to fulfill the application's processing needs. As edge and cloud layers are increasingly tightly integrated, imbalanced resource allocations and sub-optimally placed tasks can quickly deteriorate the overall system performance. This paper proposes an emulation approach for the investigation of task placements across the edge-to-cloud continuum. We demonstrate that emulation can address the complexity and many degrees-of-freedom of the problem, allowing us to investigate essential deployment patterns and trade-offs. We evaluate our approach using a machine learning-based workload, demonstrating the validity by comparing emulation and real-world experiments. Further, we show that the right task placement strategy has a significant impact on performance - in our experiments, between 5% and 65% depending on the scenario.
André Luckow, Kartik Rattan, Shantenu Jha
ICFEC3
2021 IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads
abstract
The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2–3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100 × to 1000 × improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ∼ 1011 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine.
Aymen Alsaadi, Dario Alfè, Yadu N. Babuji, Agastya Bhati, Ben Blaiszik, Alex Brace, Thomas S. Brettin, Kyle Chard, Ryan Chard, Austin Clyde, Peter V. Coveney, Ian T. Foster, Tom Gibbs, Shantenu Jha, Kristopher Keipert, Dieter Kranzlmüller, Thorsten Kurth, Hyungro Lee, Zhuozhao Li, Gerald Mathias, André Merzky, Alexander Partin, Arvind Ramanathan, Ashka Shah, Abraham C. Stern, Rick L. Stevens, Mikhail Titov, Anda Trifan, Aristeidis Tsaris, Matteo Turilli, Huub J. J. Van Dam, Shunzhou Wan, David Wifling, Junqi Yin
ICPP14
2021 Comparing workflow application designs for high resolution satellite image analysis
Aymen Alsaadi, Ioannis Paraskevakos, Bento Collares Gonçalves, Heather J. Lynch, Shantenu Jha, Matteo Turilli
Future Gener. Comput. Syst.5
2020 Parallel performance of molecular dynamics trajectory analysis
abstract
Summary The performance of biomolecular molecular dynamics simulations has steadily increased on modern high‐performance computing resources but acceleration of the analysis of the output trajectories has lagged behind so that analyzing simulations is becoming a bottleneck. To close this gap, we studied the performance of trajectory analysis with message passing interface (MPI) parallelization and the PythonMDAnalysislibrary on three different Extreme Science and Engineering Discovery Environment (XSEDE) supercomputers where trajectories were read from a Lustre parallel file system. Strong scaling performance was impeded by stragglers, MPI processes that were slower than the typical process. Stragglers were less prevalent for compute‐bound workloads, thus pointing to file reading as a bottleneck for scaling. However, a more complicated picture emerged in which both the computation and the data ingestion exhibited close to ideal strong scaling behavior whereas stragglers were primarily caused by either large MPI communication costs or long times to open the single shared trajectory file. We improved overall strong scaling performance by either subfiling (splitting the trajectory into separate files) or MPI‐IO with parallel HDF5 trajectory files. The parallel HDF5 approach resulted in near ideal strong scaling on up to 384 cores (16 nodes), thus reducing trajectory analysis times by two orders of magnitude compared with the serial approach.
Mahzad Khoshlessan, Ioannis Paraskevakos, Geoffrey C. Fox, Shantenu Jha, Oliver Beckstein
Concurr. Comput. Pract. Exp.4
2019 Performance Characterization and Modeling of Serverless and HPC Streaming Applications
abstract
Industrial and scientific streaming applications require support for different types of processing and the management of heterogeneous infrastructure over a dynamic range of scales: from the edge to the cloud and HPC, and intermediate resources. Serverless is an emerging service that combines high-level middleware services, such as distributed execution engines for managing tasks, with low-level infrastructure. It offers the potential of usability and scalability but adds to the complexity of managing heterogeneous and dynamic resources. In response, we extend Pilot-Streaming to support serverless platforms. Pilot-Streaming provides a unified abstraction for resource management for HPC, cloud, and serverless, and allocates resource containers independent of the application workload removing the need to write resource-specific code. Understanding the performance and scaling characteristics of streaming applications and infrastructure presents another challenge. StreamInsight provides insight into the performance of streaming applications and infrastructure, their selection, configuration, and scaling behavior. Underlying StreamInsight is the universal scalability law, which permits the accurate quantification of scalability properties of streaming applications. Using experiments on HPC and AWS Lambda, we demonstrate that StreamInsight provides an accurate model for a variety of application characteristics, e. g., machine learning model sizes and resource configurations.
André Luckow, Shantenu Jha
IEEE BigData2
2019 Learning Everywhere: A Taxonomy for the Integration of Machine Learning and Simulations
abstract
We present a taxonomy of research on Machine Learning (ML) applied to enhance simulations together with a catalog of some activities. We cover eight patterns for the link of ML to the simulations or systems plus three algorithmic areas: particle dynamics, agent-based models and partial differential equations. The patterns are further divided into three action areas: Improving simulation with Configurations and Integration of Data, Learn Structure, Theory and Model for Simulation, and Learn to make Surrogates.
Geoffrey C. Fox, Shantenu Jha
eScience2
2019 Understanding ML Driven HPC: Applications and Infrastructure
abstract
We recently outlined the vision of "Learning Everywhere" which captures the possibility and impact of how learning methods and traditional HPC methods can be coupled together. A primary driver of such coupling is the promise that Machine Learning (ML) will give major performance improvements for traditional HPC simulations. Motivated by this potential, the ML around HPC class of integration is of particular significance. In a related follow-up paper, we provided an initial taxonomy for integrating learning around HPC methods. In this paper which is part of the Learning Everywhere series, we discuss ``how'' learning methods and HPC simulations are being integrated to enhance effective performance of computations. This paper describes several modes --- substitution, assimilation, and control, in which learning methods integrate with HPC simulations and provide representative applications in each mode. This paper discusses some open research questions and we hope will motivate and clear the ground for MLaroundHPC benchmarks.
Shantenu Jha, Geoffrey C. Fox
eScience1
2019 Workflow Design Analysis for High Resolution Satellite Image Analysis
abstract
Ecological sciences are using imagery from a variety of sources to monitor and survey populations and ecosystems. Very High Resolution (VHR) satellite imagery provide an effective dataset for large scale surveys. Convolutional Neural Networks have successfully been employed to analyze such imagery and detect large animals. As the datasets increase in volume, O(TB), and number of images, O(1k), utilizing High Performance Computing (HPC) resources becomes necessary. In this paper, we investigate a task-parallel data-driven workflows design to support imagery analysis pipelines with heterogeneous tasks on HPC. We analyze the capabilities of each design when processing a dataset of 3,000 VHR satellite images for a total of 4~TB. We experimentally model the execution time of the tasks of the image processing pipeline. We perform experiments to characterize the resource utilization, total time to completion, and overheads of each design. Based on the model, overhead and utilization analysis, we show which design approach to is best suited in scientific pipelines with similar characteristics.
Ioannis Paraskevakos, Matteo Turilli, Bento Collares Gonçalves, Heather J. Lynch, Shantenu Jha
eScience5
2018 Enabling Trade-offs Between Accuracy and Computational Cost: Adaptive Algorithms to Reduce Time to Clinical Insight
abstract
The efficacy of drug treatments depends on how tightly small molecules bind to their target proteins. Quantifying the strength of these interactions (the so called `binding affinity') is a grand challenge of computational chemistry, surmounting which could revolutionize drug design and provide the platform for patient specific medicine. Recently, evidence from blind challenge predictions and retrospective validation studies has suggested that molecular dynamics (MD) can now achieve useful predictive accuracy ( 1 kcal/mol) This accuracy is sufficient to greatly accelerate hit to lead and lead optimization. To translate these advances in predictive accuracy so as to impact clinical and/or industrial decision making requires that binding free energy results must be turned around on reduced timescales without loss of accuracy. This demands advances in algorithms, scalable software systems, and intelligent and efficient utilization of supercomputing resources. This work is motivated by the real world problem of providing insight from drug candidate data on a time scale that is as short as possible. Specifically, we reproduce results from a collaborative project between UCL and GlaxoSmithKline to study a congeneric series of drug candidates binding to the BRD4 protein - inhibitors of which have shown promising preclinical efficacy in pathologies ranging from cancer to inflammation. We demonstrate the use of a framework called HTBAC, designed to support the aforementioned requirements of accurate and rapid drug binding affinity calculations. HTBAC facilitates the execution of the numbers of simulations while supporting the adaptive execution of algorithms. Furthermore, HTBAC enables the selection of simulation parameters during runtime which can, in principle, optimize the use of computational resources whilst producing results within a target uncertainty.
Jumana Dakka, Kristof Farkas-Pall, Vivekanandan Balasubramanian, Matteo Turilli, Shunzhou Wan, David W. Wright 0001, Stefan J. Zasada, Peter V. Coveney, Shantenu Jha
CCGrid9
2018 Building Blocks for Workflow System Middleware
abstract
We suggest there is a need for a fresh perspective on the design and development of middleware for high-performance workflows and workflow systems. We argue for a building blocks approach, outline a description of this approach and define their properties. We discuss RADICAL-Cybertools as one implementation of the building blocks concept, showing how they have been designed and developed in accordance with this approach. We discuss three case-studies where RADICAL-Cybertools have been used to develop new workflow systems capabilities and in-tegrated to enhance existing ones, illustrating the potential and promise of the building blocks approach.
Matteo Turilli, André Merzky, Vivekanandan Balasubramanian, Shantenu Jha
CCGrid4
2018 Towards Exascale Computing for High Energy Physics: The ATLAS Experience at ORNL
abstract
Traditionally, the ATLAS experiment at Large Hadron Collider (LHC) has utilized distributed resources as provided by the Worldwide LHC Computing Grid (WLCG) to support data distribution, data analysis and simulations. For example, the ATLAS experiment uses a geographically distributed grid of approximately 200,000 cores continuously (250 000 cores at peak), (over 1,000 million core-hours per year) to process, simulate, and analyze its data (todays total data volume of ATLAS is more than 300 PB). After the early success in discovering a new particle consistent with the long-awaited Higgs boson, ATLAS is continuing the precision measurements necessary for further discoveries. Planned high-luminosity LHC upgrade and related ATLAS detector upgrades, that are necessary for physics searches beyond Standard Model, pose serious challenge for ATLAS computing. Data volumes are expected to increase at higher energy and luminosity, causing the storage and computing needs to grow at a much higher pace than the flat budget technology evolution (see Fig. 1). The need for simulation and analysis will overwhelm the expected capacity of WLCG computing facilities unless the range and precision of physics studies will be curtailed.
V. Ananthraj, Kaushik De, Shantenu Jha, Alexei Klimentov, Danila Oleynik, Sarp Oral, André Merzky, Ruslan Mashinistov, Sergey Panitkin, P. Svirin, Matteo Turilli, Jack C. Wells, Sean R. Wilkinson
eScience3
2018 Pilot-Streaming: A Stream Processing Framework for High-Performance Computing
abstract
An increasing number of scientific applications utilize stream processing to analyze data feeds of scientific instruments, sensors, and simulations. In this paper, we study the streaming and data processing requirements of light source experiments, which are projected to generate data at 20 GB/sec in the near future. As beamtimes available to users are typically short, it is essential that processing and analysis can be conducted in a streaming mode. The development and deployment of streaming applications is a complex task and requires the integration of heterogeneous, distributed infrastructure, frameworks, middleware and application components written in different languages and abstractions. Streaming applications may be extremely dynamic due to factors, such as variable data rates, network congestions, and application-specific characteristics, such as adaptive sampling techniques and the different processing techniques. Consequently, streaming system are often subject to back-pressures and instabilities requiring additional infrastructure to mitigate these issues. We propose Pilot-Streaming, a framework for supporting streaming applications and their resource management needs on HPC infrastructure. Underlying Pilot-Streaming is a unifying architecture that decouples important concerns and functions, such as message brokering, transport and communication, and processing. Pilot-Streaming simplifies the deployment of stream processing frameworks, such as Kafka and Spark Streaming, while providing a high-level abstraction for managing streaming infrastructure, e. g. adding/removing resources as required by the application at runtime. This capability is critical for balancing complex streaming pipelines. To address the complexity in the development of streaming applications, we present the Streaming Mini-Apps, which supports different plug-able algorithms for data generation and processing, e. g., for reconstructing light source images using different techniques. We use the streaming Mini-Apps to evaluate the Pilot-Streaming framework demonstrating its suitability for different use cases and workloads.
George Chantzialexiou, André Luckow, Shantenu Jha
eScience3
2018 Concurrent and Adaptive Extreme Scale Binding Free Energy Calculations
abstract
The efficacy of drug treatments depends on how tightly small molecules bind to their target proteins. The rapid and accurate quantification of the strength of these interactions (as measured by 'binding affinity') is a grand challenge of computational chemistry, surmounting which could revolutionize drug design and provide the platform for patient specific medicine. Recent evidence suggests that molecular dynamics (MD) can achieve useful predictive accuracy (? 1 kcal/mol). For this predictive accuracy to impact clinical decision making, binding free energy results must be turned around rapidly and without loss of accuracy. This demands advances in algorithms, scalable software systems, and efficient utilization of supercomputing resources. We introduce a framework called HTBAC, designed to support accurate and scalable drug binding affinity calculations, while marshaling large simulation campaigns. We show that HTBAC supports the specification and execution of adaptive free-energy protocols at scale and with minimal overheads on NCSA Blue Waters. We validate the results obtained and show how adaptivity can be used to improve accuracy while reducing resource consumption of TIES, a widely used free-energy protocol.
Jumana Dakka, Kristof Farkas-Pall, Matteo Turilli, David W. Wright 0001, Peter V. Coveney, Shantenu Jha
eScience6
2018 Modeling Impact of Execution Strategies on Resource Utilization
abstract
The analysis of the hundreds of petabytes of raw and derived HEP (High Energy Physics) data will necessitate exascale computing. In addition to unprecedented volume, these data are distributed over hundreds of computing centers. In response to these application requirement, as well as performance requirement by using parallel processing (i.e., parallelism), and as a consequence of technology trends, there has been an increase in the uptake of supercomputers by HEP projects.
Alexey A. Poyda, Mikhail Titov, Alexei Klimentov, Jack C. Wells, Sarp Oral, Kaushik De, Danila Oleynik, Shantenu Jha
eScience8
2018 Task-parallel Analysis of Molecular Dynamics Trajectories
abstract
Different parallel frameworks for implementing data analysis applications have been proposed by the HPC and Big Data communities. In this paper, we investigate three task-parallel frameworks: Spark, Dask and RADICAL-Pilot with respect to their ability to support data analytics on HPC resources and compare them to MPI. We investigate the data analysis requirements of Molecular Dynamics (MD) simulations which are significant consumers of supercomputing cycles, producing immense amounts of data. A typical large-scale MD simulation of a physical system of O(100k) atoms over μsecs can produce from O(10) GB to O(1000) GBs of data. We propose and evaluate different approaches for parallelization of a representative set of MD trajectory analysis algorithms, in particular the computation of path similarity and leaflet identification. We evaluate Spark, Dask and RADICAL-Pilot with respect to their abstractions and runtime engine capabilities to support these algorithms. We provide a conceptual basis for comparing and understanding different frameworks that enable users to select the optimal system for each application. We also provide a quantitative performance analysis of the different algorithms across the three frameworks.
Ioannis Paraskevakos, André Luckow, Mahzad Khoshlessan, George Chantzialexiou, Thomas E. Cheatham, Oliver Beckstein, Geoffrey C. Fox, Shantenu Jha
ICPP8
2018 Harnessing the Power of Many: Extensible Toolkit for Scalable Ensemble Applications
abstract
Many scientific problems require multiple distinct computational tasks to be executed in order to achieve a desired solution. We introduce the Ensemble Toolkit (EnTK) to address the challenges of scale, diversity and reliability they pose. We describe the design and implementation of EnTK, characterize its performance and integrate it with two exemplar use cases: seismic inversion and adaptive analog ensembles. We perform nine experiments, characterizing EnTK overheads, strong and weak scalability, and the performance of the two use case imple-mentations, at scale and on production infrastructures. We show how EnTK meets the following general requirements: (i) imple-menting dedicated abstractions to support the description and execution of ensemble applications; (ii) support for execution on heterogeneous computing infrastructures; (iii) efficient scalability up to O(104) tasks; and (iv) task-level fault tolerance. We discuss novel computational capabilities that EnTK enables and the scientific advantages arising thereof. We propose EnTK as an important addition to the suite of tools in support of production scientific computing.
Vivekanandan Balasubramanian, Matteo Turilli, Weiming Hu 0001, Matthieu Lefebvre, Wenjie Lei, Ryan T. Modrak, Guido Cervone, Jeroen Tromp, Shantenu Jha
IPDPS9
2018 Using Pilot Systems to Execute Many Task Workloads on Supercomputers
André Merzky, Matteo Turilli, Manuel Maldonado, Mark Santcroos, Shantenu Jha
JSSPP5
2018 High-throughput binding affinity calculations at extreme scales
abstract
BACKGROUND: Resistance to chemotherapy and molecularly targeted therapies is a major factor in limiting the effectiveness of cancer treatment. In many cases, resistance can be linked to genetic changes in target proteins, either pre-existing or evolutionarily selected during treatment. Key to overcoming this challenge is an understanding of the molecular determinants of drug binding. Using multi-stage pipelines of molecular simulations we can gain insights into the binding free energy and the residence time of a ligand, which can inform both stratified and personal treatment regimes and drug development. To support the scalable, adaptive and automated calculation of the binding free energy on high-performance computing resources, we introduce the High-throughput Binding Affinity Calculator (HTBAC). HTBAC uses a building block approach in order to attain both workflow flexibility and performance. RESULTS: We demonstrate close to perfect weak scaling to hundreds of concurrent multi-stage binding affinity calculation pipelines. This permits a rapid time-to-solution that is essentially invariant of the calculation protocol, size of candidate ligands and number of ensemble simulations. CONCLUSIONS: As such, HTBAC advances the state of the art of binding affinity calculations and protocols. HTBAC provides the platform to enable scientists to study a wide range of cancer drugs and candidate ligands in order to support personalized clinical decision making based on genome sequencing and drug discovery.
Jumana Dakka, Matteo Turilli, David W. Wright 0001, Stefan J. Zasada, Vivekanandan Balasubramanian, Shunzhou Wan, Peter V. Coveney, Shantenu Jha
BMC Bioinform.8
2017 Conceptualizing a Computing Platform for Science Beyond 2020: To Cloudify HPC, or HPCify Clouds?
abstract
A primary challenge of the cyberinfrastructure research community is the need to define the Platforms for Science beyond 2020. We analyze major current trends and propose that in order to deliver the Platform for Science in 2020 the dominant research challenge is to manage the convergence of capabilities of traditional HPC systems with richness of Apache Big Data systems. In this vision paper, we purport to examine the relationship between infrastructure for data-intensive computing and that for High Performance Computing and examine possible "convergence" of capabilities.
Geoffrey C. Fox, Shantenu Jha
CLOUD2
2017 High-Throughput Computing on High-Performance Platforms: A Case Study
abstract
The computing systems used by LHC experiments has historically consisted of the federation of hundreds to thousands of distributed resources, ranging from small to mid-size re-source. In spite of the impressive scale of the existing distributed computing solutions, the federation of small to mid-size resources will be insufficient to meet projected future demands. This paper is a case study of how the ATLAS experiment has embraced Titan - a DOE leadership facility in conjunction with traditional distributed high-throughput computing to reach sustained production scales of approximately 52M core-hours a years. The three main contributions of this paper are: (i) a critical evaluation of design and operational considerations to support the sustained, scalable and production usage of Titan; (ii) a preliminary characterization of a next generation executor for PanDA to support new workloads and advanced execution modes; and (iii) early lessons for how current and future experimental and observational systems can be integrated with production supercomputers and other platforms in a general and extensible manner.
Danila Oleynik, Sergey Panitkin, Matteo Turilli, Alessio Angius, Sarp Oral, Kaushik De, Alexei Klimentov, Jack C. Wells, Shantenu Jha
eScience9
2017 Evaluating Distributed Execution of Workloads
abstract
Resource selection and task placement for distributed execution poses conceptual and implementation difficulties. Although resource selection and task placement are at the core of many tools and workflow systems, the methods are ad hoc rather than being based on models. Consequently, partial and non-interoperable implementations proliferate. We address both the conceptual and implementation difficulties by experimentally characterizing diverse modalities of resource selection and task placement. We compare the architectures and capabilities of two systems: the AIMES middleware and Swift workflow scripting language and runtime. We integrate these systems to enable the distributed execution of Swift workflows on Pilot-Jobs managed by the AIMES middleware. Our experiments characterize and compare alternative execution strategies by measuring the time to completion of heterogeneous uncoupled workloads executed at diverse scale and on multiple resources. We measure the adverse effects of pilot fragmentation and early binding of tasks to resources and the benefits of backfill scheduling across pilots on multiple resources. We then use this insight to execute a multi-stage workflow across five production-grade resources. We discuss the importance and implications for other tools and workflow systems
Matteo Turilli, Yadu N. Babuji, André Merzky, Ming Tai Ha, Michael Wilde, Daniel S. Katz, Shantenu Jha
eScience7
2017 Introducing distributed dynamic data-intensive (D3) science: Understanding applications and infrastructure
abstract
Summary A common feature across many science and engineering applications is the amount and diversity of data and computation that must be integrated to yield insights. Datasets are growing larger and becoming distributed; their location, availability, and properties are often time‐dependent. Collectively, these characteristics give rise to dynamic distributed data‐intensive applications. While “static” data applications have received significant attention, the characteristics, requirements, and software systems for the analysis of large volumes of dynamic, distributed data, and data‐intensive applications have received relatively less attention. This paper surveys several representative dynamic distributed data‐intensive application scenarios, provides a common conceptual framework to understand them, and examines the infrastructure used in support of applications.
Shantenu Jha, Daniel S. Katz, André Luckow, Neil P. Chue Hong, Omer F. Rana, Yogesh L. Simmhan
Concurr. Comput. Pract. Exp.1
2017 On the complexities of utilizing large-scale lightpath-connected distributed cyberinfrastructure
abstract
Summary In Autumn 2013, we—an international team of climate scientists, computer scientists, eScience researchers, and e‐Infrastructure specialists—participated in the enlighten your research global competition, organized to showcase advanced lightpath technologies in support of state‐of‐the‐art research questions. As one of the winning entries, our enlighten your research global team embarked on a very ambitious project to run an extremely high resolution climate model on a collection of supercomputers distributed over two continents and connected using an advanced 10 G lightpath networking infrastructure. Although good progress was made, we were not able to perform all desired experiments due to a varying combination of technical problems, configuration issues, policy limitations and lack of (budget for) human resources to solve these issues. In this paper, we describe our goals, the technical and non‐technical barriers, we encountered and provide recommendations on how these barriers can be removed so future project of this kind may succeed. Copyright © 2016 John Wiley & Sons, Ltd.
Jason Maassen, Ben van Werkhoven, Maarten A. J. van Meersbergen, Henri E. Bal, Michael Kliphuis, Sandra E. Brunnabend, Henk A. Dijkstra, Gerben van Malenstein, Migiel de Vos, Sylvia Kuijpers, Sander Boele, Jules Wolfrat, Nick Hill, David Wallom, Christian Grimm, Dieter Kranzlmüller, Dinesh Ganpathi, Shantenu Jha, Yaakoub El Khamra, Frank O. Bryan, Benjamin Kirtman, Frank J. Seinstra
Concurr. Comput. Pract. Exp.18
2016 ExTASY: Scalable and flexible coupling of MD simulations and advanced sampling techniques
abstract
For many macromolecular systems the accurate sampling of the relevant regions on the potential energy surface cannot be obtained by a single, long Molecular Dynamics (MD) trajectory. New approaches are required to promote more efficient sampling. We present the design and implementation of the Extensible Toolkit for Advanced Sampling and analYsis (Ex-TASY) for building and executing advanced sampling workflows on HPC systems. ExTASY provides Python based “templated scripts” that interface to an interoperable and high-performance pilot-based run time system, which abstracts the complexity of managing multiple simulations. ExTASY supports the use of existing highly-optimised parallel MD code and their coupling to analysis tools based upon collective coordinates which do not require a priori knowledge of the system to bias. We describe two workflows which both couple large “ensembles” of relatively short MD simulations with analysis tools to automatically analyse the generated trajectories and identify molecular conformational structures that will be used on-the-fly as new starting points for further “simulation-analysis” iterations. One of the workflows leverages the Locally Scaled Diffusion Maps technique; the other makes use of Complementary Coordinates techniques to enhance sampling and generate start-points for the next generation of MD simulations. We show that the ExTASY tools have been deployed on a range of HPC systems including ARCHER (Cray CX30), Blue Waters (Cray XE6/XK7), and Stampede (Linux cluster), and that good strong scaling can be obtained up to 1000s of MD simulations, independent of the size of each simulation. We discuss how ExTASY can be easily extended or modified by end-users to build their own workflows, and ongoing work to improve the usability and robustness of ExTASY.
Vivekanandan Balasubramanian, Iain Bethune, Ardita Shkurti, Elena Breitmoser, Eugen Hruska, Cecilia Clementi, Charles A. Laughton, Shantenu Jha
eScience8
2016 Ensemble Toolkit: Scalable and Flexible Execution of Ensembles of Tasks
abstract
There are many science applications that require scalable task-level parallelism, support for flexible execution and coupling of ensembles of simulations. Most high-performance system software and middleware, however, are designed to support the execution and optimization of single tasks. Motivated by the missing capabilities of these computing systems and the increasing importance of task-level parallelism, we introduce the Ensemble toolkit which has the following application development features: (i) abstractions that enable the expression of ensembles as primary entities, and (ii) support for ensemble-based execution patterns that capture the majority of application scenarios. Ensemble toolkit uses a scalable pilot-based runtime system that decouples workload execution and resource management details from the expression of the application, and enables the efficient and dynamic execution of ensembles on heterogeneous computing resources. We investigate three execution patterns and characterize the scalability and overhead of Ensemble toolkit for these patterns. We investigate scaling properties for up to O(1000)concurrent ensembles and O(1000) cores and find linear weak and strong scaling behaviour.
Vivekanandan Balasubramanian, Antons Treikalis, Ole Weidner, Shantenu Jha
ICPP4
2016 RepEx: A Flexible Framework for Scalable Replica Exchange Molecular Dynamics Simulations
abstract
Replica Exchange (RE) simulations have emerged as an important algorithmic tool for the molecular sciences. Typically RE functionality is integrated into the molecular simulation software package. A primary motivation of the tight integration of RE functionality with simulation codes has been performance. This is limiting at multiple levels. First, advances in the RE methodology are tied to the molecular simulation code for which they were developed. Second, it is difficult to extend or experiment with novel RE algorithms, since expertise in the molecular simulation code is required. The tight integration results in difficulty to gracefully handle failures, and other runtime fragilities. We propose the RepEx framework which is addressing aforementioned shortcomings, while striking the balance between flexibility (any RE scheme) and scalability (several thousand replicas) over a diverse range of HPC platforms. The primary contributions of the RepEx framework are: (i) its ability to support different Replica Exchange schemes independent of molecular simulation codes, (ii) provide the ability to execute different exchange schemes and replica counts independent of the specific availability of resources, (iii) provide a runtime system that has first-class support for task-level parallelism, and (iv) provide a required scalability along multiple dimensions.
Antons Treikalis, André Merzky, Tai-Sung Lee, Darrin M. York, Shantenu Jha
ICPP6
2016 Integrating Abstractions to Enhance the Execution of Distributed Applications
abstract
One of the factors that limits the scale, performance, and sophistication of distributed applications is the difficulty of concurrently executing them on multiple distributed computing resources. In part, this is due to a poor understanding of the general properties and performance of the coupling between applications and dynamic resources. This paper addresses this issue by integrating abstractions representing distributed applications, resources, and execution processes into a pilot-based middleware. The middleware provides a platform that can specify distributed applications, execute them on multiple resource and for different configurations, and is instrumented to support investigative analysis. We analyzed the execution of distributed applications using experiments that measure the benefits of using multiple resources, the late-binding of scheduling decisions, and the use of backfill scheduling.
Matteo Turilli, Zhao Zhang 0007, André Merzky, Michael Wilde, Jon B. Weissman, Daniel S. Katz, Shantenu Jha
IPDPS8
2016 Application skeletons: Construction and use in eScience
Daniel S. Katz, André Merzky, Zhao Zhang 0007, Shantenu Jha
Future Gener. Comput. Syst.4
2015 HPC-ABDS High Performance Computing Enhanced Apache Big Data Stack
abstract
We review the High Performance Computing Enhanced Apache Big Data Stack HPC-ABDS and summarize the capabilities in 21 identified architecture layers. These cover Message and Data Protocols, Distributed Coordination, Security & Privacy, Monitoring, Infrastructure Management, DevOps, Interoperability, File Systems, Cluster & Resource management, Data Transport, File management, NoSQL, SQL (NewSQL), Extraction Tools, Object-relational mapping, In-memory caching and databases, Inter-process Communication, Batch Programming model and Runtime, Stream Processing, High-level Programming, Application Hosting and PaaS, Libraries and Applications, Workflow and Orchestration. We summarize status of these layers focusing on issues of importance for data analytics. We highlight areas where HPC and ABDS have good opportunities for integration.
Geoffrey C. Fox, Judy Qiu, Supun Kamburugamuve, Shantenu Jha, André Luckow
CCGRID4
2015 Pilot-Data: An abstraction for distributed data
André Luckow, Mark Santcroos, Ashley Zebrowski, Shantenu Jha
J. Parallel Distributed Comput.4
2014 Advancing next-generation sequencing data analytics with scalable distributed infrastructure
abstract
SUMMARY With the emergence of popular next‐generation sequencing (NGS)‐based genome‐wide protocols such as chromatin immunoprecipitation followed by sequencing (ChIP‐Seq) and RNA‐Seq, there is a growing need for research and infrastructure to support the requirement of effectively analyzing NGS data. Such research and infrastructure do not replace but complement algorithmic advances developments in analyzing NGS data. We present a runtime environment, Distributed Application Runtime Environment, that supports the scalable, flexible, and extensible composition of capabilities that cover the primary requirements of NGS‐based analytics. In this work, we use BFAST as a representative stand‐alone tool used for NGS data analysis and a ChIP‐Seq pipeline as a representative pipeline‐based approach to analyze the computational requirements. We analyze the performance characteristics of BFAST and understand its dependency on different input parameters. The computational complexity of genome‐wide mapping using BFAST, amongst other factors, depends upon the size of a reference genome and the data size of short reads. Characterizing the performance suggests that the mapping benefits from both scaling‐up (increased fine‐grained parallelism) and scaling‐out (task‐level parallelism – local and distributed). For certain problem instances, scaling‐out can be a more efficient approach than scaling‐up. On the basis of investigations using the pipeline for ChIP‐Seq, we also discuss the importance of dynamical execution of tasks. Copyright © 2013 John Wiley & Sons, Ltd.
Joohyun Kim 0001, Sharath Maddineni, Shantenu Jha
Concurr. Comput. Pract. Exp.3
2013 Exploring Dynamic Enactment of Scientific Workflows Using Pilot-Abstractions
abstract
Current workflow abstractions in general lack: (a) an adequate approach to handle distributed data and (b) proper separation between logical tasks and data-flow from their mapping onto physical locations. As the complexity and dynamism of data and processing distribution have increased, optimized mapping of logical tasks to physical resources have become a necessity to avoid bottlenecks. We argue that the management of dynamic data and compute should become part of the runtime system of workflow engines to enable workflows to scale as necessary to address big data challenges and fully exploit distributed computing infrastructures (DCI). In this paper we explore how the P* model for pilot-abstractions, which proposes a clear separation between the logical compute and data units and their realization as a job or a file in some physical resource, could provide these capabilities for such a runtime environment. The Pilot-API provides a general-purpose interface to pilot-abstractions and the ability to assign compute and data resources to them. We share our experience of using the case study of a DNA sequencing pipeline, to re-implement the workflow using the Pilot-API. This first exercise, which resulted in a running application that is discussed here, illustrates the potential of this API to address (a) and (b). Our initial results indicate that the pilot abstractions (as captured by the P* model)offer an interesting approach to explore the design of a new generation of workflow management systems and runtime environments that are capable of intelligently deciding on application-aware late binding to physical resources.
Mark Santcroos, Barbera D. C. van Schaik, Shayan Shahand, Sílvia Delgado Olabarriaga, André Luckow, Shantenu Jha
CCGRID6
2013 Distributed computing practice for large-scale science and engineering applications
abstract
SUMMARY It is generally accepted that the ability to develop large‐scale distributed applications has lagged seriously behind other developments in cyberinfrastructure. In this paper, we provide insight into how such applications have been developed and an understanding of why developing applications for distributed infrastructure is hard. Our approach is unique in the sense that it is centered around half a dozen existing scientific applications; we posit that these scientific applications are representative of the characteristics, requirements, as well as the challenges of the bulk of current distributed applications on production cyberinfrastructure (such as the US TeraGrid). We provide a novel and comprehensive analysis of such distributed scientific applications. Specifically, we survey existing models and methods for large‐scale distributed applications and identify commonalities, recurring structures, patterns and abstractions. We find that there are many ad hoc solutions employed to develop and execute distributed applications, which result in a lack of generality and the inability of distributed applications to be extensible and independent of infrastructure details. In our analysis, we introduce the notion of application vectors: a novel way of understanding the structure of distributed applications. Important contributions of this paper include identifying patterns that are derived from a wide range of real distributed applications, as well as an integrated approach to analyzing applications, programming systems and patterns, resulting in the ability to provide a critical assessment of the current practice of developing, deploying and executing distributed applications. Gaps and omissions in the state of the art are identified, and directions for future research are outlined. Copyright © 2012 John Wiley & Sons, Ltd.
Shantenu Jha, Murray Cole, Daniel S. Katz, Manish Parashar, Omer F. Rana, Jon B. Weissman
Concurr. Comput. Pract. Exp.1
2013 The Impact of a Ligand Binding on Strand Migration in the SAM-I Riboswitch
abstract
Riboswitches sense cellular concentrations of small molecules and use this information to adjust synthesis rates of related metabolites. Riboswitches include an aptamer domain to detect the ligand and an expression platform to control gene expression. Previous structural studies of riboswitches largely focused on aptamers, truncating the expression domain to suppress conformational switching. To link ligand/aptamer binding to conformational switching, we constructed models of an S-adenosyl methionine (SAM)-I riboswitch RNA segment incorporating elements of the expression platform, allowing formation of an antiterminator (AT) helix. Using Anton, a computer specially developed for long timescale Molecular Dynamics (MD), we simulated an extended (three microseconds) MD trajectory with SAM bound to a modeled riboswitch RNA segment. Remarkably, we observed a strand migration, converting three base pairs from an antiterminator (AT) helix, characteristic of the transcription ON state, to a P1 helix, characteristic of the OFF state. This conformational switching towards the OFF state is observed only in the presence of SAM. Among seven extended trajectories with three starting structures, the presence of SAM enhances the trend towards the OFF state for two out of three starting structures tested. Our simulation provides a visual demonstration of how a small molecule (<500 MW) binding to a limited surface can trigger a large scale conformational rearrangement in a 40 kDa RNA by perturbing the Free Energy Landscape. Such a mechanism can explain minimal requirements for SAM binding and transcription termination for SAM-I riboswitches previously reported experimentally.
Joohyun Kim 0001, Shantenu Jha, Fareed Aboul-Ela
PLoS Comput. Biol.3
2012 P∗: A model of pilot-abstractions
abstract
Pilot-Jobs support effective distributed resource utilization, and are arguably one of the most widely-used distributed computing abstractions - as measured by the number and types of applications that use them, as well as the number of production distributed cyberinfrastructures that support them. In spite of broad uptake, there does not exist a well-defined, unifying conceptual model of Pilot-Jobs which can be used to define, compare and contrast different implementations. Often Pilot-Job implementations are strongly coupled to the distributed cyber-infrastructure they were originally designed for. These factors present a barrier to extensibility and interoperability. This paper is an attempt to (i) provide a minimal but complete model (P*) of Pilot-Jobs, (ii) establish the generality of the P* Model by mapping various existing and well known Pilot-Job frameworks such as Condor and DIANE to P*, (iii) derive an interoperable and extensible API for the P* Model (Pilot-API), (iv) validate the implementation of the Pilot-API by concurrently using multiple distinct Pilot-Job frameworks on distinct production distributed cyberinfrastructures, and (v) apply the P* Model to Pilot-Data.
André Luckow, Mark Santcroos, André Merzky, Ole Weidner, Pradeep Kumar Mantha, Shantenu Jha
eScience6
2012 Pilot abstractions for compute, data, and network
abstract
Scientific experiments in a variety of domains are producing increasing amounts of data that need to be processed efficiently. Distributed Computing Infrastructures are increasingly important in fulfilling these large-scale computational requirements.
Mark Santcroos, Sílvia Delgado Olabarriaga, Daniel S. Katz, Shantenu Jha
eScience4
2012 Towards a common model for pilot-jobs
abstract
Pilot-Jobs have become one of the most successful abstractions in distributed computing. In spite of extensive uptake, there does not exist a well defined, unifying conceptual model of pilot-jobs which can be used to define, compare and contrast different implementations. This presents a barrier to extensibility and interoperability. This paper is an attempt to, (i) provide a minimal but complete model (P*) of pilot-jobs, (ii) establish the generality of the P* Model by mapping various existing and well known pilot-jobs frameworks such as Condor and DIANE to P*, (iii) demonstrate the interoperable and concurrent usage of distinct pilot-job frameworks on different production distributed cyberinfrastructures via the use of an extensible API for the P* Model (Pilot-API).
André Luckow, Mark Santcroos, Ole Weidner, André Merzky, Sharath Maddineni, Shantenu Jha
HPDC6
2012 Distributed Application Runtime Environment (DARE): A Standards-based Middleware Framework for Science-Gateways
Sharath Maddineni, Joohyun Kim 0001, Yaakoub El Khamra, Shantenu Jha
J. Grid Comput.4
2011 Energy landscape analysis for regulatory RNA finding using scalable distributed cyberinfrastructure
abstract
SUMMARY We investigate the folding energy landscape for a given RNA sequence through Boltzmann ensemble (BE) sampling of RNA secondary structures. The ensemble of sampled structures is used to derive distributions of energies and base‐pair distances between two configurations. We identify structural features that can be utilized for RNA gene finding. Characterization of the EL through BE sampling of secondary structures is computationally demanding and has multiple heterogeneous stages. We develop the Distributed Adaptive Runtime Environment to effectively address the computational requirements. Distributed Adaptive Runtime Environment is built upon an extensible and interoperable pilot‐job and supports the concurrent execution of a broad range of task sizes across a range of infrastructure. It is used to investigate two RNA systems of different sizes, S‐adenosyl methionine (SAM) binding RNA sequences known as SAM‐I riboswitches, and the S gene of the bovine corona virus RNA genome. We demonstrate how the implementation lowers the total time to solution for increases in RNA length, the number of sequences investigated, and the number of sampled structures. The distributions of energies and base‐pair distances reveal variations in folding dynamics and pathways among the SAM riboswitch sequences. Our results for BCoV RNA genome sequences also indicate sensitivity of folding to coding‐neutral variations in sequence. We search for a characteristic motif from within the SAM‐I consensus structure – a four‐way junction, among BE sampled structures for all 2910 SAM‐I sequences identified from Rfam (the curated ncRNA family database). We find that BE sampling provides insight into the variations in conformational distribution among sequences of the same ncRNA family. Therefore, BE sampling of secondary structures is a viable pre‐processing or post‐processing tool to complement comparative sequence analysis. The understanding gained shows how appropriately designed cyberinfrastructure can provide new insight into RNA folding and structure formation. Copyright © 2011 John Wiley & Sons, Ltd.
Joohyun Kim 0001, Sharath Maddineni, Fareed Aboul-Ela, Shantenu Jha
Concurr. Comput. Pract. Exp.5
2011 Understanding application-level interoperability: Scaling-out MapReduce over high-performance grids and clouds
Saurabh Sehgal, Miklós Erdélyi, André Merzky, Shantenu Jha
Future Gener. Comput. Syst.4
2010 Efficient Runtime Environment for Coupled Multi-physics Simulations: Dynamic Resource Allocation and Load-Balancing
abstract
Coupled Multi-Physics simulations, such as hybrid CFD-MD simulations, represent an increasingly important class of scientific applications. Often the physical problems of interest demand the use of high-end computers, such as TeraGrid resources, which are often accessible only via batch-queues. Batch-queue systems are not developed to natively support the coordinated scheduling of jobs - which in turn is required to support the concurrent execution required by coupled multi-physics simulations. In this paper we develop and demonstrate a novel approach to overcome the lack of native support for coordinated job submission requirement associated with coupled runs. We establish the performance advantages arising from our solution, which is a generalization of the Pilot-Job concept - which in of itself is not new, but is being applied to coupled simulations for the first time. Our solution not only overcomes the initial co-scheduling problem, but also provides a dynamic resource allocation mechanism. Support for such dynamic resources is critical for a load balancing mechanism, which we develop and demonstrate to be effective at reducing the total time-to-solution of the problem. We establish that the performance advantage of using Big Jobs is invariant with the size of the machine as well as the size of the physical model under investigation. The Pilot-Job abstraction is developed using SAGA, which provides an infrastructure agnostic implementation, and which can seamlessly execute and utilize distributed resources.
Soon-Heum Ko, Nayong Kim, Joohyun Kim 0001, Abhinav Thota, Shantenu Jha
CCGRID5
2010 SAGA BigJob: An Extensible and Interoperable Pilot-Job Abstraction for Distributed Applications and Systems
abstract
The uptake of distributed infrastructures by scientific applications has been limited by the availability of extensible, pervasive and simple-to-use abstractions which are required at multiple levels development, deployment and execution stages of scientific applications. The Pilot-Job abstraction has been shown to be an effective abstraction to address many requirements of scientific applications. Specifically, Pilot-Jobs support the decoupling of workload submission from resource assignment; this results in a flexible execution model, which in turn enables the distributed scale-out of applications on multiple and possibly heterogeneous resources. Most Pilot-Job implementations however, are tied to a specific infrastructure. In this paper, we describe the design and implementation of a SAGA-based Pilot-Job, which supports a wide range of application types, and is usable over a broad range of infrastructures, i.e., it is general-purpose and extensible, and as we will argue is also interoperable with Clouds. We discuss how the SAGA-based Pilot-Job is used for different application types and supports the concurrent usage across multiple heterogeneous distributed infrastructure, including concurrent usage across Clouds and traditional Grids/Clusters. Further, we show how Pilot-Jobs can help to support dynamic execution models and thus, introduce new opportunities for distributed applications. We also demonstrate for the first time that we are aware of, the use of multiple Pilot-Job implementations to solve the same problem; specifically, we use the SAGA-based Pilot-Job on high-end resources such as the TeraGrid and the native Condor Pilot-Job (Glide-in) on Condor resources. Importantly both are invoked via the same interface without changes at the development or deployment level, but only an execution (run-time) decision
André Luckow, Lukasz Lacinski, Shantenu Jha
CCGRID3
2010 Exploring the Performance Fluctuations of HPC Workloads on Clouds
abstract
Clouds enable novel execution modes often supported by advanced capabilities such as autonomic schedulers. These capabilities are predicated upon an accurate estimation and calculation of runtimes on a given infrastructure. Using a well understood high-performance computing workload, we find strong fluctuations from the mean performance on EC2 and Eucalyptus-based cloud systems. Our analysis eliminates variations in IO and computational times as possible causes, we find that variations in communication times account for the bulk of the experiment-to-experiment fluctuations of the performance.
Yaakoub El Khamra, Hyunjoo Kim, Shantenu Jha, Manish Parashar
CloudCom3
2010 Abstractions for Loosely-Coupled and Ensemble-Based Simulations on Azure
abstract
Azure is an emerging cloud platform developed and operated by Microsoft. It provides a range of abstractions and building blocks for creating scalable and reliable scientific applications. In this paper we investigate the applicability of the Azure abstractions to the well-known class of loosely coupled and ensemble-based applications. We propose the BigJob API as a novel abstraction for managing groups of Azure worker roles and for remotely executing tasks on them. We demonstrate that Azure enhanced with Big Job functionality provides performance comparable to other grid and cloud offerings loosely-coupled applications.
André Luckow, Shantenu Jha
CloudCom2
2010 Exploring the RNA folding energy landscape using scalable distributed cyberinfrastructure
abstract
The increasing significance of RNAs in transcriptional or post-transcriptional gene regulation processes has generated considerable interest towards the prediction of RNA folding and its sensitivity to environmental factors. We use Boltzmann-weighted sampling to generate RNA secondary structures, which are used to characterize the energy landscape, via the distributions of energies and base-pair distances. Depending upon the length of an RNA, the number of sequences investigated, and the sample size of generated structures --- generating and analyzing sufficient samples can be computationally challenging. We introduce and develop a lightweight and extensible runtime environment that is effective across a range of RNA sizes and other parameters, as well as over a range of infrastructure -- from traditional HPC grids to clouds, without requiring any changes at the application or user level. The Adaptive Distributed Application Management System (ADAMS) is built upon an extensbile and interoperable pilot-job and supports the concurrent execution of a broad range of task sizes across a range of infrastructure. We use ADAMS to investigate the folding energy landscape for two RNA systems of different sizes: a set of S-adenosyl methionine (SAM) binding RNA sequences known as SAM-I riboswitches and the S gene of the Bovine Corona Virus (BCoV) RNA genome that comprises 4092 nucleotides. Results of the energy and base-pair distance distributions suggest different energy landscapes, implying different folding dynamics. With obtained results, we demonstrated the possibility of utilizing this protocol to explore microscopic origins for reported sequence-dependent variation of binding affinity and gene expression in the two RNA systems.
Joohyun Kim 0001, Sharath Maddineni, Fareed Aboul-Ela, Shantenu Jha
HPDC5
2010 Exploring application and infrastructure adaptation on hybrid grid-cloud infrastructure
abstract
Clouds are emerging as an important class of distributed computational resources and are quickly becoming an integral part of production computational infrastructures. An important but oft-neglected question is, what new applications and application capabilities can be supported by clouds as part of a hybrid computational platform? In this paper we use the ensemble Kalman-filter based dynamic application workflow and investigate how clouds can be effectively used as an accelerator to address changing computational requirements as well as changing Quality of Service constraints (e.g., deadlines). Furthermore, we explore how application and system-level adaptivity can be used to improve application performance and achieve a more effective utilization of the hybrid platform. Specifically, we adapt the ensemble Kalman-filter based application formulation (serial versus parallel, different solvers etc.) so as to execute efficiently on a range of different infrastructure (from High Performance Computing grids to clouds that support single core and many-core virtual machines). Our results show that there are performance advantages to be had by supporting application and infrastructure level adaptivity. In general, we find that grid-cloud infrastructure can support novel usage modes, such as deadline-driven scheduling, for applications with tunable characteristics that can adapt to varying resource types.
Hyunjoo Kim, Yaakoub El Khamra, Shantenu Jha, Manish Parashar
HPDC3
2010 Global-scale distributed I/O with ParaMEDIC
abstract
Abstract Achieving high performance for distributed I/O on a wide‐area network continues to be an elusive holy grail. Despite enhancements in network hardware as well as software stacks, achieving high‐performance remains a challenge. In this paper, our worldwide team took a completely new and non‐traditional approach to distributed I/O, calledParaMEDIC: Parallel Metadata Environment for Distributed I/O and Computing, by utilizing application‐specifictransformationof data to orders of magnitude smaller metadata before performing the actual I/O. Specifically, this paper details our experiences in deploying a large‐scale system to facilitate the discovery of missing genes and constructing a genome similarity tree by encapsulating the mpiBLAST sequence‐search algorithm into ParaMEDIC. The overall project involved nine computational sites spread across the U.S. and generated more than a petabyte of data that was ‘teleported’ to a large‐scale facility in Tokyo for storage. Copyright © 2010 John Wiley & Sons, Ltd.
Pavan Balaji, Wu-chun Feng, Heshan Lin, Jeremy S. Archuleta, Satoshi Matsuoka, Andrew S. Warren, João Carlos Setubal, Ewing L. Lusk, Rajeev Thakur, Ian T. Foster, Daniel S. Katz, Shantenu Jha, K. Shinpaugh, Susan Coghlan, Daniel A. Reed
Concurr. Comput. Pract. Exp.12
2010 Large scale computational science on federated international grids: The role of switched optical networks
Peter V. Coveney, Giovanni Giupponi, Shantenu Jha, Steven Manos, Jon MacLaren, Stephen Pickles, R. S. Saksena, Thomas Soddemann, James L. Suter, Mary-Ann Thyveetil, Stefan J. Zasada
Future Gener. Comput. Syst.3
2009 Programming Abstractions for Data Intensive Computing on Clouds and Grids
abstract
MapReduce has emerged as an important data-parallel programming model for data-intensive computing - for Clouds and Grids. However most if not all implementations of MapReduce are coupled to a specific infrastructure. SAGA is a high-level programming interface which provides the ability to create distributed applications in an infrastructure independent way. In this paper, we show how MapReduce has been implemented using SAGA and demonstrate its interoperability across different distributed platforms - Grids, Cloud-like infrastructure and Clouds. We discuss the advantages of programmatically developing MapReduce using SAGA, by demonstrating that the SAGA-based implementation is infrastructure independent whilst still providing control over the deployment, distribution and runtime decomposition. The ability to control the distribution and placement of the computation units (workers) is critical in order to implement the ability to move computational work to the data. This is required to keep data network transfer low and in the case of commercial Clouds the monetary cost of computing the solution low. Using data-sets of size up to 10GB, and upto 10 workers, we provide detailed performance analysis of the SAGA-MapReduce implementation, and show how controllingthe distribution of computation and the payload per worker helps enhance performance.
Chris Miceli, Michael Miceli, Shantenu Jha, Hartmut Kaiser, André Merzky
CCGRID3
2009 An innovative application execution toolkit for multicluster grids
abstract
Multicluster grids provide one promising solution to satisfying growing computation demands of compute-intensive applications by collaborating various networked clusters. However, it is challenging to seamlessly integrate all participating clusters in different domains into a virtual computation platform. In order to take full advantages of multicluster grids capability, computer scientists need to deal with how to collaborate practically and efficiently participating autonomic systems to execute Grid-enabled applications. We make efforts on grid resource management and implement a toolkit called Pelecanus to improve the overall performance of application execution in multicluster grids environment. The Pelecanus takes advantages of the DA-TC (Dynamic Assignment with Task Containers) execution model to improve resource interoperability and enhance application execution and monitoring. Experiments show that it can significantly reduce turnaround time and increase resource utilization for certain applications with large number of sequential jobs.
Zhifeng Yun, Zhou Lei 0001, Gabrielle Allen, Daniel S. Katz, Tevfik Kosar, Shantenu Jha, J. Ramanujam
CLUSTER6
2009 An Autonomic Approach to Integrated HPC Grid and Cloud Usage
abstract
Clouds are rapidly joining high-performance Grids as viable computational platforms for scientific exploration and discovery, and it is clear that production computational infrastructures will integrate both these paradigms in the near future. As a result, understanding usage modes that are meaningful in such a hybrid infrastructure is critical. For example, there are interesting application workflows that can benefit from such hybrid usage modes to, per- haps, reduce times to solutions, reduce costs (in terms of currency or resource allocation), or handle unexpected runtime situations (e.g., unexpected delays in scheduling queues or unexpected failures). The primary goal of this paper is to experimentally investigate, from an applications perspective, how autonomics can enable interesting usage modes and scenarios for integrating HPC Grid and Clouds. Specifically, we used a reservoir characterization application workflow, based on Ensemble Kalman Filters (EnKF) for history matching, and the CometCloud autonomic Cloud engine on a hybrid platform consisting of the TeraGrid and Amazon EC2, to investigate 3 usage modes (or autonomic objectives) - acceleration, conservation and resilience.
Hyunjoo Kim, Yaakoub El Khamra, Shantenu Jha, Manish Parashar
eScience3
2009 A Fresh Perspective on Developing and Executing DAG-Based Distributed Applications: A Case-Study of SAGA-Based Montage
abstract
Most workflow based applications currently have to adapt to available tools. While this keeps the cost of development low, it can lead to performance and flexibility tradeoffs that the application developer and deployer must make. In this paper, we use the Montage astronomical image mosaicking application as prototypical DAG-based workflow application to layout the development and deployment decisions for distributed applications. We discuss and explain the lack of simple (easy-to-use), scalable, and extensible distributed applications. We then introduce SAGA as a technology that permits the construction of abstractions that aid the development and execution of the applications, and thus addresses some of common shortcomings of traditional distributed applications development. We use Montage together with SAGA to examine how legacy applications can be made to run on distributed infrastructures, to see if our reasons are valid, and to compare potential new methods for creating distributed applications with existing technologies that are currently used. We demonstrate the ability to (i) scale-out and (ii) use different production infrastructure, while maintaining performance comparable to established systems. Our hope is that by demonstrating the simplicity of development along with other advantages (performance, scalability, extensibility, and infrastructure independence), this example will encourage others to think more broadly about how distributed applications are created and how new programming models such as Dryad can be supported in an infrastructure independent way, thus eventually leading to more applications that can seamlessly scale-out.
André Merzky, Katerina Stamou, Shantenu Jha, Daniel S. Katz
eScience3
2009 Introduction
Domenico Talia, Jason Maassen, Fabrice Huet, Shantenu Jha
Euro-Par4
2009 Using clouds to provide grids with higher levels of abstraction and explicit support for usage modes
abstract
Abstract Grids in their current form of deployment and implementation have not been as successful as hoped in engendering distributed applications. Among other reasons, the level of detail that needs to be controlled for the successful development and deployment of applications remains too high. We argue that there is a need for higher levels of abstractions for current Grids. By introducing the relevant terminology, we try to understand Grids and Clouds as systems; we find this leads to a natural role for the concept of Affinity, and argue that this is a missing element in current Grids. Providing these affinities and higher‐level abstractions is consistent with the common concepts of Clouds. Thus this paper establishes how Clouds can be viewed as a logical and next higher‐level abstraction from Grids. Copyright © 2009 John Wiley & Sons, Ltd.
Shantenu Jha, André Merzky, Geoffrey C. Fox
Concurr. Comput. Pract. Exp.1
2008 Distributed Replica-Exchange Simulations on Production Environments Using SAGA and Migol
abstract
There exists a class of scientific applications for which utilizing distributed resources is critical for reducing the time-to-solution. In this paper, we discuss a specific class of applications - Replica-Exchange simulations - where the orchestration of many distributed jobs in a dynamic and inherently unreliable distributed environment is essential for a successful completion. We describe the design, development and deployment of a unique framework for constructing fault-tolerant distributed simulations. The framework consists of two primary components - SAGA and Migol. SAGA is a high-level programmatic abstraction layer that provides a standardised interface for the primary distributed functionality required for application development. We present details of a newly developed functionality in SAGA - the Checkpoint and Recovery (CPR) API. Migol is an adaptive middleware, which supports the fault-tolerance of distributed applications by providing the capability to recover applications from checkpoint files transparently. In addition to describing the integration of SAGA-CPR with the Migol infrastructure, we outline our experiences with running a large scale, general-purpose Replica-Exchange application in a production distributed environment.
André Luckow, Shantenu Jha, Joohyun Kim 0001, André Merzky, Bettina Schnor
eScience2
2007 Design and Implementation of Network Performance Aware Applications Using SAGA and Cactus
abstract
This paper demonstrates the use of appropriate programming abstractions - SAGA and cactus - that facilitate the development of applications for distributed infrastructure. SAGA provides a high-level programming interface to Grid- functionality; Cactus is an extensible, component based framework for scientific applications. We show how SAGA can be integrated with cactus to develop simple, useful and easily extensible applications that can be deployed on a wide variety of distributed infrastructure, independent of the details of the resources. Our model application can gather and analyze network performance data and migrate across heterogeneous resources. We outline the architecture of our application and discuss how it imparts important features required of eScience applications. As a proof-of-concept, we present details of the successful deployment of our application over distinct and heterogeneous Grids and present the network performance data gathered. We also discuss several interesting use cases for such an application - which can be used either as stand-alone network diagnostic agent, or in conjunction with more complex scientific applications.
Shantenu Jha, Hartmut Kaiser, Yaakoub El Khamra, Ole Weidner
eScience1
2007 Grid Interoperability at the Application Level Using SAGA
abstract
SAGA is a high-level programming abstraction, which significantly facilitates the development and deployment of Grid-aware applications. The primary aim of this paper is to discuss how each of the three main components of the SAGA landscape - interface specification, specific implementation and the different adaptors for middleware distribution - facilitate application-level interoperability. We discuss SAGA in relation to the ongoing GIN Community Group efforts and show the consistency of the SAGA approach with the GIN Group efforts. We demonstrate how interoperability can be enabled by the use of SAGA, by discussing two simple, yet meaningful applications: in the first, SAGA enables applications to utilize interoperability and in the second example SAGA adaptors provide the basis for interoperability.
Shantenu Jha, Hartmut Kaiser, André Merzky, Ole Weidner
eScience1
2006 Using Lambda Networks to Enhance Performance of Interactive Large Simulations
abstract
The ability to use a visualisation tool to steer large simulations provides innovative and novel usage scenarios, e.g. the ability to use new algorithms for the computation of free energy profiles along a nanopore [1]. However, we find that the performance of interactive simulations is sensitive to the quality of service of the network with variable latency and packet loss in particular having a detrimental effect. The use of dedicated networks (provisioned in this case as a circuit-switched, point-to-point optical lightpath or lambda) can lead to significant (50% or more) performance enhancement. When running on say 128 or 256 processors of a high-end supercomputer this saving has a significant value. We perform experiments to understand the impact of network characteristics on the performance of a large parallel classical molecular dynamics simulation when coupled interactively to a remote visualisation tool. This paper discusses the experiments performed and presents the results from the systematic studies.
Matt J. Harvey, Shantenu Jha, Mary-Ann Thyveetil, Peter V. Coveney
e-Science2
2005 SPICE: Simulated Pore Interactive Computing Environment
abstract
SPICE aims to understand the vital process of translocation of biomolecules across protein pores by computing the free energy profile of the translocating biomolecule along the vertical axis of the pore. Without significant advances at the algorithmic, computing and analysis levels, understanding problems of this size and complexity will remain beyond the scope of computational science for the foreseeable future. A novel algorithmic advance is provided by a combination of Steered Molecular Dynamics and Jarzynski’s Equation (SMD-JE); Grid computing provides the required new computing paradigm as well as facilitating the adoption of new analytical approaches. SPICE uses sophisticated grid infrastructure to couple distributed high performance simulations, visualization and instruments used in the analysis to the same framework. We describe how we utilize the resources of a federated trans-Atlantic Grid to use SMD-JE to enhance our understanding of the translocation phenomenon in ways that have not been possible until now.
Shantenu Jha, Peter V. Coveney, Matt J. Harvey
SC1