VLDB 2026 Research / reviewers in the wild / expert
Manish Parashar
dblp:p/ManishParashar
· DBLP profile ↗
212ranked-venue papers
28as first author
28since 2021 · last 2025
0000-0003-0983-7408ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 167 · 24 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 13 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5Artificial intelligence and machine learning · 4Human-computer interaction and ubiquitous computing · 3Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Private Distributed Resource Management Data: Predicting CPU Utilization with Bi-LSTM and Federated LearningabstractArtificial intelligence is increasingly pervasive in many sectors. In this regard, IT operations are having a big deal on extracting useful information from the large amount of resources' datasets available (e.g., CPU, memory, disk, energy). The issue is bigger if we consider multiple cloud tiers. Artificial intelligence is a key technology when the main goal is to improve microservice migration through offload management. However, it struggles to facilitate distributed contexts where both data transfer needs to be reduced and data privacy needs to be increased. There is therefore a need for novel solutions that resolve the problem of prediction resource utilization (e.g. CPU) while maintaining data privacy and reducing data communication. In this paper, we present a Bi-LSTM model with attention trained in Federated Learning on CPU historical data. The dataset comes from multiple Microsoft Azure trace. The results are compared with the literature and showcase a good generalization and prediction results for metrics collected by multiple virtual machines. The model is evaluated in terms of R-squared, MSE, RMSE and MAE. Lorenzo Carnevale, Daniel Balouek-Thomert, Serena Sebbio, Manish Parashar, Massimo Villari |
CCGrid | 4 |
| 2024 | CrossPrefetch: Accelerating I/O Prefetching for Modern StorageabstractWe introduce CrossPrefetch, a novel cross-layered I/O prefetching mechanism that operates across the OS and a user-level runtime to achieve optimal performance. Existing OS prefetching mechanisms suffer from rigid interfaces that do not provide information to applications on the prefetch effectiveness, suffer from high concurrency bottlenecks, and are inefficient in utilizing available system memory. CrossPrefetch addresses these limitations by dividing responsibilities between the OS and runtime, minimizing overhead, and achieving low cache misses, lock contentions, and higher I/O performance. Shaleen Garg, Rekha Pitchumani, Manish Parashar, Sudarsun Kannan |
ASPLOS (1) | 4 |
| 2024 | Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ WorkflowabstractIn-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms. Bo Zhang 0120, Philip E. Davis, Zhao Zhang 0007, Keita Teranishi, Manish Parashar |
HiPC | 5 |
| 2023 | Message from the ChairsabstractWe are delighted to extend a warm welcome to all participants of the 2023 IEEE International Conference on Cloud Computing (CLOUD 2023) sponsored by the IEEE Technical Committee on Services Computing! Tevfik Kosar, Manish Parashar, Judy Fox, Christoph Hagleithner |
CLOUD | 2 |
| 2023 | Toward Democratizing Access to Science Data: Introducing the National Data PlatformabstractOpen and equitable access to scientific data is essential to addressing important scientific and societal grand challenges, and to research enterprise more broadly. This paper discusses the importance and urgency of open and equitable data access, and explores the barriers and challenges to such access. It then introduces the vision and architecture of the National Data Platform, a recently launched project aimed at catalyzing an open, equitable and extensible data ecosystem. Manish Parashar, Ilkay Altintas |
e-Science | 1 |
| 2023 | Optimizing Data Movement for GPU-Based In-Situ Workflow Using GPUDirect RDMA
Bo Zhang 0120, Philip E. Davis, Nicolas M. Morales, Zhao Zhang 0007, Keita Teranishi, Manish Parashar |
Euro-Par | 6 |
| 2023 | Benesh: a Framework for Choreographic Coordination of In Situ WorkflowsabstractThe growing scale of high-performance computing systems increasingly enables scientists to develop more complex applications as in situ workflows composed of coupled simulation and analysis codes. It is therefore important that workflow programming systems and runtime middleware support the composition and execution of these complex applications intuitively and efficiently. The scientific computing community has put significant effort into purpose-built coupled simulation codes that have been optimized for specialized use cases. However, the development effort involving the coupling of established codes has been largely ad hoc. The Benesh programming system was recently proposed to support the development of coupled simulation workflows from existing code bases. Benesh allows a shared data model to be defined across established codes, so that they can be interfaced in a flexible, coupled workflow. In this paper, we develop Benesh into a workflow development framework. Using Benesh, we develop workflow data model and data exchange definitions for coupled, in situ workflows. We evaluate the cost of development using Benesh in terms of development time and overhead, showing that Benesh offers development advantages without undue impact upon workflow performance. Philip E. Davis, Jacob S. Merson, Pradeep Subedi, Lee F. Ricketson, Cameron W. Smith, Mark S. Shephard, Manish Parashar |
HiPC | 7 |
| 2023 | Computing Everywhere, All at Once: Harnessing the Computing Continuum for ScienceabstractEmerging data-driven scientific workflows are increasing leveraging distributed data sources to understand end-to-end phenomenon, drive experimentation, and facilitate important decision making. Despite the exponential growth of available digital data sources at the edge, and the ubiquity of non-trivial computational power for processing this data across the edge-HPC continuum, realizing such science workflows remain challenging. In this talk I will explore how the computing continuum, spanning resources at the edges, in the core and in-between, can be harnessed to support science. I will also describe recent research in programming abstractions that can express what data should be processed and when and where it should be processed, middleware services that automate the discovery of resources and the orchestration of computations across these resources. Manish Parashar |
HiPC | 1 |
| 2023 | Computing Everywhere, All at Once: Harnessing the Computing Continuum for ScienceabstractEmerging data-driven scientific workflows are increasing leveraging distributed data sources to understand end-to-end phenomenon, drive experimentation, and facilitate important decision making. Despite the exponential growth of available digital data sources at the edge, and the ubiquity of non-trivial computational power for processing this data across the edge-HPC continuum, realizing such science workflows remain challenging. Manish Parashar |
HPDC | 1 |
| 2023 | TEMA: Event Driven Serverless Workflows Platform for Natural Disaster ManagementabstractTEMA project is a Horizon Europe funded project that aims at addressing Natural Disaster Management by the use of sophisticated Cloud-Edge Continuum infrastructures by means of data analysis algorithms wrapped in Serverless functions deployed on a distributed infrastructure according to a Federated Learning scheduler that constantly monitors the infrastructure in search of the best way to satisfy required QoS constraints. In this paper, we discuss the advantages of Serverless workflow and how they can be used and monitored to natively trigger complex algorithm pipelines in the continuum, dynamically placing and relocating them taking into account incoming IoT data, QoS constraints, and the current status of the continuum infrastructure. Therefore we presented the Urgent Function Enabler (UFE) platform, a fully distributed architecture able to define, spread, and manage FaaS functions, using local IOT data managed using the Fiware ecosystem and a computing infrastructure composed of mobile and stable nodes. Christian Sicari, Alessio Catalfamo, Lorenzo Carnevale, Antonino Galletta, Daniel Balouek-Thomert, Manish Parashar, Massimo Villari |
ISCC | 6 |
| 2023 | Adaptive elasticity policies for staging-based in situ visualization
Zhe Wang 0059, Matthieu Dorier, Pradeep Subedi, Philip E. Davis, Manish Parashar |
Future Gener. Comput. Syst. | 5 |
| 2023 | Towards elastic in situ analysis for high-performance computing simulations
Matthieu Dorier, Zhe Wang 0059, Srinivasan Ramesh, Utkarsh Ayachit, Shane Snyder, Robert B. Ross, Manish Parashar |
J. Parallel Distributed Comput. | 7 |
| 2023 | Editorial
Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Assembling Portable In-Situ Workflow from Heterogeneous Components using Data ReorganizationabstractHeterogeneous computing is becoming common in the HPC world. The fast-changing hardware landscape is pushing programmers and developers to rely on performance-portable programming models to rewrite old and legacy applications and develop new ones. While this approach is suitable for individual applications, outstanding challenges still remain when multiple applications are combined into complex workflows. One critical difficulty is the exchange of data between communicating applications where performance constraints imposed by heterogeneous hardware advantage different data layouts. We attempt to solve this problem by exploring asynchronous data layout conversions for applications requiring different memory access patterns for shared data. We implement the proposed solution within the DataSpaces data staging service, extending it to support heterogeneous application workflows across a broad spectrum of programming models. In addition, we integrate heterogeneous DataSpaces with the Kokkos programming model and propose the Kokkos Staging Space as an extension of the Kokkos data abstraction. This new abstraction enables us to express data on a virtual shared space for multiple Kokkos applications, thus guaranteeing the portability of each application when assembling them into an efficient heterogeneous workflow. We present performance results for the Kokkos Staging Space using a synthetic workflow emulator and three different scenarios representing access frequency and use patterns in shared data. The results show that the Kokkos Staging Space is a superior solution in terms of time-to-solution and scalability compared to existing file-based Kokkos data abstractions for inter-application data exchange. Bo Zhang 0120, Pradeep Subedi, Philip E. Davis, Francesco Rizzi, Keita Teranishi, Manish Parashar |
CCGRID | 6 |
| 2022 | Data-Management for Extreme Science: Experiences in Translational Computer Science ResearchabstractExtreme scale computing and data have become essential to computational and data-enabled science and engineering, promising dramatic new insights into natural and engineered systems. However, data-management challenges continue to limit the potential impact of extreme-scale application workflows. Data-intensive application workflows targeting extreme scale, high-performance, distributed computing environments can involve dynamic coordination, interactions and data coupling between multiple application processes that run at scale on different resources, and with services for monitoring, analysis and visualization and archiving, and present challenges due to increasing data volumes and complex data-coupling patterns, system energy constraints, increasing failure rates, etc. In this talk I will present my translational computer science (TCS) research experiences in addressing these challenges. TCS research bridges foundational/use-inspired research with the delivery and deployment of outcomes to the target community and supports the essential bi-direction interplays. Specifically, I will explore data sharing abstractions, managed data pipelines, data-staging services, and in-situ/in-transit data placement and processing to support extreme scale in-situ workflows. Manish Parashar |
HPDC | 1 |
| 2022 | Colza: Enabling Elastic In Situ Visualization for High-performance Computing SimulationsabstractIn situ analysis and visualization have grown increasingly popular for enabling direct access to data from high-performance computing (HPC) simulations. As a simulation progresses and interesting physical phenomena emerge, however, the data produced may become increasingly complex, and users may need to dynamically change the type and scale of in situ analysis tasks being carried out and consequently adapt the amount of resources allocated to such tasks. To date, none of the production in situ analysis frameworks offer such an elasticity feature, and for good reason: the assumption that the number of processes could vary during run time would force developers to rethink software and algorithms at every level of the in situ analysis stack. In this paper we present Colza, a data staging service with elastic in situ visualization capabilities. Colza relies on the widely used ParaView Catalyst in situ visualization framework and enables elasticity by replacing MPI with a custom collective communication library based on the Mochi suite of libraries. To the best of our knowledge, this work is the first to enable elastic in situ visualization capabilities for HPC applications on top of existing production analysis tools. Matthieu Dorier, Zhe Wang 0059, Utkarsh Ayachit, Shane Snyder, Robert B. Ross, Manish Parashar |
IPDPS | 6 |
| 2022 | A General Performance and QoS Model for Distributed Software-Defined Environmentsabstract[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing] Moustafa AbdelBaky, Manish Parashar |
SERVICES | 2 |
| 2022 | RES: Real-Time Video Stream Analytics Using Edge Enhanced CloudsabstractWith increasing availability and use of Internet of Things (IoT) devices such as sensors and video cameras, large amounts of streaming data is now being produced at high velocity. Applications which require low latency response such as video surveillance, augmented reality and autonomous vehicles demand a swift and efficient analysis of this data. Existing approaches employ cloud infrastructure to store and perform machine learning-based analytics on this data. This centralized approach has limited ability to support real-time analysis of large-scale streaming data due to network bandwidth and latency constraints between data source and cloud. We propose RealEdgeStream (RES) an edge enhanced stream analytics system for large-scale, high performance data analytics. The proposed approach investigates the problem of video stream analytics by proposing (i) filtration and (ii) identification phases. The filtration phase reduces the amount of data by filtering low-value stream objects using configurable rules. The identification phase uses deep learning inference to perform analytics on the streams of interest. The phases consist of stages which are mapped onto available in-transit and cloud resources using a placement algorithm to satisfy the Quality of Service (QoS) constraints identified by a user. We demonstrate that for a 10K element data streams, with a frame rate of 15–100 per second, the job completion in the proposed system takes 49 percent less time and saves 99 percent bandwidth compared to a centralized cloud-only based approach. Muhammad K. Ali, Ashiq Anjum, Omer F. Rana, Ali Reza Zamani, Daniel Balouek-Thomert, Manish Parashar |
IEEE Trans. Cloud Comput. | 6 |
| 2022 | EiC Editorial - Advancing Reproducibility in Parallel and Distributed Systems Researchabstract, the TPDS Reproducibility Initiative [2][3] has being exploring postpublication peer review of code associated with published articles for a few years. Authors who have published in TPDS can make their published article more reproducible and earn a reproducibility badge by submitting their associated code for post-publication peer review. To date, this pilot has largely focused on two badges: 1. Code Available: The code, including any associated data and documentation, provided by the authors is reasonable and complete and can potentially be used to support reproducibility of the published results. 2. Code Reviewed: The code, including any associated data and documentation, provided by the authors is reasonable and complete, runs to produce the outputs described, and can support reproducibility of the published results. While TPDS’ goal has always been to include badges for reproducing research results using the code and/or data provided, the nature of research in parallel and distributed systems covered by TPDS makes it challenging to evaluate code and data challenging for reproducibility. This is because such an evaluation may require access to specific hardware, system architectures and scales, OS configurations, and so on, which may not be feasible or practical. vel system/middleware services, is typically infeasible. Consequently, TPDS has piloted an alternate approach where members of the community can submit short, supplemental ‘critique’ papers that present their experiences in reproducing published results using the artifacts, and/or evaluations or experiences with published artifacts. These supplemental paper submissions are reviewed and, if accepted, are linked to the original publication and are citable, serving to help validate the reproducibility of the original publication. This approach was first implemented in a special section guest edited by Stephen Lien Harrell and Beth Plale and consisting of a primary paper and 6 critique papers that reproduce the results of the primary paper. This special section continues this effort, building on the SC20 Student Cluster Competition, which was part of the SC20 conference. It consists of 9 critique papers that reproduce the results of the primary paper. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | A General Performance and QoS Model for Distributed Software-Defined EnvironmentsabstractThe landscape for cloud services and cyberinfrastructure offerings has increased drastically over the past few years. Initially, users moved their applications to the cloud to take advantage of a pay-per-usage model and on-demand access. However, as more cloud providers joined the market, users shifted their goals for using cloud computing from cost reduction to resilience, agility, and optimization. These goals can be achieved by dynamically combining services from multiple providers, for example, to avoid data center or cloud zone outages or to take advantage of extensive offerings with different price points. However, to efficiently support application deployment in this dynamic environment, new models and tools that can measure the application performance and the Quality of Service (QoS) of different configurations are required. The goal of this work is to evaluate the application performance and the QoS of a distributed Software-Defined Environment as well as calculate the QoS of alternative configurations from the underlying pool of services. In particular, we present a mathematical model and a tool for evaluating the performance and QoS of batch application workflows in a distributed environment. We experimentally evaluate the proposed model using a bioinformatics workflow running on infrastructure services from multiple cloud providers. Moustafa AbdelBaky, Manish Parashar |
IEEE Trans. Serv. Comput. | 2 |
| 2021 | RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ WorkflowsabstractWhile in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. We experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance. Pradeep Subedi, Philip E. Davis, Manish Parashar |
CLUSTER | 3 |
| 2021 | Adaptive Placement of Data Analysis Tasks For Staging Based In-Situ ProcessingabstractIn-situ processing addresses the gap between speeds of computing and I/O capabilities by processing data close to the data source, i.e., on the same system as the data source (e.g., a simulation). However, the effective implementation of in-situ processing workflows requires the optimization of several design parameters such as where on the system workflow data analysis/visualization (ana/vis) as placed and how execution as well as the interaction and data exchanges between ana/vis are coordinated. For example, in the case of hybrid in-situ processing, interacting ana/vis may be tightly or loosely coupled depending on their placement, and this can lead to very different performance and scalability. A key challenge is deciding the most appropriate ana/vis placement, which depends on dynamic applications, workflow, and system characteristics that might change at runtime. In this paper, we present a framework to support online adaptive data analysis placement during the execution of an in-situ workflow. Specifically, the paper presents a model and architecture, and explores several data analysis placement strategies. Evaluation results show that dynamically choosing appropriate data analysis placement strategies can balance the benefits and overhead of different data analysis placement patterns to reduce in-situ processing time. Zhe Wang 0059, Pradeep Subedi, Matthieu Dorier, Philip E. Davis, Manish Parashar |
HiPC | 5 |
| 2021 | Facilitating Data Discovery for Large-scale Science Facilities using Knowledge NetworksabstractLarge-scale multiuser scientific facilities, such as geographically distributed observatories, remote instruments, and experimental platforms, represent some of the largest national investments and can enable dramatic advances across many areas of science. Recent examples of such advances include the detection of gravitational waves and the imaging of a black hole's event horizon. However, as the number of such facilities and their users grow, along with the complexity, diversity, and volumes of their data products, finding and accessing relevant data is becoming increasingly challenging, limiting the potential impact of facilities. These challenges are further amplified as scientists and application workflows increasingly try to integrate facilities' data from diverse domains. In this paper, we leverage concepts underlying recommender systems, which are extremely effective in e-commerce, to address these data-discovery and data-access challenges for large-scale distributed scientific facilities. We first analyze data from facilities and identify and model user-query patterns in terms of facility location and spatial localities, domain-specific data models, and user associations. We then use this analysis to generate a knowledge graph and develop the collaborative knowledge-aware graph attention network (CKAT) recommendation model, which leverages graph neural networks (GNNs) to explicitly encode the collaborative signals through propagation and combine them with knowledge associations. Moreover, we integrate a knowledge-aware neural attention mechanism to enable the CKAT to pay more attention to key information while reducing irrelevant noise, thereby increasing the accuracy of the recommendations. We apply the proposed model on two real-world facility datasets and empirically demonstrate that the CKAT can effectively facilitate data discovery, significantly outperforming several compelling state-of-the-art baseline models. Yubo Qin, Ivan Rodero, Manish Parashar |
IPDPS | 3 |
| 2021 | Enabling microservices management for Deep Learning applications across the Edge-Cloud ContinuumabstractDeep Learning has shifted the focus of traditional batch workflows to data-driven feature engineering on streaming data. In particular, the execution of Deep Learning workflows presents expectations of near-real-time results with user-defined acceptable accuracy. Meeting the objectives of such applications across heterogeneous resources located at the edge of the network, the core, and in-between requires managing trade-offs between the accuracy and the urgency of the results. However, current data analysis rarely manages the entire Deep Learning pipeline along the data path, making it complex for developers to implement strategies in realworld deployments. Driven by an object detection use case, this paper presents an architecture for time-critical Deep Learning workflows by providing a data-driven scheduling approach to distribute the pipeline across Edge to Cloud resources. Furthermore, it adopts a data management strategy that reduces the resolution of incoming data when potential trade-off optimizations are available. We illustrate the system's viability through a performance evaluation of the object detection use case on the Grid'5000 testbed. We demonstrate that in a multi-user scenario, with a standard frame rate of 25 frames per second, the system speed-up data analysis up to 54.4% compared to a Cloud-only-based scenario with an analysis accuracy higher than a fixed threshold. Zeina Houmani, Daniel Balouek-Thomert, Eddy Caron, Manish Parashar |
SBAC-PAD | 4 |
| 2021 | Leveraging user access patterns and advanced cyberinfrastructure to accelerate data delivery from shared-use scientific observatories
Yubo Qin, Ivan Rodero, Anthony Simonet, Charles Meertens, Daniel Reiner, James Riley, Manish Parashar |
Future Gener. Comput. Syst. | 7 |
| 2021 | Editor's NoteabstractPresents the introductory editorial for this issue of the publication. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Editor's NoteabstractPresents the introductory editorial for this issue of the publication. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Guest Editorial: Special Section on SC19 Student Cluster CompetitionabstractReproducibility is foundational to solid scientific and technical research, and the ability to repeat the research that produced is a key approach for confirming the validity of a new scientific discovery. The IEEE Transactions on Parallel and Distributed Systems (TPDS) is committed to enabling reproducible research through transparency and the availability and potential reuse of code. As part of its Reproducibility Initiative, TPDS is exploring post-publication peer review of code associated with articles published in TPDS. Authors who have published in TPDS can make their published article more reproducible and earn a reproducibility badge by submitting their associated code for post-publication peer review. However, the nature of the parallel and distributed systems research covered by TPDS makes it challenging to evaluate code and data challenging for reproducibility. This is because such an evaluation may require access to specific hardware, system architectures and scales, OS configurations, and so on, which may not be feasible or practical. Consequently, TPDS is exploring an alternate approach where members of the community can submit short, supplemental ‘critique’ papers that present their experiences in reproducing published results using the artifacts and/or evaluations or experiences with published artifacts. These supplemental paper submissions are reviewed and, if accepted, are linked to the original publication and are citable, serving to validate the reproducibility of the original publication. This special section, consisting of a primary paper and 6 critique papers that reproduce the results of the primary paper, is the initial pilot of this approach. The special section builds on the efforts of the SC19 Student Cluster Competition, which was part of the SC19 conference (https://sc19.supercomputing.org/) and represents a step forward in enabling reproducible research publications in parallel and distributed systems. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | A Distributed Multi-Sensor Machine Learning Approach to Earthquake Early WarningabstractOur research aims to improve the accuracy of Earthquake Early Warning (EEW) systems by means of machine learning. EEW systems are designed to detect and characterize medium and large earthquakes before their damaging effects reach a certain location. Traditional EEW methods based on seismometers fail to accurately identify large earthquakes due to their sensitivity to the ground motion velocity. The recently introduced high-precision GPS stations, on the other hand, are ineffective to identify medium earthquakes due to its propensity to produce noisy data. In addition, GPS stations and seismometers may be deployed in large numbers across different locations and may produce a significant volume of data consequently, affecting the response time and the robustness of EEW systems.In practice, EEW can be seen as a typical classification problem in the machine learning field: multi-sensor data are given in input, and earthquake severity is the classification result. In this paper, we introduce the Distributed Multi-Sensor Earthquake Early Warning (DMSEEW) system, a novel machine learning-based approach that combines data from both types of sensors (GPS stations and seismometers) to detect medium and large earthquakes. DMSEEW is based on a new stacking ensemble method which has been evaluated on a real-world dataset validated with geoscientists. The system builds on a geographically distributed infrastructure, ensuring an efficient computation in terms of response time and robustness to partial infrastructure failures. Our experiments show that DMSEEW is more accurate than the traditional seismometer-only approach and the combined-sensors (GPS and seismometers) approach that adopts the rule of relative strength. Kevin Fauvel, Daniel Balouek-Thomert, Diego Melgar, Pedro Silva 0007, Anthony Simonet, Gabriel Antoniu, Alexandru Costan, Véronique Masson, Manish Parashar, Ivan Rodero, Alexandre Termier |
AAAI | 9 |
| 2020 | Enhancing microservices architectures using data-driven service discovery and QoS guaranteesabstractMicroservices promise the benefits of services with an efficient granularity using dynamically allocated resources. In the current evolving architectures, data producers and consumers are created as decoupled components that support different data objects and quality of service. Actual implementations of service meshes lack support for data-driven paradigms, and focus on goal-based approaches designed to fulfill the general system goal. This diversity of available components demands the integration of users requirements and data products into the discovery mechanism. This paper proposes a data-driven service discovery framework based on profile matching using data-centric service descriptions. We have designed and evaluated a microservices architecture for providing service meshes with a standalone set of components that manages data profiles and resources allocations over multiple geographical zones. Moreover, we demonstrated an adaptation scheme to provide quality of service guarantees. Evaluation of the implementation on a real life testbed shows effectiveness of this approach with stable and fluctuating request incoming rates. Zeina Houmani, Daniel Balouek-Thomert, Eddy Caron, Manish Parashar |
CCGRID | 4 |
| 2020 | Trading Data Size and CNN Confidence Score for Energy Efficient CPS Node CommunicationsabstractIn a context of Cyber-Physical Systems (CPS), energy-efficiency is a critical factor to achieve long operational life-time. The constraint of using battery-powered devices adds degrees of complexity, especially in a hard to reach environment with scarce network and energy resources. The reporting of data consumes large amount of energy, reducing the life-time of both individual nodes and the CPS as a whole. One way to reduce the energy cost of communication is to reduce the number of Bits to transmit. However, this is a viable approach only if the transmitted data remain suitable for further analysis.In this paper, we report on the effect of reducing the image dimensions on the confidence score computed by a convolutional neural network (CNN) determining the species of animals present in images. We also report on the energy consumption of transmitting full vs. reduced dimensions of images. CPS devices and CNNs developed by the Distributed Arctic Observatory (DAO) project are used as experimental platforms.The results show that the energy needed to report the images can be reduced by up to 98% while only reducing the average confidence of determining the species correctly by 0.10%. Issam Raïs, Otto J. Anshus, John Markus Bjørndalen, Daniel Balouek-Thomert, Manish Parashar |
CCGRID | 5 |
| 2020 | Staging Based Task Execution for Data-driven, In-Situ Scientific WorkflowsabstractAs scientific workflows increasingly use extreme-scale resources, the imbalance between higher computational capabilities, generated data volumes, and available I/O bandwidth is limiting the ability to translate these scales into insights. In-situ workflows (and the in-situ approach) are leveraging storage levels close to the computation in novel ways in order to reduce the required I/O. However, to be effective, it is important that the mapping and execution of such in-situ workflows adopts a data-driven approach, enabling in-situ tasks to be executed flexibly based upon data content. This paper first explores the design space for data-driven in-situ workflows. Specifically, it presents a model that captures different factors that influence the mapping, execution, and performance of data-driven in-situ workflows and experimentally studies the impact of different mapping decisions and execution patterns. The paper then presents the design, implementation, and experimental evaluation of a data-driven in-situ workflow execution framework that leverages in-memory distributed data management and user-defined task-triggers to enable efficient and scalable in-situ workflow execution. Zhe Wang 0059, Pradeep Subedi, Matthieu Dorier, Philip E. Davis, Manish Parashar |
CLUSTER | 5 |
| 2020 | Submarine: A subscription-based data streaming framework for integrating large facilities and advanced cyberinfrastructureabstractSummary Large scientific facilities provide researchers with instrumentation, data, and data products that can accelerate scientific discovery. However, increasing data volumes coupled with limited local computational power prevents researchers from taking full advantage of what these facilities can offer. Many researchers looked into using commercial and academic cyberinfrastructure (CI) to process these data. Nevertheless, there remains a disconnect between large facilities and CI that requires researchers to be actively part of the data processing cycle. The increasing complexity of CI and data scale necessitates new data delivery models, those that can autonomously integrate large‐scale scientific facilities and CI to deliver real‐time data and insights. In this paper, we present our initial efforts using the Ocean Observatories Initiative project as a use case. In particular, we present a subscription‐based data streaming service for data delivery that leverages the Apache Kafka data streaming platform. We also show how our solution can automatically integrate large‐scale facilities with CI services for automated data processing. Ali Reza Zamani, Moustafa AbdelBaky, Daniel Balouek-Thomert, Juan J. Villalobos, Ivan Rodero, Manish Parashar |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | An edge-aware autonomic runtime for data streaming and in-transit processing
Ali Reza Zamani, Daniel Balouek-Thomert, Juan J. Villalobos, Ivan Rodero, Manish Parashar |
Future Gener. Comput. Syst. | 5 |
| 2020 | Towards autonomic data management for staging-based coupled scientific workflows
Tong Jin 0002, Fan Zhang 0004, Melissa Romanus, Hoang Bui, Manish Parashar |
J. Parallel Distributed Comput. | 6 |
| 2020 | Big Data for Cyber-Physical SystemsabstractCyber-physical systems (CPS) are characterized by deep and complex intertwining among cyber components and physical components. Due to the fast increase in system complexities, the operations of CPS involve sensing, processing and storage of massive amount of data. This nature of “big data” imposes fundamental challenges on the design and management of CPS in multiple aspects such as performance, energy efficiency, security, privacy, reliability, sustainability, fault tolerance, scalability and flexibility. Tackling these challenges necessitates innovative big data techniques for handling massive data in CPS. The articles in this special section include a few selected state-of-the-art research results on the topic of big data sensing, processing and storage for CPS, and stimulates a broad range of researchers to participate in the interdisciplinary CPS research in the future. This special issue has received a significant number of submissions while only a small portion of them are selected for publications. The selected papers showcase how interesting data analytics techniques can be leveraged to optimize different metrics in CPS, such as timing, efficiency, schedulability, power, reliability, and security, etc. Shiyan Hu 0001, Xin Li 0001, Haibo He, Shuguang Cui, Manish Parashar |
IEEE Trans. Big Data | 5 |
| 2020 | Guest Editorial: Special Section on Advances of Utility and Cloud Computing Technologies and ServicesabstractThe articles in this special section focus on advancements of utility and cloud computing technologies and services. Computing is rapidly moving towards a model where it is provided as services that are delivered in a manner similar to traditional utilities such as water, electricity, gas, and telephony. In such a model, users access services according to their requirements, without regard to where the services are hosted or how they are delivered. Several computing architectures have evolved to realize this utility computing vision, including Grid computing, Service- Oriented Architecture (SOA) and Cloud computing,which has recently shifted into the center of attention in the ICT industry. Increasing numbers of IT vendors are promising to offer applications, storage and computation hosting services with conforming Service-Level Agreements Ching-Hsien Hsu, Manish Parashar, Omer F. Rana |
IEEE Trans. Cloud Comput. | 2 |
| 2020 | Editor's NoteabstractPresents the introductory editorial for this issue of the publication. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Deadline Constrained Video Analysis via In-Transit Computational EnvironmentsabstractCombining edge processing (at data capture site) with analysis carried out while data is enroute from the capture site to a data center offers a variety of different processing models. Such in-transit nodes include network data centers that have generally been used to support content distribution (providing support for data multicast and caching), but have recently started to offer user-defined programmability, through Software Defined Networks (SDN) capability, e.g., OpenFlow and Network Function Visualization (NFV). We demonstrate how this multi-site computational capability can be aggregated to support video analytics, with Quality of Service and cost constraints (e.g., latency-bound analysis). The use of SDN technology enables separation of the data path from the control path, enabling in-network processing capabilities to be supported as data is migrated across the network. We propose to leverage SDN capability to gain control over the data transport service with the purpose of dynamically establishing data routes such that we can opportunistically exploit the latent computational capabilities located along the network path. Using a number of scenarios, we demonstrate the benefits and limitations of this approach for video analysis, comparing this with the baseline scenario of undertaking all such analysis at a data center located at the core of the infrastructure. Ali Reza Zamani, Mengsong Zou, Javier Diaz Montes, Ioan Petri, Omer F. Rana, Ashiq Anjum, Manish Parashar |
IEEE Trans. Serv. Comput. | 7 |
| 2019 | Two-Stage Framework for Big Spatial Data Analytics to Support Disaster ResponseabstractDuring disaster response, large volumes and diverse types of data sets are often continuously generated, and in many case these data sets create overwhelming burdens to data processing infrastructure and teams. At the same time, decision making during disaster response requires timely and relevant information which has to be extracted as expeditiously as possible from these large data sets. Therefore, processing of disaster related data sets is often time sensitive and requires coordination and prioritization. To accomplish this, we propose a two-stage approach to facilitate efficient and effective data processing for disaster decision support. In the first stage, a Data Envelope Analysis (DEA) model is introduced to model the articulation process about information needs such that providing a formal way of prioritizing data processing task. In the second stage, the prioritized data processing workflow is implemented on an Apache Storm based streaming processing platform in the EC2 cloud, with a focus on computational resource optimization. To validate the proposed approach, a Hurricane Sandy based use case was used to evaluate the performance of the proposed approach. Results show that our approach can compute up to 69% (three supervisor nodes) faster than a conventional serial processing approach. Eduard Gibert Renart, Manish Parashar |
IEEE BigData | 4 |
| 2019 | Optimizing Performance and Computing Resource Management of In-memory Big Data Analytics with Disaggregated Persistent MemoryabstractThe performance of modern Big Data frameworks, e.g. Spark, depends greatly on high-speed storage and shuffling, which impose a significant memory burden on production data centers. In many production situations, the persistence and shuffling intensive applications can suffer a major performance loss due to lack of memory. Thus, the common practice is usually to over-allocate the memory assigned to the data workers for production applications, which in turn reduces overall resource utilization. One efficient way to address the dilemma between the performance and cost efficiency of Big Data applications is through data center computing resource disaggregation. This paper proposes and implements a system that incorporates the Spark Big Data framework with a novel in-memory distributed file system to achieve memory disaggregation for data persistence and shuffling. We address the challenge of optimizing performance at affordable cost by co-designing the proposed in-memory distributed file system with large-volume DIMM-based persistent memory (PMEM) and RDMA technology. The disaggregation design allows each part of the system to be scaled independently, which is particularly suitable for cloud deployments. The proposed system is evaluated in a production-level cluster using real enterprise-level Spark production applications. The results of an empirical evaluation show that the system can achieve up to a 3.5- fold performance improvement for shuffle-intensive applications with the same amount of memory, compared to the default Spark setup. Moreover, by leveraging PMEM, we demonstrate that our system can effectively increase the memory capacity of the computing cluster with affordable cost, with a reasonable execution time overhead with respect to using local DRAM only. Shouwei Chen, Xueyang Wu 0002, Zhen Fan 0009, Kunwu Huang, Peiyu Zhuang, Ivan Rodero, Manish Parashar, Dennis Z. Weng |
CCGRID | 9 |
| 2019 | Distributed Operator Placement for IoT Data Analytics Across Edge and Cloud ResourcesabstractThe number of Internet of Things applications is forecast to grow exponentially within the coming decade. Owners of such applications strive to make predictions from large streams of complex input in near real time. Cloud-based architectures often centralize storage and processing, generating high data movement overheads that penalize real-time applications. Edge and Cloud architecture pushes computation closer to where the data is generated, reducing the cost of data movements and improving the application response time. The heterogeneity among the edge devices and cloud servers introduces an important challenge for deciding how to split and orchestrate the IoT applications across the edge and the cloud. In this paper, we extend our IoT Edge Framework, called R-Pulsar, to propose a solution on how to split IoT applications dynamically across the edge and the cloud, allowing us to improve performance metrics such as end-to-end latency (response time), bandwidth consumption, and edge-to-cloud and cloud-to-edge messaging cost. Our approach consists of a programming model and real-world implementation of an IoT application. The results show that our approach can minimize the end-to-end latency by at least 38% by pushing part of the IoT application to the edge. Meanwhile, the edge-to-cloud data transfers are reduced by at least 38% and the messaging costs are reduced by at least 50% when using the existing commercial edge cloud cost models. Eduard Gibert Renart, Alexandre da Silva Veith, Daniel Balouek-Thomert, Marcos Dias de Assunção, Laurent Lefèvre, Manish Parashar |
CCGRID | 6 |
| 2019 | Leveraging Machine Learning for Anticipatory Data Delivery in Extreme Scale In-situ WorkflowsabstractExtreme scale scientific workflows are composed of multiple applications that exchange data at runtime. Several data-related challenges are limiting the potential impact of such workflows. While data staging and in-situ models of execution have emerged as approaches to address data-related costs at extreme scales, increasing data volumes and complex data exchange patterns impact the effectiveness of such approaches. In this paper, we design and implement DESTINY, which is an autonomic data delivery mechanism for staging-based in-situ workflows. DESTINY dynamically learns the data access patterns of scientific workflow applications and leverages these patterns to decrease data access costs. Specifically, DESTINY uses machine learning techniques to anticipate future data accesses, proactively packages and delivers the data necessary to satisfy these requests as close to the consumer as possible and, when data staging processes and consumer processes are colocated, removes the need for inter-process communication by making these data available to the consumer as shared-memory objects. When consumer processes reside on nodes other than staging nodes, the data is packaged and stored in a format the client will likely access in future. This amortizes expensive data discovery and assembly operations typically associated with data staging. We experimentally evaluate the performance and scalability of DESTINY on leadership class platforms using synthetic applications and the S3D combustion workflow. We demonstrate that DESTINY is scalable and can achieve a reduction of up to 75% in read response time as compared to in-memory staging service for production scientific workflows. Pradeep Subedi, Philip E. Davis, Manish Parashar |
CLUSTER | 3 |
| 2019 | Addressing data resiliency for staging based scientific workflowsabstractAs applications move towards extreme scales, data-related challenges are becoming significant concerns, and in-situ workflows based on data staging and in-situ/in-transit data processing have been proposed to address these challenges. Increasing scale is also expected to result in an increase in the rate of silent data corruption errors, which will impact both the correctness and performance of applications. Furthermore, this impact is amplified in the case of in-situ workflows due to the dataflow between the component applications of the workflow. While existing research has explored silent error detection at the application level, silent error detection for workflows remains an open challenge. This paper addresses silent error detection for extreme scale in-situ workflows. The presented approach leverages idle computation resource in data staging to enable timely detection and recovery from silent data corruption, effectively reducing the propagation of corrupted data and end-to-end workflow execution time in the presence of silent errors. As an illustration of this approach, we use a spatial outlier detection approach in staging to detect errors introduced in data transfer and storage. We also provide a CPU-GPU hybrid staging framework for error detection in order to achieve faster error identification. We have implemented our approach within the DataSpaces staging service, and evaluated it using both synthetic and real workflows on a Cray XK7 system (Titan) at different scales. We demonstrate that, in the presence of silent errors, enabling error detection on staged data alongside a checkpoint/restart scheme improves the total in-situ workflow execution time by up to 22% in comparison with using checkpoint/restart alone. Shaohua Duan, Pradeep Subedi, Philip E. Davis, Manish Parashar |
SC | 4 |
| 2019 | Towards comprehensive dependability-driven resource use and message log-analysis for HPC systems diagnosis
Edward Chuah, Arshad Jhumka, Samantha Alt, Daniel Balouek-Thomert, James C. Browne, Manish Parashar |
J. Parallel Distributed Comput. | 6 |
| 2019 | Editor's NoteabstractPresents the introductory editorial for this issue of the publication. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Editor's NoteabstractPresents the introductory editorial for this issue of the publication. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Editor's Note: IEEE Transactions on Parallel and Distributed Systems (TPDS) Reproducibility Initiative, June 2019abstractPresents the introductory editorial for this issue of the publication. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Coupling Exascale Multiphysics Applications: Methods and Lessons LearnedabstractWith the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration. Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood |
eScience | 29 |
| 2018 | A View from ORNL: Scientific Data Research Opportunities in the Big Data AgeabstractOne of the core issues across computer and computational science today is adapting to, managing, and learning from the influx of "Big Data". In the commercial space, this problem has led to a huge investment in new technologies and capabilities that are well adapted to dealing with the sorts of human-generated logs, videos, texts, and other large-data artifacts that are processed and resulted in an explosion of useful platforms and languages (Hadoop, Spark, Pandas, etc.). However, translating this work from the enterprise space to the computational science and HPC community has proven somewhat difficult, in part because of some of the fundamental differences in type and scale of data and timescales surrounding its generation and use. We describe a forward-looking research and development plan which centers around the concept of making Input/Output (I/O) intelligent for users in the scientific community, whether they are accessing scalable storage or performing in situ workflow tasks. Much of our work is based on our experience with the Adaptable I/O System (ADIOS 1.X), and our next generation version of the software ADIOS 2.X [1]. Scott Klasky, Matthew Wolf, Mark Ainsworth, Chuck Atkins, Jong Choi 0001, Greg Eisenhauer, Berk Geveci, William F. Godoy, Mark Kim, James Kress, Tahsin M. Kurç, Qing Liu 0002, Jeremy Logan, Arthur B. Maccabe, Kshitij Mehta, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Eric Suchyta, Lipeng Wan 0001 |
ICDCS | 17 |
| 2018 | Edge Enhanced Deep Learning System for Large-Scale Video Stream AnalyticsabstractApplying deep learning models to large-scale IoT data is a compute-intensive task and needs significant computational resources. Existing approaches transfer this big data from IoT devices to a central cloud where inference is performed using a machine learning model. However, the network connecting the data capture source and the cloud platform can become a bottleneck. We address this problem by distributing the deep learning pipeline across edge and cloudlet/fog resources. The basic processing stages and trained models are distributed towards the edge of the network and on in-transit and cloud resources. The proposed approach performs initial processing of the data close to the data source at edge and fog nodes, resulting in significant reduction in the data that is transferred and stored in the cloud. Results on an object recognition scenario show 71\% efficiency gain in the throughput of the system by employing a combination of edge, in-transit and cloud resources when compared to a cloud-only approach. Muhammad K. Ali, Ashiq Anjum, Muhammad Usman Yaseen, Ali Reza Zamani, Daniel Balouek-Thomert, Omer F. Rana, Manish Parashar |
ICFEC | 7 |
| 2018 | Scalable Data Resilience for In-memory Data StagingabstractThe dramatic increase in the scale of current and planned high-end HPC systems is leading new challenges, such as the growing costs of data movement and IO, and the reduced mean times between failures (MTBF) of system components. In-situ workflows, i.e., executing the entire application workflows on the HPC system, have emerged as an attractive approach to address data-related challenges by moving computations closer to the data, and staging-based frameworks have been effectively used to support in-situ workflows at scale. However, the resilience of these staging-based solutions has not been addressed and they remain susceptible to expensive data failures. Furthermore, naive use of data resilience techniques such as n-way replication and erasure codes can impact latency and/or result in significant storage overheads. In this paper, we present CoREC, a scalable resilient in-memory data staging runtime for large-scale in-situ workflows. CoREC uses a novel hybrid approach that combines dynamic replication with erasure coding based on data access patterns. The paper also presents optimizations for load balancing and conflict avoiding encoding, and a low overhead, lazy data recovery scheme. We have implemented the CoREC runtime and have deployed with the DataSpaces staging service on Titan at ORNL, and present an experimental evaluation in the paper. The experiments demonstrate that CoREC can tolerate in-memory data failures while maintaining low latency and sustaining high overall storage efficiency at large scales. Shaohua Duan, Pradeep Subedi, Keita Teranishi, Philip E. Davis, Hemanth Kolla, Marc Gamell, Manish Parashar |
IPDPS | 7 |
| 2018 | Exploring Power Budget Scheduling Opportunities and Tradeoffs for AMR-Based ApplicationsabstractComputational demand has brought major changes to Advanced Cyber-Infrastructure (ACI) architectures. It is now possible to run scientific simulations faster and obtain more accurate results. However, power and energy have become critical concerns. Also, the current roadmap toward the new generation of ACI includes power budget as one of the main constraints. Current research efforts have studied power and performance tradeoffs and how to balance these (e.g., using Dynamic Voltage and Frequency Scaling (DVFS) and power capping for meeting power constraints, which can impact performance). However, applications may not tolerate degradation in performance, and other tradeoffs need to be explored to meet power budgets (e.g., involving the application in making energy-performance-quality tradeoff decisions). This paper proposes using the properties of AMR-based algorithms (e.g., dynamically adjusting the resolution of a simulation in combination with power capping techniques) to schedule or re-distribute the power budget. It specifically explores the opportunities to realize such an approach using checkpointing as a proof-of-concept use case and provides a characterization of a representative set of applications that use Adaptive Mesh Refinement (AMR) methods, including a Low-Mach-Number Combustion (LMC) application. It also explores the potential of utilizing power capping to understand power-quality tradeoffs via simulation. Yubo Qin, Ivan Rodero, Pradeep Subedi, Manish Parashar, Sandro Rigo |
SBAC-PAD | 4 |
| 2018 | Runtime Management of Data Quality for Scientific Observatories Using Edge and In-Transit ResourcesabstractModern Cyberinfrastructures (CIs) operate to bring content produced from remote data sources such as sensors and scientific instruments and deliver it to end users and workflow applications. Maintaining data quality/resolution and on-time data delivery while considering an increasing number of computing, storage and network resources requires a reactive system, able to adapt to changing demands. In this paper, we propose a modelization of such system by expressing the dynamic stage of resources in the context of edge and in-transit computing. By considering resource utilization, approximation techniques and users' constraints, our proposed engine is generating mappings of workflow stages on heterogeneous geo-distributed resources. We specifically propose a runtime management layer that adapts the data resolution being delivered to the users by implementing feedback loops over the resources involved in the delivery and processing of the data streams. We implement our model into a subscription-based data streaming framework which enables integration of large facilities and advanced CIs. Experimental results show that dynamically adapting data resolution can overcome bandwidth limitation in wide area streaming analytics. Ali Reza Zamani, Daniel Balouek-Thomert, Juan J. Villalobos, Ivan Rodero, Manish Parashar |
SBAC-PAD | 5 |
| 2018 | Stacker: an autonomic data movement engine for extreme-scale data staging-based in-situ workflows
Pradeep Subedi, Philip E. Davis, Shaohua Duan, Scott Klasky, Hemanth Kolla, Manish Parashar |
SC | 6 |
| 2018 | End-to-end energy models for Edge Cloud-based IoT platforms: Application to data stream analysis in IoT
Yunbo Li, Anne-Cécile Orgerie, Ivan Rodero, Betsegaw Lemma Amersho, Manish Parashar, Jean-Marc Menaud |
Future Gener. Comput. Syst. | 5 |
| 2018 | A computational model to support in-network data analysis in federated ecosystems
Ali Reza Zamani, Mengsong Zou, Javier Diaz Montes, Ioan Petri, Omer F. Rana, Manish Parashar |
Future Gener. Comput. Syst. | 6 |
| 2018 | Farewell EditorialabstractNo abstract available. Manish Parashar, Franco Zambonelli |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2018 | Supporting Data-Intensive Workflows in Software-Defined Federated Multi-CloudsabstractCloud computing is emerging as a viable platform for scientific exploration. Elastic and on-demand access to resources (and other services), the abstraction of “unlimited” resources, and attractive pricing models provide incentives for scientists to move their workflows into clouds. Generalizing these concepts beyond a single virtualized datacenter, it is possible to create federated marketplaces where different types of resources (e.g., clouds, HPC grids, supercomputers) that may be geographically distributed, are collectively exposed as a single elastic infrastructure. This presents opportunities for optimizing the execution of application workflows with heterogeneous and dynamic requirements, and tackling larger scale problems. In this paper, we introduce a framework to manage the end-to-end execution of data-intensive application workflows in dynamic software-defined resource federation. This framework enables the autonomic execution of workflows by elastically provisioning an appropriate set of resources that meet application requirements, and by adapting this set of resources at runtime as the requirements change. It also allows users to customize scheduling policies that drive the way resources federated and used. To demonstrate the benefits of our approach, we study the execution of two different data-intensive scientific workflows in a multi-cloud federation using different policies and objective functions. Javier Diaz Montes, Manuel Diaz-Granados, Mengsong Zou, Shu Tao, Manish Parashar |
IEEE Trans. Cloud Comput. | 5 |
| 2018 | State of the JournalabstractWelcome to 2018's first issue of the IEEE Transactions on Parallel and Distributed Systems (TPDS). I'm excited with my new role as incoming Editor-in-Chief (EIC) of TPDS and look forward to serving the community over the next few years. My goal as EIC is to continue to work on increasing the visibility and relevance and impact of TPDS, as well as the quality and timeliness of the review process, to ensure that TPDS is the premier Transactions in the field. The IEEE is a hallmark of quality for technical publication. The value TPDS brings to the international community is in its collection of the highest quality research that is relevant to academia, industry, and laboratories. I will investigate new opportunities for TPDS to capture the best research while maintaining its emphasis on highest quality papers. TPDS also needs to respond to a dynamic and rapidly evolving research and publication landscape. As EIC, I will work with the IEEE community to ensure that TPDS does respond appropriately, and will carefully work with the editorial board to revisit the scope and recruit new editorial board members as needed. An important and rapid growing conversation is related to the repeatability of published research and the submission of supplementary material such as code and data. I am part of this conversation, and will work with the EB and the community to explore how to bring these practices in meaningful and measured ways to TPDS. Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | An Unsupervised Approach for Online Detection and Mitigation of High-Rate DDoS Attacks Based on an In-Memory Distributed Graph Using Streaming Data and AnalyticsabstractA Distributed Denial of Service (DDoS) attack is an attempt to make an online service, a network, or even an entire organization, unavailable by saturating it with traffic from multiple sources. DDoS attacks are among the most common and most devastating threats that network defenders have to watch out for. DDoS attacks are becoming bigger, more frequent, and more sophisticated. Volumetric attacks are the most common types of DDoS attacks. A DDoS attack is considered volumetric, or high-rate, when within a short period of time it generates a large amount of packets or a high volume of traffic. High-rate attacks are well-known and have received much attention in the past decade; however, despite several detection and mitigation strategies have been designed and implemented, high-rate attacks are still halting the normal operation of information technology infrastructures across the Internet when the protection mechanisms are not able to cope with the aggregated capacity that the perpetrators have put together. With this in mind, the present paper aims to propose and test a distributed and collaborative architecture for online high-rate DDoS attack detection and mitigation based on an in-memory distributed graph data structure and unsupervised machine learning algorithms that leverage real-time streaming data and analytics. We have successfully tested our proposed mechanism using a real-world DDoS attack dataset at its original rate in pursuance of reproducing the conditions of an actual large scale attack. Juan J. Villalobos, Ivan Rodero, Manish Parashar |
BDCAT | 3 |
| 2017 | Towards Distributed Software-Defined EnvironmentsabstractService-based access models coupled with recent advances in application deployment technologies can support emerging dynamic and data-driven applications. However, due to evolving application requirements and the dynamicity of the underlying resources, it is necessary to support flexible and opportunistic composition of services in order to satisfy application needs. The goal of this work is to provide a programmable and dynamic framework that can support these applications. The framework uses software-defined environment concepts to drive the process of dynamically composing infrastructure services from multiple providers. The resulting distributed software-defined environment autonomously evolves over the application life cycle while meeting objectives and constraints set by the users, applications, and/or resource providers. We present two different approaches for programming resources and controlling the composition process, one that is based on a rule engine and another that leverages constraint programming. Preliminary results demonstrate the framework operation and performance using simulations and real experiments running Docker containers across multiple clouds. Moustafa AbdelBaky, Javier Diaz Montes, Manish Parashar |
CCGrid | 3 |
| 2017 | Enabling Distributed Software-Defined Environments Using Dynamic Infrastructure Service CompositionabstractService-based access models coupled with emerging application deployment technologies are enabling opportunities for realizing highly customized software-defined environments, which can support dynamic and data-driven applications. However, this requires rethinking traditional resource federation models to support dynamic resource compositions, which can adapt to evolving application needs and the dynamic state of underlying resources. In this paper, we present a programmable approach that leverages software-defined techniques to create a dynamic space-time infrastructure service composition. We propose the use of Constraint Programming as a formal language to allow users, applications, and service providers to define the desired state of the execution environment. The resulting distributed software-defined environment continually adapts to meet objectives/constraints set by the users, applications, and/or resource providers. We present the design and prototype implementation of such distributed software-defined environment. We use a cancer informatics workflow to demonstrate the operation of our framework using resources from five different cloud providers, which are aggregated on-demand based on dynamic user and resource provider constraints. Moustafa AbdelBaky, Javier Diaz Montes, Merve Unuvar, Melissa Romanus, Ivan Rodero, Malgorzata Steinder, Manish Parashar |
CCGrid | 7 |
| 2017 | Leveraging Renewable Energy in Edge Clouds for Data Stream Analysis in IoTabstractThe emergence of Internet of Things (IoT) is participating to the increase of data-and energy-hungry applications. As connected devices do not yet offer enough capabilities for sustaining these applications, users perform computation offloading to the cloud. To avoid network bottlenecks and reduce the costs associated to data movement, edge cloud solutions have started being deployed, thus improving the Quality of Service. In this paper, we advocate for leveraging on-site renewable energy production in the different edge cloud nodes to green IoT systems while offering improved QoS compared to core cloud solutions. We propose an analytic model to decide whether to offload computation from the objects to the edge or to the core Cloud, depending on the renewable energy availability and the desired application QoS. This model is validated on our application use-case that deals with video stream analysis from vehicle cameras. Yunbo Li, Anne-Cécile Orgerie, Ivan Rodero, Manish Parashar, Jean-Marc Menaud |
CCGrid | 4 |
| 2017 | TGE: Machine Learning Based Task Graph Embedding for Large-Scale Topology MappingabstractTask mapping is an important problem in parallel and distributed computing. The goal in task mapping is to find an optimal layout of the processes of an application (or a task) onto a given network topology. We target this problem in the context of staging applications. A staging application consists of two or more parallel applications (also referred to as staging tasks) which run concurrently and exchange data over the course of computation. Task mapping becomes a more challenging problem in staging applications, because not only data is exchanged between the staging tasks, but also the processes of a staging task may exchange data with each other. We propose a novel method, called Task Graph Embedding (TGE), that harnesses the observable graph structures of parallel applications and network topologies. TGE employs a machine learning based algorithm to find the best representation of a graph, called an embedding, onto a space in which the task-to-processor mapping problem can be solved. We evaluate and demonstrate the effectiveness of TGE experimentally with the communication patterns extracted from runs of XGC, a large-scale fusion simulation code, on Titan. Jong Choi 0001, Jeremy Logan, Matthew Wolf, George Ostrouchov, Tahsin M. Kurç, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Melissa Romanus, Manish Parashar, Michael Churchill, Choong-Seock Chang |
CLUSTER | 11 |
| 2017 | Online Decision-Making Using Edge Resources for Content-Driven Stream ProcessingabstractThe Internet of Things (IoT) describes the emerging paradigm that connects sensors, often located at the edge of the network, to stream processing engines located at the core of the network to enable online data-driven monitoring, management, and control. As IoT applications require increasing volumes of streaming data to be processed by complex workflows in a timely manner, it is becoming important to also leverage resources closer to the edge. Furthermore, the topology of these workflows and where theyare executed is determined not only by application objectives and available resources, but also by the content of the data streams, however, current stream processing engines do not provide this flexibility. In this paper, we present a programming framework that enables applications to specify data-driven, location- and resource-aware processing of data streams. Specifically, it provides abstractions for specifying where and how a data stream is processed based on its content, spatial and temporal characteristics. We also present an implementation of the framework using an event-driven runtime, where events are associatively described. Finally, we demonstrate the effectiveness of the solution by an evaluation of scalability and performance using a disaster response application usecase. Eduard Gibert Renart, Daniel Balouek-Thomert, Manish Parashar |
eScience | 5 |
| 2017 | Supporting Data-Driven Workflows Enabled by Large Scale ObservatoriesabstractLarge scale observatories are shared-use resources that provide open access to data from geographically distributed sensors and instruments. This data has the potential to accelerate scientific discovery. However, seamlessly integrating the data into scientific workflows remains a challenge. In this paper, we summarize our ongoing work in supporting data-driven and data-intensive workflows and outline our vision for how these observatories can improve large-scale science. Specifically, we present programming abstractions and runtime management services to enable the automatic integration of data in scientific workflows. Further, we show how approximation techniques can be used to address network and processing variations by studying constraint limitations and their associated latencies. We use the Ocean Observatories Initiative (OOI) as a driving use case for this work. Ali Reza Zamani, Moustafa AbdelBaky, Daniel Balouek-Thomert, Ivan Rodero, Manish Parashar |
eScience | 5 |
| 2017 | Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo |
Euro-Par | 21 |
| 2017 | Computing in the Continuum: Combining Pervasive Devices and Services to Support Data-Driven ApplicationsabstractThe exponential growth of digital data sources has the potential to transform all aspects of society and our lives. However, to achieve this impact, the data has to be processed promptly to extract insights that can drive decision making. Further, traditional approaches that rely on moving data to remote data centers for processing are no longer feasible. Instead, new approaches that effectively leverage distributed computational infrastructure and services are necessary. Specifically, these approaches must seamlessly combine resources and services at the edge, in the core, and along the data path as needed. This paper presents our vision for enabling an approach for computing in the continuum, i.e., realizing a fluid ecosystem where distributed resources and services are programmatically aggregated on-demand to support emerging data-driven application workflows. This vision calls for novel solutions for federating infrastructure, programming applications and services, and composing dynamic workflows, which are capable of reacting in real-time to unpredictable data sizes, availabilities, locations, and rates. Moustafa AbdelBaky, Mengsong Zou, Ali Reza Zamani, Eduard Gibert Renart, Javier Diaz Montes, Manish Parashar |
ICDCS | 6 |
| 2017 | Exacution: Enhancing Scientific Data Management for ExascaleabstractAs we continue toward exascale, scientific data volume is continuing to scale and becoming more burdensome to manage. In this paper, we lay out opportunities to enhance state of the art data management techniques. We emphasize well-principled data compression, and using it to achieve progressive refinement. This can both accelerate I/O and afford the user increased flexibility when she interacts with the data. The formulation naturally maps onto enabling partitioning of the progressively improving-quality representations of a data quantity into different media-type destinations, to keep the highest priority information as close as possible to the computation, and take advantage of deepening memory/storage hierarchies in ways not previously possible. Careful monitoring is requisite to our vision, not only to verify that compression has not eliminated salient features in the data, but also to better understand the performance of massively parallel scientific applications. Increased mathematical rigor would be ideal,to help bring compression on a better-understood theoretical footing, closer to the relevant scientific theory, more aware of constraints imposed by the science, and more tightly error-controlled. Throughout, we highlight pathfinding research we have begun exploring related these topics, and comment toward future work that will be needed. Scott Klasky, Eric Suchyta, Mark Ainsworth, Qing Liu 0002, Ben Whitney, Matthew Wolf, Jong Choi 0001, Ian T. Foster, Mark Kim, Jeremy Logan, Kshitij Mehta, Todd S. Munson, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Lipeng Wan 0001 |
ICDCS | 14 |
| 2017 | Data-Driven Stream Processing at the EdgeabstractThe popularity and proliferation of the Internet of Things (IoT) paradigm is resulting in a growing number of devices connected to the Internet. These devices are generating and consuming unprecedented amounts of data at the edges of the infrastructure, and are enabling new classes of data-drivenapplications, however, current approaches typically rely on cloud platforms located at the core of the infrastructure to process data. As the number of devices and the amount of data they generate and consume increases, such core-centric approaches are becoming increasingly inefficient as they need to transfer data back and forth between the edge and the core. Furthermore, the latencies associated with such data transfer may not be able to support applications involving time-critical data-driven decision making. In this paper, we propose an edge-based programming framework that allows users to define how data streams are processed based on the content and the location of the data. This enables the definition of data-driven reactive behaviors that can effectively exploit data patterns to dynamically drive stream processing, leveraging resources located at the edges of the infrastructure. We have implemented a prototype of the proposed approach and performed several experiments to evaluate its scalability and efficiency against a more typical single-cloud approach. Using a smart-city application usecase, we illustrate that the presented programming system can support data-driven stream processing using edge resources. In terms of scalability, our experiments show that the system can scale to hundreds of nodes while keeping overheads low. Our experiments also show that our approach can perform up to 56% better than a single cloud approach that does not consider data and user locality. Eduard Gibert Renart, Javier Diaz Montes, Manish Parashar |
ICFEC | 3 |
| 2017 | WA-Dataspaces: Exploring the Data Staging Abstractions for Wide-Area Distributed Scientific WorkflowsabstractData staging has been shown to be very effective for supporting data intensive in-situ workflows and coupling of applications. Experimental sciences are increasingly becoming collaborative among geographically distributed teams, and include experimental instruments and HPC facilities. This new way of doing science poses new challenges due to data sizes, complexity of computation, and the use of wide area networks between couplings. In this paper, we explore how the staging abstraction can be extended to support such workflows. Specifically, we develop a NUMA-like abstraction that orchestrates multiple distributed local-area staging abstractions, and provides asynchronous data put/get semantics to enable data sharing across them. To mask data movement overhead and provide in-time data access, we propose the use of predictive prefetching approaches that leverage the iterative nature of the coupling. We evaluate our prototype implementation using a fusion workflow and show that our design can effectively and transparently support widearea coupled workflows. Additionally, results show that the use of prefetching techniques leads to significant gains in data access times of data that needs to be moved over the wide area network. Mehmet Fatih Aktas, Javier Diaz Montes, Ivan Rodero, Manish Parashar |
ICPP | 4 |
| 2017 | Supporting Real-Time Jobs on the IBM Blue Gene/Q: Simulation-Based Study
Daihou Wang, Eun-Sung Jung, Rajkumar Kettimuthu, Ian T. Foster, David J. Foran, Manish Parashar |
JSSPP | 6 |
| 2017 | Computational resource management for data-driven applications with deadline constraintsabstractSummary Recent advances in the type and variety of sensing technologies have led to an extraordinary growth in the volume of data being produced and led to a number of streaming applications that make use of this data. Sensors typically monitor environmental or physical phenomenon at predefined time intervals or triggered by user‐defined events. Understanding how such streaming content (the raw data or events) can be processed within a time threshold remains an important research challenge. We investigate how a cloud‐based computational infrastructure can autonomically respond to such streaming content, offering quality of service guarantees. In particular, we contextualize our approach using an electric vehicles (EVs) charging scenario, where such vehicles need to connect to the electrical grid to charge their batteries. There has been an emerging interest in EV aggregators (primarily intermediate brokers able to estimate aggregate charging demand for a collection of EVs) to coordinate the charging process. We consider predicting EV charging demand as a potential workload with execution time constraints. We assume that an EV aggregator manages a number of geographic areas and a pool of computational resources of a cloud computing cluster to support scheduling of EV charging. The objective is to ensure that there is enough computational capacity to satisfy the requirements for managing EV battery charging requests within specific time constraints. Rafael Tolosana-Calasanz, Javier Diaz Montes, Omer F. Rana, Manish Parashar, Erotokritos Xydas, Charalampos E. Marmaras, Panagiotis Papadopoulos, Liana Cipcigan |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | In-memory staging and data-centric task placement for coupled scientific simulation workflowsabstractSummary Coupled scientific simulation workflows are composed of heterogeneous component applications that simulate different aspects of the physical phenomena being modeled and that interact and exchange significant volumes of data at runtime. As the data volumes and generation rates keep growing, the traditional disk I/O–based data movement approach becomes cost prohibitive, and workflow requires more scalable and efficient approach to support the data movement. Moreover, the cost of moving large volume of data over system interconnection network becomes dominating and significantly impacts the workflow execution time. Minimize the amount of network data movement and localize data transfers are critical for reducing such cost. To achieve this, workflow task placement should exploit data locality to the extent possible and move computation closer to data. In this paper, we investigate applying in‐memory data staging and data‐centric task placement to reduce the data movement cost in large‐scale coupled simulation workflows. Specifically, we present a distributed data sharing and task execution framework that (1) co‐locates in‐memory data staging on application compute nodes to store data that needs to be shared or exchanged and (2) uses data‐centric task placement to map computations onto processor cores that a large portion of the data exchanges can be performed using the intra‐node shared memory. We also present the implementation of the framework and its experimental evaluation on Titan Cray XK7 petascale supercomputer. Fan Zhang 0004, Tong Jin 0002, Melissa Romanus, Hoang Bui, Scott Klasky, Manish Parashar |
Concurr. Comput. Pract. Exp. | 7 |
| 2017 | Modeling and Simulating Multiple Failure Masking Enabled by Local Recovery for Stencil-Based Applications at Extreme ScalesabstractObtaining multi-process hard failure resilience at the application level is a key challenge that must be overcome before the promise of exascale can be fully realized. Previous work has shown that online global recovery can dramatically reduce the overhead of failures when compared to the more traditional approach of terminating the job and restarting it from the last stored checkpoint. If online recovery is performed in a local manner further scalability is enabled, not only due to the intrinsic lower costs of recovering locally, but also due to derived effects when using some application types. In this paper we model one such effect, namely multiple failure masking, that manifests when running Stencil parallel computations on an environment when failures are recovered locally. First, the delay propagation shape of one or multiple failures recovered locally is modeled to enable several analyses of the probability of different levels of failure masking under certain Stencil application behaviors. Our results indicate that failure masking is an extremely desirable effect at scale which manifestation is more evident and beneficial as the machine size or the failure rate increase. Marc Gamell, Keita Teranishi, Jackson R. Mayo, Hemanth Kolla, Michael A. Heroux, Jacqueline Chen, Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2017 | Feedback-Control & Queueing Theory-Based Resource Management for Streaming ApplicationsabstractRecent advances in sensor technologies and instrumentation have led to an extraordinary growth of data sources and streaming applications. A wide variety of devices, from smart phones to dedicated sensors, have the capability of collecting and streaming large amounts of data at unprecedented rates. A number of distinct streaming data models have been proposed. Typical applications for this include smart cites & built environments for instance, where sensor-based infrastructures continue to increase in scale and variety. Understanding how such streaming content can be processed within some time threshold remains a non-trivial and important research topic. We investigate how a cloud-based computational infrastructure can autonomically respond to such streaming content, offering Quality of Service guarantees. We propose an autonomic controller (based on feedback control and queueing theory) to elastically provision virtual machines to meet performance targets associated with a particular data stream. Evaluation is carried out using a federated Cloud-based infrastructure (implemented using CometCloud)-where the allocation of new resources can be based on: (i) differences between sites, i.e., types of resources supported (e.g., GPU versus CPU only), (ii) cost of execution; (iii) failure rate and likely resilience, etc. In particular, we demonstrate how Little's Law-a widely used result in queuing theory-can be adapted to support dynamic control in the context of such resource provisioning. Rafael Tolosana-Calasanz, Javier Diaz Montes, Omer F. Rana, Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Modelling and Implementing Social Community CloudsabstractAs the number of people who interact on social networks increases, and coupled with the greater capability made available within our computational devices, there is the potential to establish “Social Clouds”-a resource sharing infrastructure that enable people who have trust relationships to come together to share computational/ data services within a community. Social clouds can also provide the means to enhance multi-user collaboration and greatly stimulate the exchange of resources among participants. Recent research in the establishment and use of Social Clouds has raised significant interest by proposing an environment where users are able to trade resources mediated by a social networking mechanism. In such a cloud environment the incentives for sharing can represent a solution for improving resource utilisation and for making available additional capacity to friends and collaborators. In this paper we demonstrate how revenue can be earned within a social cloud community, by executing internal (intra community) and external (inter community) tasks. A number of different scenarios are first investigated through simulation, using the PeerSim simulator, in order to validate our approach. We use two key metrics: revenue and reputation, to evaluate how the system dynamics change as new tasks are added to one or more communities for execution, along with additional behaviours, such as nodes migrating from one community to another, or selectively reporting on the outcome of task execution. Subsequently, we develop a practical deployment using a federated cloud scenario using the CometCloud system-deployed over three sites: Cardiff (UK), Rutgers and Indiana. We show how approaches that have been simulated in PeerSim can be implemented in practice. Ioan Petri, Javier Diaz Montes, Omer F. Rana, Magdalena Punceva, Ivan Rodero, Manish Parashar |
IEEE Trans. Serv. Comput. | 6 |
| 2016 | D3W: Towards Self-Management of Distributed Data-Driven Workflows with QoS GuaranteesabstractData-driven application workflows that leverage compute capabilities and hosted services near the network edge can support latency-sensitive and critical applications in emerging areas such as Internet of Things (IoT) and smart infrastructure. However, distributed instantiation and execution of these workflows using resources across service providers and datacenters can be challenging. In this paper, we present the formulation of a decentralized workflow management approach for the autonomous instantiation and execution of dynamic data-driven workflows based on the opportunistic discovery and composition of services on-demand. Given a workflow template specification, this approach allows us to decouple workflow stages, allowing the execution of different stages to be performed by individual services, which are discovered and instantiated dynamically, and can be independently scaled as needed. These services may be geographically distributed and may be offered by different service providers using various QoS levels and cost models. The design, implementation and experimental evaluation of a decentralized workflow management framework using a live media stream application in a multi-cloud infrastructure is presented. Evaluations using a sample topology shows up to 2.5 times increase in QoS-meeting throughput when using our dynamic multi-cloud approach instead of using a fixed centralized cloud of identical capacity. Mengsong Zou, Javier Diaz Montes, Kiran Nagaraja, Nimish Radia, Manish Parashar |
CLOUD | 5 |
| 2016 | Dynamic Adaptation of Policies Using Machine LearningabstractManaging large systems in order to guarantee certain behavior is a difficult problem due to their dynamic behavior and complex interactions. Policies have been shown to provide a very expressive and easy way to define such desired behaviors, mainly because they separate the definition of desired behavior from the enforcement mechanism, allowing either one to be changed fairly easily. Unfortunately, it is often difficult to define policies in terms of attributes that can be measured and/or directly controlled, or to set adaptable (i.e. non-static) parameters in order to account for rapidly changing system behavior. Dynamic policies are meant to solve these problems by allowing system administrators to define higher level parameters, which are more closely related to the business goals, while providing an automated mechanism to adapt them at a lower level, where attributes can be measured and/or controlled. Here, we present a way to define such policies, and a machine learning model that is able to dynamically apply lower level static policies by learning a hidden relationship between the high level business attribute space, and the low level monitoring space. We show that this relationship exists, and that we can learn it producing an error of at most 8.78% at least 96% of the time. Alejandro Pelaez, Andres Quiroz, Manish Parashar |
CCGrid | 3 |
| 2016 | Evaluation of In-Situ Analysis Strategies at Scale for Power Efficiency and ScalabilityabstractThe increasing gap between available compute power and I/O capabilities is resulting in simulation pipelines running on leadership computing facilities being reformulated. In particular, in-situ processing is complementing conventional post-process analysis, however, it can be performed by using the same compute resources as the simulation or using secondary dedicated resources. In this paper, we focus on three different in-situ analysis strategies, which use the same compute resources as the ongoing simulation but different data movement strategies. We evaluate the costs incurred by these strategies in terms of run time, scalability and power/energy consumption. Furthermore, we extrapolate power behavior to peta-scale and investigate different design choices through projections. Experimental evaluation at full machine scale on Titan supports that using fewer cores per node for in-situ analysis is the optimum choice in terms of scalability. Hence, further research effort should be devoted towards developing in-situ analysis techniques following this strategy in future high-end systems. Ivan Rodero, Manish Parashar, Aaditya G. Landge, Sidharth Kumar, Valerio Pascucci, Peer-Timo Bremer |
CCGrid | 2 |
| 2016 | Management, analysis, and visualization of experimental and observational data - The convergence of data and computingabstractScientific user facilities — particle accelerators, telescopes, colliders, supercomputers, light sources, sequencing facilities, and more — operated by the U.S. Department of Energy (DOE) Office of Science (SC) generate ever increasing volumes of data at unprecedented rates from experiments, observations, and simulations. At the same time there is a growing community of experimentalists that require real-time data analysis feedback, to enable them to steer their complex experimental instruments to optimized scientific outcomes and new discoveries. Recent efforts in DOE-SC have focused on articulating the data-centric challenges and opportunities facing these science communities. Key challenges include difficulties coping with data size, rate, and complexity in the context of both real-time and post-experiment data analysis and interpretation. Solutions will require algorithmic and mathematical advances, as well as hardware and software infrastructures that adequately support data-intensive scientific workloads. This paper presents the summary findings of a workshop held by DOE-SC in September 2015, convened to identify the major challenges and the research that is needed to meet those challenges. E. Wes Bethel, Martin Greenwald, Kerstin Kleese van Dam, Manish Parashar, Stefan M. Wild, H. Steven Wiley |
eScience | 4 |
| 2016 | Scheduling and flexible control of bandwidth and in-transit services for end-to-end application workflows
Mehmet Fatih Aktas, Georgiana Haldeman, Manish Parashar |
Future Gener. Comput. Syst. | 3 |
| 2015 | Realizing the Potential of IoT Using Software-Defined EcosystemsabstractPervasive computational ecosystems that combine data sources and computing/communication resources in self-managed environments, such as the ones powered by Internet of Things (IoT) devices, have the potential to automate and facilitate many aspects of our lives, and impact a variety of applications, from the management of extreme events to the optimization of everyday processes. However, this vision remains mostly unrealized despite the fact that the technology to achieve it exists, largely because of the gap between our ability to collect data and our ability to gain insight from it. In this paper, we discuss the challenges associated with providing a pervasive computational ecosystem. We then present our vision of how to best support data-driven computational ecosystems and propose a conceptual architecture that leverages ideas from software-defined environments in order to combine data, computing, and communication resources. In addition, we show how this proposed architecture enables the execution of data-driven workflows on top of these resources. Manish Parashar, Moustafa AbdelBaky, Mengsong Zou, Ali Reza Zamani, Javier Diaz Montes |
CLOUD | 1 |
| 2015 | Investigating insurance fraud using social mediaabstractSince the social media hype started in the early 2000s, the Internet has bloomed with user-generated data. The content generated by users in social media varies from blogs, forums, social network platforms, and video sharing communities. This data has a special emphasis on the relationships among users of the community. As a consequence, social media data contains significant information about their creators and people around them. For this reason, law enforcement and insurance companies, among others, are starting to explore ways of using this data to identify unlawful or fraudulent activities. In this work, we present a solution to extract and analyze social media data in pursuit of identifying insurance fraud. We describe and evaluate an initial prototype of our solution that has been implemented on top of the CometCloud framework. We show how our solution is driven by the insights obtained from the data and it is able to extract data relevant to the investigators. Manuel Diaz-Granados, Javier Diaz Montes, Manish Parashar |
IEEE BigData | 3 |
| 2015 | Integrating Software Defined Networks within a Cloud FederationabstractCloud computing has generally involved the use of specialist data centres to support computation and data storage at a central site (or a limited number of sites). The motivation for this has come from the need to provide economies of scale (and subsequent reduction in cost) for supporting large scale computation for multiple user applications over (generally) a shared, multi-tenancy infrastructure. The use of such infrastructures requires moving data to a central location (data may be pre-staged to such a location prior to processing using terrestrial delivery channels and does not always require the use of a network-based transfer), undertaking processing on the data, and subsequently enabling users to download results of analysis. We extend this model using software defined networks (SDNs), whereby capability within the network can be used to support in-transit processing while data is in movement from source to destination. Using a smart building infrastructure scenario, consisting of sensors and actuators embedded within a built environment, we describe how an SDN-based architecture can be used to support real time data processing. This significantly influences the processing times to support energy optimisation of the building and reduces costs. We describe an architecture for such a distributed, multi-layered Cloud system and discuss a prototype that has been implemented using the CometCloud system, deployed across three sites in the UK and the US. Wevalidate the prototype using data from sensors within a Sports facility and making use of EnergyPlus. Ioan Petri, Mengsong Zou, Ali Reza Zamani, Javier Diaz Montes, Omer F. Rana, Manish Parashar |
CCGRID | 6 |
| 2015 | Exploring Failure Recovery for Stencil-based Applications at Extreme ScalesabstractApplication resilience is a key challenge that must be addressed in order to realize the exascale vision. Previous work has shown that online recovery, even when done in a global manner (i.e., involving all processes), can dramatically reduce the overhead of failures when compared to the more traditional approach of terminating the job and restarting it from the last stored checkpoint. In this paper we suggest going one step further, and explore how local recovery can be used for certain classes of applications to reduce the overheads due to failures. Specifically we study the feasibility of local recovery for stencil-based parallel applications and we show how multiple independent failures can be masked to effectively reduce the impact on the total time to solution. Marc Gamell, Keita Teranishi, Michael A. Heroux, Jackson R. Mayo, Hemanth Kolla, Jacqueline Chen, Manish Parashar |
HPDC | 7 |
| 2015 | Exploring Data Staging Across Deep Memory Hierarchies for Coupled Data Intensive Simulation WorkflowsabstractAs applications target extreme scales, data staging and in-situ/in-transit data processing have been proposed to address the data challenges and improve scientific discovery. However, further research is necessary in order to understand how growing data sizes from data intensive simulations coupled with the limited DRAM capacity in High End Computing systems will impact the effectiveness of this approach. In this paper, we explore how we can use deep memory levels for data staging, and develop a multi-tiered data staging method that spans bothDRAM and solid state disks (SSD). This approach allows us to support both code coupling and data management for data intensive simulation workflows. We also show how an adaptive application-aware data placement mechanism can dynamically manage and optimize data placement across the DRAM ands storage levels in this multi-tiered data staging method. We present an experimental evaluation of our approach using wolf resources: an Infiniband cluster (Sith) and a Cray XK7system (Titan), and using combustion (S3D) and fusion (XGC1) simulations. Tong Jin 0002, Fan Zhang 0004, Hoang Bui, Melissa Romanus, Norbert Podhorszki, Scott Klasky, Hemanth Kolla, Jacqueline Chen, Robert Hager, Choong-Seock Chang, Manish Parashar |
IPDPS | 12 |
| 2015 | Local recovery and failure masking for stencil-based applications at extreme scalesabstractApplication resilience is a key challenge that has to be addressed to realize the exascale vision. Online recovery, even when it involves all processes, can dramatically reduce the overhead of failures as compared to the more traditional approach where the job is terminated and restarted from the last checkpoint. In this paper we explore how local recovery can be used for certain classes of applications to further reduce overheads due to resilience. Specifically we develop programming support and scalable runtime mechanisms to enable online and transparent local recovery for stencil-based parallel applications on current leadership class systems. We also show how multiple independent failures can be masked to effectively reduce the impact on the total time to solution. We integrate these mechanisms with the S3D combustion simulation, and experimentally demonstrate (using the Titan Cray-XK7 system at ORNL) the ability to tolerate high failure rates (i.e., node failures every 5 seconds) with low overhead while sustaining performance, at scales up to 262144 cores. Marc Gamell, Keita Teranishi, Michael A. Heroux, Jackson R. Mayo, Hemanth Kolla, Jacqueline Chen, Manish Parashar |
SC | 7 |
| 2015 | Practical scalable consensus for pseudo-synchronous distributed systemsabstractThe ability to consistently handle faults in a distributed environment requires, among a small set of basic routines, an agreement algorithm allowing surviving entities to reach a consensual decision between a bounded set of volatile resources. This paper presents an algorithm that implements an Early Returning Agreement (ERA) in pseudo-synchronous systems, which optimistically allows a process to resume its activity while guaranteeing strong progress. We prove the correctness of our ERA algorithm, and expose its logarithmic behavior, which is an extremely desirable property for any algorithm which targets future exascale platforms. We detail a practical implementation of this consensus algorithm in the context of an MPI library, and evaluate both its efficiency and scalability through a set of benchmarks and two fault tolerant scientific applications. Thomas Hérault, Aurelien Bouteiller, George Bosilca, Marc Gamell, Keita Teranishi, Manish Parashar, Jack J. Dongarra |
SC | 6 |
| 2015 | Adaptive data placement for staging-based coupled scientific workflowsabstractData staging and in-situ/in-transit data processing are emerging as attractive approaches for supporting extreme scale scientific workflows. These approaches improve end-to-end performance by enabling runtime data sharing between coupled simulations and data analytics components of the workflow. However, the complex and dynamic data exchange patterns exhibited by the workflows coupled with the varied data access behaviors make efficient data placement within the staging area challenging. In this paper, we present an adaptive data placement approach to address these challenges. Our approach adapts data placement based on application-specific dynamic data access patterns, and applies access pattern-driven and location-aware mechanisms to reduce data access costs and to support efficient data sharing between the multiple workflow components. We experimentally demonstrate the effectiveness of our approach on Titan Cray XK7 using a real combustion-analyses workflow. The evaluation results demonstrate that our approach can effectively improve data access performance and overall efficiency of coupled scientific workflows. Tong Jin 0002, Melissa Romanus, Hoang Bui, Fan Zhang 0004, Hongfeng Yu 0001, Hemanth Kolla, Scott Klasky, Jacqueline Chen, Manish Parashar |
SC | 10 |
| 2015 | ActiveSpaces: Exploring dynamic code deployment for extreme scale data processingabstractSummary Managing the large volumes of data produced by emerging scientific and engineering simulations running on leadership‐class resources has become a critical challenge. The data have to be extracted off the computing nodes and transported to consumer nodes so that it can be processed, analyzed, visualized, archived, and so on. Several recent research efforts have addressed data‐related challenges at different levels. One attractive approach is to offload expensive input/output operations to a smaller set of dedicated computing nodes known as a staging area. However, even using this approach, the data still have to be moved from the staging area to consumer nodes for processing, which continues to be a bottleneck. In this paper, we investigate an alternate approach, namely moving the data‐processing code to the staging area instead of moving the data to the data‐processing code. Specifically, we describe the ActiveSpaces framework, which provides (1) programming support for defining the data‐processing routines to be downloaded to the staging area and (2) runtime mechanisms for transporting codes associated with these routines to the staging area, executing the routines on the nodes that are part of the staging area, and returning the results. We also present an experimental performance evaluation of ActiveSpaces using applications running on the Cray XT5 at Oak Ridge National Laboratory. Finally, we use a coupled fusion application workflow to explore the trade‐offs between transporting data and transporting the code required for data processing during coupling, and we characterize sweet spots for each option. Copyright © 2014 John Wiley & Sons, Ltd. Ciprian Docan, Fan Zhang 0004, Tong Jin 0002, Hoang Bui, Julian C. Cummings, Norbert Podhorszki, Scott Klasky, Manish Parashar |
Concurr. Comput. Pract. Exp. | 9 |
| 2015 | Incentivising resource sharing in social cloudsabstractSummary Social Clouds provide the capability to share resources among participants within a social network—leveraging on the trust relationships already existing between such participants. In such a system, users are able to trade resources between each other rather than make use of capability offered at a (centralized) data center. Although such an environment has significant potential for improving resource utilization and making available additional capacity that remains dormant, incentives for sharing remain an important hurdle limiting its effective. In this paper, we utilize the socioeconomic model proposed by Silvio Gesell to demonstrate how a ‘virtual currency’ can be used to incentivise sharing of resources within a ‘community’. We subsequently demonstrate, through simulations, the benefit provided to participants within such a community using a variety of economic (such as overall credits gained) and technical (number of successfully completed transactions) metrics. Further, we describe our implementation of such a Social Cloud using CometCloud. CometCloud is an autonomic computing engine for cloud and grid environments. It supports highly heterogeneous and dynamic federated cloud/Grid infrastructures, integration of public/private clouds and autonomic cloudbursts. We demonstrate the implementation of two designs on the basis of the master/worker approach: (i) one tuple space per cluster and (ii) one coordination tuple space and multiple transient spaces—one per each cluster. Finally, we discuss an extended version of our Social Cloud model where intermediary relay nodes take on more active roles as traders in a transaction. Copyright © 2013 John Wiley & Sons, Ltd. Magdalena Punceva, Ivan Rodero, Manish Parashar, Omer F. Rana, Ioan Petri |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Market Models for Federated CloudsabstractMulti-cloud systems have enabled resource and service providers to co-exist in a market where the relationship between clients and services depends on the nature of an application and can be subject to a variety of different Quality of Service (QoS) constraints. Deciding whether a cloud provider should host (or finds it profitable to host) a service in the long-term would be influenced by parameters such as the service price, the QoS guarantees required by customers, the deployment cost (taking into account both cost of resource provisioning and operational expenditure, e.g. energy costs) and the constraints over which these guarantees should be met. In a federated cloud system users can combine specialist capabilities offered by a limited number of providers, at particular cost bands-such as availability of specialist co-processors and software libraries. In addition, federation also enables applications to be scaled on-demand and restricts lock in to the capabilities of a particular provider. We devise a market model to support federated clouds and investigate its efficiency in two real application scenarios:(i) energy optimisation in built environments and (ii) cancer image processing both requiring significant computational resources to execute simulations. We describe and evaluate the establishment of such an application based federation and identify a cost-decision based mechanism to determine when tasks should be outsourced to external sites in the federation. The following contributions are provided: (i) understanding the criteria for accessing sites within a federated cloud dynamically, taking into account factors such as performance, cost, user perceived value, and specific application requirements; (ii) developing and deploying a cost based federated cloud framework for supporting real applications over three federated sites at Cardiff (UK), Rutgers and Indiana (USA), (iii) a performance analysis of the application scenarios to determine how task submission could be supported across these three sites, subject to particular revenue targets. Ioan Petri, Javier Diaz Montes, Mengsong Zou, Thomas H. Beach, Omer F. Rana, Manish Parashar |
IEEE Trans. Cloud Comput. | 6 |
| 2015 | Governance Model for Cloud Computing in Building Information ManagementabstractThe AEC (Architecture Engineering and Construction) sector is a highly fragmented, data intensive, project based industry, involving a number or very different professions and organisations. The industry's strong data sharing and processing requirements means that the management of building data is complex and challenging. We present a data sharing capability utilising Cloud Computing, with two key contributions: 1) a governance model for building data, based on extensive research Pand industry consultation. This governance model describes how individual data artefacts within a building information model relate to each other and how access to this data is controlled; 2) a prototype implementation of this governance model, utilising the CometCloud autonomic cloud computing engine, using the Master/Work paradigm. This prototype is able to successfully store and manage building data, provide security based on a defined policy language and demonstrate scale-out in case of increasing demand or node failure. Our prototype is evaluated both qualitatively and quantitatively. To enable this evaluation we have integrated our prototype with the 3D modelling software-Google Sketchup. We also evaluate the prototype's performance when scaling to utilise additional nodes in the Cloud and to determine its performance in case of node failures. Thomas H. Beach, Omer F. Rana, Yacine Rezgui, Manish Parashar |
IEEE Trans. Serv. Comput. | 4 |
| 2014 | Data-Driven Workflows in Multi-cloud MarketplacesabstractCloud computing is emerging as a viable platform for scientific exploration. The ideas of on-demand access to resources, "unlimited" resources as well as interesting pricing models are making scientist to move their workflows into cloud computing. However, the amount of services and different pricing models offered by the providers often overwhelm users when deciding which option is best for them. Moreover, interoperability across providers remains an open topic that forces users to develop specific solutions for each provider. In this paper, we present a service framework that enables the autonomic execution of dynamic workflows in multi-cloud environments. It also allows users to customize scheduling policies to use those resources that best fit their needs. To demonstrate the benefits of this framework, we study the execution of a real scientific workflow, with data dependencies across stages, in a multi-cloud federation using different policies and objective functions. Javier Diaz Montes, Mengsong Zou, Shu Tao, Manish Parashar |
IEEE CLOUD | 5 |
| 2014 | Cloud Supported Building Data AnalyticsabstractWith increasing availability of instrumented infrastructures in built environments, it is necessary to understand how such data will be stored, processed and analysed in a timely manner. Many "smart cities" applications, for instance, identify how data from building sensors can be combined together to support applications such as emergency response, energy management, etc. Enabling sensor data to be transmitted to a Cloud environment for processing provides a number of benefits, such as scalability and elastic provisioning of computational resources - as the total data size may not be known apriori. In this application-based case study, we describe the integration of an in-building sensor network (both for sensing and actuation) with a distributed Cloud environment. Energy optimisation in buildings represents a class of problems that requires significant computational resources and generally is a time consuming process. We describe the use of Cloud computing for efficiently running and deploying Energy Plus simulations with sensor data in order to fulfil a number of energy related objectives for buildings. We describe and evaluate the establishment of such a sensor based application using a Comet Cloud implementation with data collection from a real building pilot. Although our focus is on a single application, the general architecture and analysis carried out can be generalised to other similar scenarios. Ioan Petri, Omer F. Rana, Yacine Rezgui, Haijiang Li, Thomas H. Beach, Mengsong Zou, Javier Diaz Montes, Manish Parashar |
CCGRID | 8 |
| 2014 | Leveraging deep memory hierarchies for data staging in coupled data-intensive simulation workflowsabstractNext generation in-situ/in-transit data processing has been proposed for addressing data challenges at extreme scales. However, further research is necessary in order to understand how growing data sizes from data intensive simulations coupled with limited DRAM capacity in High End Computing clusters will impact the effectiveness of this approach. In this work, we propose using deep memory levels for data staging, utilizing a multi-tiered data staging method with both DRAM and solid state disk (SSD). This approach allows us to support both code coupling and data management for data intensive simulations in cluster environment. We also show how an application-aware data placement mechanism can dynamically manage and optimize data placement across DRAM and SSD storage levels in staging method. We present experimental results on Sith - an Infiniband cluster at Oak Ridge, and evaluate its performance using combustion (S3D) and fusion (XGC) simulations. Tong Jin 0002, Fan Zhang 0004, Hoang Bui, Norbert Podhorszki, Scott Klasky, Hemanth Kolla, Jacqueline Chen, Robert Hager, Choong-Seock Chang, Manish Parashar |
CLUSTER | 11 |
| 2014 | Exploring HPC-based scientific software as a service using CometCloudabstractThe use of in-silico simulations in experimental science can greatly increase laboratory efficiency and provide additional insights into interactions not easily described by traditional methods. Such simulations require significant amounts of computational resources, accessible only via supercom Moustafa AbdelBaky, Javier Diaz Montes, Michael Johnston, Vipin Sachdeva, Richard L. Anderson 0003, Kirk E. Jordan, Manish Parashar |
CollaborateCom | 7 |
| 2014 | Online failure prediction for HPC resources using decentralized clusteringabstractEnsuring high reliability of large-scale clusters is becoming more critical as the size of these machines continues to grow, since this increases the complexity and amount of interactions between different nodes and thus results in a high failure frequency. For this reason, predicting node failures in order to prevent errors from happening in the first place has become extremely valuable. A common approach for failure prediction is to analyze traces of system events to find correlations between event types or anomalous event patterns and node failures, and to use the types or patterns identified as failure predictors at run-time. However, typical centralized solutions for failure prediction in this manner suffer from high transmission and processing overheads at very large scales. We present a solution to the problem of predicting compute node soft-lockups in large scale clusters by using a decentralized online clustering algorithm (DOC) to detect anomalies in resource usage logs, which have been shown to correlate to particular types of node failures in supercomputer clusters. We demonstrate the effectiveness of this system by using the monitoring logs from the Ranger supercomputer at Texas Advanced Computing Center. Experiments shows that this approach can achieve similar accuracy as other related approaches, while maintaining low RAM and bandwidth usage, with a runtime impact to current running applications of less than 2%. Alejandro Pelaez, Andres Quiroz, James C. Browne, Edward Chuah, Manish Parashar |
HiPC | 5 |
| 2014 | Exploring Models and Mechanisms for Exchanging Resources in a Federated CloudabstractOne of the key benefits of Cloud systems is their ability to provide elastic, on-demand (seemingly infinite) computing capability and performance for supporting service delivery. With the resource availability in single data centres proving to be limited, the option of obtaining extra-resources from a collection of Cloud providers has appeared as an efficacious solution. The ability to utilize resources from multiple Cloud providers is also often mentioned as a means to: (i) prevent vendor lock in, (ii) to enable in house capacity to be combined with an external Cloud provider, (iii) combine specialist capability from multiple Cloud vendors (especially when one vendor does not offer such capability or where such capability may come at a higher price). Such federation of Cloud systems can therefore overcome a limit in capacity and enable providers to dynamically increase the availability of resources to serve requests. We describe and evaluate the establishment of such a federation using a CometCloud based implementation, and consider a number of federation policies with associated scenarios and determine the impact of such policies on the overall status of our system. CometCloud provides an overlay that enables multiple types of Cloud systems (both public and private) to be federated through the use of specialist gateways. We describe how two physical sites, in the UK and the US, can be federated in a seamless way using this system. Ioan Petri, Thomas H. Beach, Mengsong Zou, Javier Diaz Montes, Omer F. Rana, Manish Parashar |
IC2E | 6 |
| 2014 | Exploring Automatic, Online Failure Recovery for Scientific Applications at Extreme ScalesabstractApplication resilience is a key challenge that must be addressed in order to realize the exascale vision. Process/node failures, an important class of failures, are typically handled today by terminating the job and restarting it from the last stored checkpoint. This approach is not expected to scale to exascale. In this paper we present Fenix, a framework for enabling recovery from process/node/blade/cabinet failures for MPI-based parallel applications in an online (i.e., Without disrupting the job) and transparent manner. Fenix provides mechanisms for transparently capturing failures, re-spawning new processes, fixing failed communicators, restoring application state, and returning execution control back to the application. To enable automatic data recovery, Fenix relies on application-driven, diskless, implicitly coordinated check pointing. Using the S3D combustion simulation running on the Titan Cray-XK7 production system at ORNL, we experimentally demonstrate Felix's ability to tolerate high failure rates (e.g., More than one per minute) with low overhead while sustaining performance. Marc Gamell, Daniel S. Katz, Hemanth Kolla, Jacqueline Chen, Scott Klasky, Manish Parashar |
SC | 6 |
| 2014 | Content-based histopathology image retrieval using CometCloudabstractBACKGROUND: The development of digital imaging technology is creating extraordinary levels of accuracy that provide support for improved reliability in different aspects of the image analysis, such as content-based image retrieval, image segmentation, and classification. This has dramatically increased the volume and rate at which data are generated. Together these facts make querying and sharing non-trivial and render centralized solutions unfeasible. Moreover, in many cases this data is often distributed and must be shared across multiple institutions requiring decentralized solutions. In this context, a new generation of data/information driven applications must be developed to take advantage of the national advanced cyber-infrastructure (ACI) which enable investigators to seamlessly and securely interact with information/data which is distributed across geographically disparate resources. This paper presents the development and evaluation of a novel content-based image retrieval (CBIR) framework. The methods were tested extensively using both peripheral blood smears and renal glomeruli specimens. The datasets and performance were evaluated by two pathologists to determine the concordance. RESULTS: The CBIR algorithms that were developed can reliably retrieve the candidate image patches exhibiting intensity and morphological characteristics that are most similar to a given query image. The methods described in this paper are able to reliably discriminate among subtle staining differences and spatial pattern distributions. By integrating a newly developed dual-similarity relevance feedback module into the CBIR framework, the CBIR results were improved substantially. By aggregating the computational power of high performance computing (HPC) and cloud resources, we demonstrated that the method can be successfully executed in minutes on the Cloud compared to weeks using standard computers. CONCLUSIONS: In this paper, we present a set of newly developed CBIR algorithms and validate them using two different pathology applications, which are regularly evaluated in the practice of pathology. Comparative experimental results demonstrate excellent performance throughout the course of a set of systematic studies. Additionally, we present and evaluate a framework to enable the execution of these algorithms across distributed resources. We show how parallel searching of content-wise similar images in the dataset significantly reduces the overall computational time to ensure the practical utility of the proposed CBIR algorithms. Xin Qi 0007, Daihou Wang, Ivan Rodero, Javier Diaz Montes, Rebekah H. Gensure, Fuyong Xing, Lauri A. Goodell, Manish Parashar, David J. Foran, Lin Yang 0002 |
BMC Bioinform. | 9 |
| 2014 | Hello ADIOS: the challenges and lessons of developing leadership class I/O frameworksabstractSUMMARY Applications running on leadership platforms are more and more bottlenecked by storage input/output (I/O). In an effort to combat the increasing disparity between I/O throughput and compute capability, we created Adaptable IO System (ADIOS) in 2005. Focusing on putting users first with a service oriented architecture, we combined cutting edge research into new I/O techniques with a design effort to create near optimal I/O methods. As a result, ADIOS provides the highest level of synchronous I/O performance for a number of mission critical applications at various Department of Energy Leadership Computing Facilities. Meanwhile ADIOS is leading the push for next generation techniques including staging and data processing pipelines. In this paper, we describe the startling observations we have made in the last half decade of I/O research and development, and elaborate the lessons we have learned along this journey. We also detail some of the challenges that remain as we look toward the coming Exascale era. Copyright © 2013 John Wiley & Sons, Ltd. Qing Liu 0002, Jeremy Logan, Yuan Tian 0004, Hasan Abbasi, Norbert Podhorszki, Jong Choi 0001, Scott Klasky, Roselyne Tchoua, Jay F. Lofstead, Ron A. Oldfield, Manish Parashar, Nagiza F. Samatova, Karsten Schwan, Arie Shoshani, Matthew Wolf, Kesheng Wu, Weikuan Yu |
Concurr. Comput. Pract. Exp. | 11 |
| 2014 | Special issue on computational financeabstractSpecial issue Ruppa K. Thulasiram, Manish Parashar |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | XpressSpace: a programming framework for coupling partitioned global address space simulation codesabstractSUMMARY Complex coupled multiphysics simulations are playing increasingly important roles in scientific and engineering applications such as fusion, combustion, and climate modeling. At the same time, extreme scales, increased levels of concurrency, and the advent of multicores are making programming of high‐end parallel computing systems on which these simulations run challenging. Although partitioned global address space (PGAS) languages attempt to address the problem by providing a shared memory abstraction for parallel processes within a single program, the PGAS model does not easily support data coupling across multiple heterogeneous programs, which is necessary for coupled multiphysics simulations. This paper explores how multiphysics‐coupled simulations can be supported by the PGAS programming model. Specifically, in this paper, we present the design and implementation of the XpressSpace programming system, which extends existing PGAS data sharing and data access models with a semantically specialized shared data space abstraction to enable data coupling across multiple independent PGAS executables. XpressSpace supports a global‐view style programming interface that is consistent with the PGAS memory model, and provides an efficient runtime system that can dynamically capture the data decomposition of global‐view data‐structures such as arrays, and enable fast exchange of these distributed data‐structures between coupled applications. In this paper, we also evaluate the performance and scalability of a prototype implementation of XpressSpace by using different coupling patterns extracted from real world multiphysics simulation scenarios, on the Jaguar Cray XT5 system at Oak Ridge National Laboratory. Copyright © 2013 John Wiley & Sons, Ltd. Fan Zhang 0004, Ciprian Docan, Hoang Bui, Manish Parashar, Scott Klasky |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | Guest Editors' Introduction: Special Issue on Cloud of Cloudsabstract10.1109/TC.2014.3 Bharadwaj Veeravalli, Manish Parashar |
IEEE Trans. Computers | 2 |
| 2013 | ADIOS Visualization Schema: A First Step Towards Improving Interdisciplinary Collaboration in High Performance ComputingabstractScientific communities have benefitted from a significant increase of available computing and storage resources in the last few decades. For science projects that have access to leadership scale computing resources, the capacity to produce data has been growing exponentially. Teams working on such projects must now include, in addition to the traditional application scientists, experts in various disciplines including applied mathematicians for development of algorithms, visualization specialists for large data, and I/O specialists. Sharing of knowledge and data is becoming a requirement for scientific discovery, providing useful mechanisms to facilitate this sharing is a key challenge for e-Science. Our hypothesis is that in order to decrease the time to solution for application scientists we need to lower the barrier of entry into related computing fields. We aim at improving users' experience when interacting with a vast software ecosystem and/or huge amount of data, while maintaining focus on their primary research field. In this context we present our approach to bridge the gap between the application scientists and the visualization experts through a visualization schema as a first step and proof of concept for a new way to look at interdisciplinary collaboration among scientists dealing with big data. The key to our approach is recognizing that our users are scientists who mostly work as islands. They tend to work in very specialized environment but occasionally have to collaborate with other researchers in order to take full advantage of computing innovations and get insight from big data. We present an example of identifying the connecting elements between one of such relationships and offer a liaison schema to facilitate their collaboration. Roselyne Tchoua, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Jeremy Logan, Kenneth Moreland, Jingqing Mu, Manish Parashar, Norbert Podhorszki, David Pugmire, Matthew Wolf |
e-Science | 8 |
| 2013 | Exploring energy and performance behaviors of data-intensive scientific workflows on systems with deep memory hierarchiesabstractThe increasing gap between the rate at which large scale scientific simulations generate data and the corresponding storage speeds and capacities is leading to more complex system architectures with deep memory hierarchies. Advances in non-volatile memory (NVRAM) technology have made it an attractive candidate as intermediate storage in this memory hierarchy to address the latency and performance gap between main memory and disk storage. As a result, it is important to understand and model its energy/performance behavior from an application perspective as well as how it can be effectively used for staging data within an application workflow. In this paper, we target a NVRAM-based deep memory hierarchy and explore its potential for supporting in-situ/in-transit data analytics pipelines that are part of application workflows patterns. Specifically, we model the memory hierarchy and experimentally explore energy/performance behaviors of different data management strategies and data exchange patterns, as well as the tradeoffs associated with data placement, data movement and data processing. Marc Gamell, Ivan Rodero, Manish Parashar, Stephen W. Poole |
HiPC | 3 |
| 2013 | Exploring power behaviors and trade-offs of in-situ data analyticsabstractAs scientific applications target exascale, challenges related to data and energy are becoming dominating concerns. For example, coupled simulation workflows are increasingly adopting in-situ data processing and analysis techniques to address costs and overheads due to data movement and I/O. However it is also critical to understand these overheads and associated trade-offs from an energy perspective. The goal of this paper is exploring data-related energy/performance trade-offs for end-to-end simulation workflows running at scale on current high-end computing systems. Specifically, this paper presents: (1) an analysis of the data-related behaviors of a combustion simulation workflow with an in-situ data analytics pipeline, running on the Titan system at ORNL; (2) a power model based on system power and data exchange patterns, which is empirically validated; and (3) the use of the model to characterize the energy behavior of the workflow and to explore energy/performance trade-offs on current as well as emerging systems. Marc Gamell, Ivan Rodero, Manish Parashar, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer, Aaditya G. Landge, Attila Gyulassy, Patrick S. McCormick, Scott Pakin, Valerio Pascucci, Scott Klasky |
SC | 3 |
| 2013 | Using cross-layer adaptations for dynamic data management in large scale coupled scientific workflowsabstractAs system scales and application complexity grow, managing and processing simulation data has become a significant challenge. While recent approaches based on data staging and in-situ/in-transit data processing are promising, dynamic data volumes and distributions, such as those occurring in AMR-based simulations, make the efficient use of these techniques challenging. In this paper we propose cross-layer adaptations that address these challenges and respond at runtime to dynamic data management requirements. Specifically we explore (1) adaptations of the spatial resolution at which the data is processed, (2) dynamic placement and scheduling of data processing kernels, and (3) dynamic allocation of in-transit resources. We also exploit coordinated approaches that dynamically combine these adaptations at the different layers. We evaluate the performance of our adaptive cross-layer management approach on the Intrepid IBM-BlueGene/P and Titan Cray-XK7 systems using Chombo-based AMR applications, and demonstrate its effectiveness in improving overall time-to-solution and increasing resource efficiency. Tong Jin 0002, Fan Zhang 0004, Hoang Bui, Manish Parashar, Hongfeng Yu 0001, Scott Klasky, Norbert Podhorszki, Hasan Abbasi |
SC | 5 |
| 2013 | Distributed computing practice for large-scale science and engineering applicationsabstractSUMMARY It is generally accepted that the ability to develop large‐scale distributed applications has lagged seriously behind other developments in cyberinfrastructure. In this paper, we provide insight into how such applications have been developed and an understanding of why developing applications for distributed infrastructure is hard. Our approach is unique in the sense that it is centered around half a dozen existing scientific applications; we posit that these scientific applications are representative of the characteristics, requirements, as well as the challenges of the bulk of current distributed applications on production cyberinfrastructure (such as the US TeraGrid). We provide a novel and comprehensive analysis of such distributed scientific applications. Specifically, we survey existing models and methods for large‐scale distributed applications and identify commonalities, recurring structures, patterns and abstractions. We find that there are many ad hoc solutions employed to develop and execute distributed applications, which result in a lack of generality and the inability of distributed applications to be extensible and independent of infrastructure details. In our analysis, we introduce the notion of application vectors: a novel way of understanding the structure of distributed applications. Important contributions of this paper include identifying patterns that are derived from a wide range of real distributed applications, as well as an integrated approach to analyzing applications, programming systems and patterns, resulting in the ability to provide a critical assessment of the current practice of developing, deploying and executing distributed applications. Gaps and omissions in the state of the art are identified, and directions for future research are outlined. Copyright © 2012 John Wiley & Sons, Ltd. Shantenu Jha, Murray Cole, Daniel S. Katz, Manish Parashar, Omer F. Rana, Jon B. Weissman |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Editorial
Manish Parashar, Franco Zambonelli |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2012 | Accelerating MapReduce Analytics Using CometCloudabstractMapReduce-Hadoop has emerged as an effective framework for large-scale data analytics, providing support for executing jobs and storing data in a parallel and distributed manner. MapReduce has been shown to perform very well on large datacenters running applications where the data can be effectively divided into homogeneous chunks running across homogeneous hardware. However, the performance of MapReduceHadoop is far from ideal when either or both hardware and datasets are heterogeneous. Such heterogeneity is unavoidable in many academic computing environments that use multiple generations of hardware, and share resources among users. Heterogeneity is also unavoidable in scientific applications that process a varying number of datasets of different sizes. In these cases, the performance of MapReduce-Hadoop can be a concern. In this paper, we implement MapReduce on top of CometCloud to address the issue of heterogeneity and support applications classes that involve irregular datasets (e.g. large number of small data files or datasets of varying sizes). Furthermore, we develop an autonomic manager that can schedule MapReduce tasks based on user objective, provision resources accordingly, and support on-demand scale up and cloudbursts. These resources can be selected from a hybrid infrastructure such as local clusters, data centers, and public clouds. The performance of the developed solution is verified using a protein data mining application operating on data from the Protein Data Bank. The application is deployed, based on deadline and budget constraints, on a cluster at Rutgers and/or Amazon EC2 resources. The experimental results show that the MapReduce-CometCloud framework can effectively support applications operating on large numbers of small data files on a heterogeneous and distributed environment, and satisfy user objective autonomically using cloudbursts. Moustafa AbdelBaky, Hyunjoo Kim, Ivan Rodero, Manish Parashar |
IEEE CLOUD | 4 |
| 2012 | Mining hidden mixture context with ADIOS-P to improve predictive pre-fetcher accuracyabstractPredictive pre-fetcher, which predicts future data access events and loads the data before users requests, has been widely studied, especially in file systems or web contents servers, to reduce data load latency. Especially in scientific data visualization, pre-fetching can reduce the IO waiting time. In order to increase the accuracy, we apply a data mining technique to extract hidden information. More specifically, we apply a data mining technique for discovering the hidden contexts in data access patterns and make prediction based on the inferred context to boost the accuracy. In particular, we performed Probabilistic Latent Semantic Analysis (PLSA), a mixture model based algorithm popular in the text mining area, to mine hidden contexts from the collected user access patterns and, then, we run a predictor within the discovered context. We further improve PLSA by applying the Deterministic Annealing (DA) method to overcome the local optimum problem. In this paper we demonstrate how we can apply PLSA and DA optimization to mine hidden contexts from users data access patterns and improve predictive pre-fetcher performance. Jong Choi 0001, Hasan Abbasi, David Pugmire, Norbert Podhorszki, Scott Klasky, Cristian Capdevila, Manish Parashar, Matthew Wolf, Judy Qiu, Geoffrey C. Fox |
eScience | 7 |
| 2012 | Understanding I/O Performance Using I/O Skeletal Applications
Jeremy Logan, Scott Klasky, Hasan Abbasi, Qing Liu 0002, George Ostrouchov, Manish Parashar, Norbert Podhorszki, Yuan Tian 0004, Matthew Wolf |
Euro-Par | 6 |
| 2012 | A scalable messaging system for accelerating discovery from large scale scientific simulationsabstractEmerging scientific and engineering simulations running at scale on leadership-class High End Computing (HEC) environments are producing large volumes of data, which has to be transported and analyzed before any insights can result from these simulations. The complexity and cost (in terms of time and energy) associated with managing and analyzing this data have become significant challenges, and are limiting the impact of these simulations. Recently, data-staging approaches along with in-situ and in-transit analytics have been proposed to address these challenges by offloading I/O and/or moving data processing closer to the data. However, scientists continue to be overwhelmed by the large data volumes and data rates. In this paper we address this latter challenge. Specifically, we propose a highly scalable and low-overhead associative messaging framework that runs on the data staging resources within the HEC platform, and builds on the staging-based online in-situ/in-transit analytics to provide publish/subscribe/notification-type messaging patterns to the scientist. Rather than having to ingest and inspect the data volumes, this messaging system allows scientists to (1) dynamically subscribe to data events of interest, e.g., simple data values or a complex function or simple reduction (max()/min()/avg()) of the data values in a certain region of the application domain is greater/less than a threshold value, or certain spatial/temporal data features or data patterns are detected; (2) define customized in-situ/in-transit actions that are triggered based on the events, such as data visualization or transformation; and (3) get notified when these events occur. The key contribution of this paper is a design and implementation that can support such a messaging abstraction at scale on high-end computing (HEC) systems with minimal overheads. We have implemented and deployed the messaging system on the Jaguar Cray XK6 machines at Oak Ridge National Laboratory and the Lonestar system at the Texas Advanced Computing Center (TACC), and we present the experimental performance evaluation using these HEC platforms in the paper. Tong Jin 0002, Fan Zhang 0004, Manish Parashar, Scott Klasky, Norbert Podhorszki, Hasan Abbasi |
HiPC | 3 |
| 2012 | Exploring cross-layer power management for PGAS applications on the SCC platformabstractHigh-performance parallel computing architectures are increasingly based on multi-core processors. While current commercially available processors are at 8 and 16 cores, technological and power constraints are limiting the performance growth of the cores and are resulting in architectures with much higher core counts, such as the experimental many-core Intel Single-chip Cloud Computer (SCC) platform. These trends are presenting new sets of challenges to HPC applications including programming complexity and the need for extreme energy efficiency. Marc Gamell, Ivan Rodero, Manish Parashar, Rajeev Muralidhar |
HPDC | 3 |
| 2012 | Enabling In-situ Execution of Coupled Scientific Workflow on Multi-core PlatformabstractEmerging scientific application workflows are composed of heterogeneous coupled component applications that simulate different aspects of the physical phenomena being modeled, and that interact and exchange significant volumes of data at runtime. With the increasing performance gap between on-chip data sharing and off-chip data transfers in current systems based on multicore processors, moving large volumes of data using communication network fabric can significantly impact performance. As a result, minimizing the amount of inter-application data exchanges that are across compute nodes and use the network is critical to achieving overall application performance and system efficiency. In this paper, we investigate the in-situ execution of the coupled components of a scientific application workflow so as to maximize on-chip exchange of data. Specifically, we present a distributed data sharing and task execution framework that (1) employs data-centric task placement to map computations from the coupled applications onto processor cores so that a large portion of the data exchanges can be performed using the intra-node shared memory, (2) provides a shared space programming abstraction that supplements existing parallel programming models (e.g., message passing) with specialized one-sided asynchronous data access operators and can be used to express coordination and data exchanges between the coupled components. We also present the implementation of the framework and its experimental evaluation on the Jaguar Cray XT5 at Oak Ridge National Laboratory. Fan Zhang 0004, Ciprian Docan, Manish Parashar, Scott Klasky, Norbert Podhorszki, Hasan Abbasi |
IPDPS | 3 |
| 2012 | Combining in-situ and in-transit processing to enable extreme-scale scientific analysisabstractWith the onset of extreme-scale computing, I/O constraints make it increasingly difficult for scientists to save a sufficient amount of raw simulation data to persistent storage. One potential solution is to change the data analysis pipeline from a post-process centric to a concurrent approach based on either in-situ or in-transit processing. In this context computations are considered in-situ if they utilize the primary compute resources, while in-transit processing refers to offloading computations to a set of secondary resources using asynchronous data transfers. In this paper we explore the design and implementation of three common analysis techniques typically performed on large-scale scientific simulations: topological analysis, descriptive statistics, and visualization. We summarize algorithmic developments, describe a resource scheduling system to coordinate the execution of various analysis workflows, and discuss our implementation using the DataSpaces and ADIOS frameworks that support efficient data movement between in-situ and in-transit computations. We demonstrate the efficiency of our lightweight, flexible framework by deploying it on the Jaguar XK6 to analyze data generated by S3D, a massively parallel turbulent combustion code. Our framework allows scientists dealing with the data deluge at extreme scale to perform analyses at increased temporal resolutions, mitigate I/O costs, and significantly improve the time to insight. Janine Bennett, Hasan Abbasi, Peer-Timo Bremer, Ray W. Grout, Attila Gyulassy, Tong Jin 0002, Scott Klasky, Hemanth Kolla, Manish Parashar, Valerio Pascucci, Philippe P. Pébay, David C. Thompson 0001, Hongfeng Yu 0001, Fan Zhang 0004, Jacqueline Chen |
SC | 9 |
| 2012 | Special section on autonomic cloud computing: technologies, services, and applicationsabstractSpecial section on autonomic cloud computing: technologies, Rajiv Ranjan 0001, Rajkumar Buyya, Manish Parashar |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Energy-Efficient Thermal-Aware Autonomic Management of Virtualized HPC Cloud Infrastructure
Ivan Rodero, Hariharasudhan Viswanathan, Marc Gamell, Dario Pompili, Manish Parashar |
J. Grid Comput. | 6 |
| 2012 | Cloud federation in a layered service model
David Villegas, Norman Bobroff, Ivan Rodero, Javier Delgado, Aditya Devarakonda, Liana L. Fong, Seyed Masoud Sadjadi, Manish Parashar |
J. Comput. Syst. Sci. | 9 |
| 2012 | Design and evaluation of decentralized online clusteringabstractEnsuring the efficient and robust operation of distributed computational infrastructures is critical, given that their scale and overall complexity is growing at an alarming rate and that their management is rapidly exceeding human capability. Clustering analysis can be used to find patterns and trends in system operational data, as well as highlight deviations from these patterns. Such analysis can be essential for verifying the correctness and efficiency of the operation of the system, as well as for discovering specific situations of interest, such as anomalies or faults, that require appropriate management actions. This work analyzes the automated application of clustering for online system management, from the point of view of the suitability of different clustering approaches for the online analysis of system data in a distributed environment, with minimal prior knowledge and within a timeframe that allows the timely interpretation of and response to clustering results. For this purpose, we evaluate DOC (Decentralized Online Clustering), a clustering algorithm designed to support data analysis for autonomic management, and compare it to existing and widely used clustering algorithms. The comparative evaluations will show that DOC achieves a good balance in the trade-offs inherent in the challenges for this type of online management. Andres Quiroz, Manish Parashar, Nathan Gnanasambandam, Naveen Sharma |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2012 | Proactive thermal management in green datacenters
Indraneel S. Kulkarni, Dario Pompili, Manish Parashar |
J. Supercomput. | 4 |
| 2011 | Enabling Multi-physics Coupled Simulations within the PGAS Programming FrameworkabstractComplex coupled multi-physics simulations are playing increasingly important roles in scientific and engineering applications such as fusion plasma and climate modeling. At the same time, extreme scales, high levels of concurrency and the advent of multicore and many core technologies are making the high-end parallel computing systems on which these simulations run, hard to program. While the Partitioned Global Address Space (PGAS) languages is attempting to address the problem, the PGAS model does not easily support the coupling of multiple application codes, which is necessary for the coupled multi-physics simulations. Furthermore, existing frameworks that support coupled simulations have been developed for fragmented programming models such as message passing, and are conceptually mismatched with the shared memory address space abstraction in the PGAS programming model. This paper explores how multi-physics coupled simulations can be supported within the PGAS programming framework. Specifically, in this paper, we present the design and implementation of the XpressSpace programming system, which enables efficient and productive development of coupled simulations across multiple independent PGAS Unified Parallel C (UPC) executables. XpressSpace provides the global-view style programming interface that is consistent with the memory model in UPC, and provides an efficient runtime system that can dynamically capture the data decomposition of global-view arrays and enable fast exchange of parallel data structures between coupled codes. In addition, XpressSpace provides the flexibility to define the coupling process in specification file that is independent of the program source codes. We evaluate the performance and scalability of Xpress Space prototype implementation using different coupling patterns extracted from real world multi-physics simulation scenarios, on the Jaguar Cray XT5 system of Oak Ridge National Laboratory. Fan Zhang 0004, Ciprian Docan, Manish Parashar, Scott Klasky |
CCGRID | 3 |
| 2011 | Adaptive memory power management techniques for HPC workloadsabstractThe memory subsystem is responsible for a large fraction of the energy consumed by compute nodes in High Performance Computing (HPC) systems. The rapid increase in the number of cores has been accompanied by a corresponding increase in the DRAM capacity and bandwidth, and as a result, the memory system consumes a significant amount of the power budget available to a compute node. Consequently, there is a broad research effort focused on power management techniques using DRAM low-power modes. However, memory power management continues to present many challenges. In this paper, we study the potential of Dynamic Voltage and Frequency Scaling (DVFS) of the memory subsystems, and consider the ability to select different frequencies for different memory channels. Our approach is based on tuning voltage and frequency dynamically to maximize the energy savings while maintaining performance degradation within tolerable limits. We assume that HPC applications do not demand maximum bandwidth throughout the entire period of execution. We can use these low memory demand intervals to tune down the frequency and, as a result, applications can tolerate a reduction in bandwidth to save energy. In this paper, we study application channel access patterns, and use these patterns to determine potential additional energy savings that can be achieved by accordingly controlling the channels independently. We then evaluate the proposed DVFS algorithm using a novel hybrid evaluation methodology that includes simulation as well as executions on real hardware. Our results demonstrate the large potential of adaptive memory power management techniques based on DVFS for HPC workloads. Karthik Elangovan, Ivan Rodero, Manish Parashar, Francesc Guim 0001, Isaac Hernandez |
HiPC | 3 |
| 2011 | Moving the Code to the Data - Dynamic Code Deployment Using ActiveSpacesabstractManaging the large volumes of data produced by emerging scientific and engineering simulations running on leadership-class resources has become a critical challenge. The data has to be extracted off the computing nodes and transported to consumer nodes so that it can be processed, analyzed, visualized, archived, etc. Several recent research efforts have addressed data-related challenges at different levels. One attractive approach is to offload expensive I/O operations to a smaller set of dedicated computing nodes known as a staging area. However, even using this approach, the data still has to be moved from the staging area to consumer nodes for processing, which continues to be a bottleneck. In this paper, we investigate an alternate approach, namely moving the data-processing code to the staging area rather than moving the data. Specifically, we present the Active Spaces framework, which provides (1) programming support for defining the data-processing routines to be downloaded to the staging area, and (2) run-time mechanisms for transporting binary codes associated with these routines to the staging area, executing the routines on the nodes of the staging area, and returning the results. We also present an experimental performance evaluation of Active Spaces using applications running on the Cray XT5 at Oak Ridge National Laboratory. Finally, we use a coupled fusion application workflow to explore the trade-offs between transporting data and transporting the code required for data processing during coupling, and we characterize the sweet spots for each option. Ciprian Docan, Manish Parashar, Julian C. Cummings, Scott Klasky |
IPDPS | 2 |
| 2011 | NSF/IEEE-TCPP curriculum initiative on parallel and distributed computing: core topics for undergraduatesabstractNo abstract available. Sushil K. Prasad, Almadena Yu. Chtchelkanova, Sajal K. Das 0001, Frank Dehne, Mohamed G. Gouda, Joseph F. JáJá, Krishna Kant 0001, Anita La Salle, Richard LeBlanc, Manish Lumsdaine, David A. Padua, Manish Parashar, Viktor Prasanna 0001, Yves Robert, Arnold L. Rosenberg, Sartaj Sahni, Behrooz A. Shirazi, Alan Sussman, Charles C. Weems, Jie Wu 0001 |
SIGCSE | 13 |
| 2011 | EditorialabstractNo abstract available. Manish Parashar, Franco Zambonelli |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2010 | Experiments with Memory-to-Memory Coupling for End-to-End Fusion Simulation WorkflowsabstractScientific applications are striving to accurately simulate multiple interacting physical processes that comprise complex phenomena being modeled. Efficient and scalable parallel implementations of these coupled simulations present challenging interaction and coordination requirements, especially when the coupled physical processes are computationally heterogeneous and progress at different speeds. In this paper, we present the design, implementation and evaluation of a memory-to-memory coupling framework for coupled scientific simulations on high-performance parallel computing platforms. The framework is driven by the coupling requirements of the Center for Plasma Edge Simulation, and it provides simple coupling abstractions as well as efficient asynchronous (RDMA-based) memory-to-memory data transport mechanisms that complement existing parallel programming systems and data sharing frameworks. The framework enables flexible coupling behaviors that are asynchronous in time and space, and it supports dynamic coupling between heterogeneous simulation processes without enforcing any synchronization constraints. We evaluate the performance and scalability of the coupling framework using a specific coupling scenario, on the Jaguar Cray XT5 system at Oak Ridge National Laboratory. Ciprian Docan, Fan Zhang 0004, Manish Parashar, Julian C. Cummings, Norbert Podhorszki, Scott Klasky |
CCGRID | 3 |
| 2010 | Exploring the Performance Fluctuations of HPC Workloads on CloudsabstractClouds enable novel execution modes often supported by advanced capabilities such as autonomic schedulers. These capabilities are predicated upon an accurate estimation and calculation of runtimes on a given infrastructure. Using a well understood high-performance computing workload, we find strong fluctuations from the mean performance on EC2 and Eucalyptus-based cloud systems. Our analysis eliminates variations in IO and computational times as possible causes, we find that variations in communication times account for the bulk of the experiment-to-experiment fluctuations of the performance. Yaakoub El Khamra, Hyunjoo Kim, Shantenu Jha, Manish Parashar |
CloudCom | 4 |
| 2010 | Investigating the potential of application-centric aggressive power management for HPC workloadsabstractEnergy efficiency of large-scale data centers is becoming a major concern not only for reasons of energy conservation, failures, and cost reduction, but also because such sys tems are soon reaching the limits of power available to them. Like High Performance Computing (HPC) systems, large-scale clu ster-based data centers can consume power in megawatts, and of all the power consumed by such a system, only a fraction is used for actual computations. In this paper, we study the potential of application-centric aggressive power management of data center's resources for HPC workloads. Specifically, we consider power management mechanisms and controls (currently or soon to be) available at different levels and for different subsystems, and leverage several innovative approaches that have been taken to tackle this problem in the last few years, can be effectively used in a application-aware manner for HPC workloads. To do this, we first profile sta ndard HPC benchmarks with respect to behaviors, resource usage and power impact on individual computing nodes. Based on a power and latency model and the workload profiles, we develop an algorithm that can improve energy efficiency with little or no performance loss. We then evaluate our proposed algorithm through simulations using empirical power characterization and quantification. Finally, we validate the simulation results with actual executions on real hardware. The obtained results show that by using application aware power management, we can re-du ce the average energy consumption without significant penalty in performance. This motivates us to investigate autonomic approaches for application-aware aggressive power management and cross layer and cross function predictive subsystem level power management for large-scale data centers. Ivan Rodero, Sharat Chandra, Manish Parashar, Rajeev Muralidhar, Harinarayanan Seshadri, Stephen W. Poole |
HiPC | 3 |
| 2010 | Enabling GPU and Many-Core Systems in Heterogeneous HPC Environments Using Memory ConsiderationsabstractIncreasing the utilization of many-core systems has been one of the forefront topics these last years. Although many-cores architectures were merely theoretical models few years ago, they have become an important part of the high performance computing market. The semiconductor industry has developed Graphical Processing Units (GPU) systems that provide access to many cores (i.e: Larrabee, Fermi or Tesla) that can be used for General Purpose (GP) computing. In this paper, we propose and evaluate a scheduling strategy for GPU and many-core architectures for HPC environments. Specifically, our strategy is a variant of the backfilling scheduling policy with resource sharing considerations. We propose a scheduling strategy that considers the differences between GP processors and GPU computing elements in terms of computational capacity and memory bandwidth. To do this, our approach uses a resource model that predicts how shared resources are used in both GP and GPU/many-core elements. Furthermore, it considers the differences between these elements in terms of performance. First, it models their differences in terms of computational power and how they share the access to the node's memory bandwidth. Second, it characterizes how the processes are allocated to the GPU. Using this resource model, we design the Power Aware resource selection policy, which we combine with the LessConsume scheduling policy. Our strategy tries to allocate jobs aiming at reducing the memory contention and the energy consumption. Results show that the scheduling strategies proposed in this work are able to save over 40% of energy and improve the system performance up to 30% with respect to traditional backfilling strategies. Francesc Guim 0001, Ivan Rodero, Julita Corbalán, Manish Parashar |
HPCC | 4 |
| 2010 | DataSpaces: an interaction and coordination framework for coupled simulation workflowsabstractEmerging high-performance distributed computing environments are enabling new end-to-end formulations in science and engineering that involve multiple interacting processes and data-intensive application workflows. For example, current fusion simulation efforts are exploring coupled models and codes that simultaneously simulate separate application processes, such as the core and the edge turbulence, and run on different high performance computing resources. These components need to interact, at runtime, with each other and with services for data monitoring, data analysis and visualization, and data archiving. As a result, they require efficient support for dynamic and flexible couplings and interactions, which remains a challenge. This paper presents DataSpaces, a flexible interaction and coordination substrate that addresses this challenge. DataSpaces essentially implements a semantically specialized virtual shared space abstraction that can be associatively accessed by all components and services in the application workflow. It enables live data to be extracted from running simulation components, indexes this data online, and then allows it to be monitored, queried and accessed by other components and services via the space using semantically meaningful operators. The underlying data transport is asynchronous, low-overhead and largely memory-to-memory. The design, implementation, and experimental evaluation of DataSpaces using a coupled fusion simulation workflow is presented. Ciprian Docan, Manish Parashar, Scott Klasky |
HPDC | 2 |
| 2010 | Exploring application and infrastructure adaptation on hybrid grid-cloud infrastructureabstractClouds are emerging as an important class of distributed computational resources and are quickly becoming an integral part of production computational infrastructures. An important but oft-neglected question is, what new applications and application capabilities can be supported by clouds as part of a hybrid computational platform? In this paper we use the ensemble Kalman-filter based dynamic application workflow and investigate how clouds can be effectively used as an accelerator to address changing computational requirements as well as changing Quality of Service constraints (e.g., deadlines). Furthermore, we explore how application and system-level adaptivity can be used to improve application performance and achieve a more effective utilization of the hybrid platform. Specifically, we adapt the ensemble Kalman-filter based application formulation (serial versus parallel, different solvers etc.) so as to execute efficiently on a range of different infrastructure (from High Performance Computing grids to clouds that support single core and many-core virtual machines). Our results show that there are performance advantages to be had by supporting application and infrastructure level adaptivity. In general, we find that grid-cloud infrastructure can support novel usage modes, such as deadline-driven scheduling, for applications with tunable characteristics that can adapt to varying resource types. Hyunjoo Kim, Yaakoub El Khamra, Shantenu Jha, Manish Parashar |
HPDC | 4 |
| 2010 | Towards energy-aware autonomic provisioning for virtualized environmentsabstractAs energy efficiency and associated costs become key concerns, consolidated and virtualized data centers and clouds are attractive computing platforms for data- and compute-intensive applications. Recently, these platforms are also being considered for more traditional high-performance computing (HPC) applications. However, maximizing energy efficiency, cost-effectiveness, and utilization for these applications while ensuring performance and other Quality of Service (QoS) guarantees, requires leveraging important and extremely challenging tradeoffs. These include, for example, the tradeoff between the need to efficiently create and provision Virtual Machines (VMs) on data center resources and the need to accommodate the heterogeneous resource demands and runtimes of the applications that run on them. In this paper we propose an energy-aware online provisioning approach for HPC applications on consolidated and virtualized computing platforms. Energy efficiency is achieved using a workload-aware, just-right dynamic provisioning mechanism and the ability to power down subsystems of a host system that are not required by the VMs mapped to it. Our preliminary evaluations show that our approach can improve energy efficiency with an acceptable QoS penalty. Ivan Rodero, Juan Jaramillo, Andres Quiroz, Manish Parashar, Francesc Guim 0001 |
HPDC | 4 |
| 2010 | Object-oriented stream programming using aspectsabstractHigh-performance parallel programs that efficiently utilize heterogeneous CPU+GPU accelerator systems require tuned coordination among multiple program units. However, using current programming frameworks such as CUDA leads to tangled source code that combines code for the core computation with that for device and computational kernel management, data transfers between memory spaces, and various optimizations. In this paper, we propose a programming system based on the principles of Aspect-Oriented Programming, to un-clutter the code and to improve programmability of these heterogeneous parallel systems. Specifically, we use standard C++ to describe the core computations and aspects to encapsulate all other support parts. An aspect-weaving compiler is then used to combine these two pieces of code to generate a final program. The system modularizes concerns that are hard to manage using conventional programming frameworks such as CUDA, has a small impact on existing program structure as well as performance, and as a result, simplifies the programming of accelerator-based heterogeneous parallel systems. We also present an options pricing and an n-body simulation example program to demonstrate that programs written using this system can be successfully translated to CUDA programs for NVIDIA GPU hardware and to OpenCL programs for multicore CPUs with comparable performance. For both examples, the performance of the translated code achieved ~80% of the hand-coded CUDA programs. Manish Parashar |
IPDPS | 2 |
| 2010 | PreDatA - preparatory data analytics on peta-scale machinesabstractPeta-scale scientific applications running on High End Computing (HEC) platforms can generate large volumes of data. For high performance storage and in order to be useful to science end users, such data must be organized in its layout, indexed, sorted, and otherwise manipulated for subsequent data presentation, visualization, and detailed analysis. In addition, scientists desire to gain insights into selected data characteristics `hidden' or `latent' in these massive datasets while data is being produced by simulations. PreDatA, short for Preparatory Data Analytics, is an approach to preparing and characterizing data while it is being produced by the large scale simulations running on peta-scale machines. By dedicating additional compute nodes on the machine as `staging' nodes and by staging simulations' output data through these nodes, PreDatA can exploit their computational power to perform select data manipulations with lower latency than attainable by first moving data into file systems and storage. Such intransit manipulations are supported by the PreDatA middleware through asynchronous data movement to reduce write latency, application-specific operations on streaming data that are able to discover latent data characteristics, and appropriate data reorganization and metadata annotation to speed up subsequent data access. PreDatA enhances the scalability and flexibility of the current I/O stack on HEC platforms and is useful for data pre-processing, runtime data analysis and inspection, as well as for data exchange between concurrently running simulations. Fang Zheng 0003, Hasan Abbasi, Ciprian Docan, Jay F. Lofstead, Qing Liu 0002, Scott Klasky, Manish Parashar, Norbert Podhorszki, Karsten Schwan, Matthew Wolf |
IPDPS | 7 |
| 2010 | EFFIS: An End-to-end Framework for Fusion Integrated SimulationabstractThe purpose of the Fusion Simulation Project is to develop a predictive capability for integrated modeling of magnetically confined burning plasmas. In support of this mission, the Center for Plasma Edge Simulation has developed an End-to-end Framework for Fusion Integrated Simulation (EFFIS) that combines critical computer science technologies in an effective manner to support leadership class computing and the coupling of complex plasma physics models. We describe here the main components of EFFIS and how they are being utilized to address our goal of integrated predictive plasma edge simulation. Julian C. Cummings, Jay F. Lofstead, Karsten Schwan, Alex Sim, Arie Shoshani, Ciprian Docan, Manish Parashar, Scott Klasky, Norbert Podhorszki, Roselyne Tchoua |
PDP | 7 |
| 2010 | Enabling high-speed asynchronous data extraction and transfer using DARTabstractAbstract As the complexity and scale of applications grow, managing and transporting the large amounts of data they generate are quickly becoming a significant challenge. Moreover, the interactive and real‐time nature of emerging applications, as well as their increasing runtime, make online data extraction and analysis a key requirement in addition to traditional data I/O and archiving. To be effective, online data extraction and transfer should impose minimal additional synchronization requirements, should have minimal impact on the computational performance and communication latencies, maintain overall quality of service, and ensure that no data is lost. In this paper we present Decoupled and Asynchronous Remote Transfers (DART), an efficient data transfer substrate that effectively addresses these requirements. DART is a thin software layer built on RDMA technology to enable fast, low‐overhead, and asynchronous access to data from a running simulation, and supports high‐throughput, low‐latency data transfers. DART has been integrated with applications simulating fusion plasma in a Tokamak, being developed at the Center for Plasma Edge Simulation (CPES), a DoE Office of Fusion Energy Science (OFES) Fusion Simulation Project (FSP). A performance evaluation using the Gyrokinetic Toroidal Code and XGC‐1 particle‐in‐cell‐based FSP simulations running on the Cray XT3/XT4 system at Oak Ridge National Laboratory demonstrates how DART can effectively and efficiently offload simulation data to local service and remote analysis nodes, with minimal overheads on the simulation itself. Copyright © 2010 John Wiley & Sons, Ltd. Ciprian Docan, Manish Parashar, Scott Klasky |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Online Risk Analytics on the CloudabstractIn todays turbulent market conditions, the ability to generate accurate and timely risk measures has become critical to operating successfully, and necessary for survival. Value-at-risk (VaR) is a market standard risk measure used by senior management and regulators to quantify the risk level of a firm's holdings. However, the time-critical nature and dynamic computational workloads of VaR applications, make it essential for computing infrastructures to handle bursts in computing and storage resources needs. This requires on-demand scalability, dynamic provisioning, and the integration of distributed resources. While emerging utility computing services and clouds have the potential for cost-effectively supporting such spikes in resource requirements, integrating clouds with computing platforms and data centers, as well as developing and managing applications to utilize the platform remains a challenge. In this paper, we focus on the dynamic resource requirements of online risk analytics applications and how they can be addressed by cloud environments. Specifically, we demonstrate how the CometCloud autonomic computing engine can support online multi-resolution VaR analytics using and integration of private and Internet cloud resources. Hyunjoo Kim, Shivangi Chaudhari, Manish Parashar, Christopher Marty |
CCGRID | 3 |
| 2009 | An Autonomic Approach to Integrated HPC Grid and Cloud UsageabstractClouds are rapidly joining high-performance Grids as viable computational platforms for scientific exploration and discovery, and it is clear that production computational infrastructures will integrate both these paradigms in the near future. As a result, understanding usage modes that are meaningful in such a hybrid infrastructure is critical. For example, there are interesting application workflows that can benefit from such hybrid usage modes to, per- haps, reduce times to solutions, reduce costs (in terms of currency or resource allocation), or handle unexpected runtime situations (e.g., unexpected delays in scheduling queues or unexpected failures). The primary goal of this paper is to experimentally investigate, from an applications perspective, how autonomics can enable interesting usage modes and scenarios for integrating HPC Grid and Clouds. Specifically, we used a reservoir characterization application workflow, based on Ensemble Kalman Filters (EnKF) for history matching, and the CometCloud autonomic Cloud engine on a hybrid platform consisting of the TeraGrid and Amazon EC2, to investigate 3 usage modes (or autonomic objectives) - acceleration, conservation and resilience. Hyunjoo Kim, Yaakoub El Khamra, Shantenu Jha, Manish Parashar |
eScience | 4 |
| 2009 | Advanced risk analytics on the cell broadband engineabstractThis paper explores the effectiveness of using the CBE platform for Value-at-Risk (VaR) calculations. Specifically, it focuses on the design, optimization and evaluation of pricing European and American stock options across Monte-Carlo VaR scenarios. This analysis is performed on two distinct platforms with CBE processors, i.e., IBM Q22 blade server and the Playstation3 gaming console. Ciprian Docan, Manish Parashar, Christopher Marty |
IPDPS | 2 |
| 2009 | Enabling autonomic power-aware management of instrumented data centersabstractSensor networks support flexible, non-intrusive and fine-grained data collection and processing and can enable online monitoring of data center operating conditions as well as autonomic data center management. This paper describes the architecture and implementation of an autonomic power aware data center management framework, which is based on the integration of embedded sensors with computational models and workload schedulers to improve data center performance in terms of energy consumption and throughput. Specifically, workload schedulers use online information about data center operating conditions obtained from the sensors to generate appropriate management policies. Furthermore, local processing within the sensor network is used to enable timely responses to changes in operating conditions and determine job migration strategies. Experimental results demonstrate that the framework achieves near optimal management, and in-network analysis enables timely response while reducing overheads. Nanyan Jiang, Manish Parashar |
IPDPS | 2 |
| 2008 | GridMate: A Portable Simulation Environment for Large-Scale Adaptive Scientific ApplicationsabstractIn this paper, we present a portable simulation environment GridMate for large-scale adaptive scientific applications in multi-site Grid environments. GridMate is a discrete-event based simulator, consisting of abstractions of trace-based applications, computing resources, partitioners and schedulers, a 3D visualization tool, and user interfaces. It supports the analysis of runtime management strategies that address spatial and temporal heterogeneity in both adaptive scientific applications and geographically distributed resources in Grid computing environments. The targeted applications are a class of emerging large-scale dynamic Grid applications that require large amount of computational resources typically spanning multiple sites and exhibit long execution times. The underlying partitioning and scheduling algorithms are based on our previous work on the hybrid space-time runtime management strategy (HRMS). HRMS defines a set of flexible mechanisms and policies to adapt to state transitions of both applications and resources. The major components of GridMate are developed in Java, making GridMate highly portable and extensible. The design of GridMate and simulation results using GridMate are presented. Xiaolin Li 0001, Manish Parashar |
CCGRID | 2 |
| 2008 | In-Network Data Estimation for Sensor-Driven Scientific Applications
Nanyan Jiang, Manish Parashar |
HiPC | 2 |
| 2008 | DART: a substrate for high speed asynchronous data IOabstractLarge scale simulations of complex physics phenomena have long run times and generate massive amounts of data. Saving this data to external storage systems or transferring it to remote locations for analysis is a costly operation that quickly becomes a performance bottleneck. In this paper, we present DART (Decoupled and Asynchronous Remote Transfers), an efficient data transfer substrate that effectively minimizes the data I/O overhead on the running simulations. DART is a thin software layer built on RDMA technology to enable fast, low-overhead and asynchronous access to data from a running simulation, and support high-throughput, low-latency data transfers. Ciprian Docan, Manish Parashar, Scott Klasky |
HPDC | 2 |
| 2008 | Autonomic Management of Hybrid Sensor Grid Systems and ApplicationsabstractIn this paper, we propose an autonomic management framework (ASGrid) to address the requirements of emerging large-scale applications in hybrid grid and sensor network systems. To the best of our knowledge, we are the first who proposed the autonomic sensor grid system concept in a holistic manner targeted at non-trivial large applications. To bridge the gap between the physical world and the digital world and facilitate information analysis and decision making, ASGrid is designed to smooth the integration of sensor networks and grid systems and efficiently use both on demand. Under the blueprint of ASGrid, we present several building blocks that fulfill the following major features: (1) Self-configuration through content-based aggregation and associative rendezvous mechanisms; (2) Self-optimization through utility-based sensor selection and model-driven hierarchical sensing task scheduling; (3) Self-protection through ActiveKey dynamic key management and S3Trust trust management mechanisms. Experimental and simulation results on these aspects are presented. Xiaolin Li 0001, Xinxin Liu 0006, Huanyu Zhao, Nanyan Jiang, Manish Parashar |
ICCCN | 5 |
| 2008 | Programming support for sensor-based scientific applicationsabstractTechnical advances are enabling a pervasive computational ecosystem that integrates computing infrastructures with embedded sensors and actuators, and are giving rise to a new paradigm for monitoring, understanding, and managing natural and engineered systems - one that is information/data-driven. This research investigates programming systems for sensor-driven applications. It addresses abstractions and runtime mechanisms for integrating sensor systems with computational models for scientific processes, as well as for in- network data processing, e.g., aggregation, adaptive interpolation and assimilation. The current status of this research, as well as initial results are presented. Nanyan Jiang, Manish Parashar |
IPDPS | 2 |
| 2008 | Meteor: a middleware infrastructure for content-based decoupled interactions in pervasive grid environmentsabstractAbstract Emerging pervasive information and computational environments require a content‐based middleware infrastructure that is scalable, self‐managing, and asynchronous. In this paper, we propose associative rendezvous (AR) as a paradigm for content‐based decoupled interactions for pervasive grid applications. We also present Meteor, a content‐based middleware infrastructure to support AR interactions. The design, implementation, and experimental evaluation of Meteor are presented. Evaluations include experiments using deployments on a local area network, the wireless ORBIT testbed at Rutgers University, and the PlanetLab wide‐area testbed, as well as simulations. Evaluation results demonstrate the scalability, effectiveness, and performance of Meteor to support pervasive grid applications. Copyright © 2007 John Wiley & Sons, Ltd. Nanyan Jiang, Andres Quiroz, Cristina Schmidt, Manish Parashar |
Concurr. Comput. Pract. Exp. | 4 |
| 2008 | A framework for distributed content-based web services notification in Grid systems
Andres Quiroz, Manish Parashar |
Future Gener. Comput. Syst. | 2 |
| 2008 | Squid: Enabling search in DHT-based systems
Cristina Schmidt, Manish Parashar |
J. Parallel Distributed Comput. | 2 |
| 2007 | A computational infrastructure for grid-based asynchronous parallel applicationsabstractNo abstract available. Zhen Li 0010, Manish Parashar |
HPDC | 2 |
| 2007 | Enabling scalable parallel implementations of structured adaptive mesh refinement applications
Sumir Chandra, Xiaolin Li 0001, Taher Saif, Manish Parashar |
J. Supercomput. | 4 |
| 2007 | Hybrid Runtime Management of Space-Time Heterogeneity for Parallel Structured Adaptive ApplicationsabstractStructured adaptive mesh refinement (SAMR) techniques provide an effective means for dynamically concentrating computational effort and resources to appropriate regions in the application domain. However, due to their dynamism and space-time heterogeneity, scalable parallel implementation of SAMR applications remains a challenge. This paper investigates hybrid runtime management strategies and presents an adaptive hierarchical multipartitioner (AHMP) framework. AHMP dynamically applies multiple partitioners to different regions of the domain, in a hierarchical manner, to match the local requirements of the regions. Key components of the AHMP framework include a segmentation-based clustering algorithm (SBC) that can efficiently identify regions in the domain with relatively homogeneous partitioning requirements, mechanisms for characterizing the partitioning requirements of these regions, and a runtime system for selecting, configuring, and applying the most appropriate partitioner to each region. Further, to address dynamic resource situations for long-running applications, AHMP provides a hybrid partitioning strategy (HPS) that involves application-level pipelining, trading space for time when resources are sufficiently large and underutilized, and an application-level out-of-core strategy (ALOC), trading time for space when resources are scarce in order to enhance the survivability of applications. The AHMP framework has been implemented and experimentally evaluated on up to 1,280 processors of the IBM SP4 cluster at the San Diego Supercomputer Center. Xiaolin Li 0001, Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | Experiments with Wide Area Data Coupling Using the Seine Coupling Framework
Manish Parashar, Scott Klasky |
HiPC | 2 |
| 2006 | Salsa: Scalable Asynchronous Replica Exchange for Parallel Molecular Dynamics ApplicationsabstractThis paper presents Salsa, a novel, decentralized and asynchronous realization of the "replica exchange" algorithm for simulating the structure, function, folding, and dynamics of proteins. Salsa provides a scalable communication and interaction substrate that presents a virtual shared space abstraction and enables the dynamic and asynchronous interactions required by the simulations to be simply and efficiently implemented. The design, implementation and experimental evaluation of Salsa and replica exchange-based simulations using Salsa are presented Manish Parashar, Emilio Gallicchio, Ronald M. Levy |
ICPP | 2 |
| 2006 | Dynamic structured partitioning for parallel scientific applications with pointwise varying workloadsabstractParallel implementations of scientific applications involving the simulation of reactive flow on structured grids are challenging, since the underlying phenomena include transport processes with uniform computational loads as well as reactive processes having point-wise varying workloads. As a result, traditional parallelization approaches that assume homogeneous loads are not suitable for these simulations. This paper presents "Dispatch", a dynamic structured partitioning strategy that has been applied to parallel uniform and adaptive formulations of simulations with computational heterogeneity. Dispatch maintains the computational weights associated with pointwise processes in a distributed manner, computes the local workloads and partitioning thresholds, and performs in-situ locality-preserving load balancing. The experimental evaluation of Dispatch using an illustrative 2-D reactive-diffusion kernel demonstrates improvement in load distribution and overall application performance. Sumir Chandra, Manish Parashar, Jaideep Ray |
IPDPS | 2 |
| 2006 | Enabling efficient and flexible coupling of parallel scientific applicationsabstractEmerging scientific and engineering simulations are presenting challenging requirements for coupling between multiple physics models and associated parallel codes that execute independently and in a distributed manner. Realizing coupled simulations requires an efficient, flexible and scalable coupling framework and simple programming abstractions. This paper presents a coupling framework that addresses these requirements. The framework is based on the Seine geometry-based interaction model. It enables efficient computation of communication schedules, supports low-overheads processor-to-processor data streaming, and provides high-level abstraction for application developers. The design, CCA-based implementation, and experimental evaluation of the Seine based coupling framework are presented Manish Parashar |
IPDPS | 2 |
| 2006 | Seine: a dynamic geometry-based shared-space interaction framework for parallel scientific applicationsabstractWhile large-scale parallel/distributed simulations are rapidly becoming critical research modalities in academia and industry, their efficient and scalable implementations continue to present many challenges. A key challenge is that the dynamic and complex communication/coordination required by these applications (dependent on the state of the phenomenon being modeled) are determined by the specific numerical formulation, the domain decomposition and/or sub-domain refinement algorithms used, etc. and are known only at runtime. This paper presents Seine, a dynamic geometry-based shared-space interaction framework for scientific applications. The framework provides the flexibility of shared-space-based models and supports extremely dynamic communication/coordination patterns, while still enabling scalable implementations. The design and prototype implementation of Seine are presented. Seine complements and can be used in conjunction with existing parallel programming systems such as MPI and OpenMP. An experimental evaluation using an adaptive multi-block oil-reservoir simulation is used to demonstrate the performance and scalability of applications using Seine. Copyright © 2006 John Wiley & Sons, Ltd. Manish Parashar |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | A Decentralized Computational Infrastructure for Grid-Based Parallel Asynchronous Iterative Applications
Zhen Li 0010, Manish Parashar |
J. Grid Comput. | 2 |
| 2006 | Accord: a programming framework for autonomic applicationsabstractThe emergence of pervasive wide-area distributed computing environments, such as pervasive information systems and computational grids, has enabled new generations of applications that are based on seamless access, aggregation, and interaction. However, the inherent complexity, heterogeneity, and dynamism of these systems require a change in how the applications are developed and managed. In this paper, we present a programming framework that extends existing programming models/frameworks to support the development of autonomic self-managing applications. The framework enables the development of autonomic elements and the formulation of autonomic applications as the dynamic composition of autonomic elements. The operation of the proposed framework is illustrated using a forest fire management application. Hua Liu 0001, Manish Parashar |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2005 | Using Clustering to Address Heterogeneity and Dynamism in Parallel Scientific Applications
Xiaolin Li 0001, Manish Parashar |
HiPC | 2 |
| 2005 | Enabling self-management of component-based high-performance scientific applicationsabstractEmerging high-performance parallel/distributed scientific applications and environments are increasingly large, dynamic and complex. As a result, it requires programming systems that enable the applications to detect and dynamically respond to changing requirements, state and execution context by adapting their computational behaviors and interactions. In this paper, we present such a programming system that extends the common component architecture to enable self-management of component-based scientific applications. The programming system separates and categorizes operational requirements of scientific applications, and allows them to be specified and enforced at runtime through reconfiguration, optimization and healing of individual components and the application. Two scientific simulations are used to illustrate the system and its self-managing behaviors. A performance evaluation is also presented. Hua Liu 0001, Manish Parashar |
HPDC | 2 |
| 2005 | Autonomic runtime manager for adaptive distributed applicationsabstractFor adaptive distributed applications, the computational complexity associated with each computational region varies continuously and dramatically both in space and time throughout the life cycle of the application execution. Consequently, static scheduling techniques are inefficient for such applications. In this paper, we present an autonomic runtime manager (ARM) that uses the application spatial and temporal characteristics as well as resource status as the main criteria to self optimize the execution of distributed applications at runtime. We applied the ARM system to a wildfire simulation and our experimental results show that the performance of the wildfire simulation has been improved by 45% when compared with a static partitioning algorithm. We also evaluate the performance of ARM using two partitioning strategies: natural regions (NR) approach and a graph partitioning approach. Jingmei Yang, Huoping Chen, Salim Hariri, Manish Parashar |
HPDC | 4 |
| 2005 | A concise introduction to autonomic computing
Roy Sterritt, Manish Parashar, Huaglory Tianfield, Rainer Unland |
Adv. Eng. Informatics | 2 |
| 2005 | A simulation and data analysis system for large-scale, data-driven oil reservoir simulation studiesabstractAbstract The main goal of oil reservoir management is to provide more efficient, cost‐effective and environmentally safer production of oil from reservoirs. Numerical simulations can aid in the design and implementation of optimal production strategies. However, traditional simulation‐based approaches to optimizing reservoir management are rapidly overwhelmed by data volume when large numbers of realizations are sought using detailed geologic descriptions. In this paper, we describe a software architecture to facilitate large‐scale simulation studies, involving ensembles of long‐running simulations and analysis of vast volumes of output data. Copyright © 2005 John Wiley & Sons, Ltd. Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz, Ryan Martino, Mary F. Wheeler, Malgorzata Peszynska, Alan Sussman, Christian Hansen 0002, Mrinal K. Sen, Roustam Seifoullaev, Paul L. Stoffa, Carlos Torres-Verdín, Manish Parashar |
Concurr. Pract. Exp. | 14 |
| 2005 | Autonomic oil reservoir optimization on the GridabstractThe emerging Grid infrastructure and its support for seamless and secure interactions is enabling a new generation of autonomic applications where the application components, Grid services, resources, and data interact as peers to manage, adapt and optimize themselves and the overall application. In this paper we describe the design, development and operation of a prototype of such an application that uses peer-to-peer interactions between distributed services and data on the Grid to enable the autonomic optimization of an oil reservoir. Copyright © 2005 John Wiley & Sons, Ltd. Vincent Matossian, Viraj Bhat, Manish Parashar, Malgorzata Peszynska, Mrinal K. Sen, Paul L. Stoffa, Mary F. Wheeler |
Concurr. Pract. Exp. | 3 |
| 2005 | Enabling interactive and collaborative oil reservoir simulations on the GridabstractAbstract Grid‐enabled infrastructures and problem‐solving environments can significantly increase the scale, cost‐effectiveness and utility of scientific simulations, enabling highly accurate simulations that provide in‐depth insight into complex phenomena. This paper presents a prototype of such an environment, i.e. an interactive and collaborative problem‐solving environment for the formulation, development, deployment and management of oil reservoir and environmental flow simulations in computational Grid environments. The project builds on three independent research efforts: (1) the IPARS oil reservoir and environmental flow simulation framework; (2) the NetSolve Grid engine; and (3) the Discover Grid‐based computational collaboratory. Its primary objective is to demonstrate the advantages of an integrated simulation infrastructure towards effectively supporting scientific investigation on the Grid, and to investigate the components and capabilities of such an infrastructure. Copyright © 2005 John Wiley & Sons, Ltd. Manish Parashar, Rajeev Muralidhar, Wonsuck Lee, Dorian C. Arnold, Jack J. Dongarra, Mary F. Wheeler |
Concurr. Pract. Exp. | 1 |
| 2005 | Rule-based visualization in the Discover computational steering collaboratory
Hua Liu 0001, Lian Jiang, Manish Parashar, Deborah Silver |
Future Gener. Comput. Syst. | 3 |
| 2005 | Application of Grid-enabled technologies for solving optimization problems in data-driven reservoir studies
Manish Parashar, Hector Klie, Ümit V. Çatalyürek, Tahsin M. Kurç, Wolfgang Bangerth, Vincent Matossian, Joel H. Saltz, Mary F. Wheeler |
Future Gener. Comput. Syst. | 1 |
| 2005 | Towards autonomic application-sensitive partitioning for SAMR applications
Sumir Chandra, Manish Parashar |
J. Parallel Distributed Comput. | 2 |
| 2005 | Conceptual and Implementation Models for the GridabstractThe Grid is rapidly emerging as the dominant paradigm for wide area distributed application systems. As a result, there is a need for modeling and analyzing the characteristics and requirements of Grid systems and programming models. This work adopts the well-established body of models for distributed computing systems, which are based upon carefully stated assumptions or axioms, as a basis for defining and characterizing Grids and their programming models and systems. The requirements of programming Grid applications and the resulting requirements on the underlying virtual organizations and virtual machines are investigated. The assumptions underlying some of the programming models and systems currently used for Grid applications are identified and their validity in Grid environments is discussed. A more in-depth analysis of two programming systems, the Imperial College E-Science Networked Infrastructure (ICENI) and Accord, using the proposed definitions' structure is presented. Manish Parashar, James C. Browne |
Proc. IEEE | 1 |
| 2005 | Special Issue on Grid Computing
Manish Parashar, Craig A. Lee |
Proc. IEEE | 1 |
| 2004 | Using a Jini based desktop Grid for test vector compaction and a refined economic modelabstractTesting of very large scale integrated (VLSI) circuits is done by designing huge random sets of vectors of which only a few are useful. The process of filtering these good vectors from the overall set is called vector compaction. As the integrated circuits become denser, this problem is becoming a major bottleneck and the computation time could run into days for a single chip. In this paper we demonstrate how a distributed Grid architecture can be used for the speed-up of the problem. The architecture is based on a Jini desktop Grid. Economic models become a prime issue in such a scenario. The current economic models for minimizing cost or time are ad-hoc and entirely under the control of the broker middleware architecture, leaving the end user with little choice. In this paper we present a revised economic model that gives more choice to the user in terms of time and cost before execution. We match the high-level application layer to the available resources by utilizing a system of composite performance modeling of the available resources. We demonstrate the performance of this new architecture on some VLSI benchmark circuits. Tezaswi Raja, Manish Parashar |
CCGRID | 2 |
| 2004 | Understanding the Behavior and Performance of Non-blocking Communications in MPI
Taher Saif, Manish Parashar |
Euro-Par | 2 |
| 2004 | A Dynamic Geometry-Based Shared Space Interaction Framework for Parallel Scientific Applications
Manish Parashar |
HiPC | 2 |
| 2004 | Hierarchical Partitioning Techniques for Structured Adaptive Mesh Refinement Applications
Xiaolin Li 0001, Manish Parashar |
J. Supercomput. | 2 |
| 2004 | A Peer-to-Peer Approach to Web Service Discovery
Cristina Schmidt, Manish Parashar |
World Wide Web | 2 |
| 2003 | Dynamic Load Partitioning Strategies for Managing Data of Space and Time Heterogeneity in Parallel SAMR Applications
Xiaolin Li 0001, Manish Parashar |
Euro-Par | 2 |
| 2003 | DIOS++: A Framework for Rule-Basedn Autonomic Management of Distributed Scientific Applications
Hua Liu 0001, Manish Parashar |
Euro-Par | 2 |
| 2003 | Enabling Peer-to-Peer Interactions for Scientific Applications on the Grid
Vincent Matossian, Manish Parashar |
Euro-Par | 2 |
| 2003 | A Middleware Substrate for Integrating Services on the Grid
Viraj Bhat, Manish Parashar |
HiPC | 2 |
| 2003 | Flexible Information Discovery in Decentralized Distributed SystemsabstractThe ability to efficiently discover information using partial knowledge (for example keywords, attributes or ranges) is important in large, decentralized, resource sharing distributed environments such as computational grids and peer-to-peer (P2P) storage and retrieval systems. This paper presents a P2P information discovery system that supports flexible queries using partial keywords and wildcards, and range queries. It guarantees that all existing data elements that match a query are found with bounded costs in terms of number of messages and number of peers involved. The key innovation is a dimension reducing indexing scheme that effectively maps the multidimensional information space to physical peers. The design, implementation and experimental evaluation of the system are presented. Cristina Schmidt, Manish Parashar |
HPDC | 2 |
| 2003 | A distributed object infrastructure for interaction and steeringabstractAbstract This paper presents the design, implementation and experimental evaluation of DIOS (Distributed Interactive Object Substrate), an interactive object infrastructure to enable the runtime monitoring, interaction and computational steering of parallel and distributed applications. DIOS enables application objects (data structures, algorithms) to be enhanced with sensors and actuators so that they can be interrogated and controlled. Application objects may be distributed (spanning many processors) and dynamic (be created, deleted, changed or migrated). Furthermore, DIOS provides a control network that interconnects the interactive objects in a parallel/distributed application and enables external discovery, interrogation, monitoring and manipulation of these objects at runtime. DIOS is currently being used to enable interactive visualization, monitoring and steering of a wide range of scientific applications, including oil reservoir, compressible turbulence and numerical relativity simulations. Copyright © 2003 John Wiley & Sons, Ltd. Rajeev Muralidhar, Manish Parashar |
Concurr. Comput. Pract. Exp. | 2 |
| 2003 | Optimizing Web servers using Page rank prefetching for clustered accesses
Victor Safronov, Manish Parashar |
Inf. Sci. | 2 |
| 2002 | Latency Performance of SOAP ImplementationsabstractThis paper presents an experimental evaluation of the latency performance of several implementations of Simple Object Access Protocol (SOAP) operating over HTTP, and compares these results with the performance Of JavaRMI, CORBA, HTTP, and with the TCP setup time. SOAP is an XML based protocol that supports RPC and message semantics. While SOAP has been designed as an interoperable business-to-business protocol usable over the Internet, we believe that applications will also use SOAP for interactive web applications running within an intranet. The objective of this paper is to identify the sources of inefficiency in the current implementations of SOAP and discuss changes that can improve their performance. SOAP implementations studied include Microsoft SOAP Toolkit, the SOAP: Lite Perl module, and Apache SOAP. Dan Davis, Manish Parashar |
CCGRID | 2 |
| 2002 | A Study of Discovery Mechanisms for Peer-to-Peer ApplicationsabstractPeer-to-peer applications allow peers to connect or disconnect from a network at any time and are based on a loosely coupled resource distribution model. As a result, robust and efficient discovery mechanisms are central to the efficient functioning of such applications. In this paper we evaluate four discovery mechanisms (flooding and the forward routing algorithms CHORD, Pastry and CAN) against the requirements of three prevalent classes of peer-to-peer applications, and investigate the suitability of these mechanisms for the applications. Mandar Kelaskar, Vincent Matossian, Preeti Mehra, Dennis Paul, Manish Parashar |
CCGRID | 5 |
| 2002 | Adaptive Runtime Managementof SAMR Applications
Sumir Chandra, Shweta Sinha, Manish Parashar, Yeliang Zhang, Jingmei Yang, Salim Hariri |
HiPC | 3 |
| 2002 | Engineering an interoperable computational collaboratory on the GridabstractAbstract The growth of the Internet and the advent of the computational Grid have made it possible to develop and deploy advanced computational collaboratories. These systems build on high‐end computational resources, communication technologies and enabling services underlying the Grid, and provide seamless and collaborative access to resources, applications and data. Combining these focused collaboratories and allowing them to interoperate has many advantages and can lead to truly collaborative, multidisciplinary and multi‐institutional problem solving. However, integrating these collaboratories presents significant challenges, as each of these collaboratories has a unique architecture and implementation, and builds on different enabling technologies. This paper investigates the issues involved in integrating collaboratories operating on the Grid. It then presents the design and implementation of a prototype middleware substrate to enable a peer‐to‐peer integration of and global access to multiple, geographically distributed instances of the DISCOVER computational collaboratory. An experimental evaluation of the middleware substrate is presented. Copyright © 2002 John Wiley & Sons, Ltd. Vijay Mann, Manish Parashar |
Concurr. Comput. Pract. Exp. | 2 |
| 2002 | Special Issue: Software architectures for scientific applicationsabstractComputational Science and Engineering (CS&E) is widely accepted, along with theory and experiment, as a crucial third mode of scientific investigation and engineering design.Most scientific and engineering disciplines now rely on simulation, and scientific application and system software is rapidly becoming an accepted modality for research and instruction.Recent trends have seen the evolution of scientific application software from stand-alone monolithic programs to more complex application frameworks that support families of applications and multiple user communities.These frameworks are aimed at promoting sharing of software modules, 'plug-and-play' composition, reusability, inter-operability, maintainability, portability, and extensibility, without sacrificing high performance or implementation independence.Software architecture is the identification and organization of software components, and the interaction between these components, to meet similar requirements.While software architectures have received a lot of attention in the software engineering community, their use by computational scientists is limited.One explanation for this is the uniqueness of scientific software in its inherent complexity, its heterogeneity, its rapid evolution, and its demand for very high performance.The overall goal of the SIAM CS&E conference was to draw attention to the tremendous range of major computational efforts on large problems in science and engineering, to promote the interdisciplinary culture required to meet these large-scale challenges, and to encourage the training of the next generation of computational scientists.The Software Architectures for Scientific Applications mini-symposium was aimed at exploring software architecture used by existing scientific applications and application frameworks, and investigating the applicability of software engineering practices to these software systems.Its objective was to bring together current approaches and experiences in designing large software systems for scientific applications. Manish Parashar |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | A CORBA Commodity Grid KitabstractAbstract This paper reports on an ongoing research project aimed at designing and deploying a Common Object Resource Broker Architecture (CORBA) (ww.omg.org) Commodity Grid (CoG) Kit. The overall goal of this project is to enable the development of advanced Grid applications while adhering to state‐of‐the‐art software engineering practices and reusing the existing Grid infrastructure. As part of this activity, we are investigating how CORBA can be used to support the development of Grid applications. In this paper, we outline the design of a CORBA CoG Kit that will provide a software development framework for building a CORBA ‘Grid domain’. We also present our experiences in developing a prototype CORBA CoG Kit that supports the development and deployment of CORBA applications on the Grid by providing them access to the Grid services provided by the Globus Toolkit. Copyright © 2002 John Wiley & Sons, Ltd. Manish Parashar, Gregor von Laszewski, Snigdha Verma, Jarek Gawor, Kate Keahey, Nell Rehn |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | An Application-Centric Characterization of Domain-Based SFC Partitioners for Parallel SAMRabstractStructured adaptive mesh refinement (SAMR) methods for the numerical solution of partial differential equations yield highly advantageous ratios for cost/accuracy as compared to methods based on static uniform approximations. These techniques are being effectively used in many domains including computational fluid dynamics, numerical relativity, astrophysics, subsurface modeling, and oil reservoir simulation. Distributed implementations of these methods, however, lead to significant challenges in dynamic data-distribution, load-balancing, and runtime management. This paper presents an application-centric characterization of a suite of dynamic domain-based inverse space-filling curve partitioning techniques for the distributed adaptive grid hierarchies that underlie SAMR applications. The overall goal of this research is to formulate policies required to drive a dynamically adaptive metapartitioner for SAMR grid hierarchies capable of selecting the most appropriate partitioning strategy at runtime based on current application and system state. Such a metapartitioner can significantly reduce the execution time of SAMR applications. Johan Steensland, Sumir Chandra, Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2002 | Optimizing Web Servers Using Page Rank Prefetching for Clustered Accesses
Victor Safronov, Manish Parashar |
World Wide Web | 2 |
| 2001 | Adaptive Cluster Computing using JavaSpacesabstractWe present the design, implementation and evaluation of a framework that uses JavaSpaces [1] to support this type of opportunistic adaptive parallel/distributed computing over networked clusters in a non-intrusive manner. The framework targets applications exhibiting coarse-grained parallelism and has three key features: (1) portability across heterogeneous platforms, (2) minimal configuration overheads for participating nodes, and (3) automated system state monitoring (using SNMP) to ensure nonintrusive behavior. Experimental results presented in this paper demonstrate that for applications that can be broken into coarse-grained, relatively independent tasks, the opportunistic adaptive parallel computing framework can provide performance gains. Furthermore, the results indicate that monitoring and reacting to the current system state minimizes the intrusiveness of the framework. Jyoti Batheja, Manish Parashar |
CLUSTER | 2 |
| 2001 | Adaptive Runtime Partitioning of AMR Applications on Heterogeneous ClustersabstractThis paper presents the design and evaluation of an adaptive, system sensitive partitioning and load balancing framework for distributed structured adaptive mesh refinement applications on heterogeneous and dynamic cluster environments. The framework uses system capabilities and current system state to select and tune appropriate partitioning parameters (e.g. partitioning granularity, load per processor) to maximize overall application performance. Shweta Sinha, Manish Parashar |
CLUSTER | 2 |
| 2001 | An Evaluation of Partitioners for Parallel SAMR Applications
Sumir Chandra, Manish Parashar |
Euro-Par | 2 |
| 2001 | A Distributed Object Infrastructure for Interaction and Steering
Rajeev Muralidhar, Manish Parashar |
Euro-Par | 2 |
| 2001 | Middleware Support for Global Access to Integrated Computational CollaboratoriesabstractThe growth of the Internet and the advent of the computational "Grid" have made it possible to develop and deploy advanced computational collaboratories. These systems build on high-end computational resources and communication technologies underlying the Grid, and provide seamless and collaborative access to particular resources, services or applications. Integrating these "focused" collaboratories presents significant challenges. Key among these is the design and development of robust middleware support that addresses scalability, service discovery, security and access control, and interaction and collaboration management for consistent access. The authors first investigate the architecture of such a middleware that enables global (Web-based) access to collaboratories. They then present the design and implementation of a middleware substrate that enables a peer-to-peer integration of and global (collaborative) access to geographically distributed instances of the DISCOVER computational collaboratory for interaction and steering. Vijay Mann, Manish Parashar |
HPDC | 2 |
| 2001 | System Sensitive Runtime Management of Adaptive ApplicationsabstractThis paper presents the design and evaluation of a system sensitive partitioning and load balancing framework for distributed adaptive grid hierarchies that underlie parallel adaptive mesh-refinement (AMR) techniques for the solution of partial-differential equations. The framework uses system capabilities and current system state to select and tune appropriate distribution parameters (e.g. partitioning granularity, load per processor) to maximize overall application performance. Shweta Sinha, Manish Parashar |
IPDPS | 2 |
| 2001 | DISCOVER: An environment for Web-based interaction and steering of high-performance scientific applicationsabstractAbstract This paper presents the design, implementation, and deployment of the DISCOVER Web‐based computational collaboratory. Its primary goal is to bring large distributed simulations to the scientists'/engineers' desktop by providing collaborative Web‐based portals for monitoring, interaction and control. DISCOVER supports a three‐tier architecture composed of detachable thin‐clients at the front‐end, a network of interaction servers in the middle, and a control network of sensors, actuators, and interaction agents at the back‐end. The interaction servers enable clients to connect and collaboratively interact with registered applications using a browser. The application control network enables sensors and actuators to be encapsulated within, and directly deployed with the computational objects. The application interaction gateway manages overall interaction. It uses Java Native Interface to create Java proxies that mirror computational objects and allow them to be directly accessed at the interaction server. Security and authentication are provided using customizable access control lists and SSL‐based secure servers. Copyright © 2001 John Wiley & Sons, Ltd. Vijay Mann, Vincent Matossian, Rajeev Muralidhar, Manish Parashar |
Concurr. Comput. Pract. Exp. | 4 |
| 2000 | An Architecture for Web-Based Interaction and Steering of Adaptive Parallel/Distributed Applications
Rajeev Muralidhar, Samian Kaur, Manish Parashar |
Euro-Par | 3 |
| 2000 | An Object Infrastructure for Computational Steering of Distributed Simulations
Rajeev Muralidhar, Manish Parashar |
HPDC | 2 |
| 2000 | Interpretive Performance Prediction for Parallel Application Development
Manish Parashar, Salim Hariri |
J. Parallel Distributed Comput. | 1 |
| 1997 | A Common Data Management Infrastructure for Adaptive Algorithms for PDE SolutionsabstractThis paper presents the design, development and application of a computational infrastructure to support the implementation of parallel adaptive algorithms for the solution of sets of partial differential equations. The infrastructure is separated into multiple layers of abstraction. This paper is primarily concerned with the two lowest layersof this infrastructure: a layer which defines and implements dynamic distributed arrays (DDA), and a layer in which several dynamic data and programming abstractions are implemented in terms of the DDAs. The currently implemented abstractions are those needed for formulation of hierarchical adaptive finite difference methods, hp-adaptive finite element methods, and fast multipole method for solution of linear systems. Implementation of sample applications based on each of these methods are described and implementation issues and performance measurements are presented. Manish Parashar, James C. Browne, H. Carter Edwards, Kenneth Klimkowski |
SC | 1 |
| 1995 | Software Tool Evaluation MethodologyabstractThe recent development of parallel and distributed computing software has introduced a variety of software tools that support several programming paradigms and languages. This variety of tools makes the selection of the best tool to run a given class of applications on a parallel or distributed system a non-trivial task that requires some investigation. We expect tool evaluation to receive more attention as the deployment and usage of distributed systems increases. In this paper, we present a multi-level evaluation methodology for parallel/distributed tools in which tools are evaluated from different perspectives. We apply our evaluation methodology to three message passing tools viz Express, p4, and PVM. The approach covers several important distributed systems platforms consisting of different computers (e.g., IBM-SP1, Alpha cluster, SUN workstations) interconnected by different types of networks (e.g., Ethernet, FDDI, ATM). Salim Hariri, Sungyong Park, Rajashekar Reddy, Mahesh Subramanyan, Rajesh Yadav, Geoffrey C. Fox, Manish Parashar |
ICDCS | 7 |
| 1994 | Interpreting the performance of HPF/Fortran 90DabstractWe present a novel interpretive approach for accurate and cost effective performance prediction in a high performance computing environment, and describe the design of a source driven HPF/Fortran 90D performance prediction framework based on this approach. The performance prediction framework has been implemented as part of a HPF/Fortran 90D application development environment. A set of benchmarking kernels and application codes are used to validate the accuracy, utility, usability, and cost effectiveness of the performance prediction framework. The use of the framework for selecting appropriate compiler directives and for application performance debugging is demonstrated.> Manish Parashar, Salim Hariri, Tomasz Haupt, Geoffrey C. Fox |
SC | 1 |
| 1994 | Communication system for high-performance distributed computingabstractAbstract With the current advances in computer and networking technology coupled with the availability of software tools for parallel and distributed computing, there has been increased interest in high‐performance distributed computing (HPDC). We envision that HPDC environments with supercomputing capabilities will be available in the near future. However, a number of issues have to be resolved before future network‐based applications can fully exploit the potential of the HPDC environment. In the paper we present an architecture for a high‐speed local area network and a communication system that provides HPDC applications with high bandwidth and low latency. We also characterize the message‐passing primitives required in HPDC applications and develop a communication protocol that implements these primitives efficiently. Salim Hariri, J.-B. Park, Manish Parashar, Geoffrey C. Fox |
Concurr. Pract. Exp. | 3 |
| 1993 | A Message Passing Interface for Parallel and Distributed ComputingabstractThe proliferation of high performance workstations and the emergence of high speed networks have attracted a lot of interest in parallel and distributed computing (PDC). The authors envision that PDC environments with supercomputing capabilities will be available in the near future. However, a number of hardware and software issues have to be resolved before the full potential of these PDC environments can be exploited. The presented research has the following objectives: (1) to characterize the message-passing primitives used in parallel and distributed computing; (2) to develop a communication protocol that supports PDC; and (3) to develop an architectural support for PDC over gigabit networks.> Salim Hariri, Jong Park, Fang-Kuo Yu, Manish Parashar, Geoffrey C. Fox |
HPDC | 4 |
| 1992 | A Requirement Analysis for High Performance Distributed Computing over LANsabstractWith the proliferation of high performance workstations and the current trend towards high speed communication networks. the cumulative computing power provided by a group of general purpose workstations is comparable to supercomputers. However a number of obstacles have to be overcome before the full potential of these network-based distributed systems can be exploited. This paper investigates the requirements of current workstation clusters interconnected by local area networks (LANs) which would allow them to be used as platforms for high performance distributed computing. The blocked LU decomposition of dense matrices is used as the running example in the presented study. Performance of this algorithm is measured on the iPSC/860 hypercube and on a set of homogeneous workstations (SUN SPARCstation 1+) interconnected by Ethernet. These measures are analyzed and a set of requirements are identified which would enable a network of workstations to deliver high performance distributed computing.> Manish Parashar, Salim Hariri, A. Gaber Mohamed, Geoffrey C. Fox |
HPDC | 1 |
| 1992 | On the Parallelization of Blocked LU Factorization Algorithms on Distributed Memory ArchitecturesabstractThe authors present the parallelization of blocked algorithms for LU factorization. They isolate problems inherent in sequential blocked algorithms and provide approaches to overcome them on distributed memory architectures. The performances of the parallelized versions of three blocked algorithms suited to column oriented Fortran are compared. Experiments are performed on the iPSC/860 hypercube. It is shown that it is not intuitively clear which algorithm might perform best on a given architecture; this is dependent on the problem size and the number of available parameters.> Gregor von Laszewski, Manish Parashar, A. Gaber Mohamed, Geoffrey C. Fox |
SC | 2 |