VLDB 2026 Research / reviewers in the wild / expert
Anirban Mandal
dblp:31/6757
· DBLP profile ↗
31ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Computer networks · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AeroResQ: Edge-accelerated UAV framework for scalable, resilient and collaborative escape route planning in wildfire scenarios
Suman Raj, Radhika Mittal, Rajiv Mayani, Pawel Zuk, Anirban Mandal, Michael Zink, Yogesh L. Simmhan, Ewa Deelman |
Future Gener. Comput. Syst. | 5 |
| 2025 | A Greedy Consensus-Based Approach to Distributed Job Selection: Toward Fully-Decentralized Workload Management SystemabstractCurrent approaches to resilience for highly distributed, heterogeneous, large-scale scientific workflows are limited. Most existing workflow and resource management systems have a single point of failure and resilience strategies are often static, depend on a centralized control, and require considerable design effort from experts. The increasing scale and complexity of workflows coupled with limited resilience capabilities in centralized systems necessitates a fully decentralized, adaptive resource management approach. This paper addresses a very important slice of the overall problem by leveraging the advances in multi-agent systems (MAS). In particular, we explore the suitability of a MAS consisting of globally distributed agents to perform distributed job selection from a dynamic job pool in a truly decentralized, performant, and resilient manner. We present a novel consensus formulation of the distributed job selection problem. By introducing a cost function encapsulating the requirements and constraints of the job and resource loads, we design a novel, greedy consensus algorithm leveraging the Practical Byzantine Fault Tolerance (PBFT)-based consensus method, allowing agents to collectively select jobs in a resilient manner. We compared our algorithms with other state of the art approaches by deploying them in a network testbed infrastructure to emulate distributed job selection. Our evaluation results demonstrated that our greedy consensus algorithm employing the cost-function and PBFT-based consensus method outperforms the ones using the vanilla PBFT-based consensus method - improving scheduling latency by as much as 63.5 % and reducing resource idle time by as much as 63.8 %, with benefits increasing with higher numbers of agents emulated. Komal Thareja, Raghavan Krishnan, Anirban Mandal, Pawel Zuk, Imtiaz Mahmud, Mariam Kiran, Ewa Deelman |
CCGrid | 3 |
| 2025 | Advancing anomaly detection in computational workflows with active learning
Raghavan Krishnan, George Papadimitriou 0002, Anirban Mandal, Mariam Kiran, Prasanna Balaprakash, Ewa Deelman |
Future Gener. Comput. Syst. | 4 |
| 2024 | DISTRI: Development and Integration of Simulation Tools for Resilient InfrastructureabstractIn contemporary scientific research, data acquisition and analysis platforms have grown increasingly complex, often spanning multiple facilities with diverse internal structures. Efficiently managing the interactions between job scheduling, resource allocation, and networking across these distributed systems requires a robust simulation framework. However, existing simulators fall short in capturing the detailed interactions necessary for comprehensive analysis of large-scale distributed environments. To address this gap, we introduce DISTRI, a versatile framework specifically designed for the development and testing of distributed multi-facility workflows. DISTRI allows for customizable facility configurations and includes built-in support for distributed, resilient scheduling and resource management, alongside detailed network simulation for data communication. Key features of DISTRI encompass inter- and intra-facility resource management, agent-based distributed scheduling, and extensive performance metrics logging for both resource and network management. By providing these essential tools, DISTRI enables thorough analysis and optimization, thereby advancing research in the resilience and efficiency of multi-facility systems. Imtiaz Mahmud, Pawel Zuk, Cong Wang 0014, Mariam Kiran, Kesheng Wu, Komal Thareja, Raghavan Krishnan, Anirban Mandal, Ewa Deelman |
IEEE Big Data | 8 |
| 2024 | Large Language Models for Anomaly Detection in Computational Workflows: From Supervised Fine-Tuning to In-Context LearningabstractAnomaly detection in computational workflows is critical for ensuring system reliability and security. However, traditional rule-based methods struggle to detect novel anomalies. This paper leverages large language models (LLMs) for workflow anomaly detection by exploiting their ability to learn complex data patterns. Two approaches are investigated: (1) supervised fine-tuning (SFT), where pretrained LLMs are fine-tuned on labeled data for sentence classification to identify anomalies, and (2) in-context learning (ICL), where prompts containing task descriptions and examples guide LLMs in few-shot anomaly detection without fine-tuning. The paper evaluates the performance, efficiency, and generalization of SFT models and explores zeroshot and few-shot ICL prompts and interpretability enhancement via chain-of-thought prompting. Experiments across multiple workflow datasets demonstrate the promising potential of LLMs for effective anomaly detection in complex executions. George Papadimitriou 0002, Raghavan Krishnan, Pawel Zuk, Prasanna Balaprakash, Cong Wang 0014, Anirban Mandal, Ewa Deelman |
SC | 7 |
| 2023 | FlyPaw: Optimized Route Planning for Scientific UAVMissionsabstractMany Internet of Things (IoT) applications require compute resources that cannot be provided by the devices themselves. On the other hand, processing of the data generated by IoT devices and sensors often has to be performed in real- or near real-time, i.e., with stringent latency requirements in constrained environments (e.g., intermittent network connectivity and limited power envelopes). Examples of such scenarios are autonomous vehicles in the form of cars and drones where the processing and analysis of observational data (e.g., video feeds) need to be performed expeditiously to allow for safe operation of the vehicles and to deliver the results in a timely fashion to the stakeholders of the mission. To support the compute and timeliness requirements of such applications, it is essential to include suitable edge resources to process these workflows, and to develop an end-to-end system that can route the vehicles dynamically and process and deliver mission-critical data and analyzed results. In this paper, we develop and evaluate a dynamic scheduling approach that considers complex tradeoffs between real-time constraints, network availability, and latency sensitivity of the mission. We devise an optimized route planning and data transmission schedule for drone flights. The scheduling algorithm is encapsulated in a novel end-to-end architecture (FlyPaw) and an associated adaptive drone mission control system, which enables deployment and management of an integrated cyberphysical system (CPS) – from real drone testbed to base stations to edge-to-cloud resources. The planning algorithm takes into account measured network communication characteristics, estimated uncertainties of future data link connectivity, and data timeliness requirements of the mission to prioritize candidate decision tree solutions based on a risk metric derived from Sharpe's ratio. Our results show that for given task sets, Net Time to Retrieve, our metric describing the time required to perform end-to-end collection and downstream processing of data, can be significantly reduced compared to other naive approaches. The theoretical improvement provided by our algorithm over other naive approaches is dependent on several factors — task locations, network connectivity, processing times and available resources, and is bounded by the duration of the drone flight. Andrew Grote, Eric Lyons 0001, Komal Thareja, George Papadimitriou 0002, Ewa Deelman, Anirban Mandal, Prasad Calyam, Michael Zink |
e-Science | 6 |
| 2022 | Automating Edge-to-cloud Workflows for Science: Traversing the Edge-to-cloud Continuum with PegasusabstractIn this paper, we describe how we extended the Pegasus Workflow Management System to support edge-to-cloud workflows in an automated fashion. We discuss how Pegasus and HTCondor (its job scheduler) work together to enable this automation. We use HTCondor to form heterogeneous pools of compute resources and Pegasus to plan the workflow onto these resources and manage containers and data movement for executing workflows in hybrid edge-cloud environments. We then show how Pegasus can be used to evaluate the execution of workflows running on edge only, cloud only, and edge-cloud hybrid environments. Using the Chameleon Cloud testbed to set up and configure an edge-cloud environment, we use Pegasus to benchmark the executions of one synthetic workflow and two production workflows: CASA-Wind and the Ocean Observatories Initiative Orcasound workflow, all of which derive their data from edge devices. We present the performance impact on workflow runs of job and data placement strategies employed by Pegasus when configured to run in the above three execution environments. Results show that the synthetic workflow performs best in an edge only environment, while the CASA - Wind and Orcasound workflows see significant improvements in overall makespan when run in a cloud only environment. The results demonstrate that Pegasus can be used to automate edge-to-cloud science workflows and the workflow provenance data collection capabilities of the Pegasus monitoring daemon enable computer scientists to conduct edge-to-cloud research. Ryan Tanaka, George Papadimitriou 0002, Sai Charan Viswanath, Cong Wang 0014, Eric Lyons 0001, Komal Thareja, Chengyi Qu, Alicia Esquivel Morel, Ewa Deelman, Anirban Mandal, Prasad Calyam, Michael Zink |
CCGRID | 10 |
| 2022 | Data Integrity Error Localization in Networked Systems with Missing DataabstractMost recent network failure diagnosis systems focused on data center networks where complex measurement systems can be deployed to derive routing information and ensure network coverage in order to achieve accurate and fast fault localization. In this paper, we target wide-area networks that support data-intensive distributed applications. We first present a new multi-output prediction model that directly maps the application level observations to localize the system component failures. In reality, this application-centric approach may face the missing data challenge as some input (feature) data to the inference models may be missing due to incomplete or lost measurements in wide area networks. We show that the presented prediction model naturally allows the multivariate imputation to recover the missing data. We evaluate multiple imputation algorithms and show that the prediction performance can be improved significantly in a large-scale network. As far as we know, this is the first study on the missing data issue and applying imputation techniques in network failure localization. Yufeng Xin, Shih-Wen Fu, Anirban Mandal, Ryan Tanaka, Mats Rynge, Karan Vahi, Ewa Deelman |
ICC | 3 |
| 2021 | WIRE: Resource-efficient Scaling with Online Prediction for DAG-based WorkflowsabstractThis paper introduces WIRE that manages resources for the DAG-based workflows on IaaS clouds. WIRE predicts and plans resources over the MAPE (Monitor-Analyze-Plan-Execute) loops to: 1) Estimate task performance with online data, 2) Conduct simulations to predict the upcoming loads based on online estimates and workflow DAGs, 3) Apply a resource-steering policy to size cloud instance pools for the maximal parallelism that is consistent with low cost. We implement WIRE on Pegasus WMS/HTCondor and evaluate its performance on the ExoGENI network cloud. The results show that WIRE attains low resource cost with the performance that is typically within a factor of two of optimal. Qiang Cao 0005, Mayuresh Kunjir, Linli Wan, Jeffrey S. Chase, Anirban Mandal, Mats Rynge |
CLUSTER | 6 |
| 2021 | Predicting Flash Floods in the Dallas-Fort Worth Metroplex Using Workflows and Cloud ComputingabstractAccurate and timely prediction of flash flooding events can be a very useful tool for stormwater officials and first responders. Having lead time with which to issue evacuation directives, to close flood prone roadways, to deploy rescue gear and personnel, and to fortify areas against flooding is essential to minimize property damage and risk of casualties. In this poster, we are presenting a flash flooding prediction workflow based on the Hydrology Lab-Research Distributed Hydrologic Model (HL-RDHM). This workflow leverages cloud computing and the Pegasus Workflow Management System to provide continuous high resolution flood predictions for the Dallas-Fort Worth Metroplex area in North Texas, and can be easily expanded to other regions. Eric Lyons 0001, Dong-Jun Seo, Sunghee Kim, Hamideh Habibi, George Papadimitriou 0002, Ryan Tanaka, Ewa Deelman, Michael Zink, Anirban Mandal |
e-Science | 9 |
| 2021 | Mining Workflows for Anomalous Data TransfersabstractModern scientific workflows are data-driven and are often executed on distributed, heterogeneous, high-performance computing infrastructures. Anomalies and failures in the work-flow execution cause loss of scientific productivity and inefficient use of the infrastructure. Hence, detecting, diagnosing, and mitigating these anomalies are immensely important for reliable and performant scientific workflows. Since these workflows rely heavily on high-performance network transfers that require strict QoS constraints, accurately detecting anomalous network performance is crucial to ensure reliable and efficient workflow execution. To address this challenge, we have developed X-FLASH, a network anomaly detection tool for faulty TCP workflow transfers. X-FLASH incorporates novel hyperparameter tuning and data mining approaches for improving the performance of the machine learning algorithms to accurately classify the anomalous TCP packets. X-FLASH leverages XGBoost as an ensemble model and couples XGBoost with a sequential optimizer, FLASH, borrowed from search-based Software Engineering to learn the optimal model parameters. X-FLASH found configurations that outperformed the existing approach up to 28%, 29%, and 40% relatively for F-measure, G-score, and recall in less than 30 evaluations. From (1) large improvement and (2) simple tuning, we recommend future research to have additional tuning study as a new standard, at least in the area of scientific workflow anomaly detection. Huy Tu, George Papadimitriou 0002, Mariam Kiran, Cong Wang 0014, Anirban Mandal, Ewa Deelman, Tim Menzies |
MSR | 5 |
| 2021 | Special issue on workflows in Support of Large-Scale Science
Anirban Mandal, Raffaele Montella |
Future Gener. Comput. Syst. | 1 |
| 2021 | End-to-end online performance data capture and analysis for scientific workflows
George Papadimitriou 0002, Cong Wang 0014, Karan Vahi, Rafael Ferreira da Silva, Anirban Mandal, Zhengchun Liu, Rajiv Mayani, Mats Rynge, Mariam Kiran, Vickie E. Lynch, Rajkumar Kettimuthu, Ewa Deelman, Jeffrey S. Vetter, Ian T. Foster |
Future Gener. Comput. Syst. | 5 |
| 2020 | Detecting anomalous packets in network transfers: investigations using PCA, autoencoder and isolation forest in TCP
Mariam Kiran, Cong Wang 0014, George Papadimitriou 0002, Anirban Mandal, Ewa Deelman |
Mach. Learn. | 4 |
| 2019 | Cyberinfrastructure Center of Excellence Pilot: Connecting Large Facilities CyberinfrastructureabstractThe National Science Foundation's Large Facilities are major, multi-user research facilities that operate and manage sophisticated and diverse research instruments and platforms (e.g., large telescopes, interferometers, distributed sensor arrays) that serve a variety of scientific disciplines, from astronomy and physics to geology and biology and beyond. Large Facilities are increasingly dependent on advanced cyberinfrastructure (i.e., computing, data, and software systems; networking; and associated human capital) to enable the broad delivery and analysis of facility-generated data. These cyberinfrastructure tools enable scientists and the public to gain new insights into fundamental questions about the structure and history of the universe, the world we live in today, and how our environment may change in the coming decades. This paper describes a pilot project that aims to develop a model for a Cyberinfrastructure Center of Excellence (CI CoE) that facilitates community building and knowledge sharing and that disseminates and applies best practices and innovative solutions for facility CI. Ewa Deelman, Ryan Mitchell, Loïc Pottier, Mats Rynge, Erik Scott, Karan Vahi, Marina Kogan, Jasmine Mann, Tom Gulbransen, Daniel Allen, David Barlow, Anirban Mandal, Santiago Bonarrigo, Chris Clark, Leslie Goldman, Tristan Goulden, Phil Harvey, David Hulsander, Steve Jacobs, Christine Laney, Ivan Lobo-Padilla, Jeremy Sampson, Valerio Pascucci, John Staarmann, Steve Stone, Susan Sons, Jane Wyngaard, Charles Vardeman, Steve Petruzza, Ilya Baldin, Laura Christopherson |
eScience | 12 |
| 2019 | Toward a Dynamic Network-Centric Distributed Cloud Platform for Scientific Workflows: A Case Study for Adaptive Weather SensingabstractComputational science today depends on complex, data-intensive applications operating on datasets from a variety of scientific instruments. A major challenge is the integration of data into the scientist's workflow. Recent advances in dynamic, networked cloud resources provide the building blocks to construct reconfigurable, end-to-end infrastructure that can increase scientific productivity. However, applications have not adequately taken advantage of these advanced capabilities. In this work, we have developed a novel network-centric platform that enables high-performance, adaptive data flows and coordinated access to distributed cloud resources and data repositories for atmospheric scientists. We demonstrate the effectiveness of our approach by evaluating time-critical, adaptive weather sensing workflows, which utilize advanced networked infrastructure to ingest live weather data from radars and compute data products used for timely response to weather events. The workflows are orchestrated by the Pegasus workflow management system and were chosen because of their diverse resource requirements. We show that our approach results in timely processing of Nowcast workflows under different infrastructure configurations and network conditions. We also show how workflow task clustering choices affect throughput of an ensemble of Nowcast workflows with improved turnaround times. Additionally, we find that using our network-centric platform powered by advanced layer2 networking techniques results in faster, more reliable data throughput, makes cloud resources easier to provision, and the workflows easier to configure for operational use and automation. Eric Lyons 0001, Anirban Mandal, George Papadimitriou 0002, Cong Wang 0014, Komal Thareja, Paul Ruth, Juan J. Villalobos, Ivan Rodero, Ewa Deelman, Michael Zink |
eScience | 2 |
| 2019 | Custom Execution Environments with Containers in Pegasus-Enabled Scientific WorkflowsabstractScience reproducibility is a cornerstone feature in scientific workflows. In most cases, this has been implemented as a way to exactly reproduce the computational steps taken to reach the final results. While these steps are often completely described, including the input parameters, datasets, and codes, the environment in which these steps are executed is only described at a higher level with endpoints and operating system name and versions. Though this may be sufficient for reproducibility in the short term, systems evolve and are replaced over time, breaking the underlying workflow reproducibility. A natural solution to this problem is containers, as they are well defined, have a lifetime independent of the underlying system, and can be user-controlled so that they can provide custom environments if needed. This paper highlights some unique challenges that may arise when using containers in distributed scientific workflows. Further, this paper explores how the Pegasus Workflow Management System implements container support to address such challenges. Karan Vahi, Michael Zink, Mats Rynge, George Papadimitriou 0002, Duncan A. Brown, Rajiv Mayani, Rafael Ferreira da Silva, Ewa Deelman, Anirban Mandal, Eric Lyons 0001 |
eScience | 9 |
| 2019 | COMET: Distributed Metadata Service for Multi-cloud ExperimentsabstractA majority of today’s cloud services are independently operated by individual cloud service providers. In this approach, the locations of cloud resources are strictly constrained by the distribution of cloud service providers’ sites. As the popularity and scale of cloud services increase, we believe this traditional paradigm is about to change toward further federated services, a.k.a., multi-cloud, due to the improved performance, reduced cost of compute, storage and network resources, as well as increased user demands. In this paper, we present COMET, a lightweight, distributed storage system for managing metadata on large scale, federated cloud infrastructure providers, end users, and their applications (e.g. HTCondor Cluster or Hadoop Cluster). We showcase use case from NSF’s, Chameleon, ExoGENI and JetStream research cloud testbeds to show the effectiveness of COMET design and deployment. Komal Thareja, Cong Wang 0014, Paul Ruth, Anirban Mandal, Ilya Baldin, Michael J. Stealey |
ICNP | 4 |
| 2017 | Toward Prioritization of Data Flows for Scientific Workflows Using Virtual Software Defined ExchangesabstractRecent advances in cloud systems, on-demand circuits and software-defined networking have created new opportunities to enable complex, data-intensive scientific applications to run on dynamic networked cloud infrastructures. In this work, we present an end-to-end framework for autonomic adaptation for scientific workflows on networked cloud systems, which leverages novel network provisioning technologies. We present an application-independent controller framework called Mobius++ that includes dynamic network adaptation capabilities using Software-Defined Networking (SDN) mechanisms, which enables workflow management systems to address competing priorities of workflow operations, data movements in particular. We use a representative, data-intensive bioinformatics workflow as a driving use case to showcase the above capabilities. Experimental results show that the Mobius++ framework, in conjunction with a novel virtual Software Defined Exchange (SDX) platform, is able to dynamically prioritize bandwidths between different end-points, on-demand, and being driven by priority directives from a workflow management system. We show that data transfer jobs from two workflows with different priorities are accurately arbitrated as the relative priorities change. Anirban Mandal, Paul Ruth, Ilya Baldin, Rafael Ferreira da Silva, Ewa Deelman |
eScience | 1 |
| 2014 | Domain Science Applications on GENI: Presentation and DemoabstractMulti-tenant cloud infrastructures are increasingly used for high-performance and high-throughput domain science applications. In recent years, machine virtualization has come a long way toward supporting domain science applications. Various cloud platforms, such as Open Stack, Cloud Stack, and Amazon EC2 are attracting scientists to these platforms with the promise of customized environments with virtually infinite compute resources. At the same time, research efforts, such as NSF GENI are bringing together cloud computing with advanced network infrastructure provisioning. This paper presents work toward evaluating the use of GENI to support domain science applications. The evaluation involved two different domain science applications deployed on ExoGENI and Insta GENI. The first application is ADCIRC, a storm surge model that uses Message Passing Interface (MPI). The second is Motif network, a genomics application using the Pegasus workflow management system to manage a large data-intensive workflow. Paul Ruth, Anirban Mandal |
ICNP | 2 |
| 2012 | Dynamic network provisioning for data intensive applications in the cloudabstractAdvanced networks are an essential element of data-driven science enabled by next generation cyberinfrastructure environments. Computational activities increasingly incorporate widely dispersed resources with linkages among software components spanning multiple sites and administrative domains. We have seen recent advances in enabling on-demand network circuits in the national and international backbones coupled with Software Defined Networking (SDN) advances like OpenFlow and programmable edge technologies like OpenStack. These advances have created an unprecedented opportunity to enable complex scientific applications to run on specially tailored, dynamic infrastructure that include compute, storage and network resources, combining the performance advantages of purpose-built infrastructures, but without the costs of a permanent infrastructure. This work presents an experience deploying scientific workflows on the ExoGENI national test bed that dynamically allocates computational resources with high-speed circuits from backbone providers. Dynamically allocated bandwidth-provisioned high-speed circuits increase the ability of scientific applications to access and stage large data sets from remote data repositories or to move computation to remote sites and access data stored locally. The remainder of this extended abstract is a brief description of the test bed and several scientific workflow applications that were deployed using bandwidth-provisioned high-speed circuits. Paul Ruth, Anirban Mandal, Yufeng Xin, Ilya Baldin, Chris Heermann, Jeffrey S. Chase |
eScience | 2 |
| 2011 | Provisioning and Evaluating Multi-domain Networked Clouds for Hadoop-based ApplicationsabstractThis paper presents the design, implementation, and evaluation of a new system for on-demand provisioning of Hadoop clusters across multiple cloud domains. The Hadoop clusters are created "on-demand" and are composed of virtual machines from multiple cloud sites linked with bandwidth-provisioned network pipes. The prototype uses an existing federated cloud control framework called Open Resource Control Architecture (ORCA), which orchestrates the leasing and configuration of virtual infrastructure from multiple autonomous cloud sites and network providers. ORCA enables computational and network resources from multiple clouds and network substrates to be aggregated into a single virtual "slice" of resources, built to order for the needs of the application. The experiments examine various provisioning alternatives by evaluating the performance of representative Hadoop benchmarks and applications on resource topologies with varying bandwidths. The evaluations examine conditions in which multi-cloud Hadoop deployments pose significant advantages or disadvantages during Map/Reduce/Shuffle operations. Further, the experiments compare multi-cloud Hadoop deployments with single-cloud deployments and investigate Hadoop Distributed File System (HDFS) performance under varying network configurations. The results show that networked clouds make cross-cloud Hadoop deployment feasible with high bandwidth network links between clouds. As expected, performance for some benchmarks degrades rapidly with constrained inter-cloud bandwidth. MapReduce shuffle patterns and certain Hadoop Distributed File System (HDFS) operations that span the constrained links are particularly sensitive to network performance. Hadoop's topology-awareness feature can mitigate these penalties to a modest degree in these hybrid bandwidth scenarios. Additional observations show that contention among co-located virtual machines is a source of irregular performance for Hadoop applications on virtual cloud infrastructure. Anirban Mandal, Yufeng Xin, Ilya Baldin, Paul Ruth, Chris Heermann, Jeffrey S. Chase, Victor Orlikowski, Aydan R. Yumerefendi |
CloudCom | 1 |
| 2010 | Modeling memory concurrency for multi-socket multi-core systemsabstractMulti-core computers are ubiquitous and multi-socket versions dominate as nodes in compute clusters. Given the high level of parallelism inherent in processor chips, the ability of memory systems to serve a large number of concurrent memory access operations is becoming a critical performance problem. The most common model of memory performance uses just two numbers, peak bandwidth and typical access latency. We introduce concurrency as an explicit parameter of the measurement and modeling processes to characterize more accurately the complexity of memory behavior of multi-socket, multi-core systems. We present a detailed experimental multi-socket, multi-core memory study based on the PCHASE benchmark, which can vary memory loads by controlling the number of concurrent memory references per thread. The make-up and structure of the memory have a major impact on achievable bandwidth. Three discrete bottlenecks were observed at different levels of the hardware architecture: limits on the number of references outstanding per core; limits to the memory requests serviced by a single memory controller; and limits on the global memory concurrency. We use these results to build a memory performance model that ties concurrency, latency and bandwidth together to create a more accurate model of overall performance. We show that current commodity memory sub-systems cannot handle the load offered by high-end processor chips. Anirban Mandal, Robert J. Fowler, Allan Porterfield |
ISPASS | 1 |
| 2009 | Combined Fault Tolerance and Scheduling Techniques for Workflow Applications on Computational GridsabstractComplex scientific workflows are now Increasingly executed on computational grids. In addition to the challenges of managing and scheduling these workflows, reliability challenges arise because of the unreliable nature of large-scale grid infrastructure. Fault tolerance mechanisms like over-provisioning and checkpoint-recovery are used in current grid application management systems to address these reliability challenges. In this work, we propose new approaches that combine these fault tolerance techniques with existing workflow scheduling algorithms. We present a study on the effectiveness of the combined approaches by analyzing their impact on the reliability of workflow execution, workflow performance and resource usage under different reliability models, failure prediction accuracies and workflow application types. Anirban Mandal, Charles Koelbel, Keith D. Cooper |
CCGRID | 2 |
| 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault toleranceabstractToday's scientific workflows use distributed heterogeneous resources through diverse grid and cloud interfaces that are often hard to program. In addition, especially for time-sensitive critical applications, predictable quality of service is necessary across these distributed resources. VGrADS' virtual grid execution system (vgES) provides an uniform qualitative resource abstraction over grid and cloud systems. We apply vgES for scheduling a set of deadline sensitive weather forecasting workflows. Specifically, this paper reports on our experiences with (1) virtualized reservations for batchqueue systems, (2) coordinated usage of TeraGrid (batch queue), Amazon EC2 (cloud), our own clusters (batch queue) and Eucalyptus (cloud) resources, and (3) fault tolerance through automated task replication. The combined effect of these techniques was to enable a new workflow planning method to balance performance, reliability and cost considerations. The results point toward improved resource selection and execution management support for a variety of e-Science applications over grids and cloud systems. Lavanya Ramakrishnan, Charles Koelbel, Yang-Suk Kee, Richard Wolski, Daniel Nurmi, Dennis Gannon, Graziano Obertelli, Asim YarKhan, Anirban Mandal, T. Mark Huang, Kiran Thyagaraja, Dmitrii Zagorodnov |
SC | 9 |
| 2008 | Fault Tolerance and Recovery of Scientific Workflows on Computational GridsabstractIn this paper, we describe the design and implementation of two mechanisms for fault-tolerance and recovery for complex scientific workflows on computational grids. We present our algorithms for over-provisioning and migration, which are our primary strategies for fault-tolerance. We consider application performance models, resource reliability models, network latency and bandwidth and queue wait times for batch-queues on compute resources for determining the correct fault-tolerance strategy. Our goal is to balance reliability and performance in the presence of soft real-time constraints like deadlines and expected success probabilities, and to do it in a way that is transparent to scientists. We have evaluated our strategies by developing a Fault-Tolerance and Recovery (FTR) service and deploying it as a part of the Linked Environments for Atmospheric Discovery (LEAD) production infrastructure. Results from real usage scenarios in LEAD show that the failure rate of individual steps in workflows decreases from about 30% to 5% by using our fault-tolerance strategies. Gopi Kandaswamy, Anirban Mandal, Daniel A. Reed |
CCGRID | 2 |
| 2006 | Scalable Grid Application Scheduling via Decoupled Resource Selection and SchedulingabstractOver the past years grid infrastructures have been deployed at larger and larger scales, with envisioned deployments incorporating tens of thousands of resources. Therefore, application scheduling algorithms can become unscalable (albeit polynomial) and thus unusable in large-scale environments. One reason for unscalability is that these algorithms perform implicit resource selection. One can achieve better scalability by performing explicit resource selection independently from scheduling in a "decoupled' approach. Furthermore, we hypothesize that one can achieve similar or even better performance as with the non-decoupled approach, which we call the "one step" approach, by selecting resources judiciously. Leveraging the Virtual Grid abstraction, we demonstrate that the decoupled approach is indeed both scalable and effective in large-scale and highly heterogeneous resource environments. Anirban Mandal, Henri Casanova, Andrew A. Chien, Yang-Suk Kee, Ken Kennedy, Charles Koelbel |
CCGRID | 2 |
| 2006 | Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time predictionabstractLarge-scale distributed systems offer computational power at unprecedented levels. In the past, HPC users typically had access to relatively few individual supercomputers and, in general, would assign a one-to-one mapping of applications to machines. Modern HPC users have simultaneous access to a large number of individual machines and are beginning to make use of all of them for single-application execution cycles. One method that application developers have devised in order to take advantage of such systems is to organize an entire application execution cycle as a workflow. The scheduling of such workflows has been the topic of a great deal of research in the past few years and, although very sophisticated algorithms have been devised, a very specific aspect of these distributed systems, namely that most supercomputing resources employ batch queue scheduling software, has heretofore been omitted from consideration, presumably because it is difficult to model accurately. In this work, we augment an existing workflow scheduler through the introduction of methods which make accurate predictions of both the performance of the application on specific hardware, and the amount of time individual workflow tasks will spend waiting in batch queues. Our results show that although a workflow scheduler alone may choose correct task placement based on data locality or network connectivity, this benefit is often compromised by the fact that most jobs submitted to current systems must wait in overcommited batch queues for a significant portion of time. However, incorporating the enhancements we describe improves workflow execution time in settings where batch queues impose significant delays on constituent workflow tasks. Daniel Nurmi, Anirban Mandal, John Brevik, Charles Koelbel, Richard Wolski, Ken Kennedy |
SC | 2 |
| 2005 | Task scheduling strategies for workflow-based applications in gridsabstractGrid applications require allocating a large number of heterogeneous tasks to distributed resources. A good allocation is critical for efficient execution. However, many existing grid toolkits use matchmaking strategies that do not consider overall efficiency for the set of tasks to be run. We identify two families of resource allocation algorithms: task-based algorithms, that greedily allocate tasks to resources, and workflow-based algorithms, that search for an efficient allocation for the entire workflow. We compare the behavior of workflow-based algorithms and task-based algorithms, using simulations of workflows drawn from a real application and with varying ratios of computation cost to data transfer cost. We observe that workflow-based approaches have a potential to work better for data-intensive applications even when estimates about future tasks are inaccurate. Jim Blythe, Ewa Deelman, Yolanda Gil, Karan Vahi, Anirban Mandal, Ken Kennedy |
CCGRID | 6 |
| 2005 | Scheduling strategies for mapping application workflows onto the gridabstractIn this work, we describe new strategies for scheduling and executing workflow applications on grid resources using the GrADS [Ken Kennedy et al., 2002] infrastructure. Workflow scheduling is based on heuristic scheduling strategies that use application component performance models. The workflow is executed using a novel strategy to bind and launch the application onto heterogeneous resources. We apply these strategies in the context of executing EMAN, a bio-imaging workflow application, on the grid. The results of our experiments show that our strategy of performance model based, in-advance heuristic workflow scheduling results in 1.5 to 2.2 times better makespan than other existing scheduling strategies. This strategy also achieves optimal load balance across the different grid sites for this application. Anirban Mandal, Ken Kennedy, Charles Koelbel, Gabriel Marin, John M. Mellor-Crummey, S. Lennart Johnsson |
HPDC | 1 |
| 2004 | Scheduling workflow applications in GrADSabstractIn this work, we describe new strategies for scheduling and executing workflow applications on Grid resources using the GrADS infrastructure. Workflow scheduling is based on heuristic scheduling strategies that use combined computational and memory hierarchy application component performance models. The workflow is executed using a novel strategy to bind and launch the application onto heterogeneous resources. We apply these strategies in the context of launching EMAN, a bio-imaging workflow application, onto the Grid. Anirban Mandal, Anshuman Dasgupta, Ken Kennedy, Mark Mazina, Charles Koelbel, Gabriel Marin, Keith D. Cooper, John M. Mellor-Crummey, S. Lennart Johnsson |
CCGRID | 1 |