William E. Allcock

dblp:a/WilliamEAllcock · also Bill Allcock · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
6since 2021 · last 2022
0000-0002-7984-6847ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 6 first-author · 6 since 2021Computer networks · 1
YearPublicationVenuePosition
2022 MRSch: Multi-Resource Scheduling for HPC
abstract
Emerging workloads in high-performance computing (HPC) are embracing significant changes, such as having diverse resource requirements instead of being CPU -centric. This advancement forces cluster schedulers to consider multiple schedulable resources during decision-making. Existing scheduling studies rely on heuristic or optimization methods, which are limited by an inability to adapt to new scenarios for ensuring long-term scheduling performance. We present an intelligent scheduling agent named MRSch for multi-resource scheduling in HPC that leverages direct future prediction (DFP), an advanced multi-objective reinforcement learning algorithm. While DFP demonstrated outstanding performance in a gaming competition, it has not been previously explored in the context of HPC scheduling. Several key techniques are developed in this study to tackle the challenges involved in multi-resource scheduling. These techniques enable MRSch to learn an appropriate scheduling pol-icy automatically and dynamically adapt its policy in response to workload changes via dynamic resource prioritizing. We compare MRSch with existing scheduling methods through extensive trace-base simulations. Our results demonstrate that MRSch improves scheduling performance by up to 48 % compared to the existing scheduling methods.
Boyang Li 0018, Yuping Fan, Matthew T. Dearing, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka
CLUSTER6
2022 What does Inter-Cluster Job Submission and Execution Behavior Reveal to Us?
abstract
Modern High Performing Computing (HPC) facil-ities have multiple computing clusters that serve different pur-poses. These include large-scale computing clusters and smaller data visualization and analysis clusters, which are meant to shift the load of data analytics jobs from the large-scale systems. We perform the first in-depth characterization of cross-cluster behavior of users and jobs and provide an analysis of three inter-related systems at the Argonne Leadership Computing Facility (ALCF). Our analysis reveals interesting trends related to the resource utilization and predictability of user and job behavior across different clusters.
Tirthak Patel, Devesh Tiwari, Rajkumar Kettimuthu, William E. Allcock, Paul M. Rich, Zhengchun Liu
CLUSTER4
2022 Hybrid Workload Scheduling on HPC Systems
abstract
Traditionally, on-demand, rigid, and malleable applications have been scheduled and executed on separate systems. The ever-growing workload demands and rapidly developing HPC infrastructure trigger the interest of converging these applications on a single HPC system. Although allocating the hybrid workloads within one system could potentially improve system efficiency, it is difficult to balance the tradeoff between the responsiveness of on-demand requests, incentive for malleable jobs, and the performance of rigid applications. In this study, we present several scheduling mechanisms to address the issues involved in co-scheduling on-demand, rigid, and malleable jobs on a single HPC system. We extensively evaluate and compare their performance under various configurations and workloads. Our experimental results show that our proposed mechanisms are capable of serving on-demand workloads with minimal delay, offering incentives for declaring malleability, and improving system performance.
Yuping Fan, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka
IPDPS4
2022 DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing
abstract
Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. The experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.
Yuping Fan, Boyang Li 0018, Dustin Favorite, Naunidh Singh, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan
IEEE Trans. Parallel Distributed Syst.7
2021 Operating Liquid-Cooled Large-Scale Systems: Long-Term Monitoring, Reliability Analysis, and Efficiency Measures
abstract
The past decade has seen a rise in the use of liquid cooling due to its energy efficiency. While many previous works have helped make progress toward improving data center cooling, a vast majority of them perform studies on a small system over a short span. The computer systems and HPC community lacks a long-term study highlighting the challenges and solutions in operating a liquid-cooled large-scale data center. We conduct the first detailed characterization of a petascale supercomputer, Mira, over a span of six years. The study is enabled by systematic monitoring of the environmental metrics, and discusses new research avenues, including coolant monitor failures.
Rohan Basu Roy, Tirthak Patel, Rajkumar Kettimuthu, William E. Allcock, Paul M. Rich, Adam Scovel, Devesh Tiwari
HPCA4
2021 Deep Reinforcement Agent for Scheduling in HPC
abstract
Cluster scheduler is crucial in high-performance computing (HPC). It determines when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a novel, hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. A unique training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. The experiments with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 45%.
Yuping Fan, Zhiling Lan, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka
IPDPS5
2020 Job characteristics on large-scale systems: long-term analysis, quantification, and implications
abstract
HPC workload analysis and resource consumption characteristics are the key to driving better operation practices, system procurement decisions, and designing effective resource management techniques. Unfortunately, the HPC community does not have easy accessibility to long-term introspective work-load analysis and characterization for production-scale HPC systems. This study bridges this gap by providing detailed long-term quantification, characterization, and analysis of job characteristics on two supercomputers: Intrepid and Mira. This study is one of the largest of its kind - covering trends and characteristics for over three billion compute hours, 750 thousand jobs, and spanning a decade. We confirm several long-held conventional wisdom, and identify many previously undiscovered trends and its implications. We also introduce a learning based technique to predict the resource requirement of future jobs with high accuracy, using features available prior to the job submission and without requiring any application-specific tracing or application-intrusive instrumentation.
Tirthak Patel, Zhengchun Liu, Rajkumar Kettimuthu, Paul M. Rich, William E. Allcock, Devesh Tiwari
SC5
2019 Scheduling Beyond CPUs for HPC
abstract
High performance computing (HPC) is undergoing significant changes. The emerging HPC applications comprise both compute- and data-intensive applications. To meet the intense I/O demand from emerging data-intensive applications, burst buffers are deployed in production systems. Existing HPC schedulers are mainly CPU-centric. The extreme heterogeneity of hardware devices, combined with workload changes, forces the schedulers to consider multiple resources (e.g., burst buffers) beyond CPUs, in decision making. In this study, we present a multi-resource scheduling scheme named BBSched that schedules user jobs based on not only their CPU requirements, but also other schedulable resources such as burst buffer. BBSched formulates the scheduling problem into a multi-objective optimization (MOO) problem and rapidly solves the problem using a multi-objective genetic algorithm. The multiple solutions generated by BBSched enables system managers to explore potential tradeoffs among various resources, and therefore obtains better utilization of all the resources. The trace-driven simulations with real system workloads demonstrate that BBSched improves scheduling performance by up to 41% compared to existing methods, indicating that explicitly optimizing multiple resources beyond CPUs is essential for HPC scheduling.
Yuping Fan, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka, Brian Austin, David Paul
HPDC4
2018 A Job Sizing Strategy for High-Throughput Scientific Workflows
abstract
The user of a computing facility must make a critical decision when submitting jobs for execution: how many resources (such as cores, memory, and disk) should be requested for each job? If the request is too small, the job may fail due to resource exhaustion; if the request is too large, the job may succeed, but resources will be wasted. This decision is especially important when running hundreds of thousands of jobs in a high throughput workflow, which may exhibit complex, long tailed distributions of resource consumption. In this paper, we present a strategy for solving the job sizing problem: (1) applications are monitored and measured in user-space as they run; (2) the resource usage is collected into an online archive; and (3) jobs are automatically sized according to historical data in order to maximize throughput or minimize waste. We evaluate the solution analytically, and present case studies of applying the technique to high throughput physics and bioinformatics workflows consisting of hundreds of thousands of jobs, demonstrating an increase in throughput of 10-400 percent compared to naive approaches.
Benjamín Tovar, Rafael Ferreira da Silva, Gideon Juve, Ewa Deelman, William E. Allcock, Douglas Thain, Miron Livny
IEEE Trans. Parallel Distributed Syst.5
2017 Trade-Off Between Prediction Accuracy and Underestimation Rate in Job Runtime Estimates
abstract
Job runtime estimates provided by users are widely acknowledged to be overestimated and runtime overestimation can greatly degrade job scheduling performance. Previous studies focus on improving accuracy of job runtime estimates by reducing runtime overestimation, but fail to address the underestimation problem (i.e., the underestimation of job runtimes). Using an underestimated runtime is catastrophic to a job as the job will be killed by the scheduler before completion. We argue that both the improvement of runtime accuracy and the reduction of underestimation rate are equally important. To address this problem, we propose an online runtime adjustment framework called TRIP. TRIP explores the data censoring capability of the Tobit model to improve prediction accuracy while keeping a low underestimation rate of job runtimes. TRIP can be used as a plugin to job scheduler for improving job runtime estimates and hence boosting job scheduling performance. Preliminary results demonstrate that TRIP is capable of achieving high accuracy of 80% and low underestimation rate of 5%. This is significant as compared to other well-known machine learning methods such as SVM, Random Forest, and Last-2 which result in a high underestimation rate (20%-50%). Our experiments further quantify the amount of scheduling performance gain achieved by the use of TRIP.
Yuping Fan, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan
CLUSTER3
2017 Experience and Practice of Batch Scheduling on Leadership Supercomputers at Argonne
William E. Allcock, Paul M. Rich, Yuping Fan, Zhiling Lan
JSSPP1
2016 Consecutive Job Submission Behavior at Mira Supercomputer
abstract
Understanding user behavior is crucial for the evaluation of scheduling and allocation performances in HPC environments. This paper aims to further understand the dynamic user reaction to different levels of system performance by performing a comprehensive analysis of user behavior in recorded data in the form of delays in the subsequent job submission behavior. Therefore, we characterize a workload trace covering one year of job submissions from the Mira supercomputer at ALCF (Argonne Leadership Computing Facility). We perform an in-depth analysis of correlations between job characteristics, system performance metrics, and the subsequent user behavior. Analysis results show that the user behavior is significantly influenced by long waiting times, and that complex jobs (number of nodes and CPU hours) lead to longer delays in subsequent job submissions.
Stephan Schlagkamp, Rafael Ferreira da Silva, William E. Allcock, Ewa Deelman, Uwe Schwiegelshohn
HPDC3
2016 A data driven scheduling approach for power management on HPC systems
abstract
Modern schedulers running on HPC systems traditionally consider the number of resources and the time requested for each job that is to be executed when making scheduling decisions. Until recently this has been sufficient, however as systems get larger, other metrics like power consumption become necessary to ensure system stability. In this paper, we propose a data driven scheduling approach for controlling the power consumption of the entire system under any user defined budget. Here, “data driven” means that our approach actively observes, analyzes, and assesses power behaviors of the system and user jobs to guide scheduling decisions for power management. This design is based on the key observation that HPC jobs have distinct power profiles. Our work contains an empirical analysis of workload power characteristics on a production system, dynamic learner to estimate the job power profile for scheduling, and an online power-aware scheduler for managing the overall system power. Using real workload traces, we demonstrate that our design effectively controls system power consumption while minimizing the impact on system utilization.
Sean Wallace, Xu Yang 0009, Venkatram Vishwanath, William E. Allcock, Susan Coghlan, Michael E. Papka, Zhiling Lan
SC4
2015 Practical Resource Monitoring for Robust High Throughput Computing
abstract
Robust high throughput computing requires effective monitoring and enforcement of a variety of resources including CPU cores, memory, disk, and network traffic. Without effective monitoring and enforcement, it is easy to overload machines, causing failures and slowdowns, or underutilize machines, which results in wasted opportunities. This paper explores how to describe, measure, and enforce resources used by computational tasks. We focus on tasks running in distributed execution systems, in which a task requests the resources it needs, and the execution system ensures the availability of such resources. This presents two non-trivial problems: how to measure the resources consumed by a task, and how to monitor and report resource exhaustion in a robust and timely manner. For both of these tasks, operating systems have a variety of mechanisms with different degrees of availability, accuracy, overhead, and intrusiveness. We describe various forms of monitoring and the available mechanisms in contemporary operating systems. We then present two specific monitoring tools that choose different tradeoffs in overhead and accuracy, and evaluate them on a selection of benchmarks.
Gideon Juve, Benjamín Tovar, Rafael Ferreira da Silva, Dariusz Król 0002, Douglas Thain, Ewa Deelman, William E. Allcock, Miron Livny
CLUSTER7
2011 Understanding and improving computational science storage access through continuous characterization
abstract
Computational science applications are driving a demand for increasingly powerful storage systems. While many techniques are available for capturing the I/O behavior of individual application trial runs and specific components of the storage system, continuous characterization of a production system remains a daunting challenge for systems with hundreds of thousands of compute cores and multiple petabytes of storage. As a result, these storage systems are often designed without a clear understanding of the diverse computational science workloads they will support.In this study, we outline a methodology for scalable, con tinuous, systemwide I/O characterization that combines storage device instrumentation, static file system analysis, and a new mechanism for capturing detailed application-level behavior. This methodology allows us to quantify systemwide trends such as the way application behavior changes with job size, the "burstiness" of the storage system, and the evolution of file system contents over time. The data also can be examined by application domain to determine the most prolific storage users and also investigate how their I/O strategies correlate with I/O performance. At the most detailed level, our characterization methodology can also be used to focus on individual applications and guide tuning efforts for those applications. We demonstrate the effectiveness of our methodology by performing a multilevel, two-month study of Intrepid, a 557 teraflop IBM Blue Gene/P system. During that time, we captured application-level I/O characterizations from 6,481 unique jobs spanning 38 science and engineering projects with up to 163,840 processes per job. We also captured patterns of I/O activity in over 8 petabytes of block device traffic and summarized the con tents of file systems containing over 191 million flies. We then used the results of our study to tune example applications, highlight trends that impact the design of future storage systems, and identify opportunities for improvement in I/O characterization methodology.
Philip H. Carns, Kevin Harms, William E. Allcock, Charles Bacon, Samuel Lang, Robert Latham, Robert B. Ross
MSST3
2011 Understanding and Improving Computational Science Storage Access through Continuous Characterization
abstract
Computational science applications are driving a demand for increasingly powerful storage systems. While many techniques are available for capturing the I/O behavior of individual application trial runs and specific components of the storage system, continuous characterization of a production system remains a daunting challenge for systems with hundreds of thousands of compute cores and multiple petabytes of storage. As a result, these storage systems are often designed without a clear understanding of the diverse computational science workloads they will support. In this study, we outline a methodology for scalable, continuous, systemwide I/O characterization that combines storage device instrumentation, static file system analysis, and a new mechanism for capturing detailed application-level behavior. This methodology allows us to identify both system-wide trends and application-specific I/O strategies. We demonstrate the effectiveness of our methodology by performing a multilevel, two-month study of Intrepid, a 557-teraflop IBM Blue Gene/P system. During that time, we captured application-level I/O characterizations from 6,481 unique jobs spanning 38 science and engineering projects. We used the results of our study to tune example applications, highlight trends that impact the design of future storage systems, and identify opportunities for improvement in I/O characterization methodology.
Philip H. Carns, Kevin Harms, William E. Allcock, Charles Bacon, Samuel Lang, Robert Latham, Robert B. Ross
ACM Trans. Storage3
2010 Lessons learned from moving earth system grid data sets over a 20 Gbps wide-area network
abstract
In preparation for the Intergovernmental Panel on Climate Change (IPCC) Fifth Assessment Report, the climate community will run the Coupled Model Intercomparison Project phase 5 (CMIP-5) experiments, which are designed to answer crucial questions about future regional climate change and the results of carbon feedback for different mitigation scenarios. The CMIP-5 experiments will generate petabytes of data that must be replicated seamlessly, reliably, and quickly to hundreds of research teams around the globe. As an end-to-end test of the technologies that will be used to perform this task, a multi-disciplinary team of researchers moved a small portion (10 TB) of the multimodel Coupled Model Intercomparison Project, Phase 3 data set used in the IPCC Fourth Assessment Report from three sources---the Argonne Leadership Computing Facility (ALCF), Lawrence Livermore National Laboratory (LLNL) and National Energy Research Scientific Computing Center (NERSC)---to the 2009 Supercomputing conference (SC09) show floor in Portland, Oregon, over circuits provided by DOE's ESnet. The team achieved a sustained data rate of 15 Gb/s on a 20 Gb/s network. More important, this effort provided critical feedback on how to deploy, tune, and monitor the middleware that will be used to replicate the upcoming petascale climate datasets. We report on obstacles overcome and the key lessons learned from this successful bandwidth challenge effort.
Rajkumar Kettimuthu, Alex Sim, Dan Gunter, William E. Allcock, Peer-Timo Bremer, John Bresnahan, Andrew Cherry, Lisa Childers, Eli Dart, Ian T. Foster, Kevin Harms, Jason Hick, Jason Lee 0001, Michael Link, Jeff Long, Keith Miller 0005, Vijaya Natarajan, Valerio Pascucci, Kenneth Raffenetti, David Ressman, Dean N. Williams, Loren Wilson, Linda Winkler
HPDC4
2009 I/O performance challenges at leadership scale
abstract
Today's top high performance computing systems run applications with hundreds of thousands of processes, contain hundreds of storage nodes, and must meet massive I/O requirements for capacity and performance. These leadership-class systems face daunting challenges to deploying scalable I/O systems. In this paper we present a case study of the I/O challenges to performance and scalability on Intrepid, the IBM Blue Gene/P system at the Argonne Leadership Computing Facility. Listed in the top 5 fastest supercomputers of 2008, Intrepid runs computational science applications with intensive demands on the I/O system. We show that Intrepid's file and storage system sustain high performance under varying workloads as the applications scale with the number of processes.
Samuel Lang, Philip H. Carns, Robert Latham, Robert B. Ross, Kevin Harms, William E. Allcock
SC6
2007 GridCopy: Moving Data Fast on the Grid
abstract
An important type of communication in grid and distributed computing environments is bulk data transfer. GridFTP has emerged as a de facto standard for secure, reliable, high-performance data transfer across resources on the Grid. GridCopy provides a simple GridFTP client interface to users and extensible configuration that can be changed dynamically by administrators to make efficient data movement in the Grid easier for users.
Rajkumar Kettimuthu, William E. Allcock, Liming Lee, John-Paul Navarro, Ian T. Foster
IPDPS2
2007 A portal for visualizing Grid usage
abstract
Abstract We introduce a framework for measuring the use of Grid services and exposing simple summary data to an authorized set of Grid users through a JSR168‐enabled portal. The sensor framework has been integrated into the Globus Toolkit and allows Grid administrators to have access to a mechanism helping with report and usage statistics. Although the original focus was the reporting of actions in relationship to GridFTP services, the usage service has been expanded to report also on the use of other Grid services. Copyright © 2007 John Wiley & Sons, Ltd.
Gregor von Laszewski, Jonathan DiCarlo, William E. Allcock
Concurr. Comput. Pract. Exp.3
2005 Restricted Slow-Start for TCP
abstract
In network protocol research a common goal is optimal bandwidth utilization, while still being network friendly. The drawback of TCP in networks with large bandwidth-delay products due to its AIMD based congestion control mechanism is well known. The congestion control algorithm of TCP has two phases namely slow-start phase and congestion-avoidance phase. Many researchers have focused on modifying the congestion avoidance phase of the algorithm. In this work, we propose a modification to the slow-start phase of the algorithm to achieve better performance. Restricted slow-start algorithm is a simple sender side alteration to the TCP congestion window update algorithm
William E. Allcock, S. Hegde, Rajkumar Kettimuthu
CLUSTER1
2005 CODO: firewall traversal by cooperative on-demand opening
abstract
Firewalls and network address translators (NATs) cause significant connectivity problems along with benefits such as network protection and easy address planning. Connectivity problems make nodes separated by a firewall/NAT unable to communicate with each other. Due to the bidirectional and multi-organizational nature of grids, they are particularly susceptible to connectivity problems. These problems make collaboration difficult or impossible and cause resources to be wasted. This paper presents a system, called CODO, which provides applications end-to-end connectivity over firewalls/NATs in a secure way. CODO allows applications authorized through strong security mechanisms to traverse firewalls/NATs, while blocking unauthorized applications. This paper also formalizes the firewall/NAT traversal problem and clarifies how a traversal system fits in the overall security policy enforcement by a firewall/NAT.
Se-Chang Son, William E. Allcock, Miron Livny
HPDC2
2005 The Globus Striped GridFTP Framework and Server
abstract
The GridFTP extensions to the File Transfer Protocol define a general-purpose mechanism for secure, reliable, high-performance data movement. We report here on the Globus striped GridFTP framework, a set of client and server libraries designed to support the construction of data-intensive tools and applications. We describe the design of both this framework and a striped GridFTP server constructed within the framework. We show that this server is faster than other FTP servers in both single-process and striped configurations, achieving, for example, speeds of 27.3 Gbit/s memory-to-memory and 17 Gbit/s disk-to-disk over a 60 millisecond round trip time, 30 Gbit/s network. In another experiment, we show that the server can support 1800 concurrent clients without excessive load. We argue that this combination of performance and modular structure make the Globus GridFTP framework both a good foundation on which to build tools and applications, and a unique testbed for the study of innovative data management techniques and network protocols.
William E. Allcock, John Bresnahan, Rajkumar Kettimuthu, Michael Link
SC1
2003 Grid-enabled particle physics event analysis: experiences using a 10 Gb, high-latency network for a high-energy physics application
William E. Allcock, John Bresnahan, Julian J. Bunn, S. Hegde, Joseph A. Insley, Rajkumar Kettimuthu, Harvey B. Newman, Sylvain Ravot, Tony Rimovsky, Conrad Steenberg, Linda Winkler
Future Gener. Comput. Syst.1
2003 High-performance remote access to climate simulation data: a challenge problem for data grid technologies
Ann L. Chervenak, Ewa Deelman, Carl Kesselman, William E. Allcock, Ian T. Foster, Veronika Nefedova, Jason Lee 0001, Alex Sim, Arie Shoshani, Bob Drach, Dean N. Williams, Don Middleton
Parallel Comput.4
2002 GridMapper: A Tool for Visualizing the Behavior of Large-Scale Distributed Systems
abstract
Grid applications can combine the use of computation, storage, network, and other resources. These resources are often geographically distributed, adding to application complexity and thus the difficulty of understanding application performance. We present GridMapper, a tool for monitoring and visualizing the behavior of such distributed systems. GridMapper builds on basic mechanisms for registering, discovering, and accessing performance information sources, as well as for mapping from domain names to physical locations. The visualization system itself then supports the automatic layout of distributed sets of such sources and animation of their activities. We use a set of examples to illustrate how the system can provide valuable insights into the behavior and performance of a range of different applications.
William E. Allcock, Joseph Bester, John Bresnahan, Ian T. Foster, Jarek Gawor, Joseph A. Insley, Joseph M. Link, Michael E. Papka
HPDC1
2002 Reliable File Transfer in Grid Environments
abstract
Grid-based computing environments are becoming increasingly popular for scientific computing. One of the key issues for scientific computing is the efficient transfer of large amounts of data across the Grid. In this paper we present a reliable file transfer (RFT) service that significantly improves the efficiency of large-scale file transfer. RFT can detect a variety of failures and restart the file transfer from the point of failure. It also has capabilities for improving transfer performance through TCP tuning.
Ravi K. Madduri, William E. Allcock
LCN3
2002 Data management and transfer in high-performance computational grid environments
William E. Allcock, Joseph Bester, John Bresnahan, Ann L. Chervenak, Ian T. Foster, Carl Kesselman, Sam Meder, Veronika Nefedova, Darcy Quesnel, Steven Tuecke
Parallel Comput.1
2001 File and Object Replication in Data Grids
abstract
Data replication is a key issue in a data grid and can be managed in different ways and at different levels of granularity: for example, at the file level or the object level. In the high-energy physics community, data grids are being developed to support the distributed analysis of experimental data. We have produced a prototype data replication tool, the Grid Data Management Pilot (GDMP) that is in production use in one physics experiment, with middleware provided by the Globus toolkit used for authentication, data movement and other purposes. We present a new, enhanced GDMP architecture and prototype implementation that uses Globus data-grid tools for efficient file replication. We also explain how this architecture can address object replication issues in an object-oriented database management system. File transfer over wide-area networks requires specific performance tuning in order to gain optimal data transfer rates. We present performance results obtained with GridFTP, an enhanced version of FTP, and discuss tuning parameters.
Heinz Stockinger, Asad Samar, Koen Holtman, William E. Allcock, Ian T. Foster, Brian Tierney
HPDC4
2001 High-performance remote access to climate simulation data: a challenge problem for data grid technologies
abstract
In numerous scientific disciplines, terabyte and soon petabyte-scale data collections are emerging as critical community resources. A new class of Data Grid infrastructure is required to support management, transport, distributed access to, and analysis of these datasets by potentially thousands of users. Researchers who face this challenge include the Climate Modeling community, which performs long-duration computations accompanied by frequent output of very large files that must be further analyzed. We describe the Earth System Grid prototype, which brings together advanced analysis, replica management, data transfer, request management, and other technologies to support high-performance, interactive analysis of replicated data. We present performance results that demonstrate our ability to manage the location and movement of large datasets from the user's desktop. We report on experiments conducted over SciNET at SC'2000, where we achieved peak performance of 1.55Gb/s and sustained performance of 512.9Mb/s for data transfers between Texas and California.
William E. Allcock, Ian T. Foster, Veronika Nefedova, Ann L. Chervenak, Ewa Deelman, Carl Kesselman, Jason Lee 0001, Alex Sim, Arie Shoshani, Bob Drach, Dean N. Williams
SC1