Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Brian Tierney

dblp:60/6795 · also Brian L. Tierney · DBLP profile ↗
← Back
35ranked-venue papers
10as first author
0since 2021 · last 2018
0000-0003-1607-2909ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 8 first-authorComputer networks · 5Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
18 papers
Distributed systems · 41% High-performance computing · 39% Storage systems · 10%
Computer networks
9 papers
Transport protocols and congestion control · 34% Internet architecture and protocols · 33% Network management and operations · 14%
Network and information security
1 paper
Network security · 77% Systems and software security · 23%

Topics — the 30 heaviest of 50, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
data-intensive computing
0.222012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
A Network-Aware Distributed Storage Cache for Data Intensive Environments · HPDC 1999
High-performance computing › data transfer
wide-area data transfer
0.222012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
A Network-Aware Distributed Storage Cache for Data Intensive Environments · HPDC 1999
Transport protocols and congestion control
transport protocols
0.112012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
Distributed systems
grid computing
0.142004
The Grid2003 Production Grid: Principles and Practice · HPDC 2004
Monitoring data archives for grid environments · SC 2002
File and Object Replication in Data Grids · HPDC 2001
Transport protocols and congestion control › TCP performance enhancement
TCP tuning
0.122002
A TCP tuning daemon · SC 2002
Enabling Network-Aware Applications · HPDC 2001
Distributed systems › grid computing
data grid
0.122002
Giggle: a framework for constructing scalable replica location services · SC 2002
File and Object Replication in Data Grids · HPDC 2001
Network security › intrusion detection and prevention
intrusion detection
0.112006
S06 - Computing protection in open HPC environments · SC 2006
Network measurement and analytics
traffic characterization
0.112005
A First Look at Modern Enterprise Traffic · Internet Measurement Conference 2005
Distributed systems
workflow management
0.112005
Techniques for tuning workflows in cluster environments · HPDC 2005
High-performance computing › scientific workflow
workflow optimization
0.112005
Techniques for tuning workflows in cluster environments · HPDC 2005
Distributed systems › middleware
communication middleware
0.012012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
Distributed systems › replication
data replication
0.022002
File and Object Replication in Data Grids · HPDC 2001
Giggle: a framework for constructing scalable replica location services · SC 2002
Distributed systems
distributed caching
0.021999
A Network-Aware Distributed Storage Cache for Data Intensive Environments · HPDC 1999
High-Speed Distributed Data Handling for On-Line Instrumentation Systems · SC 1997
Cloud and datacenter computing
cluster resource management and scheduling
0.012002
Giggle: a framework for constructing scalable replica location services · SC 2002
Distributed systems › observability › distributed monitoring
distributed application monitoring
0.012002
Dynamic Monitoring of High-Performance Distributed Applications · HPDC 2002
Distributed systems
distributed coordination and fault tolerance
0.012002
Giggle: a framework for constructing scalable replica location services · SC 2002
Distributed systems
fault tolerance
0.012002
Monitoring data archives for grid environments · SC 2002
Distributed systems › replication
replica location service
0.012002
Giggle: a framework for constructing scalable replica location services · SC 2002
Performance modeling and evaluation
performance diagnosis
0.022002
The NetLogger Methodology for High Performance Distributed Systems Performance Analysis · HPDC 1998
Dynamic Monitoring of High-Performance Distributed Applications · HPDC 2002
Internet architecture and protocols › network adaptation
network-aware applications
0.012001
Enabling Network-Aware Applications · HPDC 2001
Storage systems › file systems
distributed file system
0.012001
File and Object Replication in Data Grids · HPDC 2001
High-performance computing › scientific visualization
remote visualization
0.012000
Using High-Speed WANs and Network Data Caches to Enable Remote and Distributed Visualization · SC 2000
High-performance computing
scientific visualization
0.012000
Using High-Speed WANs and Network Data Caches to Enable Remote and Distributed Visualization · SC 2000
Storage systems
distributed storage
0.021994
Using high speed networks to enable distributed parallel image server systems · SC 1994
Distributed Parallel Data Storage Systems: A Scalable Approach to High Speed Image Servers · ACM Multimedia 1994
Storage systems › distributed storage
parallel storage system
0.021994
Using high speed networks to enable distributed parallel image server systems · SC 1994
Distributed Parallel Data Storage Systems: A Scalable Approach to High Speed Image Servers · ACM Multimedia 1994
Performance modeling and evaluation
distributed system performance
0.011998
The NetLogger Methodology for High Performance Distributed Systems Performance Analysis · HPDC 1998
Distributed systems
distributed data processing
0.011997
High-Speed Distributed Data Handling for On-Line Instrumentation Systems · SC 1997
Storage systems › storage hierarchy
tertiary storage
0.011997
High-Speed Distributed Data Handling for On-Line Instrumentation Systems · SC 1997
Network measurement and analytics › traffic measurement
traffic monitoring
0.012005
A First Look at Modern Enterprise Traffic · Internet Measurement Conference 2005
High-performance computing
cluster computing
0.012005
Techniques for tuning workflows in cluster environments · HPDC 2005

Methods — techniques the papers use, named apart from their topics

task synchronization · 0.3flow control · 0.3connection management · 0.3kernel-level tuning · 0.1TCP instrumentation · 0.1parallel streams · 0.1TCP buffer tuning · 0.1netlogger · 0.1system resource analysis · 0.0algorithm analysis · 0.0relational data archive · 0.0parameterized architectural framework · 0.0data archiving · 0.0
YearPublicationVenuePosition
2018 Calibers: A bandwidth calendaring paradigm for science workflows
Fatma Alali, Nathan Hanford, Eric Pouyoul, Rajkumar Kettimuthu, Mariam Kiran, Ben Mack-Crane, Brian Tierney, Yatish Kumar, Dipak Ghosal
Future Gener. Comput. Syst.7
2018 Editorial INDIS special section FGCS
abstract
Nowadays, the cyber, social and physical worlds are increasingly integrating and merging. Especially, combining the strengths of humans and machines helps tackle increasing hard tasks that neither can be done alone. Following this trend, this paper designs a Quality aware Truthful Incentive mechanism for cyber–physical enabled Geographic crowdsensing called Geo-QTI. Different from existing work, Geo-QTI appropriately accommodates the utilities of various stakeholders: requesters, participants and the crowdsourcing platform, and explicitly takes the requesters’ quality requirements, and participants’ quality provision into account. Geo-QTI explicitly includes four components: requester selection, participant selection, pricing and allocation. Requester selection with feasible analysis removes the requesters whose job cannot be completed by all participants or suffers from the monopoly participant (without the participant’s contribution, others cannot cover requesters’ requirement), obtains winning requesters set and determines actual payments. In participant selection phase, the platform aggregates the requested tasks (submitted by all winning requesters) in the sensed geographic area, and chooses the appropriate participants satisfying the winning requesters’ quality requirements with total cost as low as possible. Pricing phase determines the payments to winning participants. The phase of allocation assigns the specific participants to minimally cover the quality requirements of those winning requesters. Rigid theoretical analysis demonstrates Geo-QTI can achieve both requesters’ and participants’ individual rationality and truthfulness, computational efficiency and budget balance for the platform. Furthermore, the extensive simulations confirm our theoretical analysis, and illustrate that Geo-QTI can reduce requesters’ expenses greatly and ensure the fairness of allocation.
Paola Grosso, Malathi Veeraraghavan, Brian Tierney, Cees T. A. M. de Laat
Future Gener. Comput. Syst.3
2018 Enabling intent to configure scientific networks for high performance demands
Mariam Kiran, Eric Pouyoul, Anu Mercian, Brian Tierney, Chin Guok, Inder Monga
Future Gener. Comput. Syst.4
2018 The medical science DMZ: a network design pattern for data-intensive medical science
abstract
OBJECTIVE: We describe a detailed solution for maintaining high-capacity, data-intensive network flows (eg, 10, 40, 100 Gbps+) in a scientific, medical context while still adhering to security and privacy laws and regulations. MATERIALS AND METHODS: High-end networking, packet-filter firewalls, network intrusion-detection systems. RESULTS: We describe a "Medical Science DMZ" concept as an option for secure, high-volume transport of large, sensitive datasets between research institutions over national research networks, and give 3 detailed descriptions of implemented Medical Science DMZs. DISCUSSION: The exponentially increasing amounts of "omics" data, high-quality imaging, and other rapidly growing clinical datasets have resulted in the rise of biomedical research "Big Data." The storage, analysis, and network resources required to process these data and integrate them into patient diagnoses and treatments have grown to scales that strain the capabilities of academic health centers. Some data are not generated locally and cannot be sustained locally, and shared data repositories such as those provided by the National Library of Medicine, the National Cancer Institute, and international partners such as the European Bioinformatics Institute are rapidly growing. The ability to store and compute using these data must therefore be addressed by a combination of local, national, and industry resources that exchange large datasets. Maintaining data-intensive flows that comply with the Health Insurance Portability and Accountability Act (HIPAA) and other regulations presents a new challenge for biomedical research. We describe a strategy that marries performance and security by borrowing from and redefining the concept of a Science DMZ, a framework that is used in physical sciences and engineering research to manage high-capacity data flows. CONCLUSION: By implementing a Medical Science DMZ architecture, biomedical researchers can leverage the scale provided by high-performance computer and cloud storage facilities and national high-speed research networks while preserving privacy and meeting regulatory requirements.
Sean Peisert, Eli Dart, William K. Barnett, Edward Balas, James A. Cuff, Robert L. Grossman, Ari Berman, Anurag Shankar, Brian Tierney
J. Am. Medical Informatics Assoc.9
2017 Enhancing ESnet's OSCARS Path Computation Engine
abstract
The Department of Energy (DOE) supports scientific collaboration in six program areas, ranging from nuclear physics, to biological and environmental research. To enable these data intensive large scale collaborations, the Energy Sciences Network (ESnet) is used to move upwards of 50 Petabytes a month between sites within the US and Europe. Bandwidth within ESnet can be requested, reserved, and utilized through the On-demand Secure Circuits and Advance Reservation System (OSCARS), which provisions network resources with guaranteed bandwidth over a known reservation schedule. Traditionally, OSCARS has only supported simple point-to-point connections between endpoints (e.g. universities, research laboratories), which has limited how efficiently the network may be harnessed by users. This paper details recent extensive enhancements prototyped for OSCARS, to be incorporated into a future release, which enable users to select novel service types including survivability, asymmetric bandwidth or routes, and anycast/manycast. We quantitatively compare several of these enhancements to the baseline point-to-point service, and find that the new services provide not only greater flexibility for the end-user, but savings for network administrators in terms of blocking and network resource usage as well.
Dylan A. P. Davis, Jeremy Plante, Evangelos Chaniotakis, Chin Guok, Vishal Sundarrajan, Brian Tierney, Inder Monga, Vinod Vokkarane
GLOBECOM6
2016 Improving network performance on multicore systems: Impact of core affinities on high throughput flows
Nathan Hanford, Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal, Mehmet Balman, Eric Pouyoul, Brian Tierney
Future Gener. Comput. Syst.7
2016 Special section on high-performance networking for distributed data-intensive science
Brian Tierney, Mehmet Balman, Cees T. A. M. de Laat
Future Gener. Comput. Syst.1
2016 The Medical Science DMZ
abstract
OBJECTIVE: We describe use cases and an institutional reference architecture for maintaining high-capacity, data-intensive network flows (e.g., 10, 40, 100 Gbps+) in a scientific, medical context while still adhering to security and privacy laws and regulations. MATERIALS AND METHODS: High-end networking, packet filter firewalls, network intrusion detection systems. RESULTS: We describe a "Medical Science DMZ" concept as an option for secure, high-volume transport of large, sensitive data sets between research institutions over national research networks. DISCUSSION: The exponentially increasing amounts of "omics" data, the rapid increase of high-quality imaging, and other rapidly growing clinical data sets have resulted in the rise of biomedical research "big data." The storage, analysis, and network resources required to process these data and integrate them into patient diagnoses and treatments have grown to scales that strain the capabilities of academic health centers. Some data are not generated locally and cannot be sustained locally, and shared data repositories such as those provided by the National Library of Medicine, the National Cancer Institute, and international partners such as the European Bioinformatics Institute are rapidly growing. The ability to store and compute using these data must therefore be addressed by a combination of local, national, and industry resources that exchange large data sets. Maintaining data-intensive flows that comply with HIPAA and other regulations presents a new challenge for biomedical research. Recognizing this, we describe a strategy that marries performance and security by borrowing from and redefining the concept of a "Science DMZ"-a framework that is used in physical sciences and engineering research to manage high-capacity data flows. CONCLUSION: By implementing a Medical Science DMZ architecture, biomedical researchers can leverage the scale provided by high-performance computer and cloud storage facilities and national high-speed research networks while preserving privacy and meeting regulatory requirements.
Sean Peisert, William K. Barnett, Eli Dart, James A. Cuff, Robert L. Grossman, Edward B. Talbot, Ari Berman, Anurag Shankar, Brian Tierney
J. Am. Medical Informatics Assoc.9
2014 Impact of the end-system and affinities on the throughput of high-speed flows
abstract
Network throughput is scaling "up" to higher data transfer rates while processors are scaling "out" to multiple cores. As a result, network adapter "offloads" and performance "tuning" have received a good deal of attention lately. However, much of this attention is focused on the "how" and not the "why" of performance efficiency. There are two types of efficiencies that we have found particularly intriguing: First, processor core "affinity," or "binding" is fundamentally the choice of which processor core or cores handle certain tasks in a network- or I/O-heavy application running on a MIMD machine. Second, Ethernet "pause frames" slightly violate the "end-to-end" nature of TCP/IP in order to perform link-to-link flow control. The goal of our research is to delve deeper into why these tuning suggestions and this offload exist, and how they affect the end-to-end performance and efficiency of a single, large TCP flow.
Nathan Hanford, Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal, Mehmet Balman, Eric Pouyoul, Brian Tierney
ANCS7
2013 The Science DMZ: a network design pattern for data-intensive science
abstract
The ever-increasing scale of scientific data has become a significant challenge for researchers that rely on networks to interact with remote computing systems and transfer results to collaborators worldwide. Despite the availability of high-capacity connections, scientists struggle with inadequate cyberinfrastructure that cripples data transfer performance, and impedes scientific progress. The Science DMZ paradigm comprises a proven set of network design patterns that collectively address these problems for scientists. We explain the Science DMZ model, including network architecture, system configuration, cybersecurity, and performance tools, that creates an optimized network environment for science. We describe use cases from universities, supercomputing centers and research laboratories, highlighting the effectiveness of the Science DMZ model in diverse operational settings. In all, the Science DMZ model is a solid platform that supports any science workflow, and flexibly accommodates emerging network technologies. As a result, the Science DMZ vastly improves collaboration, accelerating scientific discovery.
Eli Dart, Lauren Rotman, Brian Tierney, Mary Hester, Jason Zurawski
SC3
2012 Efficient data transfer protocols for big data
abstract
Data set sizes are growing exponentially, so it is important to use data movement protocols that are the most efficient available. Most data movement tools today rely on TCP over sockets, which limits flows to around 20Gbps on today's hardware. RDMA over Converged Ethernet (RoCE) is a promising new technology for high-performance network data movement with minimal CPU impact over circuit-based infrastructures. We compare the performance of TCP, UDP, UDT, and RoCE over high latency 10Gbps and 40Gbps network paths, and show that RoCE-based data transfers can fill a 40Gbps path using much less CPU than other protocols. We also show that the Linux zero-copy system calls can improve TCP performance considerably, especially on current Intel “Sandy Bridge”-based PCI Express 3.0 (Gen3) hosts.
Brian Tierney, Ezra Kissel, D. Martin Swany, Eric Pouyoul
eScience1
2012 Protocols for wide-area data-intensive applications: design and performance issues
abstract
Providing high-speed data transfer is vital to various data-intensive applications.While there have been remarkable technology advances to provide ultra-high-speed network bandwidth, existing protocols and applications may not be able to fully utilize the bare-metal bandwidth due to their inefficient design.We identify the same problem remains in the field of Remote Direct Memory Access (RDMA) networks.RDMA offloads TCP/IP protocols to hardware devices.However, its benefits have not been fully exploited due to the lack of efficient software and application protocols, in particular in wide-area networks.In this paper, we address the design choices to develop such protocols.We describe a protocol implemented as part of a communication middleware.The protocol has its flow control, connection management, and task synchronization.It maximizes the parallelism of RDMA operations.We demonstrate its performance benefit on various local and wide-area testbeds, including the DOE ANI testbed with RoCE links and InfiniBand links.
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi, Brian Tierney, Eric Pouyoul
SC6
2007 The NIDS Cluster: Scalable, Stateful Network Intrusion Detection on Commodity Hardware
Matthias Vallentin, Robin Sommer, Jason Lee 0001, Craig Leres, Vern Paxson, Brian Tierney
RAID6
2006 Intra and Interdomain Circuit Provisioning Using the OSCARS Reservation System
abstract
With the advent of service sensitive applications such as remote controlled experiments, time constrained massive data transfers, and video-conferencing, it has become apparent that there is a need for the setup of dynamically provisioned, quality of service enabled virtual circuits. The ESnet on-demand secure circuits and advance reservation system (OSCARS) is a prototype service enabling advance reservation of guaranteed bandwidth secure virtual circuits. OSCARS operates within the energy sciences network (ESnet), and has provisions for interoperation with other network domains. ESnet is a high-speed network serving thousands of Department of Energy scientists and collaborators worldwide. OSCARS utilizes the Web services model and standards to implement communication with the system and between domains, and for authentication, authorization, and auditing (AAA). The management and operation of end-to-end virtual circuits within the network is done at the layer 3 network level. Multi-protocol label switching (MPLS) and the resource reservation protocol (RSVP) are used to create the virtual circuits or label switched paths (LSP's). quality of service (QoS) is used to provide bandwidth guarantees. This paper describes our experience in implementing OSCARS, collaborations with other bandwidth-reservation projects (including interdomain testing) and future work to be done.
Chin Guok, David W. Robertson, Mary R. Thompson, Jason Lee 0001, Brian Tierney, William E. Johnston
BROADNETS5
2006 S06 - Computing protection in open HPC environments
abstract
Scientific collaboration is critical in high performance computational (HPC) environments. Computer security incidents, however, are a constant threat in today's interconnected computational environments. Many sites struggle with finding a balance between the needs of users and the threat from unauthorized usage or access. Understanding the constantly changing battlefield is critical for effective and optimal deployment of computer security resources.This tutorial will provide an underlying knowledge of the current state of computer security, its effects on HPC environments and mitigation strategies. Topics we will address include:1.HPC Security Fundamentals2.HPC Security Protection Fundamentals3.Intrusion Detection Using Bro4.The Changing Computer Protection Environment5.Cross-site Implications of Collaborative Environments6.Keeping One Step Ahead7.Risk Assessment, Mitigation and Compliance8.Components of a Good Protection ProgramWe will also show how described techniques are in use at SC06 to protect SCinet.
Stephen Lau, Scott Campbell, William T. Kramer, Brian Tierney
SC4
2005 Techniques for tuning workflows in cluster environments
abstract
An important class of parallel processing jobs on clusters today are workflow-based applications that process large amounts of data in parallel. Traditional cluster performance tools are designed for tightly coupled parallel jobs, and not as effective for this type of application. We describe how the NetLogger Toolkit methodology is more appropriate for this class of cluster computing, and describe our new automatic workflow anomaly detection component. We also describe how this methodology is being used by the Nearby Supernova Factory (SNfactory) project at Lawrence Berkeley National Laboratory.
Brian Tierney, Dan Gunter
HPDC1
2005 A First Look at Modern Enterprise Traffic
Ruoming Pang, Mark Allman, Mike Bennett, Jason Lee 0001, Vern Paxson, Brian Tierney
Internet Measurement Conference6
2004 The Grid2003 Production Grid: Principles and Practice
Ian T. Foster, Jerry Gieraltowski, Scott Gose, Natalia Maltsev, Edward N. May, Alexis A. Rodriguez, Dinanath Sulakhe, A. Vaniachine, Jim Shank, Saul Youssef, David Adams, Richard Baker 0003, Wensheng Deng, Dantong Yu, Iosif Legrand, Conrad Steenberg, M. Anzar Afaq, Eileen Berman, James Annis, L. A. T. Bauerdick, Michael Ernst, Ian Fisk, Lisa Giacchetti, Gregory E. Graham, Anne Heavey, Joseph Kaiser, Nickolai Kuropatkin, Ruth Pordes, Vijay Sekhri, John Weigand, Yujun Wu, Keith Baker, Lawrence Sorrillo, John Huth, Matthew Allen, Leigh Grundhoefer, John Hicks, Fred Luehring, Steve Peck, Robert Quick, Stephen C. Simms, George Fekete, Jan vandenBerg, Kihyeon Cho, Kihwan Kwon, Dongchul Son, Hyoungwoo Park, Shane Canon, Keith R. Jackson, David E. Konerding, Jason Lee 0001, Doug Olson, Iwona Sakrejda, Brian Tierney, Mark Green 0001, Russ Miller, James Letts, Terrence Martin, David Bury, Catalin Dumitrescu, Daniel Engh, Robert W. Gardner, Marco Mambelli, Yuri Smirnov, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009, Paul Avery, Richard Cavanaugh, Bockjoo Kim, Craig Prescott, Jorge Rodríguez 0002, Andrew Zahn, Shawn McKee, Christopher T. Jordan, James E. Prewett, Timothy L. Thomas, Horst Severini, Ben Clifford, Ewa Deelman, Larry Flon, Carl Kesselman, Gaurang Mehta, Nosa Olomu, Karan Vahi, Kaushik De, Patrick McGuigan, Mark Sosebee, Dan Bradley, Peter Couvares, Alan DeSmet, Carey Kireyev, Erik Paulson 0001, Alain J. Roy, Scott Koranda, Brian Moe, Bobby Brown, Paul Sheldon
HPDC57
2003 NetLogger: A Toolkit for Distributed System Performance Tuning and Debugging
Dan Gunter, Brian Tierney
Integrated Network Management2
2003 System capability effects on algorithms for network bandwidth measurement
abstract
A large number of tools that attempt to estimate network capacity and available bandwidth use algorithms that are based on measuring packet inter-arrival time. However in recent years network bandwidth has become faster than system input/output (I/O) bandwidth. This means that it is getting harder and harder to estimate capacity and available bandwidth using these techniques. This paper examines the current bandwidth measurement and estimation algorithms, and presents an analysis of how these algorithms might work in a high-speed network environment. This paper also discusses the system resource (hardware and software) issues that affect each of these algorithms, especially running on generic platforms built from off-the-shelf components.
Guojun Jin, Brian Tierney
Internet Measurement Conference2
2002 Dynamic Monitoring of High-Performance Distributed Applications
abstract
Developers and users of high-performance distributed systems often observe performance problems such as unexpectedly low throughput or high latency. Determining the source of the performance problems requires detailed end-to-end instrumentation of all components, including the applications, operating systems, hosts, and networks. However, one must be very careful to design the instrumentation to have extremely low overhead, and not affect the system being monitored. In this paper we present a very light-weight instrumentation system that can be dynamically activated to unobtrusively collect and aggregate detailed end-to-end monitoring information from distributed applications. We also show how emerging "web services" can be used to facilitate remote interaction with this system.
Dan Gunter, Brian Tierney, Keith R. Jackson, Jason Lee 0001, Martin Stoufer
HPDC2
2002 Giggle: a framework for constructing scalable replica location services
abstract
In wide area computing systems, it is often desirable to create remote read-only copies (replicas) of files. Replication can be used to reduce access latency, improve data locality, and/or increase robustness, scalability and performance for distributed applications. We define a replica location service (RLS) as a system that maintains and provides access to information about the physical locations of copies. An RLS typically functions as one component of a data grid architecture. This paper makes the following contributions. First, we characterize RLS requirements. Next, we describe a parameterized architectural framework, which we name Giggle (for GIGa-scale Global Location Engine), within which a wide range of RLSs can be defined. We define several concrete instantiations of this framework with different performance characteristics. Finally, we present initial performance results for an RLS prototype, demonstrating that RLS systems can be constructed that meet performance goals.
Ann L. Chervenak, Ewa Deelman, Ian T. Foster, Leanne Guy, Wolfgang Hoschek, Adriana Iamnitchi, Carl Kesselman, Peter Z. Kunszt, Matei Ripeanu, Robert Schwartzkopf, Heinz Stockinger, Kurt Stockinger, Brian Tierney
SC13
2002 A TCP tuning daemon
abstract
Many high performance distributed applications require high network throughput but are able to achieve only a small fraction of the available bandwidth. A common cause of this problem is improperly tuned network settings. Tuning techniques, such as setting the correct TCP buffers and using parallel streams, are well known in the networking community, but outside the networking community they are infrequently applied. In this paper, we describe a tuning daemon that uses TCP instrumentation data from the Unix kernel to transparently tune TCP parameters for specified individual flows over designated paths. No modifications are required to the application, and the user does not need to understand network or TCP characteristics.
Thomas H. Dunigan, Matthew Mathis, Brian Tierney
SC3
2002 Monitoring data archives for grid environments
abstract
Developers and users of high-performance distributed systems often observe performance problems such as unexpectedly low throughput or high latency. To determine the source of these performance problems, detailed end-to-end monitoring data from applications, networks, operating systems, and hardware must be correlated across time and space. Researchers need to be able to view and compare this very detailed monitoring data from a variety of angles. To address this problem, we propose a relational monitoring data archive that is designed to efficiently handle high-volume streams of monitoring data. In this paper we present an instrumentation and monitoring event archive service that can be used to collect and aggregate detailed end-to-end monitoring information from distributed applications. This archive service is designed to be scalable and fault tolerant. We also show how the archive is based on the "Grid Monitoring Architecture" defined by the Global Grid Forum.
Jason Lee 0001, Dan Gunter, Martin Stoufer, Brian Tierney
SC4
2001 File and Object Replication in Data Grids
abstract
Data replication is a key issue in a data grid and can be managed in different ways and at different levels of granularity: for example, at the file level or the object level. In the high-energy physics community, data grids are being developed to support the distributed analysis of experimental data. We have produced a prototype data replication tool, the Grid Data Management Pilot (GDMP) that is in production use in one physics experiment, with middleware provided by the Globus toolkit used for authentication, data movement and other purposes. We present a new, enhanced GDMP architecture and prototype implementation that uses Globus data-grid tools for efficient file replication. We also explain how this architecture can address object replication issues in an object-oriented database management system. File transfer over wide-area networks requires specific performance tuning in order to gain optimal data transfer rates. We present performance results obtained with GridFTP, an enhanced version of FTP, and discuss tuning parameters.
Heinz Stockinger, Asad Samar, Koen Holtman, William E. Allcock, Ian T. Foster, Brian Tierney
HPDC6
2001 Enabling Network-Aware Applications
abstract
Many high-performance distributed applications use only a small fraction of their available bandwidth. A common cause of this problem is not a flaw in the application design, but rather improperly tuned network settings. Proper tuning techniques, such as setting the correct TCP buffers and using parallel streams, are well-known in the networking community, but outside this community they are infrequently applied. In this paper, we describe a service that makes the task of network tuning trivial for application developers and users. Widespread use of this service should virtually eliminate a common stumbling block for high-performance distributed applications.
Brian Tierney, Dan Gunter, Jason Lee 0001, Martin Stoufer, Joseph B. Evans
HPDC1
2000 A Monitoring Sensor Management System for Grid Environments
abstract
Large distributed systems, such as computational grids, require a large amount of monitoring data be collected for a variety of tasks, such as fault detection, performance analysis, performance tuning, performance prediction and scheduling. Ensuring that all necessary monitoring is turned on and that the data is being collected can be a very tedious and error-prone task. We have developed an agent-based system to automate the execution of monitoring sensors and the collection of event data.
Brian Tierney, Brian Crowley, Dan Gunter, Mason Holding, Jason Lee 0001, Mary R. Thompson
HPDC1
2000 NetLogger: A Toolkit for Distributed System Performance Analysis
abstract
Diagnosis and debugging of performance problems on complex distributed systems requires end-to-end performance information at both the application and system level. We describe a methodology, called NetLogger, that enables real-time diagnosis of performance problems in such systems. The methodology includes tools for generating precision event logs, an interface to a system event-monitoring framework, and tools for visualizing the log data and real-time state of the distributed system. Low overhead is an important requirement for such tools, therefore we evaluate efficiency of the monitoring itself. The approach is novel in that it combines network, host, and application-level monitoring, providing a complete view of the entire system.
Dan Gunter, Brian Tierney, Brian Crowley, Mason Holding, Jason Lee 0001
MASCOTS2
2000 Using High-Speed WANs and Network Data Caches to Enable Remote and Distributed Visualization
abstract
Visapult is a prototype application and framework for remote visualization of large scientific datasets. We approach the technical challenges of tera-scale visualization with a unique architecture that employs high speed WANs and network data caches for data staging and transmission. This architecture allows for the use of available cache and compute resources at arbitrary locations on the network. High data throughput rates and network utilization are achieved by parallelizing I/O at each stage in the application, and by pipelining the visualization process. On the desktop, the graphics interactivity is effectively decoupled from the latency inherent in network applications. We present a detailed performance analysis of the application, and improvements resulting from field-test analysis conducted as part of the DOE Combustion Corridor project.
E. Wes Bethel, Brian Tierney, Jason Lee 0001, Dan Gunter, Stephen Lau
SC2
2000 A data intensive distributed computing architecture for "Grid" applications
Brian Tierney, William E. Johnston, Jason Lee 0001, Mary R. Thompson
Future Gener. Comput. Syst.1
1999 A Network-Aware Distributed Storage Cache for Data Intensive Environments
abstract
Modern scientific computing involves organizing, moving, visualizing, and analyzing massive amounts of data at multiple sites around the world. The technologies, the middleware services, and the architectures that are used to build useful high-speed, wide area distributed systems, constitute the field of data intensive computing. We describe an architecture for data intensive applications where we use a high-speed distributed data cache as a common element for all of the sources and sinks of data. This cache-based approach provides standard interfaces to a large, application-oriented, distributed, on-line, transient storage system. We describe our implementation of this cache, how we have made it "network aware ", and how we do dynamic load balancing based on the current network conditions. We also show large increases in application throughput by access to knowledge of the network conditions.
Brian Tierney, Jason Lee 0001, Brian Crowley, Mason Holding, Jeremy Hylton, Fred L. Drake
HPDC1
1998 The NetLogger Methodology for High Performance Distributed Systems Performance Analysis
abstract
We describe a methodology that enables the real-time diagnosis of performance problems in complex high-performance distributed systems. The methodology includes tools for generating precision event logs that can be used to provide detailed end-to-end application and system level monitoring; a Java agent-based system for managing the large amount of logging data; and tools for visualizing the log data and real-time state of the distributed system. We developed these tools for analyzing a high-performance distributed system centered around the transfer of large amounts of data at high speeds from a distributed storage server to a remote visualization client. However this methodology should be generally applicable to any distributed system. This methodology called NetLogger has proven invaluable for diagnosing problems in networks and in distributed systems code. This approach is novel in that it combines network, host, and application-level monitoring, providing a complete view of the entire system.
Brian Tierney, William E. Johnston, Brian Crowley, Gary Hoo, Christopher X. Brooks, Dan Gunter
HPDC1
1997 High-Speed Distributed Data Handling for On-Line Instrumentation Systems
abstract
The advent (and promise) of shared, widely available, high-speed networks provides the potential for new approaches to the collection, organization, storage, and analysis of high-speed and high-volume data streams from high data-rate, on-line instruments. We have worked in this area for several years, have identified and addressed a variety of problems associated with this scenario, and have evolved an architecture, implementations, and a monitoring methodology that have been successful in addressing several different application areas.In this paper we describe a distributed, wide area network-based architecture that deals with data streams that originate from on-line instruments. Such instruments and imaging systems are a staple of modern scientific, health care, and intelligence environments. Our work provides an approach for reliable, distributed real-time analysis, cataloguing, and archiving of the data streams through the integration and distributed management of a high-speed distributed cache, distributed high-performance applications, and tertiary storage systems.
William E. Johnston, William Greiman, Gary Hoo, Jason Lee 0001, Brian Tierney, Craig E. Tull, Doug Olson
SC5
1994 Distributed Parallel Data Storage Systems: A Scalable Approach to High Speed Image Servers
abstract
We have designed, built, and analyzed a distributed parallel storage system that will supply image streams fast enough to permit multi-user, “real-time”, video-like applications in a wide-area ATM network-based Internet environment. We have based the implementation on user-level code in order to secure portability; we have characterized the performance bottlenecks arising from operating system and hardware issues, and based on this have optimized our design to make the best use of the available performance. Although at this time we have only operated with a few classes of data, the approach appears to be capable of providing a scalable, high-performance, and economical mechanism to provide a data storage system for several classes of data (including mixed multimedia streams), and for applications (clients) that operate in a high-speed network environment.
Brian Tierney, Jason Lee 0001, Ling Tony Chen, Hanan Herzog, Gary Hoo, Guojun Jin, William E. Johnston
ACM Multimedia1
1994 Using high speed networks to enable distributed parallel image server systems
abstract
We describe the design and implementation of a distributed parallel storage system that uses high-speed ATM networks as a key element of the architecture. Other elements include a collection of network-based disk block servers, and an associated name server that provides some file system functionality. The implementation is based on user level software that runs on UNIX workstations. Both the architecture and the implementation are intended to provide for easy and economical scalability. This approach has yielded a data source that scales economically to very high speed. Target applications include online storage for both very large images and video sequences. This paper describes the architecture, and explores the performance issues of the current implementation.>
Brian Tierney, William E. Johnston, Hanan Herzog, Gary Hoo, Guojun Jin, Jason Lee 0001, Ling Tony Chen, Doron Rotem
SC1